NEW

NEW

Aurora 2.0

Aurora 2.0

Aurora 2.0

Our best avatar model yet. One photo, one voice track, up to a full minute.

Our best avatar model yet. One photo, one voice track, up to a full minute.

Our best avatar model yet. One photo, one voice track, up to a full minute.

AT A GLANCE

AT A GLANCE

Aurora 2.0 by the numbers.

Aurora 2.0 by the numbers.

9
.
1
4
9
.
1
4
9
.
1
4

SyncNet LSE-C · best lip sync in the comparison

2
.
0
×
2
.
0
×
2
.
0
×

faster generation than Aurora 1.0

6
0
s
6
0
s
6
0
s

continuous video from a single photo

7
¢
7
¢
7
¢

per output-second at 720p, half of Aurora 1.0

WHAT AURORA 2.0 DOES

One photo becomes a complete performance.

One photo becomes a complete performance.

Start from a person, a character, or an ad frame your team already approved. Aurora 2.0 brings the presenter to life from the supplied voice and keeps the composition ready to run.

Best-in-class lip sync

9.14 SyncNet LSE-C, the highest result in our side-by-side comparison.

Full-body expressiveness

Facial nuance, eye contact, breathing, head motion, hand gestures, and body language move together.

Exact source voice

Your original audio stays intact, with 99.97% waveform correlation and 0 ms measured offset in long-video validation.

One photo, no setup session

Start from the person, character, styling, and composition you already approved.

Built for any language

The performance follows the supplied voice track, opening the same workflow to global campaigns.

Made for creator formats

UGC ads, podcasts, explainers, localization, virtual spokespeople, and character content from the same simple input.

See the upgrade
on the same inputs.

See the upgrade
on the same inputs.

Three selected examples on the same source image and audio.
Both play in sync. Sound follows Aurora 2.0.

Three selected examples on the same source image and audio.
Both play in sync. Sound follows Aurora 2.0.

Aurora 1.0
Aurora 2.0
01 / 03Sharper mouth timingOn the fixed avatar benchmark, lip-sync accuracy rises from 8.78 to 9.14 SyncNet LSE-C — the strongest result in the comparison.

LONG-FORM

Long enough to tell the whole story.

Long enough to tell the whole story.

Up to a minute of continuous video from one photo and one complete voice track — long enough for a full ad, product story, lesson, or localized message, not just the opening line.

6
0
s
6
0
s
6
0
s

continuous video from one photo and one voice track

0
.
0
ms
0
.
0
ms
0
.
0
ms

measured audio offset in long-form validation

9
9
.
9
5
%
9
9
.
9
5
%
9
9
.
9
5
%

minimum waveform correlation against the supplied track

29.2 S · FULL SPEECH TRACK

A complete message from one portrait

Five overlapping generation windows carry the full 28.8-second speech track, with nothing trimmed.

59.7 S · ONE PHOTO

One continuous minute

A 59.7-second render with the original waveform preserved at 99.96% correlation. This validation run contains one detected camera cut near 54 seconds.

From a single asset

From a single asset

to a complete campaign.

to a complete campaign.

The same simple input — one photo and one voice track — covers performance ads, long-form explainers, and characters that hold their look for a full minute. No reshoots, no setup session.

RETAIL SKINCARE UGC · 5.2 S

Ads and UGC

Turn an approved creator image and script read into a complete one-take performance. Generate hooks, offers, and localized variants without reshooting.

ONBOARDING EXPLAINER · 40.6 S

Podcasts and explainers

Let a thought finish. Longer clips make product walkthroughs, lessons, commentary, and narrative content practical in one generation.

STYLIZED CHARACTER · 59.7 S

Virtual people and characters

Animate real, stylized, or synthetic identities from a single frame while preserving the intended look and original voice.

BENCHMARK

BENCHMARK

How Aurora 2.0 compares.

How Aurora 2.0 compares.

Aurora 2.0 reaches the strongest overall result in the comparison — a 90.5 Avatar Performance Score — at the lowest listed public price, 7¢ per output-second. Every model runs the same fixed benchmark and the same scoring pipeline.


Avatar Performance Score weights 50% lip sync, 25% perceptual quality, and 25% identity preservation, normalized to 0–100. Latency is request-to-result time recorded during the benchmark; Aurora 2.0 uses warm serving measurements, and provider queue and cold-start behavior can vary.

SPEED × QUALITYQuality at production speedHigher Avatar Performance Score and lower median generation latency are better. Latency is request-to-result time recorded during the benchmark; Aurora 2.0 uses warm serving measurements.
Aurora 2.0 by Creatify Labs | One Photo, One Minute, Your Exact Voice707580859010s20s40s80s160s320s640sMedian generation latency · seconds · log scaleAvatar Performance ScoreBEST DIRECTION ↖Aurora 2.0Aurora 1.0VEED Fabric 1.0Veo 3.1fal H3 Max Lip SyncHeyGen Avatar 4Kling Avatar ProOmniHuman 1.5
Aurora 2.090.5 · 31.6s
Aurora 1.089.3 · 63.4s
VEED Fabric 1.087.4 · 160s
Veo 3.185.6 · 63.7s
fal H3 Max Lip Sync81.9 · 12.7s
HeyGen Avatar 481.8 · 86.1s
Kling Avatar Pro75.3 · 425s
OmniHuman 1.569.9 · 359s
SPEED × QUALITYQuality at production speedHigher Avatar Performance Score and lower median generation latency are better. Latency is request-to-result time recorded during the benchmark; Aurora 2.0 uses warm serving measurements.
Aurora 2.0 by Creatify Labs | One Photo, One Minute, Your Exact Voice707580859010s20s40s80s160s320s640sMedian generation latency · seconds · log scaleAvatar Performance ScoreBEST DIRECTION ↖Aurora 2.0Aurora 1.0VEED Fabric 1.0Veo 3.1fal H3 Max Lip SyncHeyGen Avatar 4Kling Avatar ProOmniHuman 1.5
Aurora 2.090.5 · 31.6s
Aurora 1.089.3 · 63.4s
VEED Fabric 1.087.4 · 160s
Veo 3.185.6 · 63.7s
fal H3 Max Lip Sync81.9 · 12.7s
HeyGen Avatar 481.8 · 86.1s
Kling Avatar Pro75.3 · 425s
OmniHuman 1.569.9 · 359s
QUALITY × ECONOMICSAurora 2.0 moves the frontierHigher Avatar Performance Score and lower public price per output-second are better. Aurora prices are at 720p; competitors at their benchmarked public tier.
Aurora 2.0 by Creatify Labs | One Photo, One Minute, Your Exact Voice70758085905¢7¢10¢15¢20¢50¢Public price per output-second · log scaleAvatar Performance ScoreBEST DIRECTION ↖Aurora 2.0Aurora 1.0VEED Fabric 1.0Veo 3.1fal H3 Max Lip SyncHeyGen Avatar 4Kling Avatar ProOmniHuman 1.5
Aurora 2.090.5 · 7¢
Aurora 1.089.3 · 14¢
VEED Fabric 1.087.4 · 15¢
Veo 3.185.6 · 40¢
fal H3 Max Lip Sync81.9 · 8¢
HeyGen Avatar 481.8 · 10¢
Kling Avatar Pro75.3 · 11.5¢
OmniHuman 1.569.9 · 16¢
FULL SCORECARD
ModelAvatar Performance Score ↑SyncNet LSE-C ↑Q-Align / 5 ↑ID-SIM ↑Public price / output-second ↓
Aurora 2.090.59.144.880.827.00¢
Aurora 1.089.38.784.880.8414.00¢
VEED Fabric 1.087.48.674.930.7815.00¢
Veo 3.185.68.384.940.7640.00¢
fal H3 Max Lip Sync81.97.954.800.728.00¢
HeyGen Avatar 481.87.934.790.7310.00¢
Kling Avatar Pro75.36.754.880.6911.50¢
OmniHuman 1.569.96.264.740.5916.00¢
Ranked by Avatar Performance Score: 50% lip sync, 25% perceptual quality, 25% identity preservation, normalized to 0–100. Aurora 2.0, Aurora 1.0 and VEED Fabric 1.0 are priced at 720p; fal H3 Max Lip Sync at its 768p tier; other providers at their benchmarked public tier.

PRICING

Simple, per-second pricing.

Simple, per-second pricing.

Pay only for the seconds you generate. 720p is the default.

RESOLUTION

CREDITS PER SECOND

CREDITS/S

$ PER SECOND

$/S

480p

0.4

$0.04

720p

0.7

$0.07

$0.07

1080p

1.4

$0.14

2K

Up to 14 seconds of speech

2.8

$0.28

Billed per output-second. 720p is the default; 2K supports up to 14 seconds of speech.

One photo. One minute.
Your exact voice.

One photo. One minute.
Your exact voice.


One photo. One minute. Your exact voice.

One photo. One minute. Your exact voice.

MEASUREMENT NOTES

Quality metrics

SyncNet LSE-C measures audio-to-mouth alignment, Q-Align measures perceptual video quality, and ID-SIM measures identity preservation. Higher is better for every quality score shown.

Avatar Performance Score

50% lip sync, 25% perceptual quality and 25% identity preservation, normalized to 0–100. Every model runs the same fixed evaluation suite and scoring pipeline.

Examples

Visual examples are selected from the evaluation suite, not a random sample. Each 1.0 vs 2.0 pair uses the same source image, audio and evaluation window.

Latency

Recorded request-to-result time. Aurora 2.0 uses warm-serving measurements; provider queue and cold-start behavior can vary.

Pricing

Public list prices per output-second, not internal cost. Aurora 2.0, Aurora 1.0 and VEED Fabric 1.0 at 720p; fal H3 Max Lip Sync at 768p; other providers at their benchmarked public tier. Enterprise terms may differ.

Long-form

60 seconds is the public product ceiling. Audio offset and waveform correlation are measured on the long-form validation set. The 59.7-second character clip is a validation render and contains one detected camera cut near 54 seconds.

Updated September 28, 2026.

Creatify logo white

The research lab building lifelike AI video models.

linkedin logo
twitter logo

Creatify Lab • Copyright © 2026

Creatify logo white

The research lab building lifelike AI video models.

linkedin logo
twitter logo

Creatify Lab • Copyright © 2026

Creatify logo white

The research lab building lifelike AI video models.

linkedin logo
twitter logo

Creatify Lab • Copyright © 2026

Creatify logo white

The research lab building lifelike AI video models.

linkedin logo
twitter logo

Creatify Lab • Copyright © 2026