SyncNet LSE-C · best lip sync in the comparison
faster generation than Aurora 1.0
continuous video from a single photo
per output-second at 720p, half of Aurora 1.0
WHAT AURORA 2.0 DOES
Start from a person, a character, or an ad frame your team already approved. Aurora 2.0 brings the presenter to life from the supplied voice and keeps the composition ready to run.
Best-in-class lip sync
9.14 SyncNet LSE-C, the highest result in our side-by-side comparison.
Full-body expressiveness
Facial nuance, eye contact, breathing, head motion, hand gestures, and body language move together.
Exact source voice
Your original audio stays intact, with 99.97% waveform correlation and 0 ms measured offset in long-video validation.
One photo, no setup session
Start from the person, character, styling, and composition you already approved.
Built for any language
The performance follows the supplied voice track, opening the same workflow to global campaigns.
Made for creator formats
UGC ads, podcasts, explainers, localization, virtual spokespeople, and character content from the same simple input.
LONG-FORM
Up to a minute of continuous video from one photo and one complete voice track — long enough for a full ad, product story, lesson, or localized message, not just the opening line.
continuous video from one photo and one voice track
measured audio offset in long-form validation
minimum waveform correlation against the supplied track
29.2 S · FULL SPEECH TRACK
A complete message from one portrait
Five overlapping generation windows carry the full 28.8-second speech track, with nothing trimmed.
59.7 S · ONE PHOTO
One continuous minute
A 59.7-second render with the original waveform preserved at 99.96% correlation. This validation run contains one detected camera cut near 54 seconds.
The same simple input — one photo and one voice track — covers performance ads, long-form explainers, and characters that hold their look for a full minute. No reshoots, no setup session.
RETAIL SKINCARE UGC · 5.2 S
Ads and UGC
Turn an approved creator image and script read into a complete one-take performance. Generate hooks, offers, and localized variants without reshooting.
ONBOARDING EXPLAINER · 40.6 S
Podcasts and explainers
Let a thought finish. Longer clips make product walkthroughs, lessons, commentary, and narrative content practical in one generation.
STYLIZED CHARACTER · 59.7 S
Virtual people and characters
Animate real, stylized, or synthetic identities from a single frame while preserving the intended look and original voice.
Aurora 2.0 reaches the strongest overall result in the comparison — a 90.5 Avatar Performance Score — at the lowest listed public price, 7¢ per output-second. Every model runs the same fixed benchmark and the same scoring pipeline.
Avatar Performance Score weights 50% lip sync, 25% perceptual quality, and 25% identity preservation, normalized to 0–100. Latency is request-to-result time recorded during the benchmark; Aurora 2.0 uses warm serving measurements, and provider queue and cold-start behavior can vary.
| Model | Avatar Performance Score ↑ | SyncNet LSE-C ↑ | Q-Align / 5 ↑ | ID-SIM ↑ | Public price / output-second ↓ |
|---|---|---|---|---|---|
| Aurora 2.0 | 90.5 | 9.14 | 4.88 | 0.82 | 7.00¢ |
| Aurora 1.0 | 89.3 | 8.78 | 4.88 | 0.84 | 14.00¢ |
| VEED Fabric 1.0 | 87.4 | 8.67 | 4.93 | 0.78 | 15.00¢ |
| Veo 3.1 | 85.6 | 8.38 | 4.94 | 0.76 | 40.00¢ |
| fal H3 Max Lip Sync | 81.9 | 7.95 | 4.80 | 0.72 | 8.00¢ |
| HeyGen Avatar 4 | 81.8 | 7.93 | 4.79 | 0.73 | 10.00¢ |
| Kling Avatar Pro | 75.3 | 6.75 | 4.88 | 0.69 | 11.50¢ |
| OmniHuman 1.5 | 69.9 | 6.26 | 4.74 | 0.59 | 16.00¢ |
PRICING
Pay only for the seconds you generate. 720p is the default.
RESOLUTION
480p
0.4
$0.04
720p
0.7
1080p
1.4
$0.14
2K
Up to 14 seconds of speech
2.8
$0.28
Billed per output-second. 720p is the default; 2K supports up to 14 seconds of speech.
MEASUREMENT NOTES
Quality metrics
SyncNet LSE-C measures audio-to-mouth alignment, Q-Align measures perceptual video quality, and ID-SIM measures identity preservation. Higher is better for every quality score shown.
Avatar Performance Score
50% lip sync, 25% perceptual quality and 25% identity preservation, normalized to 0–100. Every model runs the same fixed evaluation suite and scoring pipeline.
Examples
Visual examples are selected from the evaluation suite, not a random sample. Each 1.0 vs 2.0 pair uses the same source image, audio and evaluation window.
Latency
Recorded request-to-result time. Aurora 2.0 uses warm-serving measurements; provider queue and cold-start behavior can vary.
Pricing
Public list prices per output-second, not internal cost. Aurora 2.0, Aurora 1.0 and VEED Fabric 1.0 at 720p; fal H3 Max Lip Sync at 768p; other providers at their benchmarked public tier. Enterprise terms may differ.
Long-form
60 seconds is the public product ceiling. Audio offset and waveform correlation are measured on the long-form validation set. The 59.7-second character clip is a validation render and contains one detected camera cut near 54 seconds.
Updated September 28, 2026.


