The voice, the audio and the 3 scripts — LocutaIA handles the rest.
EL Profile
Recommended: audio under 30 minutes. Longer files work but cost more tokens and take longer.
Original audio
Drop audio file here or click to browse
mp3, wav, m4a, flac, ogg
Scripts (3 SRT)
Drop 3 SRT files here or click to browse
Raw (EN) + Translated (ES) + Adapted (ES) — auto-classified
SRT ClassificationClear saved SRTs ×
SRT Raw (EN)—
SRT Translated—
SRT Adapted—
Voice generation mode
⚡ Normalfast
Generates segments in parallel. For intonation, ElevenLabs only uses the neighboring segment's text. Fast and cached.
🎯 Probetter tone
Generates sequentially, chaining the actual audio of the previous segment → keeps the tone more consistent. Cost: ~4× slower, no cache.
Sync & finishing
Strict sync (2 tracks)
Nothing is moved to force a fit: in-sync clips go to track B, the rest to track C (red) for Premiere. Downloads become a 2-track ZIP.
Match volume when finished
Runs the editor's Match volume automatically at the end (voice-only RMS): the MP3 and every export come out already leveled.
— optional, good defaults
ElevenLabs TTS — Speed
Controls how fast the cloned voice speaks. The pipeline uses this exact value for every segment — bad fits go to the retry phase (text rewrite or manual review).
Speed0.96?
Speaking rate applied to every segment. 1.0 = natural speed. ElevenLabs range: 0.70 - 1.20
Gap fill100%?
How much SHORT segments are slowed/stretched to fill the slot. 100% = fill the gap (default, current behavior). 0% = leave the natural silence, never slow the voice. Lower it when you leave intentional blanks after sentences (Freya). Only affects under-segments; never compresses long ones.
Voice Settings
Stability, similarity, style, and speaker boost are applied uniformly to every segment. (Per-segment Gemini prosody analysis was removed 2026-06-10 — uniform settings give a more consistent voice across the whole video.)
Stability0.50?
Voice consistency across the segment. Low (0.2-0.4) = expressive, emotional delivery. High (0.6-0.8) = calm, stable narration. Affects how much the voice varies in pitch and cadence.
Similarity0.85?
How closely the output matches your cloned voice. Higher = more faithful to the original voice sample but can sound robotic. Lower = more natural but may drift from the voice.
Style0.00?
Amplifies the natural style of the cloned voice. 0 = neutral delivery (recommended for dubbing). Higher values add exaggeration — use sparingly for dramatic moments.
Speaker boostON?
Enhances the distinctive characteristics of your cloned voice. Usually best left ON for clearer voice identity.
Finish — volume
Voice level (RMS)-15.4 dBFS?
Target level for the whole job's voice and for "Match volume when finished". -15.4 dBFS = the pipeline standard. Only move it if your mix needs the voice hotter or quieter.
Script source
🎚️ Autodefault
Uses the translated script and swaps to the adapted one only when a line won't fit. Best of both.
📝 Translatedfaithful
Always the translated script (most faithful). Long lines are fitted with speed/atempo; never swapped to adapted.
✂️ Adaptedfits best
Always the adapted script (shorter, fits timing best). Falls back to translated when a line has no adaptation.
Pronunciation dictionary
Global per-language rules applied to every job right before the voice is generated (e.g. "AAA" → "Triple A"). Shared by all users; the subtitles never change.
Re-resolve syncRe-runs fitting, placement and overlap resolution with the current policy. No TTS — free and repeatable.
Job log
Job files
Settings used for this job
Raw state.json
Edit Profile
📖 Pronunciation dictionary
Rules applied to every job right before the voice is generated — the visible subtitles never change. For brands, acronyms and words the voice mispronounces: AAA → Triple A.
Written termHow the voice should say it
Delete profile
Are you sure?
🎮
¿Perfil de voz equivocado?
El contenido no parece coincidir con el perfil seleccionado.