Realistic AI voice generation and cloning — used by indie creators and Hollywood alike.
Last updated
ElevenLabs is a synthetic voice platform, and the interesting questions about it are almost never about whether the voice sounds real. It usually does. The questions that decide whether a project ships are about permission, consistency and performance: whose voice you are allowed to use, whether the voice holds together across six hours of audio rather than six seconds, and which parts of a read a machine still cannot do.
This page is organised around those three, because they are what actually stops production.
Voice cloning is the feature that sells the product and the feature that creates legal exposure. Two different permissions are involved and they get collapsed together constantly.
The first is consent from the person whose voice it is. A recording you have the right to distribute is not the same as a recording you have the right to synthesise from. An audiobook narrator who was paid for a performance did not necessarily grant the right to generate new performances from it, and a contract signed before synthesis was practical almost certainly does not address it. If the voice belongs to an employee, a founder, a customer or a contractor, get consent in writing that names synthesis specifically, and name what it may be used for and for how long.
The second is the platform's own verification. Higher-fidelity cloning typically requires the person to record a verification statement rather than letting you upload any audio you happen to have, which is a deliberate friction rather than an oversight. Treat any workflow that routes around that friction as a warning sign about the workflow, not about the platform.
Beyond consent there is the question of commercial rights, which vary by plan tier and have been revised, and beyond that the law, which varies by jurisdiction. Several places now treat a recognisable voice as a protected attribute of a person in its own right. None of this is a reason to avoid the technology. It is a reason to decide who owns the voice before you build a production pipeline around it, because retrofitting consent onto a published catalogue is not a task anyone enjoys.
A demo is thirty seconds long and reveals almost nothing about production use. The problems in long-form work are cumulative.
Pronunciation is the obvious one. Proper nouns, character names, technical vocabulary, acronyms and anything borrowed from another language will be guessed, and the guess is stable enough that a wrong guess is wrong every single time. Any serious long-form workflow needs a pronunciation pass and a per-project dictionary, and that pass is a real, recurring cost that the per-character price does not include.
The subtler problem is drift. Generate a chapter in one session and the next chapter a week later, after a model update or with different generation settings, and the two can differ in pace, brightness or energy in a way that is hard to name but audible on a continuous listen. The defences are unglamorous: fix your settings and record them, generate in the largest coherent unit the tool allows rather than sentence by sentence so context carries, keep the source of truth in a script file rather than in pasted fragments, and regenerate whole sections rather than patching single lines.
The third is that errors compound with length. A one-in-a-hundred oddity is invisible in a product video and appears dozens of times in an audiobook. Long-form work needs a listen-through by a human, which is the cost most budgets forget.
Synthetic voice has closed most of the gap on tone, clarity and naturalness, and it has not closed the gap on interpretation. The things it still does worse are specific rather than vague.
Breath and silence are used deliberately by a good reader. A pause before a revelation, a breath that signals exhaustion, a held beat that lets a joke land — these are performance choices made from understanding the text, and a model generating plausible prosody is not making them. It can be steered towards them with markup and direction, but you are the one deciding where they go, line by line, which is slower than people expect.
Sustained emotional arc is the other gap. A model can render a sentence as sad. Carrying a character through a scene where grief turns into anger, with the change audible in the voice before it is visible in the words, is a different task. Character work in dialogue compounds this: not just different voices, but the same character sounding different when speaking to different people.
And there is the question of the listener's contract. For a personal essay, a memorial, an apology from a company, or anything where the point is that a person is speaking to you, synthesis does not fail technically — it fails at the premise. Disclosure helps, and it does not restore what was lost.
These get compared because both touch audio, and they sit on opposite sides of a clean line. ElevenLabs generates audio that was never recorded. Descript edits audio that was: it transcribes a recording and lets you change the audio by changing the transcript, including patching a few words in a recorded voice.
The practical test is whether a recording exists. If a person sat down and read the thing and you need to cut, tighten or fix it, that is an editing job. If nothing was ever recorded and the script needs a voice — or if the volume of scripts means nobody is ever going to sit down and read them — that is a generation job. Podcast and interview workflows are editing. Documentation, e-learning catalogues and in-app voices are generation. Plenty of teams run both and route by that question. For a longer walk through the output quality, we have a separate ElevenLabs review.
Do not use it for a voice you do not have documented, synthesis-specific permission to use. This is the one non-negotiable item on the list, and the awkward version of the mistake is internal: cloning a colleague's voice for a demo because it was funny is how an organisation discovers it has no policy.
Do not use it where the value of the audio is that a specific person chose to speak. Leadership messages during a crisis, condolences, anything framed as personal — the efficiency gain is real and it is not what is being bought.
Do not use it for a flagship performance without budgeting a human pass. Full-cast fiction, high-profile brand narration and anything where the read is the product will need a director's attention line by line, at which point you should compare the total cost honestly against hiring a narrator rather than against the per-character rate.
Do not use it for high-stakes short-form under time pressure with no review step. Legal disclaimers, medical instructions, safety announcements and financial terms are exactly where a mispronounced word or a dropped negation does damage, and exactly where the volume is low enough that generation was never saving you much.
And do not build an unguarded cloning feature into a consumer product. If users can upload arbitrary audio and get a usable clone, you have built an impersonation tool, and the fact that the underlying platform has verification requirements does not transfer that responsibility away from you.
Internal training, product documentation, release notes read aloud, course modules that change every quarter. The economics here were never in favour of recording, which means synthetic voice is not replacing a narrator, it is replacing silence or an unread page.
A price changed, a feature got renamed, a line was wrong. Regenerating one segment in a matched voice avoids rebooking a session for thirty seconds of audio. This is also the case where consent paperwork matters most, because the voice being matched usually belongs to someone specific.
Dubbing and multilingual narration for markets where hiring a native narrator for every update is not realistic. Budget for a native speaker to review the output rather than shipping unreviewed, since the failure mode is a confidently wrong pronunciation that no one on your team can hear.
Temporary lines so a scene can be blocked, timed and playtested before casting. Even studios that intend to record every final line use synthesis for the draft, because waiting on a booking to find out a scene does not work is expensive.
Assistants, IVR flows and accessibility readouts where the text is generated at runtime and no recording could exist. Latency and interruption handling matter more than raw naturalness here, and they are what you should actually test.
Turning articles, reports and documentation into audio for people who prefer or need to listen. The bar is comprehension over hours rather than beauty over seconds, which makes pronunciation handling the thing to evaluate.
Backlist titles, technical manuals and niche non-fiction that would never earn back a studio recording. This works, and it works best when someone listens to the whole thing before release and the listing is honest about how it was produced.
ElevenLabs bills text-to-speech through credits, where one credit corresponds to roughly one character of text on the standard models, and layers several tiers on top of that: a free tier, then Starter (around $5/mo), Creator (around $22/mo), Pro (around $99/mo), and higher volume and business tiers up to a custom Enterprise plan. The structure matters more than any individual figure. Because you are billed by characters rather than by minutes, cost scales with script length, and a long-form project is best budgeted by counting the characters in the manuscript before you start rather than by looking at the monthly price. Regeneration is not free either, so a workflow that iterates on a chapter ten times costs roughly ten chapters. Conversational voice agents are metered separately from the character allowance, on a per-minute basis. Two plan details decide more purchases than the headline rate: which tier carries commercial usage rights and what attribution is required on the free tier, and which tier unlocks higher-fidelity voice cloning. ElevenLabs has revised allowances, tier contents and usage terms more than once, so confirm the current terms on its own pricing page before committing a production budget.
Yes, and more specifically than people assume. Owning a recording is not the same as holding the right to generate new speech from it, and consent given before synthesis was practical usually does not cover synthesis. Get written permission that names voice synthesis explicitly, states what the generated audio may be used for, and states for how long. Higher-fidelity cloning generally also requires the person to record a verification statement, which is a deliberate check rather than a hurdle to route around. Several jurisdictions now treat a recognisable voice as a protected attribute of a person, so if the use is commercial, this is a question for counsel rather than for a review page.
That depends on the tier you are on, and the terms have been revised. Free usage has historically carried attribution requirements and narrower rights than paid usage. Rather than trusting any number or condition quoted second-hand, read the current usage terms for the specific plan you intend to buy before you publish, particularly if the audio will run in an advertisement or a paid product.
More often than they notice in a clip, and for cumulative reasons. Proper nouns and technical terms get a fixed wrong pronunciation that recurs on every occurrence. Chapters generated weeks apart can drift in pace and energy. Rare artefacts that are invisible in a two-minute video appear repeatedly across six hours. All of this is manageable with a pronunciation dictionary, locked generation settings, generation in large coherent units and a human listen-through before release, but that listen-through is a real cost and it is the one most budgets leave out.
Ask whether a recording exists. Descript is for audio somebody recorded: it transcribes the take and lets you edit the audio by editing the text, including patching a handful of words. ElevenLabs is for audio nobody recorded and nobody is going to. Podcasts and interviews are editing work. Documentation, e-learning and in-product speech are generation work. Teams that do both usually run both and route each job by that single question.
Full review coming soon.