There used to be a tell. Synthetic narration had a flatness to it, a way of hitting every sentence with the same energy, and the ear caught it even when the words were perfect. That tell is mostly gone, and we now have a measure of how far gone it is.
Which means the interesting questions have moved. The technology works; the problems that remain are human ones — and in the eighteen months since this piece was first published, one of them has been answered by a federal court in a way that changes how the rest should be approached.
The craft has genuinely improved
Modern voice tools handle the things that used to betray them: the breath before a long clause, the small downward drift at the end of a paragraph, the way emphasis lands on the word that carries the meaning. Long-form narration — audiobooks, documentary voiceover, podcast inserts — is now within reach of a laptop and a licence.
The strongest evidence for this is not a demo reel but a controlled study. Writing in Scientific Reports in March 2025, Sarah Barrington, Emily Cooper and Hany Farid asked 604 participants to judge clips drawn from 220 speakers, with the synthetic versions produced using ElevenLabs’ instant cloning. Asked simply whether a voice was real or generated, listeners were right about 60% of the time for the synthetic clips and 67% of the time for the genuine ones. Their conclusion is worth quoting exactly: relying on human perception to detect AI-generated voice clones “is no longer consistently reliable.”
It is worth being precise about what that does and does not say, because the looser version of this claim has been circulating for a while. Sixty per cent is not chance. Listeners are not guessing — they retain a real, if slight, edge. What has collapsed is the margin that made the ear a usable safeguard. An audience that is right three times in five is an audience that cannot be asked to police the difference, which is a different and more useful statement than saying nobody can tell.
There is a second result in the same study that matters more for this piece, and it runs the other way. Alongside the real-or-generated task, the researchers asked a different question: presented with two clips, did they sound like the same person? There, participants judged a clone to match the speaker it was built from around 80% of the time. Set the two findings side by side and the picture inverts the usual worry. A synthetic voice is only moderately good at passing as human in the abstract. It is considerably better at passing as you.
That distinction is not a technicality. The harm this article is about has never really been that audiences cannot tell a machine from a person; it is that a specific person’s voice can be worn by someone else. The measurement that tracks that harm is the identity one, and it is the higher of the two numbers.
For independent producers this is also a material change in economics. A correction that once meant booking the booth again is now a line edit. It is one of the clearest cases of the pattern we traced in the unbundling of the creative studio: a capability that used to require a room, a schedule and another person’s afternoon has become a line item that renews monthly.
A voice is not just a sound. It is a record of a person having been somewhere, having meant something. Cloning the sound is easy now. Accounting for the person is the part we keep skipping.
The part the demos skip
Every impressive clone is trained on someone. The consent question is not abstract, and it is no longer hypothetical either.
In 2019 and 2020 two professional voice actors, Paul Lehrman and Linnea Sage, took small jobs through Fiverr. They were told the recordings were for internal, academic or test purposes and would not be used commercially. The people who hired them, according to the complaint they later filed, were employees of the AI voice company Lovo. Their voices went on to be sold through Lovo’s commercial product under the names “Kyle Snow” and “Sally Coleman.” Lehrman found out the way anyone would least want to: he recognised his own voice narrating a video he had never been hired for.
That is the scenario the licensing conversation tends to describe in the abstract — a timbre for sale in a marketplace its owner never entered. It is worth knowing that it has a case name, a docket and a ruling.
What the law actually decided
On 17 July 2025 a federal judge in the Southern District of New York ruled on Lovo’s motion to dismiss, and the shape of the decision is more interesting than a simple win or loss.
The claims that survived were the ordinary ones: breach of contract, New York’s right-of-publicity statute, and the state’s consumer protection law. The claims that were dismissed were the copyright claims — including the argument that the AI-generated clones were derivative works of the original recordings. As Skadden’s summary of the opinion puts it, the court held that copyright law “protects only the original sound recordings (the fixed expression), not the abstract qualities of a voice or new recordings that merely imitate or simulate the original.”
Read that again with the argument of this piece in mind. A court, reasoning from copyright doctrine, arrived at the same place: the recording is property, the voice is not. The law protects the artefact and struggles with the person. That is not a gap in the drafting — it is what happens when a body of law built around fixed works meets a technology that copies the thing underneath the work.
Two qualifications keep this from being the last word. The dismissal of the training-data claim came with leave to amend, so the question of whether ingesting a recording to build a model is itself an infringement was deferred rather than settled. And a motion to dismiss decides only what may be argued, not who is right; the surviving claims still have to be proved. What the ruling does establish is the terrain. Anyone hoping copyright would do the heavy lifting here now has a written answer, and it is no.
Legislatures have started patching around it. Tennessee’s ELVIS Act, signed on 21 March 2024 and effective from 1 July 2024, expanded the state’s personal rights law to cover “an individual’s actual voice as well as simulations of that voice,” and reached further than most: as Skadden notes, it creates liability not only for the person who makes an unauthorised clone but for anyone who distributes a tool whose primary purpose is producing one. Whether that provision survives contact with general-purpose software is an open question — most cloning tools neither require nor prevent an authorisation check — but the direction of travel is unambiguous.
Using it honestly
There is a defensible way to work here, and it is not complicated: clone your own voice, or license one with clear, revocable consent and fair terms. Disclose synthetic narration where it matters to the audience. Keep a human in the loop for anything that carries a claim or a feeling.
What has changed since that advice was first written is that provenance is becoming a product feature rather than purely a matter of personal discipline. ElevenLabs now describes blocking the cloning of celebrity and other high-risk voices, and requiring technological verification for access to its professional cloning tool. That is a meaningful shift: the question “who agreed to this?” is starting to be asked by the tool at the point of upload rather than by the producer afterwards, if at all.
It is not sufficient. Verification at upload does nothing about a recording obtained under a false description, which is precisely what the Lovo plaintiffs allege happened to them — the consent was fraudulently procured, not absent. And the same judgement problem applies here as everywhere else in this stack: knowing which take carries the meaning is a skill the tools do not supply, in narration as much as in the working grammar of prompting a diffusion model.
The honest summary is narrower than the one this piece originally offered, and more useful for it. The tools have solved the sound. The courts have now told us they will not solve the ownership, because the thing being taken is not the kind of thing copyright was built to hold. What is left is contract, disclosure and consent — unglamorous instruments, all of them, and the only ones actually on the table. Pretending otherwise is the one shortcut this medium cannot afford.