ElevenLabs
1. The organization
ElevenLabs is an AI speech research and software company providing generative voice AI technologies, synthetic speech synthesis, and automated voice cloning tools. The organization operates globally with research and operational hubs across Europe and North America, delivering services to individual content creators, developers, and enterprise clients.
The company provides digital audio solutions through three primary service offerings:
- Creative & Publishing Tools: A web interface for text-to-speech (TTS), automated dubbing, and synthetic voice generation across dozens of languages.
- Voice Cloning Infrastructure: A service allowing users to create digital replicas of human voices using short audio samples (Instant Voice Cloning) or extensive training audio (Professional Voice Cloning).
- Developer API & Voice Agents: Developer infrastructure for integrating real-time speech-to-speech conversion and interactive conversational AI agents into third-party commercial applications and automated workflows.
As a high-profile provider of synthetic media and biometric processing tools operating within the European market, the company serves as a critical subject for data-ethical scrutiny regarding consent, transparency, and misuse prevention under frameworks like the Data Ethics Decision Aid (DEDA).
2. The AI technologies Employed
- Text-to-Speech (TTS) Synthesis: The foundational technology used to generate lifelike speech from text. Unlike traditional robotic systems, these models are context-aware, meaning they automatically analyze text to detect and project emotions like anger, sadness, or alarm, adjusting intonation and pacing to match the narrative.
- Voice Cloning:
- Instant Voice Cloning (IVC): Creates a digital replica of any voice using a sample of less than one minute.
- Professional Voice Cloning (PVC): Uses 3–6 hours of audio to build a fine-tuned model that captures the deep personality and emotional range of a speaker, often used for audiobooks and branded agents.
- Speech-to-Text (STT) — "Scribe": A flagship model for transcribing audio into text with high accuracy. It supports real-time transcription (sub-150ms latency) and features like speaker diarization and character-level timestamps.
- AI Dubbing and Localization: An automated pipeline that translates video and audio into over 90 languages. It uses "sync-aware" translation to match original starts, stops, and pacing while preserving the original speaker's performance (tone and emotional intent) across languages.
- Conversational AI / Voice Agents: A platform for building interactive agents capable of real-time, two-way dialogue. It utilizes specialized low-latency models like Flash v2.5 (~75ms latency) to support natural turn-taking.
- Generative AI for Music and Sound Effects: "Eleven Music" generates studio-quality tracks from natural language prompts. This model was built through licensing deals with major labels, allowing for full commercial usage rights.
- Audio Enhancement and Forensic Tools:
- Voice Isolator: Removes background noise from recordings to separate speech.
- Voice Changer: Transforms existing audio into a different vocal style.
- AI Speech Classifier: A detection tool designed to identify if an audio sample originated from ElevenLabs' technology to combat deepfakes and misinformation.
- Provenance Watermarking: Embeds decodable signals into synthetic speech to enable tracing and accountability.
Long-Term Research Ambition
ElevenLabs is currently working toward a Universal Audio Model, a single integrated system that can seamlessly generate and transform any kind of sound, such as converting voice directly into music or singing into sound effects.
3. Ethical concerns
ElevenLabs is clearly aware of the dual use of their software, and they have made actual efforts to protect themselves, such as the AI Speech Classifier that will identify the use of its own AI, the provenance watermarks to track its artificial audio, the consent-checked level of Professional Voice Cloning, and the deals with the record labels rather than the copyright violation through scraping. These measures are more than mere efforts. The issue here is that certain products sold by this company cause harms that their protections cannot solve.
Consent that cannot be meaningfully withdrawn
This problem has direct implications for Professional Voice Cloning (PVC) and AI Dubbing. The PVC system creates a persistent voice model of a person using only 3–6 hours of audio that can capture their "personality and emotional range"; thus, what people give their consent to is not a fixed voice recording but a potentiality of sentences they have never spoken before. informed consent is difficult in such a situation, because it is hard to consent to something that does not even exist yet. Moreover, the problem intersects with the right to erasure, because deleting the source audio does not automatically erase the voice model. Unlike a password, which can always be changed, a voice cannot be re-created once it is cloned. AI Dubbing makes matters worse by "preserving the original speaker's performance" across 90+ languages, placing a person's emotional delivery into languages they may not speak and messages they never approved.
The responsibility gap when a voice is misused
This issue is associated with Instant Voice Cloning (IVC) and the Voice Changer. IVC is capable of cloning a voice from one minute of audio sample, which is easy to access for every individual or any celebrity whose voice is online, and the Voice Changer retargets the speech in another style. When this output is used for any scam call, any form of fraud, or defamation, there will be uncertainty over the responsible person since it may involve the user of the service that cloned the voice, the platform that provided the tool, and the model that generated the speech. This is what constitutes the responsibility gap. With the help of AI Speech Classifier and watermarks, it is possible to trace the output, but tracing is different from accountability as it only traces the origin of the tool and not the guilty party.
Erosion of trust in audio itself
This concern arises less from any single product than from context-aware TTS and the Conversational AI voice agents working together at scale. These produce synthetic speech that carries convincing emotion (anger, alarm, warmth) in real-time, two-way dialogue. The harm is not one deception. It is a background effect. As synthetic voices become indistinguishable from real ones, recorded audio stops working as evidence. This damages trust as a shared resource, rather than harming one individual. Any recording can now be dismissed as fake. Any fake can be passed off as real. So "hearing it yourself" loses its force in journalism, courts, and personal life. Worse, an emotionally persuasive agent edges toward manipulation. It is a voice engineered to sound sincere with no inner state behind it. Watermarking is a genuine attempt to hold the line. But it only works if every producer adopts it, which no single company can guarantee. The cost is collective. It is borne by everyone, not just the tool's users.
4. Recommendations
-
(Consent) ElevenLabs should treat consent as ongoing, not one-time. Right now the company asks once and assumes it holds forever. That does not fit a tool that keeps generating new speech. Any cloned person needs a way to revoke their voice model at any point. Revocation must delete the trained model itself, not just the source audio. Deleting the recording means little if the model still runs. A written confirmation should follow, so the person knows it is gone. Cloning licenses also need expiry dates. Consent should lapse by default, and be renewed on purpose. Silence cannot count as agreement to continue. Before any renewal, the speaker should see where their voice has already been used. A person cannot judge future use without seeing past use. Dubbing needs separate sign-off per language and per project. A voice approved for an audiobook is not approved for an ad. A performance approved in English is not approved in forty other languages. Any use that carries the speaker's emotional delivery into unfamiliar content should be flagged for them. And people need a way to refuse specific outputs, not just the tool as a whole.
-
(The responsibility gap) The more prone to abuse a technology is, the stricter should be the requirements for it. The permission form on the IVC tool is not sufficient, as this does not prove the identity of its user. The voice should be cloned only when the user manages to prove that it belongs to him or that he received the speaker’s explicit permission for that. This would eliminate the gap through which anyone could clone any voice based on one minute of audio taken from the internet. It is also important to have a log of users who used which technology. If any abuse happens as a result of a voice cloning, it should be possible to name the user behind the act, rather than blaming the platform. This is different from tracking the tool – the watermark tells you the model, while the log tells you the user.
-
(On erosion of trust) The disclosure has to be inevitable. All outputs should come with an unavoidable watermark. A watermark that can be turned off does nothing but protect the bad guys. The mark has to withstand editing, compression, and even re-recording. A mark that gets destroyed after the moment audio becomes modified is not a security measure at all. Marks created by a single company do not work. ElevenLabs should work on introducing a universal watermark format in the industry. Detection will work only if everybody in the industry will use the same kind of the watermark. A watermark, which only ElevenLabs will be able to detect, will make most of the synthetic audio undetectable. The voice agent has to reveal itself as an artificial intelligence from the very beginning of the conversation. No person can consent to talk to the machine that pretends to be a human being.