What is this?
AI audio spans three revolutions at once. It generates music from a text prompt, voices indistinguishable from humans, and real-time conversation. It has produced chart-topping AI songs and an explosion of audiobooks and voice agents. It also brought the first major copyright verdicts.
Key tools & players
- Suno, Udio — full songs (vocals, lyrics, production) from a prompt
- ElevenLabs — the leader in lifelike speech and voice cloning
- OpenAI / Google voice modes — real-time conversational voice
- Open source: YuE2 for full songs, MusicGen (Meta), Whisper for transcription
Milestones
- 2022-2024 — Whisper made transcription free, then Suno and Udio shipped radio-quality songs and the labels sued
- 2025 — Real-time voice agents replace call-center queues; voice cloning scams rise
- 2026 — Munich court rules against Suno in Europe's first big AI music copyright case (our coverage)
- Jul 2026 — OpenAI puts DeepMind's invisible SynthID watermark in ChatGPT voice and opens a verification API (our coverage)
- Aug 2026 — ByteDance's SeedRealtime takes full-duplex audio-visual: it watches, listens and speaks at once, live in the Doubao app
- Aug 2026 — Spotify drops badged AI acts from recommendations and Australia's ARIA bars fully AI-made tracks from the charts (our coverage)
- Sep 2026 — Suno's v6 switches to licensed catalogues while open YuE2 matches it in blind tests (our coverage)
- Sep 2026 — ElevenLabs' Music v2.5 lets even free accounts sell their tracks, with credit to the tool (our coverage)
- Sep 2026 — Gemini 3.8 Live takes first on speech-to-speech and Qwen drops audio input cost by 98% (our coverage)
- Sep 2026 — ElevenLabs' Eleven v4 clones a voice from ten seconds and speaks 90+ languages natively (our coverage)
Mini glossary
- TTS: text-to-speech — turning writing into natural voice
- Voice cloning: reproducing a specific person's voice from short samples
- Stem: an isolated track (vocals, drums) within a song