Core features
Text to Speech
Convert text into lifelike speech with the
s2.1-pro, s2-pro, and deprecated s1 models.Speech to Text
Transcribe audio with
transcribe-1-pro: multi-speaker conversations, long recordings, emotion cues, and optional timestamps.Voice Cloning
Clone a voice instantly from a clip, or train a persistent model.
Realtime Streaming
Stream audio as it generates, for voice agents and live apps.
Manage Voices
List, inspect, update, and delete your voice models.
Also in the web app
These run in the browser, no code required. See the Platform guide.Professional Voice Clone
Clone a real, verified voice at studio quality, then release it.
Voice Design
Design a voice from a plain-text description. Best for original characters.
Voice Changer
Transform existing audio into a different voice.
Story Studio
Produce multi-speaker, long-form audio: audiobooks and narration.
Music & Sound Effects
Generate music and cinematic sound effects from a prompt.
Audio Separation
Split audio into stems, and related processing utilities.
Models
These text-to-speech models power most capabilities:s2.1-pro: the recommended production model, with improved quality, latency, and throughput over S2-Pro.s2.1-pro-free: the same model at $0 for testing, prototyping, development, and smaller businesses, without TTFA or DPA guarantees.s2-pro: the previous-generation S2 model, with multi-speaker and natural-language expression control.s1: the previous generation, with(parenthesis)emotion tags. S1 is deprecated and will be retired on December 31, 2026. After that date, requests that specifys1are served bys2.1-pro. See Migrate from S1.
model header on POST /v1/asr:
transcribe-1-pro: the recommended model, for everything from short clips to multi-speaker conversations (speaker markers and speaker turns) and long recordings, with inline emotion cues. Sendmodel: transcribe-1-proexplicitly.transcribe-1: general transcription of short recordings. A request whosemodelheader is missing or not an exact match is served by this model.
Pick your path
Use the web app
No code: generate audio, clone voices, and produce projects in your browser.
Build with the SDK
The Python library for your application.
Call the API
Raw REST and WebSocket endpoints for any language.
Use your AI coding agent
Install the Fish Audio skill so your agent writes correct code.

