Hey, it’s Hoda.
When people think of AI voice synthesis tools, ElevenLabs and Speechify are usually the first names that come up — but Fish Audio has quietly been building a real presence lately.
The company claims its voice sounds more natural than ElevenLabs’, and backs that up with blind test data, which is hard to ignore. Two things in particular caught my attention: the emotion tag system, and just how good the voice quality is in languages other than English — more on that below, since I put it through its paces in my own language.
I’ve already reviewed ElevenLabs and Speechify (links at the end), and Fish Audio sits in the same category as a genuine alternative worth weighing against both — so this time I’m digging into it properly.
What Is Fish Audio?

Fish Audio is an AI voice platform built and run by Hanabi AI Inc. It covers a wide range of voice-related features: text-to-speech (TTS), voice cloning, speech recognition (STT), voice translation, and audio separation.
At the core is its own Fish Audio S2 model, which recorded roughly a 60% win rate in a head-to-head public benchmark against ElevenLabs. That figure is based on more than 5,000 blind tests — evaluations where users don’t know which tool produced which sample.
It officially supports 30+ languages (some parts of the site claim “80+ languages” — that’s likely because the newest flagship model, S2.1 Pro, specifically covers 83 languages, while older models in the lineup cover a smaller set closer to 30). The voice library holds over 2 million community-created voices. And yes, Japanese is one of the supported languages — I’m Japanese myself, so I tested it extensively in my own language, and it held up well (more on that further down).
A practical example
On my YouTube channel, I walk through how to connect the Fish Audio API to Google Sheets to batch-convert text into MP3 audio files (narrated in Japanese, but easy enough to follow along with visually even if you don’t speak the language).
Core Features
Text-to-Speech (TTS)

The basic function is converting typed text into natural-sounding speech. Three things stand out here.
Fine-grained emotional control through emotion tags is the biggest differentiator. You can drop tags like [excited], [whispering], [sad], or [laughing] directly into your text, and the emotional tone shifts phrase by phrase. There are 64+ tags available, giving you noticeably finer control than ElevenLabs’ slider-based approach.
Low-latency streaming is another strength — first-output latency is under 200ms, which points to a design built for real-time voice agents and conversational AI.
Voice stability over long-form content also holds up well — the character’s voice doesn’t drift or degrade much over several minutes of narration, which matters for audiobooks or YouTube narration.
Voice Cloning

You can clone a voice from just 15 to 30 seconds of sample audio. That’s a shorter requirement than most industry standards, which is a real advantage — but noise or multiple overlapping speakers in the recording will significantly hurt clone accuracy. A clean, quiet recording is essentially a prerequisite.
▼ I cloned my own voice here. Not to toot my own horn, but it’s a little unsettling how close it sounds.
Cloned voices work across languages too — you can record a sample in Japanese, for example, and use it to generate English TTS output, which is a genuinely useful cross-lingual trick.
You can create a clone of your own voice for free, but if you want to set it to private or unlisted under the “Type” setting, you’ll need to upgrade to a paid plan. On the free plan, any clone you create has to be public — worth knowing before you start.

Speech Recognition (STT) and Other Features
Fish Audio also includes transcription, audio separation (splitting vocals from instrumentals), sound effect generation, and voice translation. TTS and voice cloning are the main draw, but having a decent chunk of the voice workflow handled in one place is genuinely convenient.
API

Fish Audio offers a REST API along with Python and JavaScript SDKs, backed by solid developer documentation. API pricing for TTS runs $15 per million UTF-8 bytes, which is significantly cheaper than ElevenLabs at the equivalent rate.
Fish Audio has also made its flagship S2.1 Pro model free for developers via API, under a fair-use policy — effectively an unlimited free text-to-speech API, at least for now. It’s a genuinely unusual move in this space, since state-of-the-art voice quality is almost always paywalled elsewhere.

Pricing
Here’s how the plans break down, based on Fish Audio’s official pricing page (USD, checked September 2026).
| Plan | Monthly | Annual (per mo.) | Billed annually | Credits/mo | Generation time |
|---|---|---|---|---|---|
| Free | $0 | $0 | $0 | 8,000 | Up to 7 min |
| Plus | $15 | $11 | $132/year | 250,000 | Up to 200 min |
| Pro | $100 | $75 | $900/year | 2,000,000 | Up to 1,620 min |
| Max | $999 | $749 | $8,988/year | 25,000,000 | Up to 6,250 min |
| Enterprise | Custom | Custom | — | — | — |
On commercial use: the free plan restricts you to non-commercial use. If you’re running a monetized YouTube channel or using it for company promotional content, you’ll need Plus or higher.
Character limit per generation: 500 characters on Free, 15,000 on Plus, 30,000 on Pro. If you’re processing long scripts in bulk, Plus or higher is the practical minimum.
-
Free plan requires no credit card to try
-
Plus plan is $11/month on annual billing — solid value
-
Annual billing saves roughly three months' worth (about 27% off)
-
API pricing is among the cheapest in the industry ($15 per million bytes)
-
Free plan caps out at 7 minutes/month — really just enough to kick the tires
-
Commercial use requires Plus or higher
-
Pro plan runs $75–100/month, pricey for most individual creators
An Honest Look at Voice Quality
Strength: emotion and naturalness
A third-party review site (DIY AI) evaluated Fish Audio against its own dataset and gave it the highest score of any category on emotional range: 8.9/10. Its overall score (8.7/10) came in just behind ElevenLabs, and both naturalness and voice-clone accuracy also scored well.
I’m Japanese, so I put the Japanese output through a real test myself — and it held up impressively well, sounding more natural than ElevenLabs’ Japanese output in several places I tried. That’s a decent signal on its own: Japanese is a genuinely hard language for TTS engines to get right, so if it can handle that reasonably well, there’s a good chance it’ll sound natural in other non-English languages too, not just English. That said, quality does vary noticeably by voice, so it’s worth previewing a few before settling on one for your use case.
Weakness: noise handling
The noise-handling score came in lower than other categories, at 7.8/10. If your source recording for a voice clone has noise or reverb, that tends to bleed into the generated output. It’s worth pairing Fish Audio with a cleanup tool (Adobe Podcast Enhancer, Rovoice, or similar) rather than feeding it raw recordings.
Fish Audio vs. ElevenLabs
| Category | Fish Audio | ElevenLabs |
|---|---|---|
| Overall score (DIY AI) | 8.7/10 | 8.9/10 |
| Emotional expression | 8.9/10 (best-in-class) | 8.7/10 |
| Voice naturalness | 8.8/10 | Industry-leading |
| Voice clone accuracy | 8.8/10 | Comparable |
| Non-English quality | Strong, per hands-on testing | Mixed — varies by voice |
| API pricing | $15 / 1M bytes (best value) | $165 / 1M bytes |
| Free plan | No credit card needed | More restrictive |
Who It’s a Good Fit For (and Who Isn’t)
Fish Audio is especially well-suited to content creators who need expressive, emotionally rich character voices or narration, creators publishing in languages other than English who care about voice quality, and engineers who want to build voice features into a product without a huge API bill.
It’s less of a fit in a couple of cases. If you need to reliably mass-produce clean, corporate narration at scale — compliance training, brand-voice content — ElevenLabs is easier to govern and manage. It’s also not the right tool if you’re planning to build voice clones from samples recorded in noisy environments.
A Few Things to Keep in Mind
Rights and likeness: cloning a celebrity’s or voice actor’s voice without permission and using it commercially is a real likeness/publicity-rights issue. Not every voice available in the community library is necessarily cleared for commercial use, so check the rights before you use one.
Disclosing AI voice use: if you use generated audio in a video or similar content, it’s good practice — and generally considered the ethical baseline — to disclose that AI voice was used.
Deleting your account: any voice models you’ve built and unused credits are forfeited when you delete your account. Export anything you need and cancel any paid plan before you do.
FAQ
Does Fish Audio handle languages other than English well? Yes. It officially supports 30+ languages, with the newest flagship model covering 83. I’m Japanese, and after testing it thoroughly in Japanese specifically, I can say it held up impressively well — Japanese is one of the harder languages for TTS engines to get right, so that’s a decent signal for quality in other non-English languages too. If you’re working in a particular language, it’s still worth previewing a few sample voices first, since quality varies somewhat by voice.
How much can I actually do on the free plan? You get 8,000 credits a month, which works out to about 7 minutes of generation. Not needing a credit card to try it is a plus, but commercial use isn’t allowed, each generation is capped at 500 characters, and enhanced voice-cloning features are limited. For anything beyond testing, you’ll want Plus or higher.
Is it better than ElevenLabs? Overall polish edges slightly in ElevenLabs’ favor, but Fish Audio has real advantages in API pricing, non-English quality, and the granularity of its emotion tags. Which one wins depends on your use case — trying both free tiers first is a reasonable way to decide.
What kind of audio do I need for voice cloning? 15 to 30 seconds of sample audio at minimum. For better results, use a clean, single-speaker recording with no background noise (MP3, WAV, or M4A).
Can I use it commercially? Yes, on Plus or any higher plan. The free plan is personal-use only — a monetized YouTube channel or any commercial use requires upgrading to a paid plan.
Fish Audio is worth putting on the shortlist alongside ElevenLabs and Speechify, especially if you’re creating in a language other than English. Its combination of pricing, emotion-tag control, and non-English voice quality gives it a genuine edge in a few specific use cases, and depending on what you need, it could easily become your primary tool.
