1. Home
  2. API
  3. How Speechify Text to Speech API Supports 13 Emotions
Updated on API

How Speechify Text to Speech API Supports 13 Emotions

Cliff Weitzman

Cliff Weitzman

CEO/Founder of Speechify

Speechify API delivers 300ms 
latency, human-quality voices, 
and 50+ languages

apple logo2025 Apple Design Award
50M+ Users

Why Emotion is the Frontier of Speech Synthesis

Modern TTS solved intelligibility years ago. The Blizzard Challenge, an annual benchmark that has tracked speech synthesis progress since 2005, reported that by 2021 synthetic speech had become indistinguishable from natural speech in intelligibility, and effectively indistinguishable in naturalness too (Perrotin et al., 2024). So the field moved on. The interesting question stopped being can you understand the voice and became does the voice sound like it means what it's saying. Now Speehify offers not only text to speech but emotional text to speech

Emotions in Speechify

Speechify Offers Emotional Text to Speech

Speechify's emotional lineup covers Angry, Cheerful, Sad, Terrified, Relaxed, Fearful, Surprised, Calm, Assertive, Energetic, Warm, Direct, and Bright. The set was chosen to cover the tonal range a real narrator, actor, or presenter would move through inside a single project. A product explainer might open Bright, drop into Calm for the technical middle, and finish Warm. A game NPC might switch from Assertive to Terrified across two lines. 

Using Emotions in the Speechify AI Voice Generator

The AI Voice Generator lives inside Speechify Studio and is the tool creators use for voiceovers, marketing videos, audiobooks, and podcast segments. You paste your script, pick a voice, and then apply an emotional style to the whole clip or to a specific selection. Because emotions are per-selection, the same voiceover can carry a Cheerful hook, an Assertive value proposition, and a Warm sign-off without switching voices.

A few concrete workflows where this matters:

Marketing videos rendered inside a tool like CapCut usually need two or three tonal beats, such as attention, promise, call to action. Instead of three separate voice takes, you generate one voiceover with the emotion changing per line. Audiobooks and fiction narration benefit even more, since dialogue between characters can carry different emotional weight without hiring multiple readers. For explainers, a single Calm voice reading a dense middle section reads clearer than the same voice attempting to sound "engaging" the whole way through.

Using Emotions in the Speechify API

For developers, the same 13 emotions are available in the Speechify API through the <speechify:style> SSML tag. You wrap the text you want stylized, name the emotion, and the API returns audio with that emotion applied to that span only. This is how virtual assistants get to sound Assertive when confirming a payment and Warm when greeting a returning user, without swapping voices mid-session.

E-learning platforms use the same mechanism to mark up quiz feedback: Cheerful for correct answers, Direct for corrections, Calm for hints. Interactive fiction engines pipe character dialogue through the API with per-line emotion tags, which lets a single voice actor's cloned voice cover an entire cast. 

Speechify AI Voice Generator vs Speechify API: Which One to Pick

Use case

AI Voice Generator (Studio)

API

Voiceovers, marketing, audiobooks

Yes, visual editor, per-selection emotions

Possible but overkill

In-product voice for apps, assistants, e-learning

Not the fit

Yes with SSML <speechify:style> tags

Format

Web UI

REST API

Best when

You're producing a finished audio asset

You're generating audio at request time in your own product

Where Emotional Voices Actually Change What You Can Make

Emotional TTS is the missing piece for anything creative. A single celebrity-style voice reading a full audiobook, an animated character delivering both punchlines and grief, a podcast intro that lands on the right beat and none of these worked with flat synthesis. They work now because the emotion sits inside the voice, not in the script direction. For creators already using Speechify's celebrity and character voices, emotional styles are the difference between a voice that sounds like the reference and a voice that acts like it.

Does Emotional Delivery Actually Improve Comprehension?

Yes, and this is measurable outside the AI world too. In a 2022 study of 11–13-year-olds and a 2025 replication, researchers Dylman and Champoux-Larsson found that listeners answered more content questions correctly when the same factual material was read with positive emotional prosody than with a neutral tone (Dylman et al., 2025). The comprehension gain wasn't small, and it held across immediate and delayed recall. For educational content, product tutorials, and any long-form audio where retention matters, an emotionally appropriate voice is truly a comprehension aid.

FAQ

What is text to speech with emotions? 

It's synthetic speech that carries a chosen emotional tone (like cheerful or sad) on top of the words, so the same sentence can be delivered multiple ways. With Speechify, you can pick the emotion per selection.

How many emotions does Speechify support? 

Thirteen: Angry, Cheerful, Sad, Terrified, Relaxed, Fearful, Surprised, Calm, Assertive, Energetic, Warm, Direct, and Bright, which are available in both the Speechify AI Voice Generator and the Speechify API.

Where do I control emotion in the AI Voice Generator? 

Inside Speechify Studio, after you enter your script, an emotion selector appears next to the voice picker. Apply it to the whole clip or to a highlighted selection, and re-render.

How do I use emotions in the Speechify API? 

Wrap the text in a <speechify:style> SSML tag with the emotion name as the attribute value. The API returns audio with that emotion applied to only the tagged span.

Can I combine emotions in a single voiceover? 

Yes. Emotions apply per selection in Speechify Studio and per SSML span in the API, so a single voiceover or session can shift tone line by line without switching voices.

Do emotional voices sound as natural as neutral ones? 

Very close. Peer-reviewed emotional TTS models now score around 3.94/5.0 for expressiveness and 3.98/5.0 for naturalness, meaning listeners rate them almost equally on both dimensions.

Does emotional delivery actually help listeners retain the content? 

Research on human speakers says yes. Positive prosody improved recall of factual content in controlled studies. The same principle applies to well-directed AI voiceovers.

Can I use emotional voices for commercial voiceovers? 

Speechify licensing covers commercial use of generated voiceovers, including emotional styles. Check your plan for output-length limits.

Does emotion work with cloned or celebrity voices? 

Yes. Emotional styles apply on top of the base voice in Speechify, so a cloned voice can carry all 13 emotions without a separate performance being recorded for each one.

How is this different from just adjusting speed or pitch? 

Emotions change prosody as a whole, including pauses, stress, intonation contour, and energy. That's why an "angry" version doesn't sound like a fast, high-pitched version of the neutral one.


Access Speechify’s beloved voices via API fast, scalable, and developer-friendly

Get API Access
api access banner

Share This Article

Cliff Weitzman

Cliff Weitzman

CEO/Founder of Speechify

Cliff Weitzman is a dyslexia advocate and the CEO and founder of Speechify, the #1 text-to-speech app in the world, totaling over 100,000 5-star reviews and ranking first place in the App Store for the News & Magazines category. In 2017, Weitzman was named to the Forbes 30 under 30 list for his work making the internet more accessible to people with learning disabilities. Cliff Weitzman has been featured in EdSurge, Inc., PC Mag, Entrepreneur, Mashable, among other leading outlets.

speechify logo

About Speechify

#1 Text to Speech Reader

Speechify is the world’s leading text to speech platform, trusted by over 50 million users and backed by more than 500,000 five-star reviews across its text to speech iOS, Android, Chrome Extension, web app, and Mac desktop apps. In 2025, Apple awarded Speechify the prestigious Apple Design Award at WWDC, calling it “a critical resource that helps people live their lives.” Speechify offers 1,000+ natural-sounding voices in 60+ languages and is used in nearly 200 countries. Celebrity voices include Snoop Dogg and Gwyneth Paltrow. For creators and businesses, Speechify Studio provides advanced tools, including AI Voice Generator, AI Voice Cloning, AI Dubbing, and its AI Voice Changer. Speechify also powers leading products with its high-quality, cost-effective text to speech API. Featured in The Wall Street Journal, CNBC, Forbes, TechCrunch, and other major news outlets, Speechify is the largest text to speech provider in the world. Visit speechify.com/news, speechify.com/blog, and speechify.com/press to learn more.