Predicted emotion
SHARE OF SEGMENTSNeutral dominates; rare labels were sought separately to give every emotion five listening examples.
A closer listen to the feeling in a voice. Explore real production segments through predicted emotions, energy bands, and a spectrogram that follows every word.
Explore the clips ↓Before balancing the listening set, we measured the natural mix in a separate sample. These charts describe that reference, not the curated clips below.
Neutral dominates; rare labels were sought separately to give every emotion five listening examples.
Empirical thirds, with ties kept together. Energy is a model score, not loudness or a probability.
Five examples in each group. Click a transcript to jump to that clip. Cleaned audio plays wherever the production cleaner was applied.
All 55 clips are from distinct completed IPTV jobs. Enhanced segments use the verified cleaned episode audio; unenhanced segments use the original. Each source and segment mapping was rechecked against its production manifest.
English-only channel metadata and English transcript language tags, with final predicted MOS ≥3.5/5. Reference candidates are 4–16 seconds with at least five words. MOS is an automated quality estimate, not a human listening score.
Emotion and energy come from Hotdog. Emotion confidence describes the model’s prediction. The emotion examples favor confident predictions and varied channels; they do not represent the natural emotion mix. Transcripts are automatic and may contain errors.
All videos use a shared −90 to 0 dB spectrogram scale, display 0–12 kHz, and retain full-band 48 kHz audio without loudness normalization. A moving playhead tracks both waveform and spectrum. Energy bands use fixed reference-sample cutoffs, with examples spanning scores within each band.