🗣️ SotoSpeech

About SotoSpeech

SotoSpeech is a free speaking gym for kids. No ads, no sign-up, nothing you say is uploaded.

What it is

Children learn to speak a language by hearing real people and copying them. SotoSpeech takes a real recording — a cartoon, a podcast, a reading contest, a documentary — and turns it into sentence-sized pieces. The child hears one sentence, then says it straight back while the microphone records them, so they can hear the difference. This is called shadowing or echoing, and it is the single most useful thing a learner can do for pronunciation, rhythm and confidence.

The three echoing sections

🇬🇧 EnglishAuthentic British voices — the BBC, podcasts, audiobooks, Downton Abbey — in Received Pronunciation. Whole sentences, never broken mid-sentence.
🇨🇳 中文Everyday spoken Mandarin. Pinyin sits above every character, an English meaning under every word, and the whole line in English under the player. Each tier can be hidden as the child improves.
🇫🇷 FrançaisPeppa Pig, Bluey and the Twirlywoos in French, plus real children reading aloud at the Comédie-Française. One short sentence per line, an English meaning under every word, and the whole line in English.

The two talk-training drills

💬 Say It Five Ways. Forty everyday sentences ("I'm really tired today.", "It just didn't work."). The child says the same thing five genuinely different ways before the examples are revealed one at a time. It trains rephrasing, being specific, adding the detail that matters, and repairing a conversation when someone says "I don't understand".

🤝 Empathy & Conversation. Thirty situations in which a friend or parent says something with a feeling behind it. The child responds first, then up to five good responses are revealed. It trains the order that makes people feel heard: acknowledge the feeling, show it makes sense, ask a question that follows, offer help only when it fits.

Both run as full-screen slide shows, read aloud by a warm British voice, and remember where the child left off.

How it is made

Every recording is transcribed, cut into sentences and checked line by line — including a second transcription with a stronger model to catch words that were misheard, and a word-by-word read-through with the English translation. Chinese lessons carry pinyin generated per character and meanings per word; French lessons carry a meaning under every word. The voice on the drills is a synthetic one; the echoing lessons are always the original human recording.

Privacy

The site has no accounts and sets no tracking cookies. When the child records their voice, the recording stays in the browser and is discarded when the lesson closes; it is never sent anywhere. Anonymous visit counts are collected with Umami (outside mainland China) or Baidu Tongji (inside), so we can see which sections are used.

Who makes it

SotoSpeech is made by SOTOBY (苏托比工作室), a small studio building learning tools for its own children first. It is free and supported by readers who buy us a coffee. Write to us at [email protected].

Start practising →