The short version: the vowels i and u go silent between two voiceless consonants or at the end of a word after one. That single rule turns です into des, ました into mashta, 好き into ski and 靴 into k'tsu. Add the four faces of ん and three consonants romaji lies about, and spoken Japanese stops being a different language from the one you read.
There is a particular disappointment most learners meet in month two. You have done the kana. You can read あ い う え お without hesitation. You say 先生、元気ですか carefully, syllable by syllable, and it sounds right to you.
Then you hear an actual Japanese person and half the syllables are simply not there. です has no “su”. ました has no “shi”. 好きです sounds like ski des. And the conclusion nearly everyone draws is the wrong one: they speak too fast, or they are slurring.
They are not. Japanese has a phonetic layer that the kana do not write down, and it is as regular as spelling rules in any language. Three systems account for almost all of the gap. None of them takes more than ten minutes to understand, and understanding them fixes both your accent and — more importantly — your listening.
System 1 · Vowel devoicing
母音の無声化 · the big oneVoiced sounds are made with the vocal cords vibrating. Put a finger on your throat and say “zzz”, then “sss” — the first buzzes, the second does not. Japanese vowels are normally voiced. But i and u lose that voicing in two specific environments, and what is left is the mouth shape without the buzz: a puff of breath where a vowel should be.
Environment A — sandwiched between two voiceless consonants. The voiceless set is k, s, sh, t, ch, ts, h, f, p.
Environment B — at the end of a word, after a voiceless consonant. This is the です / ます case.
Final u after s, at the end of the word — Environment B.
u trapped between s and k — Environment A. This is why 好きです sounds like ski des.
i trapped between sh and t. Every polite past-tense verb you will ever hear does this.
Now the pattern, on the vocabulary you already have:
| Written | What the book implies | What you actually hear | Why |
|---|---|---|---|
| です | desu | des | final u after s |
| ます | masu | mas | final u after s |
| ですか | desu ka | des ka | u between s and k |
| 好き | suki | ski | u between s and k |
| 靴 | kutsu | k'tsu | u between k and ts |
| 人 | hito | h'to | i between h and t |
| 学生 | gakusei | gaksei | u between k and s |
| 失礼 | shitsurei | sh'tsurei | i between sh and ts |
| 少し | sukoshi | s'koshi | u between s and k |
| 薬 | kusuri | k'suri | u between k and s |
| 机 | tsukue | ts'kue | u between ts and k |
| ふたつ | futatsu | f'tatsu | u between f and t |
| 聞きます | kikimasu | kikimas | final u after s |
| でしょう | deshou | desho | i between sh and ou |
System 2 · ん is four different sounds
撥音 · the one that hides word shapesYou learned ん as “n”. It is actually a chameleon that takes its shape from whatever comes next — and Japanese speakers make the switch unconsciously, without ever noticing there is a choice.
m
before p, b, m散歩 → sampo新聞 → shimbunn
before t, d, n, z, rみんな → minna案内 → annaiŋ
before k, g — the ng in singer日本語 → nihoŋgoりんご → riŋgoN
word-final, or before a vowel / s / h / w — nasal, no tongue contact本電話This is why nineteenth-century romanisation gave us shimbun and Namba — it was transcribing what people actually said. It is also why 日本 and 日本語 do not contain quite the same sound, though no Japanese speaker would say they differ.
For listening, the practical point is length. ん is a full mora. 三 is two beats; 今年 is three; 本当 is four. Learners who treat ん as a half-swallowed consonant at the end of the previous syllable end up with a rhythm that is consistently one beat short, and rhythm is what native listeners key on first.
System 3 · The consonants romaji misrepresents
four sounds that are not what the letters sayら り る れ ろ is not R and not L. It is a single tap: the tip of the tongue touches the ridge behind your upper teeth once and releases immediately. The closest thing in English is the tt in the American pronunciation of butter, or the dd in ladder. Not the held contact of L, not the bunched tongue of English R. Try 来年 with a butter-tap and it will suddenly sound correct.
ふ is not “fu”. English F puts the top teeth on the lower lip. Japanese ふ uses both lips and no teeth at all — the shape of blowing out a candle. This is why 富士山 sounds softer than Foo-ji, and why some older romanisations wrote Hujisan.
つ is the ts of “cats”, a single unit, not t followed by s. English speakers can produce it easily at the end of a word and struggle at the start, so practise from cats → tsu, borrowing the sound you already own.
を is pronounced exactly like お. The separate character survives as a grammatical marker only; in modern standard Japanese there is no w in it. Textbooks that romanise it “wo” are teaching spelling, not sound.
Why this is really a listening problem
Everything above sounds like it belongs to accent. It does not. It belongs to comprehension, and here is the mechanism.
Understanding speech means matching an incoming sound to a stored form. If the stored form is wrong, the match fails — not because you do not know the word, but because the word you know is not the word you heard. A learner listening for su-ki-de-su hears skides and finds nothing. The word was N5 vocabulary. The failure was phonetic.
And each failure costs time. Japanese arrives at roughly seven to eight moras a second; a half-second spent resolving skides means the next two words go by unprocessed. That cascade — one missed word taking three more with it — is what people describe as “Japanese is too fast”. It is rarely speed. It is a bad lookup table.
How to train it in ten minutes a day
Say the beats, not the syllables. Tap the table with one finger per mora while you speak: で·す is two taps, ま·し·た is three, さ·ん·ぽ is three. Devoice the vowel but keep the tap. This single habit fixes rhythm and devoicing at the same time.
Shadow short lines, not long ones. Take one native sentence, five to eight moras, and repeat it immediately after the speaker until your version and theirs overlap. Shadowing works precisely because it forces your production to match your perception — you cannot keep hearing de-su if your mouth keeps producing des. The common mistakes are covered in the shadowing trap.
Use dictation to test the mapping in reverse. Shadowing trains sound → mouth. Dictation trains sound → written form, which is the direction comprehension actually runs. Hearing gaksei and writing 学生 is the exact skill that was broken; ten minutes of it a day repairs faster than hours of passive listening.
This is the whole reason Kanjijo carries conversation with native audio, shadowing with a microphone, and a separate dictation track rather than one generic “listening” feature — they train three different directions of the same mapping. Alongside them sit the JLPT listening exercises with full transcripts, so that when a line does not resolve you can see precisely which syllable went missing instead of replaying it blindly.
Quick reference
- Devoicing rule
- i and u lose voicing between two voiceless consonants (k, s, sh, t, ch, ts, h, f, p) or word-finally after one.
- Canonical examples
- です→des · ます→mas · ました→mashta · 好き→ski · 靴→k'tsu · 人→h'to · 学生→gaksei · 少し→s'koshi · 薬→k'suri.
- Devoicing blockers
- The pitch-accented mora keeps its voice; adjacent devoiceable syllables usually alternate (靴下 = k'tsushita). Weaker in Kansai than in Tokyo.
- The beat never disappears
- A devoiced vowel still occupies one full mora. です = 2 beats.
- The four ん
- m before p/b/m · n before t/d/n/z/r · ŋ before k/g · a back nasal word-finally or before vowels, s, h, w. Always one full beat.
- Consonant corrections
- ら-row = a single tap (American butter) · ふ = both lips, no teeth · つ = the ts of cats · を = identical to お.
- Why listening fails
- The stored form does not match the spoken form, so the lookup misses and the next words are lost while you recover.
Related Reading on Kanjijo
Frequently Asked Questions
Vowel devoicing (母音の無声化). The final u follows s, a voiceless consonant, and sits at the end of the word, so the vocal cords stop vibrating and what remains is a puff of breath. The syllable still takes its full beat — です is two moras to a Japanese ear — so devoice the vowel but keep the timing.
Only i and u, in two environments: between two voiceless consonants (k, s, sh, t, ch, ts, h, f, p) — 好き→ski, 人→h'to, 学生→gaksei — and word-finally after one, as in です and ます. Devoicing is blocked on the pitch-accented mora, alternates when two candidates are adjacent, and is weaker in Kansai than in Tokyo.
Four, chosen by what follows: m before p/b/m (散歩 sampo, 新聞 shimbun), n before t/d/n/z/r (みんな), the ng of singer before k/g (日本語), and a back nasal with no tongue contact at the end of a word or before vowels, s, h and w (本, 電話). It always takes one full beat.
Neither — it is a single tongue tap, closest to the tt in the American pronunciation of butter. Holding the tongue in place, as English L does, is the most common accent giveaway. Romaji misleads the same way for ふ (both lips, no teeth), つ (the ts of cats) and を (pronounced exactly like お).
Because the forms stored in your head came from kana rather than from speech, so incoming sound does not match. Listening for de-su and su-ki means missing des and ski, and each miss costs the moment you needed for the next word. Fix the three systems, then shadow short lines and do dictation — the vocabulary was never the problem.
Train the sound, not just the spelling
Kanjijo attacks this from three sides: conversation with native audio, shadowing so your mouth matches what you hear, and a dictation track that forces sound back into written Japanese. Add JLPT listening practice with full transcripts, JLPT reading, complete mock tests, the full N5–N1 kanji, vocabulary and grammar path with exclusive mnemonics for every kanji and every word, an SRS engine, home, lock screen and interactive test widgets, an OCR scanner, and lessons built from any Japanese video you paste in. Free on iPhone.
Download Kanjijo FreeLearning more than Japanese? Our sister apps use the same method - free:
中 Hanzijo - Learn Chinese (HSK 1–9)한 Hanguljo - Learn Korean (TOPIK 1–6)One Premium account unlocks all three apps.