HomeBlog › Why です Sounds Like Des

Why です Sounds Like “Des”

You learned de-su. Natives say des. You are not mishearing and they are not being lazy — there is a rule, it is completely regular, and nobody taught it to you.

Published September 13, 2026 · 10 min read · Level: N5

The short version: the vowels i and u go silent between two voiceless consonants or at the end of a word after one. That single rule turns です into des, ました into mashta, き into ski and くつ into k'tsu. Add the four faces of ん and three consonants romaji lies about, and spoken Japanese stops being a different language from the one you read.

Level: N5. Everything here applies from your very first sentence, and every example uses beginner vocabulary with furigana. This is the layer textbooks skip because it belongs to phonetics rather than grammar — which is exactly why it goes unrepaired for years.

There is a particular disappointment most learners meet in month two. You have done the kana. You can read あ い う え お without hesitation. You say 先生せんせい元気げんきですか carefully, syllable by syllable, and it sounds right to you.

Then you hear an actual Japanese person and half the syllables are simply not there. です has no “su”. ました has no “shi”. きです sounds like ski des. And the conclusion nearly everyone draws is the wrong one: they speak too fast, or they are slurring.

They are not. Japanese has a phonetic layer that the kana do not write down, and it is as regular as spelling rules in any language. Three systems account for almost all of the gap. None of them takes more than ten minutes to understand, and understanding them fixes both your accent and — more importantly — your listening.

System 1 · Vowel devoicing

母音ぼいん無声化むせいか · the big one

Voiced sounds are made with the vocal cords vibrating. Put a finger on your throat and say “zzz”, then “sss” — the first buzzes, the second does not. Japanese vowels are normally voiced. But i and u lose that voicing in two specific environments, and what is left is the mouth shape without the buzz: a puff of breath where a vowel should be.

Environment A — sandwiched between two voiceless consonants. The voiceless set is k, s, sh, t, ch, ts, h, f, p.

Environment B — at the end of a word, after a voiceless consonant. This is the です / ます case.

ですde-SUdes

Final u after s, at the end of the word — Environment B.

su-kiski

u trapped between s and k — Environment A. This is why 好きです sounds like ski des.

ましたma-shi-tamashta

i trapped between sh and t. Every polite past-tense verb you will ever hear does this.

Now the pattern, on the vocabulary you already have:

WrittenWhat the book impliesWhat you actually hearWhy
ですdesudesfinal u after s
ますmasumasfinal u after s
ですかdesu kades kau between s and k
sukiskiu between s and k
くつkutsuk'tsuu between k and ts
ひとhitoh'toi between h and t
学生がくせいgakuseigakseiu between k and s
失礼しつれいshitsureish'tsureii between sh and ts
すこsukoshis'koshiu between s and k
くすりkusurik'suriu between k and s
つくえtsukuets'kueu between ts and k
ふたつfutatsuf'tatsuu between f and t
きますkikimasukikimasfinal u after s
でしょうdeshoudeshoi between sh and ou
What devoicing is not. The syllable does not disappear. Japanese timing is counted in はく — moras, equal beats — and a devoiced vowel still occupies its full beat. です is two beats to a Japanese ear, not one. If you compress it into a single English-style syllable, the rhythm breaks and you will sound foreign even with the vowel correctly devoiced. Say the beat; just do not voice it.
Two refinements that separate good from very good. Devoicing is blocked when the syllable carries the pitch accent — the accented mora keeps its voice. And when two devoiceable syllables sit side by side, typically only one gives way: 靴下くつした comes out as k'tsushita, not k'tsush'ta. Devoicing is also strongest in Tokyo and eastern Japan and noticeably weaker in Kansai, which is one of the quiet reasons Kansai speech sounds “rounder” to learners.

System 2 · ん is four different sounds

撥音はつおん · the one that hides word shapes

You learned ん as “n”. It is actually a chameleon that takes its shape from whatever comes next — and Japanese speakers make the switch unconsciously, without ever noticing there is a choice.

m

before p, b, m散歩さんぽ → sampo新聞しんぶん → shimbun

n

before t, d, n, z, rみんな → minna案内あんない → annai

ŋ

before k, g — the ng in singer日本語にほんご → nihoŋgoりんご → riŋgo

N

word-final, or before a vowel / s / h / w — nasal, no tongue contactほん電話でんわ

This is why nineteenth-century romanisation gave us shimbun and Namba — it was transcribing what people actually said. It is also why 日本にほん and 日本語にほんご do not contain quite the same sound, though no Japanese speaker would say they differ.

For listening, the practical point is length. ん is a full mora. さん is two beats; 今年ことし is three; 本当ほんとう is four. Learners who treat ん as a half-swallowed consonant at the end of the previous syllable end up with a rhythm that is consistently one beat short, and rhythm is what native listeners key on first.

System 3 · The consonants romaji misrepresents

four sounds that are not what the letters say

ら り る れ ろ is not R and not L. It is a single tap: the tip of the tongue touches the ridge behind your upper teeth once and releases immediately. The closest thing in English is the tt in the American pronunciation of butter, or the dd in ladder. Not the held contact of L, not the bunched tongue of English R. Try らいねん with a butter-tap and it will suddenly sound correct.

ふ is not “fu”. English F puts the top teeth on the lower lip. Japanese ふ uses both lips and no teeth at all — the shape of blowing out a candle. This is why 富士山ふじさん sounds softer than Foo-ji, and why some older romanisations wrote Hujisan.

つ is the ts of “cats”, a single unit, not t followed by s. English speakers can produce it easily at the end of a word and struggle at the start, so practise from catstsu, borrowing the sound you already own.

を is pronounced exactly like お. The separate character survives as a grammatical marker only; in modern standard Japanese there is no w in it. Textbooks that romanise it “wo” are teaching spelling, not sound.

One optional extra: the nasal が. In conservative Tokyo speech and among older broadcasters, が in the middle of a word is nasalised — かがみ with the ng of singer. It is called 鼻濁音びだくおん and it is genuinely receding among younger speakers, so it is worth recognising and not worth drilling.

Why this is really a listening problem

Everything above sounds like it belongs to accent. It does not. It belongs to comprehension, and here is the mechanism.

Understanding speech means matching an incoming sound to a stored form. If the stored form is wrong, the match fails — not because you do not know the word, but because the word you know is not the word you heard. A learner listening for su-ki-de-su hears skides and finds nothing. The word was N5 vocabulary. The failure was phonetic.

And each failure costs time. Japanese arrives at roughly seven to eight moras a second; a half-second spent resolving skides means the next two words go by unprocessed. That cascade — one missed word taking three more with it — is what people describe as “Japanese is too fast”. It is rarely speed. It is a bad lookup table.

The good news is how small the repair is. There is no vocabulary to learn here and no grammar. Three rules, perhaps twenty example words to re-hear correctly, and the stored forms are fixed for good. Learners who patch this usually notice the change within two weeks, and it is not a gradual improvement — sentences that were mush start arriving as words.

How to train it in ten minutes a day

Say the beats, not the syllables. Tap the table with one finger per mora while you speak: で·す is two taps, ま·し·た is three, さ·ん·ぽ is three. Devoice the vowel but keep the tap. This single habit fixes rhythm and devoicing at the same time.

Shadow short lines, not long ones. Take one native sentence, five to eight moras, and repeat it immediately after the speaker until your version and theirs overlap. Shadowing works precisely because it forces your production to match your perception — you cannot keep hearing de-su if your mouth keeps producing des. The common mistakes are covered in the shadowing trap.

Use dictation to test the mapping in reverse. Shadowing trains sound → mouth. Dictation trains sound → written form, which is the direction comprehension actually runs. Hearing gaksei and writing 学生がくせい is the exact skill that was broken; ten minutes of it a day repairs faster than hours of passive listening.

This is the whole reason Kanjijo carries conversation with native audio, shadowing with a microphone, and a separate dictation track rather than one generic “listening” feature — they train three different directions of the same mapping. Alongside them sit the JLPT listening exercises with full transcripts, so that when a line does not resolve you can see precisely which syllable went missing instead of replaying it blindly.

Quick reference

Devoicing rule
i and u lose voicing between two voiceless consonants (k, s, sh, t, ch, ts, h, f, p) or word-finally after one.
Canonical examples
です→des · ます→mas · ました→mashta · 好き→ski · 靴→k'tsu · 人→h'to · 学生→gaksei · 少し→s'koshi · 薬→k'suri.
Devoicing blockers
The pitch-accented mora keeps its voice; adjacent devoiceable syllables usually alternate (靴下 = k'tsushita). Weaker in Kansai than in Tokyo.
The beat never disappears
A devoiced vowel still occupies one full mora. です = 2 beats.
The four ん
m before p/b/m · n before t/d/n/z/r · ŋ before k/g · a back nasal word-finally or before vowels, s, h, w. Always one full beat.
Consonant corrections
ら-row = a single tap (American butter) · ふ = both lips, no teeth · つ = the ts of cats · を = identical to お.
Why listening fails
The stored form does not match the spoken form, so the lookup misses and the next words are lost while you recover.

Frequently Asked Questions

Vowel devoicing (母音の無声化). The final u follows s, a voiceless consonant, and sits at the end of the word, so the vocal cords stop vibrating and what remains is a puff of breath. The syllable still takes its full beat — です is two moras to a Japanese ear — so devoice the vowel but keep the timing.

Only i and u, in two environments: between two voiceless consonants (k, s, sh, t, ch, ts, h, f, p) — 好き→ski, 人→h'to, 学生→gaksei — and word-finally after one, as in です and ます. Devoicing is blocked on the pitch-accented mora, alternates when two candidates are adjacent, and is weaker in Kansai than in Tokyo.

Four, chosen by what follows: m before p/b/m (散歩 sampo, 新聞 shimbun), n before t/d/n/z/r (みんな), the ng of singer before k/g (日本語), and a back nasal with no tongue contact at the end of a word or before vowels, s, h and w (本, 電話). It always takes one full beat.

Neither — it is a single tongue tap, closest to the tt in the American pronunciation of butter. Holding the tongue in place, as English L does, is the most common accent giveaway. Romaji misleads the same way for ふ (both lips, no teeth), つ (the ts of cats) and を (pronounced exactly like お).

Because the forms stored in your head came from kana rather than from speech, so incoming sound does not match. Listening for de-su and su-ki means missing des and ski, and each miss costs the moment you needed for the next word. Fix the three systems, then shadow short lines and do dictation — the vocabulary was never the problem.

Train the sound, not just the spelling

Kanjijo attacks this from three sides: conversation with native audio, shadowing so your mouth matches what you hear, and a dictation track that forces sound back into written Japanese. Add JLPT listening practice with full transcripts, JLPT reading, complete mock tests, the full N5–N1 kanji, vocabulary and grammar path with exclusive mnemonics for every kanji and every word, an SRS engine, home, lock screen and interactive test widgets, an OCR scanner, and lessons built from any Japanese video you paste in. Free on iPhone.

Download Kanjijo Free

Learning more than Japanese? Our sister apps use the same method - free:

中 Hanzijo - Learn Chinese (HSK 1–9)한 Hanguljo - Learn Korean (TOPIK 1–6)

One Premium account unlocks all three apps.