How the Japanese rating is measured, including the threshold I got wrong on the first try
Every other site that tells you what a Japanese voice work is like is telling you an opinion. So is this one, in the end. The difference is that here you can check the arithmetic.
The rating comes from four numbers, and I am going to show you all four, what they mean, where I cut them, and the place I already had to move a cutoff because it was wrong.
What gets measured
The free material only. If a work has a streaming sample or a trial download, that audio gets transcribed and measured. If it has neither, it does not get a rating, and the box says so rather than guessing.
| Number | What it is | Why it matters |
|---|---|---|
| Speech ratio | Share of the runtime that is actual speech, silence and non-speech removed | The blunt version of "is she talking or not" |
| Characters per speech-minute | Transcribed characters divided by the speech time only, not the whole runtime | Distinguishes dense dialogue from occasional murmuring |
| Unique bigram ratio | How much the vocabulary actually moves | The important one, see below |
| No-speech probability | How confident the transcriber is that a segment is not speech at all | Cross-check on the first number |
The third number is the one that does the work
Here is the trap. Transcription software will happily turn moaning into text. Feed it forty minutes of breathy vocalisation and it will hand you thousands of characters. Judge by character count alone and you will call that a dialogue-heavy work, which is exactly backwards.
So the fix is to look at whether the vocabulary is moving. Chop the transcript into overlapping two-character pieces and count how many are distinct. Real speech wanders across a wide range. Non-verbal vocalisation loops on a handful of sounds and the ratio collapses.
That is what separates "she is talking" from "she is making sounds", and it is the whole reason this can be done by machine at all.
Where I put the cutoffs
Ordinary Japanese speech runs at roughly 300 to 400 characters a minute. Everything below is a fraction of that.
- Low: under 35% speech at all, or under 190 characters per speech-minute. Also low if the vocabulary ratio drops under 0.30 while speech is high, because that combination only happens with sustained non-verbal vocalisation.
- Medium: 190 to 300.
- High: 300 and above, the pace of ordinary conversation.
The one I got wrong
My first version treated a vocabulary ratio under 0.45 as the non-verbal signature.
Then I measured this:
RJ216777
Nipple Training, for People Already Hooked on Dry Orgasms
メスイキ中毒者のためのメス乳首化調教
- Circle
- Chastity Fancier 性的禁欲愛好家
- List price
- ¥847 as of Aug 2026
- Copies sold
- 17,685
- Rating
- 4.67 from 5,527 ratings
- Try before buying
- Downloadable trial (a zip from the shop, no streaming player on the page)
- Japanese needed
- Some — There is a situation to follow. Knowing the setup beforehand covers most of it.Measured, not guessed96% of the runtime is speech at 206 characters per minute of it. There is something to follow, but it is paced slowly enough that the situation carries a lot of it. 96% speech, 206.5 chars per speech-minute, 13 min of trial download, 2026-08-06. faster-whisper small, VAD filtered, density measured per speech-minute
Download the free trial on DLsite A zip you unpack yourself, free
Thirteen minutes of trial audio, and it is 96% speech. The opening line transcribes as a mistress at a club counter explaining the house rules to a customer, at length. Whatever this is, it is not moaning.
It scored 0.45. One thousandth away from being labelled the exact opposite of what it is.
So 0.45 was not the boundary between speech and vocalisation. It was a perfectly normal value for a single speaker delivering a long monologue, and I had drawn the line straight through the middle of ordinary dialogue. A ratio drops when one voice talks continuously in one register, regardless of whether words are involved. On its own it means very little.
The cutoff is now 0.30, and it only applies when the speech ratio is already high, so it can only ever fire on the case it was designed for.
I am writing this down rather than quietly editing the number because a threshold that has been tested against exactly one work is not a threshold, it is a guess with decimal places.
What three works look like side by side
Since then I have put two more through. Same measurement, same cutoffs.
| Ear-licking work | Milking driving school | Nipple training | |
|---|---|---|---|
| Speech ratio | 33% | 97% | 96% |
| Characters per speech-minute | 114 | 233 | 207 |
| Vocabulary ratio | 0.70 | 0.80 | 0.45 |
| Rating | Low | Medium | Medium |
The first column is what a genuinely language-light work looks like. Only a third of it is sound at all, and even inside that third the words come at 114 a minute, roughly a third of conversational pace. The other two are dense enough to follow.
Note the vocabulary ratio doing nothing useful here. The ear-licking work scores 0.70, higher than the dialogue-heavy one at 0.45. If I had kept my original rule, the work with almost no speech in it would have been rated identically to the one that is nothing but speech, and the monologue would have been the one flagged as non-verbal. That number only earns its place in the one narrow case it was rebuilt for.
Two things that limit this
The transcription is often wrong. This genre is full of vocabulary the model has never seen, and it produces some inventive nonsense. It mostly does not matter, because I am measuring how much speech there is, not what it says, and a garbled transcript still has the right number of characters in roughly the right places. But it does inflate the vocabulary ratio slightly, since transcription errors look like new words. Treat that column as the softest of the four.
Trial lengths vary a lot. The driving school work gave me 93 seconds to work with; the others gave twelve to thirteen minutes. A 93-second sample is a thinner basis for a judgement about a full-length work, and where that is the case it is noted on the work itself.
The measurement was wrong for a week
I had been dividing the character count by the whole runtime, including silence. That is broken in two separate ways, and it took a 24-minute work to expose both.
A work called Keigo Senpai came back at 38 characters a minute and was rated Low. The transcript is a woman explaining, at length, how to answer the office phone. Nothing about it is language-light.
The first fault: dividing by total runtime conflates "she stops talking a lot" with "she does not say much". Those are different works. Dividing by speech time only separates them.
The second fault is worse, because it scales. The longer the audio, the more the transcriber drops, and the lower the density falls, regardless of content. Lining up everything measured so far made it obvious:
| Trial length | Characters per speech-minute |
|---|---|
| 1.5 min | 233 |
| 12 min | 114 |
| 13 min | 207 |
| 24 min | 52 |
Fifty-two is not a quiet work. It is a transcript with most of it missing.
So the script now checks for that and refuses to rate anything that comes back implausibly thin. Keigo Senpai is marked not rated rather than given a number I know to be wrong.
Working out where to draw that line needed an experiment. The ear-licking work sits at 114, which is also low. Was that missing text, or a genuinely wordless work? I ran it again through a larger model: 106 became 114, a rise of eight percent. Real dropped text would have jumped. So 114 is the work, and the cutoff for "the transcript is broken" belongs below it, at 80.
That also means the ear-licking work keeps its Low, which it earned for the right reason after all: a third of it is speech, and the speech itself is sparse.
Thresholds will move again as more works go through, and when they do it will be written here.
What this cannot tell you
Whether the dialogue is any good. Whether the situation is obvious enough to follow from tone. Whether you will like the voice.
Instruction-led works in particular are much easier to follow without the language than their character count suggests, because the situation explains itself. The measurement does not know that. Where it matters, I say so in the article.
What the number does replace is the part nobody could check: someone asserting they listened and forming a view. You can run these numbers yourself on the same free trial and get the same answer, which is the only reason to trust them.
Prices, sales counts and review counts in this article are from the date shown next to them. Japanese shops run sales constantly, so check the store page for what it costs today.