Kikitai

Japanese doujin audio, picked by how much Japanese you actually need

How the Japanese rating is measured, including the threshold I got wrong on the first try

Every other site that tells you what a Japanese voice work is like is telling you an opinion. So is this one, in the end. The difference is that here you can check the arithmetic.

The rating comes from four numbers, and I am going to show you all four, what they mean, where I cut them, and the place I already had to move a cutoff because it was wrong.

What gets measured

The free material only. If a work has a streaming sample or a trial download, that audio gets transcribed and measured. If it has neither, it does not get a rating, and the box says so rather than guessing.

Number What it is Why it matters
Speech ratio Share of the runtime that is actual speech, silence and non-speech removed The blunt version of "is she talking or not"
Characters per minute Transcribed Japanese characters divided by minutes Distinguishes dense dialogue from occasional murmuring
Unique bigram ratio How much the vocabulary actually moves The important one, see below
No-speech probability How confident the transcriber is that a segment is not speech at all Cross-check on the first number

The third number is the one that does the work

Here is the trap. Transcription software will happily turn moaning into text. Feed it forty minutes of breathy vocalisation and it will hand you thousands of characters. Judge by character count alone and you will call that a dialogue-heavy work, which is exactly backwards.

So the fix is to look at whether the vocabulary is moving. Chop the transcript into overlapping two-character pieces and count how many are distinct. Real speech wanders across a wide range. Non-verbal vocalisation loops on a handful of sounds and the ratio collapses.

That is what separates "she is talking" from "she is making sounds", and it is the whole reason this can be done by machine at all.

Where I put the cutoffs

  • Low: under 35% speech, or under 110 characters per minute. Also low if the vocabulary ratio drops under 0.30 while speech is high, because that combination only happens with sustained non-verbal vocalisation.
  • Medium: everything between that and 260 characters per minute.
  • High: 260 characters per minute and above, with vocabulary spread to match.

The one I got wrong

My first version treated a vocabulary ratio under 0.45 as the non-verbal signature.

Then I measured this:

RJ216777

Nipple Training, for People Already Hooked on Dry Orgasms

メスイキ中毒者のためのメス乳首化調教

Circle
Chastity Fancier 性的禁欲愛好家
List price
¥847 as of Aug 2026
Copies sold
17,685
Rating
4.67 from 5,527 ratings
Try before buying
Downloadable trial (a zip from the shop, no streaming player on the page)
Japanese needed
Some — There is a situation to follow. Knowing the setup beforehand covers most of it.Measured, not guessed96% of the runtime is speech at 197 characters per minute. Dense enough that there is something to follow, slow enough that the situation carries a lot of it. 96% speech, 197.2 chars/min, 13 min of trial download, 2026-08-03. faster-whisper small, VAD filtered

Download the free trial on DLsite A zip you unpack yourself, free

Thirteen minutes of trial audio, and it is 96% speech. The opening line transcribes as a mistress at a club counter explaining the house rules to a customer, at length. Whatever this is, it is not moaning.

It scored 0.45. One thousandth away from being labelled the exact opposite of what it is.

So 0.45 was not the boundary between speech and vocalisation. It was a perfectly normal value for a single speaker delivering a long monologue, and I had drawn the line straight through the middle of ordinary dialogue. A ratio drops when one voice talks continuously in one register, regardless of whether words are involved. On its own it means very little.

The cutoff is now 0.30, and it only applies when the speech ratio is already high, so it can only ever fire on the case it was designed for.

I am writing this down rather than quietly editing the number because a threshold that has been tested against exactly one work is not a threshold, it is a guess with decimal places. These will move again as more works go through, and when they do I will say so here.

What this cannot tell you

Whether the dialogue is any good. Whether the situation is obvious enough to follow from tone. Whether you will like the voice.

Instruction-led works in particular are much easier to follow without the language than their character count suggests, because the situation explains itself. The measurement does not know that. Where it matters, I say so in the article.

What the number does replace is the part nobody could check: someone asserting they listened and forming a view. You can run these numbers yourself on the same free trial and get the same answer, which is the only reason to trust them.

Prices, sales counts and review counts in this article are from the date shown next to them. Japanese shops run sales constantly, so check the store page for what it costs today.