Word Frequency

Count how often each word appears. A run of Chinese, Japanese or Korean stays one token unless you split it.

Which words show up, and how often

A word counter tells you the length of the text. This page tells you which tokens repeat. The first line is a total. Each following line is a word, a tab, and a count, most frequent first. Length, sentences and reading time stay on word counter.

How to use it

  1. Paste the text. Ignore case is on, so To and to share one row.
  2. Count each CJK character is off. A run such as 北京 is one token. Turn it on to count 北 and 京 separately. Hiragana, katakana and Hangul follow the same switch.
  3. Minimum length drops tokens shorter than that many characters. The default is 1.
  4. Show example loads To be, or not to be. plus Paris Paris London. With the defaults, to and be are 2, and paris is 2.

What counts as a word

Latin words are letters and digits. An apostrophe inside a word stays, so don't is one token. Punctuation and spaces split tokens. A continuous run of Han, kana or Hangul is one token until a space or other script, unless you split CJK. Ignore case stores the lowercase form, using the locale's lowercasing. Rows sort by count, then by the word. Tokens shorter than the minimum are left out of both the list and the total. The summary line uses the labels for unique words and token count.

Where people use it

  • Seeing which words dominate a paragraph before you edit it.
  • Checking that a Chinese place name was pasted twice, as one token rather than two characters.
  • Listing terms with a minimum length so one-letter noise drops out.

Questions

北京 was one row, not two characters.

That is the default. Turn on Count each CJK character if you want a count per character. Spaces still separate runs.

The and the were merged.

Ignore case is on. Turn it off to keep capitals as a different row.