Count how often each word appears. A run of Chinese, Japanese or Korean stays one token unless you split it.
A word counter tells you the length of the text. This page tells you which tokens repeat. The first line is a total. Each following line is a word, a tab, and a count, most frequent first. Length, sentences and reading time stay on word counter.
To and to share one row.北京 is one token. Turn it on to count 北 and 京 separately. Hiragana, katakana and Hangul follow the same switch.To be, or not to be. plus Paris Paris London. With the defaults, to and be are 2, and paris is 2.Latin words are letters and digits. An apostrophe inside a word stays, so don't is one token. Punctuation and spaces split tokens. A continuous run of Han, kana or Hangul is one token until a space or other script, unless you split CJK. Ignore case stores the lowercase form, using the locale's lowercasing. Rows sort by count, then by the word. Tokens shorter than the minimum are left out of both the list and the total. The summary line uses the labels for unique words and token count.
That is the default. Turn on Count each CJK character if you want a count per character. Spaces still separate runs.
Ignore case is on. Turn it off to keep capitals as a different row.