ASCII and Unicode Codes

Show the code point of each character, or turn decimal or hex codes back into text.

Code points, not hex bytes

ASCII is the numbers 0–127. A Chinese character is not one ASCII byte; it is a Unicode code point, and in UTF-8 that point is several bytes. This page prints one number per character, in decimal or hex. It does not produce the hex dump of UTF-8 bytes. For those bytes, use string to bytes or text to binary.

How to use it

  1. Paste text, or a list of codes separated by spaces, commas or line breaks.
  2. Choose “Text to codes” or “Codes to text”, then Decimal or Hexadecimal.
  3. The example AB is 65 66 in decimal. In hex those are 41 42.
  4. A token that is not a number in the selected base, or a number that is not a character, is an error. The page does not skip the bad token.

Which number you are looking at

65 is the character A. 20013 is the character 中. Neither number is a UTF-8 byte: the UTF-8 bytes of 中 are E4 B8 AD. If another tool shows three hex pairs for one character, it is showing bytes, and this page will not match it until you switch tools. Hex here has no 0x and no U+ prefix. Spaces, commas and new lines are separators, not characters to encode, when you are going from codes back to text.

Where people use it

  • Reading a decimal code point from a spec or an error message.
  • Checking that a character is the one you think, when two look alike.
  • Building a short list of codes for a parser exercise.

Questions

Why is 中 not 228 or E4?

Those are byte values in UTF-8. This page prints the code point, which is one number per character.

An emoji became one number, not two.

The number is the code point. A single emoji can be one point above 65535, or several points if it is a sequence. Each point is one number.