Character Encoding Converter
Convert text to and from UTF-8 bytes (hex, binary or decimal), URL percent-encoding, Base64, \uXXXX escapes and HTML numeric entities.
About the Character Encoding Converter
Computers store text as bytes, and the character encoding decides which bytes stand for which character. Today that is almost always UTF-8, where ASCII characters take one byte, é takes two (c3 a9), the rupee sign ₹ takes three (e2 82 b9) and most emoji take four. This converter shows those bytes and translates text between the formats developers run into: UTF-8 bytes written in hex, binary or decimal, URL percent-encoding, Base64, JavaScript and JSON \u escapes, and HTML numeric entities.
Pick a From format and a To format. The input is first decoded to plain text and then encoded into the target, so you can go straight from one representation to another, for example from a percent-encoded URL parameter to hex bytes, or from Base64 to \u escapes. "₹100" in URL encoding is %E2%82%B9100, and the same text as hex bytes is e2 82 b9 31 30 30. Byte input is checked strictly: if the bytes are not valid UTF-8, you get an error instead of a string full of replacement characters.
This is useful for reading a URL parameter, checking what bytes a string really contains before it goes into a database or an API, or decoding escaped text from a JSON file or a log. Only UTF-8 is supported for byte formats; legacy encodings such as ISO-8859-1, Windows-1252 or Shift_JIS are not. To see one code number per character instead of bytes, use the ASCII Converter; to encode whole files as Base64, use the Base64 Encoder/Decoder.
How to use the Character Encoding Converter
- 1
Choose the From format
Select how your input is written: Plain text, Hex bytes, Binary bytes, Decimal bytes, URL encoding, Base64, Unicode escapes or HTML numeric entities.
- 2
Choose the To format
Select the format you want. The arrow button between the lists swaps the two formats and moves the current output into the input box.
- 3
Paste your input and convert
Paste the text or data and press Convert. If the input is not valid for the From format, a red message explains what is wrong.
- 4
Copy the output
Press Copy above the output box.
What it can do
Any format to any format
Eight representations, including plain text, can be converted in either direction without an intermediate step.
Strict UTF-8 checking
Byte sequences that are not valid UTF-8 are reported rather than silently turned into � characters.
Flexible byte input
Hex may be written with or without spaces, commas or 0x prefixes; binary and decimal bytes are separated by spaces or commas.
Emoji-correct escapes
\u output follows JavaScript and JSON, writing an emoji as a surrogate pair such as 😀; the decoder also reads the \u{1F600} form.
Full HTML entity decoding
Decoding HTML entities uses the browser's parser, so named entities like ♥ work as well as numeric ones.
Limitations
- Only UTF-8 is supported for byte formats. ISO-8859-1 (Latin-1), Windows-1252, UTF-16 byte output and other legacy encodings are not available.
- URL decoding does not turn + into a space. Query strings from HTML forms use + for spaces, so replace + with %20 before decoding.
- Unicode escape output escapes only characters above 127 and uses lowercase \uXXXX; \x and \U forms are not produced.
- HTML entity output uses decimal numeric entities only, such as é, never named ones.
Privacy
Everything is converted by JavaScript in your browser; the text you paste is not sent to FlexyPdf's servers.
Frequently asked questions
Which encodings does this tool support?
Plain text, UTF-8 bytes as hex, binary or decimal, URL percent-encoding, Base64 of the UTF-8 bytes, \uXXXX escapes and HTML numeric entities. Legacy code pages such as ISO-8859-1 and Windows-1252 are not supported.
Why does é turn into two bytes?
UTF-8 uses a variable number of bytes per character: one for ASCII, two for most accented Latin letters, three for symbols such as ₹ and many Indian scripts, and four for most emoji. é is c3 a9 in hex.
What does "These bytes are not valid UTF-8 text" mean?
The byte values do not form valid UTF-8. Common causes are bytes from a different encoding (é in Latin-1 is the single byte e9), a sequence cut off in the middle, or a typo in the hex.
How do I decode a query string that uses + for spaces?
Replace each + with %20 (or a space) first, then decode with URL encoding as the From format. The decoder follows the standard percent-encoding rules, where + is an ordinary character.
What is the difference between \u00e9 and é?
Both mean é. \u00e9 is a JavaScript or JSON string escape with the value in hexadecimal; é is an HTML entity with the value in decimal.
Ratings & Reviews
Rate Character Encoding Converter
Help others by sharing your experience. Your rating is shown without your name.
More Converters
View all →CSV to JSON
Convert CSV rows into a JSON array of objects
JSON to CSV
Turn a JSON array of objects into CSV for Excel or Google Sheets
JSON to YAML
Convert JSON to YAML, or common YAML config files back to JSON
Base64 Encoder/Decoder
Encode text or files to Base64 and decode Base64 back