Download

invalid byte sequence for encoding "UTF8"

PostgreSQL received bytes that aren’t valid UTF-8, usually a file saved in Latin-1 or Windows-1252 (é is 0xe9 there), or a NUL byte (0x00), which text can’t hold. Tell PostgreSQL the file’s real encoding, or convert the file, or strip the NULs.

PostgreSQL error 22021· Tested on PostgreSQL 18.6 (also 14.24)· Updated 11 October 2026

ERROR:  invalid byte sequence for encoding "UTF8": 0xe9 0x2c 0x53

What it means

Every string you send is in the client encoding, which is almost always UTF-8. PostgreSQL checks the bytes before storing them, and these weren’t valid UTF-8. The hex values are the bytes at the point where it gave up:

ERROR:  invalid byte sequence for encoding "UTF8": 0xe9 0x2c 0x53
CONTEXT:  COPY customers, line 3

0xe9 is é in Latin-1 (ISO 8859-1) and Windows-1252; in UTF-8, é takes two bytes (0xc3 0xa9), and a lone 0xe9 followed by an ordinary character is invalid. The next bytes (0x2c 0x53, a comma and an S) are only what came after it. Other common first bytes: 0x92, 0x93, 0x96 (Windows curly quotes and dashes), 0xa3 (£ in Latin-1), and 0x00.

0x00 is a NUL byte. Text in PostgreSQL can’t contain one in any encoding, so this error is the result even when the rest is perfect UTF-8.

A related message, character with byte sequence 0xe2 0x82 0xac in encoding "UTF8" has no equivalent in encoding "LATIN1", is the opposite problem: valid UTF-8 that can’t be converted to a narrower encoding (code 22P05).

Common causes

  1. A CSV exported from Excel or an older system in Windows-1252 or Latin-1, loaded with COPY or \copy as if it were UTF-8.
  2. A client that sends another encoding while telling the server it’s UTF-8, or the reverse.
  3. NUL bytes in strings: binary data put in a text column, C strings with trailing zeros, or \u0000 escapes.
  4. convert_from() on bytes in another encoding, or bytes cut in the middle of a multi-byte character (a string truncated by byte length in the application).

How to fix it

Tell COPY the file’s encoding

If you know the file is Latin-1 or Windows-1252, say so, and PostgreSQL converts it:

\copy customers FROM 'customers.csv' WITH (FORMAT csv, HEADER, ENCODING 'LATIN1')

Use 'WIN1252' for Windows files: it’s Latin-1 plus curly quotes, dashes and € in the 0x80–0x9f range. For server-side COPY … FROM '/path' the option is the same. Alternatively, set the client encoding for the session (SET client_encoding = 'LATIN1'; or \encoding LATIN1 in psql) and reset it afterwards.

Convert the file to UTF-8

On a Mac or Linux:

file -I customers.csv          # macOS (file -i on Linux): the charset it detects
iconv -f WINDOWS-1252 -t UTF-8 customers.csv > customers-utf8.csv

When in doubt between Latin-1 and Windows-1252, choose Windows-1252: it decodes everything Latin-1 does in the printable range, plus the curly quotes.

Remove NUL bytes

In the application, strip \0 before sending text (or store the value as bytea if it’s really binary). For a file: tr -d '\000' < in.csv > out.csv. In jsonb, \u0000 isn’t allowed either (unsupported Unicode escape sequence); remove it from the JSON before loading.

Decode bytes with the right encoding

SELECT convert_from('\x636166e9'::bytea, 'LATIN1');   -- café

Reproduce it

On PostgreSQL 18.6 (server and client encoding UTF8), with a CSV saved in Latin-1 whose third line is 2,José,São Paulo:

\copy customers FROM 'customers-latin1.csv' WITH (FORMAT csv, HEADER)
ERROR:  invalid byte sequence for encoding "UTF8": 0xe9 0x2c 0x53
CONTEXT:  COPY customers, line 3

The same command with ENCODING 'LATIN1' loaded both rows (COPY 2) and they read back as José and São Paulo; so did the plain command after SET client_encoding = 'LATIN1'. Directly in SQL:

SELECT convert_from('\x636166e9'::bytea, 'UTF8');
ERROR:  invalid byte sequence for encoding "UTF8": 0xe9
SELECT E'abc\x00def';
ERROR:  invalid byte sequence for encoding "UTF8": 0x00
SELECT convert_to('€ “quotes”', 'LATIN1');
ERROR:  character with byte sequence 0xe2 0x82 0xac in encoding "UTF8" has no equivalent in encoding "LATIN1"

convert_from('\x636166e9'::bytea, 'LATIN1') returned café. With \set VERBOSITY verbose, psql shows the code: ERROR: 22021: invalid byte sequence for encoding "UTF8": 0xe9. PostgreSQL 14.24 gives the same messages.

In Inlet

Inlet’s CSV import and COPY-based import load files into a table, and the inspector shows long text values in full, so you can check how accented characters came through. When a statement fails, Inlet shows the server’s error; with your own Anthropic API key, Ask Claude (Pro, ⌘L) can fix the failed statement, sending the schema, the SQL and the error, never rows.

Inlet: a database client for the Mac

One native app for PostgreSQL, MySQL, SQL Server, SQLite, MongoDB and Redis. It explains errors where they happen, holds your edits until you save them, and keeps production read-only until you say so.

Version 0.1.0 · macOS 26 Tahoe or later · Apple silicon and Intel