PEP 686: Python Finally Kills the Windows Locale Default
If you develop on Linux or macOS and run pipelines on Windows—or manage cross-platform CI matrix builds—you have almost certainly run headfirst into the Windows locale bug.
Your code looks completely clean and idiomatic. Your test suite runs without a hitch on your local machine. But the moment your runner spins up on a Windows box, the build halts with an infuriating exception:
UnicodeDecodeError: 'charmap' codec can't decode byte 0x9d in position...
The culprit is rarely your data; it is an invisible platform default that has lingered for thirty years.
with open("data.json") as f:
payload = f.read()
Without an explicit encoding="utf-8", that simple two-line block behaved differently depending on where the interpreter was running. On Linux and macOS, Python has long defaulted to UTF-8. On Windows, Python inspected the system locale and defaulted to the legacy ANSI code page—typically cp1252 in Western environments, or Shift-JIS, GBK, or cp949 elsewhere. The moment your file contained a curly apostrophe, an em dash, or an accented vowel, the code broke.
PEP 686 permanently fixes this in Python 3.15.
What Changes Under PEP 686
For years, avoiding this bug required sheer habit: remembering to pass encoding="utf-8" to every single open() call, TextIOWrapper, and file-based helper function across your entire codebase. Miss it once in an auxiliary utility, and cross-platform parity fell apart.
PEP 686 turns Python’s UTF-8 Mode (first introduced in PEP 540) into the universal default. Standard file operations and stdio streams (sys.stdin, sys.stdout, sys.stderr) now assume UTF-8 everywhere, including Windows.
If you are interfacing with older enterprise tools or legacy files that genuinely require the host system’s ANSI encoding, you no longer rely on implicit ambient defaults. Instead, you declare that intent explicitly using encoding="locale" (introduced in Python 3.10 via PEP 597):
# Explicitly opt into the host operating system's legacy code page
with open("export.csv", encoding="locale") as f:
payload = f.read()
If you want this cross-platform consistency in your pipelines right now without waiting for Python 3.15, you can opt in today by setting the environment variable PYTHONUTF8=1 or running Python with the -X utf8 flag.
The PowerShell 5.1 Trap
While Python 3.15 solves the encoding issue inside the interpreter, data engineering scripts rarely execute in a vacuum. On Windows, Python jobs are routinely triggered by scheduled tasks, build agents, and shell wrappers.
If your environment runs modern PowerShell 7 (pwsh), everything works out of the box. PowerShell 7 was built from the ground up to standardize on UTF-8 without BOM across cmdlets and streams.
The problem lies with Windows PowerShell 5.1 (powershell.exe). Version 5.1 remains the built-in, default shell across Windows and Windows Server installations. If your automated tasks run in 5.1, two silent gotchas will corrupt your data before Python even gets a chance to parse it:
-
Stdio Piping (
|): In 5.1, piping data between native binaries is governed by$OutputEncoding, which defaults to 7-bit US-ASCII. When you pipe string content into Python ($raw_json | python ingest.py), PowerShell encodes the stream as ASCII first. Any non-ASCII character is replaced with a literal question mark (?). Python 3.15 will read the stream as valid UTF-8, but the data is already irreparably mangled. -
Redirection to Disk (
>): In 5.1, the>redirection operator is syntactic sugar forOut-File, which defaults toUnicode(UTF-16 LE with BOM). If a script redirects program output to a file and you later read it in Python 3.15 with a bareopen("output.json"), it immediately fails because Python expects UTF-8, not UTF-16.
If you have legacy 5.1 automation orchestrating Python tasks, you need to normalize both stdio streaming and cmdlet redirection upfront:
# Prevent ASCII mangling when piping to external executables
$OutputEncoding = [System.Text.UTF8Encoding]::new($false)
[Console]::OutputEncoding = [System.Text.UTF8Encoding]::new($false)
# Force '>' (Out-File) to write UTF-8 rather than UTF-16 LE
$PSDefaultParameterValues['Out-File:Encoding'] = 'utf8'
Closing the Gap
It is a modest change in the language specification, but PEP 686 eliminates decades of unnecessary platform drift. Once Python 3.15 becomes your baseline, UTF-8 is simply the default language of text I/O everywhere.
Just make sure your host shell isn’t quietly converting your strings to ASCII behind the interpreter’s back.