Handbooks / Python / Chapter 6

Files & JSON

45 pages · ~84 min✓ Reviewed

Builds on Errors & Exceptions. Next up: Iterators & Generators.

Part 1 · Introduction

Files & JSON: Making Your Programs Remember

Every variable in a Python program disappears the moment the program ends. Files are how your work survives: a log of what happened, a spreadsheet someone sent you, a settings file, the response saved from a web service. Almost every real program reads data from a file, writes results to one, or does both, so this is one of the first skills that turns small scripts into useful tools.

This chapter starts with the big picture of how Python talks to files, then builds up one skill at a time. You will learn how open() modes and the with statement keep files safe, how to read and write text, and how pathlib treats paths as objects instead of fragile strings. Then come the details that cause most real-world trouble: encodings and UTF-8, the csv module for tabular data, and json for structured data that moves between programs and languages.

By the end you will be able to load a CSV or JSON file, clean it up and save the result. You will handle a missing file with a clear message instead of a crash, and stream a file that is too big for memory one line at a time. You will also recognise the most common file mistakes (forgotten closes, wrong modes, unspecified encodings, hand-built paths) and know the fix for each. The last page is a cheat sheet you can keep open while you work.

Before you start

You need Python 3.11 or newer and a plain text editor or IDE. Every example uses only the standard library, so there is nothing to install. You should already be comfortable with variables, strings, lists, dictionaries, for loops and defining simple functions. Run the examples from a scratch folder you do not mind filling with small test files.

Part 2 · Files in Python: The Big Picture

What open() gives you

Every file operation in Python starts with open(). It does not hand you the contents of the file. It returns a file object, a handle that stands between your code and the operating system. The handle keeps track of three things: the OS file descriptor (the number the operating system uses to identify the open file), a position (where the next read or write happens), and a buffer (a chunk of memory that batches data so Python does not call the OS for every character).

The example below opens a new file for writing and looks at the handle itself. Notice that nothing has been read yet. The handle starts at position 0, and writing moves the position forward.

python
import os, tempfile

folder = tempfile.mkdtemp()
path = os.path.join(folder, 'notes.txt')

f = open(path, 'w', encoding='utf-8')
print(type(f).__name__)
print(f.mode, f.closed)
print(f.tell())
print(f.write('hello'))
print(f.tell())
f.close()
print(f.closed)

The handle is an object with state: mode, position and a closed flag.

output
TextIOWrapper
w False
0
5
5
True

write() returns how many characters it wrote, and tell() reports the position. After close() the handle still exists as a Python object, but closed is True and any further read or write on it raises an error.

Text mode and binary mode

The mode string you pass to open() decides what kind of data the handle deals in. Text mode (the default) decodes the bytes on disk into str using an encoding, and translates line endings for you. Binary mode (add a b to the mode) skips all of that and gives you raw bytes exactly as stored.

Text modeBinary mode
Mode string'r', 'w', 'a''rb', 'wb', 'ab'
read() returnsstrbytes
write() acceptsstrbytes
EncodingApplied, so pass encoding='utf-8'None; you get the raw bytes
Line endingsTranslated to \n on readLeft untouched
Typical useNotes, logs, CSV, JSONImages, audio, zip files, downloads
Handle classTextIOWrapperBufferedReader or BufferedWriter

Here the same file is read both ways. The bytes version shows the b prefix on the result, which is Python's way of telling you it is not a string.

python
with open(path, 'rb') as f:
    data = f.read()
print(type(f).__name__, data)

with open(path, 'r', encoding='utf-8') as f:
    text = f.read()
print(type(f).__name__, text)

Same file, two views: bytes and str.

output
BufferedReader b'hello'
TextIOWrapper hello
Common mistake: mixing str and bytes

Writing a str to a file opened with 'wb', or bytes to a file opened with 'w', raises TypeError. Pick the mode that matches the data you actually have.

The life cycle and the three layers

Open, use, close

Every file handle follows the same life cycle. You open it, you read or write through it, and you close it. Closing matters for two reasons. It hands the file descriptor back to the operating system, which only allows a limited number of open files per process. And it pushes any data still sitting in the buffer out to the disk.

Life cycle of a file handle
  1. 1open()OS gives a file descriptor; Python builds the handle
  2. 2read / writeposition moves; writes collect in the buffer
  3. 3flush()buffer is pushed to the OS
  4. 4close()flushes, then releases the descriptor

The surprising part is the buffer. When you call write(), the data usually does not reach the file straight away. It waits in memory until the buffer fills up, until you call flush(), or until close() runs. The next example opens the same file with a second handle to prove it. The writer has written text, but a reader sees an empty file until the flush.

python
w = open(path, 'w', encoding='utf-8')
w.write('saved?')

with open(path, encoding='utf-8') as r:
    print(repr(r.read()))

w.flush()
with open(path, encoding='utf-8') as r:
    print(repr(r.read()))

w.close()

Data written but not flushed is not yet visible in the file.

output
''
'saved?'

If the program crashed between the write and the flush, the text would simply be gone. That is how skipping close() can lose data. In practice you rarely call close() by hand. The with statement, covered in the next section, closes the file for you even when an error occurs.

Common mistake: forgetting to close

Code like open(path, 'w').write('data') never closes the handle explicitly. It may work on a small script, but it can leave data unwritten or leak descriptors in a long-running program.

Three layers you will build on

Working with files in Python comes down to three layers. Each later section of this chapter uses one or more of them, so it helps to know which layer solves which problem.

open()

Built-in, no import

Reads and writes any file

Gives you a handle

pathlib

Paths as objects

Join, check, list files

Quick read_text / write_text

csv and json

Structured formats

Rows and dictionaries

Built on top of file handles

The last example touches all three layers. pathlib builds the path and does the quick read and write, and json converts between a Python dictionary and text. Under the hood pathlib still opens and closes a file handle for you.

python
from pathlib import Path
import json

p = Path(folder) / 'data.json'
p.write_text(json.dumps({'ok': True}), encoding='utf-8')
print(p.name, p.suffix)
print(json.loads(p.read_text(encoding='utf-8')))

pathlib for the path, json for the format, open() working underneath.

output
data.json .json
{'ok': True}

Interview Point

What to remember

A file object is a handle with a descriptor, a position and a buffer. Text mode gives str, binary mode gives bytes, and the buffer is only guaranteed to reach disk on flush() or close().

What does open() return?
  • A file object, not the contents. In text mode it is a TextIOWrapper; in binary mode a BufferedReader or BufferedWriter.
  • It wraps the OS file descriptor, a read/write position and a buffer.
  • You get the data by calling methods on it, such as read(), readline() or write().
f = open('notes.txt', 'w', encoding='utf-8')
print(type(f).__name__)  # TextIOWrapper
Why must a file be closed?
  • Writes sit in a buffer, and close() flushes it. Without it, data can be lost if the program ends badly.
  • It releases the OS file descriptor. Each process can only hold a limited number open.
  • On some systems an open file cannot be deleted or renamed by other programs.
  • The with statement closes the file automatically, even if an exception happens.
with open('notes.txt', 'w', encoding='utf-8') as f:
    f.write('safe')
What is the difference between text and binary mode?
  • Text mode decodes bytes into str using an encoding and translates line endings.
  • Binary mode ('rb', 'wb') returns and accepts raw bytes with no decoding.
  • Use text for human-readable files and binary for images, audio, archives and anything where exact bytes matter.
open('a.txt', 'r', encoding='utf-8').read()  # str
open('a.txt', 'rb').read()                   # bytes

Part 3 · open() Modes and the with Statement

Choosing a mode for open()

Every call to open() answers one question before anything else: what are you about to do with this file? You answer it with the second argument, the mode. It is a short string of letters, and it decides whether the file must already exist, whether old content survives, and whether you get text or raw bytes.

There are four main modes. Leave the argument out and you get 'r', so open('a.txt') means "read text from a file that must already exist".

ModeMeaningFile missingFile exists
'r'Read (the default)FileNotFoundErrorOpened, starts at the beginning
'w'WriteCreatedTruncated to empty immediately
'a'AppendCreatedKept, writes go at the end
'x'Exclusive createCreatedFileExistsError

Two more letters combine with those. + adds the other direction, so 'r+' can read and write and 'w+' can too. t means text and is the default: you work with str and Python translates line endings and encodes for you. b means binary: you work with bytes and nothing is translated. So 'rb' is read bytes, and 'r' is the same as 'rt'.

Which mode do I need?

The example below runs all four behaviours against a scratch file in a temporary folder. It reads a missing file, creates one with 'w', adds a line with 'a', reads it back, and finally tries 'x' on a file that is now there.

python
import tempfile, os

folder = tempfile.mkdtemp()
path = os.path.join(folder, 'notes.txt')

try:
    open(path)
except FileNotFoundError:
    print('r: missing file -> FileNotFoundError')

with open(path, 'w') as f:
    f.write('first\n')
with open(path, 'a') as f:
    f.write('second\n')
with open(path) as f:
    print(f.read().splitlines())

try:
    open(path, 'x')
except FileExistsError:
    print('x: file exists -> FileExistsError')

The path variable is reused in later examples.

output
r: missing file -> FileNotFoundError
['first', 'second']
x: file exists -> FileExistsError
Common mistake: forgetting that b gives bytes

In 'rb' mode f.read() returns bytes such as b'hi', not str. Mixing the two gives a TypeError. Use 'b' for images, archives and other non-text data, and plain text mode for anything you want to treat as strings.

The with statement

An open file holds an operating-system resource, and buffered writes may not reach the disk until the file is closed. The old way to be safe was try/finally: open the file, do the work in try, and call f.close() in finally. It works but it is noisy, and it is easy to forget.

The with statement does that job for you. In with open('a.txt') as f: the file is opened, bound to f, and then f.close() is called as soon as the block ends. That includes the case where an exception is raised inside the block. Both versions are shown here, and the second one raises on purpose.

python
f = open(path, 'w')
try:
    f.write('data')
finally:
    f.close()
print('try/finally closed:', f.closed)

try:
    with open(path, 'w') as g:
        g.write('partial')
        raise ValueError('boom')
except ValueError as err:
    print('caught:', err)
print('with closed:', g.closed)
output
try/finally closed: True
caught: boom
with closed: True

The exception was not swallowed: it still reached the except. What with guarantees is that cleanup runs on the way out. The name g also stays available after the block, and g.closed confirms the file is shut.

This is not a special trick of files. It is the context manager protocol. Any object with an __enter__ method and an __exit__ method can sit after with. Python calls __enter__ first and binds its return value to the name after as. When the block ends, normally or by exception, it calls __exit__. A file object simply has an __exit__ that closes it.

You can watch the protocol with a tiny class of your own. The __exit__ method receives the exception type (or None when all went well). Returning False lets the exception continue to propagate, which is what files do.

python
class Tracer:
    def __enter__(self):
        print('enter')
        return self
    def __exit__(self, exc_type, exc, tb):
        print('exit', exc_type.__name__ if exc_type else None)
        return False

with Tracer():
    print('body')

try:
    with Tracer():
        1 / 0
except ZeroDivisionError:
    print('still raised')
output
enter
body
exit None
enter
exit ZeroDivisionError
still raised
What with replaces

with open(...) as f: is a shorter, safer form of open, then try, then finally: f.close(). Prefer it every time you open a file.

w, a, x and r+ side by side

The modes differ most in what they do to data that is already in the file. This is where students lose work, so the comparison is worth keeping in your head.

ModeExisting contentIf file existsTypical use
'w'Wiped on openOverwrittenRegenerate an output file
'a'KeptAppended toLogs, adding rows
'x'Not touchedRefused (FileExistsError)Never clobber a file
'r+'Kept, no truncationRead and write in placeEdit a file that must exist
Common mistake: w destroys data on open

The old content is gone the moment open(path, 'w') returns, before you write a single character. If your program then crashes, or you open the wrong path, the data is already lost. Reach for 'a' to add to a file and 'x' when overwriting would be a bug.

The next example shows all of this on one file. It writes some text, opens it with 'r+' to read and then overwrite the start, appends with 'a', and finally opens with 'w' and writes nothing at all. After r+ and a the content is intact; after w it is empty.

python
with open(path, 'w') as f:
    f.write('precious data')

with open(path, 'r+') as f:
    print('r+ sees:', f.read())
    f.seek(0)
    f.write('PRECIOUS')

with open(path, 'a') as f:
    f.write('!')
with open(path) as f:
    print('after a:', f.read())

with open(path, 'w'):
    pass
with open(path) as f:
    print('after w:', repr(f.read()))
output
r+ sees: precious data
after a: PRECIOUS data!
after w: ''

Interview points

What does with guarantee?
  • The context manager's __exit__ runs when the block ends, so the file is closed.
  • That holds on normal exit, on return, and when an exception is raised inside the block.
  • It does not hide the exception; it is re-raised after cleanup.
  • It replaces the try/finally pattern with one line.
with open('a.txt') as f:
    data = f.read()
# f is closed here, even if read() failed
What is the difference between 'w' and 'a'?
  • 'w' truncates the file to empty as soon as it is opened, then writes from the start.
  • 'a' keeps the existing content and every write goes at the end.
  • Both create the file if it is missing.
When would you use 'x'?
  • When you must create a brand new file and never overwrite an existing one.
  • If the file exists, Python raises FileExistsError instead of destroying it.
  • The check and creation happen in one step, so it is safer than testing for existence first and then opening with 'w'.
try:
    with open('report.txt', 'x') as f:
        f.write('v1')
except FileExistsError:
    print('report already exists')
Is closing a file twice an error?
  • No. Calling close() on an already closed file does nothing.
  • Reading or writing a closed file is the error: it raises ValueError.
  • This is why an explicit close() inside a with block is harmless, though unnecessary.

The last check is easy to run yourself. The second close() passes quietly, and only the read fails.

python
f = open(path)
f.close()
f.close()
print(f.closed)
try:
    f.read()
except ValueError as err:
    print(err)
output
True
I/O operation on closed file.
Remember

Pick the mode by what must happen to existing data: 'r' leaves it alone, 'a' adds to it, 'w' erases it, 'x' refuses to touch it. Then let with handle the closing.

Part 4 · Reading Text Files

Four ways to pull text out of a file

Once open() gives you a file object, you choose how much text to take from it at a time. Python offers four reading methods, and the difference between them is what comes back and how much of the file they consume. The main choice is between one big string, a fixed number of characters, a single line, or a list of every line.

f.read() returns the whole file as one str. f.read(n) returns up to n characters, and fewer if the file runs out first. f.readline() returns one line, and that line keeps its trailing '\n'. f.readlines() returns a list of all the lines, each one still ending in '\n' (except possibly the last, if the file does not end with a newline).

The example below first creates a small file called notes.txt so that you can run it as it stands. Writing is covered properly in a later section, so for now treat that first block as setup. Each read then gets its own with block, so every method starts from the top of a freshly opened file. Printing with repr() makes the invisible '\n' characters visible.

python
with open('notes.txt', 'w', encoding='utf-8', newline='\n') as f:
    f.write('alpha\nbeta\ngamma\n')

with open('notes.txt', encoding='utf-8') as f:
    print(repr(f.read()))

with open('notes.txt', encoding='utf-8') as f:
    print(repr(f.read(7)))

with open('notes.txt', encoding='utf-8') as f:
    print(repr(f.readline()))
    print(repr(f.readline()))

with open('notes.txt', encoding='utf-8') as f:
    print(f.readlines())

newline='\n' in the setup keeps the file byte-for-byte identical on every operating system

output
'alpha\nbeta\ngamma\n'
'alpha\nb'
'alpha\n'
'beta\n'
['alpha\n', 'beta\n', 'gamma\n']

Notice that read(7) stopped in the middle of the second line. It counts characters, not lines, so it happily cuts through a '\n'. Notice also that the two readline() calls returned different lines: each call picks up where the previous one stopped. That behaviour comes from the file position, which is the subject of the next page.

MethodReturnsHow much it takesKeeps the '\n'?
f.read()one strthe whole fileyes, inside the string
f.read(n)one strup to n charactersyes, if one falls in the range
f.readline()one stra single lineyes, at the end of the line
f.readlines()a list of strevery lineyes, on each line

The file position and end of file

A file object remembers where it is. Think of a cursor that starts at the beginning when you open the file and moves forward with every read. Whatever the cursor has already passed is not returned again. This single idea explains why readline() gives you a new line each time, and why reading a file twice in a row gives a surprising result.

When the cursor reaches the end, you are at EOF (end of file). There is nothing left to read, so read() returns an empty string ''. It does not raise an error. Calling read() a second time on the same open file therefore gives '', because the first call already moved the cursor to the end.

Two methods let you inspect and control the cursor. f.tell() reports the current position, and f.seek(0) jumps back to the start so you can read the content again.

What the cursor does
  1. 1open()cursor at position 0
  2. 2f.read()returns all text, cursor moves to the end
  3. 3f.read() againreturns '' because the cursor is at EOF
  4. 4f.seek(0)cursor back at position 0
  5. 5f.read()returns all the text once more

The next example reuses notes.txt from the previous page. Follow the numbers printed by tell(): the file holds 17 characters in total (alpha\n is 6, beta\n is 5 and gamma\n is 6).

python
with open('notes.txt', encoding='utf-8') as f:
    print(f.tell())
    print(repr(f.read(6)))
    print(f.tell())
    print(repr(f.read()))
    print(f.tell())
    print(repr(f.read()))
    f.seek(0)
    print(repr(f.readline()))
output
0
'alpha\n'
6
'beta\ngamma\n'
17
''
'alpha\n'

The sequence matches the diagram. The first read(6) consumed 'alpha\n' and moved the cursor to 6. The plain read() took everything that remained and moved the cursor to 17. The next read() found nothing and returned ''. After seek(0), readline() returned the first line again.

Common mistake: reading the same open file twice

A beginner writes text = f.read() and later lines = f.readlines() on the same open file, then wonders why lines is []. The first call used up the file. Either keep the first result and work with it, call f.seek(0) before the second read, or open the file again.

Memory, iteration and the newline

read() and readlines() both load the entire file into memory before you can use any of it. For a notes file that is perfectly fine. For a log file of several gigabytes it can exhaust your RAM. The alternative is to loop directly over the file object with for line in f:. Python then hands you one line at a time and keeps only that line in memory, so the file can be as large as the disk.

read() / readlines()for line in f
Memory useThe whole file at onceOne line at a time
You getA string or a list, ready to reuseLines as they stream past
Good forSmall files you need to examine as a wholeLarge files, or a single pass over the data
Reading it twiceReuse the variableNeeds seek(0) or a fresh open()
Which reading style fits?

Every line you receive, whichever way you read it, still carries its '\n'. To get clean text, remove it. line.rstrip('\n') removes only trailing newlines, so leading spaces survive. line.strip() removes whitespace from both ends, which is handy for plain words but will also erase meaningful indentation. The example loops over the file, shows the line before and after cleaning, and then builds a clean list with a comprehension. The last lines show why the newline matters in practice.

python
with open('notes.txt', encoding='utf-8') as f:
    for line in f:
        print(repr(line), '->', repr(line.rstrip('\n')))

with open('notes.txt', encoding='utf-8') as f:
    names = [line.strip() for line in f]
print(names)

row = '3,4\n'
print(row.split(','))
print(int('42\n'))
output
'alpha\n' -> 'alpha'
'beta\n' -> 'beta'
'gamma\n' -> 'gamma'
['alpha', 'beta', 'gamma']
['3', '4\n']
42
Common mistake: forgetting the newline

Printing a raw line with print(line) adds a second newline, so the output looks double-spaced. Comparisons also fail: line == 'alpha' is False when line is 'alpha\n'. The '\n' stays glued to the last field in line.split(','), giving '4\n' above. int(line) is safe, because int() ignores whitespace around the number, but do not rely on that for other types. Strip first and your code stays predictable.

Interview Point: read, readline and readlines

File reading comes up often in Python interviews because the answers test whether you understand the file position and memory use. Practise saying each answer out loud in one or two sentences, then back it up with a short snippet.

What does read() return at end of file?
  • It returns an empty string '', not None and not an error.
  • That is how you detect EOF: an empty result means nothing is left.
  • The same holds for readline(). A blank line in the middle of a file is '\n', which is not empty, so the two cases cannot be confused.
with open('notes.txt', encoding='utf-8') as f:
    f.read()
    print(repr(f.read()))
How do read, readline and readlines differ?
  • read() returns the whole remaining file as a single string, and read(n) returns at most n characters.
  • readline() returns exactly one line, including its '\n', and is '' at EOF.
  • readlines() returns a list of every remaining line, each with its '\n'.
  • read() and readlines() hold the whole file in memory. Looping with for line in f does not.
Why does reading a file twice give an empty result?
  • The file object keeps a position that only moves forward as you read.
  • The first read() leaves it at EOF, so the second one has nothing left and returns ''.
  • Call f.seek(0) to rewind to the start, or reopen the file. f.tell() shows where you are.
with open('notes.txt', encoding='utf-8') as f:
    first = f.read()
    f.seek(0)
    second = f.read()
    print(first == second)
Remember

Reading consumes the file from front to back. The position only moves forward until you seek(), EOF shows up as '', and every line you read still carries its '\n'. For big files, iterate instead of loading everything.

Part 5 · Writing and Appending Text

write() and writelines()

Once a file is open in a writing mode, the main tool is f.write(s). It takes one str, puts it into the file at the current position, and hands back the number of characters it wrote. It adds nothing of its own: no space, no separator, and no newline. If you want your text on separate lines, the '\n' is your job.

The example below makes a scratch folder, writes three pieces, and keeps each return value. Then it reads the file back with repr() so the invisible newline characters show up as \n.

python
import tempfile
from pathlib import Path

folder = Path(tempfile.mkdtemp())
path = folder / 'notes.txt'

with open(path, 'w', encoding='utf-8') as f:
    a = f.write('alpha')
    b = f.write('beta\n')
    c = f.write('gamma\n')

print(a, b, c)
print(repr(path.read_text(encoding='utf-8')))

write() returns a count and never adds a newline

output
5 5 6
'alphabeta\ngamma\n'

alpha and beta ran together because the first write had no newline. The counts are 5, 5 and 6 because 'beta\n' is four letters plus the newline, and 'gamma\n' is five plus the newline. The return value is counted in characters, not bytes, so a non-ASCII character still counts as one.

f.writelines(items) takes any iterable of strings and writes them one after another. Despite its name, it also adds no newlines. It is only a shortcut for a loop of write() calls. Include the '\n' in each item, or add it as you generate them.

python
lines = ['red', 'green', 'blue']

with open(path, 'w', encoding='utf-8') as f:
    f.writelines(lines)
print(repr(path.read_text(encoding='utf-8')))

with open(path, 'w', encoding='utf-8') as f:
    f.writelines(line + '\n' for line in lines)
print(repr(path.read_text(encoding='utf-8')))

writelines() glues items together unless they carry their own newline

output
'redgreenblue'
'red\ngreen\nblue\n'
MethodTakesReturnsAdds newline?
f.write(s)One strNumber of characters writtenNo
f.writelines(items)An iterable of strNoneNo
print(s, file=f)Any valuesNoneYes, by default
Common mistake: everything on one line

Calling f.write('apple') then f.write('pear') produces applepear. Neither write() nor writelines() inserts line breaks, so end each line with '\n' yourself.

print() to a File, str(), and Appending

If you would rather not manage newlines, send print() to the file with its file= argument. print('text', file=f) writes the text and then a newline, exactly as it would on screen. It also converts non-strings for you, and you can still use sep= and end=.

write() is stricter. It accepts only a str. Handing it a number raises TypeError, so convert first with str(). The example mixes all three styles in one file.

python
with open(path, 'w', encoding='utf-8') as f:
    print('first line', file=f)
    print('total:', 42, file=f)
    f.write(str(42))
    f.write('\n')
    try:
        f.write(42)
    except TypeError as err:
        print(type(err).__name__)

print(path.read_text(encoding='utf-8'), end='')

print adds the newline; write needs a str

output
TypeError
first line
total: 42
42

The TypeError line comes first only because it is printed to the screen while the file text is read back at the end. The failed f.write(42) wrote nothing, so the file holds just the three lines before it.

Everything so far used mode 'w', which empties the file when it opens. To add to the end of an existing file instead, open it with 'a' (append). The file is created if it does not exist, and every write lands after the current content.

python
with open(path, 'w', encoding='utf-8') as f:
    f.write('one\n')

with open(path, 'a', encoding='utf-8') as f:
    f.write('two\n')
    print('three', file=f)

print(path.read_text(encoding='utf-8'), end='')

'a' keeps what is there and writes after it

output
one
two
three
Common mistake: opening with 'w' to add a line

Opening a log file with 'w' each time wipes its earlier content. Use 'a' when you mean to add to a file.

Buffering, flush() and Crashes

Writing to a disk one tiny piece at a time would be slow, so Python collects your text in a buffer in memory and hands it over in larger chunks. A write() call that returns successfully therefore does not mean the text has reached the file yet.

f.flush() pushes the buffer to the operating system on demand. Leaving the with block does that and also closes the file. The next example checks the file size from outside before and after a flush.

python
f = open(path, 'w', encoding='utf-8')
f.write('saved?')
print(path.stat().st_size)
f.flush()
print(path.stat().st_size)
f.close()

the text only reaches the file after flush

output
0
6

Before the flush the file was still empty, even though write() had already returned 6. If the program had been killed at that moment, those six characters would have been lost. Buffered data lives only in your process, and a crash or power cut gives it no chance to be written.

SituationIs the data safe?
write() returned, no flush yetNot guaranteed; it may still be in the buffer
After f.flush()Handed to the OS, so it survives a crash of your program
After the with block endsFlushed and closed
Program crashes before eitherThe last buffered writes can be lost
Common mistake: open() without close()

Writing to a file opened without with and never calling close() can leave the last part unwritten. Use with open(...) as f: so the flush and close happen even when an error occurs inside the block.

Remember

write() returns a character count, adds no newline, and accepts only str. print(..., file=f) adds the newline for you. Buffered text is safe only after flush() or leaving the with block.

Interview: Why does write() not add a newline? What does it return?
  • write() is a low-level call that puts exactly the string you give it into the file, so you control the layout. The same method can write a line, a fragment, or a whole block.
  • It returns the number of characters written, for example 6 for 'gamma\n'.
  • print() adds the newline because it is meant for lines of output.
with open('out.txt', 'w', encoding='utf-8') as f:
    n = f.write('hello\n')
    print(n)
Interview: Why can data be missing after a crash, and how do flush() and with help?
  • write() puts text in an in-memory buffer first, and a crash before the buffer is emptied loses what was in it.
  • flush() sends the buffer to the operating system immediately.
  • Leaving a with block flushes and closes the file, even if an exception occurs inside the block, so with is the default habit.

Part 6 · pathlib: Paths as Objects

Building paths with the / operator

Until now a path has been a plain string such as 'data/in/a.txt'. That works, but a string knows nothing about folders, extensions or separators. The pathlib module wraps a path in a Path object that does know. You build one by starting from Path('data') and joining pieces with the / operator, and Python picks the separator for the operating system it runs on: a backslash on Windows, a forward slash on Linux and macOS.

At least one side of the / must be a Path. Path('data') / 'in' works, and so does 'data' / Path('a.txt'), because the Path object handles the operator from either side. Two plain strings, 'data' / 'a.txt', raise a TypeError, since strings do not support /. In practice you create one Path at the start and keep joining from it.

Once you have a Path you can read its pieces as properties. .name is the last part, .stem is the name without its final suffix, .suffix is that final extension including the dot, .parent is the folder that holds it, and .parts is a tuple of every component. The method .with_suffix('.csv') returns a new Path with the extension swapped. It does not rename anything on disk. The example prints .as_posix() so the output uses forward slashes on every machine.

Pieces of Path('data') / 'in' / 'a.txt'
  1. 1.parentdata/in
  2. 2.namea.txt
  3. 3.stema
  4. 4.suffix.txt
python
from pathlib import Path

p = Path('data') / 'in' / 'a.txt'
print(p.as_posix())
print(p.name)
print(p.stem)
print(p.suffix)
print(p.parent.as_posix())
print(p.parts)
print(p.with_suffix('.csv').as_posix())

print(('data' / Path('a.txt')).as_posix())
try:
    'data' / 'a.txt'
except TypeError:
    print('TypeError')

Build a path, read its pieces, and see which operands the / operator accepts

output
data/in/a.txt
a.txt
a
.txt
data/in
('data', 'in', 'a.txt')
data/in/a.csv
data/a.txt
TypeError
PartWhat it givesFor data/in/a.txt
.nameLast componenta.txt
.stemName without the last suffixa
.suffixLast extension, with the dot.txt
.parentThe folder that contains itdata/in
.partsTuple of all components('data', 'in', 'a.txt')

Building a Path never touches the disk, so the folders in it do not need to exist yet. Questions about the disk are separate methods. .exists() says whether anything is there, .is_file() and .is_dir() say which kind it is, and .resolve() turns a relative path into an absolute one, following any .. parts and links. The next page uses the first three.

One-liners for reading and writing

A Path can read and write its own file. p.write_text(s, encoding='utf-8') opens the file, writes the string and closes it in a single call. p.read_text(encoding='utf-8') opens, reads everything and closes. There are no handles to forget about, so no with block is needed. Always pass encoding='utf-8' yourself, as in the earlier sections, so the result does not depend on the machine's default.

For raw bytes there is a matching pair. p.write_bytes(b) and p.read_bytes() skip text decoding entirely, which is right for images, archives and anything that is not text. When you need a real file object, for example to loop over lines, p.open() is a drop-in for open(p). It takes the same mode and encoding arguments and works in a with block.

The example below runs inside a temporary folder that disappears afterwards, so you can run it without leaving files behind. It writes one accented word, checks the file with .exists(), .is_file() and .is_dir(), then reads it back as text and as bytes. Printing with ascii() keeps the output plain even on a terminal that cannot show accents.

python
import tempfile
from pathlib import Path

with tempfile.TemporaryDirectory() as tmp:
    base = Path(tmp)
    note = base / 'note.txt'
    print(note.exists())
    note.write_text('h\u00e9llo', encoding='utf-8')
    print(note.exists(), note.is_file(), note.is_dir())
    print(ascii(note.read_text(encoding='utf-8')))
    print(note.read_bytes())

    blob = base / 'blob.bin'
    blob.write_bytes(b'\x00\x01\x02')
    print(blob.read_bytes())

    with note.open(encoding='utf-8') as f:
        print(len(f.read()))

Text helpers, byte helpers and p.open() on a temporary file

output
False
True True False
'h\xe9llo'
b'h\xc3\xa9llo'
b'\x00\x01\x02'
5

The bytes line shows why the encoding matters. The single character é is stored as two bytes, \xc3\xa9, in UTF-8, yet the text read back has five characters. read_text did the decoding for you, and it can only decode correctly if you told it the right encoding.

You wantUseOpens and closes for you?
Whole file as a stringp.read_text(encoding='utf-8')Yes
Replace file with a stringp.write_text(s, encoding='utf-8')Yes
Whole file as bytesp.read_bytes()Yes
Replace file with bytesp.write_bytes(b)Yes
A file object for loops, appending or streamingp.open(...) inside withNo, the with closes it
Common mistake: write_text replaces everything

write_text opens the file in write mode, so it erases the old content first. It cannot append. To add to the end of a file, use p.open('a', encoding='utf-8') inside a with block.

Working with folders

The same object also manages directories. p.mkdir(parents=True, exist_ok=True) creates a folder safely. parents=True creates any missing folders along the way, so reports/2026 works even when reports does not exist yet. exist_ok=True stops Python raising an error when the folder is already there, so running the line twice is harmless. Without those two flags, mkdir fails with FileNotFoundError for a missing parent and FileExistsError for an existing folder.

To look inside a folder, p.iterdir() yields every child, files and folders alike, in no promised order. p.glob('*.csv') yields only the children matching a pattern. p.rglob('*.py') does the same but searches recursively through every subfolder. All three give Path objects, not strings, and they produce results lazily, so wrap them in sorted(...) or list(...) when you need a stable order. The methods p.unlink() and p.rename(new) delete a file and move or rename it.

python
import tempfile
from pathlib import Path

with tempfile.TemporaryDirectory() as tmp:
    root = Path(tmp)
    out = root / 'reports' / '2026'
    out.mkdir(parents=True, exist_ok=True)
    out.mkdir(parents=True, exist_ok=True)

    (out / 'sales.csv').write_text('a,b', encoding='utf-8')
    (out / 'costs.csv').write_text('c,d', encoding='utf-8')
    (out / 'notes.txt').write_text('x', encoding='utf-8')
    (out / 'deep.py').write_text('', encoding='utf-8')
    (root / 'tool.py').write_text('', encoding='utf-8')

    print(sorted(c.name for c in out.iterdir()))
    print(sorted(c.name for c in out.glob('*.csv')))
    print(sorted(c.name for c in root.rglob('*.py')))

    (out / 'notes.txt').rename(out / 'notes.md')
    (out / 'costs.csv').unlink()
    print(sorted(c.name for c in out.iterdir()))

Create nested folders, list and search them, then rename and delete

output
['costs.csv', 'deep.py', 'notes.txt', 'sales.csv']
['costs.csv', 'sales.csv']
['deep.py', 'tool.py']
['deep.py', 'notes.md', 'sales.csv']

Notice that the second mkdir call did nothing and raised no error. glob('*.csv') found only the two CSV files in out, while rglob('*.py') also reached deep.py inside the nested folders from the root. After the rename and the delete, three files remain.

MethodWhat it doesWatch out for
p.mkdir(parents=True, exist_ok=True)Creates the folder and any missing parentsWithout both flags it raises on missing parents or an existing folder
p.iterdir()Yields every child of a folderOrder is not guaranteed; includes folders
p.glob('*.csv')Children matching a pattern, one levelDoes not enter subfolders
p.rglob('*.py')Matches in this folder and all below itCan be slow on huge trees
p.unlink()Deletes a fileRaises if the file is missing; it does not delete folders
p.rename(new)Renames or moves to a new pathAn existing target may be replaced or cause an error, depending on the OS
Which helper for this job?

pathlib versus strings, and interview points

Older code builds paths from strings. os.path.join(a, b) joins pieces with the right separator, and plain concatenation such as 'data' + '/' + 'in' hard-codes one. Both work on strings, so every extra step needs another function: os.path.basename, os.path.splitext, os.path.exists, and so on, all wrapped around the same string. pathlib is object-based: the path carries its own methods, you chain them left to right, and the code reads like a description of what you want.

Strings (os.path)pathlib
TypePlain strPath object
Joinos.path.join('data', 'in') or +Path('data') / 'in'
File nameos.path.basename(p)p.name
Extensionos.path.splitext(p)[1]p.suffix
Read a fileopen(p) then read() then closep.read_text(encoding='utf-8')
ReadabilityNested calls, read inside outChained, read left to right

Windows paths bring a second trap. In an ordinary string literal the backslash starts an escape sequence, so 'C:\new\table.txt' contains a real newline after C: and a real tab before able. The path is silently wrong. Prefix the literal with r to make a raw string, or let Path build the path for you. The example uses PureWindowsPath, which formats paths the Windows way on any machine, so its output is the same everywhere.

python
from pathlib import PureWindowsPath

bad = 'C:\new\table.txt'
good = r'C:\new\table.txt'
print('\n' in bad, '\t' in bad, len(bad), len(good))

win = PureWindowsPath(good)
print(win.name, win.parent)
print(PureWindowsPath(r'C:\data') / 'in' / 'a.txt')

Escape characters in a plain Windows path string

output
True True 14 16
table.txt C:\new
C:\data\in\a.txt

The plain string is two characters shorter, 14 against 16, because \n and \t each collapsed into a single control character. The raw string kept every backslash.

Common mistake: backslashes in plain strings

Writing open('C:\new\table.txt') looks fine but asks for a file whose name contains a newline and a tab. Use r'C:\new\table.txt', forward slashes, or Path. Never type a Windows path as a plain string.

Interview point: four questions

Why prefer pathlib over os.path? Because a Path is an object that bundles the path with its operations, the / operator handles separators for any OS, and the same code reads, writes, searches and renames files without a pile of string functions. What does p.stem return for 'a.tar.gz'? Only a.tar, because .stem strips just the last suffix and .suffix is .gz. The full list of extensions is in .suffixes. How do you create nested folders safely? With mkdir(parents=True, exist_ok=True). What does Path('.') resolve to? The current working directory as an absolute path, the same as Path.cwd(). The example checks the last three.

python
from pathlib import Path

q = Path('backup/a.tar.gz')
print(q.stem)
print(q.suffix)
print(q.suffixes)

print(Path('.'))
print(Path('.').resolve() == Path.cwd())

Answers to the interview questions, checked in code

output
a.tar
.gz
['.tar', '.gz']
.
True
Remember

Join with /, read the pieces with .name, .stem, .suffix and .parent, and let read_text, write_text and mkdir(parents=True, exist_ok=True) do the routine work. Keep Windows paths out of plain strings.

Part 7 · Encodings and UTF-8

Bytes, text and UTF-8

A file on disk holds only bytes, which are numbers from 0 to 255. Python's str holds text, which is a sequence of characters. An encoding is the rule that turns one into the other. str.encode() goes from text to bytes, and bytes.decode() goes back. When you call open() in text mode, Python does this conversion for you on every read and write.

strbytes
HoldsCharacters (text)Numbers 0 to 255
Example'caf\u00e9'b'caf\xc3\xa9'
Length countsCharactersBytes
Turned into the other with.encode(encoding).decode(encoding)
Used byopen(..., 'r'), printopen(..., 'rb'), sockets, raw files

UTF-8 is the encoding you should use almost everywhere. It is variable length: each character takes between one and four bytes. Plain ASCII characters take 1 byte, most accented Latin letters take 2, many other scripts and symbols such as the euro sign take 3, and emoji take 4. For example, '\u00e9'.encode('utf-8') gives b'\xc3\xa9', two bytes for one character.

The example below encodes four characters and prints how many bytes each one needs. It writes the characters as escapes and prints them with ascii(), so the output is plain ASCII and looks the same on every machine.

python
samples = ['A', '\u00e9', '\u20ac', '\U0001f600']
for ch in samples:
    data = ch.encode('utf-8')
    print(ascii(ch), len(data), data)

One character, one to four bytes

output
'A' 1 b'A'
'\xe9' 2 b'\xc3\xa9'
'\u20ac' 3 b'\xe2\x82\xac'
'\U0001f600' 4 b'\xf0\x9f\x98\x80'
Remember

Length in characters and length in bytes are different numbers. len('caf\u00e9') is 4, but the UTF-8 bytes of that word are 5 long.

Why you must name the encoding

If you call open() without encoding=, Python picks the platform default, which comes from locale.getpreferredencoding(). On many Windows machines that is cp1252, while on Linux and macOS it is usually UTF-8. The same script can therefore read the same file in two different ways on two different computers, and nothing warns you.

The next example writes a file as UTF-8 and then reads it twice, once with the right encoding and once with cp1252. The wrong read does not fail. It quietly produces mojibake, which is garbled text, because every byte of the two-byte letter is read as its own character.

python
import tempfile
from pathlib import Path

path = Path(tempfile.gettempdir()) / 'enc_demo.txt'
path.write_text('caf\u00e9', encoding='utf-8')

print(path.read_bytes())
print(ascii(path.read_text(encoding='utf-8')))
print(ascii(path.read_text(encoding='cp1252')))

Same bytes, two readings (shown with ascii() so the output is plain ASCII)

output
b'caf\xc3\xa9'
'caf\xe9'
'caf\xc3\xa9'

The bytes on disk are the same in all three lines. The second line is the correct word. The third line has five characters instead of four, because \xc3 and \xa9 were treated as two separate cp1252 characters. The fix is one short habit: always write encoding='utf-8' in every open(), read_text() and write_text() call.

Common mistake

Leaving out encoding= because the script works on your laptop. It may fail, or silently garble text, on a teammate's Windows machine or on a server.

Which encoding do I pass?

Decode errors, BOMs and line endings

A UnicodeDecodeError means the bytes in the file are not valid for the encoding you asked for. It is common with files saved by older Windows programs, which often use cp1252 or latin-1. You have three options: make the decoder forgiving, accept some loss, or find the real encoding.

OptionWhat happensRisk
errors='replace'Bad bytes become the replacement characterYou can see where damage happened
errors='ignore'Bad bytes are droppedData is lost silently
encoding='cp1252' or 'latin-1'Bytes are read with the encoding that wrote themCorrect only if you guessed right

Here the bytes come from a Latin-1 file. UTF-8 refuses them, and each option then gives a different result. The printed text uses ascii() so the replacement character shows up as the escape \ufffd.

python
data = b'caf\xe9 au lait'

try:
    data.decode('utf-8')
except UnicodeDecodeError as e:
    print(e)

print(ascii(data.decode('utf-8', errors='replace')))
print(ascii(data.decode('utf-8', errors='ignore')))
print(ascii(data.decode('latin-1')))

One byte, four outcomes

output
'utf-8' codec can't decode byte 0xe9 in position 3: invalid continuation byte
'caf\ufffd au lait'
'caf au lait'
'caf\xe9 au lait'

Only the last line recovers the real word. Prefer finding the true encoding over replace or ignore, and use ignore only when losing data is acceptable.

A BOM (byte order mark) is a few special bytes at the very start of a file. In UTF-8 it is EF BB BF. Excel on Windows looks for it to recognise a CSV as UTF-8, so a file without it may show accented letters wrongly. The codec 'utf-8-sig' writes the BOM when writing and removes it when reading. Plain 'utf-8' leaves it in as a stray character \ufeff.

python
path.write_text('caf\u00e9', encoding='utf-8-sig')
print(path.read_bytes())
print(ascii(path.read_text(encoding='utf-8')))
print(ascii(path.read_text(encoding='utf-8-sig')))

Reuses path from the earlier example

output
b'\xef\xbb\xbfcaf\xc3\xa9'
'\ufeffcaf\xe9'
'caf\xe9'

Line endings are the last encoding-related detail. In text mode, Python turns each '\n' you write into the operating system's line ending, which is \r\n on Windows. The csv module writes its own line endings, so you pass newline='' to stop Python translating them a second time and producing blank rows. The example sets newline explicitly so the result is the same on every system.

python
with open(path, 'w', encoding='utf-8', newline='') as f:
    f.write('a\nb\n')
print(path.read_bytes())

with open(path, 'w', encoding='utf-8', newline='\r\n') as f:
    f.write('a\nb\n')
print(path.read_bytes())

newline='' writes \n untouched; newline='\r\n' converts it

output
b'a\nb\n'
b'a\r\nb\r\n'
Common mistake

Opening a CSV for writing without newline=''. On Windows this adds an extra \r and you get blank lines between rows.

Interview: what is the difference between str and bytes?
  • str is text made of characters; bytes is raw numbers from 0 to 255.
  • You convert with .encode() (str to bytes) and .decode() (bytes to str), always naming an encoding.
'caf\u00e9'.encode('utf-8')   # b'caf\xc3\xa9'
b'caf\xc3\xa9'.decode('utf-8')  # 'caf\u00e9'
Interview: why should you always pass encoding='utf-8'?
  • Without it, Python uses the platform default, often cp1252 on Windows and UTF-8 elsewhere.
  • The same code would then read or write different bytes on different machines.
Interview: what causes UnicodeDecodeError, and what is a BOM?
  • UnicodeDecodeError: the bytes are not valid in the encoding you asked for. Fix it by finding the real encoding, or use errors='replace' or 'ignore' knowing the cost.
  • BOM: marker bytes at the start of a file. Use 'utf-8-sig' to read or write UTF-8 with a BOM, for example CSVs meant for Excel.

Part 8 · CSV with the csv Module

Reading and writing rows

A CSV file looks like plain text with commas, so it is tempting to read it with line.split(','). That works until a field contains a comma itself, such as the name Smith, John. The csv module exists to handle exactly these cases: it knows the quoting rules, so you get back the fields the file's author meant.

csv.writer(f) wraps an open file. Call writerow(row) to write one list, or writerows(rows) to write many lists at once. When a value contains the delimiter, a quote or a newline, the writer wraps it in double quotes for you.

python
import csv
import tempfile
from pathlib import Path

folder = Path(tempfile.mkdtemp())
path = folder / "people.csv"

rows = [
    ["name", "city", "age"],
    ["Smith, John", "Leeds", "34"],
    ["Ana Ruiz", "Lima", "29"],
]

with open(path, "w", newline="", encoding="utf-8") as f:
    writer = csv.writer(f)
    writer.writerow(rows[0])
    writer.writerows(rows[1:])

print(path.read_text(encoding="utf-8"), end="")

The comma inside Smith, John is protected by quotes

output
name,city,age
"Smith, John",Leeds,34
Ana Ruiz,Lima,29

Reading is the mirror image. csv.reader(f) is an iterator: each step gives you one row as a list of strings. Because it understands the quotes, "Smith, John" comes back as a single field, while a plain split(',') tears it in two.

python
with open(path, newline="", encoding="utf-8") as f:
    for row in csv.reader(f):
        print(row)

line = '"Smith, John",Leeds,34'
print(line.split(","))

reader keeps the name whole; split breaks it

output
['name', 'city', 'age']
['Smith, John', 'Leeds', '34']
['Ana Ruiz', 'Lima', '29']
['"Smith', ' John"', 'Leeds', '34']
Common mistake: expecting numbers

Every value from csv.reader is a str, even 34. Adding two ages gives '3429', not 63. Convert with int() or float() yourself.

Named columns, newline and when to use csv

Counting positions gets tedious. csv.DictReader(f) takes the first row as the header and yields each later row as a dict, so you write row['name'] instead of row[0]. For writing, csv.DictWriter(f, fieldnames=[...]) needs the column names up front, and you must call writeheader() before the first writerow(dict), or the header line never appears.

python
fields = ["name", "city", "age"]
with open(path, "w", newline="", encoding="utf-8") as f:
    writer = csv.DictWriter(f, fieldnames=fields)
    writer.writeheader()
    writer.writerow({"name": "Smith, John", "city": "Leeds", "age": 34})
    writer.writerow({"name": "Ana Ruiz", "city": "Lima", "age": 29})

total = 0
with open(path, newline="", encoding="utf-8") as f:
    for row in csv.DictReader(f):
        print(row["name"], type(row["age"]).__name__)
        total += int(row["age"])
print("total age:", total)

Values go in as numbers but come back as str

output
Smith, John str
Ana Ruiz str
total age: 63

One habit goes with every CSV file you open: pass newline='' (and encoding='utf-8'). The csv module handles line endings itself. If Python also translates them, writing on Windows can add a blank line between rows, and quoted fields that contain a newline can be split wrongly when read back.

Common mistake: forgetting newline=''

open('out.csv', 'w') without newline='' often gives blank lines between rows on Windows. Always write open(path, 'w', newline='', encoding='utf-8'), for reading as well as writing.

csv.readercsv.DictReader
Each row isA list of stringsA dict keyed by the header
AccessPosition: row[2]Name: row['age']
If columns moveSilently reads the wrong columnStill finds the right one
Header rowYou must skip it yourselfConsumed for you
Best forFiles with no headerFiles with a header
python
import io

old = "name,age\nSmith,34\n"
new = "age,name\n34,Smith\n"
for text in (old, new):
    first = list(csv.reader(io.StringIO(text)))[1]
    named = next(csv.DictReader(io.StringIO(text)))
    print("reader:", first[0], "| DictReader:", named["name"])

tsv = "a\tb\n1\t2\n"
for row in csv.reader(io.StringIO(tsv), delimiter="\t"):
    print(row)

The same code after the columns swap, then a tab-separated file

output
reader: Smith | DictReader: Smith
reader: 34 | DictReader: Smith
['a', 'b']
['1', '2']

After the columns swap, positional access quietly returns the age while DictReader still finds the name. For tab-separated files (TSV), pass delimiter='\t'. The csv module reads one row at a time and needs no install, so it suits light, streaming work. For heavy analysis such as grouping, joining or statistics, reach for pandas.

Why not just use split(',')?
  • A quoted field like "Smith, John" contains a comma, so split cuts it into two pieces.
  • csv also handles escaped quotes and newlines inside fields, which split cannot.
Why open CSV files with newline=''?
  • The csv module manages line endings itself; extra translation causes blank rows on Windows.
  • It also keeps quoted fields that contain newlines intact when reading.
What type are all values from csv.reader, and when would you pick DictReader?
  • Every value is a str, so convert with int() or float().
  • Pick DictReader when the file has a header: access by name survives reordered columns and reads more clearly.

Part 9 · JSON with json.load and json.dump

Four functions, one memory aid

JSON is the format most programs use to exchange structured data: settings files, API responses, saved game state. Python's standard json module converts between Python objects and JSON text. It gives you four functions, and the only choice is where the text lives: in a file object or in a string.

Reads JSON into PythonWrites Python out as JSON
File objectjson.load(f)json.dump(obj, f)
Stringjson.loads(s)json.dumps(obj)
Memory aid

The trailing s means string. loads is "load from string" and dumps is "dump to string". The versions without the s work on an open file.

Start with the string functions, because you can see their result directly. dumps turns a dictionary into JSON text, and loads turns that text back into an equal dictionary. Notice that Python's True and None appear in the text as true and null.

python
import json

data = {"city": "Zurich", "ratings": [5, 4], "open": True, "owner": None}
text = json.dumps(data)
print(text)
print(json.loads(text) == data)

Round trip through a string

output
{"city": "Zurich", "ratings": [5, 4], "open": true, "owner": null}
True

The file versions do the same job but take an open file as the second argument (for dump) or the only argument (for load). Open JSON files in text mode and pass encoding="utf-8", as you learned in the encodings section. The example below uses a temporary directory so it leaves nothing behind.

python
from pathlib import Path
from tempfile import TemporaryDirectory

with TemporaryDirectory() as tmp:
    path = Path(tmp) / "city.json"
    with open(path, "w", encoding="utf-8") as f:
        json.dump(data, f)
    with open(path, encoding="utf-8") as f:
        loaded = json.load(f)
print(loaded["ratings"])

dump writes to a file, load reads it back

output
[5, 4]
Common mistake

Passing a filename to json.load("city.json") fails, because load wants an open file object, not a path. If you only have a string of JSON, use loads. If you have a path, open it first.

What JSON can and cannot hold

JSON has only a handful of types, and each one maps onto a Python type. When you load a document, you get the Python side of this table. When you dump, Python values are converted the other way.

JSONPython
objectdict
arraylist
stringstr
numberint or float
true / falseTrue / False
nullNone
python
doc = json.loads('{"a": 1, "b": 2.5, "c": true, "d": null, "e": [1, 2], "f": "x"}')
print({k: type(v).__name__ for k, v in doc.items()})

Loading gives you the Python types from the table

output
{'a': 'int', 'b': 'float', 'c': 'bool', 'd': 'NoneType', 'e': 'list', 'f': 'str'}

Python has many types that JSON does not. A tuple is accepted but is written as an array, so it comes back as a list: JSON has no way to say "this array was a tuple". Other types, such as set, datetime and your own classes, are not accepted at all.

Can json.dumps write this value?

The first branch is the easy one: basic types are written exactly as the table shows. Only values that fall through to the second question need attention. The next example shows both outcomes: the tuple survives as a list, and the set raises a TypeError.

python
print(json.dumps((1, 2)))
print(json.loads(json.dumps((1, 2))))

try:
    json.dumps({1, 2})
except TypeError as err:
    print(err)

A tuple turns into a list; a set is rejected

output
[1, 2]
[1, 2]
Object of type set is not JSON serializable

You have two ways out. Convert the value yourself before dumping (for example sorted(my_set) gives a list), or pass default=str. The default function is called for every value JSON cannot handle, and its return value is written instead. With str, a date becomes its text form.

python
from datetime import date

record = {"day": date(2026, 1, 5), "tags": sorted({"b", "a"})}
print(json.dumps(record, default=str))

default=str handles the date; the set was converted first

output
{"day": "2026-01-05", "tags": ["a", "b"]}
Common mistake

default=str is a quick fix, not a round trip. The date is now a plain string, and json.load will not turn it back into a date. Converting a set with default=str also gives you the text "{'a'}" rather than a list, so convert sets with sorted(...) or list(...) first.

Readable output and key quirks

By default dump writes everything on one line with no spacing choices, which is compact but hard to read. Three options fix that. indent sets how many spaces to indent nested levels, which spreads the document over many lines. sort_keys writes dictionary keys in alphabetical order, which makes files stable and easy to compare. ensure_ascii controls non-ASCII characters: the default True escapes them as \uXXXX, while False writes them as they are.

The first example uses indent and sort_keys. The name contains a ü, and with the default ensure_ascii=True it is written as an escape. That keeps this printed output pure ASCII.

python
data = {"name": "Zürich", "ratings": [5, 4], "a": 1}
print(json.dumps(data, indent=2, sort_keys=True))

indent=2 and sort_keys=True

output
{
  "a": 1,
  "name": "Z\u00fcrich",
  "ratings": [
    5,
    4
  ]
}

Escapes are valid JSON, but a person opening the file sees Z\u00fcrich. With ensure_ascii=False the file holds the real character, so you must open it with encoding="utf-8" when writing. To see the difference without depending on your terminal's encoding, the next example prints the encoded bytes: ü becomes the two UTF-8 bytes \xc3\xbc.

python
raw = json.dumps({"name": "Zürich"}, ensure_ascii=False)
print(raw.encode("utf-8"))

ensure_ascii=False keeps the character itself

output
b'{"name": "Z\xc3\xbcrich"}'

A common way to combine all three for a file you want people to read is json.dump(data, f, indent=2, ensure_ascii=False, sort_keys=True).

There is one more quirk. JSON object keys are always strings, so Python dictionary keys of other types are converted to strings on the way out. They stay strings when you load the file again.

python
print(json.loads(json.dumps({1: "a"})))

The int key 1 comes back as '1'

output
{'1': 'a'}
Common mistake

After a round trip, data[1] raises a KeyError because the key is now '1'. Look up data['1'], or convert the keys back with {int(k): v for k, v in data.items()}.

Readable files

Use indent to make it readable, sort_keys to make it stable, and ensure_ascii=False plus encoding="utf-8" to keep real characters in the file.

Big files, JSON Lines and interview questions

A JSON file holds one document. To read it, json.load must parse the whole thing and build the full object in memory. That is fine for settings or a small export. It is a poor fit for logs or event data that grows forever, because you cannot append one more record to a valid JSON array without rewriting the closing bracket.

JSON Lines (often saved as .jsonl) solves this by writing one complete JSON object per line. To add a record you append a line. To read, you walk the file line by line and call json.loads on each line, so only one record is in memory at a time.

One JSON documentJSON Lines
StructureOne value for the whole fileOne object per line
Readingjson.load(f) loads it all at oncejson.loads(line) for each line
MemoryWhole file in memoryOne record at a time
Adding dataRewrite the documentAppend a line
Best forSettings, small exportsLogs, events, large or growing data
python
with TemporaryDirectory() as tmp:
    path = Path(tmp) / "events.jsonl"
    with open(path, "w", encoding="utf-8") as f:
        for i in range(1, 4):
            f.write(json.dumps({"id": i, "ok": i != 2}) + "\n")

    good = 0
    with open(path, encoding="utf-8") as f:
        for line in f:
            record = json.loads(line)
            if record["ok"]:
                good += 1
print(good)

Write one object per line, then stream it back

output
2

Each line is written with json.dumps (never with indent, since a record must fit on one line) and ends with \n. Reading is the same line-by-line loop you used for text files.

What is the difference between load and loads?
  • load(f) reads JSON from an open file object; loads(s) parses JSON from a string.
  • The same split applies to dump(obj, f) and dumps(obj).
  • The s stands for string.
Why does a tuple come back as a list?
  • JSON has only arrays, with no separate tuple type.
  • dumps writes a tuple as an array, and loads always builds a list from an array.
  • If you need a tuple, convert it yourself after loading.
json.loads(json.dumps((1, 2)))  # [1, 2]
What happens to int dict keys?
  • JSON object keys must be strings, so 1 is written as "1".
  • On loading you get the string key back, so {1: 'a'} reloads as {'1': 'a'}.
  • Convert the keys with int(k) if you need numbers again.
When would you choose JSON Lines?
  • When data is large or keeps growing, such as logs or events.
  • You can append a record without rewriting the file.
  • You can read it line by line with json.loads(line) instead of loading everything into memory.

Part 10 · Handling Missing Files and Errors

The Exceptions Files Raise

Files live outside your program, so they can fail in ways your code cannot control. A path may be wrong, a folder may be locked, or a file may hold bytes that are not valid text. Python reports each of these problems as a specific exception. If you learn the names, you can catch exactly the problem you can fix and let the rest crash loudly.

Four of the common ones belong to one family. FileNotFoundError, PermissionError, IsADirectoryError and FileExistsError are all subclasses of OSError. Writing except OSError therefore catches every one of them. The two content errors, UnicodeDecodeError and json.JSONDecodeError, are different. They come from bad data, not from the operating system, and both are kinds of ValueError.

ExceptionWhen it happensParent
FileNotFoundErrorThe file or one of its folders does not existOSError
PermissionErrorYou may not read, write or list that pathOSError
IsADirectoryErrorYou tried to open a folder as if it were a fileOSError
FileExistsErrorMode 'x' was used on a file that already existsOSError
UnicodeDecodeErrorThe bytes are not valid in the encoding you choseUnicodeError, then ValueError
json.JSONDecodeErrorThe text is not valid JSONValueError

You can check this family tree yourself. The program below asks Python for the parent of FileNotFoundError and tests each class against OSError.

python
import json

print(FileNotFoundError.__mro__[1].__name__)
for exc in (FileNotFoundError, PermissionError, IsADirectoryError, FileExistsError):
    print(exc.__name__, issubclass(exc, OSError))
for exc in (UnicodeDecodeError, json.JSONDecodeError):
    print(exc.__name__, issubclass(exc, OSError), exc.__mro__[1].__name__)
output
OSError
FileNotFoundError True
PermissionError True
IsADirectoryError True
FileExistsError True
UnicodeDecodeError False UnicodeError
JSONDecodeError False ValueError

Next, here is each error raised on purpose. A temporary folder keeps the demo safe to run anywhere. The program prints only the exception name, which is the part you will write in your own except clauses.

python
import json
import tempfile
from pathlib import Path

with tempfile.TemporaryDirectory() as tmp:
    folder = Path(tmp)
    note = folder / "note.txt"
    note.write_text("hi", encoding="utf-8")

    attempts = {
        "open missing file": lambda: open(folder / "nope.txt"),
        "mode x on existing": lambda: open(note, "x"),
        "decode bad bytes": lambda: b"\xff\xfe".decode("utf-8"),
        "parse bad json": lambda: json.loads("{oops"),
    }
    for label, action in attempts.items():
        try:
            action()
        except Exception as exc:
            print(f"{label}: {type(exc).__name__}")
output
open missing file: FileNotFoundError
mode x on existing: FileExistsError
decode bad bytes: UnicodeDecodeError
parse bad json: JSONDecodeError
Operating systems differ

Opening a folder with open() raises IsADirectoryError on Linux and macOS, but Windows often reports PermissionError instead. Because both are OSError, catching OSError works the same on every system.

Catching a Missing File, and Where Relative Paths Point

The basic way to handle a missing file is to put the with open(...) block inside a try and name the exception you expect. If the file is there, the try body runs normally. If not, Python jumps straight to the except clause and your program carries on.

There is a catch that surprises many learners. A relative path such as "note.txt" is not resolved against your script's folder. Python resolves it against the current working directory, which is the folder your terminal was in when you started the program. The same script with the same path can succeed in one terminal and fail in another.

The program below makes the effect visible. It builds two folders and puts note.txt in only one of them. It then changes the working directory and calls the same function twice.

python
import os
import tempfile
from pathlib import Path

def read_note():
    try:
        with open("note.txt", encoding="utf-8") as f:
            return f.read()
    except FileNotFoundError:
        return "note.txt not found here"

start = os.getcwd()
with tempfile.TemporaryDirectory() as home, tempfile.TemporaryDirectory() as elsewhere:
    (Path(home) / "note.txt").write_text("hello from home", encoding="utf-8")
    os.chdir(home)
    print(read_note())
    os.chdir(elsewhere)
    print(read_note())
    os.chdir(start)
output
hello from home
note.txt not found here

To make a path independent of the terminal, anchor it to the script itself. The built-in __file__ holds the location of the running script, so Path(__file__).parent / "data.txt" always points at data.txt beside the script, no matter where you launched it. Use Path(__file__).resolve().parent if you also want an absolute path.

Common mistake: it works on my machine

A script that opens "data.txt" works when you run it from its own folder, then fails from your home folder, a scheduler or an editor's Run button. When the error says the file is missing but you can see it, print Path.cwd() first. Then switch to a path anchored on Path(__file__).parent.

Look Before You Leap, or Ask Forgiveness

There are two styles for dealing with a file that might not exist. LBYL means Look Before You Leap. You check first with if p.exists(): and open only when the answer is yes. EAFP means Easier to Ask Forgiveness than Permission. You simply try to open the file and handle the exception if it fails. Python programmers prefer EAFP, and for files there is a concrete reason.

LBYL (check first)EAFP (try, then catch)
Code shapeif p.exists(): open(p)try: open(p) except FileNotFoundError
Race conditionYes, the file can vanish after the checkNo, the open itself is the test
Other failuresPermission and decoding errors still need a tryOne try covers every failure
Python styleDiscouraged for filesIdiomatic and safe

The weakness of LBYL is a race condition. Between the moment you ask whether the file exists and the moment you open it, another program, a cleanup job or the user can delete it. The check gave a true answer that went stale. The program below imitates that gap by deleting the file right after the check.

python
import tempfile
from pathlib import Path

with tempfile.TemporaryDirectory() as tmp:
    p = Path(tmp) / "job.txt"
    p.write_text("x", encoding="utf-8")

    if p.exists():
        p.unlink()  # another process removes the file here
        try:
            p.open()
        except FileNotFoundError:
            print("exists() said yes, open() still failed")
output
exists() said yes, open() still failed

Since you need the try anyway, the exists() check adds nothing. It only gives a false feeling of safety. Catch the narrowest exception you actually know how to handle. FileNotFoundError says exactly what you mean, while a wide net such as except Exception also swallows problems you never planned for. If you want both a specific and a general handler, put the specific one first, because Python uses the first clause that matches.

Common mistake: checking and then opening

Code like if os.path.exists(name): f = open(name) still crashes with PermissionError or IsADirectoryError. It can even hit FileNotFoundError under the race above. Drop the if and wrap the open in a try.

What To Do When a File Is Missing

Catching the error is only half the job. You also need to decide what a missing file means for your program. There are three honest answers, and the right one depends on the file's role.

Choosing a response to a missing file

A default on missing is right for optional files such as a config. A first run has no config.json yet, so the function returns an empty dictionary. A create on missing is right when you are about to write. The file will not exist yet, and neither might its folder, so call p.parent.mkdir(parents=True, exist_ok=True) before writing. When the file is truly required, you can do nothing useful locally. Log and re-raise with a bare raise, which keeps the original traceback for whoever can handle it.

python
import json
import logging
import sys
import tempfile
from pathlib import Path

logging.basicConfig(stream=sys.stdout, format="%(levelname)s: %(message)s")
log = logging.getLogger("files")

def load_config(path):
    try:
        with open(path, encoding="utf-8") as f:
            return json.load(f)
    except FileNotFoundError:
        return {}

def save_config(path, data):
    path = Path(path)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(data), encoding="utf-8")

def read_report(path):
    path = Path(path)
    try:
        return path.read_text(encoding="utf-8")
    except OSError as exc:
        log.error("cannot read %s: %s", path.name, type(exc).__name__)
        raise

with tempfile.TemporaryDirectory() as tmp:
    cfg = Path(tmp) / "app" / "config.json"
    print(load_config(cfg))
    save_config(cfg, {"theme": "dark"})
    print(load_config(cfg))
    try:
        read_report(Path(tmp) / "report.txt")
    except FileNotFoundError:
        print("re-raised FileNotFoundError")
output
{}
{'theme': 'dark'}
ERROR: cannot read report.txt: FileNotFoundError
re-raised FileNotFoundError

Look at how the pieces fit. load_config returns {} on the first call, because the folder does not exist yet. save_config creates the app folder and the file in one step. read_report logs the problem and then re-raises, so the caller still sees the original FileNotFoundError.

Common mistake: the bare except

Writing except: with no exception name catches everything, including typos such as f.red() instead of f.read(). It even catches Ctrl+C. The function quietly returns its fallback value and the real bug stays hidden. Always name the exception, and use except Exception only at the outermost layer when you also log the traceback.

Remember

Try the operation, catch the narrowest exception you can handle, and decide deliberately between a default, creating the file, or logging and re-raising.

Interview Points

Why is EAFP preferred over exists()?
  • An exists() check is stale the moment it returns, because the file can be deleted before you open it, which is a race condition.
  • You need a try anyway, since PermissionError, IsADirectoryError and decoding problems can still occur.
  • EAFP tests the real operation, so one try/except handles every case and the code stays shorter.
try:
    with open(path, encoding="utf-8") as f:
        data = f.read()
except FileNotFoundError:
    data = ""
What is the parent class of FileNotFoundError?
  • OSError. PermissionError, IsADirectoryError and FileExistsError share that parent.
  • So except OSError catches all of them, but a handler for the narrower FileNotFoundError should come first if you want special treatment.
  • UnicodeDecodeError and json.JSONDecodeError are not part of this family. They are ValueError subclasses.
Why does a relative path work in one terminal and fail in another?
  • A relative path is resolved against the current working directory, which is the folder the terminal was in when the program started.
  • Different terminal, scheduler or editor setting means a different working directory, so the same name points to a different place.
  • Anchor the path to the script with Path(__file__).parent / "data.txt" so it no longer depends on where you run from.

Part 11 · Processing Large Files Line by Line

Why iterating over a file is memory-safe

A file object is an iterator over its lines. When you write for line in f:, Python reads a block of bytes into an internal buffer, hands you one line at a time, and refills the buffer when it runs dry. Lines you have already processed are dropped. Memory use therefore stays roughly constant whether the file is 1 MB or 100 GB. That is O(1) memory and O(n) time, because every line is visited once.

The opposite habit is to pull everything in at once. f.read() returns one giant string and f.readlines() returns a list holding every line, so both cost O(file size) memory. On a multi-gigabyte file that can exhaust RAM, slow the machine to a crawl, or crash the program with a MemoryError.

CallWhat you getMemorySafe for huge files?
for line in fOne line per stepO(1), roughly constantYes
f.read(n)At most n characters or bytesO(n) for the chunk size you pickYes
f.readline()One line per callO(1)Yes
f.read()The whole file as one stringO(file size)No
f.readlines()A list of every lineO(file size) plus list overheadNo

To try the examples below, the first one builds a small log file in a temporary folder. Later examples reuse the names work and log from it.

python
import tempfile
from pathlib import Path

work = Path(tempfile.mkdtemp())
log = work / "app.log"
with open(log, "w", encoding="utf-8") as f:
    for i in range(1, 11):
        level = "ERROR" if i % 4 == 0 else "INFO"
        f.write(f"{i:03d} {level} request {i}\n")

count = 0
with open(log, encoding="utf-8") as f:
    for line in f:
        count += 1
print(count)

Iterating counts lines without ever holding more than one of them

output
10

For contrast, here is what readlines() does with the same file. It builds a list with every line in it. That is harmless for ten lines and dangerous for ten million.

python
with open(log, encoding="utf-8") as f:
    lines = f.readlines()
print(len(lines), type(lines).__name__)
output
10 list
The rule of thumb

If you only need to look at each line once, loop over the file object. Reach for read() or readlines() only when you really need the whole content in memory at the same time.

Counting, filtering and writing as you go

Because the loop yields one line at a time, a generator expression can count matches directly from the file. sum(1 for line in f if 'ERROR' in line) adds one for every matching line and never builds a list.

python
with open(log, encoding="utf-8") as f:
    print(sum(1 for line in f if "ERROR" in line))
output
2

When you need the matching lines themselves, resist the urge to collect them in a list. Write each result to a second file the moment you find it. Memory stays flat on the output side too, and if the program dies halfway you still have everything written so far. Opening the source with errors='replace' is a good habit for messy logs, because a few bad bytes then turn into a replacement character instead of stopping the whole run.

python
errors_file = work / "errors.txt"
written = 0
with open(log, encoding="utf-8", errors="replace") as src, open(errors_file, "w", encoding="utf-8") as out:
    for line in src:
        if "ERROR" in line:
            out.write(line)
            written += 1
print(written)
print(errors_file.read_text(encoding="utf-8"), end="")

Read one line, test it, write it, move on

output
2
004 ERROR request 4
008 ERROR request 8

Here is what errors='replace' does in practice. The file below contains a byte, 0xff, that is not valid UTF-8. With the default strict setting the loop would raise UnicodeDecodeError. With replace the bad byte becomes the character U+FFFD and the loop carries on.

python
messy = work / "messy.log"
messy.write_bytes(b"ok line\nbad \xff byte\n")
with open(messy, encoding="utf-8", errors="replace") as f:
    for line in f:
        print("\ufffd" in line)
output
False
True
Common mistake: collecting results first

Writing hits = [l for l in f if 'ERROR' in l] on a huge log trades one big list for another and brings the memory problem straight back. Stream matches to an output file, or keep only a running count.

Chunks and generator pipelines

Not every file is made of lines. A video, an archive or a database dump may have no newlines at all, or one enormous line. For binary or unstructured data, read fixed-size chunks instead. The walrus operator makes the loop compact: while chunk := f.read(1024*1024): process(chunk) reads up to 1 MiB, stops when read returns an empty value at end of file, and keeps memory bounded by the chunk size. The example uses a small 4096-byte chunk so the output is short.

python
blob = work / "data.bin"
blob.write_bytes(b"x" * 10_000)
total = 0
with open(blob, "rb") as f:
    while chunk := f.read(4096):
        total += len(chunk)
        print(len(chunk))
print(total)

The last chunk is simply whatever is left

output
4096
4096
1808
10000

Several processing steps can be chained without building any intermediate list by writing each step as a generator function that yields lines. A generator produces its next value only when the loop asks for it, so a line travels through the whole pipeline before the next line is even read from disk.

A lazy pipeline: one line at a time
  1. 1Filefor line in f
  2. 2clean_linesstrip, skip blanks
  3. 3only_errorskeep matches
  4. 4Consumerprint or write
python
def clean_lines(f):
    for line in f:
        line = line.strip()
        if line:
            yield line

def only_errors(lines):
    for line in lines:
        if "ERROR" in line:
            yield line

with open(log, encoding="utf-8") as f:
    for line in only_errors(clean_lines(f)):
        print(line)

Each function is small and testable, and none of them builds a list

output
004 ERROR request 4
008 ERROR request 8
Why generators fit

A generator keeps only its current position, not the data it has already produced. That is why a pipeline of generators uses the same tiny amount of memory as a single loop, while staying readable.

Whole file or streaming? A comparison, plus interview questions

Streaming is not always better. For a small file, reading everything at once is shorter and just as fast. The comparison below shows when each approach is the right tool.

Whole-file readStreaming
Best forSmall files, under a few MBFiles that may not fit in RAM
CodeSimplest: text = f.read()A loop, a chunk loop or a generator
MemoryO(file size)O(1), or O(chunk size)
Random access to the dataEasy, it is all in memoryYou see each piece once; reopen to go again
JSONjson.load(f) parses the entire documentjson.loads(line) per line, one record at a time
CSVlist(csv.reader(f))Looping over csv.reader(f) or csv.DictReader(f)

Two formats stream naturally. The csv reader is itself an iterator, so you can loop over it row by row. A file with one JSON object per line (often called JSON Lines) can be read with json.loads on each line. In contrast, a single whole-file json.load must read and parse the entire document before it returns anything, so it does not stream.

python
import csv
import json

events = work / "events.jsonl"
events.write_text('{"user": "a", "ms": 120}\n{"user": "b", "ms": 340}\n{"user": "a", "ms": 80}\n', encoding="utf-8")
total_ms = 0
with open(events, encoding="utf-8") as f:
    for line in f:
        total_ms += json.loads(line)["ms"]
print(total_ms)

sales = work / "sales.csv"
sales.write_text("item,qty\npen,3\nbook,5\npen,4\n", encoding="utf-8")
with open(sales, newline="", encoding="utf-8") as f:
    print(sum(int(row["qty"]) for row in csv.DictReader(f)))

Both loops hold only one record at a time

output
540
12
Which reading style should I use?

Interview Point

Why is iterating over f memory-safe?
  • The file object is an iterator that reads through a small buffer and returns one line per step.
  • Lines you are done with are discarded, so memory stays roughly constant: O(1) memory, O(n) time.
How would you process a 20 GB log on a laptop?
  • Open it once with encoding='utf-8' and errors='replace', then stream it line by line.
  • Keep only a counter or a small summary, and write any matches to an output file as you go.
  • If the file has no useful line structure, read it in fixed-size chunks.
with open("huge.log", encoding="utf-8", errors="replace") as f:
    print(sum(1 for line in f if "ERROR" in line))
What is the cost of readlines()?
  • It builds a list of every line, so memory grows with the file size: O(file size).
  • On a multi-gigabyte file it can exhaust RAM and crash with MemoryError; read() has the same problem.
How does a generator help?
  • It yields one cleaned line at a time, so you can chain steps (clean, filter, transform) without creating lists between them.
  • Only the current item is alive at once, so the pipeline keeps the same constant memory as a plain loop.

Part 12 · Common File Mistakes and Fixes

Wiping a File and Half-Written Saves

Most file bugs are not exotic. They come from a handful of habits that look fine until the day they cost you data. The first and most painful is opening a file with 'w' when you meant to keep what was in it. Mode 'w' truncates the file to zero bytes the moment open() returns, before you write a single character. Even if your program crashes on the very next line, the old contents are already gone.

ModeIf the file existsIf it does not existUse it when
'w'Erased instantlyCreatedYou truly want a fresh file
'a'Kept, new text goes at the endCreatedAdding log lines or records
'x'Raises FileExistsErrorCreatedYou must never overwrite

The fix is to pick the mode that matches your intent. Use 'a' to add to the end, and use 'x' when overwriting would be a bug, so Python refuses instead of silently destroying data. The example below keeps a note, appends to it, shows 'x' refusing, and then shows what a careless 'w' does.

python
from pathlib import Path
import tempfile

folder = Path(tempfile.mkdtemp())
notes = folder / 'notes.txt'
notes.write_text('keep me\n', encoding='utf-8')

with open(notes, 'a', encoding='utf-8') as f:
    f.write('added line\n')
print(notes.read_text(encoding='utf-8').splitlines())

try:
    with open(notes, 'x', encoding='utf-8') as f:
        f.write('new')
except FileExistsError:
    print('x refused to overwrite')

with open(notes, 'w', encoding='utf-8') as f:
    pass
print(repr(notes.read_text(encoding='utf-8')))
output
['keep me', 'added line']
x refused to overwrite
''
Common mistake

Opening the real data file with 'w' to "update" it. The truncation happens at open time, so any crash, exception or Ctrl+C before the write finishes leaves you with an empty file and no way back.

Updating safely: write a temp file, then replace

Sometimes you really do need to rewrite a whole file, for example saving a settings file. Writing straight over the original is risky because a crash halfway leaves a half-written file. The safe pattern is to write the complete new content to a temporary file next to the original, close it, and only then swap it into place with Path.replace. The swap is a single rename, so readers see either the whole old file or the whole new one, never a mixture.

Safe update
  1. 1Write new content to notes.txt.tmporiginal untouched
  2. 2Close the temp fileleaving the with block flushes it
  3. 3tmp.replace(original)one rename swaps it in
python
def save_safely(path, text):
    tmp = path.with_suffix(path.suffix + '.tmp')
    with open(tmp, 'w', encoding='utf-8') as f:
        f.write(text)
    tmp.replace(path)

save_safely(notes, 'version 2\n')
print(notes.read_text(encoding='utf-8'), end='')
print(sorted(p.name for p in folder.iterdir()))
output
version 2
['notes.txt']

If the program dies while writing the temp file, the original is still intact and you only have a stray .tmp file to delete. After a successful run no temp file is left, because replace moved it over the original.

Forgotten with, Encoding and newline

The next group of mistakes is about things you forgot to say. Each one works on your machine most of the time, which is exactly why they survive until someone else runs your code.

ForgottenWhat goes wrongFix
with / close()Buffered text may not reach the disk; files stay locked, especially on Windowswith open(...) as f: closes the file even when an error happens
encoding='utf-8'Python uses the platform default, so accents and symbols turn into garbage on another machineAlways pass encoding='utf-8' for text
newline='' for csvBlank lines between rows on Windows, or broken quoted fieldsopen(path, newline='', encoding='utf-8') for both reading and writing csv

The csv module handles line endings itself, so the file object must not translate them as well. With newline='' the writer's \r\n goes to disk exactly as written. Without it, Windows turns that into \r\r\n, which shows up as an empty row in spreadsheets.

python
import csv

rows_path = folder / 'scores.csv'
with open(rows_path, 'w', encoding='utf-8', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['name', 'score'])
    writer.writerow(['Asha', '10'])
print(rows_path.read_bytes())
output
b'name,score\r\nAsha,10\r\n'
Common mistake

Calling f = open(...) and never f.close(), or closing only on the happy path. If an exception fires between open and close, the file stays open. Use with and the problem disappears.

Reading twice gives an empty string

An open file keeps a position, like a cursor. read() moves it to the end, so a second read() starts at the end and finds nothing. It returns '' rather than raising an error, which makes this bug quiet and confusing. You can move the cursor back with seek(0), or simply open the file again.

python
with open(notes, encoding='utf-8') as f:
    first = f.read()
    second = f.read()
    f.seek(0)
    third = f.read()
print(repr(first), repr(second), repr(third))
output
'version 2\n' '' 'version 2\n'

The same thing happens when you loop over a file twice with for line in f. The second loop sees no lines. Either call seek(0) between them, or read the lines into a list once if the file is small enough.

Types and Stray Newlines

Files hold text, and Python does not guess what that text means. Three conversions trip people up again and again.

  • Every value read by the csv module is a str, even when it looks like a number. '10' + '5' joins the strings into '105'. Convert with int() or float() first.
  • write() accepts only str. Passing an int raises TypeError, so convert with str(n) or an f-string.
  • JSON object keys are always strings. A dict with key 1 comes back from a dump and load round trip with key '1', so data[1] fails with KeyError.

There is also the newline. When you iterate over a file, each line keeps its trailing '\n'. Calling line.strip() on a last line that has no newline is harmless: it simply returns the same text. Forgetting to strip is the real danger, because the stray newline stays in your data and makes comparisons fail. The last lines of the example below show a line that looks equal to 'yes' but is not until it is stripped.

python
import csv, json

with open(rows_path, newline='', encoding='utf-8') as f:
    rows = list(csv.DictReader(f))
score = rows[0]['score']
print(score + '5', int(score) + 5)

back = json.loads(json.dumps({1: 'one'}))
print(back, 1 in back, '1' in back)

try:
    notes.write_text(42, encoding='utf-8')
except TypeError:
    print('write needs str')

line = 'yes\n'
print(line == 'yes', line.strip() == 'yes')
output
105 15
{'1': 'one'} False True
write needs str
False True
Common mistake

Comparing raw lines such as if line == 'yes'. The line is really 'yes\n', so the test is False and the program quietly takes the wrong branch. Strip first, then compare.

Convert at the edge

Convert types right where data enters or leaves your program: int() straight after reading a csv cell, str() just before writing, and int(key) when loading JSON keys back. Then the rest of your code works with real numbers.

Paths, Big Files, Broad Excepts and Interview Questions

The last four mistakes are about structure. Building paths by gluing strings with + or backslashes breaks on other operating systems and invites typos. Path objects join with /, work on every platform, and give you .name, .suffix and .parent for free. A second trap is relying on the working directory: a relative path like 'data.txt' means whatever folder the program was started from. Anchor your paths to the script instead, for example Path(__file__).resolve().parent / 'data.txt'.

Loading a huge file with read() pulls every byte into memory at once. A file object is already an iterator over lines, so a for loop holds one line at a time. Finally, wrapping a whole block in except Exception hides real bugs such as typos and wrong types behind a vague message. Catch only the error you expect, like FileNotFoundError, and keep the try block small.

python
good = folder / 'data' / 'log.txt'
print(good.name, good.suffix, good.parent.name)

log = folder / 'app.log'
log.write_text('ok\nERROR disk\nok\nERROR net\n', encoding='utf-8')
count = 0
try:
    with open(log, encoding='utf-8') as f:
        for line in f:
            if line.startswith('ERROR'):
                count += 1
    with open(folder / 'missing.log', encoding='utf-8') as f:
        pass
except FileNotFoundError:
    print('missing.log not found')
print(count)
output
log.txt .txt data
missing.log not found
2
HabitBetterWhy
'C:\\data' + '\\' + namePath('C:/data') / namePortable and harder to get wrong
Relative path 'data.txt'Path(__file__).resolve().parent / 'data.txt'Independent of where you launch from
f.read() on a big filefor line in f:One line in memory at a time
except Exception: around everythingexcept FileNotFoundError: around the openReal bugs still show up
Common mistake

Writing except Exception: print('error') around a whole block. A typo in a variable name now prints the same message as a missing file, and you lose the traceback that would have told you what really broke.

Interview Point

How do you update a file safely so a crash cannot corrupt it?
  • Write the full new content to a temporary file in the same folder, never over the original.
  • Close the temp file so everything is flushed, then call tmp.replace(original).
  • The replace is one rename, so the file is always either entirely old or entirely new.
tmp = path.with_suffix('.tmp')
tmp.write_text(new_text, encoding='utf-8')
tmp.replace(path)
Why does reading a file after writing it in 'w+' mode return nothing until you seek(0)?
  • Mode 'w+' opens for both reading and writing, but there is a single position.
  • After write(), the position sits at the end of what you wrote, so read() starts at the end and returns ''.
  • Call f.seek(0) to move back to the start, then read. Note that 'w+' also truncates the file on open.
with open('t.txt', 'w+', encoding='utf-8') as f:
    f.write('hello')
    f.seek(0)
    print(f.read())
What is the first thing you check when text appears garbled?
  • The encoding: was the file written in one encoding and read in another?
  • Pass encoding='utf-8' explicitly on both the write and the read, rather than trusting the platform default.
  • If it is still wrong, find out what encoding the file really uses before trying others.
Garbled text

Part 13 · Python File Handling Cheat Sheet

Modes, Reading, Writing and Position

Every file operation starts with the mode you pass to open(). The mode decides whether the file may be read, written, created or erased, so choosing it carelessly is the quickest way to lose data. The table below is the whole vocabulary; modes combine, so rb means read bytes and w+ means write and read.

ModeWhat it doesIf the file is missing
rRead text (the default)Raises FileNotFoundError
wTruncate to empty, then writeCreates it
aAppend at the end, keep old contentCreates it
xCreate only, never overwriteCreates it; raises FileExistsError if it exists
+Add read or write on top of another mode, as in r+Depends on the base mode
bBinary: bytes instead of str, no encodingDepends on the base mode

The habit to memorise is one line: with open(p, 'r', encoding='utf-8') as f:. The with block closes the file for you even when an error happens, and the explicit encoding stops your program from depending on the machine it runs on. Leave the encoding out only for binary mode.

Common mistake

Opening an existing file with w to "add a line". Python empties the file the moment open() returns. Use a to add, or x when overwriting would be a bug.

Reading has four styles. f.read() returns everything as one string, f.readline() returns one line including its newline, f.readlines() returns a list of all lines, and looping with for line in f hands you one line at a time lazily. Writing has three: f.write(str) writes one string and does not add a newline, f.writelines(list) writes each string in a list without adding newlines either, and print(x, file=f) writes like print does, newline included.

TaskCallReturns or does
Read allf.read()One str, whole file in memory
Read somef.read(n)Up to n characters (bytes in binary)
Read one linef.readline()One line, newline kept
Read all linesf.readlines()List of lines in memory
Read lazilyfor line in fOne line at a time
Writef.write(s)Writes s, no newline added
Write manyf.writelines(items)Writes each str, no newlines added
Write like printprint(x, file=f)Writes x plus a newline
Where am If.tell()Current position
Jumpf.seek(0)Back to the start

A file keeps a position, and each read moves it forward. That is why a second f.read() gives an empty string, and why f.seek(0) is the way to start over. The example writes a file three ways, then reads it every way.

python
from pathlib import Path
import tempfile

tmp = Path(tempfile.mkdtemp())
p = tmp / 'notes.txt'

with open(p, 'w', encoding='utf-8') as f:
    f.write('one\n')
    f.writelines(['two\n', 'three\n'])
    print('four', file=f)

with open(p, 'r', encoding='utf-8') as f:
    print(repr(f.read(3)))
    print(f.tell())
    f.seek(0)
    print(repr(f.readline()))
    print(repr(f.read()))
    f.seek(0)
    print(f.readlines())
    f.seek(0)
    print([line.strip() for line in f])
output
'one'
3
'one\n'
'two\nthree\nfour\n'
['one\n', 'two\n', 'three\n', 'four\n']
['one', 'two', 'three', 'four']

pathlib, CSV, JSON and Errors

pathlib treats a path as an object instead of a string. The / operator joins parts, so Path(a) / b works the same on Windows and Linux. A path can read and write its own file with .read_text() and .write_text(), create folders with .mkdir(parents=True, exist_ok=True), list matches with .glob(), and tell you its parts through .suffix and .stem.

WantpathlibExample result
JoinPath(a) / ba/b
Read whole filep.read_text(encoding='utf-8')str
Write whole filep.write_text(s, encoding='utf-8')Character count
Make foldersp.mkdir(parents=True, exist_ok=True)No error if it exists
Find filesp.glob('*.txt')Matching paths
Extensionp.suffix'.txt'
Name without extensionp.stem'report'
python
data = tmp / 'data' / 'raw'
data.mkdir(parents=True, exist_ok=True)
report = data / 'report.txt'
report.write_text('hello\n', encoding='utf-8')
print(report.name, report.stem, report.suffix)
print(report.read_text(encoding='utf-8').strip())
print([x.name for x in sorted(data.glob('*.txt'))])
output
report.txt report .txt
hello
['report.txt']

For CSV, open the file with newline='' so the csv module controls line endings itself, then use csv.reader (rows as lists), csv.DictReader (rows as dicts keyed by the header), csv.writer or csv.DictWriter. Everything read from a CSV is a str, even numbers. For JSON, json.load(f) and json.dump(obj, f, indent=2, ensure_ascii=False) work on open files, while json.loads and json.dumps work on strings. Setting ensure_ascii=False keeps accented letters readable instead of escaping them.

python
import csv
import json

people = tmp / 'people.csv'
with open(people, 'w', newline='', encoding='utf-8') as f:
    w = csv.DictWriter(f, fieldnames=['name', 'age'])
    w.writeheader()
    w.writerow({'name': 'Asha', 'age': 31})
    w.writerow({'name': 'Ben', 'age': 27})

with open(people, newline='', encoding='utf-8') as f:
    rows = list(csv.DictReader(f))
print(rows)

out = tmp / 'people.json'
with open(out, 'w', encoding='utf-8') as f:
    json.dump(rows, f, indent=2, ensure_ascii=False)
text = out.read_text(encoding='utf-8')
print(text)
print(json.loads(text) == rows)
output
[{'name': 'Asha', 'age': '31'}, {'name': 'Ben', 'age': '27'}]
[
  {
    "name": "Asha",
    "age": "31"
  },
  {
    "name": "Ben",
    "age": "27"
  }
]
True
Common mistake

Expecting numbers back from a CSV. The ages were written as ints but come back as the strings '31' and '27'. Convert with int() yourself, and the JSON you save from those rows will hold strings too.

Python prefers EAFP: try the operation and handle the failure, rather than checking first and racing against the file system. Catch FileNotFoundError for a missing file and the broader OSError for permissions, folders and disk problems. Put the specific handler first. For big files, never call read() on the whole thing: iterate over lines, or read fixed-size chunks with a loop, so memory stays flat.

python
missing = tmp / 'nope.txt'
try:
    with open(missing, encoding='utf-8') as f:
        f.read()
except FileNotFoundError:
    print('missing file')
except OSError:
    print('other OS error')

big = tmp / 'big.txt'
big.write_bytes(b'ab\n' * 5000)

count = 0
with open(big, encoding='utf-8') as f:
    for line in f:
        count += 1
print(count)

total = 0
with open(big, 'rb') as f:
    while chunk := f.read(4096):
        total += len(chunk)
print(total)
output
missing file
5000
15000

Interview Point: CSV to JSON, Safe Overwrite, Failed Reads

Interviewers like this trio because each part tests a different habit. Start with the function: read the CSV with DictReader, then save the rows as JSON. The version below also answers the second question, because it never writes straight onto the destination.

python
import os

def csv_to_json(src, dst):
    with open(src, newline='', encoding='utf-8') as f:
        rows = list(csv.DictReader(f))
    staging = Path(str(dst) + '.tmp')
    with open(staging, 'w', encoding='utf-8') as f:
        json.dump(rows, f, indent=2, ensure_ascii=False)
    os.replace(staging, dst)
    return len(rows)

n = csv_to_json(people, tmp / 'safe.json')
print(n)
print(json.loads((tmp / 'safe.json').read_text(encoding='utf-8'))[0])
output
2
{'name': 'Asha', 'age': '31'}
How do you safely overwrite a file?
  • Never open the real file with w first, because that empties it before you know the new content is good.
  • Write the complete new content to a temporary file in the same folder.
  • Swap it in with os.replace, which replaces the destination in one step, so readers see either the old file or the new one.
  • If anything fails while writing, the original is untouched and you can delete the temporary file.
Safe overwrite
  1. 1Write new contentto dst.tmp, same folder
  2. 2Close the fileleave the with block
  3. 3os.replacetmp becomes dst

The third question is about diagnosis. When a read fails, work through five checks in order, and most failures show up in the first two.

FailureUsually meansCheck
FileNotFoundErrorWrong path or working folder1, 2
IsADirectoryError, PermissionErrorFolder given, or access denied2, 3
UnicodeDecodeErrorFile is not in the encoding you named4
JSONDecodeError, odd CSV rowsFile is empty or damaged5
Keep in mind

Use with, name the encoding, open CSV with newline='', catch the specific error, stream big files, and replace files through a temporary copy.

Part 14 · Check yourself

Quiz

Work out each answer before you open it. These questions ask you to predict what code does or to find the bug, not to recite definitions.

What does this program print, and why is the second value what it is?
  • It prints 4 ''.
  • The first f.read() returns the whole text 'a\nb\n', which is 4 characters.
  • Reading moves the file position to the end, so the second f.read() returns an empty string.
  • f.seek(0) before the second read, or reopening the file, would give the text again.
from pathlib import Path
p = Path('t.txt')
p.write_text('a\nb\n', encoding='utf-8')
with open(p, encoding='utf-8') as f:
    first = f.read()
    second = f.read()
print(len(first), repr(second))
A script is meant to add one line to the end of log.txt each run, but the log only ever holds the last line. Find the bug.
  • The file is opened with 'w', which truncates it to empty the moment it opens, before any write happens.
  • Use 'a' to keep the old content and write at the end.
  • If you want to refuse to overwrite an existing file, use 'x' instead.
with open('log.txt', 'w', encoding='utf-8') as f:
    f.write('run finished\n')
The file orders.csv has the rows apple,10 and pear,5. What does this print, and what should the code do instead?
  • It prints 105, not 15.
  • Every value from the csv module is a str, so + joins the two strings.
  • Convert first: int(rows[0][1]) + int(rows[1][1]) gives 15.
  • Open the file with newline='' and encoding='utf-8' as well.
import csv
with open('orders.csv', newline='', encoding='utf-8') as f:
    rows = list(csv.reader(f))
print(rows[0][1] + rows[1][1])
What does this print? Name the two things JSON changed on the round trip.
  • It prints {'1': [1, 2]}.
  • The integer key 1 became the string '1', because JSON object keys are always strings.
  • The tuple (1, 2) became the list [1, 2], because JSON has only arrays.
  • If you need the original types back, convert them yourself after loading.
import json
data = {1: (1, 2)}
back = json.loads(json.dumps(data))
print(back)
A teammate guards a read with if p.exists(): and says the code can no longer raise FileNotFoundError. Are they right?
  • No. The file can be deleted between the check and the open(), which is a race condition.
  • The check also does not cover PermissionError or a path that is a directory.
  • Prefer EAFP: wrap the open() in try and catch FileNotFoundError, the narrowest exception you can handle.
  • except OSError catches the whole family when you want one handler for all of them.

Summary

  • Open files with with open(p, encoding='utf-8') so they close even when an error is raised, and so the text decodes the same on every machine.
  • Mode 'w' wipes the file the moment it opens; use 'a' to add, 'x' to refuse to overwrite.
  • Reading moves the position forward, so a second read() returns '' until you seek(0); iterate with for line in f to stream big files.
  • write() takes only str and adds no newline; use pathlib (Path(a) / b, read_text, write_text) to build and handle paths.
  • CSV needs newline='' and gives back only strings; JSON turns int keys into strings and tuples into lists.
  • Catch the narrowest exception (FileNotFoundError, or OSError for the family) instead of checking first or using a bare except.