Handbooks / Python asyncio / Chapter 2
Tasks, Cancellation & Timeouts
41 pages · ~63 min✓ Reviewed
Builds on Event Loop & Coroutines. Next up: Async I/O Patterns.
Part 1 · Running Concurrent Work Safely with asyncio Tasks
Running Concurrent Work Safely with asyncio Tasks
Starting concurrent work in asyncio is easy. Keeping it under control is the hard part. A task you forget about can be garbage-collected halfway through its job. An error in one branch can leave its siblings running with nobody listening. A cancel request that you swallow can stall a whole shutdown, and a call with no deadline can hang a service forever. None of these failures announces itself with a loud crash. They show up as missing emails, half-written records and programs that will not exit.
This chapter walks through the tools Python gives you to prevent that. You start with how tasks differ from plain coroutines, then learn to keep references with create_task, fan out with gather, and move to TaskGroup for fail-fast scopes. From there you read the ExceptionGroup errors a group raises using except*, stop work cleanly with cancel() and CancelledError, protect critical writes with shield, and bound waiting with timeout() and wait_for(). The last pages tie these together into structured concurrency: every task has an owner, and nothing outlives its block.
When you finish, you can decide between gather and TaskGroup for a given job, and you can handle grouped errors without losing any of them. You can stop a task and still run its cleanup, and you can put a deadline on a single call or on a whole multi-step block. You can also build a bounded worker pool on a queue that shuts down cleanly on Ctrl+C. A cheatsheet at the end collects the key calls and the Python version each one needs.
You need Python 3.11 or newer, because TaskGroup, asyncio.timeout() and except* all arrived in 3.11. Only the standard library is used, so there is nothing to install. You should already be comfortable with async def, await and asyncio.run(main()). Run each example as a script file, and keep one open so you can change the sleep times and watch how the behavior shifts.
Part 2 · Coroutines, Tasks and the Event Loop
Calling a coroutine is not running it
A function defined with async def is a coroutine function. Calling it does not execute its body. It hands you back a coroutine object, which is a paused piece of work. None of the body runs until that object is awaited or scheduled on the event loop.
The example below makes the gap visible. The print inside fetch stays silent when we create the object, and only fires once we await it.
import asyncio async def fetch(): print("running") return 42 async def main(): coro = fetch() # nothing printed yet print(type(coro).__name__) result = await coro # body runs now print(result) asyncio.run(main())
calling is not running
coroutine
running
42Awaiting a coroutine directly runs it inline: the caller waits right there until the coroutine finishes, then moves on. So await a(); await b() runs a to completion and only then starts b. Nothing overlaps, and you get no concurrency from plain awaits, no matter how many you write in a row.
await means wait for a result right here. A Task, covered next, means run it alongside me.
Tasks: running alongside the caller
A Task wraps a coroutine and schedules it on the event loop. asyncio.create_task(coro) returns the Task immediately, and the wrapped coroutine then runs concurrently with the code that created it. The caller can carry on, and collect the result later by awaiting the Task.
Task is a subclass of Future, so it inherits the whole Future API. This is how you inspect, wait on or stop a piece of background work.
| Method | What it gives you |
|---|---|
.done() | Has the task finished yet, whether by result, error or cancel? |
.result() | The return value, or it re-raises the task's error |
.exception() | The error the task raised, or None if it succeeded |
.cancel() / .cancelled() | Request cancellation / check whether it ended cancelled |
.add_done_callback(fn) | Call fn(task) once the task is done |
import asyncio async def square(n): await asyncio.sleep(0.1) return n * n async def main(): t = asyncio.create_task(square(7), name="sq") t.add_done_callback( lambda task: print("callback:", task.get_name(), task.result())) print(t.done()) await t print(t.done(), t.result(), t.exception(), t.cancelled()) print(isinstance(t, asyncio.Future)) asyncio.run(main())
the Future API on a Task
False callback: sq 49 True 49 None False True
Right after create_task, .done() is still False because the task has only been scheduled. Once await t returns, the result, the absent exception and the not-cancelled flag are all available.
The cooperative event loop
The event loop is single-threaded and cooperative. Only one task runs at any moment, and a running task keeps the loop until it hands control back. The only place that happens is an await that actually suspends, such as waiting on a sleep, a socket or another task.
The flip side is that a blocking call freezes every task on the loop. time.sleep() or a long CPU-bound loop never gives control back, so no other task can progress until it returns. Use the awaitable alternative, or push the blocking work to a thread.
| Blocks the loop | Use instead |
|---|---|
time.sleep(1) | await asyncio.sleep(1) |
| CPU-heavy loop | await asyncio.to_thread(f) |
| Blocking I/O call | await asyncio.to_thread(f) |
The ticker below should print every 0.1 seconds. In the first run the main task calls time.sleep(0.5) and the ticker cannot move. In the second run the same sleep is offloaded with asyncio.to_thread, so the loop stays free and the ticker keeps its rhythm.
import asyncio, time async def ticker(start): for i in range(3): print(f"tick {i} at {time.perf_counter() - start:.1f}s") await asyncio.sleep(0.1) async def run(label, work): print(label) start = time.perf_counter() t = asyncio.create_task(ticker(start)) await asyncio.sleep(0) # let the ticker start await work() await t async def blocking(): time.sleep(0.5) async def offloaded(): await asyncio.to_thread(time.sleep, 0.5) asyncio.run(run("time.sleep:", blocking)) asyncio.run(run("to_thread:", offloaded))
a blocking call stalls the ticker
time.sleep: tick 0 at 0.0s tick 1 at 0.5s tick 2 at 0.6s to_thread: tick 0 at 0.0s tick 1 at 0.1s tick 2 at 0.2s
asyncio.run(main()) is the normal entry point, and it owns the loop's whole lifetime. It follows four steps.
- 1Create the loopa fresh event loop
- 2Run main()until it completes
- 3Cancel leftoverstasks still pending
- 4Close the loopnothing runs after this
Two helpers let you look at what is running. asyncio.current_task() returns the Task you are currently inside, and asyncio.all_tasks() returns the set of tasks that are not done yet. Below, the long-sleeping worker is never awaited, so asyncio.run cancels it on the way out instead of waiting ten seconds.
import asyncio async def worker(): await asyncio.sleep(10) async def main(): w = asyncio.create_task(worker(), name="worker") await asyncio.sleep(0) print(asyncio.current_task().get_name()) print(sorted(t.get_name() for t in asyncio.all_tasks())) asyncio.run(main()) print("loop closed")
introspection, then cleanup by asyncio.run
Task-1 ['Task-1', 'worker'] loop closed
Two sleeps: 2 seconds or 1 second
The clearest proof that plain awaits are sequential and Tasks overlap is timing. Both functions below perform two one-second sleeps. The first awaits them one after the other. The second wraps each in a Task, so both timers run during the same second, and awaiting the Tasks just collects them.
import asyncio, time async def main_sequential(): await asyncio.sleep(1) await asyncio.sleep(1) async def main_tasks(): t1 = asyncio.create_task(asyncio.sleep(1)) t2 = asyncio.create_task(asyncio.sleep(1)) await t1 await t2 for fn in (main_sequential, main_tasks): start = time.perf_counter() asyncio.run(fn()) print(f"{fn.__name__}: {time.perf_counter() - start:.0f}s")
same work, different shape
main_sequential: 2s main_tasks: 1s
| Style | Total time |
|---|---|
| Two plain awaits | 2s |
| Two Tasks | 1s |
Writing fetch() on its own line only creates a coroutine object. Python warns RuntimeWarning: coroutine 'fetch' was never awaited, and the work never runs at all. Either await fetch() or wrap it with asyncio.create_task(fetch()).
await a(); await b() runs a fully before b starts. If you want them to overlap, turn them into Tasks first.
time.sleep() or a heavy CPU loop inside a coroutine stalls every other task. Swap in await asyncio.sleep() or await asyncio.to_thread(...).
A coroutine does nothing until awaited or scheduled. A Task is a Future that runs it alongside you, and the single-threaded loop only switches tasks at an await that really suspends.
Part 3 · create_task and Keeping References
What create_task Does
asyncio.create_task(coro, *, name=None, context=None) takes a coroutine object, schedules it on the running event loop and hands you back a Task straight away. It does not wait for the coroutine to run, let alone finish. The Task is your handle: you can await it, ask done(), call cancel() or read result() later.
The important detail is when the body begins. The loop is cooperative and single-threaded, so nothing else can run while your current coroutine is executing. The new task only gets its first turn when the caller reaches an await that really suspends. The task starts at the caller's next await point, not inside the create_task call.
- 1create_task(job())Task made and queued
- 2Returns the Taskcaller keeps running
- 3Caller awaitsit suspends here
- 4job() startsloop picks the queued task
The example below makes the ordering visible. Right after create_task the task is not done and job has printed nothing. await asyncio.sleep(0) suspends main for one loop iteration, which lets job run to completion.
import asyncio async def job(): print('job: running') async def main(): t = asyncio.create_task(job()) print('main: task created, done =', t.done()) await asyncio.sleep(0) print('main: after the await, done =', t.done()) asyncio.run(main())
The body of job runs only when main suspends
main: task created, done = False job: running main: after the await, done = True
Assuming the task has started, or even finished its first step, on the line after create_task. If you need to know something happened, await the task or use a signal such as an asyncio.Event.
Weak References and Fire-and-Forget
Creating a task does not by itself keep it alive. The event loop stores only weak references to tasks, so a task that nothing else points to can be garbage-collected in the middle of its run. A coroutine waiting on a timer or a socket is held only by that Task object, so when the Task goes the work goes with it.
| Holder | Reference | Does the task survive? |
|---|---|---|
| Event loop | weak | No, it does not count |
| Local variable | strong | Only while the variable is in scope |
| Set with discard callback | strong | Until the task is done |
| TaskGroup | strong | Yes, until the group exits |
This produces the classic fire-and-forget bug. Writing asyncio.create_task(send_email()) and discarding the result looks harmless, and it often works in small tests. Under memory pressure a collection can run while the email is still in flight, and the task vanishes silently: no result, no error, no email.
asyncio.create_task(send_email()) as a bare statement. The return value is the only strong reference you will ever get, so throwing it away is what makes the task collectable.
The fix is to hold your own strong reference in a set and let each task remove itself when it finishes. Adding it to the set keeps it alive; add_done_callback(background.discard) stops the set from growing forever.
import asyncio background = set() def spawn(coro): t = asyncio.create_task(coro) background.add(t) t.add_done_callback(background.discard) return t async def send_email(): await asyncio.sleep(0.1) print('email sent') async def main(): spawn(send_email()) print('pending:', len(background)) await asyncio.sleep(0.2) print('pending:', len(background)) asyncio.run(main())
The set keeps the task alive and cleans up after it
pending: 1 email sent pending: 0
Dropping the reference also hides failures. If a task raises and nobody ever awaits it or reads exception(), nothing is reported at that moment. The only trace is a Task exception was never retrieved message that the event loop logs when the Task object is finally garbage-collected, which may be long after the failure or never in a short-lived program. Awaiting your tasks, or inspecting them in a done callback, is what makes errors visible on time.
Names and Context
Tasks get automatic names like Task-17, which tell you nothing when you dump asyncio.all_tasks() during an incident. Pass name= when you create the task, or call task.set_name() afterwards, and debug output, logs and dumps become readable. The main task started by asyncio.run is called Task-1.
import asyncio async def worker(): await asyncio.sleep(0.1) async def main(): a = asyncio.create_task(worker(), name='mailer') b = asyncio.create_task(worker()) b.set_name('indexer') for t in sorted(asyncio.all_tasks(), key=lambda t: t.get_name()): print(t.get_name()) await asyncio.gather(a, b) asyncio.run(main())
Both naming styles show up in all_tasks()
Task-1
indexer
mailerThe context= argument, added in Python 3.11, controls which contextvars.Context the task runs in. If you leave it out, the task receives a copy of the current context. It sees the values set so far, but anything it sets stays private to the task. Passing your own Context lets you seed a task with chosen values, and you can inspect that context afterwards.
import asyncio, contextvars request_id = contextvars.ContextVar('request_id', default='none') async def show(label): print(label, request_id.get()) request_id.set('changed-in-task') async def main(): request_id.set('r-1') await asyncio.create_task(show('default copy:')) print('main still sees:', request_id.get()) ctx = contextvars.copy_context() ctx.run(request_id.set, 'r-2') await asyncio.create_task(show('explicit ctx:'), context=ctx) print('main still sees:', request_id.get()) print('ctx sees:', ctx[request_id]) asyncio.run(main())
Default is a copy; an explicit Context is used directly
default copy: r-1 main still sees: r-1 explicit ctx: r-2 main still sees: r-1 ctx sees: changed-in-task
A name such as mailer or db-poll costs nothing and turns a wall of Task-N entries in a dump into a list you can act on.
Eager Tasks and TaskGroup
Python 3.12 added asyncio.eager_task_factory. You install it with loop.set_task_factory(asyncio.eager_task_factory) on the running loop. With it, a new task does not wait for the next loop iteration: it runs synchronously inside create_task until its first real suspension. Only then does create_task return the Task. A task that finishes without ever suspending, such as a cache hit, completes before it is ever scheduled, which saves a loop iteration per task.
| Step | Default factory | eager_task_factory (3.12+) |
|---|---|---|
| create_task is called | Task queued, body not started | Body runs right away |
| Body reaches its first suspension | Not yet reached | create_task returns the Task |
| Body starts running | At the caller's next await | Already started |
| Cost for a task that never suspends | One extra loop iteration | None, it is done on return |
Writing code that depends on the task not having started yet. Under the eager factory the first part of the body has already run when create_task returns, so anything it prints or mutates happens earlier than the default order suggests.
For most code the best answer to all the lifetime questions is not a hand-made set but a TaskGroup (3.11+). TaskGroup.create_task keeps a strong reference to every child, and leaving the async with block awaits them all for you, so no task is lost and no error goes unretrieved.
import asyncio async def job(n): await asyncio.sleep(0.05 * n) return n * 10 async def main(): async with asyncio.TaskGroup() as tg: a = tg.create_task(job(1), name='a') b = tg.create_task(job(2), name='b') print(a.result(), b.result()) asyncio.run(main())
The group holds both tasks and awaits them on exit
10 20
| Bare create_task | TaskGroup.create_task | |
|---|---|---|
| Strong reference | You must keep one | Held by the group |
| Awaiting the task | Your job | Done on block exit |
| Child error | Easy to lose | Raised from the block |
| Task outlives the scope | Yes, possible | No |
Prefer TaskGroup.create_task. Use bare create_task only for true background work, and then always store the task in a set with a discard callback and give it a name.
Part 4 · asyncio.gather
Running awaitables together
await asyncio.gather(*aws, return_exceptions=False) takes any number of awaitables, runs them concurrently, and hands you back one list of results. The call suspends the awaiting coroutine until every child has finished. Total time is therefore roughly that of the slowest child, not the sum of all of them.
You can pass bare coroutines straight in. gather wraps each one in a Task for you, so they all start on the event loop's next iteration. You don't call create_task yourself. Tasks and futures you pass in are used as they are.
The list is ordered by argument position, never by who finished first. Slot 0 always belongs to the first awaitable you passed, even if it was the slowest. The table below shows a dry run.
| Arg | Awaitable | Finishes | Slot in result |
|---|---|---|---|
| 0 | fetch(a) | 2.0 s | r[0] |
| 1 | fetch(b) | 0.5 s | r[1] |
| whole call | gather | 2.0 s | [ra, rb] |
Because the order is fixed, you can unpack the list directly: r1, r2 = await gather(fetch(a), fetch(b)) puts the first result in r1 and the second in r2. In the program below the second fetch finishes first, as the printed order shows, yet the unpacked results still line up with the arguments.
import asyncio async def fetch(name, delay): await asyncio.sleep(delay) print(f'finished {name}') return f'{name} done' async def main(): r1, r2 = await asyncio.gather(fetch('a', 0.2), fetch('b', 0.05)) print(r1) print(r2) asyncio.run(main())
b finishes first, but r1 is still the result of fetch('a')
finished b finished a a done b done
Expecting results in completion order. They are always in argument order, so never match results to inputs by when they arrived.
When a child raises
What gather does with a failure depends on return_exceptions. The default is False. The first exception to occur is re-raised to whoever awaits the gather. The surprising part is what happens to the other children. They are not cancelled. They keep running in the background, as though nothing had happened.
| Condition | What gather does |
|---|---|
| child raises, return_exceptions=False | first exception propagates to the awaiter; siblings keep running |
| child raises, return_exceptions=True | exception is placed in the result list; nothing is raised |
In this program boom() fails after 0.05 s while slow needs 0.2 s. Watch what slow is doing after the exception arrives.
import asyncio async def ok(name, delay): await asyncio.sleep(delay) print(f'{name} finished') return name async def boom(): await asyncio.sleep(0.05) raise ValueError('boom') async def main(): slow = asyncio.create_task(ok('slow', 0.2)) try: await asyncio.gather(boom(), slow) except ValueError as e: print('caught', e) print('slow cancelled?', slow.cancelled()) print('slow done?', slow.done()) await asyncio.sleep(0.3) print('slow done?', slow.done()) asyncio.run(main())
caught boom slow cancelled? False slow done? False slow finished slow done? True
The awaiter got its ValueError and moved on, but slow kept running and finished later. With return_exceptions=True nothing is raised at all. Each failure sits in the slot of the child that produced it, next to the successful values. That gives you partial results. You then have to check each slot, because a result may be an exception object instead of a value.
import asyncio async def one(): return 42 async def bad(): raise ValueError('x') async def two(): return 7 async def main(): res = await asyncio.gather(one(), bad(), two(), return_exceptions=True) print(res) for i, r in enumerate(res): if isinstance(r, BaseException): print(i, 'failed', repr(r)) else: print(i, 'ok', r) asyncio.run(main())
check isinstance(r, BaseException) before using a slot
[42, ValueError('x'), 7] 0 ok 42 1 failed ValueError('x') 2 ok 7
Using the results list without checking for exception objects. With return_exceptions=True a slot may hold an error where you expect a value.
Cancellation and gather
Cancellation passes through gather in two different directions, and the two are easy to mix up. If the gather itself is cancelled, it cancels every child that has not finished yet. If only one child is cancelled, gather treats that as the child raising CancelledError. The gather is not cancelled by this, and the other children carry on.
| Event | Effect |
|---|---|
| gather cancelled | unfinished children are cancelled |
| one child cancelled | treated as CancelledError; gather is not cancelled |
The program has two parts. First, it cancels one child while using return_exceptions=True, so the cancelled child shows up as a CancelledError in its slot while its sibling finishes normally. Then it cancels a gather and checks that both children were cancelled with it.
import asyncio async def work(n): await asyncio.sleep(n) return n async def main(): t1 = asyncio.create_task(work(0.1)) t2 = asyncio.create_task(work(0.2)) g = asyncio.gather(t1, t2, return_exceptions=True) await asyncio.sleep(0.05) t1.cancel() res = await g print([type(r).__name__ if isinstance(r, BaseException) else r for r in res]) print('gather cancelled:', g.cancelled()) a = asyncio.create_task(work(0.5)) b = asyncio.create_task(work(0.5)) g2 = asyncio.gather(a, b) await asyncio.sleep(0.05) g2.cancel() try: await g2 except asyncio.CancelledError: print('outer raised CancelledError') print(a.cancelled(), b.cancelled()) asyncio.run(main())
['CancelledError', 0.2] gather cancelled: False outer raised CancelledError True True
Thinking a cancelled child cancels the gather. It does not. The gather keeps waiting for the rest and reports the child as CancelledError.
Orphans and when gather is still right
Put the failure rule and the cancellation rule together and you get a hazard. After the first exception, the surviving siblings become orphans. They are still running, but nothing awaits them any more. If one of them fails later, that error is lost. The awaiter has already moved on, and the gather does not report it. The only way to see it is to have kept a handle on the task yourself.
import asyncio async def quick_fail(): await asyncio.sleep(0.05) raise KeyError('first') async def late_fail(): await asyncio.sleep(0.1) raise RuntimeError('late') async def main(): sibling = asyncio.create_task(late_fail()) try: await asyncio.gather(quick_fail(), sibling) except KeyError as e: print('caught', repr(e)) await asyncio.sleep(0.2) print('sibling done:', sibling.done()) print('its error:', repr(sibling.exception())) asyncio.run(main())
only the stored task handle lets us see the second error
caught KeyError('first') sibling done: True its error: RuntimeError('late')
Assuming one failure cancels the siblings. Ignoring orphans after the first exception. When you need fail-fast with sibling cancellation, use TaskGroup (next section).
gather is still a good fit for simple fan-out where you want partial results. Call gather(..., return_exceptions=True), wait for every child to finish, then inspect each slot. Nothing is orphaned because you wait for all of them, and no error is lost because each one lands in the list.
Best-effort batches: gather(..., return_exceptions=True), then check each slot for an exception object before using it.
Part 5 · TaskGroup (Python 3.11+)
A scope that owns its tasks
asyncio.TaskGroup turns a block of code into a task scope. You open it with async with, start children through tg.create_task(coro()), and every child is bound to that block. Each call returns an ordinary Task, and that task object is your handle for reading the outcome later.
The key rule is at the exit. When execution reaches the end of the async with body, the group waits for every child to finish. No task started through the group can outlive the block, so you can't lose a task or leave an error unread.
The example below starts two children and prints from the body. The body has no await, so the children haven't started yet and both report done() == False. The children finish while the block exit waits for them. Only then do we call .result() on each handle.
import asyncio async def square(n, delay): await asyncio.sleep(delay) print("finished", n) return n * n async def main(): async with asyncio.TaskGroup() as tg: a = tg.create_task(square(2, 0.02)) b = tg.create_task(square(3, 0.01)) print("body done, children running:", not a.done()) print("after block:", a.done(), b.done()) print(a.result(), b.result()) asyncio.run(main())
The block exit waits; results are read afterwards
body done, children running: True finished 3 finished 2 after block: True True 4 9
Results come from the handles, and only after the block. Each task keeps its own return value, and t.result() hands it back. Inside the block a child may still be running, and calling result() on an unfinished task raises InvalidStateError.
Calling t.result() inside the async with body, before the child is done. Wait for the block to end, then read the handles.
When something goes wrong
A TaskGroup is fail-fast. If a child raises anything other than CancelledError, the group cancels all the remaining children and also cancels the body that is still running. The exit then waits for those cancellations to finish, so every finally in the cancelled tasks runs before you see the error.
The group does not report only the first failure. It collects every child failure and raises them together as one ExceptionGroup. If any of the collected errors is a BaseException rather than an Exception, you get a BaseExceptionGroup instead. Siblings cancelled by the group are not failures and don't appear in it.
Two children below fail in the same loop step while a third is sleeping. The group cancels the sleeper, which prints from its finally. Both errors arrive together, and we split them with except*. Exception groups and except* get their own section, so here we only use them.
import asyncio async def slow(): try: await asyncio.sleep(10) finally: print("slow cleaned up") async def bad_key(): raise KeyError("k") async def bad_value(): raise ValueError("v") async def main(): try: async with asyncio.TaskGroup() as tg: s = tg.create_task(slow()) tg.create_task(bad_key()) tg.create_task(bad_value()) except* KeyError as eg: print("key:", eg.exceptions) except* ValueError as eg: print("value:", eg.exceptions) print("slow cancelled:", s.cancelled()) asyncio.run(main())
One sleeper, two failures, both reported
slow cleaned up key: (KeyError('k'),) value: (ValueError('v'),) slow cancelled: True
The failure can also start in your own code. An exception raised in the async with body cancels the children just as a failing child would, and it is wrapped in the same kind of group. Once the block has ended, the group is finished, so asking it for another task raises RuntimeError.
import asyncio async def slow(): try: await asyncio.sleep(10) finally: print("slow cleaned up") async def main(): try: async with asyncio.TaskGroup() as tg: w = tg.create_task(slow()) await asyncio.sleep(0) raise RuntimeError("body failed") except* RuntimeError as eg: print("caught", eg.exceptions) print("worker cancelled:", w.cancelled()) late = slow() try: tg.create_task(late) except RuntimeError as e: late.close() print("refused:", type(e).__name__) asyncio.run(main())
A body error cancels children; a finished group refuses new tasks
slow cleaned up caught (RuntimeError('body failed'),) worker cancelled: True refused: RuntimeError
| Event | What the group does |
|---|---|
| A child raises an Exception | Cancels the other children and the body, then raises an ExceptionGroup |
| Several children fail | Collects all of them into one group |
| A BaseException is among the errors | Raises a BaseExceptionGroup instead |
| The body raises | Cancels every child, then raises the error inside a group |
| KeyboardInterrupt or SystemExit in a child | Re-raised directly, not wrapped in any group |
| tg.create_task after the group finished | Raises RuntimeError |
Writing a plain except KeyError: around the block misses the failure, because it always arrives wrapped in a group. Use except*. Also don't swallow CancelledError inside a child: the group's exit then waits for a task that ignored its cancellation.
TaskGroup compared with gather
asyncio.gather and TaskGroup both run work concurrently, but they behave differently when something fails. With gather, the first exception propagates to the awaiter while the sibling children keep running as orphans, and their later errors are lost. A TaskGroup cancels the siblings and reports every error.
| On failure | gather | TaskGroup |
|---|---|---|
| Siblings | Keep running | Cancelled by the group |
| Errors reported | Only the first one reaches you | Every error, in an ExceptionGroup |
| Leftover work | Orphan tasks may outlive the call | Nothing outlives the block |
| Partial results | Possible with return_exceptions=True | Not offered: all must succeed |
The two also differ in how you get results back. gather returns an ordered list, with one slot per argument. A TaskGroup gives you the Task handles, and you collect the values yourself with t.result() after the block.
| Collecting results | gather | TaskGroup |
|---|---|---|
| What you get back | A list in argument order | The Task handle from each create_task |
| How you read a value | Unpack the awaited list | Call t.result() after the block |
| Where the order comes from | Argument position | Your own variables or collection |
| When values are available | When the whole call returns | Once the block has exited |
Use a TaskGroup for all-or-nothing work. Use gather(return_exceptions=True) when a best-effort batch with partial results is what you want.
Part 6 · Exception Groups and except*
What an ExceptionGroup Is
A plain exception describes one failure. Concurrent code breaks that assumption, because several tasks can fail at the same moment and each failure is unrelated to the others. ExceptionGroup (PEP 654, Python 3.11) solves this by carrying many exceptions inside a single raised object. You build one with a message and a list: ExceptionGroup(msg, [excs]). The members are stored in the .exceptions tuple, and each member is called a leaf (a member can itself be a group, as you will see later).
There are two classes. BaseExceptionGroup is the base and may hold any BaseException, including KeyboardInterrupt. ExceptionGroup also inherits from Exception, so it can only hold Exception subclasses. Trying to put a BaseException inside an ExceptionGroup raises TypeError, while BaseExceptionGroup quietly builds the right kind of group for what you give it.
eg = ExceptionGroup('batch', [ValueError('bad id'), TypeError('not str')]) print(eg) print(eg.exceptions) print(isinstance(eg, Exception)) try: ExceptionGroup('x', [KeyboardInterrupt()]) except TypeError as e: print('TypeError:', e) beg = BaseExceptionGroup('mix', [ValueError('a'), KeyboardInterrupt()]) print(type(beg).__name__)
batch (2 sub-exceptions) (ValueError('bad id'), TypeError('not str')) True TypeError: Cannot nest BaseExceptions in an ExceptionGroup BaseExceptionGroup
| BaseExceptionGroup | ExceptionGroup | |
|---|---|---|
| Role | The base class of all groups | Subclass that is also an Exception |
| May hold | Any BaseException | Exception subclasses only |
Caught by except Exception | Only if all leaves are Exceptions | Yes |
| Bad leaf | Not applicable | TypeError at construction |
Building an ExceptionGroup with a BaseException such as KeyboardInterrupt or CancelledError raises TypeError. Use BaseExceptionGroup when a leaf may not be an Exception.
Matching Leaves with except*
The except* clause is written like except, but it works on the leaves of a group. except* ValueError as eg: pulls out every ValueError leaf and binds eg to a new ExceptionGroup holding just those leaves. The original group's message is kept. A bare ValueError (not in a group) is also matched, wrapped in a group of one, so the handler body can always treat eg as a group.
Unlike plain except, which runs at most one clause, **several except\* clauses can run for a single group**, one for each type that has matching leaves. Once all clauses have run, any leaves that nobody matched are re-raised automatically as a smaller group, so nothing is lost silently.
def run(): try: raise ExceptionGroup('batch', [ValueError('a'), ValueError('b'), TypeError('t'), KeyError('k')]) except* ValueError as eg: print('values:', len(eg.exceptions), type(eg).__name__) except* TypeError as eg: print('types:', len(eg.exceptions)) try: run() except ExceptionGroup as rest: print('left:', rest.exceptions, type(rest).__name__)
values: 2 ExceptionGroup types: 1 left: (KeyError('k'),) ExceptionGroup
| Aspect | except | except* |
|---|---|---|
| Matches | One exception | Leaves inside a group |
| Clauses run | At most one | One per matching type |
| Unmatched | Propagates as it is | Re-raised as a smaller group |
| break / continue / return | Allowed | SyntaxError |
| Mixed in one try | Not with except* | Not with except |
You cannot mix except and except* in the same try statement, and an except* body cannot use break, continue or return. Both are a SyntaxError, because the block must be able to finish and let the leftover group be re-raised.
Traps, Splitting and Nesting
The classic trap comes from TaskGroup. When a child task fails, the group never raises the child's exception directly. It always wraps it in an ExceptionGroup, even if only one task failed. A plain except ValueError therefore does not match, and the whole group escapes as if you had no handler at all.
import asyncio async def boom(): raise ValueError('bad') async def plain(): try: async with asyncio.TaskGroup() as tg: tg.create_task(boom()) except ValueError: print('caught') async def starred(): try: async with asyncio.TaskGroup() as tg: tg.create_task(boom()) except* ValueError as eg: print('caught', eg.exceptions) async def main(): try: await plain() except ExceptionGroup as eg: print('escaped:', type(eg).__name__, eg.exceptions) await starred() asyncio.run(main())
escaped: ExceptionGroup (ValueError('bad'),) caught (ValueError('bad'),)
Plain except ValueError does not catch an ExceptionGroup that contains a ValueError. Wrap async with TaskGroup() in try / except*, never a plain except.
You can also take a group apart without a try statement. eg.subgroup(pred) returns a new group keeping only the leaves that match, where pred is either an exception type or a function taking a leaf. eg.split(pred) returns a (match, rest) pair covering every leaf exactly once. Either half is None when it would be empty, so check before using it.
eg = ExceptionGroup('io', [ValueError('a'), TypeError('b'), ValueError('c')]) only = eg.subgroup(lambda e: isinstance(e, ValueError)) print(only.exceptions) match, rest = eg.split(ValueError) print(len(match.exceptions), len(rest.exceptions)) print(eg.subgroup(KeyError)) match, rest = eg.split(lambda e: True) print(rest)
(ValueError('a'), ValueError('c')) 2 1 None None
Groups can nest: a leaf may itself be an ExceptionGroup, for example when a TaskGroup runs inside another TaskGroup's task. When an uncaught group reaches the top, the traceback prints a tree, with each sub-exception on its own branch under a header such as ExceptionGroup: run (2). A branch is either a leaf exception or a nested group with its own branches. Matching and splitting walk through the nesting for you, so except* TimeoutError finds a TimeoutError at any depth. If you iterate .exceptions by hand, remember that an element may be a group and recurse into it.
Putting it together, a TaskGroup often needs a different reaction per failure type. Both clauses below can fire for one group, and anything unmatched is re-raised afterwards.
import asyncio async def slow(): raise TimeoutError('t') async def bad(): raise ValueError('v') async def main(): try: async with asyncio.TaskGroup() as tg: tg.create_task(slow()) tg.create_task(bad()) except* TimeoutError: print('retry') except* ValueError as eg: print('log', eg.exceptions) asyncio.run(main())
retry
log (ValueError('v'),)Groups carry many leaves, except* handles them by type, and whatever you do not handle is re-raised for the caller.
Part 7 · Cancellation and CancelledError
Asking a Task to Stop
A task cannot be killed from outside. What you can do is call task.cancel(msg=None), which only requests cancellation. The loop later throws a CancelledError into the coroutine at the point where it is currently suspended, so the exception surfaces at its next await. Until the task actually resumes and unwinds, it is still running.
- 1t.cancel()returns True, request recorded
- 2Next loop turntask is resumed
- 3CancelledErrorraised at the await it was parked on
- 4finally / except blockscleanup runs, error re-raised
- 5Task endscancelled() is True, awaiter sees the error
The return value of cancel() tells you whether the request was accepted, not whether the task has stopped. It returns False if the task is already done, because there is nothing left to cancel. Otherwise it returns True, and the cancellation is still pending. The example below prints done() right after cancel() to make that gap visible.
import asyncio async def worker(): await asyncio.sleep(10) async def main(): t = asyncio.create_task(worker()) await asyncio.sleep(0) # let worker start and park print(t.cancel()) # request accepted print(t.done()) # but nothing has stopped yet try: await t except asyncio.CancelledError: print('worker stopped') print(t.done()) print(t.cancel()) # already done: no-op asyncio.run(main())
cancel() is a request; done() flips only after the task unwinds
True False worker stopped True False
| Call | Returns | Meaning |
|---|---|---|
| cancel() on a running task | True | request made, task not stopped yet |
| cancel() on a finished task | False | no-op, nothing to cancel |
Reading cancel() returning True as "the task has already stopped". It only means the request was recorded. To know the task has ended, await it.
CancelledError Is Not an Ordinary Exception
Since Python 3.8, CancelledError subclasses BaseException rather than Exception. That is deliberate: a broad except Exception: written for ordinary failures does not accidentally swallow a cancellation. The example checks the class hierarchy and shows that the Exception handler never fires while the finally block still does.
import asyncio print(issubclass(asyncio.CancelledError, Exception)) print(issubclass(asyncio.CancelledError, BaseException)) async def stubborn(): try: await asyncio.sleep(10) except Exception: print('caught by Exception') finally: print('finally ran') async def main(): t = asyncio.create_task(stubborn()) await asyncio.sleep(0) t.cancel() try: await t except asyncio.CancelledError: print('cancelled:', t.cancelled()) asyncio.run(main())
False True finally ran cancelled: True
If the task holds a connection, a file or a lock, it needs to release it when cancelled. There are two correct shapes: put the release in try/finally, or catch CancelledError, clean up, and always re-raise. The second form is useful when you want to do something only on cancellation. The example uses it and also passes a message to cancel(), which travels on the exception as args.
import asyncio async def poll(): try: while True: await asyncio.sleep(1) except asyncio.CancelledError: print('cleanup') raise # always re-raise async def main(): t = asyncio.create_task(poll()) await asyncio.sleep(0) t.cancel('shutdown') try: await t except asyncio.CancelledError as e: print(e.args) print(t.cancelled()) asyncio.run(main())
the msg given to cancel() rides on the CancelledError
cleanup ('shutdown',) True
| Handler inside the task | Cancel works? |
|---|---|
| try: ... finally: cleanup() | yes |
| except CancelledError: cleanup(); raise | yes |
| except CancelledError: pass | no, swallowed |
| except Exception: ... | never reached, error is a BaseException |
Swallowing the error turns the cancellation into a no-op. The task carries on, or returns normally, as if nothing had been asked of it. That is worse than it sounds, because the machinery built on cancellation, namely TaskGroup, asyncio.timeout() and application shutdown, all rely on the task really stopping. A task that ignores the request makes the group stall, the timeout hang or never fire, and the program refuse to exit cleanly.
import asyncio async def swallower(): try: await asyncio.sleep(10) except asyncio.CancelledError: print('ignoring cancel') return 'done anyway' async def main(): t = asyncio.create_task(swallower()) await asyncio.sleep(0) t.cancel() result = await t print(result, t.cancelled()) asyncio.run(main())
the task finishes normally, so cancelled() is False
ignoring cancel
done anyway FalseWriting except CancelledError: pass inside the task itself. The task ignores the cancel, and TaskGroup, timeout() and shutdown all end up waiting on work that will not stop. Catch it only to tidy up, then raise.
Seeing and Stopping a Cancelled Task
task.cancelled() is True only when the task actually ended by raising CancelledError. A request that is still pending does not count, and neither does a task that swallowed the error and returned a value, as the previous page showed. Awaiting a cancelled task also propagates the error: the awaiter receives CancelledError, exactly as if it had awaited any task that raised.
| Situation | cancelled() | await t |
|---|---|---|
| cancel() called, task not yet resumed | False | still waiting |
| task unwound with CancelledError | True | raises CancelledError in the awaiter |
| task swallowed the cancel and returned | False | returns the value |
| task finished before cancel() | False | returns the value |
This gives the standard way to stop a background task: cancel it, then await it. The await matters, because it waits until the task has finished its cleanup. Catching CancelledError around that await is the one place where swallowing is correct, since you, the owner, asked for the cancellation and expect it.
import asyncio async def ticker(): try: while True: await asyncio.sleep(0.01) finally: print('ticker closed') async def main(): t = asyncio.create_task(ticker()) await asyncio.sleep(0.05) t.cancel() try: await t # wait for the cleanup to finish except asyncio.CancelledError: pass # expected: we asked for it print('stopped:', t.cancelled()) asyncio.run(main())
cancel, then await: the standard way to stop a background task
ticker closed
stopped: TrueThe same idea scales to a list of tasks. Cancel each one, then use gather(*tasks, return_exceptions=True) so that every task is waited for and the cancellation errors are collected instead of raised.
Cancel, then await. Clean up in finally, never swallow the error inside the task, and treat except CancelledError: pass as something only the code that requested the cancel may write.
cancelling() and uncancel()
Python 3.11 added a counter to every task. task.cancelling() returns the number of cancel requests that are still pending, and task.uncancel() decrements it by one. Each cancel() call adds one, so two separate callers asking for cancellation give a count of 2.
import asyncio async def idle(): await asyncio.sleep(10) async def main(): t = asyncio.create_task(idle()) await asyncio.sleep(0) print(t.cancelling()) t.cancel() t.cancel() print(t.cancelling()) t.uncancel() print(t.cancelling()) try: await t except asyncio.CancelledError: print('cancelled:', t.cancelled()) asyncio.run(main())
the count tracks requests; uncancel() withdraws one
0 2 1 cancelled: True
The counter exists so that TaskGroup and asyncio.timeout() can tell their own cancellation apart from one that came from outside. Both call cancel() on the current task when they fire, then call uncancel() when they handle the resulting CancelledError. If the count is back to its original value, the cancel was theirs, so timeout() converts it into TimeoutError. If the count is still above that value, someone else also asked for cancellation, and the CancelledError must keep propagating.
| Step | Event | cancelling() |
|---|---|---|
| 1 | timeout() deadline hits, it cancels the task | 1 |
| 2 | shutdown also calls t.cancel() | 2 |
| 3 | timeout exits and calls uncancel() | 1 |
| 4 | count is still above zero, so CancelledError propagates instead of TimeoutError | 1 |
This is another reason never to swallow CancelledError in your own code: the bookkeeping only works when the error reaches the TaskGroup or timeout() block that is waiting for it. You will rarely call uncancel() yourself.
Part 8 · asyncio.shield
What shield protects
Normally, cancelling a task cancels whatever it is awaiting at that moment. await asyncio.shield(aw) changes that for one await. The shield sits between the outer task and the inner awaitable, and it absorbs the cancellation request so the inner work is not told to stop.
The outer task gets no such protection. Its await still raises CancelledError, so the outer task is cancelled as usual. Only the inner task keeps running, and anyone who holds a reference to it can still collect its result.
- 1outer.cancel()request arrives
- 2await shield(t)raises CancelledError
- 3inner tkeeps running
- 4t.result()work finished
In the example below, handler awaits a shielded commit and is cancelled halfway through. The handler is cancelled, but the commit still finishes, and the result is available afterwards.
import asyncio async def commit(): await asyncio.sleep(0.2) print("commit finished") return "saved" async def handler(inner): try: await asyncio.shield(inner) except asyncio.CancelledError: print("handler cancelled") raise async def main(): inner = asyncio.create_task(commit()) outer = asyncio.create_task(handler(inner)) await asyncio.sleep(0.05) outer.cancel() try: await outer except asyncio.CancelledError: print("outer cancelled:", outer.cancelled()) print("inner done yet:", inner.done()) print("inner result:", await inner) asyncio.run(main())
handler cancelled outer cancelled: True inner done yet: False commit finished inner result: saved
Shield guards one direction only
The shield only intercepts cancellation that comes from the outside through the outer task. If someone calls inner.cancel() directly, the shield has nothing to intercept, because the inner task really is cancelled. The shield's outer future then reports the cancellation to whoever is awaiting it.
async def waiter(inner): try: return await asyncio.shield(inner) except asyncio.CancelledError: print("waiter saw CancelledError") raise async def main(): inner = asyncio.create_task(commit()) outer = asyncio.create_task(waiter(inner)) await asyncio.sleep(0.05) inner.cancel() try: await outer except asyncio.CancelledError: print("outer cancelled:", outer.cancelled()) print("inner cancelled:", inner.cancelled()) asyncio.run(main())
waiter saw CancelledError outer cancelled: True inner cancelled: True
| Who is cancelled | Outer await | Inner task |
|---|---|---|
| The outer task | raises CancelledError | keeps running |
| The inner task directly | raises CancelledError | is cancelled |
Keep a reference to the inner task
shield accepts a coroutine as well as a task or future. When you pass a bare coroutine, as in shield(commit()), asyncio wraps it in a new Task for you. Nothing in your code holds that task, so you cannot await it or inspect it later.
That matters because the event loop keeps only weak references to tasks. If nothing else holds a strong reference, the garbage collector may destroy the inner task in the middle of the write you were trying to protect. So create the task yourself, keep it in a variable, and shield that.
| Form | Who holds the inner task | Safe? |
|---|---|---|
shield(commit()) | only the loop, weakly | no, it can be collected |
t = create_task(commit()) then shield(t) | your variable | yes, while t is in scope |
t stored in a set until it is done | the set | yes, even after the function returns |
await asyncio.shield(commit()) hides the inner task from you. You cannot wait for it after a cancel, and it may be garbage-collected mid-run. Always create the task first.
When it is worth using
The classic use case is a write that must not be torn half-way: a database commit, or a payment where the charge and the ledger record have to land together. If a client disconnects or a timeout fires in the middle of such a write, you want the outer request to be cancelled, but you do not want the write itself to be cut off between its two steps.
The commit pattern
Shielding on its own leaves a gap. When the outer task is cancelled, it stops waiting and moves on, and nobody checks that the commit actually completed. The standard pattern closes that gap. In the except block you wait for the inner task to finish, and only then do you re-raise the cancellation.
async def save(task): try: await asyncio.shield(task) except asyncio.CancelledError: print("cancel arrived, waiting for commit") await task print("commit landed") raise async def main(): t = asyncio.create_task(commit()) s = asyncio.create_task(save(t)) await asyncio.sleep(0.05) s.cancel() try: await s except asyncio.CancelledError: print("save cancelled:", s.cancelled()) print(t.result()) asyncio.run(main())
cancel arrived, waiting for commit commit finished commit landed save cancelled: True saved
The final raise is not optional. The caller asked for cancellation, and swallowing it would break timeouts, task groups and shutdown, which all rely on the task really ending as cancelled.
A second cancel can still tear the commit
Look at what the handler awaits. Inside the except block it writes await task, and that is the inner task awaited directly, with no shield. The shield protected only the first await. If a second cancel hits save() while it sits in await task, the cancellation goes straight into t, and the commit can be torn after all.
There are two honest options. You can accept that a second cancel means "abort now", which is often a reasonable policy. Or you can keep the wait protected by shielding again in a loop, which absorbs every extra cancel until the inner task is done.
async def save(task): cancelled = False while not task.done(): try: await asyncio.shield(task) except asyncio.CancelledError: cancelled = True print("cancel arrived, still waiting") if cancelled: raise asyncio.CancelledError return task.result() async def main(): t = asyncio.create_task(commit()) s = asyncio.create_task(save(t)) await asyncio.sleep(0.05) s.cancel() await asyncio.sleep(0.05) s.cancel() try: await s except asyncio.CancelledError: print("save cancelled:", s.cancelled()) print(t.result()) asyncio.run(main())
cancel arrived, still waiting
cancel arrived, still waiting
commit finished
save cancelled: True
savedBoth cancels were absorbed, the commit finished, and only then did save() end as cancelled. The loop is bounded because it exits as soon as the inner task is done. Do not use this shape for work that can run forever.
shield vs finally, and its limits
Shield and try/finally both come up around cancellation, but they do opposite things. finally lets the cancellation arrive and then gives you a place to tidy up. shield stops the cancellation from reaching the inner work at all.
| shield | try/finally | |
|---|---|---|
| Cancellation | hidden from the inner work | delivered to the work |
| Inner work | keeps running to completion | stops at its current await |
| Cleanup | not needed, the work finishes itself | runs after the cancel |
| Use for | must-finish writes | releasing connections, files, locks |
If every await is shielded, tasks stop responding to cancellation. They become zombies that block shutdown. A task that keeps re-shielding forever also keeps a TaskGroup from exiting, because the group waits for all its children to end. Reserve shield for short, bounded writes.
Shield only guards against cancelling the outer task. Calling cancel() on the inner task still cancels it, so keep that handle private to the code that owns the write.
Shield is not durability
When the program ends, asyncio.run cancels every task that is still pending, and that includes shielded inner tasks. The shield protects against a cancel arriving through one particular outer task, not against the loop being torn down. In the next example main returns while a slow write is still running, and the write is cancelled.
async def slow_write(): try: await asyncio.sleep(10) except asyncio.CancelledError: print("write cancelled at shutdown") raise async def hold(inner): await asyncio.shield(inner) async def main(): inner = asyncio.create_task(slow_write()) asyncio.create_task(hold(inner)) await asyncio.sleep(0.05) print("main returning") asyncio.run(main())
main returning write cancelled at shutdown
Shield keeps one write from being cut off by one cancel. A crash, a kill signal or loop shutdown will still stop it. For data that must survive, rely on a database transaction, not on shield.
Part 9 · Timeouts: asyncio.timeout vs wait_for
asyncio.timeout: one deadline for a whole block
Python 3.11 added async with asyncio.timeout(5):, a context manager that puts one deadline over everything inside the block. Every await in the body shares the same five seconds, so a fetch followed by a save is bounded as a unit. Passing None means no deadline. That is useful when you only learn the limit later, and we come back to it on the next page.
The mechanism has three steps. When the deadline passes, the event loop calls cancel() on the current task. The task sees a CancelledError at whatever await it was suspended on. When the block exits, the context manager's __aexit__ recognises that this cancellation was its own and converts the CancelledError into a TimeoutError.
- 1Deadline passesloop timer fires
- 2task.cancel()on the current task
- 3CancelledErrorraised at the suspended await, inside the block
- 4Block exitsaexit sees its own cancel
- 5TimeoutErrorraised to code outside the block
The consequence is the most important rule on this page. Inside the block you only ever see CancelledError, never TimeoutError, so an except TimeoutError placed inside the block can never match. Put the try around the async with. The example below shows both sides, and also confirms that asyncio.TimeoutError is the same object as the builtin TimeoutError since 3.11. Old code that catches asyncio.TimeoutError keeps working.
import asyncio async def slow(): try: await asyncio.sleep(10) except asyncio.CancelledError: print('inside: CancelledError') raise async def main(): try: async with asyncio.timeout(0.1): await slow() except TimeoutError: print('outside: TimeoutError') print(asyncio.TimeoutError is TimeoutError) asyncio.run(main())
The inner code sees the cancel; the outer code sees the timeout
inside: CancelledError
outside: TimeoutError
TrueInside the block: CancelledError. Outside the block: TimeoutError. Catch the timeout around the async with, and re-raise CancelledError if you ever catch it inside.
Absolute deadlines and moving the deadline
asyncio.timeout(delay) takes a number of seconds from now. When several steps must finish by one fixed moment, use asyncio.timeout_at(when) instead. It takes an absolute time on the event loop's clock, which you read with loop.time(). This is not time.time(): the two clocks have different origins, and a wall-clock value would give a deadline far in the past or future.
Both forms return a context manager object, usually bound with as cm. Calling cm.reschedule(new_when) moves the deadline, again using a loop.time() value. cm.when() reads the current deadline, and cm.expired() tells you afterwards whether the deadline was what stopped the block. Starting with timeout(None) and calling reschedule once you know the limit is a common way to set a deadline from a config you have only just loaded.
import asyncio async def main(): loop = asyncio.get_running_loop() async with asyncio.timeout(None) as cm: print(cm.when()) cm.reschedule(loop.time() + 0.3) await asyncio.sleep(0.05) print('done in time') print(cm.expired()) try: async with asyncio.timeout_at(loop.time() + 0.1): await asyncio.sleep(1) except TimeoutError: print('timeout_at fired') asyncio.run(main())
No deadline at first, a deadline set later, then an absolute one
None done in time False timeout_at fired
wait_for and wait
await asyncio.wait_for(aw, timeout) is the older, single-purpose tool. It wraps one awaitable. On timeout it cancels that awaitable, waits for the cancellation to finish, and only then raises TimeoutError. So when the exception reaches you, the awaitable has already stopped and its finally blocks have run. Since Python 3.12, wait_for is implemented on top of asyncio.timeout, so the two behave the same way underneath. wait_for is simply the shorthand for a single call.
asyncio.wait(tasks, timeout=..., return_when=...) is different in kind. It never cancels anything. It watches a collection of tasks and returns a (done, pending) pair of sets when the timeout passes or the return_when condition is met. FIRST_COMPLETED returns as soon as any one task finishes. Whatever is in pending is still running, and stopping it is your job. Note that wait takes tasks (or futures), not bare coroutines.
import asyncio async def job(n, delay): await asyncio.sleep(delay) return n async def main(): try: await asyncio.wait_for(job(1, 1), timeout=0.1) except TimeoutError: print('wait_for: TimeoutError') tasks = [asyncio.create_task(job(1, 0.05)), asyncio.create_task(job(2, 0.5))] done, pending = await asyncio.wait( tasks, timeout=0.2, return_when=asyncio.FIRST_COMPLETED) print(len(done), len(pending)) print([t.cancelled() for t in pending]) for t in pending: t.cancel() await asyncio.gather(*pending, return_exceptions=True) print([t.cancelled() for t in pending]) asyncio.run(main())
wait_for cancels for you; wait leaves the slow task running until you cancel it
wait_for: TimeoutError 1 1 [False] [True]
The three tools differ in what they cover and in what they do to the work when time runs out:
| Tool | Covers | On timeout | Cancels the work? |
|---|---|---|---|
asyncio.timeout() | A block of several awaits | TimeoutError after the block exits | Yes, the current task |
asyncio.wait_for() | One awaitable | Waits for the cancel, then TimeoutError | Yes, that awaitable |
asyncio.wait() | Many tasks | Returns (done, pending), raises nothing | No, never |
Nested timeouts and the swallowed-cancel trap
Timeouts can be nested, and the rule is that the innermost expired deadline raises. Each timeout() context counts the cancel requests on the task through task.cancelling(). When it exits after its own expiry it calls uncancel() to take its request back. If the count is still above what it was when the block started, someone else (for example an outer timeout) also cancelled the task, so the CancelledError is left to propagate instead of becoming a TimeoutError. The outer block then turns it into its own TimeoutError. That is how an outer expiry is told apart from an inner one.
import asyncio async def main(): async with asyncio.timeout(1): try: async with asyncio.timeout(0.05): await asyncio.sleep(1) except TimeoutError: print('inner expired') print('outer still running') try: async with asyncio.timeout(0.05): async with asyncio.timeout(1): await asyncio.sleep(2) except TimeoutError: print('outer expired') print(asyncio.current_task().cancelling()) asyncio.run(main())
Inner expiry is recoverable; outer expiry passes through the inner block
inner expired
outer still running
outer expired
0In the first half the inner deadline fires, is caught, and the outer 1-second budget carries on. In the second half the outer deadline fires first. The inner block sees a cancel that is not its own, so it lets the CancelledError through, and the outer block converts it. Every uncancel() has been paired with its cancel, so the final cancelling() count is 0.
A timeout works only by cancelling the task, so any code that catches CancelledError and does not re-raise it defeats the timeout. If the code returns normally, the block exits with no exception and no TimeoutError ever appears. If it loops and keeps going, the block may hang past its deadline. Clean up in finally, or catch CancelledError, tidy up and raise.
import asyncio async def stubborn(): try: await asyncio.sleep(1) except asyncio.CancelledError: print('swallowed') return 'finished anyway' async def main(): try: async with asyncio.timeout(0.1): print(await stubborn()) except TimeoutError: print('TimeoutError') print('block ended without TimeoutError') asyncio.run(main())
The cancel was swallowed, so the timeout never surfaces
swallowed finished anyway block ended without TimeoutError
Do not put except TimeoutError inside the async with block, because it can never match there. Do not pass time.time() to timeout_at. Do not assume wait() cancels slow tasks, since pending is yours to cancel.
Use timeout() for a block with one budget, wait_for() for a single awaitable, and wait() when you want to watch many tasks without cancelling them. Catch TimeoutError outside the block, never swallow CancelledError inside it, and cancel the pending tasks from wait() yourself.
Part 10 · Structured Concurrency in Practice
Scopes that own their tasks
Structured concurrency means a task's lifetime is bounded by a lexical scope. You open a block, start children inside it, and the block does not exit until every child has finished. The indentation of your code then shows who owns what, the same way nested function calls show who called whom.
Python 3.11's asyncio.TaskGroup is the tool for this. The async with block is the scope, and tg.create_task(...) starts a child that belongs to it. Leaving the block waits for all children, so no task can slip out past the end of the scope.
- 1async with TaskGroup()scope opens
- 2tg.create_task(...)children start, owned by the scope
- 3body finishesblock reaches its end
- 4wait for childrenexit blocks until all are done
- 5scope closednothing outlives the block
What the scope guarantees
Three guarantees follow from tying tasks to a scope. Together they are the reason to prefer it over loose tasks.
| Guarantee | What it means | What goes wrong without it |
|---|---|---|
| No orphans | Every child finishes or is cancelled before the block exits | A forgotten task keeps running after its caller moved on |
| Errors reach the parent | A child failure surfaces at the async with block, as an exception group | The error is logged late or lost |
| Cancellation flows down | Cancelling the parent cancels the whole subtree | Children keep running after the parent is stopped |
Where the idea comes from
The idea comes from the Trio library, whose nursery is a block that owns its child tasks. Nathaniel J. Smith argued for it in his essay 'Go statement considered harmful': a bare spawn call, like a goto, lets control flow escape its block. TaskGroup is asyncio's nursery, added in Python 3.11.
If a task is started inside a block, it must also be finished inside that block.
Own every task: helpers and services
A scope only protects you if every task is started through it. When a helper function needs to start work, give it the group as a parameter instead of letting it call a bare create_task. The caller then decides which scope owns the tasks, and the helper cannot leak one.
import asyncio async def work(i): await asyncio.sleep(0.01 * i) print(f'worker {i} done') def spawn_workers(tg, n): for i in range(n): tg.create_task(work(i)) async def main(): async with asyncio.TaskGroup() as tg: spawn_workers(tg, 3) print('all workers finished') asyncio.run(main())
The helper is a plain function: it only needs the group to attach tasks to it.
worker 0 done worker 1 done worker 2 done all workers finished
The last line prints only after all three workers are done, because leaving the block waits for them. If spawn_workers had called asyncio.create_task instead, the event loop would hold only a weak reference to those tasks. Nothing would wait for them, and an error in one could disappear.
Background services belong in main()
Long-lived services such as a heartbeat, a metrics reporter or a cache refresher also need an owner. Open one TaskGroup in main() and start the services inside it. They then run for as long as the app does, and when main()'s block exits or is cancelled they are cancelled and their finally blocks run. That is what gives you a graceful Ctrl+C: asyncio.run cancels main(), the group cancels its children, and nothing is left behind.
A helper that calls bare asyncio.create_task creates an orphan. Nobody awaits it, nobody sees its exception, and it can outlive the scope that called the helper. Pass tg in instead.
Keep the loop responsive
Tasks switch only at an await, so a blocking call such as time.sleep or a long CPU loop freezes every task on the loop. Send that work to a thread with await asyncio.to_thread(fn). The thread runs the function while the loop keeps serving its other tasks.
import time def crunch(): time.sleep(0.05) # stands in for blocking or CPU work return 'crunched' async def ticker(ticks): for _ in range(3): await asyncio.sleep(0.01) ticks.append('tick') async def main(): ticks = [] async with asyncio.TaskGroup() as tg: tg.create_task(ticker(ticks)) job = tg.create_task(asyncio.to_thread(crunch)) print(job.result(), len(ticks)) asyncio.run(main())
The ticker keeps ticking while the blocking function runs in a thread.
crunched 3Fan-out limits and producer/consumer
A group will happily start ten thousand tasks if you ask it to. To cap how many run at once, share an asyncio.Semaphore(n) and wrap the real work in async with sem:. All the tasks exist, but at most n are past the semaphore at any moment.
active = 0 peak = 0 async def fetch(u): global active, peak active += 1 peak = max(peak, active) await asyncio.sleep(0.01) active -= 1 return u * 2 async def get(sem, u): async with sem: return await fetch(u) async def main(): sem = asyncio.Semaphore(3) async with asyncio.TaskGroup() as tg: tasks = [tg.create_task(get(sem, u)) for u in range(10)] print(sum(t.result() for t in tasks), 'peak', peak) asyncio.run(main())
Ten tasks, but never more than three inside fetch at once.
90 peak 3
Producer/consumer with a queue
When work arrives over time, use an asyncio.Queue and a fixed set of worker tasks inside one group. The producer fills the queue, then await queue.join() returns once every item has been matched by a task_done() call. Workers still wait on queue.get() at that point, so you stop them with a sentinel: one None per worker, sent after join().
async def worker(q, results): while (item := await q.get()) is not None: results.append(item * item) q.task_done() q.task_done() # account for the sentinel itself async def main(): q = asyncio.Queue() results = [] n = 3 async with asyncio.TaskGroup() as tg: for _ in range(n): tg.create_task(worker(q, results)) for x in range(1, 7): q.put_nowait(x) await q.join() for _ in range(n): q.put_nowait(None) print(sorted(results)) asyncio.run(main())
Results are sorted because the workers finish in no fixed order.
[1, 4, 9, 16, 25, 36]
Starting thousands of tasks with no semaphore floods the target. Forgetting the sentinels leaves workers blocked on q.get() forever, so the group never exits. Forgetting task_done() makes join() hang.
Choosing a tool and the checklist
gather(..., return_exceptions=True) and TaskGroup answer different questions. With gather, a failure is just another slot in the result list and the other calls carry on. With a group, a failure is an emergency: the siblings are cancelled and the error propagates. The example below runs the same failing call both ways.
async def ok(n): await asyncio.sleep(0.01) return n async def boom(): raise ValueError('bad row') async def slow(): await asyncio.sleep(1) async def main(): res = await asyncio.gather(ok(1), boom(), ok(3), return_exceptions=True) print([type(r).__name__ if isinstance(r, Exception) else r for r in res]) try: async with asyncio.TaskGroup() as tg: s = tg.create_task(slow()) tg.create_task(boom()) except* ValueError as eg: print('caught', len(eg.exceptions), 'slow cancelled:', s.cancelled()) asyncio.run(main())
gather keeps the good results; the group cancels slow() as soon as boom() fails.
[1, 'ValueError', 3] caught 1 slow cancelled: True
| gather(return_exceptions=True) | TaskGroup | |
|---|---|---|
| On an error | Keeps going | Fail-fast |
| Siblings | Keep running | Cancelled |
| Result | Partial list, errors in their slots | All results, or an ExceptionGroup |
| Handle errors with | An isinstance check on each slot | except* |
| Best for | Best-effort batch | All-must-succeed work |
| Python version | Any | 3.11+ |
Checklist before you ship
- No bare
create_task: spawn through atg, or keep a strong reference yourself. - Re-raise
CancelledErrorafter cleanup; swallowing it breaks groups, timeouts and shutdown. - Use
shieldsparingly, only for work such as a commit or a payment that must not tear. - Catch errors from a group with
except*, never a plainexcept. - Cap fan-out with a
Semaphore, and stop queue workers with one sentinel each. - Send blocking and CPU work through
asyncio.to_thread.
A plain except ValueError does not match the ExceptionGroup that a TaskGroup raises, even when a ValueError is inside it. Use except* ValueError.
Part 11 · asyncio Tasks Cheatsheet
Spawn, Keep References, gather vs TaskGroup
The default way to start concurrent work is a TaskGroup. Open it with async with asyncio.TaskGroup() as tg: and start each child with tg.create_task(c()). The call returns a Task handle straight away. Leaving the block waits for every child, so you read t.result() after the block, when each task is known to be done.
Sometimes you really do want fire-and-forget work with no scope around it. The event loop keeps only weak references to tasks, so a task nobody holds can be garbage-collected mid-run. Park the task in a set and let it remove itself when it finishes with add_done_callback(set.discard). The program below shows both patterns.
import asyncio async def square(n): await asyncio.sleep(0.01) return n * n background = set() def spawn(coro): t = asyncio.create_task(coro) background.add(t) t.add_done_callback(background.discard) return t async def mail(): await asyncio.sleep(0.01) print('mail sent') async def main(): async with asyncio.TaskGroup() as tg: a = tg.create_task(square(3)) b = tg.create_task(square(4)) print(a.result(), b.result()) spawn(mail()) print('pending:', len(background)) await asyncio.sleep(0.05) print('pending:', len(background)) asyncio.run(main())
TaskGroup for owned work, a set for fire-and-forget
9 16 pending: 1 mail sent pending: 0
When several awaitables must run together, the real choice is between gather and TaskGroup. They differ in what happens on failure, and that decides which one you want.
| gather | TaskGroup | |
|---|---|---|
| Results | List in argument order | Task handles, t.result() after the block |
| One child fails | Error raised to the awaiter; siblings keep running | Siblings are cancelled (fail-fast) |
| Errors raised | The first one | All of them, in an ExceptionGroup |
| Partial results | return_exceptions=True puts errors in their slots | None: all or nothing |
| Best for | Best-effort batches | All-must-succeed work |
Calling asyncio.create_task(job()) and dropping the result. The loop holds only a weak reference, so the task can vanish mid-run, and a failure shows up only as a late Task exception was never retrieved log line.
Fail-Fast Groups, Cancel and Shield
A TaskGroup is fail-fast. When one child raises, the group cancels its siblings and the body, waits for the cancellations to finish, and then raises an ExceptionGroup holding every failure. A plain except ValueError does not match a group, so the error sails past it. Use except* ValueError as eg: instead. eg is itself a group that holds only the matching leaves, and eg.exceptions is the tuple of those leaves.
The program below also shows that gather with return_exceptions=True returns the error object in its slot, so you get partial results and nothing is raised.
import asyncio async def ok(n): return n async def bad(): await asyncio.sleep(0.01) raise ValueError('x') async def slow(): try: await asyncio.sleep(10) finally: print('slow cancelled') async def main(): res = await asyncio.gather(ok(42), bad(), ok(7), return_exceptions=True) print(res) try: try: async with asyncio.TaskGroup() as tg: tg.create_task(slow()) tg.create_task(bad()) except ValueError: print('plain except caught it') except ExceptionGroup as g: print('plain except missed:', type(g).__name__) try: async with asyncio.TaskGroup() as tg: s = tg.create_task(slow()) tg.create_task(bad()) except* ValueError as eg: print(len(eg.exceptions), eg.exceptions[0]) print(s.cancelled()) asyncio.run(main())
Partial results with gather, fail-fast with TaskGroup
[42, ValueError('x'), 7] slow cancelled plain except missed: ExceptionGroup slow cancelled 1 x True
To stop a task, call t.cancel(), which is only a request. Then await t inside except asyncio.CancelledError so you know it has really ended. The error is thrown into the task at its next await, runs its finally blocks, and finishes with t.cancelled() true.
- 1t.cancel()returns True: request made
- 2Next awaitCancelledError thrown inside
- 3Cleanupfinally or except, then re-raise
- 4await tcaller sees CancelledError
CancelledError has been a BaseException since Python 3.8, so except Exception never sees it. Catch it only to tidy up, and always re-raise. Swallowing it breaks TaskGroup, timeout() and app shutdown.
For a write that must not tear half-way, keep a reference to the inner task and wait on it through asyncio.shield(t). If the outer task is cancelled, the shield raises CancelledError in the waiter but the inner task keeps running. In the handler, await t to let the write land, then re-raise.
import asyncio async def worker(): try: await asyncio.sleep(10) except asyncio.CancelledError: print('cleanup') raise async def commit(): await asyncio.sleep(0.05) print('committed') return 'ok' async def handler(inner): try: await asyncio.shield(inner) except asyncio.CancelledError: await inner print('finished before exit') raise async def main(): t = asyncio.create_task(worker()) await asyncio.sleep(0) print(t.cancel()) try: await t except asyncio.CancelledError: print('cancelled:', t.cancelled()) inner = asyncio.create_task(commit()) h = asyncio.create_task(handler(inner)) await asyncio.sleep(0.01) h.cancel() try: await h except asyncio.CancelledError: print('handler cancelled:', h.cancelled()) print(inner.result()) asyncio.run(main())
Cancel then await, and a shielded commit
True cleanup cancelled: True committed finished before exit handler cancelled: True ok
Writing except CancelledError: pass inside a task makes it ignore the cancel. Shielding without keeping a reference to the inner task lets it be garbage-collected. Shielding everything creates tasks that cannot be stopped, and they block shutdown.
Deadlines, Blocking Work and Fan-Out
Use async with asyncio.timeout(s): to put one deadline over a whole block, where every await inside shares the budget. At the deadline the current task is cancelled, and the block converts that into TimeoutError on exit. So you catch TimeoutError outside the block, never inside it. For a single awaitable, await asyncio.wait_for(aw, s) cancels it on timeout and raises the same error. asyncio.wait(tasks, timeout=s) is different: it cancels nothing and returns (done, pending), and cancelling the pending tasks is your job.
import asyncio async def main(): try: async with asyncio.timeout(0.05): await asyncio.sleep(1) except TimeoutError: print('deadline hit') try: await asyncio.wait_for(asyncio.sleep(1), 0.05) except TimeoutError: print('wait_for timed out') fast = asyncio.create_task(asyncio.sleep(0.01)) slow = asyncio.create_task(asyncio.sleep(1)) done, pending = await asyncio.wait({fast, slow}, timeout=0.1) print(len(done), len(pending)) for t in pending: t.cancel() await asyncio.gather(*pending, return_exceptions=True) asyncio.run(main())
Three deadline tools, three behaviours
deadline hit wait_for timed out 1 1
The loop is single-threaded, so a blocking call or a CPU-heavy loop freezes every task. Hand such work to a thread with await asyncio.to_thread(fn). To bound how many jobs run at once, wrap each one in async with asyncio.Semaphore(n). Here five jobs run with at most two in flight.
import asyncio, time def blocking(n): time.sleep(0.05) return n * 2 async def job(n, sem, stats): async with sem: stats['now'] += 1 stats['peak'] = max(stats['peak'], stats['now']) r = await asyncio.to_thread(blocking, n) stats['now'] -= 1 return r async def main(): sem = asyncio.Semaphore(2) stats = {'now': 0, 'peak': 0} async with asyncio.TaskGroup() as tg: ts = [tg.create_task(job(n, sem, stats)) for n in range(5)] print([t.result() for t in ts], stats['peak']) asyncio.run(main())
Threads for blocking work, Semaphore for bounded fan-out
[0, 2, 4, 6, 8] 2
| Goal | Call | Effect |
|---|---|---|
| Deadline for a block | async with asyncio.timeout(s) | Cancel, then TimeoutError on exit |
| Deadline for one awaitable | await asyncio.wait_for(aw, s) | Cancels it, raises TimeoutError |
| Wait without cancelling | asyncio.wait(aws, timeout=s) | Returns (done, pending) |
| Blocking or CPU work | await asyncio.to_thread(fn) | Runs in a thread, loop stays live |
| Bounded fan-out | asyncio.Semaphore(n) | At most n in flight |
Putting except TimeoutError inside the async with asyncio.timeout() block, where it never matches. Assuming wait() cancels the slow tasks. Calling time.sleep inside a coroutine. Fanning out over thousands of items with no Semaphore.
Version Map
Much of this toolkit depends on the Python version, so check which interpreter you target before you reach for a feature.
| Python | What changed |
|---|---|
| 3.8 | CancelledError becomes a BaseException, so except Exception no longer catches it |
| 3.11 | TaskGroup, asyncio.timeout() and except* with ExceptionGroup arrive |
| 3.12 | eager_task_factory is added, and wait_for is rebuilt on top of timeout |
Spawn inside a TaskGroup and read results after the block. Hold fire-and-forget tasks in a set. Use gather with return_exceptions=True for partial results. Handle groups with except*. Cancel, then await, and always re-raise CancelledError. Shield only a referenced inner task. Catch TimeoutError outside the block.
Part 12 · Check yourself
Quiz
Read each program, decide what it does, and only then open the answer underneath it. Each answer comes after the code it explains.
1. A cancel that meets except Exception
import asyncio async def worker(): try: await asyncio.sleep(10) except Exception: print('caught') finally: print('cleanup') async def main(): t = asyncio.create_task(worker()) await asyncio.sleep(0) t.cancel() try: await t except asyncio.CancelledError: print('cancelled', t.cancelled()) asyncio.run(main())
What does the program above print, and in what order?
- It prints
cleanup, thencancelled True. The wordcaughtnever appears. sleep(0)lets the worker start and suspend insidesleep(10), socancel()throws CancelledError into it there.- CancelledError is a BaseException, so
except Exceptiondoes not match it. Thefinallyblock still runs, and the task ends as cancelled. - Awaiting a cancelled task raises CancelledError in the awaiter, which is why
mainprints the second line.
2. Spot the bug in the error handling
import asyncio async def boom(): raise ValueError('bad') async def main(): try: async with asyncio.TaskGroup() as tg: tg.create_task(boom()) except ValueError: print('handled') asyncio.run(main())
The author expects to see handled. Does it print, and what is the fix?
- It does not print. The program crashes with an ExceptionGroup traceback that shows the ValueError inside it.
- TaskGroup wraps every child failure in an ExceptionGroup, and a plain
except ValueErrordoes not match a group that merely contains one. - Fix: write
except* ValueError as eg:so the clause matches the ValueError leaf, and readeg.exceptionsif you need the errors.
3. What happens to the sibling?
import asyncio async def fail(): await asyncio.sleep(0.1) raise RuntimeError('x') async def slow(): await asyncio.sleep(0.3) print('slow finished') async def main(): try: await asyncio.gather(fail(), slow()) except RuntimeError: print('gather raised') await asyncio.sleep(0.5) print('end') asyncio.run(main())
What does the program above print? Then say what would change if gather were replaced by a TaskGroup.
- It prints
gather raised, thenslow finished, thenend. - With the default
return_exceptions=False, the first exception reaches the awaiter, butslow()is not cancelled. It keeps running as an orphan and finishes at 0.3s whilemainis still sleeping. - With a TaskGroup,
slow()would be cancelled whenfail()raises, soslow finishedwould never print. The error would also arrive as an ExceptionGroup, so the handler must beexcept* RuntimeError.
4. Where does the TimeoutError appear?
import asyncio async def main(): try: async with asyncio.timeout(0.1): try: await asyncio.sleep(1) except TimeoutError: print('inside') except TimeoutError: print('outside') asyncio.run(main())
Which word does the program above print, inside or outside, and why?
- It prints
outside. - At the deadline the timeout cancels the current task, so code inside the block sees CancelledError, not TimeoutError. The inner
except TimeoutErrornever matches. - The conversion to TimeoutError happens when the
async withblock exits, so the handler belongs outside the block.
Summary
- Keep a strong reference to every task you create, or better, spawn it with
TaskGroup.create_taskso the group owns it. gatherleaves siblings running after a failure, while a TaskGroup cancels them and raises every error together as an ExceptionGroup.- Handle TaskGroup errors with
except*, because a plainexceptdoes not match a group. - CancelledError is a BaseException: clean up in
finallyorexcept CancelledError, then re-raise it, and stop a task witht.cancel()followed byawait t. shieldprotects inner work only from the outer task's cancellation, so keep a reference to the inner task and use it sparingly.asyncio.timeoutbounds a block and raises TimeoutError outside it,wait_forbounds one awaitable, andwaitcancels nothing.- Blocking or CPU-heavy work goes through
asyncio.to_thread, and bounded fan-out uses a Semaphore inside a TaskGroup.