Bounding Concurrency in Go: errgroup, Worker Pools, and Backpressure

Cap the work in flight, not just cancel it.

Page content

Unbounded fan-out is the default failure mode of Go concurrency: every go statement is one more task in flight, with no ceiling and no owner. This article is about giving both.

Cancellation and bounding are separate controls. Cancellation decides when work stops; bounding decides how much of it exists at once. A service can cancel perfectly and still fall over because ten thousand goroutines each decided, independently, that now was a good time to call the database.

glass funnel throttling a flow of glowing particles through a regulated gate

The pages around this one cover the other half of the problem. Go context.Context Done Right treats context as the control plane for stopping work; here the subject is the capacity plane — errgroup.SetLimit, semaphore channels, worker pools, and bounded queues — plus the one leak shape that survives errgroup.Wait even when every cancel path looks wired up.

Five primitives, one decision table

Primitive Partial results Fail-fast Backpressure Dynamic work Best for
sync.WaitGroup + errors.Join yes no no no wait-all batches where every result counts
errgroup.WithContext no yes no no fan-out where the first error ends the request
errgroup.SetLimit(n) no yes yes (blocks Go) no bounded fail-fast fan-out
Semaphore channel yes manual yes (blocks acquire) yes custom admission control
Worker pool yes manual yes (queue fills) yes steady streams, long-lived workers

Two questions pick the row. First: when one task fails, do you still want the others’ results? If yes, fail-fast group semantics discard work you needed. Second: does work arrive continuously, or is it a fixed batch you can count before launching? Groups handle batches; pools and queues handle streams.

errors.Join is what makes the plain WaitGroup row viable — it collects every worker’s error instead of keeping only the first, which pairs with the boundary-translation rules in Go Error Handling Architecture.

The leak that survives errgroup.Wait

The best-known errgroup failure mode has nothing to do with bounding. It is a cancellation gap, and it hides behind code that looks correct:

g, gctx := errgroup.WithContext(ctx)
results := make(chan int) // unbuffered

go func() { // producer — nobody owns this goroutine
    for i := 0; i < 100; i++ {
        results <- i
    }
    close(results)
}()

g.Go(func() error { // consumer — group member
    for r := range results {
        if r == 5 {
            return fmt.Errorf("save failed")
        }
    }
    return nil
})

err := g.Wait() // returns at the consumer's error

When the consumer returns, gctx cancels and Wait returns. The producer is parked on results <- i — and context cancellation does not unblock a channel send. Nobody owns that goroutine, so it stays parked for the life of the process. Run this shape in a loop and the leak is exactly linear:

Variant Live goroutines after 500 iterations
Send without a ctx.Done arm 501 (+500 leaked, one per iteration)
Send wrapped in select with <-gctx.Done() 502 (+1 baseline)
// the fix: every channel operation inside a cancellable task
// gets a Done arm
select {
case results <- i:
case <-gctx.Done():
    return
}

The same gap exists on the receive side. The rule is mechanical: every channel operation inside code that can be cancelled gets a <-ctx.Done() arm, and every goroutine has an owner that waits on it or cancels it. goleak in CI catches the cases review misses — it slots into the same automated quality gate as the tools in Go Linters: Essential Tools for Code Quality.

Two related shapes are worth naming. If the producer is a group member, the same code produces not a leak but a deadlock — Wait blocks forever waiting for the parked send, which at least fails loudly. And a producer that sends with select but receives on a channel whose consumer has quit leaks the same way; Done arms belong on both sides.

errgroup.SetLimit: bounded fail-fast fan-out

SetLimit turns an errgroup into an admission-controlled group. Each Go call beyond the limit blocks until a slot frees:

g, ctx := errgroup.WithContext(ctx)
g.SetLimit(8)
for _, item := range items {
    g.Go(func() error {
        return process(ctx, item)
    })
}
err := g.Wait()

Three semantics matter in production. Go blocks, so the loop above applies backpressure to the producer — that is usually what you want, but it means the caller’s goroutine can stall, so a request handler feeding a large batch needs its own timeout budget around the whole loop. SetLimit(0) blocks every Go call forever. And changing the limit while goroutines are active panics — the limit is fixed for the life of the group.

SetLimit inherits group semantics: the first error cancels the context and Wait returns the first error, discarding results of work still in flight. It is the right primitive when failing the whole operation on first failure is correct — fetching one page of data, calling a set of independent health checks, fanning a request out to shards.

Semaphore channels: backpressure as a feature

A buffered channel of capacity n is a semaphore with no dependencies:

sem := make(chan struct{}, 8)
var wg sync.WaitGroup
for _, item := range items {
    wg.Add(1)
    sem <- struct{}{} // blocks when 8 tasks are in flight
    go func() {
        defer wg.Done()
        defer func() { <-sem }()
        _ = process(ctx, item)
    }()
}
wg.Wait()

The acquire happens before the go statement, so goroutines are only ever created when a slot exists. The buffer size is the contract: it is simultaneously the concurrency ceiling and the queue of pending admissions.

The primitive earns its place when you need admission control the groups do not express — context-aware acquisition, per-downstream pools, or partial results on failure:

select {
case sem <- struct{}{}: // slot acquired
case <-ctx.Done():      // caller gave up before getting in
    return ctx.Err()
}

Measured on the same workload, the semaphore channel and SetLimit(8) finish within 1.5% of each other — the choice between them is semantics, not performance.

Worker pools: for streams, not batches

A fixed pool of long-lived workers separates admission (the queue) from execution (the workers):

func pool(ctx context.Context, workers int) chan<- func() {
    jobs := make(chan func())
    var wg sync.WaitGroup
    for i := 0; i < workers; i++ {
        wg.Add(1)
        go func() {
            defer wg.Done()
            for {
                select {
                case f, ok := <-jobs:
                    if !ok {
                        return
                    }
                    f()
                case <-ctx.Done():
                    return
                }
            }
        }()
    }
    return jobs
}

Pools fit work that arrives continuously — queue consumers, polling loops, request workers — where a per-batch group would be created and torn down endlessly. Since Go 1.25, sync.WaitGroup.Go removes the Add/Done pairing boilerplate for the common shape. The shutdown contract is the part to get right: closing jobs terminates workers after the queue drains; cancelling ctx abandons queued work. Both are legitimate, but only one can be the documented behavior.

Bounded queues: block or drop

A buffered channel is also a queue with a ceiling. Once the buffer fills, the producer blocks — and that moment is the design decision. Blocking propagates pressure upstream to whoever feeds the queue; dropping sheds load at the cost of lost work. A monitoring pipeline can drop; an order pipeline must block or spill to durable storage.

select {
case queue <- event: // admitted
default:            // queue full: drop policy lives here
    dropped.Add(1)
}

Whatever the policy, make it explicit and measured. A silent default branch is where events go to disappear.

What bounding costs — measured

The same workload, 10,000 tasks of 5 ms simulated I/O each, Go 1.27.1:

Strategy Peak goroutines Wall time
Unbounded go per task ~10,000 launched 18 ms
errgroup.SetLimit(8) 8 6.9 s
Semaphore channel (cap 8) 8 6.6 s

The unbounded run wins on wall time because nothing in the experiment pushes back — ten thousand timer goroutines simply sleep in parallel. That is exactly the trap. The cost of unbounded fan-out is never CPU in your own process; it is ten thousand simultaneous connections, a downstream that starts timing out under the burst, and a memory graph that follows the queue depth of whoever you are calling. On a real dependency, the unbounded run does not finish in 18 ms — it times out. The bounded runs pay 6.6 seconds because 10,000 tasks ÷ 8 slots × 5 ms is 6.25 s of unavoidable serialization, and the cap is the point: it turns an uncontrolled burst into a predictable 6.25-second drain.

One refinement matters when the cap is higher than a handful of slots: keep the cap at or below what the downstream can sustain concurrently, and derive it from that limit (connection pool size, rate limit, worker capacity) rather than from a feeling about “reasonable” parallelism.

Measuring and testing bounded code

Three signals cover most of it. Goroutine count trends (/sched/goroutines:goroutines exported via the Prometheus Go collector) catch leaks as a slope. Scheduler latency (/sched/latencies:seconds) catches saturation — runnable goroutines waiting for CPU time is real pressure regardless of the count. A pprof delta of two goroutine dumps pinpoints where the parked goroutines live.

For the tests, bounded workers are the case Testing Concurrent Go Code with synctest was built for: fake time makes a full 10,000-task drain run in milliseconds, synctest.Wait replaces sleep-and-hope for quiescence, and a bubble fails fast when a worker stays durably blocked — the leak repro above dies in a test bubble instead of in production.

At service scale, the same bounding question reappears between services, where the semaphore becomes a rate limit and the circuit breaker guards the downstream — Circuit Breaker Pattern in Go covers that layer, and Go Microservices for AI/ML Orchestration shows the queue-backed forms the orchestration tier takes.

Choosing

flowchart TD A[Work to run concurrently] --> B{Fixed batch or continuous stream?} B -- continuous stream --> C[Worker pool or bounded queue] B -- fixed batch --> D{Partial results needed
on first failure?} D -- yes --> E[WaitGroup + errors.Join
or semaphore channel] D -- no --> F[errgroup.WithContext] F --> G{Need a concurrency ceiling?} G -- yes --> H[errgroup.SetLimit] G -- no --> I[group without limit
only when the batch is provably small] C --> J{Full queue: block or drop?} J -- block --> K[Backpressure upstream] J -- drop --> L[Shed load, count the drops]

The table and the flowchart compress to one habit: before writing go, name the ceiling and the owner. The ceiling is SetLimit, a semaphore, or a bounded queue; the owner is a Wait, a ctx.Done arm on every channel operation, or both. Every goroutine-creation site that can answer those two names in one glance is one that will not page anyone at 3 a.m.

Subscribe

Get new posts on AI systems, Infrastructure, and AI engineering.