Hey {{first name | there}}. I promised I was going to speak on something other than Kubernetes this week.

One thing I have been meaning to explore is what happens when a goroutine starts, blocks on a channel or a lock, and never gets woken up. The process can carry on serving traffic, so you might not notice anything until the goroutine count and memory usage start climbing. From there, a goroutine profile can show you where they are waiting, but you still need to work out whether something will eventually wake them up.

This is where the goroutine leak profiler introduced in Go 1.27 comes in.

In today's technical notes:  

  • What counts as a leaked goroutine

  • How Go identifies a leaked goroutine

  • What the profile still cannot see

🧠TOGETHER WITH EVERYTHINGDEVOPS

Cloud Engineer. Platform Engineer. SRE.

Three roles that can look almost identical from a job description, until you look at what the engineers actually spend their time doing.

Prince and Amarachi get into the differences between them, what each role tends to involve, and how to think about the path that fits the kind of engineering work you want to do.

Subscribe: @EverythingDevOpsHQ for deeper dives on agents, Kubernetes, MLOps, and AI infrastructure.

📰Technical notes: Understanding goroutine leaks

The profiler first appeared as an experiment in Go 1.26, where you had to build with GOEXPERIMENT=goroutineleakprofile to enable the goroutineleak profile.

With Go 1.27, the Go blog describes how to use it in a running service. If you already have net/http/pprof set up and serving its handlers, you can collect it from the following endpoint.

/debug/pprof/goroutineleak

Before looking at how this works, I think it is useful to establish what counts as a leak here.

A goroutine is leaked when it is blocked, and the conditions needed to unblock it can no longer be met. As these goroutines accumulate, they continue to hold memory, including anything they still reference, which also leaves more work for the garbage collector.

The difficulty is that being blocked is a normal part of how goroutines work.

The usual way this happens

In his writeup, Vlad Saioc uses an example of a pattern found in real services, including ones at Uber. Work is split between several goroutines, and each one sends its result back through an unbuffered channel.

ch := make(chan result)

for _, w := range ws {

    go func() {

        res, err := processWorkItem(w)

        ch <- result{res, err}

    }()

}

for range len(ws) {

    r := <-ch

    if r.err != nil {

        return nil, r.err

    }

}

Because the channel is unbuffered, each send waits for a receiver. This works while the loop is collecting results, but notice what happens when one result contains an error. processWorkItems returns early, the loop stops receiving, and any remaining worker that reaches ch <- has nobody left to receive its result.

This is the operation the leak profile identifies. In the example from the blog, go tool pprof reported 116 leaked goroutines at that send. As the program continued to repeat the work, more goroutines ended up waiting there.

For this particular example, giving the channel a buffer of len(ws) provides enough room for every worker to send its one result, even if the receiving function returns early. We still need to make that change ourselves, but the profile gives us a specific operation to investigate.

There are already tools that help catch this during testing. goleak can fail a test when goroutines are left running, and Go 1.25 introduced synctest to help test concurrent behaviour more reliably. Those tests are useful, but a service running in production may encounter conditions we did not reproduce in the test. The new profile gives us a way to investigate some of those leaks while the service is running.

How Go identifies a leaked goroutine

A natural next question is how the runtime knows that nothing will wake a goroutine up. To answer that, it looks at which goroutines can still reach the channel or other concurrency primitive the blocked goroutine is waiting on.

For this check, a goroutine is considered live if it is not blocked on one of the operations the profiler tracks, or if another live goroutine can reach something that could unblock it. Starting with the first group, the runtime follows references to channels and sync values, then includes the goroutines waiting on those values. It repeats this as long as it finds more goroutines that could still be woken up.

Once that process finishes, the remaining goroutines are blocked on operations that nothing live can reach. Those are the ones reported as leaked.

If this sounds familiar, it is because the garbage collector already does much of the work needed to follow references through memory. A normal collection starts with every goroutine as a root, along with global variables. The leak check changes which goroutines it starts from, then expands that group as it finds more that could be unblocked. Globals still count as reachable, which matters for the limitations we will come to shortly.

The research behind this calls it partial deadlock detection. In Go, it is used to identify and report the goroutines, so collecting the profile does not cancel them or recover their memory for you.

A small way to look at one

To try this, I would start with a service built using Go 1.27 that already exposes pprof. Assuming it is listening locally on port 6060, you can collect and open the profile with the following commands.

go tool pprof leak.prof

From there, use list followed by the function name to see the source around the blocked operation. For the example above, that would be list processWorkItems. The output follows the same format as a goroutine profile, with the results limited to the goroutines identified by the leak check.

If the service is still on Go 1.26, it needs to have been built with GOEXPERIMENT=goroutineleakprofile for that endpoint to be available.

I would keep goleak in the tests as well. Catching a leak before a change reaches production is still useful, and the runtime profile gives us another way to investigate the ones that make it through.

What this changes in practice

If you have been watching a goroutine count climb, this gives you another place to start looking. You can collect the leak profile and see whether the runtime has identified any goroutines that can no longer be unblocked.

For the worker example, that takes us from a growing count to the send operation left waiting after the receiving function returned. We still need to understand why the function returned and decide how to handle the remaining work, but we have a much smaller part of the program to investigate.

The same check will not help with every goroutine stuck in I/O, so there is still a place for the regular goroutine profile and some time spent reading the code.

🌍IN THE ECOSYSTEM

Goroutine Leak Profiles walks through Vlad Saioc’s worker example, how to collect the profile and how the runtime decides which goroutines are leaked.

Go 1.26 release notes: experimental goroutine leak profile covers the initial experiment, enabled with GOEXPERIMENT=goroutineleakprofile.

Dynamic Partial Deadlock Detection and Recovery via Garbage Collection is the ASPLOS 2025 paper behind the technique. Its research prototype, Golf, also explored recovering from leaks, while the Go profile focuses on detecting them.

⏱️UNTIL NEXT TIME

That’s it for this issue. If you have a service on a supported Go build with pprof enabled, try collecting /debug/pprof/goroutineleak after it has been running for a while. I would be interested to hear whether it points to something you already suspected, or what you found when you followed the reported stack back to the code.

And if someone on your team has been trying to work out why their goroutine count keeps growing, send this their way.

Know an engineer who will find this helpful? Share this link with them

Jubril Oyetunji
CTO, EverythingDevOps

HOW DID WE DO?

Login or Subscribe to participate