Almost every attempt to make code faster starts in the wrong place. We look at the code, pick the part that seems slow and rewrite it. Then nothing changes, because the time was somewhere else. A profiler replaces the guess with a measurement. It is the most valuable tool that most developers never open.
The rule
Do not optimise anything you have not measured. It sounds obvious. In practice, it is ignored daily, by experienced people, because guessing is quicker than measuring and feels like progress.
Performance work has three steps: measure, change one thing, measure again. If the second number is not better, undo the change.
Start with the clock
Before you reach for a profiler, find out which part is slow. Simple timing is enough.
# file: report.py
import time
t0 = time.perf_counter()
rows = load_rows()
t1 = time.perf_counter()
totals = summarise(rows)
t2 = time.perf_counter()
render(totals)
t3 = time.perf_counter()
print(f"load {t1 - t0:.2f}s summarise {t2 - t1:.2f}s render {t3 - t2:.2f}s")load 0.31s summarise 8.74s render 0.12sNow you know where to look. You have also learned where not to look, which saves even more time.
Sampling profilers
A sampling profiler interrupts the program many times a second and records which function is running. After a few seconds, it has a statistical picture of where the time goes. The overhead is low, so you can use it on real workloads.
$ py-spy record -o profile.svg -- python report.pyEvery major language has one: py-spy for Python, perf for native code, the built-in profiler in Node and in browsers, pprof for Go.
Read a flame graph
The output is usually a flame graph. It looks complicated, and the reading rules are short.
- Each box is a function.
- The width of a box is the share of time spent in that function and everything it calls.
- Boxes stack: a function sits on top of its caller.
- The colours mean nothing.
Look for wide boxes near the top. A wide plateau is a function that does a lot of work itself. That is your target.

A flame graph answers one question: if I could make one function free, which one would save the most?
Priya Nair
What you usually find
After years of profiling, the same few causes appear again and again.
Work inside a loop that belongs outside
# Before: compiles the pattern 200,000 times
for line in lines:
if re.match(r"^\d{4}-\d{2}", line):
count += 1
# After
pattern = re.compile(r"^\d{4}-\d{2}")
for line in lines:
if pattern.match(line):
count += 1The wrong data structure
Testing membership in a list reads the whole list. In a set, it is one step.
seen = set(existing_ids) # not a list
new = [r for r in rows if r.id not in seen]With 100,000 rows, this change alone turned 40 seconds into 0.05.
One query per item
The application asks the database for a list, then asks again for each item in it. A page with 50 items makes 51 queries. Fetch the related data in one query, and the page is ten times faster.
Doing the same work twice
A function is called with the same arguments many times. Cache the result.
from functools import lru_cache
@lru_cache(maxsize=1024)
def exchange_rate(currency: str, day: str) -> float:
return fetch_rate(currency, day)Memory is time too
A program that allocates a lot spends its time in the allocator and the garbage collector. If the profile shows wide boxes with names like malloc or gc, look at what you create in hot loops. Reuse buffers, stream large files and avoid building a list you read only once.
| Symptom in the profile | Likely cause |
|---|---|
| One wide box in your own code | An expensive loop |
| Wide boxes in a library | You call it too often |
| Time in the garbage collector | Too many short-lived objects |
Time in read or recv | Waiting on disk or network |
| Nothing is wide | The program is waiting, not computing |
Waiting is not computing
A CPU profiler shows where the processor is busy. If your program spends its time waiting for a database or a network call, the profile looks empty. Use a wall-clock profiler or tracing to see the waits. For a web request, a trace that shows each query and each outgoing call is often more useful than a flame graph.
Know when to stop
Set a target before you start: this page loads in 200 ms, this job finishes in ten minutes. When you reach it, stop. Code that is tuned past its target is harder to read, and nobody benefits.
- Decide what "fast enough" means
- Measure to find the largest cost
- Fix that one thing
- Measure again
- Stop when you reach the target
Should I profile in production?
Are micro-benchmarks useful?
The habit
The next time something is slow, do not open the editor first. Open the profiler. In ten minutes, you will know what is slow. It will surprise you, and the fix will usually be smaller than the one you had planned.