
Most Go developers know how to write efficient Go code. But very few know what happens to that code after go build.
Consider:
func add(a, b int) int { return a + b}func main() { x := add(10, 20) fmt.Println(x)}
How many function calls happen here?
The obvious answer is one.
But the real answer may be zero.
The Go compiler can inline add, turning the call conceptually into:
X := 10 + 20
It can then perform constant folding:
x := 30
The code you write and the code your CPU executes can therefore be very different.
That is what compiler optimization is about.
1. What Does a Compiler Do?
At the simplest level:
Source Code ↓Compiler ↓Machine Code
But modern compilers do much more than translation. They analyze your program and transform it into a more efficient equivalent program.
A useful high-level model is:
SOURCE PROGRAM
│
▼
FRONTEND
│
AST / IR
│
▼
MIDDLE-END
│
Optimized IR / SSA
│
▼
BACKEND
│
▼
MACHINE CODE
Frontend
The frontend understands the source program.
It performs tasks such as: Lexing, Parsing, Type checking, Building the AST, Building compiler IR
The question it answers is:
“What does this program mean?”
Middle-end
This is where many important optimizations happen.
Examples: Inlining, Escape analysis, Constant propagation, Dead-code elimination, Devirtualization, SSA transformations
The question is:
“How can this program be represented and optimized?”
Backend
The backend turns the optimized representation into instructions for a specific CPU architecture.
It handles things such as: Instruction selection, Register allocation, Stack layout, Architecture-specific optimizations, Machine-code generation
The question is:
“How should this optimized program run on this CPU?”
2. A Brief History of Compilers
Early computers were programmed using machine code and assembly.
High-level languages introduced a new abstraction:
Human-friendly source code ↓ Compiler ↓ Machine cod
As languages became more complex, compilers became more sophisticated.
They introduced techniques such as:
- Constant folding
- Register allocation
- Dead-code elimination
- Function inlining
- Data-flow analysis
- Common-subexpression elimination
- Static analysis
Modern compilers are therefore not just translators.
They are program transformation systems.
3. The Go Compiler
The standard Go compiler is commonly called gc, meaning the Go compiler — not garbage collector.
Its source is primarily under:
src/cmd/compile
The Go compiler itself is written in Go.
This was a major change introduced with Go 1.5, when the compiler and runtime were moved from C to Go.
Go later adopted an SSA-based compiler backend in Go 1.7.
SSA became a major foundation for Go’s optimization pipeline.
4. SSA — The Important Part
SSA means:
Static Single Assignment
The basic idea is that each SSA value is assigned only once.
Normal code:
x := 10x = x + x = x * 2
Conceptually becomes:
x1 = 10x2 = x1 + 1x3 = x2 * 2
Why is this useful?
Because it makes relationships between values much easier for the compiler to analyze.
SSA enables or simplifies optimizations such as:
- Constant propagation
- Common-subexpression elimination
- Dead-code elimination
- Bounds-check elimination
- Nil-check elimination
- Register allocation
This is one of the most important concepts to understand when studying the Go compiler.
5. Optimization #1 — Inlining
Consider:
func add(a, b int) int { return a + b}func main() { x := add(10, 20)}
The compiler may replace:
add(10, 20)
with:
10 + 20
This is function inlining.
It removes the function-call overhead and, more importantly, exposes the function’s body to additional optimizations.
For example:
add(10, 20) ↓ 10 + 20 ↓ 30
You can inspect inlining decisions with:
go build -gcflags="-m=2"
You may see:
can inline addinlining call to add
6. Optimization #2 — Escape Analysis
Go automatically decides whether values need to live on the heap.
Consider:
func foo() int { x := 10 return x}
x does not need to survive the function.
Now:
func foo() *int { x := 10 return &x}
The function returns a pointer to x.
The compiler determines that x must survive after the function returns, so it may move the allocation to the heap.
Conceptually:
Does value escape? │ ┌─────┴─────┐ No Yes │ │ Stack Heap
Check this with:
go build -gcflags="-m=2"
You may see:
moved to heap: x
This is why the statement:
“Pointers always go to the heap”
is incorrect.
The compiler performs escape analysis to determine where the value needs to live.
7. Optimization #3 — Bounds Check Elimination
Go arrays and slices are memory-safe.
For:
numbers[i]
the runtime semantics require checking that i is valid.
But consider:
for i := 0; i < len(numbers); i++ { total += numbers[i]}
The compiler can prove:
i < len(numbers)
Therefore, it can eliminate a redundant bounds check.
This is called:
Bounds Check Elimination (BCE)
The important idea is:
Safety check required ↓Compiler proves it is unnecessary ↓Check removed
Go therefore keeps its safety guarantees while still generating efficient code.
8. Optimization #4 — Dead Code Elimination
Consider:
if false { expensiveOperation()}
The compiler knows this branch can never execute.
It can remove it.
Similarly:
x := 10y := 20return x + yprintln("never reached"
The unreachable code can disappear.
This is Dead Code Elimination (DCE).
Benefits include:
- Less machine code
- Smaller binaries
- Less runtime work
9. Optimization #5 — Constant Folding
The compiler can calculate constant expressions during compilation.
Instead of:
x := 10 * 20
it can effectively produce:
x := 200
Similarly:
x := 20
can become conceptually:
y := 15
This involves:
- Constant folding
- Constant propagation
These optimizations are often combined with other optimizations.
For example:
function call ↓Inlining ↓constant expression ↓constant folding ↓dead code elimination
One optimization can therefore create opportunities for another.
10. Optimization #6 — Common Subexpression Elimination
Consider:
a := x * yb := x * y
If the compiler can prove that x and y haven’t changed, it can calculate:
x * y
once and reuse the result.
Conceptually:
x * y
/ \
a b
instead of performing the multiplication twice.
This is called:
Common Subexpression Elimination (CSE)
11. Devirtualization
Go interfaces can involve dynamic method dispatch.
For example:
type Speaker interface { Speak()}func run(s Speaker) { s.Speak()}
If the compiler can determine the concrete type behind the interface, it may convert the indirect call into a direct call.
Conceptually:
Interface call ↓Compiler proves concrete type ↓Direct call ↓Possible inlining
Again, one optimization can enable another.
12. Register Allocation
The CPU has a small number of extremely fast registers.
The compiler tries to keep frequently used values in them rather than repeatedly accessing memory.
Conceptually:
Memory → CPU → Memory
is generally more expensive than:
Register → CPU → Registe
But registers are limited.
The compiler therefore performs register allocation to decide which values should stay in registers.
This is one of the final steps before machine-code generation.
13. Profile-Guided Optimization
Traditional optimization mainly uses information available from the source code.
But the compiler doesn’t necessarily know what happens in production.
For example:
function A → called 10 million timesfunction B → called 10 times
PGO allows production profiling information to influence compiler decisions.
The workflow is:
Application ↓CPU Profile ↓default.pgo ↓Go Compiler ↓Optimized Binary
Go introduced PGO experimentally in Go 1.20 and made it production-ready in Go 1.21.
PGO can improve decisions such as inlining based on actual application behaviour.
14. Common Myths
“Small functions are slow.”
Not necessarily.
They may be inlined.
“Pointers are always faster.”
No.
Pointers can introduce indirection and influence escape behaviour.
“Every slice access has a bounds check.”
Not necessarily.
The compiler can eliminate checks it can prove unnecessary.
“Returning structs is expensive.”
Not necessarily.
The compiler can keep values on the stack or in registers and eliminate unnecessary work.
“The source code is what the CPU executes.”
Definitely not.
The compiler may transform it substantially.
15. The Right Way to Think About Go Performance
Don’t think:
Go code ↓ CPU
Think:
Go source ↓Parser / Type Checker ↓Compiler IR ↓Inlining ↓Escape Analysis ↓SSA ↓Constant Propagation ↓Dead Code Elimination ↓Bounds Check Elimination ↓Devirtualization ↓Register Allocation ↓Machine Code ↓ CPU
Your source code is an abstraction.
The compiler is free to change the implementation as long as the required behaviour is preserved.
16. Recommended Resources
Official Go Compiler Documentation
The best starting point for understanding the actual compiler pipeline:
Go Compiler Introduction
https://go.dev/src/cmd/compile/README
Go Compiler Source
Explore the implementation:
cmd/compile
https://go.dev/src/cmd/compile/
Especially:
cmd/compile/internal/inlinecmd/compile/internal/escapecmd/compile/internal/ssacmd/compile/internal/walk
Compiler Optimizations
Go Compiler and Runtime Optimizations
https://go.dev/wiki/CompilerOptimizations
Go Assembly
A Quick Guide to Go’s Assembler
https://go.dev/doc/asm
Profile-Guided Optimization
PGO in Go
https://go.dev/doc/pgo
Go 1.7 and SSA
Go 1.7 Release Blog
https://go.dev/blog/go1.7
Go 1.5 and the Self-Hosting Compiler
Go 1.5 Release Blog
https://go.dev/blog/go1.5
SSA Research
Efficiently Computing Static Single Assignment Form
https://research.ibm.com/publications/efficiently-computing-static-single-assignment-form-and-the-control-dependence-graph
Conclusion
The Go compiler does much more than convert Go code into machine code.
It asks questions such as:
Can this function disappear?Can this allocation disappear?Can this bounds check disappear?Can this calculation happen at compile time?Can this interface call become direct?Can this value stay in a register?Is this code unreachable?
If the compiler can prove that an optimization is safe, it can transform the program.
