Notes · Dissecting Real Systems
growing
Utils Is Where Modularity Goes to Die
Module boundaries should follow the dependency graph, not your folder intuitions — and "optimal" can be defined precisely.
…organizations which design systems (in the broad sense used here) are constrained to produce designs which are copies of the communication structures of these organizations.
— Melvin E. Conway, "How Do Committees Invent?", Datamation (1968)
Cite this
Mangalapilly, Y. J. (2026, June). Utils Is Where Modularity Goes to Die. Saṃhitā Notes. https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/ @online{mangalapilly2026utils,
author = {Yesudeep Jose Mangalapilly},
title = {Utils Is Where Modularity Goes to Die},
journal = {Sa\d{m}hit\=a Notes},
year = {2026},
month = {June},
url = {https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/},
urldate = {2026-08-20},
} Yesudeep Jose Mangalapilly. “Utils Is Where Modularity Goes to Die.” Saṃhitā Notes, 2026. https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/. TY - ELEC
AU - Mangalapilly, Yesudeep Jose
TI - Utils Is Where Modularity Goes to Die
T2 - Saṃhitā Notes
PY - 2026
UR - https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/
Y2 - 2026-08-20
ER - A companion to the build-systems series: where it asked how build tools work, this asks how to write code so they work well. By the end you'll be able to define "optimal incremental build" as a precise objective function, state the common refactorings as deterministic algorithms, explain — formally — why the utils dumping ground is slow as well as ugly, and design against the forces that fill it in the first place.
Every codebase has the file. It's called utils.py, or helpers.js, or common.go, or Util.java. It started honest — one date-formatting function someone didn't know where else to put. Now it's fourteen hundred lines of datetime math, string parsing, number formatting, a retry decorator, and three functions that touch the DOM. You have written this file. You have grepped it at two in the morning.
It's easy to say utils is ugly — a junk drawer, low cohesion, the usual code-review tut-tutting. That's true and it's not the interesting part. The interesting part is that utils is slow, in a way you can measure and prove, and the proof tells you exactly how to fix it.
Cohesion is a build-time property, not an aesthetic one.
Feel the cost before reading the model of it. The figure below is not a picture of a split; it is the split, running in this page. It starts at — one utils, twenty-four consumers — where any edit rebuilds every consumer. Press Edit a function and watch the whole right column light up. Then slide to 4, edit again, and watch the same edit reach roughly a quarter of them. Nothing about the code changed; only where the boundaries are. The rest of the piece says why that number is the objective function and how to move it on purpose.
utils on the left, and k under your control: at k = 1 every edit rebuilds all twenty-four, and each extra split narrows what one edit can reach. Press Edit a function to land the next edit in another module.The model: a build is a graph, change is a distribution
From the survey we have the first half: a build is a directed graph whose nodes are units of compilation (files, modules) and whose edges mean " depends on " — change and may need rebuilding. A build system, given a changed node, rebuilds its dependents: everything reachable by following edges forward.
The second half is the part most discussions of modularity skip. Code doesn't change uniformly. Some modules churn constantly (business policy, UI copy); others are touched once a year (a hashing primitive, a parser for a frozen format). The rebuild cost is driven by the change probability of the edited module — and on a real project you don't have to guess : it is the file's commit frequency in git log.
Dependents — the set of nodes reachable from by following dependency edges forward; everything that must be reconsidered when changes. The opposite direction (what relies on) is its dependencies. Learn more.
With those two pieces — the graph and the change distribution — "optimal" stops being a matter of taste.
What "optimal" means, precisely
Let be the set of dependents of module , and let be the cost of rebuilding (compile time, say). When an edit lands in , the build system pays to rebuild and everyone downstream of it. The expected rebuild cost of the whole codebase, per unit time, is the sum over every module of "how often it changes" times "what that change costs to rebuild":
Reading the symbols: says "add up, over every module "; is "the module itself, plus everything downstream of it"; and , when it appears below, is the count of that downstream set. The formula is the sentence above it, written compactly — nothing in it is beyond arithmetic. Learn more.
Read it in plain terms: a module hurts in proportion to how often it changes multiplied by how much sits downstream of it. A hashing-library node has a large downstream cone but a tiny — it's cheap, because it never changes. A business-rules node changes daily but sits at a leaf with no dependents — also cheap. The expensive nodes are the ones that are both frequently changed and widely depended on. That product, , is the rebuild tax of a module.
Think of each module as paying an insurance premium: how often it causes an accident, times how many neighbors are in the blast zone. A hazmat depot that never has incidents is cheap to insure; a fireworks stand out in an empty field is cheap too. utils is a fireworks stand in a crowded market — small fires, constantly, with everyone downwind.
Because the tax is a product, it can be drawn honestly: one factor is a height, the other a width, and the cost is an area. Here are those three characters to scale — drag either factor and watch what the area does:
utils (both at once). The slivers are cheap; only the square is expensive, because only the square grows in both factors.An optimal set of module boundaries is one that minimizes over the ways you could partition the same functions into modules — given that you can't change what the code does, only how it's grouped and who depends on whom. Refactoring for incremental builds is exactly the search for that minimum.
Why utils is provably slow
Now the pathology falls out of the formula. A utils module fuses functions that change for unrelated reasons — the datetime helper, the parser, the formatter. Two consequences, both bad for :
- Its change probability is the union of its parts. A module changes if any of its functions changes, so — approximately the sum, because each reason fires rarely. (Precisely: the chance that at least one part changes is — one minus the chance that every part stays quiet, where multiplies those stay-quiet chances together. The plain sum is always at least that, and nearly equals it when each is small — the union bound.) Lumping raises .
- Everyone depends on all of it. Because file-granularity build systems track dependencies per file, every consumer of any function in
utilsis a dependent of the whole file. So is the union of the dependents of every function inside it — large.
Union bound — the chance that at least one of several rare things happens is at most the sum of their chances. Two 1-in-100 risks: the exact combined chance is ; the lazy sum says . The sum only overcounts the overlap — both firing at once — which for rare events is tiny. Learn more.
A module with both a high (everything's reasons, summed) and a large (everyone, unioned) is precisely the high-rebuild-tax node the objective tells us to eliminate. Edit the one-line datetime helper and the parser, the formatter, and everything downstream of the entire file rebuilds — work that the change-reasons had nothing to do with.
Before. One utils node. is the sum of every reason its functions change; is everyone who uses any of them. A datetime edit rebuilds the parser's dependents too.
After. Split into datetime, parse, format. Each has one reason to change (small ) and only its own dependents (). A datetime edit rebuilds only datetime's consumers.
Slide the split again, now that the tax has a name — the same figure from the top of the piece, twenty-four consumers, one utils, and under your control:
The refactorings, as algorithms
"Improve cohesion" is advice; an algorithm is something you can run. Here are the moves that lower , each stated as a deterministic procedure over the graph.
1 — Split by change-reason
The core move. Don't group "all the helpers"; group "everything that changes for the same reason."
Split a module by change-reason.
SPLIT-BY-CHANGE-REASON(M, H)
Input: module M with functions f[1:n], history H
Output: partition of M into cohesive modules
groups ← EMPTY-PARTITION
for i ← 1 to n
reason ← ESTIMATE-REASON(f[i], H)
groups ← ADD(groups, reason, f[i])
return EMIT-MODULES(groups)This lowers exactly when the reasons' consumers differ: each new module carries the change probability of one reason, not their sum, and its dependent set shrinks to the consumers of that reason's functions. The condition is the point — if every consumer used every reason, each piece keeps the full dependent set, the sum reassembles unchanged, and the split bought nothing but module count. Splitting by change-reason pays where usage was already partitioned; the algorithm just makes the partition official.
2 — Invert the dependency at a volatile boundary
When a stable module is forced to depend on a volatile one, the volatility leaks upward: the stable module inherits the volatile module's . Break it with an interface.
Invert a volatile dependency.
INVERT-VOLATILE-DEPENDENCY(stable, volatile)
Input: stable → volatile; volatile changes often
Output: volatile → I ← stable via interface I
required ← USED-OPERATIONS(stable, volatile)
I ← DEFINE-INTERFACE(required)
stable′ ← DEPEND-ON(stable, I)
volatile′ ← IMPLEMENT(volatile, I)
return (stable′, volatile′, I)This is the dependency-inversion principle, derived from the cost function: it removes a high- node from the dependency cone of a widely-used stable one.
Imagine a house wired with the lamps soldered straight to the fuse box. Every time you want a different lamp, an electrician has to open the fuse box. Put a socket on the wall instead: the fuse box now depends only on "a socket, of this shape," which never changes, and the lamp depends on the socket. Swap lamps all day and nobody touches the fuse box. The interface is the socket.
3 — Delete the barrel file
A barrel — index.ts that re-exports everything, or __init__.py that imports the world — silently re-fuses what you just split. Importing one symbol through the barrel makes you a dependent of the whole barrel, which depends on everything it re-exports. It reconstructs the utils fan-out you worked to remove.
Delete a barrel file.
DELETE-INTERNAL-BARREL(B, consumers)
Input: barrel B, modules M[1:k], internal consumers
Output: consumers importing from defining modules
for each consumer C in consumers
for each symbol s that C imports from B
owner ← DEFINING-MODULE(s, M)
C ← REWRITE-IMPORT(C, s, owner)
if B is not a public API boundary
DELETE(B)
return consumers4 — Keep hubs stable (the leaf/hub asymmetry)
The cost function says the dangerous nodes are high . A hub — a module many things depend on — has a large by definition, so it can only afford a tiny . Volatility belongs in leaves, where and a high costs nothing.
Triage by rebuild tax (a check, not a rewrite).
RANK-BY-REBUILD-TAX(M)
Input: modules M[1:n], change rates p, dependents D
Output: modules in descending rebuild-tax order
for i ← 1 to n
tax[i] ← p(M[i]) · |D(M[i])|
return STABLE-SORT-DESCENDING(M, tax)Why the drawer fills anyway
The four moves fix the utils you have. They don't explain why you'll have another one next year. Every codebase grows the drawer, which means the forces that fill it are stronger than the review culture guarding it — and a fix that doesn't name those forces is treating the symptom.
- The name is a deferral. Nothing lands in
utilsbecause it belongs there; it lands at the moment its author decides that figuring out where it belongs isn't worth the interruption. The module is named for what the code isn't — not parsing, not UI, not the domain — and "not" is not a change-reason. In the objective's terms:utilsis where functions with an unlabeled accumulate, and unlabeled reasons are exactly the ones that end up fused. - The friction is asymmetric. Appending to
utils.pyis one edit. A new module is a name to invent, a file, an import path, sometimes a build rule — and a review conversation about all four. Under deadline the cheap path wins every time; the drawer isn't a lapse of discipline, it's the equilibrium of that asymmetry. - Review pressure points the wrong way. The same reviewer who will spend three comments on a new module's name waves a helper into
utilswithout one. Each wave-through teaches the next author where code goes to avoid questions.
It's 4:50 on a Friday and the fix needs a parse_retry_after helper. There are two homes for it: a new http-headers module — a name, a file, an import path, and a reviewer asking whether it shouldn't be http/headers — or line 1401 of utils.py, where nobody will say anything. You know which one ships. So did everyone before you; that is how the file reached line 1400.
The countermeasures are the forces, inverted — one each:
- Park it next to its only caller. A helper with one caller needs no home: it lives in the caller's file and inherits the caller's change-reason, and the build graph sees nothing new. Move it out on the second unrelated caller — a cousin of the rule of three — because two real callers are the first evidence of what the helper actually is: what they share names the module.
- Ban the names, not the helpers.
utils,helpers,common,miscas a path segment is a confession that the change-reason is missing — so make the confession loud. A lint rule that rejects those path names turns a silent deferral into a visible decision at review time. The helper is welcome; the label "unlabeled" is not. - Make the right thing as cheap as the wrong one. The drawer wins on friction, so lower the other side's: a one-command module scaffold, and a norm that a one-function module with an honest name is normal. The graph doesn't mind small files; it minds fused reasons (the over-splitting caveat below still stands — stop at cohesive, not atomic).
- If you keep an inbox, give it a zero policy. Where a landing zone genuinely helps, treat it like an inbox: things may land, nothing may live. Cap its size in lint and drain it by move 4's triage — highest first.
The common thread is a feeling, and the feeling is data. The discomfort you carry about utils is the deferred decision — the open loop of a naming you skipped. Hating the module is a defense mechanism in the strict sense, and it's load-bearing: the hate is what remembering a deferral feels like, and it protects you from the real failure mode, which is comfort. A team at peace with its junk drawer has stopped noticing that it defers, and from there the tax compounds quietly. The formula gives the instinct its number; the instinct is still what fires first.
How to find your own utils problem
You don't have to guess which module to fix. The cost function hands you the procedure: compute the rebuild tax of every file and look at the top.
- — count commits touching the file over some window:
git log --oneline -- path/to/file | wc -l. - — the number of files that depend on it (your build system's dependency graph, or a static import scan).
- Rank by the product. The high-tax, top-right quadrant — frequently changed and widely depended on — is your
utils, whatever it's called. Split those by change-reason first; the long tail of low-tax modules can wait forever.
Ask the build system, not the import graph
An import scanner sees the language's edges. Your build system's edges are the ones that actually gate work, and where the two disagree is exactly where the interesting defects live — a file can import nothing unusual and still sit inside a build unit that wakes the world.
Bazel will hand you the whole graph, and the answer for any single file:
# Every node and edge, one JSON object per line.
bazel query 'deps(//...)' --output=streamed_jsonproto
# What one file invalidates — and just the tests among them.
bazel query 'rdeps(//..., //path/to:file.ext)' --output=label_kind
bazel query 'kind(".*_test", rdeps(//..., //path/to:file.ext))'
# Why an edge exists at all, when a dependent looks wrong.
bazel query 'somepath(//consumer:target, //suspect:target)'Other systems expose the same relation under other names: Buck2's cquery 'deps(...)', Cargo's --unit-graph, Go's go list -deps -json ./..., Ninja's -t graph, and GNU Make's --print-data-base. Any of them gives you without guessing at it.
A cheaper first pass: count the answers, not the sizes
Ranking by tax needs change history. There is a check that needs none, and it tends to find the same modules first. For one directory, count how many distinct dependent sets its files have.
Call that count the directory's invalidation classes. If it is one, every file in that directory invalidates exactly the same work, and no edit inside it can ever be cheaper than any other — the build cannot tell a leaf data structure from the module that everything imports. That is utils observed directly rather than inferred. "Unions its dependents" is the cause; one invalidation class is what the cause looks like from outside.
The check needs one guardrail. A single class is a defect only where the tool admits a finer unit. Rust compiles a crate at a time; Go compiles a package at a time. A crate whose files all share one dependent set is sitting on its floor, not in a drawer. Establish the compilation unit for each language before calling one class a finding, or the check will confidently indict the healthiest code you own.
The measure that hides the worst offender
counts dependents. That is the right question for a file many things import, and the wrong question for the opposite shape: one enormous unit that consumes everything.
A single test target holding three thousand tests contributes exactly one to every fan-out count in the repository. It may be the most expensive thing in the build, and it ranks last on the very metric meant to find expensive things. So compute the dual as well — for each action, how many inputs it declares:
bazel aquery 'deps(//...)' --output=jsonprotoHigh fan-out means one file wakes many units. High fan-in means one unit wakes for many files. Both are fusion. The quadrant only draws the first.
Check whether you're rebuilding at all
quietly assumes that everything invalidated is rebuilt. Under a content-addressed cache that assumption is false: an invalidated node whose inputs have been seen before is restored, for the price of a hash lookup and a copy. The honest cost carries a term the formula leaves out:
Content-addressed — stored and retrieved by a hash of the bytes themselves, so identical inputs always name the same entry, whichever machine or checkout produced it. Learn more.
where is the chance that 's inputs already resolve to a stored result.
The two factors are independent levers. Splitting a module shrinks . Turning on a cache shrinks the multiplier on every you already have — including all the ones you were never going to refactor. One is a line of configuration; the other is weeks of restructuring.
That ordering is uncomfortable for an essay about module boundaries, so let me state it plainly: if your builds are slow and nothing persists between them, your boundaries are not the first problem. A cache that does not survive a server restart, a branch switch, or a second checkout of the same commit is a cache you are not getting paid for. Check for one before you reach for the refactor — the check takes a minute, and it can save you the month.
None of which rescues utils. A cache lowers the cost of rebuilding what you invalidate; it does not lower what you invalidate. And a drawer that changes daily generates misses by construction, because its key changes whenever any one of its unrelated functions does — the same fusion, now expressed as a cache-miss rate instead of a rebuild count. The levers compose: cache what recurs, split what doesn't. Pull the cheap one first, then measure again before pulling the expensive one.
The honest limits
The model is a lamp, not a law. Three caveats:
- It needs estimated change probabilities, which means predicting the future from the past.
git logis a good prior, not a guarantee; a module quiet for a year can suddenly become hot. - Over-splitting has its own cost the formula doesn't capture. A thousand one-function modules minimize rebuild tax and maximize navigation and ceremony pain. The objective bounds the build cost, not the human cost; stop splitting when modules are cohesive, not when they're atomic.
- Granularity has two bounds, and your toolchain is rarely the binding one. The floor is the finest unit your build system can observe — the file for most (C++ does so even for headers, as the survey noted), the crate for Rust, the package for Go. Function-level incremental compilers move the floor down, and you can't refactor below it. But there is a second bound above it: the ceiling, the coarsest unit your own build files impose on top of that floor — and the ceiling is what usually binds. A toolchain that computes per-object dependencies still wakes everything if every test links one archive. A language with per-module tracking still fuses the lot if one build target globs the whole source tree. In those cases the finer granularity provably exists: the tool computed it, and the build definition threw it away. Find out which bound is binding before concluding you've hit a limit.
Lessons
- A build is a graph; change is a probability distribution over its nodes. Module design is the act of shaping both.
- Optimal has a definition: minimize . A module's rebuild tax is .
utilsis provably expensive: lumping unrelated functions sums their change probabilities and unions their dependents — high and large at once, the worst corner of the objective.- The refactorings are deterministic algorithms: split by change-reason, invert volatile dependencies, kill barrel files, keep hubs stable.
- To find what to fix, rank files by (git churn dependents) and start at the top. Take from the build system's graph, not an import scan — the build's edges are the ones that gate work.
- Check the cache before you split anything. assumes invalidated means rebuilt; a content-addressed cache makes it mean restored, multiplying every dependent cone by . Shrinking and shrinking the multiplier are independent levers, and the second one is a line of configuration.
- Fan-out alone can't see the worst unit. A single action holding three thousand tests counts as one dependent. Compute the dual — inputs per action — or the coarsest thing you own will rank last on the metric built to find it.
- One invalidation class means the unit is indivisible. Counting the distinct dependent sets in a directory needs no change history and costs one query. But normalize by compilation unit first: a Rust crate or a Go package with one class is at its floor, not in a drawer.
- Your build files set a ceiling above the toolchain's floor, and the ceiling usually binds. Per-object compilation fused into one archive, or per-module sources fused by a whole-directory glob, throws away granularity the tool already computed.
- The drawer fills because deferral is cheap and naming is not. Invert the friction: park a helper beside its only caller, move it out on the second unrelated caller, reject
utils-family path names in lint, and give any landing zone a zero-resident policy. Distrust comfort with the drawer more than you distrust the drawer.
Practice
References
- M. Conway. “How Do Committees Invent?.” 1968. — the original Conway's Law paper
- “Single-responsibility principle.” — and dependency inversion — the same ideas, stated without the build graph
- “Cohesion.” — and coupling, for the classic vocabulary
- “Rule of three.” — when a helper has earned a home of its own
- “The Build Is Proportional to the Change.” — where the build-graph model comes from
How to cite
Mangalapilly, Y. J. (2026, June). Utils Is Where Modularity Goes to Die. Saṃhitā Notes. https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/ @online{mangalapilly2026utils,
author = {Yesudeep Jose Mangalapilly},
title = {Utils Is Where Modularity Goes to Die},
journal = {Sa\d{m}hit\=a Notes},
year = {2026},
month = {June},
url = {https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/},
urldate = {2026-08-20},
} Yesudeep Jose Mangalapilly. “Utils Is Where Modularity Goes to Die.” Saṃhitā Notes, 2026. https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/. TY - ELEC
AU - Mangalapilly, Yesudeep Jose
TI - Utils Is Where Modularity Goes to Die
T2 - Saṃhitā Notes
PY - 2026
UR - https://yesudeep.com/blog/utils-is-where-modularity-goes-to-die/
Y2 - 2026-08-20
ER - Webmentions
Annotations
Thank you — your note is held for review and will appear once approved.
Thank you — your note is published.
Please sign in below to leave a note.
