There was one word I could never say at Amazon.
Not a profanity. Just an ordinary engineering term (one syllable) describing something my teams deliberately built and quietly benefited from for years. It didn’t appear on some list of banned vocabulary, and it didn’t need to. You learned within one or two review cycles which words helped you out in front of an executive audience and which ones sank the room. This particular word sank the room immediately, every time.
I will tell you the word in a moment, but first I want to show you what it does, because I am currently sitting inside a demonstration of it.
I am writing this on a train from Munich to Salzburg, and I am watching the minutes to my connection with a wariness I earned honestly.
Last year, while traveling through Germany, my wife and I missed connections that should have been easy. The issues came from the fact that the Deutsche Bahn network has spent years being asked to do more than it was built for. The trains run close together, and the timetable assumes nothing goes wrong. So when anything does go wrong anywhere in the system, the delay doesn’t stay where it started. It propagates. A ten-minute problem in one city becomes a missed connection three cities away, and the passengers absorb what the system could not.
The Swiss, famously, run their trains differently. Their timetable builds in margin: a little room at connections, a little breathing space between services. The trains are on time not because Swiss trains are magic, but because the system does not require perfection to deliver punctuality. When something slips, the slack absorbs it, and the failure stays local instead of cascading across the country.
Two networks. Two philosophies. One runs at the edge of its capacity and transmits every shock to every passenger, while the other holds something in reserve and delivers the thing everyone actually wants: arriving when you said you would.
The word for what the Swiss build into their timetable is “slack”: deliberately reserved capacity, held back on purpose, so the system can absorb what the plan did not predict.
That is the word I could never say at Amazon.
In front of an executive audience, especially an Amazon audience, “slack” does not sound like engineering. It sounds like slacking. In a company with “frugality” printed on the wall as an operating principle, where every headcount was fought for in writing and every dollar was defended in a document, the suggestion that you were deliberately not running at full capacity was close to heresy. Anyone who mentioned it would be quietly put on trial by their audience.
Now, you could build it. You could benefit from it. Smart teams did this. But you could not name it, because the name confessed to holding something in reserve instead of optimizing every inch of the process. And reserve sounded like waste to people whose job was hunting waste.
Amazon ran under what I think of as “The German System,” but other organizations I worked for had a more “Swiss” approach. I spent thirty years watching these two philosophies run inside software organizations. I know which one I always built, and I also know what it cost me to defend it in rooms where its true name was unsayable.
At Amazon, I ran the Appstore’s backend services. Around us, plenty of teams lived in a familiar rhythm: dates slipping, operational load climbing, engineers paged at two in the morning, heroics celebrated in the retrospective, and then the same fire the following month.
In contrast, my teams hit their dates. They aligned cleanly with partner teams, and they were not woken up in the middle of the night. All because we had spent real engineering effort making sure of it.
None of that was luck, and none of it was because my engineers were smarter than anyone else’s. It was because we deliberately built what the Swiss build into their train system: margin. We invested in infrastructure before it was on fire. We automated the operational work that was eating people. We fixed root causes instead of celebrating the person who patched the symptom at 3am. We did not schedule ourselves to one hundred percent of capacity, because we knew something would arrive that was not on the plan, and something always did.
When the surprise came, and it always did, it landed on the margin instead of on someone’s weekend. The delay stayed “local.” It did not cascade.
Here is what I learned about systems that run at their limit, and it applies to railways, to service architectures, and to teams: a system with zero slack does not fail gracefully. It fails everywhere, all at once. Anyone who has watched a queue at full utilization knows the shape of it; the math does not degrade politely; it goes vertical. The organization that schedules every person to capacity and every date to the optimistic case has not eliminated waste. It has arranged for the next surprise to become a catastrophe.
Building in “slack” required serious translation work. Actually building the thing was the easy half. But funding it, review after review, required speaking in code.
Instead of the word-that-must-not-be-named, I said “resilience.” I said “operational excellence.” I said “capacity headroom,” “sustainable pace,” “reduced operational load,” “availability engineering,” etc.
All of these words were true; all of these words were accurate. But all of these words were a euphemism tax I paid so that the room would never hear that one dreaded syllable that would have killed the investment in our work. The same document, with “slack” written where “resilience” stood, would have been dead by the second page. The word for deliberately reserved capacity and the word for laziness are close enough cousins that the concept becomes indefensible the moment you name it honestly.
So, I learned to camouflage it in a more palatable vernacular.
“Slack” was not alone on the unwritten banned list, either. “Technical debt” lived nearby, distrusted for the same underlying reason: both words admit out loud that the optimistic plan was never the whole truth. That one deserves its own essay someday.
It is strange to spend effort disguising good engineering as something other than what it is, but I understood the instinct I was working around. In a culture that measures delivered results, reserved capacity looks like waste, because its entire value lives in the “what if?” The outage that hasn’t occurred yet, the date that hasn’t slipped. There is no line on a dashboard for disasters that have yet to arrive. But arrive they do. They always do.
This is why, in my experience, it is so often the quiet leader who ends up building the slack. Building margin requires thinking in second-order effects, caring about what happens after the launch instead of at it, and being willing to advocate for a future that nobody else in the room can see. It also requires being at peace with a hard truth: if you do this well, it will look like nothing.
That is the trap: Prevention is invisible, and firefighting is loud.
The team that ships heroically at 2am generates war stories, visible urgency, and promotion-document material. The team that quietly built enough margin so that nothing catches fire generates a calm that gets misread. I watched it happen to my own teams. Our smoothness was sometimes mistaken for having it easy, as if the absence of crisis meant the absence of difficulty. As if the lack of disaster was a product of pure luck rather than intentional engineering.
Some organizations, without ever intending to, end up rewarding the arsonist with the firehose. The person whose corner of the system is always dramatically on fire, and who always dramatically extinguishes it, out-narrates the person whose corner never burns. If your promotion process cannot tell those two people apart, your best builders of margin will eventually leave, or worse, learn to start fires.
The wrinkle here is that too much slack really is waste. A system padded everywhere is a system that has stopped making important choices, and an organization can absolutely hide laziness inside the vocabulary of “resilience” and “readiness.” I have seen padded estimates that were just fear, and buffers that were just an unwillingness to commit.
So the answer is not a number. There is no universal amount of slack to build in. Anyone who gives you a percentage is selling something. The answer is simply a question you can continually ask about your own system, whether it is a service architecture, a team’s calendar, or a national railway:
“When the last surprise arrived, what absorbed it?”
If the answer is designed margin, the shock stayed local, and the system kept its promises, you have slack. If the answer is someone’s weekend, someone’s health, or a cascade of missed commitments three teams away, you have a system running at its limit, and the efficiency you are so proud of is a loan against the next surprise, at an interest rate you have not calculated.
The Swiss trains are on time because the timetable does not demand that every train be perfect. The German trains are perfectly on time until one slip throws them into utter chaos.
There is a version of one of these that lives in your organization.
Which one is it? And what is the cost?











