Most designs die after they ship.
Not in a meeting. In production. Under retries, bad counts, and people who are tired. I care about that part. The part where a system either stays honest or slowly starts lying.
I'm Debanjan. Senior engineer, 8+ years in backends. Right now I work on email deliverability and sequences at Apollo.io. Here I write the long form: outages, early-career mistakes, calm under pressure, and why building still matters.
Shorter thoughts, same obsessions: @theybanjan. Production scars, career questions, how to keep building without becoming a hero cosplay.
Outages next to people stuff. That's the job.
When retries become the outage
One slow shard did not take us down. Our retries did.
Calm is a production skill
Heroics feel like leadership in the moment. They usually make the next incident worse.
The questions that made me less junior
I stopped asking only 'does it work?' and started asking what happens twice, late, and when it fails halfway.
Why care about building at all
Code is temporary. The trust people put in what you shipped is not.
At Snapdeal I drove 100% source-to-target message accounting, cut AWS cost about 21%, redesigned WhatsApp as a bot-of-bots, and lifted SonarQube quality scores 78%. Writeups on counting messages and retry storms. Full path on the timeline.