Project Phoenix: The rewrite everyone tells you not to do
Rebuilding a working system from scratch is the one thing every experienced engineer will tell you not to do, and for most teams they are right. It was right for us on the code alone. What made us do it anyway was a data model that could not describe the customers we wanted.
The constraint was the model, not the code
We wanted to move more than corporate commuters. School runs, major events, factories, pet transport, non-emergency medical transport. Different transportation models, different markets, and we could not take them.
Every time we scoped one the estimate came back wrong the same way. Our data model knew one kind of customer: a company, with employees, going from home to a fixed workplace on a weekday.
A school run needs a guardian attached to a passenger who can't consent for themselves. An event needs capacity for one Saturday and none the week after. Pet transport needs a passenger who never logs in.
Each of those breaks a different assumption in the same tables. A passenger was a row in a company's employee list, so it could not exist without an employer, could not have someone else consent on its behalf, and could not be a dog. A trip was a recurring weekday commute between two fixed points, so a single Saturday with 400 seats was not a scheduling edge case, it was a different object. And the payer was always the employer, which is not true for a parent, an event organiser, or a hospital. We could have added nullable columns for all of it. What we could not do was change what the tables meant while live data depended on the old meaning.
The screens on top assumed that single customer as well, so the front ends came with it.

Our system was a building of the kind you find all over Istanbul: Roman foundation, Byzantine arches, Ottoman stonework, a Republic-era apartment on top, still occupied. Now try putting in an elevator. Every era in those walls has its own rules, so the answer is no, or yes if you scratch nothing. Fine for a historical object. We needed a foundation to build on.
The floors above were what you would expect: practices frozen at 2018, two abandoned refactors and their scaffolding, no tests, no observability, repeated code, a user management system per application. Firefighting took more of each sprint than new work.
Underneath all of it sat the data model, and it was tired. There is no refactoring your way out of a foundation.
Most teams do not have that problem, which is why "never rewrite" is good advice and Joel Spolsky's case for it has held up since 2000. It is for teams who want to rewrite because the code offends them. If that is yours, it holds.
We were not the first to set it aside. Dropbox rebuilt its sync engine from scratch and said openly the call would have been wrong elsewhere. Uber rewrote its fulfillment platform because a model built around trips and supply could not carry new verticals.
So we made the case for replacing the foundation rather than maintaining it or repairing it, because refactoring toward a nicer version of a model that assumes the wrong thing gets you a nicer version of the wrong thing. The codebase decided what the rewrite cost. It did not decide whether to do one.
Four questions decided it:
1. Is the constraint in the code, or in the model the code encodes? Code you refactor in place. A model you cannot, once live data depends on it. Be strict here: "the model is wrong" is the most flattering story a team can tell about a codebase it finds annoying.
2. Can both systems run side by side long enough to move customers one at a time? If they cannot, cutover has to be one event, and that is the version of a rewrite that kills companies.
3. Do we know what done looks like? Rewrites without a definition never fail, they just never finish. Ours: every customer on the new core, old core off, one system running.
4. Can we afford both systems for the full window, and how long is it? Without a number here you cannot tell a migration from a second platform you now own.
There is a fifth test we would apply now, and it is about who is asking. On-call teaches you which failures are weekly and which ugly code is ugly for a reason. If the people arguing for a rewrite have never carried that pager, their estimate is missing the parts that cost the most, and that alone sends us back to refactoring.
Greenfield build, phased cutover
Building from scratch was forced on us. The seam a strangler fig needs runs through the API surface, and ours ran through the tables: the same rows had to mean one thing to the old core and another to the new one. So we could not intercept our way across.
What made it survivable was the shape of the switch.
A gateway sat in front of both systems and routed each customer to whichever core owned them. The old core stayed authoritative until that customer moved, and we migrated one at a time.
We started where a bad day was survivable: small operations, weekday-only service, and a coordinator we had worked with long enough to call at seven in the morning. That was not sandbagging the curve, it was buying the right to be wrong twice before it mattered.
Two mobile apps talked to the old core, one for passengers and one for drivers, and an app update could not be part of a cutover. Passengers do not update on our schedule and some never update at all. So the greenfield core learned to speak the old core's API, versions and quirks included, and answered on the endpoints the apps already called. Our engineers spent weeks writing endpoints in the shape of a system we were deleting. The apps never knew which core answered.

Each cutover moved that customer's coordinators onto the new internal platform the same day, so data and people moved together. It was one-way. Planners build routes in whichever platform they are in, so reversing a cutover meant replanning a live operation on the old system within hours, and nobody was going to do that at 6am on a Tuesday.
We had no rollback and knew it going in. What we had instead was a small blast radius, an order we chose, and about six incidents that customers noticed across 47 cutovers. We went forward every time because forward was the only direction available.
The first cutover took around two weeks and taught us more than the eleven months of building before it. The other 46 took nine.

None of that was luck. Our engineers wrote endpoints in the shape of a system we were deleting, sat with planners the day their tools changed, and fixed forward under load for eleven weeks. At the peak we moved eight customers in a day with no way back for any of them.
Rewrites die at the switch far more often than at the build.
What running two systems cost us
Two systems cost more than twice one, and we wrote that down before starting rather than discovering it in month four.
Every bug fix landed twice, or landed once and left a divergence to track. Through the build we carried a second set of environments. Through the cutovers we ran two production systems, which cost about 60% more than one rather than double, because the new core was leaner than what it replaced. The security surface doubled and the review load with it.
The front end doubled the same way: two internal platforms, two web apps, two visual languages, and an operations team working in both throughout.
The expensive part never shows up on an invoice. On-call carried two mental models. Everyone who joined that year learned both, including the one we were deleting.
So we bounded it from the start: the last customer migrating, the old core dark. Eleven months building, two weeks on the first cutover, nine weeks on the other 46. Thirteen months to the old core going dark.
What we got wrong
The standard criticism of greenfield rewrites applies to us. For 11 months the new system had no users and nobody to tell us a workflow was wrong.
The new data model was held. Some things around it did not: shortcuts planners expected, features we never thought to build, screens right in principle and awkward in use. We shipped those in the weeks after the first cutovers, fixing bugs as they came.
Closing
Since the cutover we have been taking on work the old core could not touch. Pet transport and school runs are live, and the same foundation lets us sell the platform itself, with a first customer already on it. The rest of the list at the top of this post is now a scheduling question.
The bet was expensive and hard to reverse. We priced the downside up front, stayed with it until the old system went dark, and came out of it with a foundation we can build on and a team that has done this once already.
Plenty of teams make a big technical bet. Fewer get to the end and switch off the system they replaced. If that is the kind of work you want to do, we are hiring.
Open roles at Volt Lines.