James Okafor joined Pinnacle Enterprise Solutions as CTO in 2019 when the company had 80 engineers and a monolith that had been quietly accumulating debt for seven years. Two years later, a catastrophic outage on Black Friday would reshape how the company thought about technical risk, team topology, and the real meaning of ownership. We sat down with him at Pinnacle's Portland headquarters to talk through the incident, the 22-month recovery, and the leadership principles he formed along the way.

When you joined Pinnacle as CTO in 2019, what did the engineering organization look like?

Honestly? It was a beautiful mess. We had around 80 engineers split across three product teams, all sharing what I can only describe as a monolith with ambitions. The system had grown organically over seven years. There were parts of it that literally nobody understood anymore — services without owners, integrations without documentation, configuration files that had not been touched since 2014 because everyone was afraid of what might break.

We had one architect — brilliant, invaluable, irreplaceable — who was effectively the single point of failure for most architectural decisions. The bus factor for our core systems was one. If he got on a plane, we held our breath. That's not a technical problem. That's an organizational failure that gets expressed as a technical risk.

"The bus factor for our core systems was one. If he got on a plane, we held our breath. That's not a technical problem — it's an organizational failure expressed as technical risk."

You've spoken publicly about what you call the "near-death experience" in 2021. What actually happened?

Black Friday, 2021. We were processing roughly twelve times our normal transaction volume — it was our biggest commercial season and we had invested in marketing to drive that. The system fell over. Not partially degraded — completely down for four hours and twenty minutes. We lost somewhere in the range of eight million dollars in revenue in a single afternoon, and the downstream damage to customer trust was harder to quantify but absolutely real. We had retailers calling our CEO directly. That is not a conversation you want to have.

The root cause was a database connection pool we had been meaning to refactor for eighteen months. There was even a ticket for it. Multiple engineers knew about it. We had simply never prioritized it — it was always below the line in sprint planning, always something we'd get to in the next quarter. The technical debt was real, visible, and documented. We just chose not to pay it.

That incident changed everything. It was devastating in the moment, but it was also clarifying. It turned every abstract conversation about technical debt into something extremely concrete. I didn't have to explain what "accumulated risk" meant anymore. Everyone had just lived it.

How did you approach the recovery?

The first thing I did was resist the temptation to immediately launch a "rewrite everything" initiative. That's the natural instinct after a catastrophic failure — you want to make a dramatic gesture, show the organization you're taking decisive action. But big-bang rewrites have a genuinely terrible track record. They take three times longer than planned, they require you to understand the old system well enough to replicate its behavior including the implicit assumptions nobody documented, and they leave you in a two-year window where you have two systems to maintain and neither is quite production-ready.

Instead, we identified the five critical paths in the system — the ones that touched revenue directly — and we drew a hard line: nothing else ships until these are stabilized. It was painful. We stopped two features that were ninety percent complete. Product leadership hated us for about three months. But we finished those stabilizations, and when we came out the other side, we had something we could actually reason about.

From there, we used what's often called a strangler fig pattern — gradually replacing components of the monolith with well-bounded services, routing traffic incrementally, proving each piece at scale before moving on. We never flipped a switch. Everything was a slow migration with rollback capability. It took twenty-two months from the incident to what I'd call architectural stability. But by the end of 2023, we could handle ten times our 2021 peak load without breaking a sweat.

"We stopped two features that were ninety percent complete. Product leadership hated us for about three months. But those were the right three months to spend."

What does the industry generally get wrong about scaling?

Two things. First, teams scale architecture too early or too late — almost never at the right time. The ones who go microservices-first on day one are burning engineering cycles on infrastructure problems before they've validated their product. The coordination overhead alone can kill a small team. Then you have teams on the other end — growing fast, clearly needing to evolve, but so terrified of disruption that they keep adding layers of abstraction on top of the monolith until it becomes structurally unsound. Both failure modes are common. The timing of the transition is genuinely hard.

The second thing — and I don't hear enough people talk about this — is that teams conflate technical scaling with organizational scaling. You can have the most beautiful distributed architecture in the world and still ship slowly if you haven't figured out your team topology. Conway's Law is real. Your system will reflect your organizational structure whether you planned it that way or not. If you want loosely coupled services, you need loosely coupled teams. If your teams are organized around functional layers instead of product domains, your architecture will express that, and it will be expensive.

What's the most underrated skill in a CTO?

Narrative. Not storytelling in the marketing sense — I mean the ability to give your engineers a compelling explanation of why the work matters. The why behind the technical direction you're choosing. Engineers are fundamentally motivated by context and clarity. You can get people to execute without it. But you cannot get them to be genuinely creative, to spot the problems before they become incidents, to feel real ownership over outcomes. That requires narrative. It requires your team to be able to finish the sentence: "We're building this because..."

The best engineers I've ever worked with don't just want to build the thing. They want to understand the theory of change behind the thing. What's the problem we're actually solving? What is the shape of the solution space? Why did we make this tradeoff instead of that one? When people have that context, the quality of their decisions goes up dramatically, even at the junior level.

What do you tell engineers who aspire to CTO roles?

Start leading before you have the title. Find a technical problem that nobody owns and own it — not by doing all the work yourself, but by organizing others around solving it, documenting the decision-making, following up. The transition from senior engineer to staff to principal is about expanding your scope of impact while staying hands-on. The transition to CTO is something qualitatively different: it's about learning to work entirely through people.

If you are still the best coder on your team as a CTO, something has gone wrong. Either you haven't hired well enough, or you haven't been able to let go of the thing that gave you your identity as an engineer. Both are real failure modes and I've seen them take down genuinely talented people. The job is to create an environment where great engineers can do great engineering. That is a different job than being a great engineer yourself, and you have to be honest with yourself about whether you actually want it.

Looking back, what would you have done differently?

I would have invested in observability earlier. Not monitoring in the 2010s sense — dashboards of infrastructure metrics — but genuine observability: distributed tracing, structured logging, the ability to ask arbitrary questions of the system without deploying new instrumentation. We were flying partially blind, and that Black Friday incident was partly a consequence of that. We knew the system was slow in certain conditions. We didn't know exactly where it was slow, or what the failure mode would look like under specific load patterns. Good observability would not have prevented the incident, but it would have dramatically reduced our time to diagnosis and recovery.

The other thing I'd change is how I ran the post-mortems. We did blameless post-mortems, which was the right instinct, but I let them become process artifacts rather than learning tools. We'd generate a list of action items, assign owners, track completion. What I didn't do was ensure the learning actually spread through the organization — that the engineer on team B understood what team A had discovered about database connection pool behavior under load. Knowledge needs to move. Post-mortems that stay in Confluence are just expensive documentation.