Skip to content
Salun MarvinBook a call
← Blog9 October 2026

How to triage tech that keeps breaking

When your startup's tech keeps breaking and slowing growth, this triage playbook pinpoints the root cause and restores stability.

[ SHORT ANSWER ]

Start with triage, not tooling. For every incident that repeats, ask why it was possible, why it was not detected sooner, why recovery took as long as it did, and what stops it happening again. The answers show whether the cause is infrastructure, process or ownership. Each has a different fix, and a rewrite or a new hire fixes none of them by default.

Most founders treat a broken system like a technical problem. At the root it almost never is. The real problem is diagnostic: you are trying to fix something you have not properly identified yet. The instinct to act fast makes sense, but acting before you understand the source of the failure is how you end up paying for new engineers, a rewrite or a cloud migration and finding yourself in the same place a few months later.

This is the playbook I run through before touching any infrastructure. No technical background is required. It is a system a founder can use, not an engineering deep dive. Steps you can run alone are marked throughout, and steps that need an artifact from your engineering team are noted where they appear.

The real reason your startup's tech keeps breaking

Recurring outages are rarely a pure technology problem. They typically fall into one of three failure categories: infrastructure fragility, process debt, or organizational misalignment. Most founders skip straight to "we need better engineers" or "we need to rewrite the backend" without running any triage. That is expensive and usually wrong. Each failure type has a distinct fingerprint, and each needs a different fix.

Infrastructure failures are sudden, environment-specific and often correlated with traffic spikes or deployment events. Process debt shows up differently: recurring incidents with no postmortem, repeated symptoms, a team that fixes the same bug three times in a row. People and role problems surface as slow recovery, unclear ownership, and engineers who escalate every decision upward rather than resolving it. I look at the production side of this in why startup tech stacks fail in production.

The four questions to ask about every recurring incident

For every incident that happens more than once, run this sequence. Why was the failure possible? Why was it not detected sooner? Why did recovery take as long as it did? What prevents it from happening again? If you keep getting the same answers across different incidents, such as missing tests, unclear ownership or manual steps with no runbook, the problem is systemic. Your engineering team should be able to pull incident timelines to answer each question.

One incident is a bug. Three identical incidents are a system. When the same failure fingerprint appears repeatedly, you are no longer dealing with bad luck. You are dealing with a structural condition that individual fixes will not resolve.

Red flags the problem is organizational, not technical

Watch for these behavioral signals as a non-technical founder. Estimates that are consistently off by a wide margin. Engineers who cannot clearly explain what they are working on this week. Incidents that get "fixed" but recur soon after. A delivery backlog that never gets shorter however many engineers join the team. These are team health signals, not infrastructure signals, and they need a different response than swapping out your database or upgrading your cloud tier.

Short-term fixes that stop the bleeding

Once you have established that infrastructure fragility is a real contributor, your engineering team should have a clear priority order for reliability work. Knowing that order lets you check the team is working on the right things.

The highest-impact reliability changes

Explicit timeouts come first. Without them, a single slow dependency can hang requests, exhaust threads and bring the whole system down. Next come bounded retries with exponential backoff, which recover from brief failures without amplifying them into a retry storm. Circuit breakers follow, stopping repeated calls to a broken dependency so a partial outage does not cascade into a total one. Rate limiting and load shedding protect the system from its own traffic spikes. Finally, caching stable reads keeps high-traffic paths available when dependencies are slow or down.

A focused sprint on these five patterns, applied to the most failure-prone dependencies in your system, can reduce recurring incidents, particularly when those incidents trace back to dependency timeouts and retry behavior rather than core application logic. Frame each change in business terms: what failure mode it prevents, not how it works internally.

Not sure which dependencies to target first? That is a good thing to sort out before the sprint starts. Book a call.

Common "quick fixes" that make things worse

Unbounded retries feel like resilience but turn a partial outage into a full one. A single global circuit breaker that covers unrelated services ties them together unnecessarily, so a failure in one triggers protective shutdowns across parts of the system that were fine.

Caching everything, including data that changes frequently, hides freshness problems and can serve incorrect information to users. Adding monitoring dashboards with no actionable alerts attached is the most common trap of all. It creates the illusion of observability while producing nothing that helps anyone respond faster. These anti-patterns feel like progress and often produce a short lull before making the next incident worse. Know what to push back on when your team proposes them.

What a triage with an outside diagnostician looks like

Here is the general pattern I see. A founder arrives with a roadmap that has slipped, a team that has grown, and delivery that got slower after the hiring rather than faster. Incidents are more frequent. The founder suspects a people problem but cannot tell whether it is the team lead, the architecture, the process or some combination. That uncertainty, where the problem is real but the source is unclear, is where an outside read earns its place.

What the initial diagnosis tends to reveal

A team diagnosis works by observing how work flows from idea to deployment, identifying where decisions stall, reviewing incident patterns, and sitting in on planning and standup rituals. Common findings are several engineers working on the same problem without knowing it, critical services with no clear owner, and a team lead whose calendar is so full of meetings that there is no time to review code or unblock anyone. No amount of new hires fixes any of that.

The role audit and lean process changes

A role audit maps actual responsibilities against job titles, identifies redundant seats and missing positions, and produces a plain-language report the founder can act on. The process recommendations are minimal on purpose. A single weekly async status update can replace several recurring meetings. An incident ownership rule means every incident has one named engineer responsible for resolution and the postmortem. A lightweight deployment checklist replaces the informal "wing it and see" approach.

The fix is organizational before it is technical, and no tooling change produces the same result. How much changes depends on the root causes actually present, but when the diagnosis points to ownership gaps and process debt, targeted process changes tend to outperform infrastructure fixes.

Metrics and questions every founder should track

You do not need a full engineering dashboard to know if your team is performing. You need five numbers and the ability to ask the right questions. Tracked weekly, these give a reliable picture of system stability and team health without reading a line of code.

  • Mean time to detect: how long before the team knows something is broken
  • Mean time to recover: how long to get back to normal after an incident
  • Deployment frequency: how often code ships to production
  • Change failure rate: what share of deployments cause an incident
  • Incident recurrence rate: what share of incidents repeat a prior one

If incident recurrence is high and deployment frequency is low, you have a systemic reliability problem, not a staffing problem. If time to recover is long and change failure rate is high, your team lacks both confidence in its own system and the ability to get back to normal quickly. Trends over several weeks tell you more than any single data point. For how these compare, see DORA metrics vs cycle time.

Questions to ask your engineers that get honest answers

Ask directly. "What is the most fragile part of the system right now?" "What would you fix first if you had a free week?" "Who owns the service that caused the last incident?" "When did we last run a postmortem, and what changed afterward?" These reveal ownership gaps, morale problems and process failures faster than any metric.

Engineers who cannot answer clearly are pointing to a problem with clarity and structure. Engineers who have strong answers but are not being heard are pointing to a leadership problem, often a more urgent one. Both are worth acting on.

How to prioritize fixes without freezing your roadmap

The most common mistake is letting reliability work compete on equal terms with feature development in the same queue. It almost always loses. Reliability work has no reach, no user growth and no visible output, so scoring systems like RICE undervalue it unless it is framed correctly. Reliability is risk reduction and avoided cost, not a feature with zero business value.

Apply a "must-do" gate first. Anything blocking a contractual obligation, preventing a severe outage or creating compliance risk gets done regardless of backlog ranking. For everything else, score reliability work by expected loss avoided: probability of failure multiplied by impact, divided by effort. Score feature work with a simplified RICE model. The two queues need different scoring logic.

Reserve a fixed share of engineering capacity for reliability, maintenance and technical debt, and do not let it flex downward to meet a feature deadline. If reliability work keeps consuming more than that share, stop treating it as a backlog item. It is a systemic issue that needs a different response. Teams with frequent outages may need to run higher temporarily until the system stabilizes.

When to refactor, rewrite, or bring in outside eyes

Most startups that talk about "the rewrite" are dealing with an organizational problem a rewrite will not fix. A rewrite under the same team structure, with the same unclear ownership and the same absent incident process, produces a different codebase and the same incidents. Incremental refactoring is almost always the right starting point. Some signals do indicate a deeper structural fix is warranted.

Signals that point to a structural fix, not a process one

Delivery has not improved after a sustained stretch of process changes. The same engineers are bottlenecks on every feature. A single service failure brings down unrelated parts of the product. The team cannot deploy confidently without a manual verification ritual. New engineers take months to ship anything. Any two of these together warrant a deeper diagnostic. All five together warrant a restructure conversation. I cover that decision in repair or rebuild.

The minimum observability setup that changes everything

Start with three components. An external uptime check that runs independently of the application's own infrastructure, so you hear about outages before your users tell you. Application-level error grouping so engineers move from alert to root cause without hunting through logs. A central place to correlate metrics, logs and deployment events. UptimeRobot, Sentry and Grafana Cloud cover this at low cost with minimal setup. Better Stack for uptime monitoring and the Prometheus, Grafana and Loki stack for self-hosting are worth evaluating too.

The goal is actionable alerts that go to a person with clear ownership, not a dashboard everyone watches and nobody acts on. That matters more than which tools you choose.

Getting an outside read before a costly decision

If you are weighing an expensive decision, such as restructuring the team, replacing leadership or authorizing a rewrite, one clear-eyed outside perspective is worth having before you act. I offer a free 30-minute call. No pitch. We work out whether outside help would actually fit your situation, and if so what a targeted version would look like. Book it before the next costly decision, not after.

Triage first, then fix

When your startup's tech is constantly breaking and slowing down growth, the answer is not a bigger team or a faster rewrite. It is a structured diagnosis. Identify the failure source. Apply targeted quick fixes in the right order. Track five numbers. Ask your engineers the honest questions. Protect reliability time from feature pressure.

Predictable delivery depends less on how many engineers you have than on how clearly the team understands what it is responsible for and how fast it learns from failure.

[ FAQ ]

Questions founders ask

Why does my startup's tech keep breaking?
Recurring outages usually trace back to one of three things: fragile infrastructure, process debt, or unclear ownership. Founders tend to jump to a rewrite or more hires before finding out which one it is. Triage first, because each cause needs a different fix.
How can I tell if the problem is the team or the technology?
Look at the pattern. Infrastructure failures are sudden and tied to traffic or deployments. Team and process problems show up as incidents that recur, no postmortems, unclear owners and estimates that are consistently wrong. One incident is a bug, three identical ones are a system.
Which metrics should a non-technical founder track?
Five are enough: mean time to detect, mean time to recover, deployment frequency, change failure rate and incident recurrence rate. Track them weekly and watch the trend rather than any single number. A high recurrence rate with low deployment frequency points to a systemic reliability problem.
How do I fix reliability without freezing the roadmap?
Give reliability work its own queue and score it by the loss it avoids, not by reach. Anything that blocks a contract, risks a severe outage or creates compliance risk gets done regardless of ranking. Protect a fixed share of engineering capacity for it so it does not lose every planning meeting to features.
When should I rewrite instead of fix?
Rarely as a first move. A rewrite under the same ownership gaps and the same missing incident process produces a new codebase with the same incidents. Consider a structural fix when process changes have not helped over a sustained stretch, the same engineers block every feature, or one service failure takes down unrelated parts of the product.

[ THE OFFER ]

Still trying to figure out if the problem is the team or the tech? That's the call. Book it.