The What, When & How framework of feature testing
A short and practical guide that helps small teams build a simple testing system that prevents shipping critical bugs.
How it can help you
Every early-stage startup ships bugs. That doesn’t mean you’re doing something wrong, it's simply the consequence of a small team moving fast. Everyone is already doing three jobs, and dedicated testing resources are a luxury that comes much later, if at all. There are a lot of software testing models, theories and frameworks out there. The problem is that most of them are build for teams with proper testing resources, meaning they are an overkill for small, agile startup teams.
What is it based on?
The framework borrows useful elements from these established frameworks, but combines them in a way that makes is actually useful for small teams. The most prominent existing practices used include:
- Risk-Based Testing (RBT) by Stale Ringen
- The Testing Pyramid by Mike Cohn
- Shift-left (Larry Smith) & General Continuous Delivery practices
The What, When & How framework
For every feature you ship, ask yourself three questions in order:
- What needs the most testing effort? (Risk assessment)
- When should we test it — internally, or can we let customers find the bugs? (Test timing)
- How should we test it? (Test type)
The order matters. You start with risk because it determines how much effort to invest at all. Then you decide timing, because sometimes the right answer is to ship and observe rather than test exhaustively. And only then do you think about method, because if you decide to ship without heavy internal testing, the “how” question becomes moot.
For each of the three steps, we borrow from frameworks that already do their job well, we just sequence them in a way that works for a founding team without a testing function.
What to test: Start with risk
Not every feature is equally dangerous to ship with a bug. A bug in your payment flow is not the same as a bug in your dashboard date-filter. Risk-Based Testing (RBT) gives you a principled way to make that distinction.
RBT is a well-established testing approach that originated in software engineering and safety-critical industries. The core idea is straightforward: not all features carry equal risk, so testing effort should be allocated in proportion to risk rather than spread evenly across everything. Instead of asking “did we test this?”, RBT asks “did we test the things most likely to cause real harm if they break?” For a founding team with limited time, that shift in question alone is valuable.
You don’t need a complex implementation of it. What varies is how you assess risk. At the lighter end, this can be a quick judgment call — high, medium, or low. At the more structured end, you can create a scoring system and assign each feature a score. The important thing is to keep it fast. This should take a maximum of two minutes per feature, not twenty.
When to test: Internal first, or let the customer test?
Once you know the risk level, the next question is timing. In testing theory, this is known as the shift-left vs. shift-right decision. Shifting left means testing heavily before release, shifting right means releasing sooner and relying more on real-world exposure. But rather than a binary choice, think of it as a spectrum. The question is: Given your context, how much of the testing do you have to do internally, versus how much can you let your customers do the testing for you? Context means the specific situation your product and company is in. Are you selling to enterprises with low bug tolerance, or to early adopters who expect rough edges? Is your product mission-critical for your customers (something they depend on daily to run their business) or more of a nice-to-have? Do you have a trusted customer who's willing to beta test in exchange for early access? Those answers tell you where on the left-right spectrum you can place the given feature. This is both a company-level and a feature-level decision. At the company level, your context sets a general threshold — a dev tool selling to engineers can shift further right than a fintech product handling other people's money. At the feature level, you apply that threshold to the specific risk score from step one and decide where on the spectrum to land: heavy internal testing and a broad launch, a quick sanity check followed by a feature-flagged beta, or something in between.
How to test: The pyramid as a routing guide
If you've decided that internal testing is needed, the next question is what type of tests to use. This is where the Testing Pyramid comes in. Originally developed by Mike Cohn, the pyramid has three levels: automated unit and integration tests at the lower levels, which test the code itself, and manual end-to-end tests at the top, which simulate a real user moving through the product (though these can be automated too, it requires more effort and is less common at an early stage). The logic of the pyramid model is simple: the lower you test in the stack, the cheaper and faster it is to catch a problem. Given the risk level from step one and the release strategy from step two, the pyramid helps you decide what type of tests make most sense for this specific feature, what coverage is actually needed, and what can be automated vs. done manually.
Unit tests (bottom of the pyramid): Test individual functions or components in the code in isolation. Fast to write and run. For a billing feature, this might mean testing that your pricing calculation function returns the correct amount for a given set of inputs.
Integration tests (middle): Test how components work together. For the same billing feature, this might mean testing that your payment service correctly talks to your billing database and sends the right confirmation event.
End-to-end tests (top): Test the full user journey through the product. While it can be automated in some cases, it’s mostly done manually. It’s slow but essential for the most critical paths. For billing, this means simulating an actual customer going through checkout, including edge cases like failed cards or expired sessions.
Of course, good developers are already writing unit and integration tests. The purpose of this last step is not to make developers understand why these tests are needed, they already know that. The gap is a different one.
Developer-written tests tend to be developer-driven: they cover the code paths the developer is uncertain about, the edge cases they noticed while building, the things they want a safety net for. That’s valuable. But it’s not the same as customer-driven testing, asking which failures would hurt a real user most, and making sure those paths are covered.
The Testing Pyramid helps you make that shift. It’s not a checklist to run through exhaustively for every feature. Think of it as a quick prompt: have we thought about all three layers from the customer’s perspective, even if the answer for some layers is “not needed here”?
An Example
Let's say you're shipping a new feature that lets customers export their data as a CSV. What: The impact of a broken export is moderate — customers lose time, not data (2). CSV generation has some edge cases, so likelihood of something going wrong is medium (2). And if the export is broken, customers will notice immediately (1). Score: 2 × 2 × 1 = 4. Low risk. This feature doesn't need heavy testing attention. When: Your context is mid-market B2B customers who use exports for weekly reporting — so bug tolerance is moderate and the feature will be used regularly. But the impact of a bug is reversible (a bad CSV doesn't corrupt anything), and you have a beta customer who has agreed to test new features early. Call it: a quick internal sanity check, then release to your beta customer before a full rollout. How: Since you're doing some internal testing, run through the pyramid quickly from the customer's perspective. Are there customer-critical failures at each layer that your existing tests wouldn't catch? At the unit level: test the CSV generation logic for edge cases like empty datasets or special characters. At the integration level: confirm the export is triggered correctly and lands in the right format. At the E2E level: quick sanity check to confirm the basics are working, the data looks right, the rest will be done by your beta customer. Total decision time: five minutes. Total additional clarity: significant.