Quality Assurance

    Quality has two clocks: first use and the long run

    Score initial quality and long-term dependability separately, with numeric failure thresholds that deadlines cannot argue away.

    'Good enough to ship' is usually decided by opinion and deadline pressure. This framework replaces opinion with two scored thresholds: initial quality, covering the first days of use, and dependability, covering stability and relevance over the product's life. Retention lives in the second score.

    Ask a room whether a release is good enough and you get opinions, calibrated to the deadline. This framework replaces the debate with two numbers and a rule: below the threshold, the release is a failure. Initial quality measures the first moments and days of use. Can people use it easily, complete their task, and understand what they see? Dependability measures the long run: performance stability, continued relevance to customer tasks, and whether known problems are quietly compounding into CX debt.

    Why it matters to the business

    The two clocks pay differently. Initial quality decides adoption; dependability decides retention, and retention is where the money is. Research by Peter Kriss published in Harvard Business Review found customers with the best past experiences spent 140% more than those with the worst, and in a subscription business a great experience lifted one-year retention from 43% to 74%. The downside is equally sharp: PwC found 32% of customers will leave a brand they love after one bad experience. A release that scores well on day one and erodes by month six converts marketing spend into churn.

    How to use it

    • Set a numeric failure threshold for each score before the release date is set, so the bar cannot move under pressure.
    • Measure initial quality in the first days through observed tasks, not opinions: ease of use, task completion, clarity.
    • Measure dependability quarterly: stability, continued task relevance, and the trend of unresolved CX debt.
    • Correlate score drops with business metrics. A fall in initial quality showing up in call-center resolution time is an argument no roadmap meeting can wave away.
    • Include accessibility in both thresholds; retrofitting costs multiples of building it in.

    Where teams get it wrong

    Teams measure quality once, at launch, then move on. Dependability erodes silently, a slow screen here, a stale workflow there, until churn shows up two quarters later with no obvious cause. Averaging both clocks into a single score hides the erosion; the launch glow subsidizes the decay.

    Ask your team

    • What numeric score would make our next release a failure, and who set it?
    • How has dependability trended over the last four quarters for our core product?
    • Which business metric moves when our quality score drops, and have we checked?

    Set a failing threshold for each phase, not a vague aspiration.

    Apply this

    Reading about quality has two clocks: first use and the long run is one thing. Seeing where it applies in your journey is the useful part.

    Related signals