Skip to content
Requirements Dialogue
Part 4

What eight months of building alone teach that no theory can say

A framework on paper is convincing as long as nobody tries to actually build with it. What happens when you run your own methodology against a real project for eight months — and what actually went wrong along the way.


The first three articles in this series were theory: why Scrum is reaching its end, why SAFe solves the wrong problem, which three axes really distinguish one method from another. This article is no longer a framework. It is a field report.

The difference between an idea and a tool that is tested on itself

A methodology that exists only on paper can sound as convincing as you like. The real test only begins once you apply it to an actual project — and not to a sample project built specially for illustration, but to one whose success genuinely depends on the methodology delivering what it promises.

That has been my daily work for eight months: a software project that comes about by its own rules while those rules are still being sharpened. No team to help me place things in context. No colleagues to talk the approaches through with in detail, because the subject is still too new for most people. One person, one AI partner, and a project that has to be honest enough not to hide its own mistakes.

Three mistakes the theory would not have predicted

After 48 requirements worked through in one long-running session, three real errors occurred in quick succession: a fabricated source reference, a wrongly attributed rationale, and a logical error that made no sense on closer inspection. None of these was dramatic on its own. Taken together they revealed a pattern: beyond a certain context length, reliability drops noticeably, even though none of it becomes visible from the outside as long as nobody goes looking for it.

The genuinely interesting part is not that these errors happened. Errors happen, with or without AI, with or without a methodology. The interesting part is that a second instance, checking independently and working in a fresh context, caught all three — not because it was more disciplined, but because it structurally had no prior knowledge for an error to attach itself to. That is the difference between "we check carefully" and "a check that, by construction, cannot be anything other than unbiased".

Why a fixed rhythm mattered more than expected

One of the practical lessons from these eight months: having a good checking instance is not enough. You also need a fixed point at which a long, productive working session is deliberately ended and restarted — not only once an error has already surfaced, but preventively, after a fixed number of items worked through. In practice this rhythm proved itself both reactively, in response to errors found, and preventively, without anything having gone wrong beforehand. Both were necessary. A good checking instance on its own would have found the errors, but only after they had already happened several times.

What this means for the methodology — and what it does not

These errors are not an argument against the approach. They are the reason the approach is credible at all. A methodology claiming never to hit its own limits would not be serious. A methodology that documents its own limits openly and shows that its own checking mechanism reliably catches exactly those limits is something else — it proves itself at the point where it actually counts.

What these eight months do not prove: that the approach works equally well with several real people holding different and genuinely conflicting interests, rather than a single decision-maker. That is an open question, not one already answered.

The next article is about what actually comes out of these eight months: a tool that does not merely describe the methodology but makes it checkable against itself.

Back to the overview