The dashboard shows every control in green. In the same quarter, a leaver's account was used three months after their last working day. Both statements are true, and that is exactly the problem: "implemented" and "effective" are two different claims, and most programmes only ever collect the first.
Coverage is a statement about the list
An implementation percentage answers whether a control exists. It leaves three other questions untouched: whether it actually runs in production, whether it reaches everything it is supposed to reach, and whether it finds or prevents what it was put in place for.
The distinction is not academic. Access revocation on exit can be documented as a procedure, approved, and configured in the tooling, and still fail to reach three of twenty systems, because those systems were never connected to the process. On the list it is the same control. On the dashboard it is the same colour.
Three questions a control has to answer
| Question | What it tests | Typical evidence |
|---|---|---|
| Does it run? | existence in live operation | log, system extract, operating record covering a period |
| Does it cover the scope? | coverage | reconciliation between the inventory and the objects actually captured |
| Does it work? | outcome | test, sample, exercise, analysis of real events |
Most programmes stop after the first row, and the reason is obvious: it is the only one answerable from a document. The second needs a defensible inventory. The third needs effort that shows up in the annual plan.
The second row is the most underrated of the three. Coverage is not a question of diligence but of arithmetic: a control reaching eighty per cent of the estate is not eighty per cent effective; the remaining fifth is entirely unprotected, and that is where an attacker looks.
The difference between a record and a test
A record documents that something happened. A test creates a situation the control has to survive. The whole difference lies between them.
A backup is not tested by a backup report. It is tested by a restore, onto a system resembling the target environment, with the elapsed time measured. An access review is not tested by the existence of a signed list; it is tested by planting a known, deliberately wrong entitlement in the review population and observing whether it is caught. A leaver process is tested by picking three departures from last quarter and checking each individual system. A reporting chain is not tested by an organisation chart but by an exercise with a clock running.
All four are cheaper than their reputation. What makes them look expensive is not the effort. It is the risk that they find something.
The statute says "effective"
Under § 30(1) BSIG, besonders wichtige Einrichtungen and wichtige Einrichtungen are required to take "appropriate, proportionate and effective technical and organisational measures". Effectiveness is therefore not a maturity attribute and not a nice-to-have. It is an element of the duty itself. A control that exists but does not work does not satisfy the wording.
The catalogue of measures goes a step further: § 30(2) sentence 2 BSIG names, at number 6, concepts for assessing the effectiveness of risk-management measures, a control about controls. That entry corresponds to point (f) of Article 21(2) of Directive (EU) 2022/2555; the ten numbered items of sentence 2 map one-to-one and in the same order onto points (a) to (j), so a cross-reference between the two texts always has to be converted.
The rank of that list matters more than its content. Sentence 2 is not a set of areas you may orient yourself against; it is a binding minimum, because the measures have to cover at least what it names. Assessing effectiveness is therefore not the finishing touch on a mature programme but one of the items without which the catalogue is not met.
ISO/IEC 27001 requires, in clause 9.1, that you determine what is monitored and measured, by what methods, when and by whom, and that the results are evaluated. The management review then processes exactly those results. Here too the subject matter is not the state of implementation but the effect.
Three instruments, the same demand. What satisfies it is not a status field.
Why the green report is so persistent
It is worth naming the causes plainly, because they are structural rather than personal.
Coverage is cheap to produce; effect is expensive. The person who introduced a control usually also reports on it, an arrangement in which a red cell is a statement about their own work. A red cell also demands an explanation, and no reporting cycle has time budgeted for explanations. And historically, assessments accepted documents, so documents were produced.
The counter-measure is not exhortation but separation: whoever operates a control should not be the sole judge of whether it works. Internal audit exists precisely for that, and so does the thinking behind the three lines model. Where separation is not organisationally possible (in a small organisation it frequently is not), an occasional external assessment substitutes for it far better than a self-assessment made in good faith.
Where to start when you cannot test everything
Nobody tests every control. The selection should therefore follow a clear rule: test first the controls whose failure you would have to report. That is usually fewer than a dozen: access protection on the externally reachable entry points, recoverability, separation of administrative accounts, logging at the points where an incident would become visible at all, and the reporting chain itself.
Everything else may remain, for now, a statement about coverage, and should say so explicitly. A programme that honestly tests six controls and transparently reports only implementation status for the rest is more defensible than one that reports green everywhere.
What changes in the report
One column becomes three: implementation, coverage, and last evidence of effectiveness with a date and a method. The third column is uncomfortable because it exposes an age nobody previously knew. That is exactly why it is useful. A cell that has never carried a date is the first finding, and more precise than any maturity score. Holding controls, evidence and the age of that evidence in one place is the subject of the compliance management page.
How completeness and effectiveness can be told apart
Tools rarely model the difference, because both statements have the same shape: a tick. In Rizzqo a computation rule separates them. An object counts as fully covered only once every attached requirement has been closed as fulfilled or as not applicable, and an answer marked fulfilled but not yet finalised still counts as a gap. The finalisation is the point at which somebody owns the statement by name and timestamp, and it demands a reason and an attachment.
Two further properties keep the number honest. The object's own workflow state does not enter the calculation, and neither does an approved risk assessment: a consciously carried risk documents a gap and names who carries it, but does not make the object compliant. What is not implemented stays visible instead of disappearing into the metric.