ZillaAI

Writing

Not asked is not broken

When a status signal reports failure without a question, it trains operators to ignore the signal. We updated the control logic to distinguish between silence and failure.

Published · written with AI assistance

A status signal that stays red when nothing is wrong trains the person watching it to ignore it. This is not a theory. It is an operational hazard. When the system cries wolf quietly for weeks, the one real failure arrives unread.

On the date stated above, the monitoring surface reported one of its own primary inputs as failing. It had been reporting this state on every check for weeks. The input was fine. It was carrying tens of thousands of records an hour. The same response that declared the input degraded also carried the count proving it was not.

The discrepancy was not in the traffic. It was in the definition of health.

The logic of absence

The cause was a single function within the status layer. Two optional health probes had never been configured. Their addresses were simply absent from the configuration store. The code returned the same "not healthy" value for "never asked" as it did for "asked and it failed".

A question nobody had put was being reported as an answer of bad news.

This is a failure of the control logic, not the infrastructure. The control is responsible for interpreting the state of the system. When the control cannot distinguish between a missing probe and a broken probe, it loses its authority. The surface was lying quietly. It looked exactly like a surface telling the truth, but it was not.

The same mistake was found twice more on the same day.

An analysis component reported its own processing as degraded whenever it had correctly decided a result needed further confirmation. Doing its job was being counted as a symptom. The component was waiting for data it did not yet possess, and the status layer interpreted that wait as a stall.

A work queue presented items for human approval with every identifying field empty. The records held the values the whole time, but the layer assembling the queue looked for them one level too high and found nothing. Absence looked like a quiet system rather than a broken one.

In each case, the system was functioning. The data was present. The status was wrong.

Interception through state separation

The fix changed the shape of the status model. "Not configured" became its own state, distinct from "failing". The gap is still reported. An unasked question is a finding, not a zero. It simply no longer claims that something healthy is sick.

This change is the interception. The control logic now catches the ambiguity of missing data before it propagates as an incident. By separating the states, the system prevents the false signal from triggering an operational response where none is required.

The value of this change is not in the bug fix. It is in the preservation of trust. Operators rely on the dashboard to tell them where to look. If the dashboard points to a healthy system as broken, the operator spends time investigating nothing. That is waste. If the dashboard points to nothing, the operator looks at the right place when the real failure happens.

The expensive part was never the bug. It was that the surface had been lying quietly for weeks.

The cost of noise

Alarm fatigue is a measurable degradation of response time. When a team sees a red alert that turns out to be a configuration gap, the next red alert is treated with suspicion. The cost of that suspicion is time. It is the time spent verifying the alert before acting.

By making the report honest, we reduce the noise. The status layer now reports the truth: the probes are not active. It does not claim the service is down. It claims the check is incomplete. This distinction allows the operator to prioritise the work. They can ignore the missing probe if the traffic is flowing. They can address the missing probe if the traffic is not.

The control has caught the lie. The system is no longer broken; it is merely incomplete. There is a difference.

Honest gaps

There is work that remains. The two probes are still not configured. The change made the report honest, not the coverage complete. Saying "we have not checked" out loud is the beginning of that work, not the end of it.

We could have configured the probes immediately to silence the warning. We chose not to. A green light for an unconfigured probe would have hidden the gap again. The goal is not a dashboard that is always green. The goal is a dashboard that is always true.

Until the probes are active, the status will show the gap. That is the intended behaviour. It is better to see the gap than to see a lie.

The safety net is not the absence of errors. It is the ability to identify them before they become incidents. This update ensures that silence is not mistaken for failure. The control logic now intercepts the assumption that unasked questions are broken answers.

We are demonstrating the safety net, not the fall. The system is stable. The report is accurate. The work continues.

Review

Published, and not yet reviewed by a human. This note updates when it has been.

← All writing