MSAI Blog | Insights on Predictive Maintenance and Industrial AI

Condition Monitoring Alert Fatigue Is a Design Problem, Not a Tuning Problem

Written by Séan Allen | October 06 2026

TL;DR

  • Tuning thresholds to cut false alarms works for a while, then starts hiding the faults the system was built to catch.
  • A threshold is one number standing in for a judgment call, and it can't see context like time of day or an asset's history. The trade-off is familiar: turn it down and drown, turn it up and miss.
  • Solar Reflection Analysis takes a different approach: it evaluates each ambiguous alert against its own context and attaches a confidence read-out that travels with the alert.
  • Confirmed reflection is logged and closed. Confirmed faults and uncertain events go straight to the team, nothing closed is deleted, and the operator still makes the call.
  • Reflection is a solar example of a broader pattern. Some alerts in any condition monitoring program will always be ambiguous, which makes alert fatigue a design problem, not a tuning problem.

Why Why False Alarms Keep Coming Back

Solar Reflection Analysis taught us something bigger than solar: The fix for too many false alarms was never a smarter threshold. 

Somewhere in every reliability program's history, someone was told the same thing: if your monitoring system throws too many false alarms, tune the thresholds. Raise the bar so the noise stops. 

It works, for a while. Then the same tuning that quieted the noise starts hiding the faults the system was built to catch, and nobody notices until one of them turns into an outage. This isn't a flaw in any particular vendor's software. It's a structural limit of treating a threshold as the only lever a monitoring system has for explaining itself. 

We ran into a sharp version of this problem in solar. Solar panel glass reflects heat specularly, meaning a warm nearby object or the sun sitting low in the sky can bounce a signature into a thermal camera that looks exactly like a developing fault. A fixed camera can't re-shoot the same spot from another angle to check, the way a technician on a drone flight can. So the system either flags every ambiguous case and floods the operator with noise, or it raises the threshold and risks missing a real fault along with it. 

Solar Reflection Analysis, which we included inside Solar Performance Monitoring earlier this year, is our answer to that specific problem. If you missed the launch, we wrote about the case that convinced us the fix belonged inside the system's alerting, not on a technician's dispatch list (link: How We Taught Solar Monitoring to Stop Mistaking Reflection for a Fault). 

But the more interesting result of building it wasn't the drop in false alarms. It was proof of a bigger point: The fix for alert fatigue isn't a smarter threshold. It's a system that can score its own confidence and show its work, so the operator decides with more information instead of less. 

Why thresholds were always the wrong lever 

A threshold is a single number standing in for a judgment call the system can't actually make: is this specific event, right now, likely enough to matter that a person should look at it? That judgment depends on context a threshold can't see, time of day, the asset's own history, and whether the same pattern has shown up before and meant nothing. Compress all of that into one number and you get the trade-off every reliability team already knows by heart: Turn it down and drown, turn it up and miss. 

Solar Reflection Analysis works differently: 

  • Every ambiguous alert gets evaluated against its own context, the time of day and the sun's position

  • That evaluation produces a confidence read-out that travels with the alert instead of replacing it

  • Confirmed reflection is logged and closed

  • Confirmed faults and uncertain events go straight to the team, and closed events aren't deleted, every alert stays logged and reviewable 

The operator still makes the call, just with a case built for them instead of a single dial to fight with. 


Why this is bigger than solar 

Reflection is a solar-specific example of a pattern that shows up everywhere condition monitoring gets deployed: Some fraction of alerts will always be genuinely ambiguous, and no amount of sensor quality fixes that, because the ambiguity isn't a sensing problem. It's an interpretation problem. Vibration analysts have their own version of this with resonance and looseness signatures that mimic each other. Thermal teams watching electrical gear have their own version, as well, with load-driven heating that looks identical to resistance-driven heating until someone checks the load curve. 

The industry's habit has been to treat interpretability as a tuning exercise, something a customer success team helps a customer dial in over their first few months. We think that ceiling is closer than most vendors admit, and the next real differentiator in this category won't be who has the most sensor types. It will be who can build a system that explains its own uncertainty well enough that a lean team can trust it without babysitting it. 

Where to go from here?

Solar Reflection Analysis is one proof point. We built it because the ambiguity on a solar array was too common to ignore, and the fix couldn't just be filtering. The same design principle, evaluate the context, score the confidence, keep the person in the decision, is the direction we intend to keep building in, on other asset types where a threshold alone has never been good enough.

If your team is still choosing between too much noise and missed faults on any of your monitored assets, that's not a tuning problem. It's a design problem, and it's worth asking your vendor how they're solving it.

What next? Talk to an engineer, or read the full Solar Performance Monitoring datasheet. 

FAQs 

Why doesn't raising the alert threshold fix false alarms?

It works for a while. Then the same tuning that quieted the noise starts hiding the faults the system was built to catch, and nobody notices until one of them turns into an outage.

What does it mean that alert fatigue is a design problem, not a tuning problem?

A threshold is a single number standing in for a judgment call. That judgment depends on context a threshold can't see, like time of day, the asset's own history, and whether the same pattern has meant nothing before. Fixing it means giving the system a way to explain itself, not nudging the number.

What is a confidence read-out?

It's a score showing how closely an ambiguous alert matches known behavior, such as reflection. It travels with the alert instead of replacing it, so the operator sees the evidence along with the flag.

Does Solar Reflection Analysis decide whether an alert is a real fault?

No. It evaluates each ambiguous alert against its own context, time of day and sun position, and attaches a confidence read-out. Confirmed faults and uncertain events go straight to the team, and the operator still makes the call, just with more information.

Are alerts closed as reflection deleted?

No. Confirmed reflection is logged and closed, and every alert, resolved or escalated, stays logged and reviewable in the account's history.

Is this only a solar problem?

No. Reflection is a solar-specific example of a broader pattern: some share of alerts in any condition monitoring program will be genuinely ambiguous, because the issue is interpretation, not sensing. Solar Reflection Analysis is where we've applied this approach so far.