Osmos Global Publication · Osmos Perspective
Predictive Maintenance Must Beat a Fair Baseline
A model is valuable only if it improves the decision compared with a credible existing alternative.

Do not confuse novelty with improvement
Predictive-maintenance proposals often begin with an impressive model score. The decision-maker still needs to know what was predicted, how far in advance, which failures were included and what the existing maintenance approach would have achieved. A score without this comparison cannot establish whether the new system deserves investment or operating authority.
Hong and Li recommend explicit baseline comparisons and suitable validation for building-AI studies [HONG].
JLL’s survey analysis provides a separate reminder that technology programmes face business scrutiny [JLL].
Osmos translates these ideas into procurement: ask suppliers to demonstrate incremental decision value, not merely a functioning algorithm.
Define the event and the useful warning
A failure event needs a consistent definition. A work order, alarm, inspection finding and component replacement may refer to different stages of the same problem. Historical labels should distinguish planned replacement from breakdown, and genuine failure from an administrative closure. Otherwise the model may learn maintenance habits rather than equipment condition.
Lead time matters as much as detection. A warning that arrives after the operator already knows the equipment is failing adds little. A very early warning may also be unhelpful if the condition is too uncertain to justify intervention. Define the window in which staff can inspect, plan access and obtain parts.
Figure 1. Four checks before accepting a performance claim Original Osmos Global framework informed by Hong and Li (2026), §§3.1–3.5, CC BY 4.0. No source diagram reproduced. Prepared 1 September 2026.
What this means: A strong algorithm cannot compensate for an invalid comparison.
Choose a credible comparator
Compare the proposed model with the actual maintenance practice and, where feasible, a simple alternative such as an engineering threshold or a trend rule. Give the comparator appropriate tuning and the same information boundary. Deliberately weak baselines make a sophisticated model look better without helping the buyer.
Keep future information out of historical tests. Records created after a fault was confirmed can inadvertently reveal the answer. Time-based separation and independent review help reveal this problem. Document exactly what data would have been available when the model supposedly issued its warning.
Price the errors in operational terms
False alarms consume investigation time and can erode trust. Missed events can create disruption or safety exposure. Their consequences are not symmetrical, so an overall accuracy percentage is insufficient. Review performance by relevant asset type, failure mode and operating condition; some error categories may be unacceptable even when the average looks strong.
The economic comparison should include the added inspection effort, platform cost, integration, training and any intervention prompted too early. Avoid assigning a full replacement value to every alert. An avoided-cost claim requires a defensible explanation of what would otherwise have happened, supported by maintenance and finance reviewers.
Illustrative decision rehearsal
Suppose a supplier demonstrates a model that predicts replacement events from maintenance records. A reviewer notices that a field added when a replacement was authorised is included in the model input. The algorithm may be recognising an administrative decision already made, not predicting an impending failure.
This hypothetical example shows why a strong score can coexist with little operational value.
Reconstruct the timeline for a sample of events. Identify when each input became available, when staff first recognised the problem and when intervention was possible. Remove information that would not have existed at the claimed prediction time. Re-run the comparison before interpreting performance as a useful warning.
For the first operating cycle, ask engineers to review a bounded set of recommendations without changing the maintenance strategy automatically. Record whether each warning was new, relevant, timely and actionable. The review should include examples of known problems the model missed; examining only successful alerts creates a favourable but incomplete picture.
The final decision should distinguish analytical promise from demonstrated maintenance value. A model may deserve further testing even when the economic case is not yet established. It may also reveal that the organisation needs better failure labels before advanced modelling is worthwhile. Both conclusions are useful if they prevent the procurement team from paying for confidence that the evidence cannot support.
Earn the right to scale
Begin in an advisory mode with engineers reviewing output before it changes maintenance commitments.
Record rejected recommendations and the reasons. Extend the test into the conditions that matter, including seasonal modes and assets not used during model development. Expansion should follow evidence of transfer, not a successful demonstration on the easiest equipment.
Neither source establishes that predictive maintenance always outperforms preventive practice. Sparse failures, poor labels and changing asset conditions can make validation difficult. A useful outcome may be a better inspection process rather than a reliable failure forecast. A fair baseline preserves the organisation’s ability to choose that simpler result.
Source notes
[HONG] Tianzhen Hong and Han Li. Good practices for documenting AI-based studies on energy and buildings. Energy & Buildings / Elsevier; author copy hosted by Lawrence Berkeley National Laboratory, 2026-01-20. Sections 2, 3.1–3.6 and 4; pp. 1–4. DOI: 10.1016/j.enbuild.2026.117043. Accessed 1 September 2026. https://eta-publications.lbl.gov/sites/default/files/2026-06/1-s2.0-s0378778826001039-main.pdf [JLL] Yuehan Wang. Reality check: The true pace and payoffs of AI adoption in corporate real estate. JLL, 2025-10-27. Key highlights; AI pilot selection; Lessons learned. Accessed 1 September 2026. https://www.jll.com/en-hk/insights/global-real-estate-cre-technology-survey
Editorial and visual note
This is original Osmos Global analysis informed by the cited publications. Reported findings are distinguished from Osmos recommendations and illustrative scenarios. Source findings and trademarks remain attributable to their owners. Original visual designs do not imply endorsement by source organisations. The content is general research and does not replace site-specific professional advice.
Cite this
Osmos Global Research & Knowledge Centre (2026). Predictive Maintenance Must Beat a Fair Baseline. Osmos Perspective, Osmos Global. https://www.osmosglobal.org/articles/predictive-maintenance-must-beat-a-fair-baseline
Keep reading

Smart-Building Pilots Need an Operating Owner
An experiment becomes useful only when someone owns the decision it is meant to improve.
1 Sept 2026 · Osmos Global Research & Knowledge Centre · 5 min read

An Alert Is Not a Maintenance Outcome
Analytics creates value through verified correction, not the number of faults displayed.
1 Sept 2026 · Osmos Global Research & Knowledge Centre · 5 min read

Sensor Coverage Is Not Data Quality
Connected points need identities, context and a known level of trust before they can support decisions.
1 Sept 2026 · Osmos Global Research & Knowledge Centre · 5 min read
Download this paper
The full PDF, formatted for circulation. Downloads are for members, so that we know who our research reaches.
Discussion
Add what you are seeing on the ground.
