Osmos
Research

Osmos Global Publication · Osmos White Paper

Beyond the SLA

Designing FM contracts for reliability, experience and long-term value

Osmos Global Research & Knowledge Centre19 min readSign in to download

Executive summary

FM contracts often contain extensive service levels, reports and remedies, yet still leave a decisive question unanswered: what evidence should convince a client that an obligation has actually been delivered? A score can be mathematically compliant while repeated disruption, temporary restoration, weak asset data or unresolved user impact remains outside the reported result. This white paper develops an original Osmos Global obligation-to-proof model. It connects five management acts: define the outcome, specify acceptable evidence, assign acceptance authority, investigate exceptions and keep authorised changes traceable. The model is informed by 2025–2026 sources including UK government contract-management guidance, an Australian public-sector audit, a Malaysian qualitative study of FM contract documents, and limited industry context. These sources differ in purpose and cannot be combined into a universal failure rate. The strongest practical conclusion is that contract assurance should be designed before mobilisation. Tender documents need evidence schedules; service clocks need explicit definitions; inherited defects need disposition decisions; monthly reviews need decision records; and exit clauses need retrieval tests. Service credits can support accountability, but they do not create recovery capability or repair an unusable evidence chain. For India and GCC environments, the model is presented as an application framework rather than a claim that UK or Australian public-sector rules apply. Leaders should adapt the controls to the governing contract, local law, asset risk and retained operating capability. The objective is not more reporting. It is a smaller set of evidence that can withstand challenge and lead to timely, accountable decisions. Acceptance remains a shared operating discipline, not merely a reporting task.

Reader’s route

Start with the evidence boundaries, then use the obligation-to-proof model to redesign tender evidence, mobilisation acceptance, service measurement, recovery, change control and exit assurance.

1. The problem behind a green scorecard

An SLA defines how a service result will be measured. It does not, by itself, guarantee that the measure captures reliability, user impact, lifecycle risk or the quality of recovery. A target can be met because attendance was quick even when restoration was temporary. Portfolio averages can be green while one location repeatedly fails. A closed work order can record completed technician activity while the underlying condition remains.

The resulting problem is not necessarily inaccurate data. It is incomplete decision evidence. Contract governance becomes weak when leaders treat a reported score as the end of inquiry rather than a summary that may require assurance. The appropriate level of inquiry depends on consequence, recurrence, uncertainty and the decision being made.

Package 6 addressed the FM sourcing and governance model. Package 10 deliberately moves closer to execution. It asks how parties define acceptable proof, how they decide whether to accept delivery, how they preserve exceptions and how they keep the baseline current when operations change.

Figure 1. The obligation-to-proof model Original Osmos Global conceptual framework, 2026. No measured dataset. Prepared 1 September 2026.

What this means: A reported result becomes decision-ready only when the proof and acceptance route are explicit.

2. Evidence reviewed and methodology

Osmos Global reviewed sources published within the programme window of 1 April 2025 to 26 August 2026.

The core evidence comprises the UK Cabinet Office’s March 2026 Contract Management Playbook; the Australian National Audit Office’s March 2026 audit of procurement and contract management in the DFAT Security Enhancement Program; and a June 2026 Planning Malaysia qualitative study of facilities-management and maintenance contract documents for Malaysian office buildings.

The sources are not equivalent. The UK publication is practice guidance, not an impact study. The Australian audit is a bounded examination of 14 engagements involving mixed security works and services. [2, paras. 1.17–1.18] The Malaysian paper analyses one public contract form and uses six purposively selected participants, five of whom are contractors. [3, pp. 108–109] JLL and CBRE provide industry context but do not replace the primary documents.

Osmos Global extracted only selected claims that could be traced to the issuing page or original document.

We did not treat source examples as global benchmarks, combine incompatible samples or reproduce source figures. The frameworks and operating tools in this paper are Osmos analysis. They should be tested against the actual contract, law, site condition and professional requirements.

3. What the evidence supports

The UK playbook supports a lifecycle view: define responsibilities, track obligations, maintain accurate records, use multiple performance-information sources, manage underperformance, control change and plan exit. These practices provide a credible structure for contract management, but the publication does not quantify the improvement an organisation will achieve by adopting them. [1, pp. 45–49, 53–55, 67] The ANAO audit offers a case-specific warning about relying on formal contract existence without effective delivery control. Four of the 14 engagements examined achieved the original scope, timeframe and estimated cost together. The audit methodology included records review, a sample of contracted engagements, contract-database analysis and procurement information. This is evidence about one programme. It is not evidence that 10 out of every 14 FM contracts fail. [2, summary paras. 11, 19; paras.

1.17–1.18]

The Malaysian study connects implementation challenges to tender instructions, contract conditions and appendices. Reported issues include unclear information, missing or ambiguous provisions, variation practices, incomplete technical documentation, inconsistent work-order practices and asset-data gaps.

Because the sample is small and contractor-heavy, the findings should be used to frame assurance questions rather than estimate prevalence. [3, pp. 110–114]

4. The obligation-to-proof model

The first act is definition. An obligation should state the operating outcome with enough clarity that client and provider teams can recognise acceptable delivery. Activities and frequencies remain necessary, but they should be connected to the condition or decision they are intended to support.

The second act is evidence design. Parties should agree the minimum record, source, timing, retention and access needed to support acceptance. Evidence requirements should be proportionate: a safety-critical test and a routine low-risk task do not need identical proof.

The third act is authority. Someone must be able to accept, reject or conditionally accept the evidence.

Responsibility matrices help only when the named role has the information and authority to act.

The fourth act is exception handling. Missing records, disputed definitions, recurrence and failed verification should remain visible until resolved or explicitly accepted by an authorised person.

The fifth act is traceability. Approved variations must update scope, evidence, price and performance assumptions. Otherwise, reporting continues against a baseline that operations no longer use.

Figure 2. Keep the contract baseline traceable Original Osmos Global conceptual framework, 2026. No empirical measurements. Prepared 1 September 2026.

What this means: An authorised change is incomplete until evidence and reporting reflect the new baseline.

Worked application: a service obligation under pressure

Consider a hypothetical cooling service serving a business-critical workspace. The contractual activity might require inspection and response, while the operating outcome is a usable environment within the agreed service boundary. An obligation-to-proof record would connect the relevant asset identifier, the event history, the action taken and the verification required for acceptance. It would not declare that every closed ticket proves reliable cooling, nor would it introduce a new technical tolerance without the appropriate specification and authority.

The evidence schedule could distinguish an inspection record from an incident record. An inspection would identify the asset, method, observed condition and exception route. An incident would preserve notification, containment, restoration and any continuing restriction. If a temporary reset returns service, the record should identify it as temporary and link the permanent corrective action. The acceptance owner can then distinguish an operationally useful interim result from final correction. This illustration is an Osmos design example, not a reported finding or a model contract clause.

Now introduce a landlord-controlled dependency. The provider may be able to confirm local temperatures and inspect tenant equipment but need building-management approval to investigate central plant. The record should show when the request was made, who received it and what interim control applies. A contractual pause may affect the response calculation, but it should not erase the continuing business impact. Commercial measurement and operational consequence remain related but separate views of the same event.

The monthly review should turn that record into a decision. If the fault recurs, the participants need to decide whether further diagnosis, a changed operating arrangement or an asset intervention is appropriate.

The decision owner should state the evidence relied upon, the uncertainty remaining and the test that will show whether the action worked. A later reviewer should be able to understand why a temporary solution was accepted without assuming that temporary acceptance established permanent adequacy.

This example also exposes the cost of assurance. Collecting a record is useful only if somebody can interpret it and act. Before adding photographs, measurements or approval layers, ask what decision each item changes. Use existing reliable records wherever possible, automate stable identifiers rather than judgement, and reserve deeper review for consequential or recurring conditions. If a new requirement takes frontline capacity away from essential work, redesign the workflow or resource it explicitly rather than concealing the burden.

The same logic can be applied to a different service without copying the engineering evidence. Cleaning acceptance may depend on the specified condition and an inspection route; workplace support may require confirmation that the request was resolved; a specialist test may require a competent person’s record. The common architecture is outcome, proof, authority, exception and traceability. The technical content remains service-specific. That distinction makes the framework portable without presenting one checklist as a universal standard for every FM obligation.

5. Design assurance before contract award

An evidence schedule should accompany material tender obligations. For each obligation it identifies the expected record, creator, custodian, submission timing, acceptance test, retention requirement and decision owner. This makes bids more comparable because providers price the evidence workflow rather than discover it after award.

Staffing proposals should also be tested against operating scenarios. The purpose is not to impose a preferred headcount. It is to expose assumptions about simultaneous calls, leave, travel, specialist access, subcontractors and escalation authority. A bidder may propose technology, mobile coverage or shared expertise; the evaluation should reveal how that model protects the required outcome.

Bid evaluation should separate base compliance from alternatives. Where a provider proposes a different method, the organisation should assess lifecycle value, evidence quality and risk without corrupting the common comparison. Clarifications and assumptions need a durable record because they often become the first source of disagreement during mobilisation.

Figure 3. A mobilisation or recovery decision gate Original Osmos Global conceptual framework, 2026. Not a validated scoring model. Prepared 1 September 2026.

What this means: Calendar pressure does not remove the need to own unresolved risk.

6. Mobilisation as an acceptance decision

Mobilisation is where contractual intention meets asset condition, systems, people and inherited data. A checklist alone does not establish readiness. The final gate should evaluate critical evidence and result in an explicit go, conditional-go or hold decision.

Inherited defects require two parallel decisions. The operational disposition determines immediate mitigation, inspection, maintenance, investigation, replacement or monitored acceptance. The commercial process determines who owns the cost and whether the condition sits inside the agreed baseline.

Disagreement about the second question should not suspend necessary action on the first.

Conditional acceptance can be responsible when the affected scope is understood, fallback controls work and a named owner has a credible completion date. It becomes unsafe when conditions are vague, high-consequence controls are unresolved or deadlines continually roll forward. Every condition needs evidence that will close it.

7. Measure response without distorting reality

Response performance contains several clocks: notification, acknowledgement, attendance, containment, restoration, permanent correction and verification. Collapsing them into one target invites disputes and can reward temporary activity. The contract should define which events matter for each service and which system or actor creates the timestamp.

Pause rules require equal discipline. A status such as “awaiting client” should identify the decision requested, the authorised owner and the time at which the request was made. Reopen rules should prevent recurring failure from being reset into a series of apparently successful tickets.

Work-order closure should match task risk. A routine inspection, statutory test, temporary repair and critical restoration need different proof. High-consequence, recurring or disputed work may require reinspection or user confirmation. The goal is not maximum documentation; it is enough evidence to justify the specific conclusion.

Figure 4. One incident, several service clocks Original Osmos Global conceptual framework, 2026. Sequence is illustrative; no target times implied. Prepared 1 September 2026.

What this means: One timestamp should not be forced to represent every stage of recovery.

8. Reconcile contract data with experience

User complaints and SLA results answer different questions. Complaints express experienced disruption; SLAs describe events through contractual definitions. A useful reconciliation preserves location, time, service and unresolved consequence, then tests whether apparently separate records describe the same condition.

This is particularly important where averages hide concentration. A portfolio can meet response targets while one floor, building or business function repeatedly suffers. The correct response is not to treat perception as a breach automatically. It is to investigate whether contract measures captured the experience and whether maintenance, asset or service design needs to change.

Complaint data carries reporting and privacy limitations. Low complaint volume may reflect silence or channel friction rather than satisfaction. Matching should use only the personal data needed for operational investigation and should preserve uncertainty where links cannot be established.

9. Manage recovery, suppliers and interfaces

Service credits arrive after underperformance. They cannot restore a critical service, replace unavailable expertise or create a fallback that was never tested. Recovery capability requires roles, authority, communications, alternatives, evidence of restoration and joint exercises.

Tabletop exercises should test decisions and handoffs under a credible scenario. They are not proof of physical capacity; engineering tests and drills may also be necessary. Their value lies in exposing unavailable decision-makers, unclear priorities, fragile subcontract dependencies and weak verification.

Where subcontractors deliver critical obligations, the prime provider should retain accountability while holding source evidence the client may inspect. Where the landlord controls a shared asset, an interface record should identify who detects, reports, grants access, authorises action, communicates and verifies restoration. Contract boundaries should not become operating blind spots.

10. Keep change and decisions traceable

Facilities operations evolve constantly. Informal instructions and emergency responses become dangerous when they quietly replace the contractual baseline. A variation log should capture the trigger, immediate direction, scope, cost, risk, data and KPI effect, approval route and implementation evidence.

Monthly reviews should use the same discipline. A presentation is not a governance outcome. Material exceptions should end with an authorised decision, owner, due date and verification test. Unresolved items need ageing and escalation rather than a refreshed narrative.

Accurate records support handover, dispute avoidance and organisational learning. They should show not only what was reported but why a decision was made, what uncertainty remained and whether the chosen action later worked.

11. Renewal and exit are evidence tests

Renewal decisions should examine material exceptions across the term. Classify causes by provider control, client dependency, asset condition, interface or external factor. Review recurrence, recovery quality and whether corrective action changed later performance. This creates a more defensible comparison of renewal, negotiated correction, competition and transition than a single average score.

Exit readiness should be rehearsed before the outgoing team leaves. The receiving team should retrieve and interpret representative asset, incident, compliance and defect records. A folder is not a successful handover when identifiers conflict, permissions fail or open actions sit outside the delivered system.

The evidence test should remain proportionate and secure. Sampling cannot prove completeness, and access must respect privacy and security. Its purpose is to find correctable gaps while knowledge and authority still exist.

12. Application to India and GCC operations

India’s GCC environments can combine rapid mobilisation, premium workplaces, multi-party property interfaces and high expectations for continuity. These conditions make evidence design important, but they do not make foreign public-sector guidance locally binding. The governing agreement, Indian law, building requirements and site-specific professional judgement remain decisive.

GCC leaders should focus on four applications. First, define evidence standards before scale creates inconsistent local practices. Second, distinguish landlord, client and provider authority for shared services.

Third, keep critical asset and defect histories usable across transitions. Fourth, ensure workplace-experience signals can be reconciled with operational events without inappropriate use of personal data.

The objective is a common decision language across portfolio, FM, procurement, workplace, finance and providers. Standardisation should improve traceability while allowing site-level controls to reflect consequence and operating reality.

13. Recommendations for leaders

FM leaders should select a small number of material obligations and test the complete evidence chain from event to acceptance. Procurement teams should publish evidence schedules, test resourcing assumptions and preserve audit and exit rights. CRE and GCC leaders should resolve landlord–client–provider interfaces before critical incidents occur.

Workplace teams should retain enough location and time context to reconcile experience with service events. Finance teams should distinguish a temporary cost avoidance from a verified recurring benefit and recognise when deferred capital constrains provider performance. Service providers should be able to reproduce reported results and explain exceptions without relying on informal knowledge.

Start with one high-consequence service and one recurring workplace friction point. Pilot the controls, record definition disputes, simplify evidence that does not affect decisions and scale only after the workflow survives operational review. The destination is not a larger monthly pack. It is faster, more honest and more accountable decision-making.

14. A 90-day implementation roadmap

During days 1–30, define the pilot. Select one critical service, one recurring experience issue and a small group of obligations whose evidence can be reconstructed. Name the operational owner, commercial owner and acceptance authority. Record current definitions, systems, access constraints and known data gaps. The output is a bounded test plan, not a redesigned contract.

During days 31–60, test the evidence chain. Sample reported results, map incident clocks, trial the obligation-to-proof schedule and conduct one decision exercise. Include both provider and client dependencies. Record every point at which reviewers cannot reproduce a result, identify an owner or distinguish temporary control from permanent resolution. Simplify requirements that generate volume without changing decisions.

During days 61–90, decide what to institutionalise. Confirm definitions, acceptance roles, exception escalation, change control and evidence retention. Update the contract-management plan or operating procedures through the applicable authority. Train the people who create and accept records. Choose a second service or site only after another team can use the workflow without relying on the pilot designers.

Elapsed time is not evidence of readiness. A high-consequence unresolved gap may justify extending or stopping the pilot; a low-risk imperfection may be accepted with a named owner and date. The roadmap is an Osmos planning framework, not an empirical implementation benchmark.

15. Measures that show whether the model is working

Implementation should be assessed through the quality of decisions, not the quantity of evidence collected.

Useful operational indicators include the proportion of sampled results that can be reproduced, the ageing of unresolved acceptance exceptions, recurrence after reported closure, retrieval success for critical records and the time taken to obtain an authorised decision.

These measures require careful denominators. A reproducibility rate should state the population, sample method and evidence test. Exception ageing should distinguish low-risk administrative gaps from conditions with safety, compliance or continuity implications. Reopen counts should not punish appropriate monitoring or honest temporary restoration.

Qualitative signals remain important. Client and provider teams should be able to explain the current baseline consistently; decision records should show why actions were chosen; and receiving teams should be able to interpret critical evidence. Improvement means greater traceability and better outcomes, not simply fewer recorded exceptions.

No universal target is proposed. Baselines, risk tolerance, systems and operating models differ. Leaders should establish a credible starting point, disclose limitations and compare change over time without altering definitions merely to improve the score.

16. Limitations and unresolved questions

The evidence base does not establish a universal causal relationship between the recommended controls and FM performance. The UK playbook is guidance. The ANAO audit concerns one public programme and a small engagement sample. The Malaysian study is qualitative, small and contractor-heavy. Industry articles provide context but may reflect commercial perspectives.

Contract evidence can also create new risks. Excessive collection consumes frontline capacity; intrusive workplace data can undermine trust; shared systems create cybersecurity and access concerns; and poorly designed audit rights can blur commercial boundaries. Proportionality is therefore a core design principle rather than a reason to avoid assurance.

Some questions require contract-specific professional judgement. These include legal effect of instructions and variations, allocation of landlord and tenant obligations, statutory evidence, safety-critical acceptance, privacy, data retention and the use of subcontractor information. This paper does not supply clauses or replace qualified advice.

Further research should test whether explicit evidence schedules, differentiated service clocks, mobilisation gates and retrieval rehearsals improve outcomes across different FM models and regions. Until comparable field evidence is available, organisations should treat the Osmos frameworks as transparent hypotheses to pilot, measure and refine.

17. Conclusion

The credibility of an FM contract does not rest on the number of SLAs it contains. It rests on whether the parties can recognise delivery, reproduce the evidence, make an authorised decision and keep exceptions visible until the operational consequence is resolved or consciously accepted.

That discipline begins before award and continues through mobilisation, performance review, change, renewal and exit. It asks more of both parties: the provider must demonstrate outcomes honestly, and the client must maintain definitions, decision authority, asset choices and retained capability.

Beyond the SLA is therefore not an argument for abandoning measures. It is an argument for completing them. A metric becomes management evidence only when its meaning, source, authority and limitations are understood. Organisations that build those connections will be better placed to protect reliability, workplace experience and long-term value.

Decision checklist

• Can the obligation be recognised as an operating outcome, not only an activity? • Is the minimum evidence, source, timing, retention and access route explicit? • Does the acceptance owner have authority to reject or condition delivery? • Will recurrence, failed verification and missing records remain visible? • Do approved changes update scope, price, evidence and performance assumptions? • Can another team retrieve and interpret critical records before exit?

Source notes

[1] Cabinet Office. The Contract Management Playbook. UK Government, 2026-03-25. March 2026 edition. Printed pp.45–49, 50–55 and 67; PDF page index = printed page + 3. Accessed 1 September 2026. https://www.gov.uk/government/publications/the-contract-management-playbook Contains public sector information licensed under the Open Government Licence v3.0. https://www.nationalarchives.gov.uk/doc/open-government-licence/version/3/ [2] Australian National Audit Office. Procurement and Contract Management by the Department of Foreign Affairs and Trade for its Security Enhancement Program. Australian National Audit Office, 2026-03-18. Auditor-General Report No.25 of 2025–26. Summary paras.11, 19; audit scope/method paras.1.17–1.18; recommendation6; chapter4. Accessed 1 September 2026. https://www.anao.gov.au/work/performance-audit/procurement-and-contract-management-by-dfat-for-security-enhancement-program [3] Syarifah Nur Shaqina Syed Shabahar; Haryati Mohd Isa; Nor Suzila Lop; Hussain Ismail. THE CHALLENGES OF FACILITIES MANAGEMENT AND MAINTENANCE CONTRACT DOCUMENT IMPLEMENTATION FOR OFFICE BUILDINGS IN MALAYSIA. Planning Malaysia / Malaysian Institute of Planners, 2026-06-03. DOI 10.21837/pm.v24i42.2033; volume24 issue3, pp.102–116. Methodology pp.108–109; Table2 p.109; findings pp.110– 113; limitations p.114. Accessed 1 September 2026.

https://www.planningmalaysia.org/index.php/pmj/article/view/2033

[4] Wei Xie (publisher-page byline). Global State of Facilities Management Report 2025. JLL, 2025-11-12. Page headline: Future-proof facilities management as a core engine for competitive advantage. Introduction / research context. Accessed 1 September 2026. https://www.jll.com/en-us/insights/global-state-of-facilities-management-report [5] CBRE procurement professionals; no individual byline displayed. Risk & Resilience: Navigating Your Facilities Management Supply Chain Amid Uncertainty. CBRE, 2025-07-16. Sections01,02 and05. Accessed 1 September 2026. https://www.cbre.com/insights/articles/risk-and-resilience-navigating-your-facilities-management-supply-chain-amid-uncertainty

Editorial and visual note

This publication is original Osmos Global analysis informed by the cited sources. Reported findings are distinguished from Osmos recommendations and illustrations. Source findings and trademarks remain attributable to their owners. The content is general research and does not replace contract-specific, legal, engineering, safety or other professional advice.

Cite this

Osmos Global Research & Knowledge Centre (2026). Beyond the SLA. Osmos White Paper, Osmos Global. https://www.osmosglobal.org/knowledge/beyond-the-sla

Keep reading

Download this paper

The full PDF, formatted for circulation. Downloads are for members, so that we know who our research reaches.

Discussion

Tell us where this matches what you see in your portfolio, and where it does not. Replies are welcome.

Members can add their input here.

Comments appear under your own name and company.

Join Osmos