Osmos
Articles

Osmos Global Publication · Osmos Perspective

Keep Generative AI Out of Uncontrolled Building Commands

A useful assistant needs a clear boundary between explaining an action and executing it.

Osmos Global Research & Knowledge Centre5 min readSign in to download

One interface can hide very different risks

A conversational assistant that locates a maintenance manual is performing a different task from a system that changes a building setpoint. Both may appear in the same chat interface, but the consequences of error are not equivalent. The first can misinform a user; the second can directly alter physical conditions. A familiar interface must not obscure that difference.

Hong and Li call for documentation and domain-specific evaluation of foundation-model applications in building research [HONG]. NIST places incident preparation and recovery within broader risk management [NIST]. Osmos applies these ideas by separating information access, recommendations, approved execution and bounded automation.

Define the authority boundary

Start with read-only assistance where the task permits it. The assistant can retrieve approved material, identify a source and state uncertainty. It should not gain write access merely because it has performed well on a set of question-answering tasks. Answer quality and control safety require different tests.

If the business later proposes executable actions, define an explicit allowlist, authorised users, operating limits, independent checks and rollback arrangements. These controls belong outside the language model as well as in its instructions. A request phrased politely or confidently should not bypass established permissions.

Figure 1. Choose the authority boundary explicitly Original Osmos Global conceptual framework, 2026. Categories are not a maturity score. Prepared 1 September 2026.

What this means: More accurate predictions do not automatically justify more control authority.

Treat source documents as evidence, not authority

Retrieval can improve grounding, but a retrieved document may be obsolete, irrelevant to the asset or contain text that should never govern system behaviour. Show the document version and applicable equipment. When approved sources conflict, the assistant should surface the conflict rather than synthesise a confident new procedure.

Testing should include missing manuals, ambiguous asset names, outdated instructions and requests outside the permitted role. Review whether the system refuses safely and directs the user to the responsible engineer. An assistant that always produces an answer can be less useful than one that recognises the boundary of its evidence.

Make approval meaningful

Human approval is not a safeguard if the reviewer cannot see the proposed change, its reason, affected equipment and rollback plan. Avoid approvals that become routine clicks under time pressure. The reviewer should have the competence and operational context to challenge the action, not merely the credentials to press a button.

Log the recommendation, supporting evidence, approving person, execution result and any override. If the model or retrieval system changes, re-evaluate the relevant tasks. A vendor’s general improvement claim does not establish that the building-specific workflow remains safe.

Illustrative decision rehearsal

In an illustrative scenario, a user asks an assistant to reduce overnight cooling. The assistant finds an old schedule and proposes a change without recognising a newly commissioned critical room. The language can sound reasonable even when the evidence is incomplete. The risk comes from connecting that answer to execution before the asset scope and current requirements have been checked.

A controlled workflow would identify the affected equipment, retrieve the current approved requirement and route the proposal to the responsible engineer. If the evidence is missing or inconsistent, the action should pause. The assistant can explain the gap without inventing a complete procedure or treating the user’s urgency as authority.

The first evaluation cycle should include adversarial and ordinary mistakes: similar asset names, documents that contradict each other, users requesting actions outside their role and temporary loss of a source system.

Record whether the assistant stayed within its permitted task, not only whether its prose was persuasive.

Acceptance should be based on the actual deployment configuration. A demonstration using curated documents and unrestricted administrator access does not establish reliability in a live estate. Keep the model, retrieval corpus, permissions and action interface versioned together. When any of them changes, the owner needs to know which previous tests no longer justify the current level of authority.

Keep the ambition proportionate

Generative AI may help teams search records, draft work descriptions and explain trends. It should not be represented as a qualified engineer or a replacement for approved safety procedures. Technical specialists must determine which control pathways are appropriate for a particular facility.

This article proposes governance principles, not a control-system design or an instruction to modify live equipment. The sources do not certify any specific assistant for autonomous operation. The defensible progression is from useful information to demonstrably controlled action, with the ability to stop, investigate and recover at every stage.

Source notes

[HONG] Tianzhen Hong and Han Li. Good practices for documenting AI-based studies on energy and buildings. Energy & Buildings / Elsevier; author copy hosted by Lawrence Berkeley National Laboratory, 2026-01-20. Sections 2, 3.1–3.6 and 4; pp. 1–4. DOI: 10.1016/j.enbuild.2026.117043. Accessed 1 September 2026. https://eta-publications.lbl.gov/sites/default/files/2026-06/1-s2.0-s0378778826001039-main.pdf [NIST] Alexander Nelson, Sanjay Rekhi, Murugiah Souppaya and Karen Scarfone. Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile. National Institute of Standards and Technology, 2025-04-03. Section 2; Table 2 GV.SC-05/08; Table 3 RC.RP. DOI: 10.6028/NIST.SP.800-61r3. Accessed 1 September 2026.

Editorial and visual note

This is original Osmos Global analysis informed by the cited publications. Reported findings are distinguished from Osmos recommendations and illustrative scenarios. Source findings and trademarks remain attributable to their owners. Original visual designs do not imply endorsement by source organisations. The content is general research and does not replace site-specific professional advice.

Cite this

Osmos Global Research & Knowledge Centre (2026). Keep Generative AI Out of Uncontrolled Building Commands. Osmos Perspective, Osmos Global. https://www.osmosglobal.org/articles/keep-generative-ai-out-of-uncontrolled-building-commands

Keep reading

Download this paper

The full PDF, formatted for circulation. Downloads are for members, so that we know who our research reaches.

Discussion

Add what you are seeing on the ground.

Members can add their input here.

Comments appear under your own name and company.

Join Osmos