A confident answer is not necessarily a reliable answer
A project manager asks:
What is the approved completion date for this project?
The AI answers immediately:
The approved completion date is 30 September.
But where did that answer come from?
It could have used:
- The original contract
- A superseded programme
- An unapproved extension request
- A meeting transcript
- An outdated management report
- An email discussing a possible date
- A current approved programme
The answer may sound precise while using the wrong authority.
At StructuredLayer, we do not judge an AI system by how confidently it responds. We judge whether it used the correct project, current authorized sources, permitted information, defined business rules, and reviewable evidence.
This is a representative operating scenario. It is not a named client case study, legal interpretation, contractual decision, or professional certification.
What a trusted answer must preserve
A reliable project answer should expose:
- The project and record identifiers used
- The question the system interpreted
- The sources retrieved
- The document type and revision
- The source-authority status
- The date the information became effective
- The permissions applied
- Any conflicting evidence
- The confidence or uncertainty
- The exact citations supporting the answer
- Whether human approval is required
- When the answer was generated
- Which model and workflow version produced it
Trust does not come from the model alone. It comes from the complete operating system around the model.
The StructuredLayer three-layer model
Layer 01: The Data Layer
The Data Layer establishes what the AI is allowed to know.
We connect projects, companies, contacts, contracts, drawings, specifications, RFIs, programmes, change events, approvals, emails, and reports using stable identifiers.
A document record may include:
| Field | Example |
|---|---|
| Project ID | PRJ-00428 |
| Document ID | DOC-01842 |
| Document type | Construction programme |
| Revision | P06 |
| Status | Approved |
| Effective date | 18 July 2026 |
| Supersedes | DOC-01691 |
| Source system | Project document platform |
| Source URL | Controlled record link |
| Authority owner | Project director |
| Permission class | Commercial team |
| Extraction status | Human validated |
Without these fields, an AI system may retrieve the right words from the wrong document.
Layer 02: The Workflow Layer
The Workflow Layer determines how the question moves through the system.
A controlled answer workflow can include:
- Receive the user’s question and identity.
- Interpret the project, record type, date, and intended use.
- Apply permissions before retrieving information.
- Retrieve candidate evidence from approved sources.
- Filter superseded, draft, unrelated, or unauthorized records.
- Compare potentially conflicting sources.
- Draft an answer using the permitted evidence.
- Cite the supporting record, page, section, or field.
- Escalate uncertainty or material conflict.
- Record the answer, sources, cost, latency, and reviewer outcome.
The workflow must know when to stop. An answer should not be generated as fact when the system finds two competing approved records or cannot determine which source is authoritative.
Layer 03: AI and Automation Layer
AI can assist with:
- Understanding the user’s question
- Classifying its project and subject
- Extracting information from approved files
- Comparing relevant passages
- Summarizing evidence
- Identifying contradictions
- Drafting a source-linked response
- Explaining uncertainty in plain language
Deterministic rules should continue controlling:
- Project identity
- Document status
- Revision precedence
- Permissions
- Required citations
- Approval routing
- Confidence thresholds
- Prohibited actions
- Record retention
- Audit history
We use AI where interpretation is useful and rules where predictability is required.
Business Outcomes
Business outcomes sit above the operating model. They are the usable and measurable results produced by structured records, controlled workflows, bounded automation, and accountable human review.
The output should be more than a paragraph.
A controlled answer may contain:
Answer
The current approved completion date recorded in Programme Revision P06 is 30 September 2026.
Source
Programme P06, approved 18 July 2026, section 2.1.
Authority
Approved programme maintained by the project-controls owner.
Related evidence
Extension request EOT-04 remains under review and has not changed the approved completion date.
Confidence
High confidence based on one current approved source and no competing approved revision.
Required action
Commercial reliance or external communication requires project-director review.
This gives the user an answer they can inspect instead of an unsupported conclusion.
How we evaluate the answer
Approved historical examples
We begin with questions that qualified project professionals have already answered.
Historical examples may include:
- Current contract value
- Approved completion date
- Latest drawing revision
- RFI response status
- Outstanding change-event value
- Approved subcontractor
- Most recent cost forecast
- Current design responsibility
- Latest client instruction
- Missing closeout documents
Each example includes an expected answer and the evidence that makes it correct.
Accuracy thresholds
Different tasks require different acceptance thresholds.
A low-risk internal document classification may tolerate occasional human correction. A commercial answer about scope, payment, delay, or approval may require substantially stronger evidence and mandatory review.
We define thresholds for:
- Correct project matching
- Correct source retrieval
- Correct revision selection
- Field-extraction accuracy
- Citation accuracy
- Answer completeness
- Human-reviewer agreement
- Appropriate refusal or abstention
- False-positive and false-negative rates
False positives and false negatives
The consequences are not equal.
A false positive may incorrectly state that a change order is approved.
A false negative may fail to find an approved instruction that exists.
Both matter, but they create different operational and commercial risks. The evaluation plan must measure them separately.
Citation accuracy
A citation is not useful merely because it opens a document.
We test whether the citation:
- Opens the correct source
- Refers to the correct revision
- Supports the specific claim
- Points to the relevant page, section, or record
- Remains available to the authorized user
- Does not expose restricted information
Human-reviewer agreement
Two qualified reviewers may interpret evidence differently. We record reviewer agreement, disagreement, escalation, and final resolution.
This helps distinguish a model error from an underlying business ambiguity.
Human Control across every layer
Human Control operates across every layer. People define authoritative sources, approve sensitive actions, resolve exceptions, validate outputs, and remain accountable for professional and commercial decisions.
Questions that should require human approval
Human review should normally remain mandatory before AI output is used to:
- Accept contractual terms
- Approve payment
- Determine legal liability
- Approve a change order
- Confirm entitlement to additional time
- Issue a client commitment
- Reject a subcontractor
- Make a safety determination
- Certify completed work
- Replace professional design judgement
- Submit a bid
- Delete or overwrite authoritative records
The AI may retrieve, compare, summarize, draft, or recommend. The accountable professional makes the binding decision.
Production monitoring
Passing a pilot does not prove that the system will remain reliable forever.
We monitor:
- Answer accuracy
- Citation accuracy
- Retrieval failures
- Permission failures
- Unresolved conflicts
- Human overrides
- Escalation frequency
- Model and prompt versions
- Cost per completed answer
- Latency per answer
- Changes in source formats
- Unexpected output patterns
- Incidents and corrective actions
A model, prompt, retrieval rule, document template, or source system change can alter results. Regression testing should run after material changes.
Minimum acceptance test
Before production use, the system should demonstrate that it can:
- Match the question to the correct project.
- Enforce the requesting user’s permissions.
- retrieve the current authoritative records.
- exclude superseded and unapproved material.
- expose conflicting evidence.
- provide citations that support every material claim.
- refuse to guess when evidence is insufficient.
- route sensitive answers to an accountable person.
- record the model, workflow, sources, cost, and latency.
- continue passing approved historical test cases after changes.
- stop or roll back when production performance falls below the accepted threshold.
- preserve a reviewable history of answers and corrections.
Practical recommendation
Do not begin by giving AI access to every project folder.
Start with:
- One defined question type
- One project or controlled historical dataset
- One agreed source hierarchy
- One permission group
- One accountable reviewer
- One measurable evaluation set
- One clear approval boundary
At StructuredLayer, we build trust from evidence outward:
Connected records → controlled retrieval → source-linked answer → human authority → monitored production use
The objective is not to make AI answer everything. It is to make the system answer the permitted question from the correct evidence and stop when it cannot do so reliably.
Sources and further reading
- NIST AI Risk Management Framework
- NIST Generative AI Profile
- NIST AI Resource Center
- OpenAI: Practices for Governing Agentic AI Systems
Related StructuredLayer resources
- AI Readiness for Construction Data
- What Construction Companies Should Structure Before Automation
- Can You Trust Your AI Outputs?
Evaluate one construction AI use case
Tell us what the AI should answer, which sources should control the answer, who may access it, and what must remain subject to human approval.
