AI Technology Brief
Should construction AI preserve document structure before retrieval?
Docling can preserve a structured document tree and support hierarchy-aware or hybrid chunking. Construction buyers should compare retrieval methods on contracts, specifications, tenders, O&M manuals, planning documents, and technical reports while keeping source files and qualified reviewers authoritative.

01 / Independently verifiable claims
Begin with what the technology and standards actually support.
- Docling converts supported documents into a structured DoclingDocument representation containing a document tree and typed items rather than only an undifferentiated text stream.
- Docling documents hierarchical chunking that uses the document structure and hybrid chunking that applies token-aware refinement on top of hierarchical chunks.
- Docling supports PDF and other document conversion, layout analysis, reading order, tables, pictures, formulas, and export formats, subject to model, format, document, and configuration limits.
- Official Docling documentation supports structured document representation and hierarchy-aware or hybrid chunking. It does not document a complete built-in chunkless construction RAG product.
- Custom tree navigation, parent-child retrieval, MCP tools, graph traversal, reranking, answer generation, permissions, and accepted business state remain separately engineered parts of an operating system.
- A reconstructed tree, table, heading, reading order, or text span can be wrong. The identified source file and revision remain authoritative, and consequential claims require page-level evidence and review.
02 / The practical distinction
Choose retrieval architecture by the document and decision, not one universal chunking rule.
Construction contracts, specifications, tender documents, O&M manuals, planning records, drawings, schedules, and technical reports have different structure and evidence requirements.
Complete
Full-document context
Useful when the source fits, permissions permit it, and cross-document volume is low. Cost, distraction, context limits, and weak pinpoint citations can grow quickly.
Simple
Fixed or semantic chunks
Efficient for local passages but can separate headings, definitions, qualifications, tables, exceptions, and cross-references from the text they govern.
Structured
Layout-aware chunks
Preserve page regions, headings, lists, tables, captions, and coordinates so retrieval can use more than flattened text.
Linked
Parent-child retrieval
Retrieve a focused child passage while supplying its governing section, heading, table, or document context.
Navigated
Document-tree navigation
Let a bounded tool inspect branches and relationships on demand. This is custom architecture, not a documented built-in Docling RAG mode.
Combined
Hybrid retrieval
Use exact terms, metadata, semantics, structure, reranking, and direct source access together when evaluation supports the added complexity.
03 / Operating architecture
Keep the source, derived structure, retrieval run, answer, and accepted decision as separate records.
The document remains authoritative. Parsing and retrieval create versioned candidate evidence for a named task and reviewer.
Source control
Document ID, project, type, revision, status, issue date, issuer, permissions, hash, supersession, and authoritative location.
Structured conversion
Converter and model versions, settings, pages, headings, lists, tables, captions, coordinates, reading order, confidence, errors, and artifact hash.
Bounded retrieval
Question, purpose, identity, filters, query variants, methods, evidence IDs, scores, parent context, omissions, conflicts, and latency.
Review and acceptance
Generated claim, exact source evidence, page or region, reviewer correction, decision, accepted output, expiry, and withdrawal.
04 / Required records
Stable source and evidence IDs make retrieval inspectable and correctable.
Source document
Source ID, native system ID, project, title, document number, revision, status, date, issuer, file hash, permissions, and superseded-by link.
Document element
Element ID, source ID and revision, type, parent, children, page, coordinates, text, table cells, caption, order, confidence, and parser version.
Retrieval unit
Unit ID, element links, hierarchy, contextualized text, token count, metadata, embedding version, lexical index version, and expiry.
Retrieval run
Run ID, user and purpose, question, source scope, filters, methods, versions, returned units, ranking, conflicts, omissions, and cost.
Evidence claim
Claim ID, exact passages or cells, source pages, interpretation, confidence, unsupported status, reviewer, correction, and accepted use.
Correction
Affected source, structure, chunk, index, answer, and downstream output; owner, reason, replacement, reprocessing status, and withdrawal evidence.
05 / Construction example
Structure matters when a nearby heading, table, definition, or exception changes the meaning.
These are evaluation scenarios, not claims that one parser or retrieval method will solve every document.
Contract clause
Retrieve the obligation with its definitions, exceptions, schedules, amendments, document priority, and exact controlling revision.
Technical specification
Keep section hierarchy, product requirements, execution clauses, referenced standards, submittals, and quality requirements connected.
Tender comparison
Preserve bidder, package, line, heading, unit, quantity, rate, total, qualification, exclusion, alternative, and source quote revision.
O&M manual
Connect equipment identity, procedure, warning, parts table, maintenance interval, image caption, model, and applicable manual revision.
Planning document
Retain policy hierarchy, site designation, map or table references, conditions, dates, authority, and superseded status.
Technical report
Separate observed facts, assumptions, methods, results, limitations, recommendations, appendices, figures, and professional conclusions.
06 / Deterministic controls
Treat every derived structure as reviewable context, not a replacement source.
Freeze source identity
Do not parse an anonymous latest.pdf. Record native ID, revision, status, hash, permissions, and supersession first.
Preserve page evidence
Keep page, region, table cell, heading path, and exact text so a reviewer can inspect the source rather than trust a summary.
Evaluate by document class
Test representative clean, scanned, multi-column, table-heavy, amended, image-led, multilingual, and poor-quality sources.
Expose conflicts
Do not silently choose between revisions, addenda, contract documents, schedules, correspondence, and extracted values.
Permission every layer
Apply access at source, element, retrieval, model, cache, log, answer, export, and reviewer boundaries.
Correct and withdraw
Reprocess affected structures and downstream answers when a source, parser, index, permission, or accepted interpretation changes.
07 / Failure analysis
Structure-aware retrieval can fail before, during, and after search.
Wrong reading order
Columns, footnotes, headers, stamps, and sidebars are assembled into a plausible but incorrect sequence.
Broken table
Merged cells, units, headers, continuation pages, subtotals, and notes detach from values.
False hierarchy
A heading or numbered clause is attached to the wrong parent, changing the apparent scope.
Retrieval omission
The relevant exception, definition, amendment, drawing note, appendix, or cross-reference is not returned.
Citation drift
The answer cites a nearby passage that does not substantiate the material claim or uses a superseded source.
Derived artifact becomes authority
A Markdown export, tree, chunk, embedding, cache, or generated answer is treated as the governing project record.
08 / Deployment and cost
Docling can be one conversion component inside local, private, or managed retrieval architecture.
Local conversion
Run approved models and dependencies in a controlled environment with bounded files, resources, logs, patches, and output validation.
Queued document service
Use source IDs, immutable inputs, job state, timeouts, retries, idempotency, quarantine, versioned artifacts, and reprocessing.
Retrieval service
Implement permission filters, metadata, lexical and semantic indexes, hierarchy, reranking, caches, direct tools, citations, and evaluation.
Operating workflow
Connect candidate evidence to review, correction, acceptance, source-system action, audit history, expiry, and withdrawal.
- Source inventory, access, revision reconciliation, scanning, OCR, and document preparation
- Conversion compute, parser and model operation, storage, derived artifacts, reprocessing, and monitoring
- Chunking, embeddings, lexical indexes, hierarchy, reranking, query generation, and retrieval infrastructure
- Model use, context assembly, tools, caching, citations, retries, and complete cost per accepted answer
- Representative evaluation, expert review, correction, permissions, security, licences, maintenance, and incident response
- Integration, operating records, change control, handover, archive, expiry, withdrawal, and replacement
09 / Evaluation
Compare accepted evidence and reviewer effort across retrieval methods.
- Source identity, revision, permission, page, heading path, table, caption, and coordinate preservation
- Element and reading-order accuracy across representative document classes
- Retrieval recall, precision, ranking, conflict detection, and citation support for material claims
- Answer groundedness, abstention, unsupported additions, and consequence-weighted critical failures
- Reviewer navigation time, correction effort, accepted-answer rate, latency, and complete cost
- Parser, model, embedding, index, permission, and source-change regression behavior
- Reprocessing, correction propagation, cache invalidation, withdrawal, recovery, and audit completeness
10 / Controlled pilot
Prove the operating boundary before expanding it.
Choose two document classes
Use one structured specification or contract and one table-heavy tender or O&M manual.
Create answerable cases
Define exact evidence, governing revision, expected answer, conflicts, and required abstentions with qualified reviewers.
Build four baselines
Compare full context, flat chunks, hierarchy-aware chunks, and a justified hybrid or tree-navigation method.
Test failure documents
Include scans, columns, continuation tables, amendments, superseded revisions, missing pages, and restricted material.
Measure accepted outcomes
Track evidence coverage, critical errors, unsupported claims, correction time, latency, and complete cost.
Set stop conditions
Stop for lost source identity, revision confusion, permission leakage, unsupported contractual or professional conclusions, or untraceable citations.
11 / StructuredLayer recommendation
Preserve document structure where it improves evidence retrieval, but keep the identified source and accountable reviewer authoritative.
Start with representative construction documents and compare simple retrieval against hierarchy-aware methods. Add custom tree navigation or hybrid architecture only when measured evidence coverage and reviewer effort justify its permissions, maintenance, latency, and failure surface.
12 / Primary sources
Capability, governance, and implementation claims remain inspectable.
Docling Project
Docling documentation
Docling Core
DoclingDocument
Docling Project
Chunking
Docling Project
Hybrid chunking
Docling Project
Document conversion
IBM
Docling repository
IBM Research
Retrieval-Augmented Generation
NIST
AI Risk Management Framework
Sources reviewed 12 August 2026. Technology capabilities, laws, guidance, terms, and pricing can change.
