AI Technology Brief
StructuredLayer routes vision models by task - not by leaderboard rank
A vision-model catalogue can help discover options, but it cannot decide which model is safe or effective for a construction workflow. StructuredLayer separates general visual reasoning, multimodal processing, open deployment, object detection, and image generation; verifies model identity and lifecycle; then evaluates eligible candidates against the client's drawings, documents, site images, charts, screenshots, and video frames.

01 / Independently verifiable claims
Begin with what the technology and standards actually support.
- OpenAI officially documents image input for GPT-5, GPT-4o, GPT-4.1 mini, and GPT-4o mini, and now documents GPT-5.6 Luna, GPT-5.6 Terra, and GPT-5.6 Sol as official model tiers with image input. The supplied GPT-5.4 label was not found in the official model documentation reviewed and requires verification before use.
- Google officially documents image understanding and video understanding for Gemini models. Gemini 3 Flash Preview is documented but has a lifecycle and replacement path; Gemini 2.5 Flash is a distinct model generation, so aliases must not be treated as permanent IDs.
- Anthropic officially documents image input for Claude models, including Claude Sonnet 4.5. Claude Sonnet 4 and Claude 3.7 Sonnet are earlier generations and should not be selected from a static catalogue without current lifecycle review.
- Qwen2.5-Omni-7B supports text, image, audio, and video inputs, while Qwen2-VL-7B-Instruct and Qwen-VL-Chat are visual-language families with different generations, capabilities, and deployment requirements.
- Llama 3.2 Vision has official Llama 3.2 Vision 11B and Llama 3.2 Vision 90B model families. Names prefixed by a local runtime or hosting account are packaging identifiers, not independent model families.
- Moondream 2 is a compact visual-language model intended for efficient deployment; smaller size can support private or edge use but does not establish task accuracy.
- LLaVA 13B, LLaVA 1.6 Vicuna 13B, LLaVA 1.6 Mistral 7B, MiniGPT-4, CogVLM, InternLM-XComposer, BakLLaVA, UForm-Gen, and Qwen-VL-Chat are open research or model families with different ages, licences, hardware needs, context behavior, and maintenance status.
- OWL-ViT is an open-vocabulary object detector that returns labels, scores, and regions; it is not a general conversational visual-reasoning model.
- Kosmos-G is primarily an interleaved image-and-text generation system, so listing it beside image-understanding models without qualification confuses generation with inspection.
- Run counts, popularity, catalogue rankings, and labels such as fastest, best, or most capable are not substitutes for version-specific evaluation on representative construction evidence.
02 / The practical distinction
General reasoning, specialist detection, and image generation solve different problems.
A governed registry records each model's actual role. This prevents a convenient catalogue label from routing drawings, private site media, or proposal assets to an unsuitable system.
Hosted reasoning
General multimodal models
GPT-5, GPT-4o, GPT-4.1 mini, GPT-4o mini, Gemini 3 Flash, Gemini 2.5 Flash, Claude Sonnet 4.5, Claude Sonnet 4, and Claude 3.7 Sonnet can be candidates for image question answering, extraction, comparison, and explanation, subject to current official support.
Open deployment
Vision-language models
Qwen2.5-Omni-7B, Qwen2-VL-7B-Instruct, Llama 3.2 Vision 11B, Llama 3.2 Vision 90B, Moondream 2, LLaVA variants, MiniGPT-4, CogVLM, InternLM-XComposer, BakLLaVA, UForm-Gen, and Qwen-VL-Chat offer different self-hosting, adaptation, hardware, and maintenance trade-offs.
Specialist role
Detection or generation
OWL-ViT supports open-vocabulary object detection. Kosmos-G generates images in context. Neither should be treated as a drop-in replacement for a general evidence-analysis model.
03 / Operating architecture
Use a client-owned model registry and routing policy instead of hard-coding one model.
The registry verifies official identity, capability, lifecycle, deployment, security, cost, and evaluation status. The workflow selects only eligible models and preserves the exact version used for every output.
Approved visual input
Identified drawing, document, site photo, chart, screenshot, proposal image, or video frame with source, project, revision, permission, purpose, and fingerprint.
Deterministic router
Classify modality, task, sensitivity, quality, required evidence, latency, volume, budget, deployment boundary, and consequence before selecting a candidate.
Eligible vision model
Use a verified model ID and version from the registry with approved prompt, resolution, output schema, tool access, limits, retention setting, and fallback.
Controlled result
Validate structure and source references, expose uncertainty and conflicts, route exceptions, preserve the run record, and require human approval for consequential use.
04 / Required records
Every vision result needs source identity, model identity, and evaluation state.
Visual source
Project, record, file, page or frame, fingerprint, revision, status, capture time, location where appropriate, author, permission, sensitivity, and authoritative owner.
Task policy
Question, expected schema, allowed evidence, prohibited inference, decision consequence, required regions or citations, confidence handling, reviewer, and release condition.
Model registry
Official family, exact model ID, version, provider or deployment owner, modalities, context, resolution, lifecycle, licence, data controls, region, cost, and approved tasks.
Processing run
Source IDs, router decision, model version, prompt version, parameters, resolution, output, regions, usage, latency, warnings, retries, and fallback.
Validation
Schema, source grounding, object count, numeric checks, revision checks, unsupported claims, cross-model disagreement, ground-truth comparison, and exception state.
Review and outcome
Reviewer correction, evidence inspected, accepted fields, rejected claims, downstream record or draft, approval, timestamp, cost, and audit history.
05 / Construction example
Route each construction visual task to the narrowest evaluated capability.
The model prepares observations or candidate fields. It does not determine contractual meaning, safety status, payment, professional compliance, or permission to publish.
Drawings
Sheet and detail review
Ask source-bounded questions, identify candidate notes or regions, and compare visible changes. Drawing number, revision, status, and professional interpretation remain controlled.
Site media
Photo and video triage
Classify approved images, locate candidate objects, group repeated conditions, and prepare issue drafts. A qualified person verifies facts, severity, safety, and action.
Documents and charts
Visual extraction
Read tables, charts, scanned forms, dashboards, and mixed-layout pages into candidate structured data with page or region evidence and deterministic validation.
Digital operations
UI and proposal workflow
Analyze screenshots for support triage, compare approved proposal images, and prepare presentation content. Identity, consent, brand, and release remain governed records.
06 / Deterministic controls
Route with deterministic policy and test every model-version-task combination.
Verified model identity
Accept only official model IDs or approved self-hosted artifacts with documented provenance, version, checksum where applicable, lifecycle, licence, and owner.
Input and purpose allowlist
Restrict projects, record types, file types, image sources, video duration, resolution, sensitivity, user permissions, intended task, and prohibited inferences.
Resolution policy
Control page rendering, crops, tiling, frames, media resolution, and visual token use. Preserve the original source and transformations used.
Source-grounded output
Require page, frame, region, bounding box, or source-record references where supported. Unsupported descriptions remain candidate claims, not facts.
Model and rule fallback
Use normal software for known extraction or validation, specialist detection for bounded object tasks, and a second evaluated model only when policy permits.
Human authority
Block safety, contractual, commercial, employment, payment, compliance, design, and external-publication decisions until an accountable person reviews the source.
07 / Failure analysis
A large model list increases choice, but also increases lifecycle and evaluation failure.
Unverified catalogue name
A friendly name may not map to an official model ID. GPT-5.6 Luna, Terra, and Sol are now officially documented, while the supplied GPT-5.4 label still requires official verification before any configuration or claim.
Deprecated or preview model
Gemini 3 Flash Preview and older Claude generations illustrate why a static recommendation can age quickly. Lifecycle must be checked at selection and before every change.
Wrong capability class
An object detector, image generator, captioner, and visual reasoning model can all accept images but produce fundamentally different outputs and evidence.
Confident visual error
The model may miss a small note, misread a dimension, invent an object, confuse similar sheets, count incorrectly, or describe a condition that is not visible.
Lost source context
Cropping, tiling, frame extraction, compression, page rendering, or missing revision identity can remove the evidence needed to interpret the output correctly.
Sensitive media exposure
Site images, employee or client photos, screens, addresses, documents, and video can expose personal, commercial, security, or project-sensitive information.
Open-model operating burden
Self-hosting can improve control but adds hardware, serving, scaling, patching, monitoring, model-file security, licence review, and regression responsibility.
Benchmark transfer
General vision benchmarks and catalogue popularity do not predict performance on the client's drawing standards, scans, cameras, languages, UI, charts, or site conditions.
08 / Deployment and cost
Hosted, private, edge, and specialist routes require different operating evidence.
Hosted general models
Fastest to evaluate for broad reasoning. Confirm exact model, lifecycle, region, retention, training controls, subprocessors, limits, pricing, availability, and exit.
Private open models
Qwen, Llama Vision, Moondream, LLaVA, MiniGPT-4, CogVLM, InternLM-XComposer, BakLLaVA, UForm, and Qwen-VL families may support controlled deployment, subject to licence, hardware, quality, and maintenance evidence.
Edge or compact models
Moondream and other compact families can reduce transfer or latency for narrow tasks, but require representative evaluation and may need escalation for difficult evidence.
Specialist services
Use OWL-ViT-style detection for bounded labels and regions, deterministic OCR where suitable, and generation systems only for approved creative tasks separated from evidence analysis.
- Image, page, video, audio, and document ingestion, rendering, frame extraction, cropping, storage, transfer, and retention
- Visual tokens, reasoning tokens, output tokens, requests, retries, fallbacks, parallel comparisons, and provider minimums
- GPU hardware, serving, autoscaling, monitoring, patching, model storage, energy, and specialist support for private deployment
- Ground-truth preparation, prompt and schema design, model routing, evaluation, regression testing, and lifecycle review
- Validation rules, source-region preservation, exception queues, human review, correction, and professional approval
- Security, privacy, contracts, licences, incident response, backup, deletion, portability, and model replacement
09 / Evaluation
Evaluate accepted business outcomes by task, source condition, and model version.
- Field, object, count, chart, table, drawing-note, screenshot, and visual-question accuracy against approved ground truth
- Unsupported additions, missed evidence, wrong-page or wrong-frame references, region accuracy, and source-link completeness
- Performance by drawing family, scan quality, page density, image source, camera, lighting, crop, language, UI, chart, and video condition
- Repeatability and disagreement under fixed source, task, prompt, model version, resolution, and parameters
- Latency and complete cost per accepted page, image, frame, issue, extracted record, proposal section, or reviewed answer
- Permission, retention, deletion, sensitive-data handling, failure, timeout, truncation, fallback, and recovery behavior
- Human correction time, reviewer agreement, critical error severity, and errors reaching authoritative records or external documents
- Regression after model, alias, prompt, schema, resolution, preprocessing, routing, deployment, or source-system change
10 / Controlled pilot
Prove the operating boundary before expanding it.
Choose one visual outcome
Use one recurring source family and accepted output, such as extracting a drawing register field or triaging approved site photos.
Verify candidate models
Confirm official model IDs, current lifecycle, capability class, deployment, data controls, licence, limits, and cost. Quarantine unverified catalogue labels.
Build representative evidence
Include normal, low-quality, dense, ambiguous, revised, multilingual, conflicting, and adversarial examples with accepted ground truth.
Compare routes
Test at least one appropriate baseline and candidate on the same sources, prompt, schema, and acceptance measures; include deterministic or specialist alternatives.
Keep output non-authoritative
Use read-only processing and review queues. Do not permit autonomous design, safety, payment, contract, record, submission, or publication decisions.
Set change gates
Require re-evaluation for model alias, version, prompt, resolution, preprocessing, routing, deployment, pricing, terms, or source changes.
11 / StructuredLayer recommendation
Maintain a verified vision-model registry and route each task to the smallest, safest evaluated capability that meets the accepted outcome.
Do not select a vision model because a catalogue calls it fastest, deepest, most capable, budget, open, or popular. First separate reasoning, detection, captioning, multimodal streaming, and generation. Then evaluate official model versions against representative construction evidence, complete cost, data controls, failure behavior, source grounding, and human review. Preserve provider-neutral records so models can be replaced without losing the data layer, workflow rules, evidence, or approvals.
12 / Primary sources
Capability, governance, and implementation claims remain inspectable.
OpenAI
Images and vision
OpenAI
Models
Claude Platform
Vision
Claude Platform
Models overview
Google AI for Developers
Image understanding
Google AI for Developers
Video understanding
Google AI for Developers
Gemini models and deprecations
Qwen
Qwen2.5-Omni-7B model card
Qwen
Qwen2-VL-7B-Instruct model card
Meta
Llama 3.2 Vision model card
Moondream
Moondream 2 model card
LLaVA
LLaVA repository and model zoo
OWL-ViT model card
Vision-CAIR
MiniGPT-4 repository
Unum
UForm repository
Microsoft Research
Kosmos-G research page
Z.ai
CogVLM repository
InternLM
InternLM-XComposer repository
Qwen
Qwen-VL repository
NIST
Artificial Intelligence Risk Management Framework 1.0
Sources reviewed 22 July 2026. Technology capabilities, laws, guidance, terms, and pricing can change.
