AI Technology Brief
Interfaze combines specialist AI models - but deterministic output does not mean guaranteed correctness
Interfaze is a hosted beta model and API combining specialist OCR, document, speech, visual, web, browser, code, and structured-output capabilities. It could reduce the number of separate services needed in a construction processing pipeline. Its output still has to be evaluated as candidate evidence, not accepted as an authoritative business record because it is structured or described as deterministic.

01 / Independently verifiable claims
Begin with what the technology and standards actually support.
- The documented model identifier is `interfaze-beta`, exposed through an OpenAI-compatible Chat Completions API at `https://api.interfaze.ai/v1`.
- Interfaze documents text, image, audio, file, and video inputs; OCR, object detection, web search and scraping, speech-to-text, translation, code sandboxing, and guardrails.
- Structured-output examples use JSON Schema, Zod, or Pydantic, while the nonstandard `precontext` field can contain evidence metadata such as confidence scores and bounding boxes.
- Published limits are 1 million context tokens, 32,000 output tokens, 20 MB file objects, 80 MB URL inputs, 50 requests per second, and five minutes per request.
- The current methodology article reports 79.5% structured-output value accuracy while the leaderboard snapshot reviewed on 21 July reports 80.5%; neither is 100%, and the discrepancy requires version-specific clarification.
- The paper, leaderboard, blog, and benchmark repository are provider-associated materials. They do not establish independently reproduced construction performance.
- Pricing currently states $1.50 per million input tokens and $3.50 per million output tokens; preprocessing, media, browser, and sandbox activity can affect billed token usage.
02 / The practical distinction
Fixed structure, value accuracy, and business correctness are three different tests.
A response can satisfy a schema while containing the wrong value, and a correctly extracted value can still be mapped to the wrong project, revision, scope, or decision.
Format
Structured output
The response conforms to a requested JSON or task-specific structure. This supports downstream parsing but does not prove field correctness.
Extraction
Value accuracy
Candidate values match known ground truth under a defined benchmark. Interfaze's own published structured-output score remains below 100%.
Operation
Business correctness
The right value comes from the controlling source, is mapped to the right record, passes rules, exposes uncertainty, and receives required review.
03 / Operating architecture
Use the platform as a bounded processing layer inside a governed workflow.
Interfaze can prepare candidate fields and evidence. Client-owned records, rules, permissions, exceptions, approvals, and downstream state remain outside the model.
Approved inputs
Identified emails, quotes, forms, documents, images, audio, URLs, and permitted web pages.
Specialist processing
OCR, structured extraction, visual detection, speech, translation, browser or scraper, code sandbox, and vector retrieval.
Evidence output
Schema values plus available confidence, bounding boxes, timestamps, raw metadata, model, usage, and request identity.
Controlled use
Field validation, source mapping, exception routing, reviewer correction, approval, and client-owned system update.
04 / Required records
Every candidate value needs identity, evidence, validation, and review state.
Request
Workflow, project, purpose, model, task, prompt version, schema, input IDs, limits, start, completion, and usage.
Source
Source system, file or URL, fingerprint, revision, status, permission, page, timestamp, and authoritative owner.
Candidate field
Field name, raw value, normalized value, expected type, confidence, bounding box or timestamp, and source link.
Validation
Required, format, range, unit, arithmetic, duplicate, revision, relationship, and cross-source results.
Exception
Missing, unreadable, uncertain, conflicting, rejected, timed out, limit exceeded, assigned owner, and resolution.
Approval
Reviewer correction, evidence inspected, accepted value, downstream record, approval, timestamp, and audit history.
05 / Construction example
One pilot can test documents, voice, and portal inputs without automating decisions.
The useful question is whether one bounded workflow produces accepted records with visible evidence and exceptions at an acceptable complete cost.
Quote extraction
Subcontractor comparison
Extract candidate prices, inclusions, exclusions, alternates, and evidence regions; estimator decides scope equivalence and commercial use.
Voice
Daily report draft
Transcribe site audio with timestamps and structure observations; field author verifies identity, facts, quantities, safety, delay, and issue wording.
Document control
Register preparation
Extract drawing and document metadata; rules validate project, number, revision, status, fingerprint, and superseded relationships.
Browser
Portal monitoring
Read allowlisted pages and permitted files; workflow validates identity and changes, then stops before replies, acceptance, or submission.
06 / Deterministic controls
Require deterministic checks around every probabilistic or provider-defined result.
Input allowlist
Restrict source systems, URLs, file types, sizes, modalities, projects, permissions, and processing purpose.
Schema and field rules
Validate required fields, types, controlled values, units, arithmetic, relationships, and source references independently.
Confidence routing
Calibrate confidence by field and document family; never treat provider confidence as correctness proof.
Source provenance
Persist original evidence, fingerprint, page or timestamp, bounding box, extraction version, and request ID.
Usage budgets
Limit input and output tokens, files, request time, retries, browser steps, sandbox output, concurrency, and task cost.
Human authority
Block commercial, contractual, professional, safety, payment, permission, and external actions until approved.
07 / Failure analysis
A combined platform reduces integration count but concentrates vendor and workflow dependencies.
Valid JSON, wrong value
The schema passes while a price, date, name, quantity, status, or revision is extracted incorrectly.
Evidence unavailable
`precontext` is untyped in documented TypeScript examples and its complete schema or availability across every task is not defined.
Benchmark transfer
Provider-published OCR, speech, visual, or structured-output results do not predict performance on the client's construction records.
Hosted-service dependency
Processing location, subprocessors, model changes, availability, retention, deletion, exit, and data portability require due diligence.
Audit gap
Detailed observability and logging are listed as coming soon, limiting production investigation unless the client creates its own run records.
Unexpected billing
Multipass preprocessing and infrastructure activity can increase billed tokens beyond the content directly supplied.
Limit mismatch
Context, file, URL, output, or time limits can truncate, reject, or change processing behavior on large project packages.
Security overstatement
Offers to discuss SOC 2 or HIPAA requirements are not evidence of current certification, attestation, BAA, or independently verified compliance.
08 / Deployment and cost
Hosted beta, VPC, and self-hosted options require different evidence before use.
Hosted API
Fastest route to a pilot using the documented endpoint and token pricing. Requires review of location, subprocessors, retention, security, availability, changes, and exit.
Enterprise or VPC
Pricing lists VPC deployment, SLAs, volume terms, and compliance agreements as available by contact. Exact scope and evidence must be obtained contractually.
Self-hosted
Pricing lists self-hosted availability for custom plans but does not publish architecture, hardware, maintenance, update, licence, or support details.
Client-controlled wrapper
Keep identity, source files, validation, run logs, exception state, approvals, and downstream writes in client-owned systems regardless of deployment.
- Input and output tokens, including multipass preprocessing
- Binary media conversion, browser, scraper, sandbox, and generated output usage
- Source storage, transfer, retention, and client-side evidence preservation
- Validation rules, record matching, exception queues, approvals, and audit history
- Security review, contracts, data protection, monitoring, backup, and incident response
- Human correction, maintenance, regression testing, vendor change, and exit planning
09 / Evaluation
Test accepted construction outcomes, not provider benchmark rank.
- Field accuracy by document family, field, language, scan quality, and source condition
- Source-link, bounding-box, timestamp, confidence, and provenance completeness
- False accepted values, missed fields, wrong mappings, and unsupported additions
- Repeatability under identical request, task, schema, and source conditions
- Failure, timeout, truncation, retry, recovery, and exception behavior
- Human correction time and critical errors reaching downstream records
- Security, permission, retention, deletion, and sensitive-data handling evidence
- Latency and complete cost per accepted record, document, audio report, or portal check
10 / Controlled pilot
Prove the operating boundary before expanding it.
Choose one outcome
Use one recurring document or audio family and one accepted structured record.
Build ground truth
Create representative accepted, rejected, ambiguous, low-quality, multilingual, and edge-case examples.
Preserve evidence
Store original inputs, fingerprints, request configuration, output, precontext, usage, validation, corrections, and approval.
Keep read-only
Do not allow autonomous system updates, messages, submissions, payments, or commitments during the pilot.
Set critical failures
Treat wrong-project mapping, false accepted values, missing material qualifications, evidence loss, and unauthorized transmission as stop conditions.
Review vendor evidence
Confirm current limits, score snapshots, data location, subprocessors, retention, deletion, security, observability, SLA, pricing, and exit before expansion.
11 / StructuredLayer recommendation
Pilot Interfaze as a candidate-processing layer, not as an authoritative construction record or autonomous decision-maker.
Start with one narrow task and representative ground truth. Require source-linked fields, deterministic validation, visible exceptions, human approval, client-owned run records, and complete cost measurement. Expand only after critical errors, vendor dependencies, security evidence, observability, and recovery remain within agreed acceptance boundaries.
12 / Primary sources
Capability, governance, and implementation claims remain inspectable.
Interfaze
Interfaze documentation
Interfaze
Limits
Interfaze
Security and privacy
Interfaze
Pricing
Interfaze
Leaderboards
Interfaze
Architecture and benchmark methodology
arXiv
Interfaze: The Future of AI is built on Task-Specific Small Models
InterfazeAI GitHub
Interfaze benchmark runners
NIST
Artificial Intelligence Risk Management Framework 1.0
RICS
Responsible use of artificial intelligence in surveying practice
Sources reviewed 21 July 2026. Product capabilities, models, APIs, terms, and pricing can change.
