Back to products
Kursiv - Document Intelligence

Documents in. Trusted structured data out.

Kursiv converts PDFs, scans, images, and Office documents into validated structured data. It picks the least expensive reliable processing path, applies profiles or caller-defined schemas, returns field-level confidence and source evidence, and routes uncertain fields to human review.

5
Stage cost-aware route
6
Core backend components
6+
Reusable document profiles
0
Cross-tenant reuse target
Product vision

Broader than OCR, narrower than a workflow engine.

Kursiv is a secure document registry, profile and schema service, processing orchestrator, provider-abstraction layer, confidence and evidence engine, review workflow, and asynchronous job system - behind one stable REST and MCP contract.

The product boundary stops at trusted structured data. Kursiv does not decide which CRM record matches a document, update accounting systems, act as the knowledge base, or generate the final report. That is why the same extraction engine can serve healthcare forms, invoices, purchase orders, supplier documents, and contracts without becoming tied to one business process.

Overview

Schema-driven extraction with reviewable output.

Accept a document once

Validate type, MIME signature, size, page count, and organization ownership, then store a secure file reference and content hash.

Profiles and schemas

General extraction, versioned document profiles, and caller-defined schemas describe exactly which fields a workflow requires.

Confidence and evidence

Every business field returns a normalized value, type, confidence, and page or source evidence with coordinates where available.

Cost-aware by design

Cache and native text resolve simple documents. OCR is page-selective and strong AI is reserved for fields still unresolved.

Cost-aware processing route

Pay for the cheapest stage that still gets it right.

Each stage runs only on what the previous stage could not resolve. Stage-level cost and latency are exposed as metrics rather than buried in an infrastructure bill.

01

Approved cache

Content hash plus profile and processing version resolves repeat documents at no provider cost.

02

Native parsing

Readable native text is extracted first, before any paid OCR call is made.

03

Selective OCR

Only pages that remain unresolved go to OCR, page-selective rather than whole-document.

04

Selective AI

Fields still unresolved after cheaper stages escalate to a suitable AI extraction model.

05

Human review

Fields below the confidence threshold create a review task instead of being silently guessed.

Backend components

One contract, replaceable providers.

Provider quality, price, or availability can change without rewriting the product contract callers depend on.

Document gateway

Secure upload and registration, tenant validation, file metadata, content hashing, and file-safety checks.

Processing orchestrator

Select the profile, check cache, route native, OCR, and AI stages, and coordinate background jobs.

Provider adapters

Native parsing, OCR vendors, and AI models stay replaceable behind normalized interfaces.

Profiles and schemas

Known document types, required fields, field types, and customer-specific extraction contracts, all versioned.

Confidence and evidence

Attach field confidence, page and source evidence, and review state to every structured output.

Review and result service

Manage low-confidence corrections, final approved results, secure storage, and result references.

Business use

Wherever the work starts with a document.

A company defines or selects a profile, submits documents through an application or API, receives structured fields with evidence, and lets downstream services decide what to do.

Healthcare forms
Extract patient, provider, order, and date fields from scanned or digital forms before validation or downstream workflow.
Invoices and receipts
Extract supplier, invoice, line-item, tax, and amount data for finance workflows or anomaly checks.
Purchase orders and supplier files
Normalize PO, product, quantity, pricing, and supplier fields for onboarding or comparison.
Contracts and business documents
Extract requested parties, dates, terms, or identifiers into a caller-defined schema.
Reconciliation inputs
Turn uploaded usage forms or commission documents into structured fields before CRM comparison and Reporting.
Document APIs for software products
External clients submit files and consume validated JSON without building their own provider orchestration.
Security, cost control, and scale

Tenant-safe caching and measurable unit economics.

The largest risks in document processing are overclaimed accuracy, uncontrolled AI cost, and silent extraction error. Each one is answered with an explicit control.

Organization ownership applies to documents, jobs, results, evidence, reviews, and cache entries alike.

Authenticated request context determines organization scope - never the model.

Approved-result caching keys on content hash plus profile, schema, and processing version, preventing unsafe cross-tenant reuse.

Large documents and batch workloads run through worker queues rather than long interactive sessions.

Provider cost, cache and native hit rate, human-review rate, and cost per completed document are operating metrics, not hidden infrastructure expense.

Confidence, evidence, and review thresholds are product controls rather than optional UI metadata.

Benchmark framework

A single headline accuracy number is not enough.

Benchmarks run on labeled document sets before production migration, separate quality by profile and stage, use the same ground-truth fields across providers, and report accuracy alongside operating cost.

Field accuracy

Exact and normalized match, with precision, recall, and F1 broken down by profile and field.

Straight-through rate

Share of documents completed without human review at an agreed quality threshold.

Evidence coverage

Share of returned business fields that include usable page or source evidence.

Latency

p50 and p95 end-to-end time by pages, file type, and processing path, with queue time reported separately.

Cost per document

Native, OCR, AI, storage, and review cost with stage-level attribution and cache savings.

Escalation rate

Percentage resolved by cache and native text, selective OCR, AI, and human review.

Reliability is measured too: cross-tenant isolation, provider-failure fallback, job recovery, and schema-validation failure handling. Old and new flows run side by side on the same files and must show equal or better field accuracy, latency, and cost before cutover.

Market comparison

Orchestration around extraction, not another commodity OCR engine.

Azure, Google, AWS, and ABBYY all ship mature extraction. Kursiv competes above the provider layer on routing, evidence, profiles, and reuse across the Zelvora portfolio.

Azure AI Document Intelligence

Established strength. Read and layout models, prebuilt and custom extraction, typed fields, classification, and a broad SDK ecosystem.

Kursiv position. Compete through provider-neutral orchestration and workflow governance, not by claiming a better OCR model.

Google Cloud Document AI

Established strength. Enterprise OCR, prebuilt processors, custom extraction, validation, and model evaluation capabilities.

Kursiv position. Focus on one stable result and evidence contract with cost-aware routing across interchangeable providers.

Amazon Textract

Established strength. Managed OCR plus forms, tables, queries, expense and identity workflows, and customizable query adapters.

Kursiv position. Differentiate above the provider layer: profiles, selective escalation, review, and cross-product result reuse.

ABBYY Vantage

Established strength. Mature enterprise IDP with pretrained skills, low-code customization, human-in-the-loop, and flexible deployment.

Kursiv position. Aim for a lighter composable service integrated with the Zelvora stack rather than a full low-code IDP suite.

Differentiation worth proving

  • Cost-aware escalation across cache, native extraction, OCR, AI, and human review.
  • Provider independence: select extraction technology by document type, cost, availability, or measured quality.
  • Field-level evidence: structured value, confidence, page and source location, and review status.
  • Reusable profiles for healthcare forms, invoices, purchase orders, usage forms, contracts, and custom schemas.
  • Cross-product reuse into Brain, Curagentic, Reporting, and reconciliation workflows.
Product information

A governed document-understanding service.

Primary users
Healthcare operations, finance and AP teams, supplier onboarding, and software products needing a document API
Core workflow
Register a document, select a profile or schema, route through cache, native text, OCR, and AI, then review and deliver structured fields
Output contract
Normalized value, type, confidence, page and source evidence, and review status for every business field
Out of scope
CRM matching decisions, accounting updates, organizational knowledge, and final report rendering
Differentiator
Provider-neutral orchestration with cost-aware escalation, versioned profiles, and evidence consistency across products