Accept a document once
Validate type, MIME signature, size, page count, and organization ownership, then store a secure file reference and content hash.
Kursiv converts PDFs, scans, images, and Office documents into validated structured data. It picks the least expensive reliable processing path, applies profiles or caller-defined schemas, returns field-level confidence and source evidence, and routes uncertain fields to human review.
Kursiv is a secure document registry, profile and schema service, processing orchestrator, provider-abstraction layer, confidence and evidence engine, review workflow, and asynchronous job system - behind one stable REST and MCP contract.
The product boundary stops at trusted structured data. Kursiv does not decide which CRM record matches a document, update accounting systems, act as the knowledge base, or generate the final report. That is why the same extraction engine can serve healthcare forms, invoices, purchase orders, supplier documents, and contracts without becoming tied to one business process.
Validate type, MIME signature, size, page count, and organization ownership, then store a secure file reference and content hash.
General extraction, versioned document profiles, and caller-defined schemas describe exactly which fields a workflow requires.
Every business field returns a normalized value, type, confidence, and page or source evidence with coordinates where available.
Cache and native text resolve simple documents. OCR is page-selective and strong AI is reserved for fields still unresolved.
Each stage runs only on what the previous stage could not resolve. Stage-level cost and latency are exposed as metrics rather than buried in an infrastructure bill.
Content hash plus profile and processing version resolves repeat documents at no provider cost.
Readable native text is extracted first, before any paid OCR call is made.
Only pages that remain unresolved go to OCR, page-selective rather than whole-document.
Fields still unresolved after cheaper stages escalate to a suitable AI extraction model.
Fields below the confidence threshold create a review task instead of being silently guessed.
Secure upload and registration, tenant validation, file metadata, content hashing, and file-safety checks.
Select the profile, check cache, route native, OCR, and AI stages, and coordinate background jobs.
Native parsing, OCR vendors, and AI models stay replaceable behind normalized interfaces.
Known document types, required fields, field types, and customer-specific extraction contracts, all versioned.
Attach field confidence, page and source evidence, and review state to every structured output.
Manage low-confidence corrections, final approved results, secure storage, and result references.
A company defines or selects a profile, submits documents through an application or API, receives structured fields with evidence, and lets downstream services decide what to do.
The largest risks in document processing are overclaimed accuracy, uncontrolled AI cost, and silent extraction error. Each one is answered with an explicit control.
Organization ownership applies to documents, jobs, results, evidence, reviews, and cache entries alike.
Authenticated request context determines organization scope - never the model.
Approved-result caching keys on content hash plus profile, schema, and processing version, preventing unsafe cross-tenant reuse.
Large documents and batch workloads run through worker queues rather than long interactive sessions.
Provider cost, cache and native hit rate, human-review rate, and cost per completed document are operating metrics, not hidden infrastructure expense.
Confidence, evidence, and review thresholds are product controls rather than optional UI metadata.
Benchmarks run on labeled document sets before production migration, separate quality by profile and stage, use the same ground-truth fields across providers, and report accuracy alongside operating cost.
Exact and normalized match, with precision, recall, and F1 broken down by profile and field.
Share of documents completed without human review at an agreed quality threshold.
Share of returned business fields that include usable page or source evidence.
p50 and p95 end-to-end time by pages, file type, and processing path, with queue time reported separately.
Native, OCR, AI, storage, and review cost with stage-level attribution and cache savings.
Percentage resolved by cache and native text, selective OCR, AI, and human review.
Reliability is measured too: cross-tenant isolation, provider-failure fallback, job recovery, and schema-validation failure handling. Old and new flows run side by side on the same files and must show equal or better field accuracy, latency, and cost before cutover.
Azure, Google, AWS, and ABBYY all ship mature extraction. Kursiv competes above the provider layer on routing, evidence, profiles, and reuse across the Zelvora portfolio.
Established strength. Read and layout models, prebuilt and custom extraction, typed fields, classification, and a broad SDK ecosystem.
Kursiv position. Compete through provider-neutral orchestration and workflow governance, not by claiming a better OCR model.
Established strength. Enterprise OCR, prebuilt processors, custom extraction, validation, and model evaluation capabilities.
Kursiv position. Focus on one stable result and evidence contract with cost-aware routing across interchangeable providers.
Established strength. Managed OCR plus forms, tables, queries, expense and identity workflows, and customizable query adapters.
Kursiv position. Differentiate above the provider layer: profiles, selective escalation, review, and cross-product result reuse.
Established strength. Mature enterprise IDP with pretrained skills, low-code customization, human-in-the-loop, and flexible deployment.
Kursiv position. Aim for a lighter composable service integrated with the Zelvora stack rather than a full low-code IDP suite.