All insights
AI & Data8 Min ReadJuly 18, 2026

Building Production AI Pipelines: Inbound Email Intelligence & OCR

A production teardown of inbound email intelligence: MIME parsing, attachment OCR, structured LLM extraction, confidence routing, and safe ERP side effects—without shipping a demo that collapses on week two.

R

Revilen Engineering

AI Systems · Revilen

Every company has an inbox that quietly runs the business. Purchase orders arrive as PDFs. Invoices hide in forwarding chains. Clients reply with “please update the shipment” in a thread that also contains a vacation photo. Teams hire people to read all of it. That works—until volume doubles and accuracy drops.

Generative AI looks like the obvious fix. Paste an email into ChatGPT, get a JSON blob, push it into the ERP. That demo wins meetings. It also fails in production for predictable reasons: malformed MIME, scanned attachments, vendor name collisions, duplicate messages, and models that invent line items when confidence is low.

This article is how we build inbound intelligence at Revilen when the output must touch money, inventory, or client commitments.

Start from the operating loop, not the model

Before choosing a model, map the human process you are replacing. Who reads the email today? What do they extract? What system do they update? When do they escalate? What happens on ambiguity?

If you cannot answer those questions, an LLM will not save you. It will amplify chaos. AI belongs inside a defined loop:

  1. Ingest — capture the raw message and attachments immutably.
  2. Normalize — decode MIME, isolate body text, classify message type.
  3. Extract — pull structured fields with OCR + LLM as needed.
  4. Validate — check schema, business rules, and confidence thresholds.
  5. Act — write to systems of record or queue for human review.
  6. Learn — log outcomes so prompts, parsers, and thresholds improve.

Production AI is not a prompt. It is an operations pipeline with retries, audits, and a human escape hatch.

Ingest like a lawyer, not like a chatbot

Store the raw RFC822 / provider payload before you transform anything. When extraction is wrong three weeks later, you need the original bytes—not a cleaned string someone mutated in memory.

  • Persist message-id, thread-id, received-at, mailbox, and provider metadata.
  • Hash attachments; store them in object storage with virus scanning where appropriate.
  • Deduplicate on message-id + attachment hash so retries do not create double work.
  • Never delete raw ingest on “success.” Soft-archive after retention policy.

Normalization is where most “AI projects” actually die

Real email is ugly. HTML newsletters, reply quotes, signatures, disclaimers, embedded images, and “On Tue, Jane wrote:” chains. If you feed all of that into a model, you waste tokens and raise error rates.

Build deterministic cleaners first: strip quoted history when safe, extract visible text from multipart MIME, detect language, and classify intent with a cheap model or rules (invoice, PO, support, spam, other). Only escalate expensive multimodal OCR when the classifier says the attachment matters.

OCR and LLMs: separate jobs, clear contracts

OCR turns pixels into text. LLMs turn text into structure. Blurring those jobs creates undebuggable systems. Run OCR as its own stage with its own confidence score. Pass text + layout hints into an extraction prompt that demands a strict schema.

Extraction contract (conceptual)
{
  "documentType": "invoice",
  "vendorName": "string",
  "invoiceNumber": "string",
  "issueDate": "YYYY-MM-DD",
  "currency": "USD",
  "lineItems": [{ "sku": "string?", "description": "string", "qty": 1, "unitPrice": 0 }],
  "total": 0,
  "confidence": 0.0,
  "needsReview": false,
  "rationale": "short reason for low confidence fields"
}

Require JSON. Reject free-form answers. If parsing fails, retry once with a repair prompt; then fail to review—not to “best effort invent a total.”

Ground the model in your world

Pass known vendors, open POs, and recent SKUs as retrieval context when available. Fuzzy-match extracted vendor names against your master list. Cross-check totals against line items. Compare invoice numbers against history to catch duplicates before finance sees them.

Confidence is a product feature

Every automated write should carry a confidence score and a needsReview flag. Define thresholds with operators, not engineers alone:

  • High confidence → auto-post with audit log.
  • Medium → draft in ERP / queue for one-click approve.
  • Low → human review UI with highlighted uncertain fields.

This is how you earn trust. Teams adopt AI when it reduces workload without creating silent financial risk.

Side effects must be idempotent

Email systems retry. Workers crash halfway. Providers redeliver. If your “create bill” webhook is not idempotent, you will double-book.

  1. Derive an idempotency key from message-id + document type + invoice number.
  2. Write an outbox record before calling the ERP.
  3. Mark the outbox complete only after a confirmed success response.
  4. Dead-letter failures with enough context for a human to replay safely.

Audit every action: who/what decided, model version, prompt hash, confidence, and downstream transaction id. When finance asks “why did this post?”, you answer in seconds—not with a shrug.

Human review is not a failure mode

Design the review queue as a first-class product surface. Show the original email, the extracted fields, diffs against known records, and one-click accept/edit/reject. Measure time-to-review. If reviewers constantly fix the same field, your prompt or master data is wrong—not your staff.

Observability for AI systems

Track pipeline SLIs the same way you track API latency: ingest lag, OCR failure rate, schema parse failure rate, auto-post rate, review backlog age, and post-hoc correction rate. Corrections are gold—they are labeled data for improving extraction.

What we ship for clients

Revilen builds these pipelines as part of broader operational platforms—not as orphaned notebooks. The inbox connects to workforce tools, chat systems, and the “everything app” integration layer so extracted work becomes tasks, messages, and system updates people already live in.

If you are evaluating AI for inbound operations, ask vendors (including us) for the boring details: raw retention, idempotency, review UX, and correction metrics. The model brand on the slide matters less than whether the pipeline survives a Monday morning surge of invoices.

Demos impress. Pipelines compound. Build the second one.