Skip to content

What we automate

Document automation and data extraction

Reading documents, pulling out the fields that matter, filing them, and keeping records in step across systems.

The short answer

Document automation removes the retyping. A document arrives — an invoice, an intake form, a timesheet, a scanned certificate — and instead of a person reading it and entering the same facts into another system, the fields are extracted, validated against your own records, and filed. Every field carries a confidence score, and anything below the threshold you set goes to a person with the document open beside the uncertain value. The retyping is where most errors enter a business, so this usually improves accuracy at least as much as it saves time. That is the number we would rather be measured on.

What kinds of documents?

The common cases below cover most of what we are asked for. The test is not the format; it is whether the same fields are wanted from it every time.
Invoices and purchase orders
Supplier, date, line items, totals and tax pulled out, matched against what was ordered, and pushed into accounting without a second entry.
Intake and application forms
Customer-completed forms read into structured records, with missing or contradictory answers flagged rather than silently accepted.
Timesheets and job sheets
Field-completed paperwork turned into billable records, including the ones that arrive as a photograph taken in a van.
Certificates and compliance records
Expiry dates extracted and tracked, so the renewal reminder comes from the document rather than from someone's calendar.

How a document pipeline works

Six stages. The fifth one — routing on confidence rather than pretending to certainty — is what makes the difference between a system your team trusts and one they quietly work around.
  1. 01

    Capture

    Documents arrive the way they already arrive — an inbox, a shared folder, a portal upload, a scanner — and are picked up without anyone forwarding them somewhere special.

    No new habits

  2. 02

    Classify

    The system works out what the document is before trying to read it, because an invoice and a delivery note want different fields and mixing them is the usual source of silent errors.

    Type first

  3. 03

    Extract

    The fields that matter are pulled out, with a confidence score attached to each one rather than a single score for the page.

    Per field

  4. 04

    Validate

    Extracted values are checked against rules and against your own records: does this supplier exist, does the total match the lines, is this date plausible, have we seen this invoice number before.

    Rules, not vibes

  5. 05

    Route

    Anything above the confidence threshold and passing validation files itself. Anything below it goes to a person with the document and the uncertain field side by side.

    Threshold you set

  6. 06

    File and reconcile

    The record lands in the systems that need it, the original is stored where it can be found again, and the two stay linked so an auditor can get from one to the other.

    Audit trail kept

What actually reads well, and what does not

Extraction quality depends far more on your paperwork than on the model doing the reading. We test against your real documents during the audit instead of quoting a benchmark figure.
Clean PDFs
Digitally generated invoices and statements are the reliable end of the range. This is where most of the payback is.
Good scans
Flat, in-focus scans of printed documents work well. Quality of the scan matters more than the model reading it.
Block capitals on a form
Structured handwriting in defined boxes is usually workable, with a human check on the fields that matter most.
Free-form cursive
Often not worth automating. We will tell you that during the audit rather than after you have paid for a build.

Where your data goes

This is the question Canadian businesses raise most, so it gets a section rather than a footnote.

Where a document goes to be read is a design decision, made with you before anything is built. Depending on your requirements, processing can be kept within a chosen region, kept away from third-party model providers for specified fields, or restricted to documents you have explicitly routed to it. Where a workflow is sensitive enough that we would not be comfortable sending it to an external model at all, we say so and build it differently. That is a scoping conversation, not a reassurance.

13.4%

Cybersecurity or privacy concerns were the barrier most often reported as limiting AI use by Canadian businesses, ahead of cost at 10.6%.

Source: Statistics Canada, AI use by businesses, Q2 2026

Questions we get asked

What is document automation?
It is reading a document, pulling out the fields that matter, checking them, and filing them where they belong — without a person retyping. Invoices, purchase orders, intake forms, timesheets, certificates, reports and scanned paperwork are the common cases. The retyping is usually where errors enter a business, so removing it improves accuracy as much as it saves time.
Can it read scanned paper and handwriting?
Clean scans and typed PDFs are reliable. Handwriting varies: block capitals on a structured form generally work, free-form cursive on a crumpled sheet often does not. We test against your actual documents during the audit rather than promising a rate, because accuracy depends far more on your paperwork than on the model.
How accurate is it, and how would we know?
We do not quote a headline accuracy figure, because the number that matters is the one measured on your documents, not on a benchmark. What we do is set a confidence threshold: anything below it goes to a person to check, and everything is logged. You get a running view of how much is passing cleanly, which is the number worth managing.
Where does our data go when a document is processed?
That is decided during design, and it is one of the questions we raise first. Depending on your requirements, processing can stay within a specific region, avoid third-party model providers entirely for sensitive fields, or run only against documents you have explicitly routed to it. Cybersecurity and privacy are the barrier Canadian businesses cite most often, so we treat it as a design input rather than an afterthought.
Do we need to change the forms our customers send us?
No, though sometimes it helps. Automation that only works if everyone upstream changes their behaviour tends not to survive contact with reality. We build for the documents you actually receive. If a small change to your own form would remove a large amount of ambiguity, we will point it out and let you decide.

Related

Send us the document you retype most.

We will read it, tell you honestly how well it extracts, and show you what the pipeline would look like.

Book a workflow audit