Northwind Logistics
420 Harbour Way, Example City
Invoice
INV-2024-0841
Invoice date
14 Mar 2024
PO number
PO-55121
Free to use.No sign-up needed.
PDFGrid is a visual rule engine that runs entirely in your browser. Mark what you need — once — and extract the same data from every document that follows. Your files never leave your machine.
No sign-up to try it. No credit card. Nothing uploaded.
420 Harbour Way, Example City
Invoice
INV-2024-0847Invoice date
14 Mar 2024PO number
PO-55120next document · no rules redrawn
Built for finance, accounting and operations teams
The cost of manual entry
Invoices, bank statements, remittances, purchase orders — the numbers you need are right there on the page, but getting them into a spreadsheet means copying, pasting and double-checking. It is slow, it is error-prone, and it never ends.
Cell by cell
manual entry
Copy a figure. Check it. Copy the next one. One invoice at a time, one field at a time.
Weeks later
when errors surface
A mistyped digit that stays invisible until a reconciliation refuses to balance.
Every month
the same work
The same suppliers, the same layouts, the same afternoon — done again from scratch.
Privacy by architecture
PDFGrid processes your invoices inside your browser, without an upload or transfer step. PDFGrid processes everything inside your browser. There is no transfer step, no storage, and nothing for anyone to breach.
We have never seen a document.We store the templates you save — field names, rules and regions. PDF files and extracted page data stay on your device.
verify this yourself
Here's a ten-page invoice. Press run and PDFGrid pulls the fields out of it, right here in this tab.
A file can only leave your computer through a network request — so watch the counter while it works. It will say zero.
Northwind Logistics
420 Harbour Way, Example City
Invoice
INV-2024-0841
Invoice date
14 Mar 2024
PO number
PO-55121
Previewing page 1 of 10. The run extracts all ten pages.Download the sample PDF
Requests while extracting
0requests · 0 B
Rows will appear here as extraction completes.
What this measures. Every HTTP request this page made between the document entering memory and extraction finishing, counted by your browser's own Resource Timing API. Loading the engine is counted separately, above.
Resource Timing does not observe WebSocket frames, WebRTC data channels, service-worker requests, or traffic from your own browser extensions. PDFGrid opens no sockets, no data channels, and registers no service worker — open your network panel and check alongside this.
Sent nowhere
Documents are opened and read inside your browser. There is no document-upload endpoint to send them to.
Nothing retained
We store the templates you save — field names, rules and regions — not PDF files or extracted page data.
Built for precision
PDFGrid treats extraction like engineering: precise regions, explicit rules, repeatable results. No prompts to babysit, no models to second-guess.
Draw a box around any field — an invoice number, a date, a total. PDFGrid records it as a rule and applies it to every matching document automatically. What you see is exactly what you get.
Scan · Select Area
A rule engine, not a guess. The same document always returns the same values — no drift, no surprises, no model to second-guess.
When a rule cannot find its target, it returns empty rather than inventing a value. You see that before export.
Export clean, structured rows to Excel or CSV, ready to drop into whatever you already use.
Beyond the box
A selection works when a value stays static. When it moves, PDFGrid can find a phrase and select text relative to it instead of relying on fixed coordinates.
Two invoice layouts place the Total label and its value at different heights. The same phrase-anchored rule finds each value.
Find a phrase that stays meaningful — such as Total, Invoice No or Account number — then select the word, segment or line around it. The value can move as a document grows and the rule can still locate it.
That lets a template survive different document lengths without treating every page as an identical layout.
Nine ways to say what you want.
Returns all scanned text from the selected page or area.
Best for:Whole captures
Returns the first, last or N-th word, segment or line in reading order.
Best for:Fixed order
Finds a phrase, then returns the word, segment or line before or after it in reading order.
Best for:Moving labels
Finds a word, segment or line that matches the text you enter.
Best for:Known text
Returns the captured words between a start phrase and an end phrase.
Best for:Bounded ranges
Finds text by where it sits on the page — same row, same column, or within a distance of a phrase.
Best for:Nearby values
Finds a word, segment or line inside a fixed page-coordinate tolerance.
Best for:Headers, footers
Returns true or false depending on whether a phrase is found.
Best for:Markers, status
Matches a pattern in captured text; in a chain, it searches the previous result.
Best for:Structured values
What you capture
PDFGrid outlines every word, segment and line before you select anything. Choose the level a rule works at and see precisely what it will return.
The same invoice row is shown as individual words, as two segments split by a wide gap, and as one complete line. The visible outlines are labelled so grouping is not conveyed by colour alone.
Say you want a column of totals.
Captured as a line
Total $12,480.00
The label rides along and the result is text. That column won't sum.
Captured as a word
$12,480.00
Just the value, typed as a number. That column sums.
One setting per rule decides which.
This is why the geometry has to be right. The boxes you see are the boxes the rule uses — the same measurements drive the outlines on screen and the values in your export.
How it works
Drop in a single PDF or a stack of them. They're read directly in your browser — nothing is sent anywhere.
Box a field, or anchor to a phrase that always appears. Each becomes a named, reusable rule — no code, no prompts.
Every value appears in a review grid before it reaches your spreadsheet. Correct anything that needs it, then export to Excel or CSV.
Next time, open the documents and apply the template — step two is already done.
At scale
Add a stack of PDFs and every file runs through the same rules, giving you one clean table at the end. Because it all happens on your machine, there is no transfer wait and no per-page cost. For a practical walkthrough, see how to extract data from multiple PDFs to Excel.
Check before you commit
Extracted data lands in a results grid, not straight into a file. Every value is shown next to the document it came from, and anything that looks wrong can be corrected in place before you export.
FAQ
Everything you need to know before your first extraction. Still curious?
PDFGrid is a deterministic rule engine, not a prediction model. You define exactly which regions to extract, so the same document always returns the same values. There is nothing to babysit and every result is visible before you export it.
Nowhere. They are opened and processed inside your browser using a PDF engine compiled to WebAssembly. Nothing is transmitted to us or to anyone else — there is no upload endpoint to send them to. We store the templates you create, which contain field names, rules and regions, not PDF files or extracted page data.
Text-based PDFs with consistent structure — invoices, bank and credit-card statements, remittance advices, purchase orders, financial and lab reports.
No. Rules are created visually by drawing regions on the document. There is no scripting, no formula syntax and no prompt to tune.
Export to Excel or CSV.
Rules anchored to a phrase keep working when a value moves — a total anchored to the word Total is found whether the invoice has three line items or thirty. A rule that genuinely cannot find its target returns empty rather than guessing, and you see that immediately in the review grid.
Nothing at the moment. PDFGrid is free while I'm finding the first users, and there's no card to enter. If I add paid plans later, everyone already using it hears from me first and a free tier stays. I'd rather learn what this is worth from the people using it than guess before anyone has.
Nothing is metered on my side. Extraction runs in your browser, so there is no account limit on how many documents you process or how often.
There is one practical limit. Each time you apply a template, PDFGrid scans up to 500 new pages. Pages it has already read stay available, so applying the template again scans the remaining pages rather than re-reading the completed ones. If a batch reaches the limit, PDFGrid keeps the rows it completed, tells you pages remain, and marks any export made from it as partial.
Very large documents are limited by your browser's memory rather than by PDFGrid, so how far you can go depends on your machine and how dense the pages are.
Once the page has loaded, yes. The extraction engine and your document are both on your machine, so you can disconnect your wi-fi and carry on working. You need a connection to sign in and to save a template, and for nothing else.
PDFGrid is built and run by me, independently, in the UK. There's no team behind a support address — if you email, I'm the one who answers.
Your account and templates are stored with Supabase in Ireland (eu-west-1). That means an email address, and the field names and regions you create. Your documents are never stored anywhere, by me or by anyone else.
PDFGrid reads text-based PDFs. Scanned documents and photographs of paper are not supported — if a file has no readable text, we will tell you rather than guess. There is no API yet, and no automatic detection of tables that span multiple pages.
Open a PDF, set your first rule, and watch clean data land in your spreadsheet.
No sign-up to try it · No credit card · Nothing uploaded