Prism

Nobody reads all 200 pages.

Prism does. It rebuilds every table from the geometry of the page, checks the arithmetic in code, and shows you the exact spot on the exact page each figure came from. When it cannot read something, it says so instead of guessing.

The live demo is PIN gated and runs on published annual reports only.

Muhibbah Engineering, annual report 2025, page 68
Contract work1,070,036
Interest income74
Shipyard740,201
Cranes22,548
Other121,404
Total revenue as stated1,954,263

Reconciled. Six figures, one total, checked in code.

Long documents get approved by people who cannot read all of them.

So they sample. A credit committee samples. A design reviewer samples. An examiner, an auditor, a procurement officer, a supervisor signing off a thesis: all of them sample, and none of them are being lazy. Two hundred pages times forty documents is not a week of work, it is a quarter of one.

What survives sampling is the quiet stuff. A subtotal that does not add up. An appendix that contradicts the summary three chapters earlier. A figure that changed in revision four and did not change in the table that quotes it. None of it looks like an error. It looks like a page you have already skimmed.

The expensive failure is not a document that is wrong. It is a document that is wrong in a place nobody was ever going to look.

Why this is not another document ingestion system.

Almost every tool sold for reading documents works the same way. Split the document into chunks, turn the chunks into vectors, retrieve whichever ones resemble the question, and have a language model read those and write an answer. That pipeline is built to answer one question well: what does this document say?

Prism is built for a different question: does this document hold up?

Scroll the table sideways to see both columns.

The same document, two approaches
  Retrieval and a language model Prism
Where a number comes from The model writes it, from text it retrieved. Nothing generates it. The figure is read off the page and the arithmetic is code.
If the document is wrong The error is retrieved and repeated, fluently. The totals do not reconcile, and it says which figures it added.
Table structure Taken from a flattened text dump, where columns have already been lost. Rebuilt from where the words physically sit on the page.
Evidence A chunk of text, if you ask for the source. Page number and coordinates. Click it and the marked figure is on screen.
What it missed Silent. A page that failed to parse produces no chunk, and a chunk that does not exist cannot report itself absent. Every page is accounted for as read, scanned, blank or failed, and the count has to reconcile.
When it cannot tell Answers anyway. Confidence is a tone of voice, not a measurement. Declines to judge, and states the reason.
Run twice Two answers, worded differently. The same answer. There is no temperature on a subtraction.

A number produced by a language model is a plausible number. It is usually right, which is worse than always wrong, because nothing on the screen tells you which time it is not.

Retrieval is genuinely useful and Prism includes it, for the questions that are actually questions. It just never gets to decide whether something adds up.

What it does.

No model touches a number

Figures are parsed as exact decimals and the arithmetic runs in code. This is the reason the whole thing can be trusted near a real decision, and the reason it can run on modest hardware inside your own network.

Every figure carries its coordinates

Not a page number, a box. The result knows where on the page the total sits and where each component it added sits. A claim you can open is a claim you can check, which is the only kind worth making.

A discrepancy is a question, not a verdict

When something fails to reconcile, a second reader takes the page again as pixels, through OCR, rebuilds the table from scratch and redoes the arithmetic. Only if both independent readings agree does the discrepancy stand. If they disagree it is withdrawn. The third layer is a person.

Declining to judge is a result

Roughly half the totals in a dense financial report are declined, each with a stated reason. A tool that reports only what it managed to check is telling you a flattering story, and the people reading it will find the gap eventually.

Coverage has to reconcile

Pages declared, pages attempted, pages read from the text layer, pages recovered by OCR, pages blank, pages failed. If the ledger does not balance, that is reported as loudly as a failed sum.

It runs where your documents already are

On your own hardware, with no outbound connection. Nothing is sent to a third party API, the page loads no external code or fonts, and the whole thing works with the network unplugged.

Where a document has to be checked against itself.

The capability is not about finance. It applies wherever a long document is assembled by many hands, states figures in more than one place, and is approved by someone who cannot read all of it.

Corporate credit and lending
Statements that have to add up, ratios that have to match the figures they are derived from, an appendix that has to agree with the summary quoting it. This is the first deployment, for the corporate division of a Malaysian bank, and it is the most built out.
Engineering and technical documentation
Maintenance manuals, specifications and service bulletins, where a tolerance in a table has to agree with the same tolerance in the text, and a revision that changed one has to have changed the other.
Construction, tenders and bills of quantities
Priced schedules that have to total correctly, rates that have to be consistent between sections, and addenda that move a figure in one place and leave it standing in another.
Research, academic and clinical reporting
Figures quoted in the discussion that have to match the results table, totals in the table that have to match their own columns, and sample counts that have to survive being restated four sections later.
Regulatory and compliance submissions
Dossiers assembled from many contributors under deadline, where the summary and the annex it summarises were written weeks apart by different people.
Insurance, claims and loss adjusting
Schedules of loss and supporting documentation, where the arithmetic is the claim and the supporting pages are where it is either evidenced or quietly is not.

The verification engine is the same in every case. What changes is the vocabulary of the document and which cross references matter, and that part is built with you rather than guessed.

What it will not do.

This section is on the page for the same reason the tool declines to judge out loud. You will find these out anyway, and it is better that you find them here.

What it did on documents it had never seen.

Five annual reports published on Bursa Malaysia, read cold. Nothing was prepared, cleaned or chosen for being easy, and every figure below can be checked against the filings themselves.

Scroll the table sideways to see every column.

Measured 8 August 2026
Filing Pages Checks Reconciled Did not reconcile Declined
Muhibbah Engineering 20251733832140169
Kerjaya Prospek 20241962551080147
Pavilion REIT 2024257192830109
AHB Holdings 20241926440024
V.S. Industry 202479877080
Total8979814520529

All 897 pages are accounted for in the coverage ledger and none were unreadable. 452 totals reconciled, 529 were declined with a stated reason, and nothing was reported as a discrepancy.

A run with no failures means nothing on its own. It could equally mean the alarm is broken.

So a sixth document is kept alongside the five, with an error planted in it deliberately. It still fails, and it is still marked unverified, on every run. A zero is only worth reading next to a control that proves the detector is awake, and that control is in the demo where you can trigger it yourself.

One honest note about the last row. The V.S. Industry file is the sustainability chapter rather than the full accounts, so it contains few totals to check. It is listed because leaving out the weakest result would make the other four mean less.

Tell us what you have to check.

The useful first conversation is about one specific document and what goes wrong when it is wrong. If there is a fit we will say so, and if there is not we will say that too.

Or write to concierge@techadvantage.io.

Questions we get asked.

What does Prism actually do?

It reads every page of a document, rebuilds the tables from where the words physically sit on the page, and checks the arithmetic in code. It reports what reconciled, what did not, and what it declined to judge, and every result carries the page and the coordinates it came from.

How is this different from a chatbot over our documents?

A retrieval chatbot answers what a document says. It finds passages that resemble your question and has a language model write an answer from them. If the document contains an error, the answer contains the error. Prism checks whether the document is consistent with itself, and the checking is done by code rather than by a model.

Does a language model produce the numbers?

No. Figures are parsed as exact decimals and every sum, subtotal and comparison runs in ordinary code. A model is used for wording and for answering questions in plain language, and it is never allowed near the arithmetic.

Can it run on our own servers, with no internet?

Yes, and that is the intended deployment. Optical character recognition runs locally, the interface loads no external code, fonts or images, and the content policy on the page forbids outbound requests. It works with the network unplugged.

Do you train on our documents?

No, and there is nothing that could be trained. The checks are deterministic rules written in code. Your documents are not used to improve anything, because there is no model in the verification path to improve.

What happens when it cannot read a page?

It says so. An unreadable page is recorded as failed and appears in the coverage ledger, which has to reconcile against the page count the file itself declares. A page that produced no result is never quietly treated as a page with nothing on it.

What does it do when it finds a discrepancy?

It re-reads the page independently before saying anything. The second reading discards the text layer entirely, takes the page as an image, runs optical character recognition, rebuilds the table and redoes the arithmetic. The discrepancy is only reported if both readings agree, and it is reported as a question for a person rather than as a finding.

How much of a document does it actually verify?

On dense financial reports, roughly half the totals it identifies are checked and the rest are declined with a stated reason. That ratio is published rather than hidden because a tool reporting only its successes is not telling you anything you can act on.

Is this only for banks?

No. Corporate credit is the first deployment and the most built out, but the underlying job is checking a long document against itself. That applies to technical manuals, tender documents, research reports, regulatory submissions and claim files just as directly.

What does it take to get started?

One real document and a conversation about what going wrong costs you. Everything after that depends on the document, and we would rather look at yours than describe ours.