Human-in-the-Loop AI for Carbon Accounting: Why It's the Only Approach That Holds Up to Assurance
Fully automated AI carbon accounting is fast and cheap. It is also not assurance-ready. Here is why the human approval step is not an optional add-on — it is the structural requirement that makes AI extraction useful rather than merely convenient.
The Promise — and the Hidden Trap — of AI Carbon Accounting
The appeal of AI-assisted carbon accounting is immediate and concrete. An AI system can read a utility bill, extract the billing period, identify the consumption quantity in kilowatt hours, match it to the Singapore Energy Market Authority grid emission factor for the relevant year, and produce a tCO2e entry in seconds. For a business that is processing dozens of bills across multiple sites, multiple utility types, and multiple suppliers every quarter, this is genuinely useful. The manual alternative — downloading bills, transcribing figures into a spreadsheet, manually looking up the correct emission factor, checking the arithmetic — takes hours per quarter and introduces its own transcription errors.
The trap is not that the AI is inaccurate, though it sometimes is. The trap is that fully automated carbon accounting produces numbers that are fast, cheap, and not assurance-ready. From FY2029, SGX Main Board and Catalist companies are required to obtain limited assurance on their Scope 1 and Scope 2 greenhouse gas disclosures. An assurer does not ask whether a process was fast. They ask a different set of questions entirely: Who verified this extraction? What was the original source document? Was there a human reviewer with accountability for this entry, and can you produce a record of their approval?
Full automation — AI extracts, AI commits, no human reviews the result — cannot answer these questions. The speed advantage of full automation is real; the assurance gap it creates is equally real. Businesses that understand this trade-off early can design a process that captures the speed benefit of AI-assisted extraction while preserving the accountability trail that assurance requires. Those that treat full automation as the goal will discover the gap only when an assurer asks a question the system has no answer for.
What AI Extracts Well — and Where It Goes Wrong
Well-designed AI extraction handles routine cases with high accuracy. Given a standard Singapore electricity bill from a major utility — SP Group, Senoko Energy, or one of the licensed retailers — a properly trained extraction model will identify the billing period, extract the consumption figure in kWh, recognise the utility type as grid electricity, and pass clean structured data to the emission calculation engine. On clean, standard-format documents from familiar suppliers, this works consistently and the extracted values are reliable.
The problems appear at the edges, and the edges are more common in real business operations than a controlled demonstration suggests. A single bill covering two billing periods — which occurs when there has been a meter reading delay, a not uncommon situation — may produce an incorrect month attribution if the AI reads the issue date rather than the billing period end date. A credit memo for an overbilled period looks superficially similar to a standard charge and may be extracted as positive consumption rather than a reduction. A handwritten fuel consumption log from a logistics vehicle or a site generator is structurally different from a digital invoice and may be parsed with lower accuracy. A combined invoice covering multiple sites under a single account number may produce a consolidated consumption figure when site-level granularity is required for accurate Scope 1 and Scope 2 attribution.
Each of these errors is individually small and often inconsequential in isolation. A one-month misattribution shifts consumption from one reporting period to the adjacent one but does not change the annual total. A credit memo incorrectly extracted adds phantom consumption to a single month. But across a full year of real-world documents from an active Singapore business — a factory with multiple sub-meters, a retail operator across several outlets, a logistics company with a mixed vehicle fleet — edge cases accumulate. A report that is correct in its annual aggregate but wrong in a meaningful proportion of individual entries will surface problems when an assurer traces a specific line item back to its source document and finds that the numbers do not match. That is the scenario assurance-ready processes are designed to prevent.
The Confidence Score: AI's Own Uncertainty Signal
A well-designed AI extraction system does not only produce extracted values — it produces a confidence score for each extracted field. The confidence score represents the system's own assessed certainty that the value it has extracted is correct. A high-confidence extraction of a standard electricity bill from a utility with a familiar invoice format is unlikely to be wrong. A low-confidence extraction of an unusual document type — a handwritten log, a foreign-language invoice, a non-standard combined bill — deserves human attention regardless of whether the extracted value happens to be numerically correct in that instance.
Confidence scores are a meaningful signal because they turn the AI's own uncertainty into actionable human review prioritisation. Rather than a reviewer examining every single extraction with equal attention — which would be time-consuming and largely unnecessary for routine documents — a well-designed interface directs the reviewer's attention to low-confidence extractions and flagged anomalies, while allowing rapid approval of high-confidence routine entries. This is the correct use of human attention within an AI-assisted process: not replacing human review entirely, which would eliminate the accountability trail, but directing human attention where it adds the most value and can catch the errors the system itself signals it is uncertain about. A system that produces confidence scores and routes low-confidence items to active human review is more trustworthy than a system claiming universal accuracy with no uncertainty signal — because the former is honest about what it does not know.
What “Human-in-the-Loop” Actually Means in Practice
The phrase human-in-the-loop is used loosely across the AI industry, applied to everything from a final sign-off on a batch process to a continuous active review step. In the context of carbon accounting and sustainability assurance, it has a specific and assurance-relevant meaning that goes beyond a nominal approval checkbox.
The correct implementation works as follows. The AI system processes the source document — the utility bill, the fuel log, the refrigerant recharge record — and proposes an emission entry. The proposed entry includes the source document reference (so the original document is retrievable), the extracted billing period start and end date, the consumption quantity and unit, the identified utility type, the emission factor matched to that utility type and reporting period with its publication source cited, and the calculated tCO2e. A human reviewer — the sustainability manager, the ESG officer, or whoever holds accountability for the sustainability data — sees all of this in a single review interface. They are not asked to recalculate from scratch; they are asked to verify that the AI's extraction and calculation are correct.
The reviewer checks each field. Is the billing period correctly identified? Is the consumption quantity correct and in the right unit? Is the utility type correctly classified? Is the emission factor from the right source and the right year? Does the calculated tCO2e follow arithmetically from the quantity and factor? If something is wrong, the reviewer corrects it and records the reason for the correction. If everything is correct, the reviewer approves the entry. The approval event is logged: who approved, at what time, and what — if anything — was changed from the AI's original extraction.
Only after this approval does the entry commit to the evidence vault. The vault is append-only: an approved entry cannot be silently edited after the fact. If a correction is needed later — an error discovered during the following month's review, or a restatement following a supplier correction — a new superseding entry is created, with its own review and approval record, and the original entry is marked as superseded rather than deleted. This architecture creates an unbroken chain of custody for every emission figure: from source document, through AI extraction, through human review and approval, to committed entry, to the aggregated totals that appear in the sustainability report.
Why This Matters for IFRS S2 Assurance
The assurance requirement coming into effect from FY2029 for SGX-listed companies is not a theoretical regulatory concern. It is a third-party auditor — typically an accounting firm with sustainability assurance accreditation — examining your climate disclosures under IFRS S2 and forming an opinion on whether those disclosures are free from material misstatement. The process is systematic, document-intensive, and specifically designed to test the evidence chain behind reported figures.
The question an assurer will ask is very specific. For a given emission entry — say, 12.4 tCO2e attributed to your March 2025 electricity consumption at your Jurong facility — can you produce the original bill, the emission factor used and its source, and evidence that a human reviewer approved this entry before it entered the reported figures? With human-in-the-loop in place, the answer to all three parts of that question is yes and immediate. The source document is linked to the entry in the evidence vault and retrievable in seconds. The emission factor is version- controlled with its publication source — EMA's grid emission factor for the relevant year — cited directly. The approval record shows who reviewed the entry, when, and whether any correction was made to the AI's original extraction.
Without human-in-the-loop, the answer is substantively different. You have a final tCO2e figure in a spreadsheet, a scanned bill somewhere in a shared folder, and no recorded connection between them that demonstrates human verification. The assurer will ask who verified the extraction, and the answer will be that the system did it automatically, and there is no record of human review. That answer does not satisfy the assurance requirement, regardless of whether the final figure happens to be numerically correct. Assurance is an evidence discipline. The distinction between assurance-ready and not-assurance-ready is not primarily about accuracy — it is about the evidence chain. Businesses that build the chain before they need it will find the assurance process straightforward. Those that attempt to reconstruct it under assurance pressure in FY2029 will find it difficult, expensive, and possibly impossible if the original documents are not systematically retained.
When an assurer traces a tCO2e figure back to its source, they need to know: who approved this entry, on what date, and what was the original source document? A system with human-in-the- loop review answers this question immediately. A fully automated system cannot.
Speed vs. Accuracy: The Real Trade-Off
The perceived cost of human-in-the-loop review is time. This deserves honest examination rather than dismissal. For most Singapore SMEs with reporting obligations, Scope 1 and Scope 2 data involves a manageable number of source documents per quarter. A light-manufacturing firm might have two electricity accounts, one diesel fuel log for on-site equipment, and refrigerant recharge records if it operates cold-chain storage. A services business might have electricity bills for one or two office locations and a small fleet fuel account. A quarter of Scope 1 and Scope 2 data might represent 15 to 30 distinct extraction events in total.
Reviewing 30 AI-proposed entries — checking billing period, consumption quantity and unit, utility type, emission factor, and calculated tCO2e for each — takes approximately 15 to 20 minutes for a competent reviewer working with a well-designed review interface that presents all relevant information in a single view. That is 15 to 20 minutes per quarter, not per month, and not per document. The assurance-readiness benefit of that review is substantial and directly addresses the FY2029 requirement. The time cost is modest.
The honest comparison is between 20 minutes per quarter of structured human review and the alternative: attempting to reconstruct an evidence chain from incomplete records — a spreadsheet, a shared folder of scanned bills with no systematic linking to entries, no record of who verified what — under assurance pressure from an external auditor. Measured against that alternative, the trade-off is not close. The cost of building the right process now is small. The cost of not building it, discovered three years from now, is substantially larger.
The Architecture That Gets Better Over Time
Human-in-the-loop is not a static design with a fixed cost. As AI extraction systems accumulate operational data on the specific document types they encounter in real deployments — the invoice formats used by Singapore utilities, the layout variations across different fuel suppliers, the specific structures of combined multi-site bills — extraction accuracy improves and confidence scores calibrate more reliably. The system learns from the documents it processes, and from corrections made during human review, and becomes more accurate on the document types it has encountered before.
Over time, the proportion of extractions requiring active correction decreases. A system that required hands-on correction of 20 percent of extractions in its first quarter of operation — as it encountered new document formats and edge cases — may require hands-on correction of 5 percent by the end of its second year, as those formats become familiar and extraction accuracy for them improves. The remaining 95 percent of extractions can be approved in seconds, with the reviewer confirming that the AI's extraction looks correct rather than checking each field in detail. The human step remains, preserving the accountability trail, but its active burden decreases as the system matures.
This is the right trajectory for an AI-assisted sustainability data process: AI handles the routine with increasing competence, humans handle the exceptions and provide the accountability record, and the exceptions become rarer as the system learns the document landscape of the specific business. The important distinction from full automation is that this improvement compounds within a system that maintains the assurance requirement throughout. When a new document type appears — a new supplier, a new invoice format, a new type of fuel record — the confidence score drops appropriately, human attention is directed there, the system learns, and confidence rises. The audit trail persists across all of it. A fully automated system that removes human review to reduce cost has no mechanism to catch the edge cases it does not know it does not know — and those are precisely the cases that surface during assurance.
Frequently Asked Questions
What is human-in-the-loop AI for carbon accounting?
Human-in-the-loop AI means that the AI system extracts data from source documents — utility bills, fuel logs, refrigerant records — and proposes emission entries, but a human reviewer must approve each entry before it is committed to the evidence vault. The human sees the source document, the extracted fields, the proposed emission factor, and the calculated tCO2e, and either approves or corrects the entry. Only approved entries enter the final record.
Is AI carbon data extraction accurate enough for assurance?
AI extraction accuracy has improved significantly, but no system is perfectly accurate on diverse real-world documents. The more important question is whether the system provides an audit trail that assurance requires. An AI that extracts with 95% accuracy and provides a confidence score and human approval record is assurance-ready. An AI that claims 100% accuracy with no evidence is not — regardless of its actual accuracy.
Why does AI extraction need human review?
AI extraction can fail on edge cases: bills with multiple periods, credit memos, handwritten invoices, non-standard formatting, or ambiguous units. Each individual error is small, but errors compound across a full year of data. Human review catches these errors before they enter the record, and the review event itself creates the approval trail that assurance requires.
What should a human reviewer check when approving AI carbon extractions?
Check that the date range matches the billing period (not the issue date), that the quantity and unit are correct (e.g., kWh not MWh), that the utility type is correctly identified (electricity vs gas vs diesel), that the emission factor applied matches the correct source and year, and that the calculated tCO2e is arithmetically correct. For unusual documents — a credit memo, a partial-period bill, a combined multi-site invoice — review more carefully and document any adjustment made.
Can AI fully automate sustainability reporting?
AI can automate most of the extraction and calculation work, but assurance requires that a human has reviewed and approved each emission entry. Full automation — AI extracts and commits without human review — produces numbers that are fast and cheap but cannot satisfy an assurer's question: “Who verified this?” The right architecture is AI for speed, human for accountability.
AI Speed. Human Accountability. Assurance-Ready Evidence.
VerityOS combines AI extraction with a structured human approval gate — so every emission entry links to its source document, emission factor, and named approver. The result is a carbon accounting process that is faster than manual entry and holds up to assurance scrutiny.