Matrak Connect System Card · Addendum to Experiment 1: Estimation and Takeoff

A Re-evaluation of Frontier Language Models on Materials Takeoff: Six Projects, Audited Ground Truth, and Agentic Conditions

July 2026 · Read the original article (PDF)

Abstract. Before a construction project can be priced, someone must read its technical drawings and count every component they specify, then match each one to a purchasable product. This task, known in the industry as takeoff (Section 1), is unusually demanding: the same physical item appears from different angles across many drawing sheets and must be recognised as one item, classified against a regulated catalogue, and counted exactly, across hundreds of interdependent lines. It combines visual, spatial, classificatory, and arithmetic work, and it carries direct commercial weight, because every organisation bidding for construction performs it, by hand, many times over. The original system card (January 2026) found that no standalone frontier model could produce a takeoff accurate enough to use commercially. This addendum re-runs that test on the current model generation across six live commercial projects, drawn from the pipeline of a global construction supplier (anonymised); it adds conditions in which the models act as autonomous agents with full tool access, not only through chat; and it grades every system against a ground truth rebuilt by independent audit.

We report two findings. First, the Matrak system produced a commercially usable result, meaning a quote an estimator could act on with only minor correction, on a measure where no other system did, including the human estimating benchmark. Second, six months of model progress and full agent autonomy still leave every frontier model well short of commercial usability. The cause is not the interface or the tools; it is a persistent weakness in spatial reconciliation and exact counting, which drives the models to over-count systematically, the more costly direction in which to be wrong.

1. Background

1.1 What a takeoff is, and why it is hard

Before any construction project is priced, someone must count every material item shown on the drawings: every component, connector, and fixture, at every size and specification. This process is called takeoff, and the itemised list it produces, the bill of quantities, is the foundation of every quote, order, and delivery on the project. An error in the takeoff propagates into an error in the order.

BeforeA dense electrical construction drawing as issued, before takeoff markup.
AfterThe same electrical drawing after takeoff, with every counted item marked and identified.
Figure 1. What a takeoff markup looks like, shown on an electrical drawing carrying more than 800 elements (a Matrak AI Takeoff example). Left: the drawing as issued. Right: the same drawing after takeoff, with every counted item marked and identified. On a drawing this dense the markup is what makes the count auditable: without it, a reviewer cannot tell which items were counted and which were missed.

Takeoff is difficult by its nature, because it demands several distinct capabilities at once. It requires visual interpretation of technical drawings; spatial reasoning to reconcile the same physical item seen from different angles across multiple sheets and detail views, and to recognise it as a single item rather than several; classification of each item against a regulated product catalogue; accurate counting; and long sequences of dependent arithmetic as counted items are expanded into orderable components. Because these judgements compound, two experienced estimators working the same drawings, even within the same firm, will reach quantities that differ. Complete agreement is not the norm even among experts.

The scale of this work is substantial. In the original card we estimated that roughly 279,000 dedicated quantity surveyors and cost estimators work across Australia, the United Kingdom and the United States, spending 50 to 80 per cent of their time on manual takeoff. A further 257,000 estimating staff perform the same task on the supply side, at manufacturers, distributors, and suppliers. These are conservative figures: they count only formally titled roles. Takeoff is also performed part-time within many other construction jobs, and a growing share of it is outsourced offshore. Once these are included, the total full-time-equivalent effort devoted to takeoff across the three markets plausibly reaches a ceiling of over 900,000.*

The total is large for a structural reason: takeoff is not performed once per project. Every organisation that bids for a job, whether or not it wins, must complete its own takeoff to produce a quote, and many bid five to ten times for each job they secure. The winners then repeat the exercise for every design change and drawing revision through construction. The same drawings are counted, independently, many times over.

This combination of demands, visual, spatial, classificatory, arithmetic, and sustained over hundreds of interdependent lines, makes takeoff a demanding test for automated systems, and a particularly informative one for frontier language models, which have historically been strong on language and classification but weaker on spatial reconciliation and long, exact quantitative reasoning.

1.2 Measuring a takeoff, and what the original card found

To compare systems we need a single measure of takeoff quality, and we use the one defined in the original card, which we call Output Usability. A takeoff is delivered as a quote: a list of products, each with a quantity. We grade that quote line by line against a correct reference. A line counts as usable only if it names the right product and its quantity is within 10 per cent of the correct figure; Output Usability is the percentage of lines that meet both conditions, averaged across projects. In plain terms, it is the share of the quote an estimator could accept without correcting it.

The benchmark in this addendum was built with a single industry partner: a global supplier to the construction industry that prepares takeoffs across thousands of major projects each year, anonymised throughout this paper. Its projects, product catalogue, and reference estimates all come from live commercial work. The partner sets 85 per cent as the level at which an AI-produced quote becomes worth using. Above it, an estimator corrects a few lines and proceeds; below it, checking and fixing the draft costs more than starting again from the drawings. This makes usability a cliff rather than a slope. A score of 40 per cent does not deliver 40 per cent of the value: because the estimator cannot tell which lines are wrong without re-checking all of them, a quote that far below the bar is of near-zero commercial value.

On this measure, the original card reported that the best standalone frontier model reached 22.6 per cent on a single project, and that the strongest model of the time invented 19 product codes that did not exist. Six months later, a new generation of models is available. This addendum asks the same question of them, on a larger and more carefully graded benchmark, and adds conditions in which the models act as autonomous agents rather than only through chat.

2. Method

2.1 Projects and ground truth

Six real projects were drawn from the same partner and discipline as the original experiment, anonymised here as Projects A to F. They range from 46 to 75 unique reference lines each (the original benchmark project had 31); Project C is the most standard configuration in the set, and Project D the most complex, involving works that integrate with existing infrastructure.

Reflecting the challenging nature of the task, a takeoff's quality is not trivial to verify. The ground truth in our original card was our industry partner's takeoff, created by a senior estimator with almost 20 years of experience in the role. However, this artefact is itself the output of an expert working under commercial time pressure, and our audit found that it contains genuine errors. For this addendum we therefore constructed a stronger reference. Two independent human estimating teams and AI reviewers each produced or audited takeoffs for all six projects, and every disagreement was adjudicated line-by-line against the drawings until consensus. Working independently, each human team found real items the other had missed, and the audit confirmed the Matrak system had found items both teams missed. The resulting audited ground truth carries a documented residual error of approximately 2 per cent. We note the asymmetry itself as a finding: complete takeoff is difficult enough that even expert consensus retains a margin of error, consistent with the original card's caution that two experienced estimators may produce legitimately different, commercially valid takeoffs.

All systems in this addendum, including the human estimating benchmark, are scored against this audited ground truth.

2.2 Conditions

Five current frontier models ran the original card's standardised four-task chat protocol: itemise the drawings; select products from the partner's catalogue, supplied as roughly 300,000 characters of text; expand selections into orderable line items; and produce a final consolidated quote. This ran over a scripted multi-turn conversation with no tools of any kind. Two flagship models additionally ran a fully agentic condition: unrestricted autonomy inside an isolated workspace containing only the drawings, the catalogue, a quote template, and an identical written brief, with file access, code execution, and free choice of method. Workspace input files were hash-verified as unmodified after each run, and the agents had no access to any Matrak system or process.

Chat models ran at vendor-default reasoning settings, with one exception: Claude Opus 4.8 ships with reasoning disabled by default, so its vendor-recommended adaptive mode was enabled explicitly. Models tested: Claude Fable 5, Claude Opus 4.8, GPT 5.6, Gemini Flash 3.5, and Kimi K2.6 (chat); Claude Fable 5 via Cowork and ChatGPT 5.6 Sol via ChatGPT Work (agentic).

3. Results

3.1 Output Usability

Figure 2. Mean Output Usability across the six projects, scored against the audited ground truth, on a scale from 0 to 100 per cent. The dashed line marks the partner-defined 85% usability threshold: scores in the red band below it are not commercially usable, because correcting them costs more than estimating manually, while the green band above it is the usable range. The Matrak row is an independent, blind-scored run of the production system. The human estimating benchmark is the partner's shipped estimates scored against the same audited truth, included as the industry's operating baseline.

The Matrak system reached 88.9%, the only system above the threshold. The human estimating benchmark reached 66.5%. No frontier condition reached the threshold on any single project; the best frontier condition, Claude Fable 5 under full agentic autonomy, reached a mean of 64.4%, and the best single-project frontier result was 82%.

The relationship between the threshold and the human benchmark warrants care, because the two measure different things. The 85% threshold is a property of draft review: it is the accuracy at which correcting an AI-generated draft becomes cheaper than estimating from scratch. The human benchmark's 66.5% is a property of shipped output: it is the accuracy of estimates that leave the estimator and enter the supply chain, where their residual errors are absorbed by established downstream mechanisms, client relationships, the ability to respond to a site query or a mis-order, and the routine practice of ordering contingency. These are not the same standard, and a shipped-output accuracy of 66.5% does not imply that a 66.5% draft is usable. As Section 3.4 shows, the direction of the errors differs as well, and for an automated draft that direction raises, rather than lowers, the accuracy an estimator should require before trusting it.

One qualification applies to the human benchmark specifically. The estimates scored here are first-pass output. Every error the audit identified was subsequently caught by the human estimating teams on review, and a standard quality-assurance step applied to first-pass estimates would be expected to raise the human figure. The human benchmark should therefore be read as a floor on human performance rather than a ceiling; it is the industry's operating baseline, not a limit on what a careful estimator can achieve. We report the figures our experiment produced, and record this qualification alongside them.

System, conditionUsabilityBest projectWorst projectPhantom linesUsable?
Matrak system (independent run)88.9C: 100A: 7233YES
Human estimating benchmark (baseline)66.5C: 94D: 3741BASELINE
Claude Fable 5, agentic (Cowork)64.4C: 82D: 3547NO
Claude Fable 5, chat58.5C: 82D: 3059NO
Gemini Flash 3.5, chat49.0A: 64F: 3747NO
Claude Opus 4.8, chat39.9A: 46C: 2821NO
ChatGPT 5.6 Sol, agentic37.1C: 45D: 2022NO
GPT 5.6, chat36.1C: 59D: 2022NO
Kimi K2.6, chat27.1A: 52F: 1516NO
Table 1. All systems, scored against the audited ground truth. "Phantom lines" counts products from the partner's catalogue that appear in a system's quote but not in the audited reference, pooled over six projects (349 reference lines). Across all eight frontier conditions and 48 scored deliverables we observed exactly one fabricated product, meaning an identifier that exists nowhere in the supplied catalogue. The original card recorded 19 fabrications on a single project by the strongest 2025 model; that failure mode has effectively disappeared.

3.2 Effect of agentic autonomy

Two models ran both the chat and agentic conditions on identical inputs, which bounds the contribution of tool scaffolding for these systems. Full autonomy changed mean usability by +5.9 points for Claude Fable 5, from 58.5 to 64.4, and by +1.0 point for GPT 5.6 relative to its agentic Sol condition. Both remain far below the 85% threshold. Claude Fable 5's improvement concentrated on the more complex projects, where method freedom allowed it to cross-reference the drawings more thoroughly, while the flagship's agentic score was essentially unchanged from chat. The evidence indicates that tool access and method freedom, on their own, do not close the gap; they modestly raise a score that is already well short of usable.

Figure 3. Mean Output Usability for the two models tested in both conditions. Open markers: chat protocol. Filled markers: agentic condition. Dashed line: 85% threshold.

3.3 Relationship to model tier

Model tier did not predict performance on this task. A mid-tier model (Gemini Flash 3.5, 49.0) outscored a flagship (GPT 5.6, 36.1) under the chat protocol, and also outscored that flagship's agentic condition (37.1). The best chat-only result (58.5) likewise exceeded every agentic condition except that of its own model. General capability benchmarks appear to be weak predictors of performance on sustained, exact, domain-constrained selection and quantification.

3.4 The direction of error is commercially significant

Item identification is largely successful for the stronger models; quantity is where they fail. Figure 4 shows the final-quote decomposition for the best frontier condition: on several projects it misses almost nothing, yet fails a large share of lines on quantity alone.

passed (product and quantity within 10%) correct product, quantity out of tolerance missed entirely
Figure 4. Final-quote decomposition per project for the best frontier condition (Claude Fable 5, agentic). On Project F no reference line was missed entirely, yet 14 of 59 lines failed on quantity.

The direction of those quantity errors is not random, and it separates the three kinds of system cleanly. Figure 5 compares the human benchmark, the best frontier model, and the Matrak system by whether their erroneous lines under-count or over-count relative to the audited truth.

Figure 5. Direction of quantity error, pooled over six projects (349 reference lines). Left of centre: lines under-counted (below 90% of truth). Right of centre: lines over-counted (above 110%). The human benchmark and the frontier model fail in opposite directions; the Matrak system's errors are both fewer and close to balanced.

The two failure profiles have very different commercial consequences. The human benchmark under-counts: 72 lines under against 18 over, with a further 25 omitted. These are quiet errors. On a live project they are absorbed by standard practice, an estimator orders a margin of contingency, responds to a shortfall when it surfaces on site, and the relationship with the client accommodates the correction. The frontier models fail in the opposite direction: the best over-counts on 67 lines against 38 under. Over-ordering is the more expensive error, and endemic over-ordering is expensive systematically: it ties up capital in material that is delivered but not needed, and incurs restocking, storage, or disposal. An estimate that consistently over-orders cannot simply be trusted and adjusted with contingency; it must be corrected line by line before use, which is precisely the review cost the 85% threshold measures. This is why the threshold for an automated draft should be read as at least as high as 85%, and higher for a system whose errors skew toward over-counting, rather than being relaxed toward the human benchmark's shipped-output figure. The Matrak system, by contrast, both errs less often and errs symmetrically, 16 lines under against 10 over, and shows no systematic directional bias.

The mechanism behind the frontier over-counting is visible in the errors themselves, and it points to the spatial demand described in Section 1.1. When the same run of linear material or the same assembly appears on more than one drawing sheet, or in both a schedule table and a plan view, a human estimator recognises the two depictions as one physical item. The frontier models frequently do not: they count each depiction, and the quantity is inflated. In the clearest instances a model reported a total length in metres where the drawing schedule counts material supplied in fixed lengths, multiplying the true quantity several-fold. Spatial reconciliation, seeing the same item from several angles across several pages and recognising it as the same item, is a routine part of human takeoff and a persistent weakness of current models, and it produces a consistent bias toward over-counting rather than random noise.

4. Discussion

Three observations follow from the results. First, the failure mode that dominated the original card, fabricated product identifiers, has effectively disappeared from current frontier models: one instance across eight conditions, against 19 on a single project in the 2025 cohort. The remaining frontier errors are quantity errors on real products, and they skew toward over-counting through failures of spatial reconciliation. This is the error-propagation mechanism the original card described, now with a clearer cause and a clearer commercial consequence.

Second, the agentic comparison bounds how much of the gap is explained by interaction modality. Under identical inputs, unrestricted autonomy raised mean usability by at most 6 points and left every condition well below the threshold. Whether future model generations close the remainder is an empirical question; the present evidence indicates that tool access and method freedom alone do not, because they do not address the spatial and quantitative reconciliation that the task requires.

Third, the comparison with the human benchmark should be read as a comparison of operating baselines, not of upper bounds. The shipped estimates were produced under commercial time pressure and carry errors the industry absorbs downstream; their 66.5% is not a ceiling on human capability. The relevant standard for an automated draft is the usability threshold, and on that standard the audited results place the Matrak system above it and every frontier condition below it, with the human baseline itself below it. These results are consistent with the original card's central claim that structured organisational coordination, rather than model capability alone, is the binding constraint on takeoff automation.

These error patterns also have a direct economic reading. Under-counting, the characteristic error of the human baseline, tends to surface as a material shortfall during construction, which stalls work until the missing items are re-ordered and delivered. Over-counting, the characteristic error of the frontier models, ties up capital in material that is bought and never used. A takeoff that is both more accurate and directionally balanced, which the audited results show is achievable, therefore offers a twofold opportunity relative to standard human performance: lower ordering and holding costs from fewer surplus items, and fewer supply-driven delays from fewer shortfalls. Because takeoff is performed at the scale described in Section 1.1, repeated across every bid and every drawing revision, even a modest per-project improvement in accuracy compounds into a substantial reduction in wasted procurement and schedule risk across a supplier's portfolio.

5. Scope and limitations

6. Further work

Planned extensions: multi-discipline and multi-region validation; evaluation of visual auditability, whether a system can mark up the drawings to show what it counted; and re-evaluation as new model generations release. We expect parts of this addendum to age, and will publish updated results.

* The 279,000 and 257,000 headcounts are reconstructable from national occupational statistics: the US Bureau of Labor Statistics, the UK Office for National Statistics, and the Australian Bureau of Statistics with Jobs and Skills Australia. They count only staff whose formal occupation is estimating or quantity surveying, and so understate the true total. They exclude trades and subcontractors who prepare their own bids, project and site staff who take off materials to place orders, and the fast-growing offshore outsourcing of the task, none of which appear in those occupational codes. A fully layered estimate, adding each of these groups with its assumptions stated, reaches a high-end ceiling of over 900,000 full-time equivalents across the three markets.