Operations & Productivity

How Manufacturers Are Using AI to Boost Productivity

In a casework or millwork plant, the machines are the fast part. A CNC router holds tolerances all day. The slow part sits upstream, at the desk where someone opens a set of architectural drawings and starts counting.

That step decides how many jobs the business can quote. A set of moderate complexity takes an estimator two to six hours to work through. A large one takes longer. Because a bid depends on it, the count has to be right, which means it can’t be rushed. When invitations to bid arrive faster than that desk can process them, the plant doesn’t run out of capacity. The company just answers fewer of them.

When manufacturers talk about AI, the conversation usually stays on the floor: predictive maintenance, scheduling, vision systems on the line. Those are real. But for a make-to-order shop, the gain that arrives fastest tends to sit upstream of all of it, in the paperwork that decides which jobs come in at all.

Over the past two years, manufacturers have started handing parts of that reading step to AI. Some of it works well, some of it doesn’t, and the difference matters if you’re planning an operation around it.

Where the hours go

Break the step into its parts and it’s clear which ones a machine can take.

Finding the right sheets. A set can run to hundreds of pages. Most aren’t relevant to casework. Someone still has to page through and decide.

Counting and classifying. Every cabinet, countertop and elevation gets tallied by type. This is the bulk of the hours, and it’s mechanical.

Searching the specification. The spec document routinely runs past 200 pages. The clauses that change the price — hardware, finishes, tolerances — are scattered through it. Estimators used to spend 30 minutes or more hunting for one answer.

Judgment. What the drawing implies but doesn’t show, which assembly the detail really calls for, what to flag as a risk. This is the part that needs an experienced person.

Three of those four are searching and counting. That’s what the technology is genuinely good at.

What current AI models can and can’t do on a drawing

It’s worth being precise here, because vendor claims in this space run well ahead of reality.

We ran a benchmark of 11 AI models on architectural drawings: 119 pages, 1,430 objects marked up by hand, each model asked to find the objects and mark where they are.

The best model located 92% of them correctly. The second best managed 70%, and most of the field came in below that. Large views like floor plans were easy for nearly everything we tested. The small stuff — thin countertop lines, callout bubbles a few pixels wide — is where models fell apart, and those are exactly the items a takeoff depends on.

One more result matters for planning: the same model scored anywhere from 73% to 100% depending on which drawing set we gave it. Dense, tightly drafted sets pull the numbers down. So a single accuracy figure from a vendor describes their test drawings, not the ones your customers send you.

The practical conclusion is that a general-purpose model on its own doesn’t do a takeoff. What works in production is a system built around one: routing to the relevant sheets, cropping the views, reading them at full resolution, exporting structured data, and putting the result in front of a person.

What it looks like in production

Stevens Industries, the biggest commercial casework and architectural millwork producer in the United States, has been running that kind of system in daily production since 2025.

Drawing review that took two to six hours now takes about ten minutes. Detection accuracy runs at about 90%, measured against hand-marked drawings. The largest set processed so far ran to 470 sheets. Output is exported in a structured format the quoting system imports directly, so nobody retypes numbers. Specification lookups that took half an hour come back in under 30 seconds, with the page reference attached.

The system paid for itself in under a month. Headcount stayed where it was: each estimator now gets through ten to twenty times as much work, which for a company that wins jobs by bidding means more invitations answered, not a smaller team.

The number operations should actually track

The instinct is to manage this by accuracy percentage. That’s the wrong metric to run an operation on.

Detection accuracy on that system was deliberately stopped at about 90%. Pushing to 97% was possible and would have taken weeks of additional work, and it wouldn’t have changed the process. An estimate never leaves the building on numbers nobody has verified, so that check happens no matter how good the detector gets. What changes is how long it takes.

So the number to track is review minutes per set. At 90%, an estimator scans the output, fixes what’s wrong and moves on in about ten minutes. That’s the figure that sets throughput, and it’s the one worth putting on a dashboard.

A second thing worth measuring is the direction of the errors. A system that flags too much costs deletions, which are fast. A system that silently skips things costs a full manual check, which defeats the point. Two systems can report the same accuracy and produce completely different review times.

How to pilot it without guessing

If you’re evaluating this for your own operation, four steps are enough to get a real answer.

Take fifty of your own sets, including the dense and badly scanned ones. Vendor demo files will flatter any system.

Time the review by hand, before and after, on the same sets. Don’t estimate it from a percentage.

Count the two error types separately: what it invented and what it missed. Ask any vendor for both, not for one blended score.

Re-test when models change. New versions ship constantly, and results move in both directions. An evaluation from six months ago describes software that no longer exists.

What doesn’t change

The judgment stays with the estimator. Reading dimensions in context, spotting what a drawing implies, deciding what’s risky — none of that is automated today, and our benchmark didn’t measure it. What the technology takes off the desk is the counting, the page-turning and the searching.

That’s usually most of the clock. Across the manufacturing deployments we’ve worked on, taking it away shifts the limit from the quoting desk to production capacity, which is familiar ground for anyone running a plant.

The strange part, for a business that measures everything on the floor, is that the quoting desk usually has no numbers at all. Time it for a week before you decide what to do about it. The measurement alone tends to settle the argument.

Written by Morgan Ellis

Contributing writer covering productivity, operations, and small-business workflows for Flowster.