Can ChatGPT Do a Steel Takeoff? How to Test It
ChatGPT can read parts of a structural drawing, but it can't give you a takeoff you can bid from. It will list member sizes it sees and often sounds sure of itself. What it won't give you is a complete count across every sheet, the same answer twice, or a way to check where each number came from. For bid work, that last part is the deal-breaker.
That's the honest short answer, and it holds for the other general chat tools too (Gemini, Claude, Copilot). They're useful to an estimator, just not for the counting. Rather than hand you our own claims, this post shows you how to run a fair test on your own drawings in about an hour, what to look for, and where chat AI earns a place in your workflow.
Can ChatGPT read a structural steel drawing?
Partly. Upload a PDF or an image of a framing plan and ask "what steel members are on this sheet?" You'll usually get back a list of designations like W16X26, W12X19 and HSS6X6X1/4, sometimes with a count.
On a clean, simple sheet, that list can look right. The problems start when you check it:
- Sheet coverage. On a multi-page set, chat tools often summarize a few pages and skip others without saying so. Small text on a 36x24 sheet is easy to misread or miss.
- Counts. The count is a number in a sentence. You can't click it and see the 14 beams it counted.
- Sizes. Look-alike characters cause real errors. A W18X35 read as W18X55 is 57 percent heavier per foot. An L3X3X1/4 read as L3X3X1 is a different section entirely.
- Lengths. Chat tools rarely produce reliable member lengths, because that needs scale, grid dimensions and judgment about where each beam frames.
- Made-up items. When unsure, chat models can invent plausible-looking members that aren't on the drawing.
Quotr's write-up on ChatGPT for construction estimating reaches a similar conclusion for general trades: helpful for writing and explaining, unreliable for quantity takeoff, with no audit trail back to the drawings. It's worth reading their reasoning, and then testing it on your own work rather than taking anyone's word for it, ours included.
How do you run a fair test on your own drawings?
Here's a test any estimator can run. Pick a job you've already bid, where you trust your takeoff.
Step 1: Pick the drawings. Use 5 to 15 structural sheets from one job: a couple of framing plans, a column schedule, and a sheet of details. Include at least one sheet with a lot of repeated beams.
Step 2: Write down your answer key. From your original takeoff, list the count by section for each sheet. This is your ground truth.
Step 3: Run each chat tool the same way. Use the same prompt for each tool, for example: "List every structural steel member on these drawings, with the section size, the count, and the sheet number for each one. Output a table." Upload the same PDF.
Step 4: Run it again. Start a new chat and repeat the exact same prompt and file. Do it three times per tool.
Step 5: Score it. Use a sheet like this:
| What to measure | How to score it |
|---|---|
| Sheet coverage | Which sheets did it report on? Any skipped? |
| Sections found | Of the sections in your answer key, how many did it list? |
| Sections invented | How many sizes did it list that aren't on the drawings? |
| Count by section | For the top 5 sections, how far off is each count? |
| Repeatability | Do the three runs agree with each other? |
| Lengths | Did it give lengths? Are any right? |
| Traceability | Can you tell which spot on which sheet each item came from? |
Step 6: Do the same with a purpose-built tool. Run the same PDF through a steel takeoff tool and score it on the same sheet. On SteelFlo's free plan you can do this on 3 takeoffs at no cost.
A few rules keep the test honest. Don't coach the chat tool between runs ("you missed sheet S-3"), because you won't be there to coach it on a real bid. Use the drawings as issued, not cropped. And score the misses as carefully as the hits, since a missed member costs you more on a bid than an extra one you delete.
If you want to skip ahead to the comparison, run the same drawing through SteelFlo. Every member gets a box you can click and check, which is the part a chat answer can't give you.
Where does chat AI actually help estimators?
Chat tools are good at language work, and estimating has a lot of language work. Places they genuinely save time:
| Task | How chat AI helps | What you still check |
|---|---|---|
| Scope letters | Drafts inclusions, exclusions and clarifications from your notes | Every exclusion matches your actual scope |
| RFIs | Turns a scribbled question into a clear, polite RFI | The question is technically right |
| Spec reading | Summarizes a long spec section, pulls out coating and testing requirements | The summary against the spec itself |
| Spreadsheet formulas | Writes the formula to extend weight or nest lengths | The formula on a few rows by hand |
| Explaining terms | Explains an unfamiliar callout or standard | Against the manual or code |
| Emails | Follow-ups, bid cover notes, change order letters | Tone and facts |
Notice the pattern: these are tasks where you can read the output and know if it's right in a few seconds. Takeoff isn't like that. Checking a chat-generated count means doing the count yourself, which defeats the purpose.
One caution on spec and drawing uploads: read your bid documents' confidentiality terms and your chat tool's data settings before uploading a client's drawings to a consumer chatbot.
Where does it break: page coverage, counts, repeatability, audit trail?
If you run the test above, these are the failure patterns to watch for.
Page coverage. A 40-sheet structural set is a lot of input. Chat tools may summarize instead of reading every sheet, and they rarely tell you which pages they skipped. On a takeoff, a skipped sheet isn't a small error. It's every member on that sheet.
Counts on dense sheets. A framing plan with 60 W16X26 beams is hard to count from an image. Expect the count to drift between runs.
Repeatability. Ask the same question twice and you may get two different answers. That's how these models work. For a bid number, it means you can't trust any single run.
Typical details and schedules. "Typical at all openings," a beam schedule keyed by mark, or "see detail 4/S501" requires cross-referencing that chat tools handle unevenly.
No audit trail. This is the big one. When the GC asks why you carried 46 W12X19s, "the chatbot said so" isn't an answer. You need to point to each piece on the drawing.
Why isn't a takeoff you can't check really a takeoff?
A takeoff is a list you're willing to sign a bid on. Every line on it has to survive two questions: where is this on the drawings, and did I get all of them?
A chat answer can't answer the first question at all, and it can't answer the second without you redoing the work. So in practice, a chat-generated takeoff is a rough draft that still needs a full manual takeoff to verify. That's slower than just doing the takeoff.
The same logic applies to any tool, including purpose-built ones. If software gives you a number you can't trace, treat it with the same suspicion. The guide to AI accuracy in steel takeoffs goes deeper on how to judge any AI tool's output, and can AI read steel drawings covers what kinds of drawings are easier or harder for any software.
What do purpose-built steel takeoff tools do differently?
From the user's side, the differences are about checking, not just reading. With SteelFlo, for example:
- Every member label on every page gets a box. The count is the number of boxes. Seven boxes labeled W10X15 means seven pieces, and you can click each one.
- You review instead of trusting. Confirm boxes, reject wrong ones, add anything missing. The finished takeoff is yours, checked against the drawing.
- Sizes are steel sizes. Designations are matched to real sections in AISC, BS/IS, AS/NZS, EN, GB or cold-formed SSMA, more than 4,500 in total, with weights attached.
- Revisions you can compare. Run a revised set and compare its counts by section against the original takeoff, with every changed count tied to boxes on specific sheets.
- Output you can use. Export a cut list as CSV or Excel, an Order Sheet nested to stock lengths, or price it in "Price This" with your own material, fab and erection rates.
It works best on vector PDFs exported from CAD, which is how most engineers issue drawings today. And it's still a tool for estimators: scope, typical details and pricing judgment stay with you.
For a broader look at the category, see AI structural steel takeoff software, manual vs AI steel estimating, and the AI steel estimating software buyer's guide.
The job itself isn't going away. The BLS describes cost estimators as people who read blueprints and technical documents to prepare estimates. Tools change how much of that time is spent counting. The judgment part stays with the estimator, and so does the reference material: the AISC Shapes Database for section properties and AISC 303 for what's in a structural steel scope.
Frequently Asked Questions
Can ChatGPT read blueprints?
ChatGPT can read some text and labels from an uploaded drawing, especially on clean, simple sheets. It struggles with large multi-sheet sets, small text and dense plans, and it can't show you where each item came from. Use it for questions about a drawing, not for a count you'll bid on.
Is ChatGPT accurate enough for a steel takeoff?
Not for bid work. Its counts can vary from run to run, it may skip sheets or invent members, and there's no way to trace a number back to the drawing without redoing the takeoff. Test it on a job you've already bid to see for yourself.
What can ChatGPT do for a steel estimator?
It's useful for language tasks: drafting scope letters and RFIs, summarizing spec sections, writing spreadsheet formulas and explaining unfamiliar terms. Those are tasks where you can check the output quickly.
What's the difference between ChatGPT and AI takeoff software?
Chat AI gives you an answer in text. Purpose-built takeoff software like SteelFlo puts a box on every member it finds, on every sheet, so you can review and confirm each piece. The result is a takeoff you can trace and defend, not just a list.
How do I test an AI takeoff tool fairly?
Use a job you've already bid, write down your counts by section as the answer key, and run the same PDF through each tool more than once. Score sheet coverage, sections found and missed, invented items, count accuracy, repeatability and traceability.
Bottom line
Chat AI is a good assistant for the writing side of estimating and a poor tool for counting steel. Run the one-hour test on a job you know, and you'll see the difference for yourself. Then put the same drawing through SteelFlo's takeoff software on the free plan and check each box against the sheet.