Artificial intelligence for plan review, automated code checking, machine learning applications, and limitations. Covers maintaining human oversight and quality assurance.
2
hours
0.2
CEUs
Administrative, Legal & Management
1.7.4
Artificial intelligence for plan review, automated code checking, machine learning applications, and limitations. Covers maintaining human oversight and quality assurance.
Format
On-Demand Online
Delivery
Self-Paced
Access
24/7 After Enrollment
Certification
Certificate of Completion
Have questions about this course or our platform?
Contact our support teamEvaluate AI and automation tools for code compliance applications
Before a department can evaluate any "AI" product, it needs a vocabulary that separates two things marketing language routinely blurs together. The first is deterministic automation: software that follows rules a human wrote and produces the same output every time for the same input. The second is AI-class tooling: software that produces probabilistic output from patterns learned in training data, which means it can be impressively right and confidently wrong on the same afternoon. The risks and supervision requirements differ, so the first procurement question is which one is actually being offered.
Deterministic automation is mature and already carries a large share of the workload in modern permitting operations. Completeness screening at intake — checking that an application has an owner name, a legal description, a contractor license number, a site plan attachment — is a rule check. Fee calculation from an adopted fee schedule is arithmetic. Routing a mechanical permit to the mechanical reviewer, sequencing a plan through building, fire, and zoning review, and sending automatic status notifications are workflow rules. The *Building Department Administration* textbook devotes a full chapter to this kind of information technology and frames its purpose plainly: reduce permitting time, eliminate duplicate data entry, let customers apply and pay online, improve record keeping, and free scarce staff hours for work that actually requires judgment. None of that is artificial intelligence, and none of it should be evaluated as if it were — when these functions fail, the failure is traceable to a rule and fixable.
AI-class tools are appearing in a narrower set of roles, and an honest evaluation names them precisely. Plan-review assist tools scan drawings and flag possible issues — a door that may not meet clearance, an assembly that may lack a required rating — for a human examiner to confirm or dismiss. Chatbots answer routine process questions ("Do I need a permit for a water heater replacement?") from the department's own published FAQs and handouts. Document classification and data extraction tools sort incoming applications and pre-fill fields from the documents themselves. Transcription and summarization tools turn a recorded meeting or a dictated inspection note into draft text. Every item on that list produces a draft, a flag, or a suggestion — an input to a human decision, not a decision.
Evaluating a specific product honestly means refusing to accept the vendor demonstration as evidence — a demo runs on plansets the vendor chose. The only test that matters is a pilot on your own submissions: pull already-reviewed plansets where your examiners' findings are known, run the tool against them, and count what it flagged correctly, what it flagged wrongly, and — most important — what it missed. Ask what data trained the tool and whether it resembles the construction types, code editions, and local amendments in your jurisdiction; a tool trained on one region's commercial work may perform very differently on your housing stock. Ask how it behaves when uncertain: does it say so, or present a guess with the confidence of a certainty? The textbook's broader counsel on new technology applies directly — the profession has a responsibility to consider new tools, but each must be intelligently and objectively evaluated before it is relied on, and the reason for adopting any system should be a specific problem it demonstrably solves, not the technology itself.
A department with a six-week plan-review backlog is offered a plan-review assist tool. The demonstration is flawless. Instead of signing, the building official asks for a ninety-day pilot: staff select twenty recently completed plansets spanning the department's real mix — custom homes, tenant improvements, a small mixed-use project — where every examiner finding is documented, and compare the tool's output line by line against the human record. The results are instructive: the tool reliably catches missing information and certain repetitive geometric checks, produces steady false flags on anything involving local amendments, and misses issues that depend on project context. The department adopts it for what the pilot proved — a first-pass screen whose every flag is verified by an examiner — writes that limitation into procedure, and documents the pilot as the basis for the decision. That evidence-based scoping is the difference between adopting a tool and surrendering to one.
The most common evaluation mistake is buying the label instead of the function — paying an AI premium for rule-based workflow software, or conversely trusting probabilistic output as if it were a deterministic rule check. The correction is to require the vendor to state in writing which functions are rule-based and which are model-based, and supervise accordingly. A second mistake is evaluating against the vendor's demonstration rather than the department's own plansets; the correction is a structured pilot with known-answer submissions. A third is accepting quoted accuracy figures without asking what they were measured on — a statistic derived from another jurisdiction's submissions or another code edition tells you nothing about performance at your counter. A fourth is measuring only what the tool flags and never what it misses; a quiet tool is not necessarily a right one. Finally, departments skip the boring questions — data ownership, records retention, what happens to uploaded plans, exit costs if the vendor folds — that IT and legal advisors should answer before any submission data leaves the building.
Implement automated code checking with human validation
One principle governs every implementation decision, and it is non-negotiable: the code official's judgment cannot be delegated to software. The approval of a permit, the interpretation of a code provision, and the result of an inspection are acts of public authority performed by accountable human beings. An AI tool may flag, sort, draft, and suggest; a qualified person decides. No jurisdiction operates on any other basis — a vendor claiming its product "approves permits" is describing something that does not exist in responsible practice. Implementation is the discipline of wiring that principle into daily workflow so it survives busy weeks and staffing shortages, not just policy documents.
In practice, human validation means the automated check runs first and a qualified reviewer disposes of every flag with one of three outcomes: confirmed against the adopted code and written into a correction notice in the examiner's own analysis; dismissed as a false positive with a brief note of why; or escalated as a genuine interpretation question. The record that goes to the applicant is the examiner's, not the machine's — an examiner who signs a correction notice must be able to defend it at the counter and before the board of appeals, and "the software said so" is not a defense anyone should ever have to offer.
Implementation also requires written staff-use policies before the first login, and the essential ones are short. First, no confidential or personally identifiable information goes into public AI tools — a planset, an applicant's personal data, a complaint file — because anything pasted into an outside service may be retained beyond the department's reach; tools handling submission data must be under an agreement the jurisdiction's IT and legal advisors have reviewed. Second, every code citation an AI tool produces is verified against the adopted code before it is used — always, without exception. Language models are known to generate plausible-looking citations to sections that do not exist or do not say what is claimed; the habit of opening the book is the entire defense. Third, staff must understand the records implications: an AI-generated draft that enters the project file is public record like anything else, and the department should be able to disclose, transparently, when automated tools contributed to a document. Quiet reliance staff would be embarrassed to disclose is a signal the use itself is wrong.
Finally, implementation is staged: run a new checking tool in parallel with the existing process for a defined period, compare outputs, and expand its role only as actual performance — on this jurisdiction's work — earns it. The *Building Department Administration* textbook's guidance on adopting permitting technology fits exactly: research and evaluate before committing, phase complex systems in deliberately, and expect that procedure, not the purchase itself, is where implementations succeed or fail.
Two moments from the same week in a department piloting a plan-review assist tool show both sides of the ledger. On Tuesday, the tool flags a corridor on a tenant-improvement planset as deficient in width. The examiner does not copy the flag into a correction notice; she pulls the adopted code, confirms the occupant load calculation, and finds the tool applied a threshold that does not govern this occupancy — a false positive, dismissed with a one-line note. Passed through unverified, that flag would have become an incorrect correction that burned the applicant's time and the department's credibility. On Thursday, the same tool flags a missing rated assembly detail on the fortieth sheet of a large submittal — the kind of repetitive check that slips past a tired examiner late in a long review. The examiner verifies it against the code, confirms it is real, and writes the correction. Both moments are the validation loop working as designed: the tool is a second set of eyes that never gets tired and is sometimes wrong, and the examiner is the decision-maker who must stay sharp enough to tell the difference.
The gravest implementation error is rubber-stamping — nominally requiring human review while workload pressure turns "validate every flag" into "forward every flag." The correction is structural: dispositions are recorded, false-positive rates tracked, and supervisors audit samples, because a validation step nobody checks will quietly stop happening. A second error is letting AI-drafted text reach applicants unedited; outgoing documents must carry the examiner's own verified analysis, whatever tool produced the first draft. A third is staff pasting sensitive material into public tools because no one told them not to — a policy gap, corrected by a short written policy issued before access is granted. A fourth is skipping the parallel-run period, which discovers the tool's failure modes on live applicants instead of on a controlled comparison. A fifth is keeping no records of how the tool is used, leaving the department unable to explain its own process when an appeal or records request asks.
Understand limitations and liability of automated systems
The liability frame for automated systems is simpler than vendors and worried staff both assume, and it cuts one way: the jurisdiction owns the decision regardless of the tool. When a permit is issued, a plan approved, or an inspection passed, the responsibility rests where it always has — with the code official and the jurisdiction — whether the supporting analysis came from a checklist, a spreadsheet, or a machine-learning model. No procurement contract transfers the duty to enforce the code onto a software company. This is not a reason to avoid the tools; it is the reason the human-validation discipline of the previous module is mandatory rather than optional.
Understanding limitations starts with the characteristic error modes of AI-class systems, which differ from ordinary software bugs. The first is false confidence: these systems present their weakest guesses with the same fluency as their strongest conclusions, so the output itself gives no reliable signal of when to be suspicious. The second is fabricated specificity — most dangerously, hallucinated code citations that name a section, sound authoritative, and are wrong or nonexistent. The third is brittleness outside the training data: a tool tuned on conventional construction may fail quietly on an unusual assembly, an alternative design, or a local amendment it never saw. The fourth is staleness — a model trained on one code edition does not know what a later edition changed. None of these failure modes announces itself, which is precisely why verification must be unconditional rather than reserved for outputs that "look wrong" — looking wrong is what these errors specifically fail to do.
The workforce question deserves an honest answer rather than a slogan. Some routine work genuinely is being automated — completeness checks, data entry, document sorting, drafting boilerplate — and pretending otherwise insults the staff who can see it happening. But the realistic near-term picture for building departments is augmentation of scarce examiners, not replacement: most departments cannot hire enough qualified plans examiners and inspectors as it is, and tools that compress the routine fraction of the work extend the reach of the professionals a department already has. The judgment work — interpretation, alternative-means evaluation, the conversation at the counter, the call in the field — is exactly what the tools cannot do, and it is what justifies the profession's authority.
The final limitation lives in the staff rather than the software: skill atrophy. An examiner who has leaned on an assist tool for years must still be able to review a planset unassisted — tools fail, contracts lapse, unusual projects fall outside what the tool can see, and a professional who can no longer independently do the work cannot meaningfully validate a machine's version of it. Treat unassisted competence as a maintained skill: training, certification, and periodic reviews done the long way are the foundation that makes supervised automation safe.
A jurisdiction two years into using an assist tool faces an appeal. The applicant argues a correction notice was invalid because "a computer wrote it." The department's position is comfortable, because its procedures anticipated the question: the record shows the tool produced a flag, a named plans examiner verified it against the adopted code, the notice carries the examiner's own analysis, and the examiner is at the hearing to defend the determination — which she can, because she made it. The board upholds the correction. The counterfactual is the lesson: had the department been unable to show a human determination behind the notice, the hearing would have gone very differently, and the tool's actual accuracy would not have mattered. The authority of the department's decisions rests on the demonstrable exercise of qualified human judgment; the tools are welcome exactly as far as they support that, and not one step further.
A recurring mistake is believing, or letting staff believe, that using a well-known tool dilutes responsibility — "the system flagged it" as a shield. The correction is cultural and procedural: every decision has a name attached, and staff are evaluated on their determinations, not the tool's suggestions. A second mistake is trusting fluent output because it reads confidently; fluency is not evidence, and verification applies even to outputs that look obviously right. A third is deploying a tool trained on a different code edition or building stock without testing for the mismatch. A fourth is managing the workforce transition by silence while staff privately fear replacement — the honest augmentation case, openly made, reflects reality and preserves morale. A fifth is letting unassisted skills decay unmeasured; the correction is deliberately scheduled unassisted work and continued investment in certification and code training, so the humans validating the machine remain genuinely qualified to overrule it.
This course provides comprehensive professional development in ai and automation tools for code compliance. Participants learn to distinguish deterministic automation — completeness screening, fee calculation, routing, notifications — from AI-class tools that produce probabilistic flags and drafts requiring human validation. The governing principle: the code official's judgment cannot be delegated. AI flags, humans decide, and the approval, the interpretation, and the inspection result remain human acts of authority for which the jurisdiction is accountable regardless of the tool. Participants develop practical skills in honest tool evaluation (pilots on known plansets rather than vendor demonstrations), staff-use policy (protecting confidential information, verifying every AI-provided citation against the adopted code, understanding records implications), and workforce stewardship — using automation to extend scarce examiners while maintaining the unassisted competence that makes supervision meaningful.