AI-Assisted NDT Report Generation: What It Actually Automates (and What It Doesn't)

AI can draft a report narrative and flag inconsistent fields in seconds. It cannot interpret an indication or sign your name to it. Here's the actual boundary.

By Anoop Rayavarapu, ASNT NDT Level III ·

The pitch and the reality don't match, and that's worth being precise about

Every software vendor in the NDT space, Atlantis included, has some version of an AI-assisted reporting pitch right now, and it's worth an ASNT Level III being blunt about what that actually means in a field where a wrong call on a report has consequences measured in structural integrity, not just customer satisfaction. AI-assisted report generation is real and it's useful. It is also, right now, fundamentally a drafting and consistency-checking tool — not an inspector, not a Level II, and not a substitute for the judgment a written practice requires a certified person to apply. Understanding exactly where that line sits matters more in this industry than in almost any other software category, because the artifact being generated is a certified technical record that someone's signature stands behind.

What AI actually does well in a reporting workflow

Strip away the marketing language and the genuinely useful applications cluster around a handful of specific, bounded tasks — all of which involve organizing or restating information a qualified person has already generated, not creating new technical judgments.

Boilerplate automation: pulling from records instead of retyping

A large share of any NDT report is administrative, not interpretive: equipment serial numbers and calibration dates, technician certification numbers, procedure and revision numbers, weather conditions for outdoor RT or MT work, reference standard IDs. None of that requires judgment — it requires accurate retrieval from records that already exist elsewhere, typically in an equipment log, a personnel certification record, or a procedure library. Automating that retrieval so a technician doesn't manually retype a calibration block ID that's sitting in an equipment database three systems away is a straightforward, low-risk automation that eliminates a genuinely common source of transcription error, not a judgment call being delegated to software.

Draft narrative from field notes and structured readings

Turning a technician's shorthand field notes, or a set of structured UT thickness readings, into a readable narrative paragraph — "Examination performed on the six circumferential welds of Vessel V-2101 per procedure UT-114 Rev. 3, no reportable indications noted" — is a genuinely good use of current generative AI. It saves real time on a task that is more clerical than technical: turning data that already exists into report language that reads clearly, without the technician spending twenty minutes wordsmithing a paragraph that says the same thing a table already says.

Consistency and completeness checking

This might be the single highest-value application, and it's underappreciated because it's not flashy. A system that checks a completed report against the cited procedure and flags contradictions — an acceptance statement that doesn't match a recorded indication amplitude that exceeds the stated acceptance criteria, a calibration date entered that falls after the examination date, a required field left blank — catches exactly the kind of error that a rushed reviewer misses at 6 p.m. on the last day of a turnaround. This isn't AI making a technical call; it's AI checking arithmetic and logical consistency against rules a human already defined, which is a fundamentally different and much safer category of automation.

Where the line sits: interpretation and disposition

Here is the boundary that matters, and it isn't subtle. Deciding whether an ultrasonic indication is geometric reflection or a genuine planar flaw, whether a magnetic particle indication is a relevant discontinuity or surface condition that doesn't require rejection, whether a radiographic density variation represents porosity within acceptable limits or a linear indication requiring rejection under ASME Section VIII acceptance criteria — these are interpretive judgments that a written practice under ASNT SNT-TC-1A specifically assigns to personnel certified and qualified to make them, generally at Level II or above, with a Level III accountable for the procedure and technique those judgments are made against. No current AI system should be making that call, and no inspection company should be structuring its workflow so that a model's suggestion becomes the de facto disposition because a rushed technician accepts whatever the software drafted without independently verifying it against the actual data.

This isn't a temporary limitation waiting on a better model. It's a structural feature of how NDT accountability works: a certified person's signature on a report is a professional and often legal attestation that they applied their qualified judgment to that specific examination. Software drafting language around a decision doesn't change who made the decision or who's accountable for it — and a report that reads as if AI made an interpretive call, however well-written, is a report where accountability has gotten fuzzy in exactly the place it can't afford to.

The hallucination risk is not hypothetical here

Generative AI models are, by design, prediction engines that produce plausible-sounding text — which is exactly the property that makes them good at drafting narrative language and dangerous if used carelessly around technical data. A model asked to summarize a set of readings can, under the wrong conditions, produce a summary that sounds authoritative and is subtly wrong — smoothing over an outlier reading that should have been flagged, or generating acceptance language that doesn't actually match the numbers in the table next to it. This is precisely why AI-assisted drafting needs a human review step built into the workflow as a hard gate, not an optional courtesy — the reviewer isn't just checking grammar, they're verifying the generated language actually reflects the underlying data before anyone signs anything.

The liability question: who signs, who's accountable

Every report, however it was drafted, gets signed by a specific person who is professionally and often legally accountable for its accuracy. That doesn't change because a model wrote the first draft of the narrative paragraph. What it does change is the review discipline required before signing: a technician or Level II reviewing an AI-drafted report needs to verify it against the source data with the same rigor as reviewing a report a colleague drafted by hand — arguably more rigor, until the technology and the workflow around it have a longer track record inside a given company's operations. Treating an AI draft as "probably fine, it's just formatting" is the exact failure mode that turns a productivity tool into a liability.

Companies building AI into their reporting workflow well tend to structure it explicitly: AI drafts, populates, and flags; a qualified human reviews, corrects, and disposes; a Level II or III signs. The AI never sits in the disposition chain itself — it sits alongside it, doing the parts of the job that were always clerical rather than technical.

Where this is genuinely heading versus where the hype says it's heading

The realistic near-term trajectory is better boilerplate automation, better consistency checking, and better transcription from field data capture — mobile and offline data entry feeding directly into structured reports without a re-typing step, which is where the real time savings compound across a full inspection program. What isn't realistic, and shouldn't be promised by any vendor, is AI making acceptance/rejection calls independent of qualified human review in the near term. The technical, regulatory, and — frankly — insurance and liability structures around certified NDT reporting aren't built to accommodate that, and treating a drafting tool as if it were shouldn't be how any inspection company evaluates reporting software, however impressive the demo looks.

Structuring human review as a hard gate, not a courtesy

The difference between an AI reporting feature that helps and one that quietly introduces risk usually comes down to workflow design, not model quality. A hard gate means the report cannot reach a client-visible or finalized state without an explicit, logged review action from a qualified person — not a passive assumption that "someone probably looked at it." Concretely, this looks like: the system clearly marks every AI-generated field or narrative section as unverified until a human reviewer has opened and confirmed it, the reviewer's confirmation is logged with their name, credential level, and timestamp the same way a signature is, and the system does not allow a report to route to final approval with any AI-generated content still in an unconfirmed state. This is a deliberately higher bar than reviewing a colleague's manually drafted report, and it should be, until an inspection company has enough internal track record with AI-assisted drafting on its own reports to calibrate exactly how much independent verification each type of AI-generated content actually needs.

Training reviewers to catch AI-specific error patterns

Reviewers who've spent years catching the kinds of mistakes humans make — a transposed digit, a missed field, an inconsistent date — aren't automatically primed to catch the kinds of mistakes AI-assisted drafting introduces, which tend to look different: a narrative sentence that's grammatically perfect and reads authoritatively but subtly overstates what the data actually shows, or a summary that smooths over an outlier reading into something that sounds more consistent than it is. Training reviewers specifically on what AI-generated content errors tend to look like — plausible-sounding but unverified claims, rather than obvious typos — meaningfully improves how effective the human review gate actually is in practice, rather than assuming existing review habits transfer automatically to a new kind of draft.

A realistic checklist for evaluating an AI reporting claim

Inspection companies being pitched AI-assisted reporting features, whether from Atlantis or any other vendor, are well served asking a short list of specific questions rather than accepting "AI-powered" as self-explanatory:

  • Exactly which fields or sections does the AI populate, and which remain entirely human-entered?
  • Is there a hard technical gate preventing a report from finalizing with unreviewed AI content, or is review merely recommended?
  • Does the system log what was AI-generated versus human-entered, distinguishably, for future audit purposes?
  • Does the vendor make any claim, explicit or implied, about AI performing interpretation or disposition — and if so, does that claim survive a direct question about who's professionally accountable when it's wrong?
  • How does the system handle a case where AI-generated narrative language doesn't match the underlying structured data — does it flag the mismatch, or does it just produce fluent text regardless?

The honest way to evaluate an AI reporting feature is to ask what specific clerical task it removes, not whether it can "generate a report." Generating a report was never the hard part of this job. Getting the interpretation right, and being accountable for it, always was — and that part still belongs to a qualified person, not the software drafting around them.

How this plays out differently for Level I, Level II, and Level III staff

AI-assisted drafting changes each certification level's daily work differently, and it's worth an inspection company thinking through that distinction explicitly rather than treating "the technicians" as one undifferentiated group. For a Level I technician performing examinations under direct supervision, AI-assisted transcription and boilerplate population is close to pure upside — it removes clerical burden without touching any judgment they weren't making independently in the first place, since Level I work is already supervised and reviewed by definition. For a Level II making interpretation and initial disposition calls, the risk profile is different: a Level II working fast, under schedule pressure, reviewing an AI-drafted summary that reads smoothly, is in exactly the position where it's easiest to rubber-stamp a draft rather than independently re-derive the interpretation from the raw data. This is the level where explicit training on slowing down for AI-assisted reports — treating the draft as a first guess requiring the same independent verification as any other technician's early-career interpretation, not as a peer's finished work — matters most.

For a Level III, the responsibility doesn't change at all, and that's worth stating plainly rather than leaving implicit: the written practice, the procedures, and the final technical authority over disposition remain a Level III function regardless of what drafting tools sit underneath the reporting workflow. If anything, a Level III's role expands slightly to include periodic auditing of how well the human review gate is actually functioning across the team — not just reviewing individual reports, but checking whether AI-assisted drafts are getting genuinely verified or quietly waved through, which is a pattern that only shows up by sampling across many reports over time, not by reviewing any single one.

A note on client disclosure

Clients increasingly ask directly whether AI was involved in generating a report they're being asked to accept, and inspection companies are better served answering that question plainly than treating it as a competitive secret. Explaining exactly what the AI assisted with — boilerplate population, narrative drafting, consistency checking — and exactly what a qualified human independently verified before signing, tends to build more confidence than either overselling the technology's role or being evasive about using it at all. Clients who understand the actual workflow, gate, and accountability structure generally have no objection to AI-assisted drafting; what erodes trust is discovering, after the fact, that a claimed human review was more of a formality than a real check.

Atlantis NDT Products & Services

Atlantis NDT pairs field expertise with software: NDT inspection management software — Atlantis ERP, a digital twin platform for asset integrity, and NDT reporting software. Build your team with NDT training & certification (ASNT SNT-TC-1A) and ASNT certification pathways, or bring in ASNT Level III consulting. Affordable, accessible, fully customizable — book a free consultation.

When report turnaround is the bottleneck

Most inspection companies lose more hours to report formatting than to inspection. NDT reporting software compares the options for issuing the same dataset in several client formats without re-keying, the NDT inspection software buyer’s guide separates the four product categories that all get called “NDT software”, and the free evaluation checklist sets out the tests that actually separate marketing from capability.

Atlantis NDT Products & Services

Atlantis NDT pairs field expertise with software: NDT inspection management software — Atlantis ERP (certification tracking, work orders, method-specific reporting on every business app you need), a digital twin platform for asset integrity (3D corrosion mapping, API 581 RBI, API 579 FFS), and NDT reporting software. Build your team with NDT training & certification (ASNT SNT-TC-1A) and ASNT certification pathways, or bring in ASNT Level III consulting for RBI, FFS, and written practices — plus independent inspection data review on API 510/570/653-governed assets. Capture as-built reality with 3D laser scanning services. Affordable, accessible, fully customizable — book a free consultation.