Sorting a Backlog by the Outage It Can Actually Be Done In
In water and wastewater the plant never stops, so a turnaround is one basin, one clarifier or one reservoir out while redundancy carries the load. Deficiency tracking has to sort the backlog by outage window and by access dependency — dewatering, confined space entry, bypass pumping, coating cure, disinfection clearance — because those, not labor hours, set the critical path.
A refinery turnaround is negotiated against production economics. A utility outage is negotiated against a permit. Influent keeps arriving whatever the schedule says, and a treatment train taken out of service has to be carried by the remaining trains without breaching effluent limits, which is why wet-weather season, peak summer demand and regulatory reporting periods close windows that engineering cannot reopen. The result is that a water utility's deficiency backlog is not one queue but many small ones, each attached to an asset that can only be reached during a specific and often infrequent outage. A finding on a buried reservoir may wait five years for its window. When that window opens, the constraint is rarely the repair itself. It is dewatering, ventilation and atmosphere clearance for entry, coating cure time at the temperature actually achievable inside the structure, disinfection under AWWA C652, and two consecutive passing bacteriological sample sets before the asset can be returned to service.
Source: Sources: AWWA D100 Welded Carbon Steel Tanks for Water Storage and AWWA D101 for inspection and repair of steel water tanks; AWWA C652 Disinfection of Water-Storage Facilities; AWWA C301 and C304 for prestressed concrete cylinder pipe design and inspection; AWWA M28 for water main rehabilitation; NSF/ANSI/CAN 61 for materials in contact with potable water; AMPP (formerly NACE and SSPC) standards for surface preparation, coating inspection and holiday detection, including SSPC-PA 2 for dry film thickness measurement; NASSCO PACP and MACP defect coding for sewers and manholes; OSHA 29 CFR 1910.146 permit-required confined spaces and 29 CFR 1910.119 process safety management, with chlorine and anhydrous ammonia threshold quantities listed in Appendix A; US EPA NPDES permitting under 40 CFR Part 122.
| Step in the outage | What actually sets the duration | Who controls it | What breaks if the plan counts labor hours only |
|---|---|---|---|
| Take out of service and dewater | Basin volume, available pump capacity and where the water is permitted to go | Operations and the discharge permit | Crew mobilizes to a full structure and stands down for a day or more |
| Ventilate, monitor and permit entry | Atmosphere clearance under the confined space program, and continuous monitoring for the duration | Safety and the entry supervisor | Entry is planned as a formality; H2S or oxygen deficiency delays the first shift |
| Clean, inspect and scope | Grit and residue removal before any surface is inspectable, and scope growth once it is | Contractor and inspection | Discovered scope has no approved funding path and stalls mid-outage |
| Surface preparation and coating | Ambient temperature, dew point and substrate condition inside an enclosed structure | Weather and the coating data sheet | Application halts on dew point; the window is consumed by waiting, not working |
| Cure to immersion service | Manufacturer's cure schedule at the temperature actually achievable, not at the datasheet's reference | Chemistry and the structure | Refill is scheduled on application completion and the coating is damaged on first fill |
| Holiday detection and thickness verification | Coverage of the full coated area, not a sample | Coating inspector | Defects are found after refill, and the outage repeats |
| Disinfection and bacteriological clearance | AWWA C652 method, contact time, and lab turnaround on consecutive sample sets | Laboratory and the state primacy agency | Return to service is delayed by days with the crew already released |
The plant never stops; only a train comes out
Every framework for planning inspection work assumes a shutdown. Water and wastewater utilities do not have one. Influent arrives at whatever rate the weather and the population dictate, and finished water demand has to be met on the day it occurs. The nearest equivalent to a turnaround is taking one process unit out — a clarifier, an aeration basin, a filter, a digester, a storage reservoir, one side of a duplicated pump station — and asking the remaining units to carry the load.
That single difference reshapes the deficiency backlog. Work is not sorted by priority alone; it is sorted by which unit it sits on and when that unit can be released. Two findings of identical severity on adjacent structures may be separated by four years of scheduling because one sits on a basin that can be isolated any month of the year and the other sits on a reservoir that can only come down in a narrow shoulder season.
So a backlog presented as a single ranked list is actively misleading in this sector. What the maintenance planner needs is the backlog sliced by outage opportunity: everything that can be done during the next filter outage, everything waiting on a reservoir drawdown, everything that requires bypass pumping on a force main. Findings that are not attached to a feasible window are not scheduled work. They are a wish list.
The window is set by permit and by season
The limits on a utility outage are external. An NPDES permit sets effluent quality that must be maintained continuously, and taking a treatment train out reduces the margin available to meet it. Wet weather closes windows entirely at many plants, because peak flow events need every basin in service and the forecast horizon is shorter than the outage. On the potable side, peak summer demand closes reservoir outages for months at a stretch.
This makes window availability a fact to be discovered rather than negotiated, and it changes who owns the schedule. In an industrial turnaround, engineering proposes and operations reacts. In a utility, operations declares what is possible and engineering fits the scope inside it. A deficiency system that models outages as project schedules, freely movable, misrepresents how the decision is actually made.
The practical requirement is that the register knows the outage calendar, and that findings attach to outages rather than to dates. When operations moves a basin outage from October to April because of a wet forecast, everything scoped into it should move with it, and anything that no longer fits should surface immediately as a re-deferral decision rather than being discovered on the morning the crew arrives.
Cure, disinfection and clearance own the critical path
Ask a planner how long it takes to recoat the interior of a steel water storage tank and you will usually get an answer in working days for blasting and application. That answer is the smaller half of the outage. Coatings for immersion service cure on a schedule set by the manufacturer's data sheet at a reference temperature, and the temperature inside an enclosed structure in a cold month is not the reference temperature. Cure that reads as five days on paper can be twice that in practice, and applying immersion service before full cure damages the coating on first fill.
Application itself is weather-bound in a way that surprises people who plan indoor work. Surface preparation and coating are governed by ambient temperature, relative humidity and the margin between substrate temperature and dew point. In a covered reservoir with limited ventilation, dew point control can consume more of the window than application does, and dehumidification equipment is itself a mobilization item that has to be scoped in advance.
Then comes clearance. Dry film thickness verification under SSPC-PA 2 and holiday detection cover the whole coated area, not a sample. Disinfection of a potable water storage facility follows AWWA C652 with a defined method and contact time, and return to service depends on bacteriological results that come back on the laboratory's clock, usually as consecutive sample sets. None of these is labor. All of them are duration, and a plan built on crew hours contains none of them.
Access dependency is the field most backlogs are missing
Look at a typical utility deficiency record and you will find asset, description, severity, estimated cost and a status. What you will not find is how the work is physically reached. Yet access is what determines whether two findings can share an outage, what the mobilization costs, and how long the window has to be.
The dependencies worth making into structured fields are few and highly repetitive. Does this require dewatering, and of what volume. Does it require permit-required confined space entry. Does it require bypass pumping or temporary treatment. Does it require scaffolding or rope access. Does it need materials certified to NSF/ANSI/CAN 61 because the surface contacts potable water. Does it require disinfection and clearance before return to service. Six yes-or-no fields, answered when the finding is raised by the person who saw the asset.
The payoff is immediate and arithmetic. With those fields populated, the planner can ask the backlog which findings share a dewatering event, which ones justify a single confined space setup, and which ones will extend the window regardless of how small the work is. Without them, the same questions require reading every record's free text, which is why in practice nobody asks them and outages get scoped from memory.
PACP grades, and what a grade does not tell you
Collection systems already solved part of the normalization problem that plagues other sectors. NASSCO's Pipeline Assessment Certification Program gives CCTV operators a defined defect vocabulary and a one-to-five severity scale, separating structural defects from operations and maintenance conditions and rolling individual defects into segment ratings. Two contractors coding to PACP produce comparable output, which is more than most inspection disciplines can claim.
The limitation is that a PACP grade describes the pipe, not the consequence of the pipe failing. A grade 5 structural defect in a short run under an open field and a grade 4 under a hospital access road or crossing a creek are not the same problem, and no coding standard can tell them apart because consequence is not visible in the video. Utilities that rank rehabilitation purely by grade end up funding the wrong segments and can defend the decision, which is worse than not being able to.
A deficiency register earns its keep here by carrying the second dimension: criticality of the segment, based on diameter, depth, surface use, proximity to receiving water, redundancy and customer impact. That value is assigned once, at the asset, not repeatedly at each finding. Combining a coded defect grade with a stored criticality gives a rehabilitation ranking that survives a council presentation, and it is reproducible year to year as new CCTV comes in.
Deferred findings and an aging clock that counts outages
Everything in a utility backlog is deferred at some point, and most deferrals are correct. The reservoir cannot come down until spring, the basin is needed through the wet season, the funding is in next year's program. Deferral is not a failure of discipline; it is the normal operation of a system with fixed windows and an annual budget.
What fails is the tracking of it. A finding deferred into the next outage is typically written into a note, and the note is read by whoever plans that outage, if they know to look. When staff change, when a consultant does the planning, or when the outage is delayed twice, the note stops being read. Findings from the previous internal inspection of a reservoir routinely resurface in the next one, five years later, as new findings, because nothing carried them forward.
The fix is to make deferral a first-class action that requires a target — the next outage of a named type on a named asset — and to age findings in windows rather than days. A finding that has missed two qualifying outages is escalated automatically and visibly, regardless of its original severity, because missing two windows means something in the planning process is repeatedly discarding it. That single metric surfaces the systemic problem that a calendar-based aging report never will.
Fiscal years, project numbers and emergency declarations
A private operator can authorize a repair the day it is scoped. A public utility usually cannot. Capital work moves through a capital improvement program adopted annually by a council, board or commission, with a budget line and a project number, and the approval calendar has nothing to do with when an inspector finds a problem. Operating funds cover small work; anything of size waits for the cycle or requires an emergency declaration that carries its own scrutiny.
This means a utility deficiency register has a financial dimension that industrial registers do not. A finding needs to know which funding path it is on — operating budget, an existing CIP project, a future CIP request, or an emergency route — because that determines when it can be executed at least as strongly as its severity does. A backlog that shows only technical priority will repeatedly propose work that has no money attached, and the credibility cost of that lands on the engineering group.
It also creates a real analytical need at budget time. The engineering manager preparing next year's CIP request needs to show the backlog grouped by proposed funding year, with the consequence of each deferral stated in terms a non-technical board can weigh. A register that can produce that grouping directly turns budget preparation from a month of spreadsheet work into a query, and it makes the request defensible because every line traces to a dated inspection finding with evidence behind it.
Evaluating a system before the next window opens
Test it against a real outage you have already run, preferably one that overran. Load the findings that were in scope, the access dependencies each one carried, and the actual durations. Ask the system to reproduce the window. If it plans the outage on labor hours and returns a duration close to the original optimistic estimate, it will make the same error on the next one.
Then test the dependencies and the deferrals. Ask where dewatering, confined space entry, bypass pumping, cure and clearance live, and confirm they are structured fields the planner can filter and aggregate on. Ask what happens when operations moves an outage by six months. Ask what a finding deferred at the last reservoir inspection looks like today, and whether the planner of the next reservoir outage will see it without being told to look.
Finally, test the money. Ask whether a finding can carry a funding year and a project reference, and whether the backlog can be grouped that way for a budget submission. Atlantis builds deficiency and recommendation tracking as a module of an Odoo-based NDT ERP, alongside inspection reporting, digital twin and 3D laser scanning services for structures where access is the constraint. To walk a real outage through the model before your next window opens, contact info@atlantisndt.com.
Why is a water plant turnaround different from a refinery turnaround?
A refinery can stop making product. A treatment plant cannot stop receiving influent or supplying demand, so the outage is partial by definition: one train, one basin, one reservoir, carried by redundancy. That makes the constraint N-1 capacity against a permit rather than production loss against a margin, and it means windows are dictated by season and hydraulics. Scope that cannot be finished inside the window does not overrun; it is abandoned and re-deferred.
What access dependencies have to be fields rather than notes?
At minimum: does this require dewatering, does it require permit-required confined space entry, does it require bypass pumping or temporary treatment, does it require a coating cure period, does it require disinfection and bacteriological clearance before return to service, and does it require potable-contact materials certified to NSF/ANSI/CAN 61. Each is a yes or no that changes the outage arithmetic. Buried in a comment field, none of them can be planned or aggregated across a backlog.
How should a deferred finding be aged when the next outage is years away?
Not in days. A finding on a buried reservoir with a five-year internal inspection cycle will always look ancient on a calendar clock, and the alarm stops meaning anything. Age it in windows: how many outages of the required type have opened and closed since this finding was raised without it being executed. A finding that has missed two windows is a genuine escalation. One that has missed none is simply waiting.
How do PACP grades fit into a deficiency register?
NASSCO PACP grades individual defects one to five and rolls them into segment ratings, separating structural from operations and maintenance conditions. The value is that it is already a normalization layer, so CCTV work from different contractors is comparable. The limit is that a grade describes a defect's severity, not its consequence: a grade 4 in a low-consequence lateral and a grade 4 under a hospital access road are not the same problem, and the register has to carry the second dimension.
What does confined space entry change about scheduling?
More than most plans allow. Permit-required entry under 29 CFR 1910.146 needs an attendant, continuous atmospheric monitoring, a rescue capability on standby and a written permit for each entry, and in wastewater structures hydrogen sulfide and oxygen deficiency make clearance genuinely uncertain rather than procedural. Entry productivity is a fraction of open-air productivity, and every break, shift change and equipment retrieval is a re-entry. Estimating entry work at surface rates is the most common source of window overrun.
How does the municipal budget cycle change the workflow?
A private operator can authorize a repair the day it is scoped. A utility usually cannot. Capital work runs through a capital improvement program approved on an annual cycle by a council or board, so a finding raised in March may have no funding path until the next fiscal year unless it qualifies as an emergency. The register therefore needs a funding year and a project reference on the finding itself, or the backlog will contain work that is technically approved and financially impossible.
Built for any business that runs on operations
Most companies do not fail at their craft. They lose time, margin and goodwill in the gaps between the tools they use to run the place — a quoting spreadsheet that does not talk to the job sheet, a job sheet that does not reach accounts, and a compliance folder nobody can search when a client asks. Atlantis closes those gaps by putting the whole operation on one platform, so information is entered once and everything downstream stays in step.
What you can run on it
- Sales and CRM — leads, quotes, follow-ups and the pipeline that tells you what next month looks like.
- Projects and job costing — plan the work, track the hours and materials against it, and see the margin while the job is still live rather than at final account.
- Field and service teams — dispatch, schedules, mobile capture that works with no signal, and sign-off from site.
- Inventory and purchasing — stock, suppliers, reorder points and goods receipt, joined to the jobs that consume them.
- People — records, qualifications and licences with renewal reminders, timesheets, leave and payroll.
- Quality and documents — procedures and forms under revision control, with the audit trail an inspection or accreditation body actually asks for.
- Accounts — invoicing, expenses, multi-currency and the reporting your accountant stops chasing you for.
Affordable, accessible, fully customizable — and we mean each word
Affordable because the whole suite is included rather than sold to you a module at a time, and because implementation is done by people who have run operations rather than by a chain of subcontractors. Accessible because it runs in a browser and on a phone, works for a small team on day one, and does not need a specialist on staff to keep it alive. Fully customizable because your process is the thing that makes you competitive — the software should bend to it, not the other way round.
Industries we configure for
Service businesses and contractors, manufacturing and fabrication, trading and distribution, laboratories and testing houses, engineering consultancies, construction and facilities, and asset owners across energy, marine, aerospace and infrastructure. Inspection and testing is where we started, and it remains the sector we go deepest in — but the platform underneath is general-purpose, and most of what it does has nothing to do with inspection at all.
What happens when you get in touch
A short conversation, not a sales sequence. We ask how the business runs today and where it hurts, show you the platform doing that work, and send a written quote shaped to your region, your team size and the scope you actually need. No obligation, nothing to install first, and no pressure to decide on the call. Reach out and tell us what you are trying to fix.
Related: business management platform · inspection management software · choosing the right category of software · modules · by industry · asset integrity platform. Book a free consultation.