Wise Hustlers — Digital Product & App Development Studio Logo
Get Consultation
By Wise Hustler Admin•9/23/2026•12 min read

Incident Investigation and Corrective Actions: The Report Nobody Reopens

Incident Investigation and Corrective Actions: The Report Nobody Reopens

# Incident Investigation and Corrective Actions: The Report Nobody Reopens

TL;DR: most HSE systems don't fail at the investigation stage — they fail halfway through, when a corrective action gets assigned to someone with no real verification date, and the nonconformity that triggered it lives in a separate record that nobody ever goes back to close.

A well-written incident investigation report, with a clear root cause and assigned actions, is the kind of document that looks good in an audit and is worth nothing six months later if nobody can tell, without asking three people, whether action 14 from that report was actually implemented. This is not a discipline problem among field teams. In most cases an engineering team runs into when designing these systems, it is a data model problem: the way the system records the finding, the nonconformity, and the corrective action as three separate objects instead of a single record with states.

Where the process breaks: the action that stays "assigned" forever

The pattern is always the same. An HSE inspection, an internal audit, or an incident investigation identifies a deviation — an expired fire extinguisher, a lockout-tagout (LOTO) procedure that wasn't followed, a missing guardrail on a work platform. Someone logs the finding, assigns a corrective action to a responsible person, sets a due date. Up to this point, the process works.

The problem shows up afterward. The action sits in "assigned" or "in progress" indefinitely, because:

  • the due date passed and nobody was notified, since the notification depends on someone opening the module and scanning the list;
  • the person who was supposed to close the action changed roles, and the action was never reassigned;
  • the action was physically executed — the extinguisher was replaced, the procedure was corrected — but it was never verified by someone other than the person who implemented it, so the system still shows it as open;
  • or, the most common case, the action was marked "completed" by the same person who implemented it, with no evidence attached, and there is no way to distinguish that from an action that is genuinely closed.

None of these cases is a matter of bad faith. They are the direct consequence of treating "closing a corrective action" as a status field anyone can change, rather than a transition that requires a second actor and evidence.

The fork nobody designed on purpose: one finding, two records

There is a second, subtler problem that shows up systematically in HSE systems built by accretion — an inspection checklist module added first, a "nonconformity" module added later to satisfy an ISO 45001 or operator-client audit requirement. When this happens, the checklist item marked "nonconforming" is not, itself, the nonconformity. Someone has to open a second record, in a different module, to formally describe the nonconformity and attach a corrective action to it.

The practical result: two entries for the same fact. The checklist item shows "nonconforming" and stays frozen there, because the closure workflow lives in the other module. The nonconformity record moves forward, gets investigated, the corrective action gets closed — and nobody goes back to the original checklist to update its status. During an audit, an assessor cross-referencing the two modules finds an inconsistency that doesn't exist in operational reality: the inspection still "reports" a deviation that was actually fixed months ago.

This is not a cosmetic detail. ISO 45001 Clause 10.2 treats exactly this as a single process — report, investigate, act — not as three independent activities that each live on their own screen. When the software splits apart what the standard treats as one continuous process, it creates manual reconciliation work that, in practice, nobody has time to do consistently.

Why the response to the checklist item must BE the finding

The structural fix is conceptually simple, even if it is rarely trivial to implement on a system already in production: the record created when a checklist item is marked nonconforming should not spawn a second object — it should be the finding. A single record, with a unique identifier, born the moment the inspector marks "nonconforming," carrying forward through its lifecycle every subsequent state:

1. Open — finding created from the checklist item or the investigation, with mandatory photographic evidence and description.

2. Root cause assigned — after an analysis (5 Whys, Ishikawa diagram, or a more formal cause tree for serious incidents), with the method used recorded, not just the conclusion.

3. Corrective action defined — with an owner, a due date, and an explicit verification criterion (what has to be observed to confirm the action addressed the cause, not just the symptom).

4. Action implemented — marked by whoever carried it out.

5. Verified — marked by a second person, typically the HSE lead or the auditor who opened the finding, with new evidence.

6. Closed — only reachable after the "Verified" state.

This design eliminates the fork because there is no second record to fall out of sync. The checklist item and the nonconformity are the same record seen from two angles: the inspection list always shows the current status of the finding that came from it, because it's a reference to the same database row, not a copy.

What this means for the data model

In schema terms, this implies that the "findings" table has an optional foreign key pointing to the checklist item or the incident investigation record that spawned it — never the reverse, and never a parallel "nonconformities" table with its own independent lifecycle. The state machine stays centralized in a single status field, with transitions validated server-side (not only in the UI), and every transition writes an immutable row to the audit trail — who changed the state, when, with what evidence attached. Without that server-side transition validation, someone will always find a way to jump straight from "Open" to "Closed" through a poorly protected API call or a bulk import, and the entire two-person verification design becomes worthless.

This split between client-side and server-side state validation is exactly the kind of architectural decision that only a system built around a team's actual operating rules can enforce consistently — an off-the-shelf checklist package solves data capture but rarely enforces the full state machine without extensive configuration. Designing that state machine and permission model around a real operating workflow is part of what custom software development work for industrial operations actually looks like.

Root cause analysis: record the method, not just the conclusion

A related problem, rarely handled with the same care as closing the action: root cause analysis is often logged as a free-text field — "cause: lack of preventive maintenance" — with no record of which method produced that conclusion, and no record of what alternatives were considered and ruled out. This has two practical consequences. First, it makes it impossible to audit the quality of the investigation later — there's no way to tell a well-substantiated root cause from an after-the-fact guess. Second, it makes it impossible to aggregate root causes over time to spot patterns, because one investigator's free text never matches another's, word for word, for the same type of cause.

The alternative doesn't require a complex system: a structured field recording the method (5 Whys, Ishikawa, cause tree) as an enumerated type, the intermediate analysis steps as sub-records linked to the finding, and only then the final root cause as a field classified from a controlled list (equipment failure, procedure failure, competency failure, supervision failure, and so on). This structure is what later lets someone ask "how many findings in the last twelve months have a root cause classified as procedure failure" without rereading every report.

Closing a corrective action is not checking a box

This is worth restating plainly, because it's the point where most systems — and the manual Excel processes that precede them — fail in practice: closing a corrective action cannot be a button that the person responsible for the action also controls. Verification by a second person is not extra bureaucracy; it is the only thing that makes the closure record mean anything. A system where the same user who implements the action also closes it is indistinguishable, for audit purposes, from a system with no verification at all — even if a "verified by" field technically exists.

This connects directly to the system's permission design: separation of duties between whoever implements and whoever verifies is only real if the system prevents, at the permission level, the same account from executing both transitions on the same finding — not merely by convention documented in a procedure nobody rereads during an audit.

Indicators worth watching

The point of tracking operational HSE indicators is not to produce a good-looking number for a monthly report — it's to detect, early, that action closure is falling behind schedule before that becomes visible in an external audit or, worse, in a repeat incident. Without inventing benchmark figures that don't exist publicly and verifiably for the Angolan context, the structural indicators a well-designed system can produce directly from the data model described above include:

  • Average age of open actions, by owner and by operating unit — not just the total count of open actions, which hides old actions behind new ones.
  • Percentage of actions closed within the defined due date, calculated from the verification date, not the date someone marked "completed."
  • Root cause distribution by structured category, over time, to identify whether the same type of failure (procedure, equipment, competency) is recurring across different findings that, in isolation, would look like independent events.
  • Findings with no corrective action assigned past an internally defined number of days — the most direct symptom of an action that got lost between investigation and assignment.

None of these indicators requires sophisticated statistics. They require that the underlying data model not let a finding "disappear" between modules — the same argument that underpins the rest of this article.

In Angola, the obligation to organize occupational safety and health services is not an optional best practice — it is set out in the General Labour Law (Lei n.º 12/23, of 27 December 2023) and detailed in specific instruments: Decree n.º 31/94, of 5 August, which establishes the principles of the occupational safety, hygiene and health system, Executive Decree n.º 6/96, which approves the General Regulation of Occupational Safety and Hygiene Services in Companies, and Presidential Decree n.º 179/24, of 1 August, which regulates the licensing of occupational safety, hygiene, and health services within companies. For petroleum operations specifically, Decree n.º 38/09, of 14 August, approves the Regulation on Safety, Hygiene and Health in Petroleum Operations.

None of these instruments prescribes the design of a software system — but all of them implicitly assume that a consistent record exists of findings and the actions taken to correct them. A system where an inspection shows an unresolved deviation from six months ago, when the corrective action was actually implemented and verified in a different module, does not hold up well under an Inspecção Geral do Trabalho inspection or an audit from an international operator demanding evidence of the full closure cycle. It's the same reasoning already covered when discussing inspections, PPE, and incident records with an auditable trail: the value of an HSE system isn't in producing good-looking reports, it's in maintaining a record that still makes sense six months after it was created.

For teams still running this cycle on shared spreadsheets, it's worth looking at the same structural problems discussed in migrating petroleum operations from spreadsheets to an ERP: the fork between checklist and nonconformity described above is, structurally, the same fork between the inspections sheet and the corrective actions sheet that any team that has ever tried to reconcile the two manually recognizes immediately.

Frequently asked questions

Does a "nonconforming" checklist item and a formal "nonconformity" always have to be the same record?

Not in every case — a nonconformity can originate from an external audit, a client complaint, or an incident investigation without ever passing through a checklist. But when the origin IS a checklist item, that item should generate the finding record directly, by reference, not by duplication. The structural rule is: one origin, one record, multiple states — never a second parallel object for the same origin.

Who should be allowed to close a corrective action?

Never the same person who implemented it, and ideally someone with a distinct role in the permission system (HSE lead, internal auditor, or a second, independent technician), with that restriction enforced by the system itself and not just by written procedure.

Is it worth migrating a legacy system with this problem right away, or patching it in place?

It depends on how much debt has accumulated. Fixing the state machine on top of an existing schema is possible when the problem is mostly procedural (missing second verification); when the problem is structural — separate tables for checklist, nonconformity, and corrective action with no shared foreign key — a sustainable fix almost always means reworking the data schema, not adding one more manual synchronization field.

Does this apply only to HSE, or also to quality and local content audits?

The same fork pattern shows up in any compliance module that combines checklists with a separate nonconformity workflow — quality audits, local content verifications, or EPC contractual requirements. The principle — the response to the item is the finding, not a second record — carries over directly.

Sources

Related articles