← All writing
Software Strategy Oct 2026 15 min read

The Safety Record Moved Three Times. AI Gave It a New Writer.

Every new home for the safety record fixed one thing the last one couldn't do and broke one it could. The buyer saw the fix. The site absorbed the break.

The Safety Record Moved Three Times. AI Gave It a New Writer.
Illustration: AI-generated · How this site is made

A paper near-miss card has a margin. A worker can draw an arrow showing where the forklift came round the racking and write "blind corner — mirror gone since March" beside it. The same event in an EHS platform, typed on a phone after a run of mandatory dropdowns, reads: Category: Vehicle / Mobile Plant. Location: Warehouse B. Description: near miss forklift. The mirror detail never got typed. A photo of the corner can be attached, but no report counts it.

Vendors now sell AI for almost everything. For the record itself, the pitch is specific: speech turned into text, a photo turned into filled fields, a category picked automatically. That promises to recover both losses. Where the worker did type the detail, AI can scan the free-text box no dashboard counts. Where they didn't, a voice note can capture it again. The detail can come back. The review work moves to whoever has to check what the AI filled in.

Every safety record has three jobs. Capture: get down what happened, from the person who did the work. Find: get it back when you need it. Trend: see the pattern no single record shows. Paper, the spreadsheet, and the platform each gave the record a new home, and each new home did one job better and another worse. AI is the first change that leaves the record where it is. It gives the record a new writer instead.

Paper: first-hand, but you cannot find anything

On paper, the person who did the work writes the record: the worker's near-miss card, the supervisor's inspection checklist, the risk assessment, the trainer's attendance sheet. They write it at or near the place it happened, in their own words. That is its strength: paper does the capture job better than the spreadsheet or the platform that replaced it. Nobody between that person and the record decides what matters.

Most paper forms never said much. Most near-miss cards said "near miss forklift" too; most checklists were a column of ticks. But paper had room for the one person who wanted to say more: a note beside a ticked item, a sketch on the back.

The bottleneck is finding it again. The form goes into a tray, then a binder, then an archive box. To find every near miss at that corner, someone reads every card. Trending breaks with it. The monthly tally counts cards by type. It does not show that three of them were at the same corner. That depends on whoever happens to remember.

Who pays: the manager setting priorities, who never sees that the corner came up three times; the worker whose card went into a tray and produced no action; and the investigator, after the serious event, going through boxes to rebuild a history nobody could see while it was building.

Spreadsheet: searchable, but you cannot trust it

The spreadsheet moved the record into a file and made finding fast. Filter a column and every forklift event appears.

Capture split in two. Desk records went straight into the file: the training matrix, the action tracker, the chemical register, monitoring results from the lab. Field work did not change. The inspection still runs on a paper checklist; the near miss still goes on a card. For those, the record now lives twice. Someone types the paper version into the file later: the inspector or a supervisor at the end of the shift, an administrator at the end of the week. Often they type from memory, not from the form. The sketch does not survive the retyping. The filter is fast, but it searches the retyped copy. The margin note is still in the binder, where no filter reaches it.

The category column now holds "Forklift," "FLT," "fork lift," and "MHE." Tick "Forklift" in the filter and you get one of the four.

Nothing in a spreadsheet stops this by default. You can add a dropdown list to a column, but pasting in a row from another file wipes it out, and nothing tells you. No rule says a row must point to a real location. Nothing keeps it to one file either. On a shared network drive, when one person has "Incident Log FINAL v3" open, the next person gets a read-only copy, saves it as "Incident Log FINAL v3 (2)", and adds their rows there.

This is not a safety-specific weakness. Field audits compiled by Raymond Panko found errors in 88 percent of 113 spreadsheets audited between 1995 and 2007. Those were business spreadsheets built by professionals. A safety log is simpler, with fewer formulas to get wrong. But the few it has are the ones that reach the board: a SUM formula that stops one row short drops a lost-time case from the report. Excel marks the cell with a small green triangle in its corner. Nobody reading the printed report ever sees it. And the log is updated by a supervisor at the end of a twelve-hour shift, not by an analyst at a desk.

Who pays: the EHS manager compiling the monthly report, who reconciles two versions of the file and still cannot be sure the count is right; and the board, deciding on a lost-time count that is one case short.

Platform: consistent, but you cannot change it

The platform moved the record into a database, in the vendor's cloud or on your own servers. Corporate wanted one count across every site, and auditors wanted to see who changed what. The platform delivered both, and finding and counting finally worked. Dropdowns force one spelling. Mandatory fields stop a record being saved half-empty. Every incident links to a real site and a real date, and to a real person whenever that person is on the payroll. For the first time, the count at the top of the dashboard matches the records underneath it.

The price is that the structure is fixed. The project team decides the data model—the set of fields, categories, and links every record must fit—at implementation, usually in workshops, usually before anyone has used the system on site. Some changes are quick: an administrator can add a dropdown value in minutes. But the category list usually belongs to corporate and is shared by every site, because the group dashboard compares sites on it. One site cannot add a category without asking to change everyone's. New fields and links go through the vendor or an implementation partner. They get scoped, priced, and scheduled into a release. The reports built on the old field list then need retesting.

So the category list stays as it was at launch, and the work moves on without it. A new hazard type has no category, so workers pick "Other" or the nearest wrong one. The form has become an error trap: the wrong answer is the only one it offers. The free-text description field still exists, but no dashboard tallies it. A photo can be attached, but no report counts an image either. Trends stop at the dropdowns.

Field capture got better tools and a rigid form. A phone can do what paper did and more: the record is made at the job, with a photo, a voice note, the location, and the time attached, and it goes straight into the platform with no retyping. The limit is the form, not the phone. It was designed in those same workshops, and it often asks for the category, the subcategory, and the severity before it asks what happened. A worker in gloves, at the end of a task, taps through the mandatory fields and types the minimum in the description. The record saves complete. And the phone does not go everywhere paper went: some sites ban personal phones, and hazardous areas need certified devices. Where the phone could not go, people went back to paper, and an administrator typed up the backlog at the end of the week.

Who pays: the worker on shift, who fights the form; the analyst, whose categories no longer describe what is happening on site; and the manager reading a dashboard that is consistent and wrong.

AI: a new writer, not a new home

AI does not move the record. It still lives in the platform, or in the spreadsheet someone exports into a chatbot. What changes is who writes it. Picture the AI built into your EHS platform: the worker submits a note or a photo, the software fills in the incident fields, and a manager approves the result. The buyer wants the hours back. A supervisor who spent an evening writing up an investigation now gets a draft in a minute.

Across the first three homes, the record was written by the person who did the work, then by whoever typed it up, then by the person on a phone or an administrator clearing the backlog. The person who did the work stopped being the author long before AI. A supervisor copying scrap-paper notes or a WhatsApp thread into the form at the end of a shift is a copy typist, not a witness. But a copy typist only drops detail. AI adds conclusions: a root cause and a corrective action that nobody on site wrote. Now AI writes the record and a person approves it. Approving is a different task from writing. An author has to decide what happened. An approver only has to find nothing obviously wrong.

AI helps all three jobs without moving the record. Capture: the form fills itself. Whether the raw input is a worker's voice note, an unformatted field note, or a ten-second camera clip of a forklift near miss, AI can turn it into a category from your existing list, a summary, a draft root cause, and a draft corrective action. It fixes the rigid form, not the rigid list. The worker can say "mirror gone since March" again, and the system will capture it. The margin is back, as long as the worker's own words are kept next to the summary. Find: you can search by meaning, not by the exact word in the category column. Trend: AI scans the free text no dashboard could count.

The field still sets limits. A voice note recorded beside a running forklift, or through a respirator, comes back with words missing. A contractor speaking a second language gets transcribed into something close to what they said. The summary is built on whatever the transcript holds, gaps included.

Speech also fixes only one reason the detail went missing. The missing mirror is an equipment fault, and nobody gets blamed for reporting it. A worker who left a colleague's name out of a typed report left it out because writing it down had a cost. Saying it costs the same, and a voice note is spoken out loud, often within earshot of the supervisor or the crew it describes. AI brings back the detail the form squeezed out. It does not bring back the detail the worker decided to keep out.

AI also pulls in whatever the site's existing systems hold, and most sites run all of them at once: permits on paper, risk assessments and the training matrix in spreadsheets, incidents in the platform. Every crossing between them is a retype or a paste, and whatever was wrong on one side arrives intact on the other.

Every one of those outputs is a head start, not a finding. Someone has to check it before it becomes one. The hours the buyer wanted back are the hours the check needs. That is the new bottleneck: checking what the AI wrote. It breaks down at two scales.

1. One record: reading takes a minute, checking takes the whole investigation

The output is fluent and plausible. Modern interfaces try to speed up review by highlighting source sentences next to the generated summary. That makes reading faster, but it does not verify the findings. Checking whether the AI's root cause is right still means finding the closed permit in the permit file, opening the maintenance system, checking the training spreadsheet, and cross-examining witness statements that were never fed into the software. None of those sources talk to each other. Reading a three-paragraph draft takes under a minute. Checking it means gathering the same evidence the investigation needed. That takes as long as the investigation the draft was meant to replace. And an approver cannot hear what the transcript dropped unless the audio is linked to the draft.

Approval itself is not new. Managers have countersigned supervisor investigations for decades, and plenty of those signatures went on without anyone reading the report. What is new is that the draft no longer shows you where it is thin. A half-filled card looked half-filled. A rushed write-up read as rushed. A fluent AI summary never does: the grammar is the same whether the facts behind it are complete or missing.

Human factors research has a name for what happens next. Parasuraman and Manzey's 2010 review of automation bias (people accepting an automated system's output without checking it) found it in novices and experts alike. Training and instructions did not prevent it. The same review found that its close relative, complacency (no longer checking a system that is usually right), was strongest when the person was handling several tasks at once and the automated task competed with the others for attention. An EHS manager approving AI-drafted investigation summaries between site walks, audits, and a contractor induction is working in exactly those conditions.

The original studies came from cockpits and lab tasks. The pattern holds in medicine too. Goddard and colleagues' 2012 systematic review in the Journal of the American Medical Informatics Association found that when decision-support software gave wrong advice, clinicians made more wrong decisions than clinicians working without it. An EHS approval queue has no feature that makes it immune.

Adding a second AI pass to check the first does not remove this bottleneck. The second pass only sees the inputs the first one was given, so it misses the same operational context. Automated field checks catch a blank field. They cannot catch a wrong root cause in a fully completed record.

A vendor will say the AI can pull the permit and the maintenance log itself. On most sites it cannot: the permit is a paper sheet hanging at the job, and the maintenance system was never connected to the EHS platform. Where the connection does exist, the software still pulls in a permit number someone retyped and a category someone picked because the right one was missing.

2. Many records: the list and the trend look right and are not

Start with search. An investigator rebuilding the history of the loading dock asks for every near miss there and gets four. The fifth was written up as "aisle 7, by the roller door." A search by meaning often finds it. When it does not, the list looks just as complete. Nobody can check a list for the record that is not on it.

Trends fail the same way. A platform's categories are fixed before the incident happens. AI assigns the category when the record is submitted. The vendor updates the AI, or someone tweaks the instructions behind it, and the same kind of report starts landing somewhere else. A near miss on a wet floor by the loading dock went to "Slips and trips" last quarter. This quarter it goes to "Housekeeping." Both are defensible. Every approver who looks at one record sees a reasonable category and signs. The slips line on the board report drops, and nobody slipped less.

People disagree about categories too. Two supervisors will file the same wet floor differently. But people disagree in scattered ways, and their habits change slowly. An AI update moves every record in the same direction on the same day. When corporate changes the category list, it gets announced. An AI update often is not.

Checking records one at a time cannot catch either failure, however careful the approver is. Only a check across records can: the same question asked of the search and of a plain keyword filter; the same reports run through the old AI setup and the new one, side by side.

Who pays: the approver, who signs for an investigation they did not do; the next investigator, who relies on an approved record nobody checked; and the manager reading a trend that moved because the software changed.

One person pays every time

Where the record livesCaptureFindTrendWho pays
PaperCard, binder, boxStrong: first-hand, own wordsWeak: someone reads every cardWhoever remembersThe manager setting priorities; the worker who reported; the investigator
SpreadsheetA file on a shared driveSplit: desk records typed straight in; field forms still paper, then retypedFast, but the filter misses variantsCounts you cannot trustThe EHS manager compiling reports; the board reading them
PlatformA database, cloud or on-premiseBetter tools, rigid form: media and location attached, dropdowns firstStrong: consistent and linkedDropdown counts onlyThe worker on shift; the analyst; the manager reading the dashboard
AIStill the platformThe form fills itself from voice, photo, clip; nobody sees what the transcript droppedSearch by meaning; you cannot see what it missedScans free text; moves when the AI updatesThe approver; the next investigator; the manager reading the trend

Read the table by row and each home trades one job for another. Read the last column and one person never leaves it: whoever decides from the numbers. Paper hid the pattern from them. The spreadsheet gave them a count that was one case short. The platform gave them a consistent count of the wrong categories. AI gives them a trend that moves when the software does. The buyer is usually that person. They saw the fix, and they pay for the break later, in the numbers.

AI also changes the size of a job that was always there. When a supervisor wrote the investigation, they gathered the evidence and the approver reviewed their work. When AI writes it, nobody has gathered the evidence, so the approver has to. Each new system's sales pitch describes the bottleneck it removes and stays silent on the one it adds.

What to specify before AI writes your records

You cannot remove the check. You can make it visible and size it honestly. Most EHS platforms are sold as one product to every customer, so the audit trail is whatever the vendor built. No steering committee can make the vendor add a field for one customer. What your organisation does control is whether the AI feature is switched on, and when: it is usually a separate module or a setting the customer turns on. Before it is, get the vendor to state in writing which of the records below the product keeps today. Every "no" is a check you will run by hand, or a reason to leave the feature off until a release adds it. Requirements 5 and 6 need nothing from the vendor. Only someone who has run an investigation knows that the root cause usually sits in the permit and the maintenance log, not in the incident form, and that is where the check has to reach. Timing matters: a record that was not stamped with the AI version and approver when it was written cannot be stamped later, so every record created without that data stays unauditable.

For one record

  1. Record who wrote each field. Every field the AI fills must be stored with the software version that generated it, the person who approved it, and the approval timestamp. A record where you cannot tell AI-written text from human findings cannot be audited.
  2. Keep the raw input as the anchor, never replaced. The raw input is the record. The AI draft is only one reading of it. The voice recording, the photo, and the worker's original words must stay attached to the AI output for as long as the output is kept. The margin note survives even when the AI summary gets it wrong.
  3. Log approval time and what was opened. Record the time between the draft opening on screen and the Approve click, and which source records linked from the draft the approver opened. An approval faster than the evidence can be read was not checked. A slower one proves nothing on its own, and evidence on paper never shows up in the log. That is what requirement 6 is for.

For many records

  1. Lock the AI version for trend data. Every category the AI assigns must carry the version that assigned it, so the trend chart shows where the software changed. When an updated AI reclassifies old records, that is a new trend line. Label it with the new software version, and leave the old line as it was. If the product does not stamp the version, put advance notice of every AI update into the contract, and mark each one on the trend chart. Before an update goes live, run it on last quarter's records and compare its categories with the old version's.
  2. Check the search against a plain filter. Each quarter, ask the AI search and a plain keyword filter the same questions an investigator would ask. Every record the filter finds and the search misses is a record the search will miss again.
  3. Compare against investigations people already do in full. Your standard already requires a full investigation for serious and high-potential events. Use those as the sample. The investigator works from the evidence and does not open the AI draft until the investigation closes; then the two findings are compared. The disagreement rate is your real measure of system reliability. The sample is not random: it leans towards serious events. That is where a wrong root cause costs most, and the investigation was happening anyway, so the only new work is the comparison.

These requirements cost hours. That is the point: they put the check back into the budget the AI was bought to cut. Put each one in a named person's job, not in the project plan. The project team leaves after go-live, and the first shutdown drops every check nobody owns.

AI gives the record a new writer, faster than any before it, and leaves the accountability with the person who clicks Approve. Their click becomes someone else's number. If your specification does not test whether approved records hold up when someone rebuilds them from the evidence, you have bought the fastest engine yet built for producing records that nobody verified.

Serhat Demirkol
Serhat Demirkol

A decade running management systems on-site, then seven years leading product for enterprise EHS software. Builds the tools, then writes about why most of them fail.

Keep reading