Guide

How to set up alarm escalation in a plant so warnings reach someone who acts

Updated · 10 min read · By the ThinklytixAI team

Alarm escalation is a simple rule: if the first person told about a problem does not respond within a set time, the next person is told, and so on until someone takes ownership. This guide explains why alarms get missed, the building blocks of an escalation policy for machine alerts, how to design an alarm escalation matrix with an example, and how to keep alarm fatigue under control.

Why alarms get missed in plants

When a machine fails after a warning hours earlier, the alarm usually existed. Nobody acted on it. The causes are familiar.

  • Too many alarms. If a screen shows dozens of active alarms on a normal shift, operators learn to ignore them.
  • Nobody owns the alarm. An alarm that goes to "everyone" or to "maintenance" in general goes to no one. Each person assumes someone else has seen it.
  • Alarms live only on an HMI screen. The HMI or SCADA screen is useful when someone is standing in front of it. At night or during a break, it often is not.
  • Shift handover gaps. An alarm raised at the end of one shift is acknowledged to silence it, and the next shift never hears about it. The acknowledgement hides the problem instead of passing it on.
  • No follow-up. Even when someone acknowledges, nothing records what was done. The same alarm returns next week and nobody knows whether it was checked.

Escalation deals with the second, third and fourth causes directly. The first and fifth need alarm management discipline, which is covered later in this guide.

Alarm vs alert vs notification

These words are often used as if they mean the same thing. It helps to separate them before you design anything.

TermWhat it meansWhere it usually appearsNeeds action?
AlarmA process or equipment condition that needs an operator response within a defined timeHMI, SCADA, DCS, control panel, horn or beaconYes
AlertA message sent to a named person to make sure an alarm or abnormal reading gets a responseSMS, WhatsApp, Slack, Teams, phone call, emailYes, by a person who may not be at the machine
NotificationInformation worth knowing but not needing a response, such as a batch completed or a shift reportEmail, dashboard, daily summaryNo

A good rule: if nobody needs to do anything, it is not an alarm. Moving pure information out of the alarm list is one of the quickest ways to make real alarms stand out.

The building blocks of an escalation policy

Every working escalation policy for machine alerts has the same parts. If one is missing, the chain breaks.

Priority or severity levels

Group alarms into a small number of priorities, for example critical, high and low. Base it on the consequence of not responding and the time available.

An owner for every alarm

Each alarm, or each group of alarms, needs a named role that owns the first response: the line operator, the shift engineer, the electrical technician. Use roles, not individual names, and keep a roster that says who holds each role on each shift.

Acknowledgement

Acknowledging means "I have seen this and I am taking it." It should be a deliberate action by a person, and it should be recorded with who and when.

Timed tiers

Each tier is a group of people who are told at a set time after the alarm starts, if nobody has acknowledged. Tier 1 is told at once. Tier 2 is told after a delay. Tier 3 after a longer one.

Channels

Each tier has a channel: a screen, a messaging app, a phone call. Later tiers often need louder channels, because the reason they are being told is that the quieter ones did not work.

Stop on acknowledge

Once someone acknowledges, escalation must stop. Otherwise senior people get woken for problems already in hand, and soon ignore the messages.

An audit trail

Every alarm should leave a record: when it started, who was told, who acknowledged, when, and what was done.

How to design an alarm escalation matrix, step by step

An alarm escalation matrix is the table that brings these parts together.

  1. List the alarms that matter. Start with alarms that stop production, damage equipment or affect quality. A short, reliable list beats a long, ignored one.
  2. Assign a priority to each. Ask two questions: what happens if nobody responds, and how long before it happens?
  3. Name the first owner. Pick the role closest to the machine who can actually act. For each role, name a backup for leave and absence.
  4. Set tier delays from the response time. Each delay must be shorter than the time the process gives you. If a bearing temperature alarm gives you roughly an hour before damage, a 30-minute wait before tier 2 is already too long.
  5. Choose the channel per tier. Tier 1 can be the HMI plus a message. The last tier should be hard to miss.
  6. Decide what counts as resolved. Write down whether the owner must add a note, and whether the alarm must clear before the record is closed.
  7. Test it on a quiet shift. Trigger a test alarm, let it escalate without acknowledging, and check each person actually received it.
  8. Review after real events. After the first month, look at alarms that reached tier 2 or tier 3 and ask why tier 1 did not respond.

Example escalation matrix (illustrative)

The table below is an example only. The delays and roles will be different in your plant, depending on your process, staffing and the time each fault gives you.

PriorityTier 1 (at 0 min)Tier 2Tier 3
Critical (e.g. compressor trip, boiler pressure high)Line operator and shift engineer: HMI, WhatsAppMaintenance lead after 5 min: WhatsApp and phone callPlant head after 15 min: phone call
High (e.g. motor temperature rising, low hydraulic oil)Shift engineer: HMI, Slack or TeamsMaintenance lead after 20 min: WhatsAppProduction manager after 60 min: WhatsApp
Low (e.g. filter differential pressure, lubrication due)Maintenance planner: Slack or TeamsMaintenance lead next working day: daily summaryNone

Keep safety-critical alarms separate

Escalation software sits alongside the plant's control system. It supplements, and never replaces, the plant's safety instrumented systems, interlocks, trips and the written operator response procedures for safety alarms.

Safety-critical alarms must be handled where they are designed to be handled: on the control and safety systems, with the automatic actions and local responses they already have. A WhatsApp message or a phone call can tell a manager that a safety alarm happened, but the plant must never depend on that message reaching someone for the process to be safe. If unsure whether an alarm is safety-critical, treat it as one and check with your process safety owner.

Choosing channels: HMI, SMS, WhatsApp, Slack or Teams, phone call

The right channel is the one your people actually see during their shift. Each has strengths and weaknesses.

ChannelGood forWatch out for
HMI or SCADA screenOperators at the machine; full process contextOnly works when someone is watching the screen
SMSWorks on basic phones and weak data coverageEasy to miss among other messages; hard to acknowledge back into a system
WhatsAppAlready on most phones on the shop floor; read quicklyPersonal phones on silent; group messages get buried if the group is noisy
Slack or TeamsOffice and engineering teams; shared channels keep a historyMany floor staff do not have it open; not ideal at night
Phone callThe last tier of critical alarms; hardest to ignoreIntrusive, so it must be reserved for alarms that truly need it

Two practical tips. First, send alerts to named people or small groups rather than a large group where everyone assumes someone else will reply. Second, make the message readable in one glance: machine, tag, value, limit and how long it has been active.

Avoiding alarm fatigue in plants

Escalation makes a bad alarm system louder, not better. If operators already ignore alarms on the screen, sending the same noise to their phones will train them to ignore their phones. Alarm fatigue has to be fixed at the source.

  • Rationalise the alarm list. Go through each alarm and ask: what should the operator do when this fires, and how long do they have? If there is no clear answer, change it to a notification or remove it.
  • Add deadbands and short delays. A reading that hovers around its limit will set and clear an alarm again and again. A small deadband, or a short on-delay before the alarm fires, stops this chattering without hiding a real problem.
  • Suppress consequential alarms. When a pump trips, low flow, low pressure and several downstream alarms follow. Show the root alarm and suppress the ones that are only a consequence of it while it is active.
  • Review the top 10 noisy alarms every week. Rank alarms by how often they fired. Fix, retune or remove the worst few each week.
  • Handle stale alarms. Alarms that stay active for days become background. Find out why they are still active and either fix the cause or change the alarm.

Recognised alarm management standards and guidance, such as ISA-18.2 (also published internationally as IEC 62682) and EEMUA 191, describe a full lifecycle for alarm systems, including philosophy, rationalisation, monitoring and review. Your control system supplier or process safety team can help you apply them.

Closing the loop after an alarm

An escalation that ends at acknowledgement is only half the job. The other half is learning from it.

  • Record what was done. Ask the person who acknowledged to add a short note: what they found, what they did, and whether a follow-up is needed.
  • Review unacknowledged alarms. Any alarm that ran through every tier without a response is a failure of the system, not just of a person. Check whether the owner, the delay or the channel was wrong.
  • Look at alarms that reached tier 2 or 3. If the same alarm keeps escalating, the tier 1 owner may be overloaded, or the alarm may not be theirs to act on.
  • Carry open alarms across shift handover. Put active and recently acknowledged alarms on the handover sheet, so the incoming shift knows what is still in progress.
  • Feed repeat alarms into maintenance. An alarm that fires every week is a maintenance problem in disguise. Our guide to reducing unplanned downtime shows how to track stops and fix the biggest cause first, and predictive vs preventive maintenance explains when a warning sign should turn into a planned check.

Where software helps with alarm escalation

A small plant can run escalation with a roster, a phone tree and a logbook. With more machines and shifts, keeping it consistent by hand gets hard.

MIE, the Manufacturing Intelligence Engine, handles this part. Live today, it takes the readings your machines already send over HTTPS, MQTT, OPC-UA or Modbus, with no new sensors, and sends alerts on Slack, WhatsApp or webhook. Each rule has a list of tiers, each with a channel, recipients and a delay. The alert moves to the next tier until someone acknowledges it, then escalation stops and the acknowledgement is recorded. Plant managers can also use Ask to put plain-English questions to their plant data. Your control system keeps running the plant; MIE watches the data and escalates problems to people.

Next on the MIE roadmap is predictive maintenance early warnings, which will flag a reading drifting from normal before it becomes an alarm. That is coming soon and is not live yet.

Whatever tools you use, the method is the same: fewer, better alarms, a named owner for each, timed tiers that stop on acknowledgement, and a weekly review of what was missed.

Questions

Common questions.

What is an alarm escalation matrix?
It is a table that says, for each alarm priority, who is told first, through which channel, and who is told next if nobody acknowledges within a set time. It turns an alarm from a screen message into a chain of named people.
How long should you wait before escalating an alarm?
It depends on how much time the process gives you to respond. Set the delay for each tier shorter than the time it takes for the problem to cause damage or stop production, and review it after real events.
What is the difference between an alarm and an alert?
An alarm tells an operator that a process condition needs a response, usually on the HMI or control system. An alert is a message sent to a person, often off the control system, to make sure someone acts. A notification is information that needs no action.
How do you reduce alarm fatigue in a plant?
Remove or fix alarms that need no action, add deadbands and short delays to stop chattering, suppress alarms that are only a consequence of another alarm, and review the ten noisiest alarms every week.
Can escalation software replace safety alarms?
No. Safety-critical alarms must stay on the plant's control and safety systems with their own operator response procedures. Escalation software only adds a way to reach people; it does not protect the process.
Should machine alerts go to WhatsApp or SMS?
Use the channel your team actually reads on shift. Messaging apps suit routine escalation well, while a phone call is harder to miss for the highest tier. Many plants use one channel for the first tier and a louder one for the last.
See it on a real plant

Ask the demo plant your own question.

We'll walk you through MIE on a live plant in 30 minutes.

Or call +91 98288 93692 · +91 88519 85656 · info@thinklytixai.com

Talk to us