Guide

How to reduce unplanned downtime in manufacturing and process plants

Updated · 9 min read · By the ThinklytixAI team

To reduce unplanned downtime, record every stop with a reason, rank the reasons by the output they cost you, and fix the top few one at a time. In most plants that means making sure warnings reach a person who acts on them, moving critical assets to condition-based checks, standardising changeovers and reviewing the downtime log every week.

This guide is for plant owners and production, maintenance and operations managers in power and steam, process and discrete manufacturing. It covers what counts as unplanned downtime, the causes that come up most often, how to measure it with MTBF, MTTR and a Pareto chart, and a seven-step plan to reduce machine downtime that you can start this month.

What is the difference between planned and unplanned downtime?

Planned downtime is time the equipment is not running because you decided so in advance: preventive maintenance, statutory inspections, cleaning, planned changeovers, shift breaks or a scheduled shutdown. You can plan labour, spares and customer orders around it.

Unplanned downtime is every stop you did not schedule: a motor trips, a boiler feed pump fails, the line waits for packaging material, compressed air pressure drops, or a changeover runs two hours over. It is expensive because it arrives without warning, pulls people off other work and often causes knock-on losses such as scrap, restart losses and overtime.

The line between the two matters. If an overrunning changeover is logged as "planned", the overrun disappears from view. A simple rule helps: anything beyond the standard time for a planned activity counts as unplanned.

What causes unplanned downtime?

The causes of unplanned downtime differ by plant, but most stops fall into six groups.

  • Equipment failure. Bearings, seals, pumps, motors, valves, drives and instruments wear out or fail. Poor lubrication, misalignment and running outside design limits speed this up.
  • Early warning signs that nobody noticed. Most mechanical failures give signs first: rising vibration or temperature, a slow drift in pressure or current, longer cycle times. If nobody is watching the trend, the first sign anyone sees is the breakdown.
  • Alarms nobody acts on. An alarm that goes to an empty control room, a shared inbox or a phone on silent is the same as no alarm. Too many nuisance alarms teach people to ignore the real ones.
  • Changeovers. Product, size or recipe changes that depend on one experienced person, lack a checklist, or wait for tools and parts regularly overrun.
  • Material and utility shortages. The machine is fine but it has nothing to run on: raw material, packaging, steam, compressed air, cooling water or power.
  • Operator and process errors. Wrong settings, skipped steps, untrained cover on a shift, or a process running too close to its limits.

Notice that only the first group is purely a maintenance problem. The others are about information, planning and routine, which is why downtime is a whole-plant issue and not only a maintenance team target.

How do you track downtime properly?

You cannot reduce what you do not record. Good downtime tracking needs four things for every stop:

  1. Which machine or line stopped.
  2. Start and end time, so you get the duration.
  3. A reason from a fixed list, not free text. Keep the list short, around 10 to 15 reasons, grouped under the six causes above, with an "other" option that you review and split when it grows.
  4. What was done and by whom, so the log also records the fix.

Where machines already send run status to a PLC, SCADA or historian, the start and end times can be captured automatically and people only add the reason. Where they do not, a paper or spreadsheet log at each line is a perfectly good start. Short stops of a minute or two are easy to miss by hand, so decide a minimum duration for manual logging and stick to it.

If you want to go further and link downtime to overall performance, our guide on how to calculate OEE shows where availability losses sit next to speed and quality losses.

How do you calculate MTBF and MTTR?

Two numbers tell you a lot about a critical asset.

  • MTBF (mean time between failures) = total operating time ÷ number of failures. Higher is better. It tells you how often the asset breaks.
  • MTTR (mean time to repair) = total repair (down) time ÷ number of repairs. Lower is better. It tells you how long each failure costs you.
  • Availability ≈ MTBF ÷ (MTBF + MTTR).

The two numbers point to different fixes. A low MTBF says the asset fails too often: look at root causes, condition monitoring and operating limits. A high MTTR says each failure takes too long to fix: look at how quickly people find out, spares on the shelf, access, and repair instructions. Much of MTTR is often not the repair itself but the time before anyone knows the machine has stopped.

How do you find which downtime reasons matter most?

Once you have a few weeks of tagged stops, total the lost time (or better, lost output) by reason and sort from largest to smallest. This is a Pareto analysis. It usually shows that a handful of reasons account for most of the loss.

Reason (illustrative, one line, one month)Hours lostShareCumulative
Waiting for packaging material1230%30%
Feed pump trips1025%55%
Changeover over standard time820%75%
Compressed air pressure low615%90%
Other410%100%
Total40100%

In this example the top three reasons cover 75% of lost time, and only one of them is a machine fault. When lines run at different rates, convert hours to units lost (hours × output per hour) before ranking, so an hour on your bottleneck counts for more than an hour on a line with spare capacity.

How to reduce unplanned downtime: a seven-step plan

1. Measure and tag every stop with a reason

Set up the log described above on every critical line and asset. Agree the reason list with operators and maintenance together, so both sides trust it. Aim for every stop above your minimum duration to have a reason by the end of each shift.

2. Rank reasons by lost output

Run the Pareto every week or month. Pick the top one to three reasons and give each an owner and a date. Resist working on everything at once.

3. Fix alerting so warnings reach a person

For each important alarm, decide who owns it, on which channel they will see it (phone, WhatsApp, Slack, SMS) and what they should do. Then add escalation: if the first person does not acknowledge within a set time, it goes to the next person, then the supervisor or plant manager, until someone takes it. Record who acknowledged and when. At the same time, remove or retune nuisance alarms so the ones that remain are worth answering. This step alone often cuts MTTR, because the gap between a stop and someone arriving shrinks.

4. Move critical assets to condition-based or predictive checks

List the assets whose failure stops the plant: boilers and their feed pumps, compressors, main drives, the bottleneck machine. For these, watch the readings that drift before failure (vibration, temperature, pressure, current, flow) and act on the trend, not just the trip limit. Many plants already collect these readings in a PLC or historian and simply never look at the trend. Our comparison of predictive and preventive maintenance explains when each approach fits.

5. Standardise changeovers

Write down the best known changeover as a checklist. Separate tasks that can be done while the machine is still running (preparing tools, parts, material, settings) from tasks that need it stopped. Set a standard time, log actual time against it, and treat overruns as unplanned downtime.

6. Close the loop by recording outcomes

For every warning and repair, record what was found: fixed, scheduled for later, or nothing found. This tells you which warnings are useful and which are noise, and builds the history you need for better maintenance decisions. Without it, the same failure repeats and nobody can prove what worked.

7. Review weekly

Hold a short weekly review with production, maintenance and stores. Look at total downtime, the Pareto, MTBF and MTTR for critical assets, open actions and whether last week's fixes held. Keep it to 30 minutes and end with named owners and dates.

Cause, early sign and what to do

CauseEarly signWhat to do
Bearing or motor failureRising vibration or temperature over days or weeksTrend the reading, set a warning well below the trip limit, plan the replacement
Pump or boiler feed troubleDrifting pressure, flow or motor currentCompare against normal running, inspect seals and strainers, keep critical spares
Alarms not answeredUnacknowledged alarms, stops found at shift handoverName an owner per alarm, add timed escalation, remove nuisance alarms
Changeover overrunsActual time creeping above standard, results vary by shiftChecklist, prepare while running, log actual vs standard
Material shortageLow stock at line side, late supplier deliveriesReorder points for line-side stock, daily check with stores and planning
Utility shortfall (steam, air, power)Header pressure dipping at peak load, compressor running flat outMonitor utility headers, fix leaks, alert before pressure reaches the limit
Operator or process errorRepeated stops after shift change, settings outside rangeStandard settings, short training, interlocks or checks on critical parameters

Quick wins this month vs longer-term work

This month:

  • Start a downtime log with a fixed reason list on your most critical line.
  • Give every critical alarm a named owner and a backup.
  • Set up escalation for alarms that stop production.
  • Check that critical spares are actually on the shelf.
  • Write a checklist for your most frequent changeover.
  • Book a weekly 30-minute downtime review.

Over the next few months:

  • Capture run and stop status automatically from PLCs, SCADA or a historian. If you are choosing how to connect machines, see OPC-UA vs MQTT vs Modbus.
  • Trend key readings on critical assets and move them to condition-based checks.
  • Track MTBF and MTTR per critical asset and set targets.
  • Build a history of warning outcomes to decide where predictive methods are worth it.
  • Link downtime to production goals, so the plant knows early when a month is slipping.

Where software helps

None of these steps need special software to start. Software becomes useful when the number of machines, alarms and people grows beyond what a sheet and a phone tree can handle.

MIE, the Manufacturing Intelligence Engine, is built for this. Live today, it takes the readings your machines already send over HTTPS, MQTT, OPC-UA or Modbus, and sends alerts on WhatsApp, Slack or webhook through timed escalation tiers that climb until someone acknowledges, then records who did. Plant managers can also ask plain-English questions of their plant data.

Next on the MIE roadmap is predictive maintenance and production goal tracking: early warnings when a reading drifts, outcomes recorded as fixed, scheduled or nothing found, stop alerts that show the units being lost, a warning when the monthly goal is at risk, and downtime reasons tagged on each stop. These are coming soon, not live yet.

Whatever tools you use, the method stays the same: record every stop with a reason, fix the biggest reason first, make sure warnings reach a person, and review every week.

Questions

Common questions.

What is the difference between planned and unplanned downtime?
Planned downtime is scheduled in advance, such as preventive maintenance, cleaning or a planned changeover. Unplanned downtime is any stop you did not schedule, such as a breakdown, a trip, a material shortage or a utility failure.
What are the most common causes of unplanned downtime?
Equipment failure, early warning signs that nobody noticed, alarms that reached no one or were ignored, changeovers that overran, shortages of material or utilities such as steam, air and power, and operator or process errors. The mix is different in every plant, which is why you need your own downtime log.
How do you calculate MTBF and MTTR?
MTBF is total operating time divided by the number of failures. MTTR is total repair time divided by the number of repairs. Availability can then be estimated as MTBF divided by the sum of MTBF and MTTR.
How do I start downtime tracking without new software?
Start with a shared sheet or paper log at each line with the machine, start time, end time, a reason picked from a fixed list, and what was done. Review it weekly and rank the reasons by lost output.
Can predictive maintenance eliminate unplanned downtime?
No approach removes it completely, but condition-based and predictive checks on critical assets let you catch many failures early and plan the repair. It works best alongside good alerting, standard changeovers and a steady supply of material and utilities.
How quickly can a plant reduce unplanned downtime?
Some quick wins, such as a reason list, alarm owners and escalation, can be in place within a month. Larger gains from condition monitoring and changeover work build over several months as you collect data and fix the top causes one by one.
See it on a real plant

Ask the demo plant your own question.

We'll walk you through MIE on a live plant in 30 minutes.

Or call +91 98288 93692 · +91 88519 85656 · info@thinklytixai.com

Talk to us