Running a Predictive Maintenance Pilot: A 30/60/90-Day Guide
Running a Predictive Maintenance Pilot: A 30/60/90-Day Guide
In one line: A predictive-maintenance pilot succeeds or fails on three decisions you make before it starts: the right assets, success criteria defined up front, and a bounded timeline. Run it as a 30/60/90-day plan (days 1–30 to set up and baseline, 30–60 to detect and validate, 60–90 to prove value and decide) and you get a clear go/no-go answer instead of an open-ended experiment that quietly fizzles. This is how to structure one that actually converts.
Why most pilots fail before they start
The predictive-maintenance pilot is where the technology stops being a demo and has to work on your machines. It's also where a surprising number of programs die, not because the technology failed, but because the pilot had no shape.
An open-ended "let's try it and see" pilot has no finish line, so it never crosses one. Nobody defined what success looked like, so nobody can say whether it worked. Too many assets got connected at once, so attention scattered. And the one question that decides everything (what do we need to see to say yes?) never got asked until the budget conversation, when it was too late.
A pilot that converts is the opposite: a small number of the right assets, explicit success criteria written down before day one, and a bounded 30/60/90-day arc with a decision at the end. The 90 days aren't the point; the structure is.
Before day one: the three decisions that decide it
Get these right and the rest is execution. Get them wrong and no amount of ML rescues the pilot.
- Pick the right assets: a few, not many. Choose three to five machines that are (a) critical enough that catching a failure matters, (b) instrumented or easy to instrument, and (c) ideally have a known, recurring failure mode. Resist the urge to "monitor everything." A focused pilot on the assets that matter beats a diffuse one across the whole plant.
- Define success criteria up front, in writing. Decide now what a "yes" looks like: detection quality on your equipment, explainable alerts your engineer trusts, a workflow that fits, an ROI case. Write it down before the vendor connects anything, so the D90 decision is a measurement, not an argument.
- Get data access sorted early. The fastest pilot-killer is an OT-security review that stalls for a quarter. This is where a read-only architecture pays off: a monitoring tool that subscribes to your existing sensors over read-only OPC-UA, with no write path to your control system, clears review far faster than one that wants to touch the process.
Days 1–30: set up and baseline
Goal: data flowing, a healthy baseline forming, and the first signals appearing.
Connect the platform to your existing sensors and get telemetry flowing. A good platform starts cold-start detection from day one: a fast anomaly model is watching within the first day, while a deeper model learns your machine's specific normal behavior over the following couple of weeks. You're not training anything by hand; you're watching the baseline form.
What to watch in the first month:
- Data quality. This is the real first-30-days work. Loose sensor mounts, low sampling rates, and connection dropouts show up as poor data quality, and no model can rescue bad data. A good platform scores data quality continuously and tells you which sensor to check. Get this healthy before you judge the predictions.
- The first anomalies, and a warning. Early alerts are a mix of real small changes and noise, and the false-alarm rate is highest in the first weeks before the deeper model matures. This is normal. Don't judge accuracy on day 10; judge it once the baseline is established.
- An early checkpoint. Around day 30, review: is data flowing cleanly, is the baseline formed, are alerts starting to make sense? This is a health check, not the verdict.
By the end of month one you should have clean data, a formed baseline, and the beginnings of real detections: not a final answer, but a foundation.
Days 30–60: detect and validate
Goal: confirm the system catches real issues, explains them, and fits your workflow.
Now the models are mature and the pilot earns its keep. The work shifts from setup to validation:
- Validate alerts against reality. When the platform flags an asset, put your reliability engineer on it. Was it a real developing issue? Did the alert say which sensors drove it, clearly enough to act on? This is the core test: not "did it fire," but "was it right, and could we act on it?"
- Wire the workflow. A prediction only matters if it becomes action. Confirm the path works: alert → a work order in your CMMS with the asset, likely fault, and sensor evidence attached → a planned repair. A pilot that proves detection but not workflow hasn't proven the thing you're buying.
- Tune out the noise. Feed false positives back so thresholds sharpen. By the end of month two, alert quality should be materially better than in week two, and your engineer should trust it as a genuine second opinion.
Around day 60 you have the real conversation: the system works on your data. Do we scale it? That's a commercial discussion built on evidence, not a leap of faith.
Days 60–90: prove value and decide
Goal: measure against the criteria you set on day zero, and make a go/no-go call.
The final stretch is about evidence and decision:
- Measure against the pre-set criteria. Pull out the success criteria you wrote before day one and score against them honestly. Detection quality, explainability, workflow fit, ROI: did the pilot clear the bar you set, not a bar invented after the fact?
- Build the business case. Tally what the pilot caught and what it plausibly avoided, and set it against the cost. The U.S. Department of Energy's O&M work at PNNL frames the savings (8–12% over preventive, 30–40%+ over reactive maintenance) and Deloitte's analysis puts uptime gains at 10–20%; your pilot turns those industry figures into a number for your plant. Weigh it against what the platform costs.
- Decide, and scale deliberately. Convert and expand to the next tier of assets, or walk away with a clear reason. Either way you have a decision backed by evidence, which is the entire point of running a pilot instead of arguing about one.
The honest part: what a pilot can and can't prove
Here's the caveat most vendors won't volunteer, and it's the one that keeps a pilot honest: a pilot proves the system works on your data; it can't guarantee it catches a specific catastrophic failure inside the window.
If none of your piloted assets happens to degrade in 90 days, you won't get a dramatic "it saved us" moment, and that's not a failure of the tool. Judging a pilot solely on "did it predict a big failure" is a trap, because failure timing doesn't cooperate with your calendar. Measure what you can control: detection quality on known-degrading assets, whether alerts are explainable and actionable, whether false positives are manageable, and whether the workflow fits. Those tell you if the system will work when a real failure does develop.
The other honest failure modes, all avoidable:
- No success criteria up front: the pilot can't be judged, so it drifts.
- Too many assets: attention scatters and nothing gets validated properly.
- No one assigned to act: alerts pile up unactioned and the pilot proves nothing.
- Ignored false positives: alert fatigue sets in, and alert fatigue is the single most-cited reason predictive-maintenance programs fail. A pilot that floods the team with noise converts no one, no matter how clever the model.
What this looks like with Prevly
Prevly is built to make a pilot fast to start and honest to judge:
- Fast setup, so the pilot starts producing signal quickly. Read-only OPC-UA onto your existing sensors, no new hardware, and cold-start detection from day one: the setup that stalls other pilots is the part Prevly is designed to clear quickly.
- An 8-week pilot SLA. A bounded commitment on your critical assets, not an open-ended engagement: you get to a real decision on a known timeline.
- Run by your team, with explainable alerts. Every alert reports which sensors drove it, so your reliability engineer validates it with domain knowledge, no data scientist required.
- Predictions become work orders. Detection plus remaining-useful-life (reported as conformal prediction intervals) plus attribution becomes a drafted work order in your existing CMMS, so the pilot proves the whole loop, not just the model.
- Transparent by design. Published pricing and honest criteria, so the D90 decision is a measurement against what you set out to prove.
Frequently asked questions
How long should a predictive maintenance pilot be? Long enough to set up, validate on real data, and decide: typically a 30/60/90-day arc (setup, validation, decision). Prevly commits to an 8-week pilot SLA on your critical assets. The key isn't the exact length; it's a bounded timeline with a decision at the end, rather than an open-ended trial that never concludes.
What assets should you pick for a PdM pilot? Three to five machines that are critical enough that catching a failure matters, already instrumented or easy to instrument, and ideally with a known recurring failure mode. Resist monitoring everything: a focused pilot on the right assets validates faster and cleaner than a diffuse one across the plant.
How do you measure whether a predictive maintenance pilot succeeded? Against criteria you set before it started: detection quality on your equipment, whether alerts are explainable and actionable, whether the prediction-to-work-order workflow fits, and the ROI case. Don't judge it solely on catching a catastrophic failure: failure timing may not cooperate with a 90-day window.
What makes a predictive maintenance pilot fail? Usually not the technology. The common causes are no success criteria defined up front, too many assets diluting focus, no one assigned to act on alerts, and unmanaged false positives causing alert fatigue, the most-cited reason PdM programs fail. A stalled OT-security review is another; read-only architecture avoids it.
Do you need to catch a failure during the pilot for it to succeed? No. A pilot proves the system works on your data; it can't guarantee a specific failure occurs in the window. Judge it on detection quality, explainability, workflow fit, and false-positive rate on your known-degrading assets; those predict whether it'll catch a real failure when one develops.
Run a pilot with a real finish line
A predictive-maintenance pilot shouldn't be an open-ended experiment. Pick a few critical assets, set your success criteria up front, and run a bounded plan to a decision.
Request a Prevly demo and we'll scope a pilot (assets, criteria, and timeline) around a decision you can actually make.
Related reading: How to choose a PdM platform · How much does predictive maintenance cost? · Predictive maintenance without a data scientist · The ROI of predictive maintenance · Read-only OPC-UA monitoring