The Reliability Specialist
The scarcest resource in industrial maintenance isn't software — it's reliability engineering skill. Every IREAMS deployment comes with it already on staff: thirteen specialist agents that read the maintenance data you already have, run the same deterministic engines a reliability engineer would run, and turn what they find into decisions your engineers, your crews and your leadership can act on this week. They draft the work. Your people approve it.
Thirteen agents.
One Specialist.
A reliability department is not one job — it is a dozen. Each agent has a defined brief, a fixed set of the platform's sixteen deterministic tools, and an autonomy ceiling written into it. None of them can exceed that ceiling, and none of them invent a number: if the tool returns nothing, the agent says so.
Find what failure is costing you
Bad Actor Hunter
Tier 2 · DraftsRanks your worst-performing assets by real maintenance spend and reads the Pareto curve out loud — how few assets are driving most of the cost. For the top offenders it drafts a defect-elimination task with a concrete root cause and proposed fix, grounded in that asset's own record.
PM Optimizer
Tier 1 · AdvisesCuts maintenance cost by finding preventive programmes that are too frequent, ineffective, or redundant — measured against what the failure history actually shows, not against the interval someone typed in years ago. Reports the PM events per year you get back.
Warranty Recovery
Tier 1 · AdvisesScans completed work against active warranty windows and surfaces maintenance money the business can claim back from OEMs and vendors. Usually the fastest cash return in the whole engagement — and the one that pays for the rest of it.
Corrosion Sentinel
Tier 1 · AdvisesMechanical integrity to API 510 / 570 / 653. Reads your thickness history, computes short- and long-term corrosion rates and the governing rate, and names the pressure equipment and piping approaching t-min — with code-capped next inspection dates.
Do the engineering
Weibull Analyst
Tier 2 · DraftsCharacterises how an asset actually fails — censored Weibull fit with β, η and B10 life, assets still running treated as suspensions — then converts the statistics into a defensible PM interval and drafts the interval change for approval. The reasoning is printable and reviewable.
RCA Copilot
Tier 1 · AdvisesSits inside a live investigation and facilitates it with your team — proposing the next question, structuring the Physical → Human → Latent causal ladder, and tagging each cause node evidenced or assumed. It co-drives; the humans decide, and every change takes a human click.
RCA Challenger
Tier 1 · AdvisesThe sceptic in the room. Give it a proposed root cause and it stress-tests it against the actual failure record: evidence gaps, logical leaps, alternative hypotheses, and stopping too early at a convenient symptom. Evidence is graded fact > inference > opinion > hearsay, and it ends with a verdict.
Root Success Analyst
Tier 1 · AdvisesThe mirror of RCA, from the PSC framework. Failure analysis asks why the worst asset fails; this asks why the best one succeeds — positive deviance within an equipment class — and proposes propagating the practice. The cheapest improvement mechanism a plant already owns.
Keep the operation moving
Reliability & Integrity Digest
Tier 1 · AdvisesThe Monday-morning briefing, written for a reliability or maintenance manager and sent by email on a schedule. Emerging bad actors, overdue inspections, backlog load, and the trend since baseline — cited, and leading with what to act on this week.
Manual Reader
Tier 1 · AdvisesThe part of the Specialist that has actually read your documentation. Torque specs, clearances, lubrication intervals, commissioning steps, alarm meanings — answered from your own indexed OEM manuals and SOPs, with the document name and page number attached.
CMMS Analyst
Tier 1 · AdvisesThe one that gets you started. Hand it an export from SAP PM, Maximo, MaintainX, eMaint, Limble, Fiix, UpKeep or a spreadsheet and it proposes how those columns map onto the IREAMS schema — flagging what is ambiguous. The mapping only becomes real when a human confirms it.
Assessment Narrator
Tier 1 · AdvisesWrites the executive summary of your assessment report — over findings the engines have already computed deterministically. It narrates numbers; it never produces them. That ordering is the whole reason the report survives scrutiny.
Specialist Supervisor
Tier 1 · AdvisesThe colleague you actually talk to. Ask a fleet or asset question in plain language and it reaches for the right tools rather than answering from memory, chaining the other agents when a question needs more than one of them. This is where "work up asset P-101 end to end" starts.
Working in a day — before you migrate anything
You don't have to move onto IREAMS to see what the Specialist finds. Point it at an export from your current system and it starts there.
Send a maintenance-history export
From SAP PM (IW38 / IH06), IBM Maximo, MaintainX, Limble, eMaint, Fiix, UpKeep — or a plain spreadsheet. Each source ships with its own export instructions, so your planner can produce the file without involving IT.
Your Specialist reports back
A full reliability assessment: what failure is costing you, which assets are driving it, what the life data says about PM intervals, and where money is recoverable. Every figure traced to a work order in your own file.
Then it runs continuously
A Monday briefing, nightly watchdogs on emerging bad actors and PM drift, drafted work waiting for approval, and a value ledger measuring what the decisions actually saved. Move onto IREAMS fully and it gets sharper — because it's reading a live, governed record instead of a periodic export.
What your Specialist finds in week one
This is the actual contents of the report — not a maturity questionnaire, not a scorecard. Findings in dollars, computed from your records.
Where the money is going
Pareto bad-actor ranking by real maintenance spend, with the cumulative concentration curve — usually a handful of assets driving the majority of cost.
Life data on your worst offenders
Censored Weibull fits with β, η and B10 life — and the PM interval the data actually supports, with confidence bounds. Assets still running are treated as suspensions, so risk is never overstated.
PM waste
Over-maintenance, ineffective tasks, and redundant routines identified across the fleet — where you are spending labour on work the failure data does not justify.
Recoverable warranty money
Work performed on assets that were still under warranty, surfaced in dollars — often the fastest cash return in the whole report.
Integrity red flags
Where thickness data exists: API 570 short- and long-term corrosion rates, the governing rate, remaining life, and code-capped next inspection dates.
A data-quality appendix
An honest account of what your record can and cannot support. Where the history is thin, your Specialist says so rather than inventing a number.
Analysis nobody acts on
is worth nothing.
Most reliability tools stop at a finding and leave the hard part — getting a decision made, by the right person, in time — to the plant. Your Specialist carries each finding all the way to the desk where it can be actioned, in the form that desk needs it.
The Engineer
A finding becomes a study
Reliability engineers get the working, not the conclusion. Every fit, ranking and critique is reproducible, cited to specific work orders, and savable as a versioned study on the asset — so the analysis becomes a record the next engineer inherits, not a screenshot in someone's inbox.
- Proposals queue — drafted work with its evidence chain attached
- Saved studies — Weibull fits versioned against the asset
- Challenger review — the sceptical second opinion before you commit
The Team
A decision becomes the week's work
Supervisors and crews see the outcome, not the analysis. Approved proposals become drafted work orders and revised routines; wins land back in the crew's own thread on the asset they fixed. Reinforcement, rather than reporting — the difference between a programme people believe in and one they tolerate.
- Weekly meeting pack — auto-drafted agenda: wins, stuck decisions, night signals
- Crew threads — approvals posted back to the people who did the work
- Offline-first mobile — the work reaches the floor, signal or not
Management
The week becomes a number you can defend
Leadership gets the one slide a reliability engineer keeps their job with: what changed, what it cost, and what it saved. Assessment snapshots are append-only, so every run is comparable to the last — the trend is measured, not asserted, and it is printable for the board pack.
- Return on Reliability — measured value, stated separately from identified
- Decision latency — how long proposals wait, tracked as a culture KPI
- Strategy coverage — % of critical assets on a deliberate strategy, trended
A reliability engineer covers one plant, quarterly.
This runs every night.
Below the agents sits a watchdog that needs no prompting and no language model at all — five deterministic checks against your whole fleet, every night, queueing proposals while nobody is looking.
- Emergent bad actor An asset whose corrective run-rate steps sharply above its own prior-year baseline — caught as it turns, not at the next quarterly review.
- PM-effectiveness drift An active programme whose asset keeps taking corrective hits anyway. The routine is not defending what it was written to defend.
- Big-failure RCA A major loss event opens a draft investigation with the context prefilled — the same draft a good engineer opens the morning after, while the evidence is fresh.
- Golden-Spot drift An asset sliding out of its optimal band queues a restore-the-optimum proposal — acting on sub-optimal drift before critical departure.
- Data-quality regression Failure-code or cost coverage falling away against its own trailing average, noted before it quietly undermines every analysis above it.
It is idempotent by design: a check stays quiet while a matching proposal is pending, and for 30 days after a human has decided. Your Specialist does not nag.
It doesn't do the math.
The engines do.
Most "AI" in this market is a language model talking about your data. Ours is a language model driving verified engines. The agent decides which analysis to run and writes the explanation; every number it quotes comes from a deterministic engine that a reliability engineer could reproduce by hand.
That is the difference between a chatbot that sounds like an engineer and an employee whose calculations would survive a reliability review. Ask your Specialist to show its math — it can, because it didn't do the math.
The deterministic engines
- ✔️ Censored Weibull — median-rank regression with Johnson-adjusted ranks; suspensions handled properly; conditional mean residual life.
- ✔️ Monte Carlo RAM — discrete-event simulation, inverse-CDF Weibull sampling, Box-Muller lognormal repair times, convergence-checked.
- ✔️ Spectral diagnosis — Hann-windowed FFT and Hilbert envelope analysis, matched against BPFO / BPFI / BSF / FTF bearing frequencies.
- ✔️ Integrity — API 570 short/long-term corrosion rates, governing rate, remaining life; API 510/653 interval caps.
- ✔️ RCM & FMEA — SAE JA1011-aligned decision logic that generates PM tasks from failure modes.
- ✔️ Standards-first limits — ISO 20816-3 vibration zones and learned baselines, every alarm band citing its source.
It reasons like an engineer who knows this asset
Detection tells you something changed. Diagnosis tells you what is failing and why — deterministically, before any language model is involved.
Named failure modes, not scores
Hypotheses are ranked against the ISO 14224 failure taxonomy with evidence citations attached — for example: "envelope tone 87.2 Hz matches BPFO 87.4 Hz (6205, drive end) — outer-race defect signature."
Rotating ≠ static
A heat exchanger is not judged like a pump. Health assessment branches on equipment class — rotating is vibration-led, static is integrity-led, electrical is thermal, instruments are calibration-led.
It knows this asset's history
Hypotheses matching a documented FMEA mode, or a mode this asset has actually failed with before, outrank generic matches. "Failed this way 3× since 2024" is evidence, and it is weighted as such.
Alarms that don't chatter
ISA-18.2-style deadband and persistence logic means a drifting value produces one actionable alert with a cause and a recommended action — not a stream your team learns to ignore.
Stop asking only what fails.
Start protecting what works.
Classical reliability is failure-centric: it maps the road to breakdown. The PSC (Percentage of Success Centred) framework mirrors it on the success side — measuring how long an asset sustains optimal performance, and defending that.
Where RCM tracks the D-I-P-F curve toward failure, PSC tracks D-I-S-G — Design, Installation, Success, Golden Spot. Your Golden Spot is the performance envelope where an asset is genuinely doing its job well, derived directly from the alarm bands your site already maintains. No new data entry, no new instrumentation.
From that, your Specialist measures MTOP (Mean Time of Optimal Performance — the success-side complement to MTBF), MTTRg (Mean Time To Restore Golden spot), and Success Rate — availability measured against optimal performance rather than mere functioning. It warns you the moment an asset drifts sub-optimal, long before it trips an alarm.
Relantern is the reference implementation of PSC — built by the person who published it.
Olorunfemi (2026), "A Success-Centric Evolution of Reliability-Centered Maintenance in Modern Asset Management," Science, Technology & Public Policy.
The success layer
- ✔️ Golden Spot residency — how much of the observed period each asset spent genuinely performing, not merely running.
- ✔️ Success Rate — SR = MTOP / (MTOP + MTTRg). Target ≥ 90%; world class ≥ 95%.
- ✔️ Sub-optimal drift watchdog — flags assets sliding out of the envelope, with the limiting parameter named.
- ✔️ SMEA — Success Mode & Effects Analysis: the value-centric complement to FMEA, capturing the conditions that sustain performance.
- ✔️ SPN — Success Priority Number = Value × Sustainability × Monitorability. Where FMEA rewards what's easy to detect, SMEA rewards what can be actively sustained.
PSC is an additive success layer. FMEA, RCA and classical RCM stay fully in place and remain authoritative for safety-critical cases.
Every vendor claims savings.
Almost none measure them.
The usual pattern is to total up the savings a recommendation could deliver and call it ROI. Your Specialist reports that figure too — labelled identified, and never mixed with the other one.
Measured value is different. For every asset where a proposal was approved, it compares the corrective-cost run rate in the year before against the rate since, and reports the difference as it accrues. Negative results are shown as "no measurable change yet", never quietly dropped. An asset counts once, no matter how many actions touched it, and only after a 30-day maturity window.
No figure on that statement is estimated by a language model. It is arithmetic over your own cost records, printable, with the method stated on the page — because the plant manager who presents it will be asked how it was calculated.
The Return on Reliability statement
- ✔️ Measured value to date — before/after corrective run rate on the assets your Specialist actually touched.
- ✔️ Identified value — the estimate carried on drafted work, kept visibly separate.
- ✔️ Snapshot corroboration — total plant spend trend across append-only assessment runs.
- ✔️ Proposals accepted — and how long each waited for a decision.
- ✔️ Stated caveats — run-rate deltas are attribution-adjacent, not causal proof, and the page says so.
It drafts. You decide.
Your Specialist may analyse freely and draft work for your approval. It never changes your plant data by itself. Every proposal lands in a review queue with its reasoning and evidence attached, and a human clicks Apply.
An alert nobody actions is worth nothing — which is why the loop closes into drafted work rather than a dashboard. But the authority to commit that work stays with your engineers, permanently.
Of the thirteen agents, eleven are read-only. Two may draft. None can write. The ceiling is enforced in the code that runs the agent, not in the prompt that instructs it — an agent asking for a tool above its tier is refused before the tool runs.
Autonomy, explicitly bounded
- ✔️ Tier 1 — Advise. Read-only analysis, ranking and answers. No data is changed.
- ✔️ Tier 2 — Draft. Produces a proposal a human approves before anything is written. This is the ceiling for every agent.
- ✔️ Tier 3 — Act. Reserved and not enabled. No agent writes to your record autonomously.
- ✔️ Citations required. Every claim references specific work orders, failure records or analyses.
- ✔️ Re-checked on the way out. Nothing is delivered to a CMMS until the server re-reads the proposal and confirms a human approved it — a browser cannot deliver work that was never signed off, whatever it claims.
- ✔️ Immutable audit. Every agent run is written to an append-only log — who, what, when, and on what evidence.
What your Specialist is not
Reliability engineers are paid to be sceptical. Here is what we will not claim.
- It is not machine learning. There is no trained black-box model anywhere in the product. Every result comes from published statistics, signal processing and rules — which is precisely why it can show its working.
- The agents do not compute. They choose which deterministic tool to run and explain what came back. No number on any screen was produced by a language model, and no agent can exceed the autonomy tier written into it.
- It does not need sensors — and it is not a sensor platform. The core analysis runs on work-order history alone. Condition and meter readings, CSV import and REST/historian polling are supported and additive; streaming OPC-UA and MQTT are not available today.
- It exchanges files, not live two-way sync. History comes in as an export. Approved work goes back out as a CMMS-shaped package — SAP PM, Maximo and MaintainX formats, as CSV/XLSX or delivered to an endpoint you configure. What we don't claim is a live bidirectional integration that keeps two systems continuously in step.
- Screening is labelled as screening. Risk-based inspection is an API 580/581-inspired prioritisation aid, not a full quantitative API 581 analysis. Remaining-useful-life estimates are conditional life-data estimates, not certified prognostics.
- Where your data is thin, it says so. A capability appears only when the data behind it exists. You will get an honest empty state before you get an invented number.
See what one day of your Specialist looks like
Send a maintenance-history export from whatever system you run today. Your Specialist reports back with what failure is costing you, which assets are driving it, and the work that would stop it — free, and yours to keep whether or not you go further.
- Any system: SAP PM, Maximo, MaintainX, Limble, eMaint, Fiix, UpKeep — or a spreadsheet.
- No migration: nothing to rip out, nothing to reconfigure.
- Findings in dollars: not a maturity score.