Home / Selected work / Production monitoring case study
Case study · Enterprise AI

From firefighting to prevention

How an AI agent with event management helps a manufacturer catch a problem before it stops the line - architecture, timeline and indicative cost.

A note on transparency: This case is illustrative - it summarises a typical situation and approach I encounter repeatedly with manufacturing companies. The company name and specific figures are fictional, meant to demonstrate a realistic scenario without tying it to a specific client.

About the company

A mid-size manufacturer, roughly 300 employees, running continuous operation. One hall, three production lines, an in-house IT team of two people who handle everything from the network to the ERP.

Starting point: just firefighting

The operations team found out about a problem the moment the line had already stopped. A sensor that had been reporting increased vibration on one of the machines a week earlier was never monitored - the data was collected, but nobody reviewed it regularly. A network outage between the hall and the server room struck without warning, because monitoring existed only as a dashboard people looked at once something had already gone wrong.

A typical week looked like this: two to three unplanned interventions, with at least one usually having a real impact on production. The IT team spent most of its time reactively - diagnosing after the incident, not preventing it.

What leadership wanted above all was to stop hearing about problems from operators on the line.

What we built: a monitoring agent with event management

The goal wasn't "another dashboard", but a system that evaluates what's happening on its own and raises a flag before it becomes an incident.

01

Data collection

Connecting to existing sensors (vibration, temperature, machine load) and network telemetry. No new hardware - just bringing online what the company already had but wasn't using.

02

Classic anomaly detection

Statistical models that learned what "normal" operation looks like for each machine, and flag deviations.

03

An agent on top

Evaluates deviations in context - is this a one-off blip, or a trend over the last week? It opens an event in the internal ticketing system itself, assigns a priority and writes a clear summary, not just "value out of range".

04

Escalation

Low-priority events are just logged and tracked further. Higher priority triggers an alert to a specific person. Critical cases (risk of the line stopping within 24h) are escalated immediately with a recommended next step.

A person still decides what to do about it - the agent doesn't take over the work, it just names it earlier and more clearly.

Timeline and indicative cost

6 weekspilot deployment on one line
12 weeksfull rollout across all three lines
230,000-520,000 CZKdepending on scope - pilot, or full rollout

The price mainly depends on whether the company wants to validate the approach on one line first, or deploy across the whole operation right away. Phase breakdown at a 10,000 CZK/day rate:

  • Analysis and architecture design: 5-8 days
  • Sensor and data pipeline integration: 8-10 days
  • Anomaly-detection model and tuning: 6-8 days
  • Pilot deployment and iteration: 4-6 days
  • Agent - context evaluation, ticketing, escalation: 8-10 days
  • Rollout to remaining lines: 8-10 days
230,000-320,000 CZKbasic package - one AI agent, pilot on a single line (phases 1-4)
390,000-520,000 CZKextended package - full rollout across all three lines (phases 1-6)

The basic package corresponds to a pilot deployment on a single line, including data collection and the anomaly-detection model - without the agent layer with ticketing and escalation, which makes more sense to add once the pilot has been validated. The extended package covers the complete implementation - the agent as well as rollout to the remaining lines. Cost further depends on how much sensor infrastructure the company already has - in this case, most of the data was available but unused, so the bulk of the budget went into the agent and integration, not new equipment. The architecture design alone, without implementation, tends to be significantly cheaper and faster - similar projects typically run 80,000-120,000 CZK.

What it changed in operations

  • The IT team stopped spending time manually scanning dashboards - they now only get what genuinely needs attention.
  • An event history built up, making it possible to trace which pattern typically precedes a specific type of fault - useful for planning maintenance, not just after a breakdown.
  • Unplanned interventions affecting production dropped from 2-3 a week to roughly 1-2 a month. The exact figure varies by operation, but the direction is always the same: from reactive firefighting to planned maintenance.
  • A side effect companies often don't expect: for the first time, the IT team has data to justify investing in a specific machine or network component, instead of it staying a vague sense that something should be addressed.

Where the limits are

  • The agent doesn't replace maintenance - it just says sooner where to look. The decision and the physical intervention stay with people.
  • Without a reasonable amount of historical data, the model can't recognise what's normal - for a new machine or a new line it takes weeks to months before detection becomes reliable.
  • It doesn't make sense where a company has no sensor data available at all - there, the first step is investing in data collection itself, not an AI layer on top of it.

Why I'm writing about this

This isn't about "deploying AI" as an end in itself. It's that companies often already have data they aren't using - and the difference between firefighting and prevention tends to be a question of who (or what) watches that data consistently, more than a question of new technology.

Frequently asked questions

Is this case study a real client?

No, it's an illustrative scenario. It summarises a typical situation and approach I encounter repeatedly with manufacturing companies - the company name and specific figures are fictional, meant to demonstrate a realistic approach without tying it to a specific client.

What does a monitoring agent with event management do?

It evaluates sensor-data anomalies in context, opens an event in the ticketing system itself, assigns a priority and writes a clear summary - for critical cases it escalates immediately with a recommended next step. Decisions and intervention stay with people.

When doesn't this type of solution make sense?

Where a company has no sensor data available at all - the first step there has to be investing in data collection, not an AI layer on top of it. Without enough historical data, anomaly detection takes weeks to months to become reliable.

What's the difference between the basic and extended package?

The basic package (230,000-320,000 CZK) is a pilot deployment of one AI agent on a single line - data collection and the anomaly-detection model. The extended package (390,000-520,000 CZK) adds the agent layer with ticketing and escalation, plus rollout to all three lines.

Dealing with something similar?

I'll tell you plainly whether this would make sense for you too and what it would actually involve.

Book a free consultation →