AI projects usually expose data problems that were already there
When a manufacturing organisation starts an AI project, the first request is often for a model. After a few weeks, the project team discovers that the temperature tag has changed name three times, timestamps come from different clocks, machine states are encoded differently across PLCs and production orders cannot be reliably correlated with process data.
None of these are AI problems. They are manufacturing data problems. Artificial intelligence simply makes them impossible to ignore.
Start with physical meaning
A useful data pipeline begins with a clear description of the process. For every signal, know what it represents, its unit, source system, sampling characteristics and operational meaning.
Do not treat a tag name as sufficient documentation. A value called Temp_01 is nearly meaningless without knowing where the sensor is located, whether it is raw or filtered, and which process state makes the value relevant.
Time is the backbone of industrial data
Industrial AI is often fundamentally about sequences. A vibration event may happen before a motor fault. A valve response may occur seconds after a command. A batch record may describe a process phase that spans several hours.
Use consistent timestamps and record enough information to reconstruct event order. Where possible, preserve source timestamps rather than applying a generic ingestion timestamp to everything. Clock synchronisation across systems is therefore not glamorous, but it is important.
Separate raw data from interpreted data
A good architecture keeps the original measurement available while also producing derived features and contextual tables. Raw data provides forensic value. Derived data provides usability.
For example, keep the original motor current and also calculate rolling averages, load states and start-stop events. Keep the PLC state code and also map it into human-readable operating modes. This layered approach allows models to be rebuilt when assumptions change.
Context is what turns signals into information
Sensor data alone rarely explains why the machine behaved as it did. Product type, recipe, shift, operator intervention, maintenance status and ambient conditions may all matter.
A useful manufacturing dataset might join PLC data with MES order information, maintenance events and quality results. This is where OT/IT integration becomes directly relevant to AI. The better the context, the more meaningful the patterns.
Data quality controls should be explicit
Detect missing values, constant signals, impossible ranges, timestamp gaps and sudden changes in sampling frequency. Treat these conditions as data-quality events rather than silently cleaning everything and forgetting that the issue happened.
Models trained on unexamined data can learn the behaviour of the data-collection system rather than the machine itself.
Streaming is not mandatory
Not every AI problem requires a real-time pipeline. Predictive maintenance may use hourly or daily feature batches. Quality analysis may run after a batch is complete. Energy optimisation may need minute-level data.
Choose the architecture based on the decision latency the use case actually requires. Real-time infrastructure is expensive and operationally demanding. Build it where the business decision genuinely depends on it.
The practical architecture
A common pattern is PLC and SCADA sources feeding a historian or operational data layer, with MES and contextual data joined upstream of the modelling environment. An event or message bus can support decoupling, while a governed analytical store provides historical access.
The important architectural principle is that AI should not become a hidden alternate source of truth. The source systems remain authoritative for control and production records. The AI layer consumes curated data and produces derived intelligence.
When the pipeline is ready for AI
You do not need perfect data. You need data whose limitations are understood. Engineers should be able to explain where a value came from, what it means and which periods are missing or unreliable.
Once those questions have good answers, model development becomes considerably more straightforward. The difficult work has not disappeared; it has simply been done where it creates the most long-term value.
Tag semantics are an AI asset
One of the least visible assets in a manufacturing organisation is the meaning encoded in engineering naming conventions. A well-structured tag model can dramatically reduce the work required to prepare analytics. A poorly structured one forces every project to rediscover the same relationships.
Define equipment identifiers, units, engineering limits, state definitions and relationships consistently. Where possible, make metadata available alongside the values. This is especially useful when multiple sites need to feed a common analytics platform.
Event data deserves first-class treatment
Many industrial pipelines focus on periodic samples and treat events as another tag. That can lose important context. Start, stop, mode changes, recipe transitions, alarms, operator actions and maintenance events are naturally event data.
For some AI problems, the event stream is the key to segmenting time-series data correctly. A model trained without the operating-state boundaries may learn the transition between modes as an “anomaly.”
Design for reprocessing
AI workflows evolve. A feature that seemed useful during the first model may be replaced later. Keeping raw data available and preserving the transformation logic allows the organisation to rebuild derived datasets without collecting everything again.
This is one reason to separate ingestion, storage, feature engineering and modelling instead of putting the entire pipeline into one script.
Data contracts reduce long-term friction
A data contract states what a producer promises: names, units, types, frequency, timing expectations and quality behaviour. In manufacturing this can be applied to the interfaces between PLC/SCADA, MES, historians and analytics services.
Once these contracts are explicit, changes become easier to govern. An AI model should not silently break because a PLC programmer renamed a status bit.
Where the pipeline should be intentionally boring
Ingestion and storage are not the place for cleverness. They should be predictable, observable and easy to troubleshoot. If a sensor stops reporting, the data platform should make that fact visible rather than filling the gap with a plausible value without recording what happened.
Build a minimum viable data product
Pick one asset family and one decision. Define the signal set, context fields, data-quality checks and update frequency. Deliver a small, well-understood dataset that engineers can inspect themselves. That dataset becomes the foundation for later models rather than another black box owned by a single project team.
Make lineage inspectable
When an AI result is challenged, the team should be able to trace it back from model output to feature, raw signal and source system. Lineage does not need to be complicated to be useful; even a clear table of transformations and ownership prevents many debugging sessions from becoming archaeology.