What is Industrial DataOps?
Industrial DataOps is the practice, and the software platform supporting it, by which an organisation collects data from operational technology, adds context, governs it and delivers it at scale to IT consumers such as analytics, MES and AI. The idea is borrowed from generic DataOps but adapted to the quirks of plant-floor data: thousands of loose tags, dozens of protocols and data locked away in historians and controllers. Industrial DataOps therefore forms the data layer of OT convergence.
π― Why is OT data so hard to use?
Anyone trying to analyse data from a plant quickly runs into the same problems:
-
Tags without context β a controller reports a value such as
FIC101.PV = 42.7, but says nothing about the machine, unit, product or order it belongs to - Many protocols β OPC UA, Modbus, PROFINET, EtherNet/IP and proprietary drivers coexist, often differing from line to line
- Data in silos β time series live in the historian, orders in the MES, maintenance in a CMMS and master data in the ERP, with no common key
- Inconsistent naming β every brownfield line has its own tag names; the same pump has three different names at three sites
- Point-to-point integrations β each new application gets its own connection, producing an unmanageable web of interfaces
As a result, data scientists spend a large share of their time searching for and cleaning data. Industrial DataOps tackles this structurally: context is added once, close to the source, and then reused by every consumer.
π How does it differ from DataOps and MLOps?
The term DataOps first appeared in 2014 and was developed from 2015 onwards into a discipline for delivering analytics-ready data faster and more reliably, modelled on DevOps. The industrial variant was popularised in the late 2010s by vendors such as HighByte and the analyst firm LNS Research, which use the term Industrial DataOps for OT-specific data platforms.
| Characteristic | DataOps | Industrial DataOps | MLOps |
|---|---|---|---|
| Goal | Reliable data pipelines for analytics | Connect, contextualise and deliver OT data to IT | Train, deploy and monitor machine learning models |
| Typical source | Databases, SaaS, log files | PLCs, SCADA, historians, sensors | Prepared datasets |
| Data type | Mainly records and tables | Mainly high-frequency time series | Features and models |
| Main challenge | Quality and version control | Missing context and protocol diversity | Model drift and reproducibility |
| Location | Data centre or cloud | Edge and cloud | Cloud or data centre, sometimes deployed at the edge |
Industrial DataOps is thus a prerequisite for MLOps on the plant floor: without clean, contextualised data there is little for Industrial AI to learn from.
π§ How does Industrial DataOps work?
An Industrial DataOps platform combines five core capabilities:
| Capability | What it does | Example |
|---|---|---|
| Connectivity | Reads data from controllers, historians and databases | Drivers for OPC UA, Modbus, SQL, REST |
| Contextualisation | Links raw values to asset, unit, order and product | Tag FIC101.PV becomes Line 2 / Mixer / Flow (mΒ³/h) |
| Data modelling | Defines reusable models per asset type | One Pump model with pressure, flow, status and running hours |
| Data pipelines | Transforms, filters, aggregates and routes data | One-minute averages to the cloud, events to the MES |
| Governance | Sets ownership, quality, versioning and access | Who may change which model, which consumer receives which data |
For the models, many organisations use the equipment hierarchy from ISA-95 (internationally IEC 62264): Enterprise β Site β Area β Work Centre β Work Unit. Every data stream then has a recognisable place, regardless of the controller brand underneath.
π What role do the Unified Namespace and Sparkplug B play?
The Unified Namespace (UNS) is an architectural pattern popularised by Walker Reynolds: every system
publishes its current state to a single, hierarchically organised broker, and every consumer subscribes to
it. Point-to-point integrations give way to a hub-and-spoke structure. In practice, the topic structure
often follows ISA-95, for example enterprise/site/area/line/machine.
MQTT is the usual transport layer. The protocol was created in 1999 for pipeline monitoring, became an OASIS standard in 2014 and ISO/IEC 20922 in 2016; version 5.0 followed in 2019. MQTT Sparkplug B adds conventions that plain MQTT lacks:
-
Fixed topic structure β
spBv1.0/<group>/<message type>/<edge node>/<device> - Compact payload β encoded with Google Protocol Buffers
- State management β birth and death messages (NBIRTH, DBIRTH, NDEATH, DDEATH) so consumers always know whether a source is online
- Report by exception β only changed values are sent, saving bandwidth
Sparkplug was developed by Cirrus Link in 2016, moved in 2018β2019 under the stewardship of the Eclipse Foundation and, as version 3.0, was published in 2023 as ISO/IEC 20237. A common caveat is that Sparkplug topics are technical in nature, so many architectures combine a Sparkplug layer for connectivity with a separate ISA-95-structured UNS for contextualised data.
βοΈ Where does processing happen: edge or cloud?
Industrial DataOps is almost always hybrid. At the edge, close to the machines, data is collected, normalised and given context. This keeps latency low, limits network traffic and lets local applications keep running if the connection drops. In the cloud or central data centre, data from multiple sites is combined for reporting, benchmarking and model training. A useful rule of thumb: what the process needs stays local; what the organisation needs may go central.
π What kinds of tools and vendors are there?
| Category | Role | Examples |
|---|---|---|
| Edge DataOps / data hub | Connect, model and route at or near the site | HighByte Intelligence Hub, Litmus Edge |
| Industrial data platform | Contextualise OT, IT and engineering data at enterprise level | Cognite Data Fusion |
| Historian / data infrastructure | Store and serve time series | AVEVA PI System (formerly OSIsoft, acquired in 2021) |
| MQTT broker | Core of the Unified Namespace | HiveMQ, EMQX, Mosquitto |
| General-purpose streaming | Data streams in the IT environment | Apache Kafka |
HighByte was founded in 2018 and focuses specifically on industrial DataOps; Cognite was established in Norway in 2016 out of the oil and gas industry. The right choice depends mainly on scale, the existing historian and how much you want to manage yourself.
π How do you keep the OT-to-IT data flow secure?
A DataOps platform is by definition a bridge between networks and deserves the same attention as any other conduit under IEC 62443:
- Through the IDMZ β data never travels directly from the control network to the office or cloud; a broker or replica in the IDMZ is the handover point
- Outbound connections only β the OT side initiates the connection (publishing outwards); no inbound ports are opened towards OT
- One-way where possible β for critical processes, a data diode physically enforces that data can only leave
- No writes without a reason β disable commands (such as Sparkplug NCMD/DCMD) by default or authorise them strictly
- Access control per topic β TLS, per-client certificates and authorisation per topic or data model
- Governance β record who owns each model, which consumer may see which data and how changes are approved
π§ A step-by-step approach to Industrial DataOps
- Start with a use case β for example calculating OEE automatically for one line, not βall data to the cloudβ
- Inventory your sources β which controllers, historians and databases deliver the values you need, and over which protocol?
- Agree a naming standard β an ISA-95 hierarchy and fixed models per asset type
- Build the edge layer β connectivity and contextualisation on site, with buffering in case of outages
- Design the secure handover β IDMZ, outbound connections, a data diode where justified
- Serve the consumers β MES, dashboards, the data lake and AI applications subscribe to the same models
- Scale out β reuse the models for the next line and site; that is where the real payoff lies
In the European context, the Data Act (Regulation (EU) 2023/2854) also matters. Since 12 September 2025, users of connected products have had the right to access the data those products generate, and connected products placed on the market after 12 September 2026 must be designed so that this data is accessible by default. For manufacturers and machine builders, well-modelled and documented data is therefore no longer a luxury but an obligation.
β Frequently asked questions
Is Industrial DataOps the same as a historian?
No. A historian stores time series efficiently, whereas Industrial DataOps handles connectivity, context, modelling and delivery to multiple consumers. A historian is often one of the sources or destinations within an Industrial DataOps architecture.
Do I need a Unified Namespace for Industrial DataOps?
A Unified Namespace is not mandatory, but it is a widely used pattern for implementing Industrial DataOps. Its advantage is that each source publishes once and each consumer subscribes on its own, which sharply reduces the number of integrations.
What is the difference between Industrial DataOps and an MES?
An MES directs and records production: orders, batches, quality and performance. Industrial DataOps supplies the contextualised data on which an MES, as well as dashboards and AI applications, depend. The two complement each other.
Why is contextualisation so important in Industrial DataOps?
Without context, a tag value is just a number without a unit, asset or order. Contextualisation in Industrial DataOps makes data understandable for people and machines alike, so that analyses and AI models become comparable across lines and sites.
Is Industrial DataOps a risk to OT security?
Industrial DataOps increases the number of connections between OT and IT, and therefore the attack surface if it is poorly designed. With an IDMZ, outbound-only connections, one-way traffic where possible and strict per-topic access control, that risk remains manageable.
Which standards are used in Industrial DataOps?
Common standards include ISA-95 (IEC 62264) for the asset hierarchy, OPC UA for connectivity, MQTT and Sparkplug (ISO/IEC 20237) for publish/subscribe transport, and IEC 62443 for securing the data flows.
π In summary
Industrial DataOps turns loose OT tags into reliable, contextualised information that IT applications can use at scale. If you start with a single use case, reuse ISA-95 models and secure the data flow through an IDMZ, you lay the foundation for MES, OEE analysis and industrial AI.
