- OT priority order is safety, then availability, then integrity, then confidentiality, the reverse of the usual IT thinking
- Standard IT containment can sever the telemetry operators need to run a process safely, turning a contained breach into an operational incident
- Effective response uses separate, protocol-aware playbooks for IT and OT that coordinate through a shared incident register, not one converged playbook triggered by a single tool
- Forensic readiness has to exist before the incident: most PLCs and RTUs do not log in a format a SIEM can ingest, and generic memory-acquisition tools can crash a real-time operating system
Why an IT playbook does not survive contact with a live process
An IT network is a set of hosts moving data. An OT network is an industrial process that happens to run over Ethernet. That distinction sounds academic until you look at what each side optimises for. IT incident response is built around the CIA triad: confidentiality first, then integrity, then availability. OT engineering has run on the opposite ordering for decades, usually written as safety, availability, integrity, confidentiality. A control engineer will accept degraded confidentiality without blinking. They will not accept a control loop losing its feedback signal, because that is how a pressure vessel overpressurises or a pump runs dry.
This is not a philosophical difference, it changes the mechanics of response. In IT, you isolate a compromised host and worry about the fallout later. In OT, the historian, the HMI, and the PLC are often on the same broadcast domain by design, because the control system was engineered assuming a trusted, flat, low-latency network. Segmenting that network the way you would quarantine an infected workstation can cut the link between a historian and the controllers it polls, which blinds the operator watching tank levels or line pressure. The breach gets contained. The process does not.
Most Modbus, DNP3, and IEC 61850 deployments were designed with no authentication and no encryption, because the threat model at the time was a technician with a laptop, not a remote attacker. Deep packet inspection or an inline IDS that adds even a few milliseconds of latency can break time-sensitive functions like breaker synchronisation or GOOSE messaging in a substation. The security control becomes the outage. This is why a firewall rule that is safe on a corporate VLAN can be the wrong call on a control network, and why response teams need to know which protocol they are touching before they touch it.
Convergence claims deserve scrutiny, not adoption
A recurring pitch in this market is the single XDR or SOAR platform that "responds across IT and OT" with one playbook. Treat that claim carefully. A playbook written for IT assumes the asset it is acting on is disposable: reimage it, disable the account, block the IP. Applied to OT without process-aware context, the same automated action can isolate a network segment that includes an emergency shutdown system, or disable a workstation that a safety instrumented system depends on for a scheduled function test. The platform did exactly what it was told. What it was told was wrong for that environment.
That does not mean IT and OT response should run in silos with no communication. It means the coordination happens above the playbook layer, not inside a single automated action. The IT team runs its own detection and containment for phishing, credential theft, and lateral movement on the corporate side. The OT team runs its own detection and containment for protocol anomalies, unauthorised writes, and device tampering, using controls that respect process constraints: rate limiting rather than a hard segment cut, protocol whitelisting rather than a blanket firewall block, a read-only shadow HMI rather than taking the primary interface offline. Both feed a shared incident register so the organisation has one timeline and one set of facts, even though the response actions on each side look nothing alike. That shared war room needs visibility across both IT logs and OT telemetry to work, gathered passively rather than through anything that puts an agent on a PLC, without forcing a single automated playbook across both domains. Orchestration tools such as FortiSOAR earn their keep here as the case-management and coordination layer, not as the thing executing containment on the control network directly.
What a real OT incident response capability needs
An asset inventory with process context, not just IP addresses
A spreadsheet of IPs and MAC addresses tells you nothing about impact. What you need for every OT asset is its Purdue level, its role (engineering workstation, historian, HMI, PLC, safety instrumented system), its firmware version, the vendor and support contract behind it, and what physical process depends on it. Without that, an incident responder cannot answer the first question that matters: if we isolate this device, what stops working, and is that survivable for the next hour, the next shift, or the next three days.
Protocol-aware monitoring, not repurposed IT tooling
A NetFlow-based IT SIEM has no concept of a Modbus function code or a DNP3 timestamp anomaly. Detecting reconnaissance or tampering in OT means monitoring at the protocol layer: a spike in write-single-register commands, unexpected polling of safety interlocks, engineering-station traffic outside a maintenance window. This is a different skill set from IT log analysis, and it takes sustained exposure to a specific protocol and plant to build, not a week of cross-training.
Containment that respects the process
The default IT move, cut network access, is often the wrong move in OT. Rate limiting, protocol-aware filtering that blocks specific function codes rather than all traffic, and read-only shadow replicas of an HMI that stay online during an incident preserve operator situational awareness while an investigation runs. The goal is to remove the attacker's ability to act while keeping the operator's ability to see.
OEM support agreements negotiated before, not during, an incident
Forensic imaging of a Siemens, ABB, Rockwell, Honeywell, or Yokogawa controller usually needs vendor involvement, because the tooling and the risk of disrupting a live process are specific to that platform. If the incident response plan does not already name a contact, an SLA, and a secure method for sharing diagnostic data with the OEM, that negotiation happens for the first time during the incident, while legal reviews a data-sharing agreement and the clock runs.
Forensic readiness on devices that were never built to be forensicked
Most PLCs and RTUs do not support syslog, keep only a shallow local log, or lose their log entirely on a power cycle. Waiting until an incident to figure out how you will collect evidence from a controller is how evidence gets lost, or worse, how the collection itself causes an outage. Standard memory-acquisition tools built for general-purpose operating systems can crash a device running a real-time OS, because they assume timing tolerances that industrial firmware does not have.
Build this before you need it: packet capture from mirrored taps positioned to see engineering and control traffic, write-once storage for historian data so it cannot be altered after the fact, and a tested procedure for imaging removable media from RTUs and safety controllers without interrupting the process they run. The 2017 TRITON/HatMan incident against a Saudi petrochemical facility's Triconex safety instrumented system is the reference case for why this matters: the attacker's goal was not data theft, it was manipulating the safety layer itself, and the incident only became public because the SIS triggered a safe shutdown rather than failing silently. That is what forensic and detection investment in OT is actually buying: the difference between a safety system that fails safe and one that fails invisibly.
Tabletops that test the actual constraint
A tabletop exercise that lets the team assume "we can take the system offline" is not testing anything real. The constraint that matters is the one operations will actually impose: the process cannot stop for 48 hours, the compressor cannot be shut down mid-cycle, the safety function test window will not move. Design exercises around that constraint. A useful scenario: a suspicious write command appears on a control network segment that cannot be isolated without halting a live process. What does the team do in the next fifteen minutes, and who has the authority to accept the operational risk of either action or inaction? If nobody in the room can answer that, the plan has a gap that a real incident will find for you.
What GCC compliance actually expects
NESA's UAE Information Assurance Standards and Saudi Arabia's NCA Essential Cybersecurity Controls both require an incident response capability for critical infrastructure operators, and both get satisfied on paper by a document that was never built for OT. Assessors increasingly ask for evidence, not just a policy: show the OT-specific detection use case, show the last tabletop that included an operations stakeholder, show the OEM contact list with actual response-time commitments attached. Map each control to a real procedure rather than treating the framework as a checklist to close out. For further background on building the underlying document, see this incident response plan template, and for a wider view of where GCC OT programmes tend to fall short, see this assessment of OT and ICS risk in GCC power and utilities.
The decision rule
If your incident response plan uses the word "reimage" for a PLC, the same containment steps for IT and OT, or does not name a specific person with engineering knowledge who owns the OT side, it is an IT plan wearing an OT label. Fix it in this order: build the asset inventory with process context first, because nothing else works without it; stand up protocol-aware detection second; negotiate OEM support agreements third, before you need them, not during; and only then invest in orchestration to coordinate the two sides through a shared incident register. Safety first, availability second, everything else after. That ordering should show up in the plan, not just in the intro paragraph.