In plain terms — A sound strategy can still lose money when an order is wrong, data are incomplete, a system stops responding, or a provider interrupts service. “Pay more attention” is not a control framework: the process needs safeguards, tested alternatives, and recovery procedures.
The Basel Framework defines operational risk as the risk of loss resulting from inadequate or failed internal processes, people, and systems, or from external events. The definition includes legal risk and excludes strategic and reputational risk. It is a banking taxonomy, but the distinction is also useful in trading because it identifies the loss mechanism before a control is chosen.
Operational risk is not the same as systemic risk. A failure may remain local or spread through shared dependencies and infrastructures; potential propagation does not change the category of the originating risk.
Where it can originate
| Source | Examples | Evidence to monitor |
|---|---|---|
| Processes | duplicate instruction, missing reconciliation, incorrect approval | exceptions, bypassed controls, unreconciled positions |
| People | quantity error, inadequate skills, improper access | approval logs, segregation of duties, training |
| Systems and data | unavailable API, stale feed, incorrect clock or mapping | data integrity and freshness, alerts, capacity and failover |
| External events | cyberattack, provider outage, inaccessible site | third-party dependencies, scenarios, actual recovery times |
| Legal matters | unenforceable contract, deficient mandate or authorisation | terms, responsibilities, documentation and jurisdiction |
A loss may have several causes. For example, faulty external data becomes an internal incident when validation, blocking, and reconciliation are missing. An incident record should therefore separate cause, impact, failed control, and corrective action instead of reducing every event to “human error”.
Controls before, during, and after an event
- Prevent: least-privilege access, segregation of duties, size and notional limits, instrument and price validation, change management.
- Detect: data and order monitoring, reject and latency alerts, and reconciliation across portfolio, broker, bank, and custodian.
- Respond: escalation roles, authority to stop automation, alternative channels, and prepared communications.
- Recover: priorities for critical operations, recoverable data, controlled manual procedures, and state validation before restart.
- Learn: root-cause analysis, control remediation, and a test proving that the corrective action works.
A checklist reduces variability but does not replace automated safeguards or supervision. A limit that can be silently ignored is not equivalent to a technical block with approval and an audit trail.
Operational resilience and continuity
Operational resilience is the ability to deliver critical operations through disruption. Business continuity prepares the response and recovery. It does not promise that incidents will never happen; it establishes which services must continue, what disruption can be tolerated, and how to return to a controlled state.
A verifiable plan identifies at least:
- critical operations, their owners, and the people, data, systems, and third parties on which they depend;
- escalation triggers and impact tolerances appropriate to the context;
- data copies and procedures for recovering systems and alternative sites;
- alternative channels to staff, clients, intermediaries, counterparties, and relevant authorities;
- controlled ways to reduce risk, suspend orders, or obtain access to funds and assets;
- severe but plausible scenarios, exercises, findings, remediation, and retesting.
A second account is not a continuity plan by itself. It must be accessible, support the required instruments and permissions, be funded as planned, and be tested without creating unintended exposures. Likewise, a backup is useful only when restoration and data integrity have been tested.
The BCBS, FINRA, and IOSCO materials below address specific regulated firms. They provide an authoritative control framework, while applicable duties and review frequencies depend on the entity and jurisdiction.
Typical mistake — Treating an incident as isolated, restarting the system, and failing to verify dependencies, data, outstanding orders, and the cause of the failed control.
Sources
- Basel Committee on Banking Supervision — Basel Framework, OPE10: Definitions and application
- Basel Committee on Banking Supervision — Revisions to the Principles for the Sound Management of Operational Risk
- Basel Committee on Banking Supervision — Principles for operational resilience
- FINRA — Rule 4370: Business Continuity Plans and Emergency Contact Information
- IOSCO — Mechanisms for Trading Venues to Effectively Manage Electronic Trading Risks and Plans for Business Continuity
- IOSCO — Market Intermediary Business Continuity and Recovery Planning