Skip to content
Learning path Gold Professional operator

Operational risk

Operational risk arises from inadequate or failed processes, people, systems, or external events. Resilience and continuity help critical operations withstand or recover from disruption.

In plain terms — A sound strategy can still lose money when an order is wrong, data are incomplete, a system stops responding, or a provider interrupts service. “Pay more attention” is not a control framework: the process needs safeguards, tested alternatives, and recovery procedures.

The Basel Framework defines operational risk as the risk of loss resulting from inadequate or failed internal processes, people, and systems, or from external events. The definition includes legal risk and excludes strategic and reputational risk. It is a banking taxonomy, but the distinction is also useful in trading because it identifies the loss mechanism before a control is chosen.

Operational risk is not the same as systemic risk. A failure may remain local or spread through shared dependencies and infrastructures; potential propagation does not change the category of the originating risk.

Operational resilience: before, during and after disruption People, processes, systems and third parties can turn an error into financial loss Operational resilience: before, during and after disruption People, processes, systems and third parties can turn an error into financial loss RISK SOURCESpeople and processessystems and cyberbrokers and third parties PREVENTIVE CONTROLSaccess and segregation ofdutieslimits, reconciliation andtestingchange and continuitymanagement DISRUPTION a critical operation isunavailable orperformed incorrectly RESPONSE AND RECOVERYdetect and containcontinue throughalternativesrecover, reconcile andlearn A control reduces likelihood or impact; it does not eliminate residualrisk. Cyclepedia · source-checked visual explainer
Prevent where possible, absorb disruption, recover critical operations, and learn from the incident.

Where it can originate

Source Examples Evidence to monitor
Processes duplicate instruction, missing reconciliation, incorrect approval exceptions, bypassed controls, unreconciled positions
People quantity error, inadequate skills, improper access approval logs, segregation of duties, training
Systems and data unavailable API, stale feed, incorrect clock or mapping data integrity and freshness, alerts, capacity and failover
External events cyberattack, provider outage, inaccessible site third-party dependencies, scenarios, actual recovery times
Legal matters unenforceable contract, deficient mandate or authorisation terms, responsibilities, documentation and jurisdiction

A loss may have several causes. For example, faulty external data becomes an internal incident when validation, blocking, and reconciliation are missing. An incident record should therefore separate cause, impact, failed control, and corrective action instead of reducing every event to “human error”.


Controls before, during, and after an event

  1. Prevent: least-privilege access, segregation of duties, size and notional limits, instrument and price validation, change management.
  2. Detect: data and order monitoring, reject and latency alerts, and reconciliation across portfolio, broker, bank, and custodian.
  3. Respond: escalation roles, authority to stop automation, alternative channels, and prepared communications.
  4. Recover: priorities for critical operations, recoverable data, controlled manual procedures, and state validation before restart.
  5. Learn: root-cause analysis, control remediation, and a test proving that the corrective action works.

A checklist reduces variability but does not replace automated safeguards or supervision. A limit that can be silently ignored is not equivalent to a technical block with approval and an audit trail.


Operational resilience and continuity

Operational resilience is the ability to deliver critical operations through disruption. Business continuity prepares the response and recovery. It does not promise that incidents will never happen; it establishes which services must continue, what disruption can be tolerated, and how to return to a controlled state.

A verifiable plan identifies at least:

  • critical operations, their owners, and the people, data, systems, and third parties on which they depend;
  • escalation triggers and impact tolerances appropriate to the context;
  • data copies and procedures for recovering systems and alternative sites;
  • alternative channels to staff, clients, intermediaries, counterparties, and relevant authorities;
  • controlled ways to reduce risk, suspend orders, or obtain access to funds and assets;
  • severe but plausible scenarios, exercises, findings, remediation, and retesting.

A second account is not a continuity plan by itself. It must be accessible, support the required instruments and permissions, be funded as planned, and be tested without creating unintended exposures. Likewise, a backup is useful only when restoration and data integrity have been tested.

The BCBS, FINRA, and IOSCO materials below address specific regulated firms. They provide an authoritative control framework, while applicable duties and review frequencies depend on the entity and jurisdiction.

Typical mistake — Treating an incident as isolated, restarting the system, and failing to verify dependencies, data, outstanding orders, and the cause of the failed control.


Sources