Solutions Series·DSIB

Operational Excellence: The Design Safety Chapter

Bringing Design Safety, Quality, Productivity and Enterprise Risk into One View

Built for engineering firms, this independent assessment brings design-safety and technical-risk evidence together with quality, productivity and management performance. It helps leaders spot recurring issues, understand their wider impact, and decide what needs attention at enterprise level. As handover approaches, it also gives asset owners a clearer view of how design-safety risks have been managed.

Independent assessment KPI & KRI integration Holistic view
01 The solution

What we integrate

In an engineering company running many projects at once, design-safety and technical-risk information is often distributed across several management channels. Project controls report hours, progress, schedule and cost; the quality management system records reviews, rework, NCRs and change; design-safety and assurance processes hold assumptions, hazard findings, deviations, assurance actions and readiness evidence. We bring these evidence streams together so that design-safety issues can be read alongside the functions they affect, compared across the portfolio and escalated into enterprise-risk governance when their scale or consequence requires company-level attention.

The work does not replace project risk management, design safety management or the company's quality management system. It provides the clarity needed for company-wide decisions through a holistic KPI and KRI monitoring framework. Project and engineering evidence is interpreted through a technical-risk and design-safety lens; recurrence is tested across projects, offices and business lines; management-system causes are translated into operational-excellence action; and only exposures with sufficient reach or consequence move to enterprise level. Where one capital project needs intervention, the issue remains project-specific. This framework addresses what the combined project portfolio is telling the engineering company about its own operating model.

Enterprise risk
Not every open design-safety action, NCR or late technical decision belongs in the enterprise risk register. An issue becomes an enterprise candidate when its recurrence, consequence or cross-project reach can affect delivery capability, legal or contractual compliance, readiness for safe operation, reputation, strategic capacity or the reliability of the company's management system. The framework applies this escalation test while preserving a traceable line back to the originating projects and evidence.
Design Safety Management
Design safety is assessed through the evidence produced by SCE lifecycle steps, specific engineering activities, safety reviews and audits, together with the company objectives, standards, accountabilities and management controls that govern them. The KPI and KRI set covers technical execution, action quality, gate-pass decisions, management issues and escalation. It shows whether a weakness is limited to one project, recurring across several projects, concentrated in an office or business line, or significant enough for company-level attention.
Quality
How the mechanisms that govern engineering quality perform in practice: review and approval discipline, design changes, holds, NCRs, technical queries, rework and the action-management arrangements behind them. We add early-warning indicators to the existing KPI set and examine whether the quality-management mechanisms required for design-safety management are defined, used and evidenced consistently. The reading connects these controls to correct engineering execution and to the financial and strategic outcomes the company expects from it.
Productivity
How man-hours, design-safety staff and software are converted into usable design-safety outcomes. We compare budgets and forecasts with actual MxH, utilisation and workload; software licence and usage data; action-closeout effort; interface burden; rework; and the cost and schedule effects of safety-related change. The reading shows which project types or activities repeatedly overrun, where design-safety capacity is overloaded or underutilized, whether tools are used effectively, which interfaces absorb disproportionate effort, and whether investment in the design-safety management system produces measurable gains in gate-pass success, SCE performance criteria, lower rework and stronger readiness.
Figure 1
Closure rate is not the same as risk reduction
On-time action closure is a useful KPI. It becomes meaningful only when it is read with evidence that the closure worked and with the rate at which the same issue returns in later reviews or projects.
Office A
94%
65%
22%
Office B
87%
82%
6%
Office C
96%
58%
28%
Office D
81%
78%
7%
Actions closed by due date Closure confirmed effective Same issue repeated within 12 months
Illustrative output, anonymised. Office C appears strongest when only the closure KPI is read, yet it has the lowest rate of closures confirmed as effective and the highest recurrence. Office B reports a lower headline closure rate but achieves far more durable results. The combined reading shifts the management question from “How many actions were closed?” to “Which risks were actually reduced?”
Figure 2
The root cause of post-IFC safety-related change determines the management response
A single change total does not show which cause produced it. Connecting root cause, rework and schedule effect shows whether the response belongs in project control, quality management, interface governance or commercial change.
Root cause Projects affected Average rework MxH per project Average schedule effect after IFC Management read
Unstable design basis or late process decisions1446041 daysA major rework source recurring across many projects, with significant schedule deviations. Review philosophy preparation schedules, control points and the related procedures; many philosophy and design-basis decisions were found to be deferred to safety gates.
SRS and SIS / FGS cause-and-effect interface misalignment932024 daysThe issue sits across process, safety and control disciplines. Clarify input maturity, interface ownership, cause-and-effect review points and final approval authority.
Late vendor or package information827536 daysThe delays are major and typically concentrate in a few equipment and technology types. Handover of SCE-relevant data remains under vendor dominance, and acceptance processes were found to provide weak control in this respect.
Client scope change after IFC654068 daysManage legitimate client change through project and commercial controls, but analyze whether internal weaknesses — late clarification, weak scope boundaries or incomplete design assumptions — are being recorded as client change.
Illustrative output, anonymised. A single total for post-IFC safety-related change cannot distinguish a recurring company-control weakness from a client-driven event. Linking root cause to rework, interface impact and schedule delay shows which actions belong at project level and which require company-wide attention.

The distinction is the integrated reading: design-safety evidence remains technical, while quality, productivity, schedule and cost data show how its effects move through the company.

02 Method

How the work typically runs

Engineering companies often use similar management systems, but the way project controls, quality, risk, design assurance and enterprise reporting work in practice can vary. We therefore tailor the assessment to the decisions the company needs to make rather than applying a fixed dashboard. The sequence below is the typical path.

Scope, period and objectives are agreed
Which projects, entities, offices, business lines and years are included, and what the assessment has to serve — a board or risk-committee decision, an engineering recovery programme, a QMS review, operational-excellence planning, readiness improvement or due diligence. The boundary between project-specific intervention and company-level analysis is made explicit.
Indicator definitions are fixed
KPIs and KRIs are developed around three evidence groups: design safety, quality management and productivity. They are then examined for their operational-excellence, strategic and enterprise implications. Each indicator definition, data source, owner, frequency, threshold, evidence requirement and escalation rule is agreed before comparison, so similar-looking numbers do not carry different meanings across the company.
Data is collected and checked
Project controls, timesheets, schedules, budgets, design-deliverable and review records, technical queries, assumptions, deviations, change registers, NCR and rework records, all risk registers, action registers, assurance studies and commissioning or handover evidence are brought together and validated for accuracy, consistency and whether the same measures can be maintained across reporting cycles.
Report output design
We agree how the portfolio will be broken down — completed against ongoing, by office, department, discipline, business line, contract model and project type — and which KPI–KRI pairings will carry each finding. Every executive-level signal remains traceable to the projects and evidence beneath it, and the report distinguishes isolated events, recurring patterns and systemic exposure.
Strategy and company-wide targets
Company strategy and company-wide expectations are translated into clear design-safety, quality and productivity targets. We agree what good performance should look like, how it will be measured and which results require management action. The report then makes clear how these targets support the company's broader objectives for delivery predictability, engineering quality, compliance, client confidence, financial resilience and strategic growth.
Interim reviews and working sessions
Findings are linked to the risk factors behind them and set out in clear tables and charts. Progress is reviewed with project, engineering, quality and management representatives in interim meetings; focused workshop sessions are held where a finding requires root-cause analysis, additional evidence or management judgement.
Building our recommendations
Root causes are worked through in the workshop sessions and turned into actions with an owner, timing and target KPI or KRI. Recommendations are separated into immediate corrections, management-system improvements and company-wide decisions, then reported once for the teams implementing them and once for senior management.
03 Deliverables

Two documents, two audiences

The findings are written twice, for two different readers: in full for the teams who will implement them, and in short for the executives who will decide.

Deliverable 1
Operational Excellence: The Design Safety Chapter

The complete analysis, structured for office, department, business-line and project owners.

See contents
  • Integrated KPI and KRI architecture, definitions, thresholds and data-confidence statement
  • KPI and KRI readings separately for each office, department, business line and strategic pillar
  • Multi-year and cross-project trend analysis, with completed and ongoing work kept apart
  • Project- and department-level health check on a single scale, with ongoing and completed work compared separately
  • Root-cause findings from the workshop sessions, with the supporting evidence attached
  • Traceability from operational outcomes to quality, productivity, technical-risk and design-safety evidence
  • A full action register, each item with an owner, duration and target indicator
Deliverable 2
Executive brief — written for the CEO and executive leadership team

A short, standalone document an executive can read in one sitting and act on immediately. It follows a fixed structure.

See contents
  • Current situation — company objectives and targets read against the existing KPI and KRI baseline
  • Action plan and classification — actions grouped by control mechanism, urgency, owner and expected result
  • Transformation needs by department, business line and office — where the operating model must change, evidenced from the relevant data
  • Alarming points — the company-wide signals requiring immediate senior-management attention
  • Company-wide performance comparisons — external references where available, multi-year trend analysis and office-to-office comparison
  • Corporate improvement roadmap — expected gains, milestones and KPI or KRI targets
  • Senior-management support and expectations — the decisions, sponsorship, resources and escalations required from leadership
  • Main programme risks and opportunities — the factors that may obstruct, accelerate or broaden the improvement programme
Figure 3
Indicator board — extract
Performance indicators show how the engineering system is working; risk indicators show where technical exposure is likely to emerge next.
Indicator Type Movement Status Reading
Safety gate-pass milestone adherenceKPI▼ to 76%At riskSafety reviews and required evidence are not consistently aligned with project milestones
Safety-critical design actions closed by the agreed scheduleKPI▲ to 81%WatchOverall closeout is improving, but the high-significance subset remains late
Safety-engineering staff billabilityKPI▼ −9 ptsWatchWorkload and design-safety capacity imbalance is becoming visible before revenue is affected
Design-safety budget adherence in ongoing projectsKRI▼ to 62%WatchLate change and rework are eroding the design-safety MxH forecast
Design-safety schedule adherence in ongoing projectsKRI▼ to 68%At riskCross-discipline inputs are arriving after committed safety-engineering milestones
Hours spent reworking safety philosophies and related designKRI▲ +37%At riskRepeated correction is concentrated in design criteria, safeguarding and downstream interfaces
Projects completing alarm rationalisation before pre-commissioningKPI▲ to 71%WatchSeveral projects still enter configuration and testing before alarm decisions are stable
RBPSMS audit-score improvementKPI▲ +6 ptsStableOverall progress is visible, with two management-system elements still below target
Projects meeting functional-safety assurance gatesKPI▼ to 64%At riskGate exceptions are increasing and compressing verification and assessment effort
Average engineering rework hours triggered by FSA findingsKRI▲ to 146 hAt riskAssessment findings are requiring redesign rather than confirming mature evidence
Average schedule slippage from post-IFC safety-critical changesKRI▲ to 1 monthAt riskLate safety change is disrupting package release, procurement support and integrated planning
Average rework duration triggered by PSSR recommendationsKRI▲ to 40 daysWatchPSSR findings are triggering late design and handover work that should have been resolved earlier
Recurring NCR root causes across projectsKRI▲ 3 consecutive quartersAt riskThe same root causes have increased for three consecutive quarters; isolated closure is not removing the shared management-system cause
Illustrative output. Each indicator is defined with an owner, definition, data source, frequency, threshold and escalation route.
04 The action register

From findings to actions

Actions are grouped according to the control mechanism they are intended to strengthen: design assurance and gate-pass success; technical-risk governance and reporting; quality, change and engineering-development control; and commissioning and operational readiness. Each action is sequenced against a defined implementation horizon — 0–3, 3–6 or 6–9 months — according to current exposure, dependencies and the time required for the control to become effective. The register names the owner and the KPI or KRI that will demonstrate the result.

0–3 months
Stabilise

Contain current exposure and restore the reliability of engineering-risk information, so gate and executive decisions rest on evidence.

Read moreClose
  • Validate high-significance open actions, deviations and temporary assumptions
  • Reconcile change, NCR, rework, technical-query and commissioning-critical registers
  • Review projects crossing a design or approval gate with unresolved major technical risks
  • Isolate and minimise company-wide alarming points and their escalation potential
  • Re-establish the scope, authority, evidence quality and closeout discipline of Design Safety Reviews
  • Identify and control recurring safety-related changes generating more than 5,000 rework MxH across the portfolio
3–6 months
Rebuild

Repair the controls that allow technical risk to emerge late, recur across projects or disappear between functions.

Read moreClose
  • Phase-gate criteria, evidence requirements and risk-based approval routes
  • Cross-discipline interface, design-change and engineering-development control
  • Recurring NCR and rework root-cause control, with independent verification where needed
  • Risk-based assurance mechanisms and detailed MxH estimates systematically embedded in project schedules and plans
6–9 months
Position

Embed enterprise technical-risk governance into the engineering operating model and make readiness for safe operation a normal management output.

Read moreClose
  • Portfolio-wide technical-risk dashboard, thresholds and escalation rules
  • Engineering assurance capability, competency and independent-review model
  • Readiness-for-safe-operation requirements integrated into the quality management system
Figure 4
Action register at a glance
Actions grouped by category and implementation period.
Design assurance & gate-pass success
28
Technical-risk governance & reporting
21
Quality, change & engineering-development control
18
Commissioning & operational readiness
13
Months 0–3
Stabilise
Current exposure validated, critical actions controlled and unreliable reporting corrected. Roughly 60% of immediate actions closed or under control.
Months 3–6
Rebuild
Gate criteria, evidence rules, interface controls and change/NCR mechanisms in use. Cumulative 90%.
Months 6–9
Position
Enterprise technical-risk governance, assurance capacity, QMS integration and safe-operation readiness embedded. Cumulative 100%.
Illustrative output from one assessment: 80 actions, each with an owner, a duration and the KPI or KRI that will measure it. The distribution differs from company to company.
05 Where the conversation starts

Some patterns we usually begin talking about

These are the questions we ask at the beginning of an assessment because they reveal whether design-safety issues are being handled as isolated project events or understood as patterns affecting engineering quality, productivity and enterprise risk. Some companies answer them quickly; questions that require evidence from several functions are often where the most useful findings begin.

Figure 5
Completed work against work in execution
In this portfolio, completed projects met their gates more consistently. The ongoing portfolio shows higher post-freeze change, more actions crossing gates and rising rework — before the full schedule and cost effect is visible.
Post-freeze safety-critical design changes
Completed
4%
Ongoing
18%
High-significance actions carried across gate
Completed
6%
Ongoing
27%
Rework hours / engineering hours
Completed
3%
Ongoing
11%
Illustrative output, anonymised. The company's combined figures looked acceptable because completed work carried them; the ongoing portfolio was accumulating technical risk. Comparisons are stage-normalised and repeated by office, department and discipline wherever the difference is large enough to matter.
Do you see a recurring or worsening trend in projects passing a design gate with high-significance technical actions still open?
A conditional pass can be valid. It becomes a management concern when the acceptance basis, owner, due date, downstream effect or closure evidence is unclear, or when the same pattern repeats. In that case, the gate has transferred the risk rather than removed it.
Which departments carry the most safety-related rework, and can you track it consistently?
A single rework total does not show whether the cause is late process definition, an interdisciplinary interface, a safety-review finding, vendor information or an internal quality failure. Tracking root cause, department, project type and MxH trend shows where the company is paying repeatedly for the same correction.
How and when do you identify safety-related NCRs, and what counts as an NCR in your company?
If the definition varies, the same condition may be recorded as an NCR in one office, a review comment in another and an informal action elsewhere. A common definition and detection route are needed before closure rates or recurrence trends can be compared.
Does functional-safety evidence mature with the design, or is it assembled near commissioning?
When lifecycle planning, safety requirements, verification records and assessment inputs arrive late, validation is compressed and design assumptions are harder to correct. The issue is not only missing documents; the design may have advanced without timely challenge.
When does a repeated project-level technical issue become an enterprise risk?
The threshold is not a fixed number of occurrences. It depends on consequence, recurrence, cross-project reach, controllability, compliance, impact on safe commissioning and operation, and whether the issue reveals a weakness in a company-wide standard, capability or governance process.
Do the projects under the greatest schedule pressure also carry the most post-IFC safety-related change and rework?
If the patterns align, the cause may lie in design maturity, review timing or interface control rather than planning alone. It is also worth testing whether safety and quality controls are the first things relaxed under schedule pressure. If safety-related rework is concentrated in the same projects, that case becomes stronger.
Do SRS, SIS and FGS cause-and-effect, 3D safety review and alarm-database activities repeatedly create interface problems — and who owns the solution?
These activities cross process, safety, instrumentation, control, layout, operations and sometimes vendor boundaries. Repeated problems usually point to unclear input maturity, interface ownership or approval authority. The key question is not only who closes each issue, but who prevents it from recurring.
How does your company define readiness for safe operation, and which procedures, safety reviews and design-assurance mechanisms support it?
The contract should make the required scope and evidence clear, and the company should apply that basis consistently across projects. During design, SCE lifecycle integrity requirements must be understood, together with commissioning, operation and maintenance needs that can still be addressed through design. The framework should also show which reviews create the evidence, who makes the readiness decision, how exceptions are approved and how minimum requirements are applied across offices and business lines.
At start-up, do you know how many alarms your design will create in the control room?
An alarm flood at start-up is not only an operations problem. It may show that rationalisation was completed too late, or that cause-and-effect logic, set points, shelving, routing, HMI design or testing were not mature before configuration and commissioning. Reviewing alarm rationalisation before those decisions become expensive to reverse — and feeding start-up experience back into design — is what turns the event into organisational learning.
In short

We bring the engineering KPIs and technical-risk KRIs that sit in different projects, disciplines and assurance processes into one framework, read them across several years and lifecycle stages, and turn what they show into a management roadmap.

Senior management sees how recurring project-level issues accumulate into enterprise exposure — where they originate, how they move between functions, which lifecycle gates they cross and whether they are weakening readiness for safe operation. Each office and department receives the indicators that show whether its own way of working is producing the intended result, while executive leadership receives the thresholds, trends and escalation decisions that require company-level action.

The actions are specific, not generic, and written to sit inside the company's own quality management system with an owner, a date and a target KPI or KRI. The assessment can be repeated quarterly or half-yearly to keep the picture current. DSIB principals lead the work directly, using the organisation's own data.

Design Safety Intelligence Bureau
Solutions Series
© Design Safety Intelligence Bureau. All rights reserved. This document and its contents may not be copied, reproduced, distributed, downloaded or printed, in whole or in part, without prior written permission.
This document is not available for printing. © Design Safety Intelligence Bureau. All rights reserved.