Skip to content

Services

Sixteen engineering methods, grouped by the decision they support

Each method below is a defined piece of engineering work with a stated purpose, a reason a data-center operator needs it, and a deliverable you can hand to your maintenance, operations or finance team. Most engagements combine three or four of them.

Sixteen engineering methods, grouped by the decision they support

01

Assess and prioritise reliability risk

Establish an independent baseline: what the facility actually consists of, how it can fail, and which assets carry the consequence.

01 — Assess and prioritise reliability risk
01Assess and prioritise reliability risk

Data Center Reliability Assessment and Improvement Roadmap

See where availability is truly at risk

  • Structured review of critical electrical, mechanical, cooling, control, fire-protection, fuel, water and supporting systems.
  • Document review, site inspection, staff interviews, maintenance-history analysis, alarm and event review, and evaluation of operating procedures.
  • Examination of equipment condition, redundancy, system dependencies, maintenance practice, staffing, vendor support, deferred work and operational readiness.
Read the full method

Technical Hierarchy and Asset Register Development

A trusted map of the assets behind every critical service

  • Define asset boundaries and parent-child relationships from portfolio or campus down to systems, equipment, subassemblies and maintainable items.
  • Reconcile single-line diagrams, P&IDs, as-builts, equipment schedules, CMMS/EAM records, DCIM, BMS and EPMS against physical site verification.
  • Set naming conventions, asset classes, mandatory attributes and ownership rules.

Asset Criticality Analysis

Focus maintenance and investment on the assets that matter most

  • Rank assets and systems by the consequence of failure: service availability, safety, regulatory compliance, environmental impact, financial loss, customer commitments, reputation, redundancy, recoverability and repair time.
  • Assess each asset's role within the complete power, cooling or control chain — not the equipment in isolation.
Read the full method

02

Engineer the right maintenance strategy

Move from equipment lists and OEM defaults to maintenance that is derived from the failures you are trying to prevent.

02 — Engineer the right maintenance strategy
02Engineer the right maintenance strategy

Failure Modes and Effects Analysis — FMEA / FMECA

Understand how critical systems fail before service is affected

  • Define required functions, then identify functional failures, failure modes and causes, local and system-level effects, existing controls and detection methods, consequences, current maintenance response and recommended risk reduction.
  • Where appropriate, add criticality or risk ranking to produce an FMECA.
Read the full method

Reliability-Centred Maintenance — RCM

Match every important failure mode with the most effective response

  • Apply structured RCM decision logic to each significant failure mode: condition-based maintenance, predictive monitoring, time- or usage-based tasks, scheduled restoration or replacement, failure-finding tests for protective and standby functions, run-to-failure where risk is acceptable, operating changes, or engineering redesign.
  • Link every task to a defined failure mode with frequency or trigger, responsible role, work instruction, acceptance criteria and follow-up.
Read the full method

Preventive Maintenance Optimisation — PMO

Remove maintenance that adds work without adding reliability

  • Review the existing PM programme against criticality, equipment history, failure modes, operating context, task effectiveness and applicable manufacturer or regulatory requirements.
  • Identify duplicate tasks, tasks with no identifiable failure mode, missing activities, frequencies that are too short or too long, poor instructions, missing acceptance criteria, bundling opportunities, candidates for condition-based maintenance, and intrusive work that creates more risk than it controls.

Condition-Based and Predictive Maintenance

Intervene when condition indicates risk, not when the calendar says so

  • Identify the critical failure modes detectable through condition or performance: vibration and bearing condition, infrared thermography, lubricant and fluid condition, UPS and battery performance, power quality, temperature and differential pressure, chiller, pump and fan performance, cooling-water condition, leak detection, alarm and event patterns, and existing BMS, EPMS, DCIM or controller data.
  • Define monitoring routes, data requirements, alarm limits, escalation rules, diagnostic responsibilities and work-order triggers.

03

Strengthen resilience and eliminate repeat failure

Test whether redundancy survives contact with real maintenance, real failures and real people — and close out the failures that keep coming back.

03 — Strengthen resilience and eliminate repeat failure
03Strengthen resilience and eliminate repeat failure

RAM, Single-Point-of-Failure and Dependency Analysis

Understand how the whole service chain behaves when equipment fails

  • Map functional dependencies between utilities, electrical systems, generators, UPS, batteries, distribution, cooling plant, controls, fuel, water, communications and supporting infrastructure.
  • Assess reliability, availability and maintainability; redundancy and standby arrangements; single points of failure; hidden dependencies; common-cause failure; isolation and maintainability; planned maintenance states; repair and restoration time; degraded capacity; and failover and recovery scenarios.
  • Quantitative availability modelling where data supports it; structured qualitative or semi-quantitative scenario analysis where it does not.

Root Cause Analysis and Defect Elimination

Stop recurring failures instead of repeatedly restoring service

  • Investigate significant incidents, repeat failures, near misses and chronic problems using physical evidence, event logs, alarms, work orders, interviews, procedures, design information and operating history.
  • Examine technical, design, installation, operational, maintenance, organisational, vendor and human-performance contributors; assign owners, dates and effectiveness checks to corrective actions.

Outage, Change and Maintenance-Window Planning

Execute intrusive work without unnecessarily exposing the live load

  • Develop or independently review plans for maintenance, switching, testing and modification work: system states, isolation boundaries, temporary configurations, contingency and abort criteria, and restoration steps.
  • Review MOPs, SOPs and EOPs, roles and authorisations, vendor involvement, spares and equipment readiness, and communications.

04

Build the asset data foundation

Engineering decisions only survive if they land in the systems your team uses every day, in a structure that supports analysis.

04 — Build the asset data foundation
04Build the asset data foundation

CMMS / EAM Data Structure and Implementation Support

Put the engineering into the system your team actually uses

  • Structure and load master data: hierarchy, asset register, equipment attributes, job plans, task libraries, BOMs, spares links, failure coding aligned to ISO 14224, work types and priority rules.
  • Support configuration, data migration, quality control, acceptance testing, user procedures and data governance.
Read the full method

Reliability KPIs and Asset Health Framework

Measure emerging risk, not only yesterday's downtime

  • Define a balanced set of asset, maintenance and reliability indicators appropriate to the facility and the data available: availability and interruption, failure frequency and repeats, repair and restoration time, emergency and reactive work, PM and inspection effectiveness, condition findings, backlog and schedule compliance, asset condition, capacity margin, obsolescence exposure, data quality and corrective-action completion.
  • Where useful, combine condition, performance, failure history, criticality, obsolescence and supportability into an asset-health index.

Critical Spares and Obsolescence Management

Have the right part available when restoration time matters

  • Determine spares requirements from criticality, failure modes, installed population, expected demand, lead time, repairability, interchangeability, storage requirements, vendor support and restoration objectives.
  • Assess equipment, firmware, software, controls and components approaching end of support or technological obsolescence.

05

Make better lifecycle and asset-management decisions

Connect day-to-day asset decisions to service commitments, capital planning and the organisation's objectives.

05 — Make better lifecycle and asset-management decisions
05Make better lifecycle and asset-management decisions

Lifecycle, Renewal and Capacity Planning

Replace, refurbish, extend or defer — with evidence

  • Assess assets on condition, health, performance, criticality, maintenance cost, energy performance, capacity, utilisation, supportability, obsolescence and estimated remaining useful life.
  • Compare continued operation, targeted maintenance, refurbishment, life extension, technology upgrade, capacity expansion, replacement or decommissioning using life-cycle cost and total cost of ownership.

ISO 55001 Asset Management System, SAMP and Asset Management Plans

Connect daily asset decisions to business and service objectives

  • Assess and develop the governance, strategy, processes, information, roles and performance framework required for effective asset management.
  • ISO 55001 maturity and gap assessment, asset-management policy, scope and stakeholder requirements, value and decision criteria, Strategic Asset Management Plan (SAMP), objectives, Asset Management Plans, roles and RACI, risk and opportunity management, asset-information requirements, performance evaluation, management review and a continual-improvement roadmap.
Read the full method

Operational Readiness and Asset Handover

Make the facility operable and maintainable before go-live

  • Verify that people, processes, information, systems, materials and vendor arrangements are ready before operational acceptance.
  • Cover commissioning and test evidence, asset register and hierarchy, CMMS loading, maintenance plans and job instructions, BOMs and critical spares, SOPs/MOPs/EOPs, training and competency, vendor and service agreements, warranty requirements, operating limits and alarm settings, drawings, manuals and configuration records, and open defects.

Discuss scope

Tell us the facility, the systems that concern you, and the decision you are trying to make. We will tell you which engagement fits.