The number of hours a data centre will operate without interruption is determined not when the first server is installed, but during the weeks in which the capacity schedule is prepared, the redundancy level selected and the airflow path designed. Most outages we encounter in the field arise not from equipment failure, but from decisions that were omitted or made incorrectly during installation. The following seven topics address, in sequence, the issues that the project team must clarify at an early stage.

1. Plan capacity by power per rack, not by rack count

A data-centre requirement is often described as “a room for twenty racks”. Yet it is not the rack itself that determines the infrastructure, but the power per rack. The same twenty racks represent a 60 kW facility at 3 kW per rack, while at 15 kW they require an entirely different 300 kW power and cooling architecture. The difference changes every element, from panel selection and cable sizing to the cooling method.

Three figures must be stated separately during planning:

  • Current average load: The power that will actually be drawn on the day of commissioning.
  • Design load: The target power that the infrastructure can support continuously.
  • Growth scenario: The capacity expected to be added within three to five years and the stage at which it will be commissioned.

Failing to distinguish these three figures creates two typical errors. Either the infrastructure is built around today's load and the entire room must be reconsidered at the first expansion, or it is oversized for a distant target, inflating the initial investment and leaving equipment inefficient at low load. The correct approach is to design the backbone for target capacity and phase the commissioning: cable routes, main distribution and cooling infrastructure are built to accommodate growth, while equipment is added progressively.

2. Select the redundancy level together with the maintenance scenario

When redundancy is discussed, the usual question is “What happens if something fails?” The more decisive question is: Do I have to stop operations to perform maintenance? A system may tolerate a failure, but if it loses redundancy during maintenance, every planned maintenance activity becomes a window of risk.

LevelWhat it meansDuring maintenance
NA single path sized exactly to meet demandAn outage is required
N+1One additional redundant component beyond demandOne component can be maintained, but the system operates without redundancy
2NTwo fully independent pathsOne path operates with its redundancy intact while the other is maintained

When making this selection, remember that redundancy must be sustained across the entire chain. If the uninterruptible power supply is designed as 2N but the output panel remains a single point, the facility's actual level is determined by that single panel. Redundancy must therefore be decided along the complete path from the incoming supply to the rack outlet, not equipment by equipment.

The true test of redundancy is not a failure, but maintenance day. If you cannot maintain a system without shutting it down, that system operates without redundancy several planned times each year.

3. Solve cooling through airflow, not equipment alone

A significant proportion of cooling problems is caused not by insufficient capacity, but by cold air circulating in the wrong places. If temperatures rise in the upper sections of racks even though the combined capacity of the cooling units is sufficient, the problem lies in the airflow path rather than the cooling plant.

The practical installation considerations are:

  • Hot- and cold-aisle separation: Racks must be arranged with their intake faces towards the same aisle. Otherwise, hot air discharged by one rack enters the intake of its neighbour.
  • Blanking unused rack space: If empty rack units are not closed with blanking panels, cold air escapes to the rear without passing through the equipment.
  • Underfloor organisation: Congested cabling beneath a raised floor obstructs airflow; vent positions must be planned together with load distribution.
  • Operating range: Cooling the room more than necessary increases energy consumption. Accepted temperature and humidity ranges for IT equipment should be the starting point for the design.

A properly designed airflow path also increases the number of days on which free cooling can be used, directly reducing operating costs. The purpose of cooling automation is not to operate each unit independently, but to balance load among the units, bring the standby unit online at the correct time and manage the transition from mechanical to free cooling when outdoor conditions permit.

4. Do not leave the monitoring layer until later

Monitoring is one of the elements most frequently deferred during installation: “Get the room running first; measurements can be added later.” Yet adding power analysers, temperature and humidity sensors, and equipment communications retrospectively increases cabling costs and loses the baseline data from the facility's first months. Those initial months provide the very baseline against which future capacity decisions will be compared.

The minimum measurement set is:

  • Energy consumption at the incoming supply, uninterruptible-power output and rack-row levels
  • Cold-aisle intake temperature and relative humidity at multiple points along rack height
  • Cooling-unit operating status, discharge temperature and fault signals
  • Status information from generators, transfer switches and battery banks

Once this data is consolidated within a common monitoring layer, two capabilities become possible: tracking facility efficiency through PUE over time and making capacity decisions from measurements. The answer to “Can we add another rack to this row?” is no longer an estimate, but the row's measured load curve. This is the data-centre equivalent of the approach to turning operational data into a decision-making resource.

Alarm management is different from an alarm list

A common mistake in monitoring projects is to turn every available signal into an alarm. Hundreds of equal-priority alarms begin to be ignored within a few weeks. Events should be classified into those requiring intervention, those that need only monitoring and those recorded for information; every intervention alarm must state who is responsible for doing what.

5. Separate the automation network from the corporate network

A data centre's building and power automation belongs to a different world from the IT systems it hosts: devices have long lifecycles, their software is updated infrequently and many communicate using protocols without authentication. Connecting these devices directly to the corporate network creates a path through which an office-side incident can reach cooling controls.

The measures required during installation are not expensive, but implementing them retrospectively is:

  • Define the automation network as a separate zone and route communications through controlled points
  • Provide maintenance suppliers with remote access that includes authentication, time limits and session recording
  • Back up controller and monitoring-server configurations and record all changes

For a detailed treatment of these topics, explore our OT cybersecurity solution. Data-centre automation requires the same security discipline as the manufacturing floor.

6. Do not reduce commissioning to a “does it work?” test

Confirming that each component operates individually is not commissioning. Outages usually emerge from the combined behaviour of systems that work correctly in isolation: when utility power fails, the uninterruptible supply engages, the generator starts and transfer occurs—but all cooling units restart simultaneously, and the resulting inrush current places stress on the transfer. Or the generator comes online, but the cooling plant is not connected to the standby supply, causing the room to heat up within minutes.

Commissioning must therefore proceed in stages:

  1. Component tests: Validate each item of equipment independently.
  2. System tests: Test the power and cooling systems as complete systems in their own right.
  3. Integrated systems test: Rehearse scenario-based outages under real load—including utility failure, failure of the generator to start and loss of a single cooling unit.

The objective of integrated testing is not merely to pass the system, but to discover weaknesses before handover. These tests also allow the operations team to experience the system under real conditions for the first time. During many subsequent incidents, the decisive factor is that the team has encountered that scenario before.

7. Write the operating rules before handover

Operations begin where installation ends, and this is where availability is ultimately secured or lost. The handover file must include current single-line diagrams, panel layouts, labelling schedules, alarm-to-response mappings and maintenance intervals. Without them, the infrastructure becomes undocumented as soon as the first change is made.

The following must be clear before handover:

  • Which maintenance task will be performed, how often and by whom?
  • Are critical spare parts held in stock, and what are their lead times?
  • How does the escalation chain operate when an alarm occurs, and what is the target response time?
  • What approval and record-keeping process governs work such as adding a rack or changing a power feed?

Change records are particularly important: a significant share of outages occurs when an undocumented change meets an intervention made months later on the basis of an incorrect assumption. After-sales engineering support and regular reviews are integral to installation if the infrastructure is to retain its day-one standard.

In summary

A data-centre installation is not the sum of independent equipment purchases. Capacity, redundancy, airflow, monitoring, network security, commissioning and operating rules are parts of the same architecture; leaving one incomplete also limits the benefit delivered by the others. Making these decisions under a single engineering responsibility is the most effective way to avoid paying a much greater price later.

You can find full details of the installation scope, operating model and the work delivered by our own teams on our data-centre installation and infrastructure automation solution page.