An improved power management method

By relocating high-priority applications and managing server operations using state information, the method addresses data centre downtime during power outages, ensuring continued functionality and efficient power usage.

WO2026115136A1PCT designated stage Publication Date: 2026-06-04BRITISH TELECOM PLC

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BRITISH TELECOM PLC
Filing Date
2025-11-28
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Data centres experience downtime due to power outages, as existing backup methods like UPS only provide auxiliary power for a limited time, rendering servers inactive until mains power is restored, leading to inefficiencies.

Method used

A method and system utilizing state information to relocate high-priority applications and manage server operations during power failures, including actions like shutting down low-priority applications, reducing CPU frequencies, and moving applications to other data centres based on latency and priority, ensuring continued operation on reduced power.

Benefits of technology

Enables data centres to maintain functionality during power outages by prioritizing application relocation and power management, minimizing downtime and optimizing resource usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025084752_04062026_PF_FP_ABST
    Figure EP2025084752_04062026_PF_FP_ABST
Patent Text Reader

Abstract

There is herein described a method of operating a data centre in response to a reduction in a power supply to the data centre, the method comprising using state information to relocate one or more high priority applications running on the one or more servers, away from the one or more servers, where the state information comprises information relating to one or more servers located at the data centre and / or one or more applications running on the one or more servers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] An improved Power Management Method

[0002] Data centres can suffer from power outages, i.e. a failure of the mains power supply to the data centre.

[0003] In order to prevent data being lost, data centres are provided with an Uninterruptible Power Supply which provides auxiliary power to the data centre. This auxiliary power is a lower level of power than the mains supply and only lasts a limited time. During this time, the servers at the data centre are backed up and gracefully shut down. Although data loss is prevented, this method has the disadvantage that the severs are inactive until the mains power supply has been re-established.

[0004] It would be desirable to overcome and / or substantially mitigate some or all of the above- mentioned and / or other disadvantages of the prior art.

[0005] According to a first aspect of the invention there is provided a method of operating a data centre in response to a reduction in a power supply to the data centre, the method comprising: using state information to relocate one or more high priority applications running on the one or more servers, away from the one or more servers, where the state information comprises information relating to one or more servers located at the data centre and / or one or more applications running on the one or more servers.

[0006] The step of using state information to relocate the one or more high priority applications may comprise inputting the state information to a control algorithm. The control algorithm may identify the high priority applications and send an instruction to a management module to relocate one or more high priority applications.

[0007] The state information may comprise a latency requirement for one or more of the applications. A selection of which of the one or more high priority applications is to be relocated from the one or more servers may be carried out in dependence on the latency requirements for the applications. In a case where an application is to be moved to a further data centre, a selection of which of a plurality of data centres to host the moved application is made in dependence on a latency metric for each of the plurality of data centres with respect to the application to be moved.

[0008] The present technique is therefore particularly applicable when used with edge sites for hosting workloads that may have a geographic constraint, perhaps due to latency sensitivity.

[0009] The state information may comprise a geographical indication for a user of one or more of the applications, and a destination server to which one of the applications is to be moved may be selected in dependence on a geographical relationship between the user and the target server. The geographical relationship may be one or both of a spatial distance and a data distance (for example a number of links within a network between the user and the target server, or a total distance of those links between the user and the target server).

[0010] State information may comprise information identifying the one or more servers. The state information may further comprise data relating to the resources contained in each of the one or more servers. That data may comprise one or more of:

[0011] CPU count;

[0012] RAM size; storage capacity; chip manufacturer identity.

[0013] State information may comprise information identifying the one or more applications running on the servers. State information may further comprise data indicative of one or more of: relative priorities of the one or more applications, hardware dependencies of the one or more applications; dependencies on other workloads of the one or more applications; sensitivities to resource constraints of the one or more applications.

[0014] State information may further comprise: data identifying which of the one or more servers, each of the one or more applications is running on. the total current utilisation of server resources; system usage statistics; process statistics; power consumption by each of the one or more servers; and power input state to the data centre.

[0015] State information may be obtained from the one or more servers using telemetry.

[0016] A management module may provide the state information to the control algorithm and may do so on a regular basis. The control algorithm may be performed by a control algorithm module. The reduction in the power supply may be a failure of the mains power supply to the data centre.

[0017] In the event of a power failure, an Uninterruptible Power Supply may provide auxiliary power to the data centre for a limited time. Furthermore, in response to the power failure, the Uninterruptible Power Supply may send a trigger signal to the control algorithm module, indicating that power has been lost. In response to the trigger signal, the control algorithm module may perform the control algorithm using the most recent values of the state information it has received from the management module as inputs. The control algorithm may also use historical state information as input. The control algorithm may output an instruction to the management module to perform one or more of the following actions:

[0018] Discontinue the running on one or more servers of one or more applications that the state information indicates to be as low-priority and shut down the one or more servers upon which they run;

[0019] Move high-priority applications onto a small number of servers, eg one single server, and shut down the other servers;

[0020] Reduce the core frequency of the CPUs in one or more servers;

[0021] Reduce the uncore frequency of the CPUs in one or more servers;

[0022] Reduce the memory bandwidth of one or more servers;

[0023] Discontinue one or more applications running on the one or more servers and instruct a further data centre to run the one or more applications, and shut down the one or more servers;

[0024] Back up the one or more servers;

[0025] Shut down the one or more servers. The data centre may be located at the edge of a telecommunications network.

[0026] According to a second aspect of the invention there is provided a power management system for managing the operation of a data centre in response to a reduction in a power supply to the data centre, the system comprising:

[0027] A management apparatus adapted to use state information to relocate one or more high priority applications running on the one or more servers, away from the one or more servers, where the state information comprises information relating to one or more servers located at the data centre and / or one or more applications running on the one or more servers.

[0028] The management apparatus may comprise a control algorithm module and a management module. The features noted above in relation to the first aspect of the invention are applicable to the second aspect of the invention.

[0029] An embodiment of the invention will now be described in detail, for illustration only, with reference to the appended drawings, in which:

[0030] Fig 1 is a schematic view of a data centre and backup power system according to the prior art;

[0031] Fig 2 is a schematic view of a data centre and backup power system according to embodiments of the invention.

[0032] Fig 1 shows a schematic view of a data centre and backup power system according to the prior art. The data centre is located at an edge location. This means it is located closer to the network edge (ie the customer) than a centralised data centre. This is significant because, by bringing computing resources closer to the edge of the network, the processing takes place nearer to where data is generated and consumed (the customer). As a result, the user experiences a lower latency. The data centre comprises servers 15, 16 and 17 on rack 14. Infrastructure management component 18 manages the operation of the servers in a manner that would be familiar to the person skilled in the art. The mains power 1 1 supplies the rack 14 via Uninterruptible Power Supply (UPS) 12. In the event of a power outage, power form the mains power 11 is lost. UPS 12 detects this outage and sends a trigger signal to power generator 13. The trigger signal causes power generator 13 to provide a temporary power supply to UPS 12, which in turn directs that temporary power supply to the server rack 14. During the time that the temporary power supply lasts, the infrastructure management component 18 instructs the servers to back up their systems and perform a graceful shutdown.

[0033] This ensures that all data is saved. However, this approach has the disadvantage that the data centre is then inactive until power returns.

[0034] Fig 2 is a schematic view of a data centre and backup power system according to embodiments of the invention. Fig 2 has many like components with Fig 1. These components are numbered in a like fashion. In addition to the components of Fig 1 , Fig 2 also has a rescue algorithm component 29. The rescue algorithm component 29 is electrically connected to both the UPS 22 and the infrastructure management component 28. Infrastructure management component 28 provides state information to rescue algorithm component 29.

[0035] The events following a power outage using the arrangement of Fig 2 will now be described. In particular, if power from the mains power supply 21 is lost, this outage is detected by UPS 22. UPS 22 then sends a trigger signal to power generator 23. The trigger signal causes power generator 23 to provide a temporary power supply to UPS 22, which in turn directs that temporary power supply to the server rack 24. In contrast to the method of the prior art, the servers 15, 16, 17 are not backed up and shut down automatically.

[0036] In normal operation, infrastructure management component 28 transmits state information to the rescue algorithm component 29. State information comprises i) hardware information and ii) workload information.

[0037] Hardware information comprises information relating to the current hardware at the data centre. For example hardware information comprises an inventory of the servers along with their respective resource eg CPU count, RAM, storage capacity, and chip manufacturer identity. Workload information comprises an inventory of the applications running on the servers. This contains information such as the relative priorities of the different applications, their hardware dependencies, their dependencies on other workloads, and their sensitivities to resource constraints. Workload information also contains workload-related information obtained in real time eg through telemetry. This information comprises, eg the identity of the server an application is running on, and the total current utilisation of server resources, system usage statistics, process statistics, power consumption and power input state.

[0038] The workload information may also comprise an indication of a latency requirement for one or more of the applications. This may be expressed in a number of ways, including an expectation of communication time or round trip time between a server on which the application is running and the (computer of) the end user of the application. It may also be expressed geographically, in terms of a spatial distance (for example in km), or a data distance (for example a number of communication links, or a combined length of the the communication links, or a measure of the combination of both) between the server and the end user. By having a latency requirement for an application, the system is able to prioritise actions in event of a power outage in a manner which results in an application being retained (not moved) if it has a particular latency requirement only served by that server, or in a manner which results in an application moving to a specific data centre / server selected because it too satisfies the latency requirement for the application.

[0039] In response to the power state signal from UPS 22 indicating that power has been lost, the rescue algorithm component 29 inputs the latest values of the state information it has received from the infrastructure management component 28 into a rescue algorithm. The rescue algorithm outputs one or more power-conservation actions. Some examples of these actions are:

[0040] Discontinue the running on servers 25, 26 and 27 of applications identified by the state information as low-priority and shut down the servers upon which they run;

[0041] Move high-priority applications onto a small number of servers, eg one single server, and shut down the other servers;

[0042] Reduce the core frequency of the CPUs in servers 25, 26 and 27;

[0043] Reduce the uncore frequency of the CPUs in servers 25, 26 and 27; Reduce the memory bandwidth of the servers 25, 26 and 27;

[0044] Discontinue one or more applications running on the servers 25, 26 and 27 and instruct a different data centre (not shown in Fig 2) to run the one or more applications on its servers. The servers 25, 26 and 27 upon which the applications ran are then shut down.

[0045] As explained above, certain applications may have specific latency requirements. This means that they can only effectively be hosted at a server (or data centre) satisfying those latency requirements. This may mean that only data centres at an edge (that is, close to the customer) may be able to satisfy these requirements. There are several consequences of this recognition.

[0046] Firstly, when identifying which application to be shut down at the server / data centre experiencing the power outage, applications having less strict latency requirements may be offloaded to a different data centre (further from the edge / customer) in preference to applications having stricter ones.

[0047] Secondly, when identifying to which other data centres particular applications are to be relocated, an application may be relocated to another data centre only if that data centre satisfies the latency requirements for that application. In other words, a selection of a particular data centre for receiving and hosting the application to be moved is made on the basis of the capabilities of data centre in satisfying latency requirements.

[0048] Rather than considering latency requirements directly, the latency requirement may be expressed geographically, as explained above. This is equivalent, since latency is a function of the relative location of the end user and the server hosting the application being accessed by the end user.

[0049] It will be appreciated that applications may be selected to be moved (or not moved) based both on latency and priority (as discussed herein). These considerations may be combined in any appropriate manner depending on use case.

[0050] As these actions result in the shutting down of some of the servers, and / or running the remaining servers on reduced power, the datacentre can continue to operate on lower power than normal. As the algorithm takes the current power consumption into account, the instructed actions take account of the remaining backup power. If it is about to run out, the remaining active servers are backed up and a graceful shut down is implemented.

Claims

9Claims1 .A method of operating a data centre in response to a reduction in a power supply to the data centre, the method comprising: using state information to relocate one or more high priority applications running on the one or more servers, away from the one or more servers, where the state information comprises information relating to one or more servers located at the data centre and / or one or more applications running on the one or more servers.

2. A method as claimed in claim 1 , wherein the state information comprises a latency requirement for one or more of the applications.

3. A method as claimed in claim 2, wherein a selection of which of the one or more high priority applications is to be relocated from the one or more servers is carried out in dependence on the latency requirements for the applications.

4. A method as claimed in claim 2 or claim 3, wherein in a case where an application is to be moved to a further data centre, a selection of which of a plurality of data centres to host the moved application is made in dependence on a latency metric for each of the plurality of data centres with respect to the application to be moved.

5. A method as claimed in any preceding claim, wherein the state information comprises a geographical indication for a user of one or more of the applications, and wherein a destination server to which one of the applications is to be moved is selected in dependence on a geographical relationship between the user and the target server.

6. A method as claimed in claim 5, wherein the geographical relationship comprises one or both of a spatial distance and a data distance.

7. A method as claimed in any preceding claim , wherein the state information comprises information identifying the one or more servers.

8. A method as claimed in any preceding claim, wherein the state information further comprises data relating to the resources contained in each of the one or more servers, the data comprising one or more of:CPU count;RAM size; storage capacity; chip manufacturer identity.

9. A method as claimed in any preceding claim, wherein the state information comprises data indicative of one or more of: relative priorities of the one or more applications, hardware dependencies of the one or more applications; dependencies on other workloads of the one or more applications; sensitivities to resource constraints of the one or more applications.

10. A method as claimed in any preceding claim, wherein the state information further comprises one or more of: data identifying which of the one or more servers, each of the one or more applications is running on; the total current utilisation of server resources; system usage statistics; process statistics; power consumption by each of the one or more servers; and power input state to the data centre.1 1. A method as claimed in any preceding claim, wherein the state information is obtained from the one or more servers using telemetry.

12. A method as claimed in any preceding claim, wherein the step of relocating one or more high priority applications running on the one or more servers, away from the one or more servers, comprises: moving high-priority applications onto a small number of servers, eg one single server, and shut down the other servers; and / or11 discontinuing one or more applications running on the one or more servers and instruct a further data centre to run the one or more applications, and shut down the one or more servers.

13. A method as claimed in any preceding claim, wherein the step of using state information to relocate one or more high priority applications running on the one or more servers, away from the one or more servers comprises performing a control algorithm using the state information as inputs, and instructing a management module to perform the step of relocating one or more high priority applications running on the one or more servers, away from the one or more servers.

14. A method as claimed in any preceding claim, wherein the control algorithm outputs an instruction to the management module to perform one or more of the following actions: discontinue the running on one or more servers of one or more applications that the state information indicates to be as low-priority and shut down the one or more servers upon which they run; move high-priority applications onto a small number of servers, eg one single server, and shut down the other servers; reduce the core frequency of the CPUs in one or more servers; reduce the uncore frequency of the CPUs in one or more servers; reduce the memory bandwidth of one or more servers; discontinue one or more applications running on the one or more servers and instruct a further data centre to run the one or more applications, and shut down the one or more servers; back up the one or more servers; shut down the one or more servers.

15. A power management system for managing the operation of a data centre in response to a reduction in a power supply to the data centre, the system comprising: a management apparatus adapted to use state information to relocate one or more high priority applications running on the one or more servers, away from the one or more servers, where the state information comprises information relating to one or more servers located at the data centre and / or one or more applications running on the one or more servers.