Heat sensing data center energy consumption and service life optimization method and system based on health degree and storage medium thereof

By constructing a health rating model and a temperature fluctuation memory window, and combining two-level time-scale scheduling, the problems of insufficient health assessment and lack of temperature fluctuation memory in existing technologies are solved, and the synergistic optimization of data center energy consumption and lifespan is achieved.

CN121835387APending Publication Date: 2026-04-10中邮建技术有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies fail to effectively integrate information such as the server's long-term operating history, temperature environment, and fault records, resulting in insufficient health assessment, lack of long-term memory of temperature fluctuations, and a single scheduling time scale, making it difficult to achieve coordinated optimization of energy consumption and lifespan.

Method used

A continuously updated server health rating model is constructed, a temperature fluctuation memory window is introduced, and a multi-time period joint optimization model is used. Combined with a two-level time scale scheduling strategy, IT energy consumption, cooling energy consumption and lifespan loss costs are optimized. Health interval division and business binding are adopted to achieve coordinated optimization of energy consumption and lifespan.

Benefits of technology

By explicitly reflecting the health status of servers through health scores and temperature fluctuation indicators, the damage to hardware caused by frequent temperature shocks is reduced. This approach balances overall planning with real-time response, thereby optimizing the overall energy consumption and lifespan of the data center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121835387A_ABST
    Figure CN121835387A_ABST
Patent Text Reader

Abstract

The invention discloses a heat sensing data center energy consumption and service life optimization method and system based on the health degree and a storage medium thereof, and belongs to the technical field of data center energy saving and equipment service life management. A sustainable updating server health degree scoring model is constructed, a temperature fluctuation memory window is introduced to describe a thermal fatigue effect, IT energy consumption, cooling energy consumption and life loss cost converted from health degree reduction are uniformly incorporated into a multi-time-period joint optimization model, and a two-stage time scale scheduling strategy is adopted to optimize the health degree of a server. On the premise that service requirements and temperature safety are guaranteed, the server operation environment and the load state are guided to evolve towards the direction of low energy consumption, low temperature impact and low failure rate, and collaborative optimization of the total energy consumption of the data center and the service life of the server is achieved. On the premise of meeting business requirements and temperature safety, comprehensive energy consumption is reduced, health degree attenuation is slowed down, the service life of the server is prolonged, and long-term operation stability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data center energy saving and equipment life management, in particular to a heat-aware data center energy consumption and life optimization method and system based on health degree and a storage medium thereof, which is suitable for green operation and operation and maintenance management of cloud data centers, edge data centers and industry private data centers. BACKGROUND

[0002] With the rapid development of cloud computing, big data and artificial intelligence business, the number of large and medium-sized data centers continues to increase, and the problems of data center energy consumption and carbon emissions are becoming increasingly prominent. Public data shows that data center electricity consumption accounts for an increasing proportion of overall electricity consumption, and a considerable part of energy consumption is on information technology equipment (IT equipment) and cooling systems, and the operating and energy costs continue to rise.

[0003] To reduce data center energy consumption, existing technologies mainly research and practice from the following aspects:

[0004] 1. Server-side energy saving

[0005] By means of server virtualization, integrated deployment, dynamic voltage frequency adjustment, server active shutdown or standby, etc., the business is concentrated on part of the servers to run, and the other servers enter a low-power state, thereby reducing IT energy consumption.

[0006] 2. Cooling side energy saving

[0007] By means of cold and hot aisle isolation, optimizing air flow organization, increasing chilled water supply temperature, fine control of air supply, using variable frequency fans, etc., the energy consumption of the refrigeration system is reduced, and the power utilization efficiency is improved.

[0008] 3. Heat-aware load distribution

[0009] On the basis of existing server integration and cooling energy saving, some schemes introduce cabinet temperature, server inlet temperature and other heat information to perform heat-aware scheduling on the working load, trying to avoid some areas from becoming "hot spots", thereby reducing cooling energy consumption under the premise of ensuring temperature safety.

[0010] 4. Multi-time period energy consumption optimization

[0011] There are also schemes that take multi-time period load distribution, server start-stop and cooling control as a whole optimization problem, and under the premise of guaranteeing business demand and temperature red line constraints, use mathematical programming, heuristic algorithms and other methods to optimize energy consumption in multiple time scales.

[0012] 5. Equipment health degree evaluation and operation and maintenance

[0013] In power systems, building operation and maintenance, and some data center scenarios, existing technologies assess the health of equipment based on operating history, fault records, and other information to guide maintenance planning, fault warning, or asset management. However, such health is mostly used for operation and maintenance reports, alarms, and maintenance decisions, and is less directly used in dispatch optimization models, and is not integrated with cooling control, load dispatch, and life cost for joint modeling.

[0014] Although the above technologies have achieved certain results in energy saving and operation and maintenance, there are still the following deficiencies:

[0015] 1. Lack of explicit and sustainable "health" dispatch state

[0016] Existing energy saving schemes usually only focus on power consumption and temperature at the current time, and some add server start-stop times, load migration times, etc. as penalty terms in the objective function, but do not integrate server long-term operation history, temperature environment, fault records, etc. into a sustainable "health" state, which cannot reflect the operation and maintenance experience that "some servers run for a long time at high intensity and need proper rest", and cannot directly constrain the health change to the dispatch decision.

[0017] 2. Only control instantaneous temperature, lack of memory and quantitative description of temperature fluctuations

[0018] Traditional thermal sensing methods generally constrain the server inlet temperature not to exceed a fixed red line value, but lack description of frequent and large changes in temperature within a certain time window. In fact, frequent up and down jumps in temperature within a certain time window will exacerbate the thermal fatigue of devices such as solder joints and fan bearings, reducing hardware life. Controlling only the upper limit of instantaneous temperature cannot reflect this long-term cumulative damage. Existing schemes generally lack methods to remember and quantify temperature fluctuation trajectories.

[0019] 3. Single time scale or rough division of dispatch, difficult to balance global and real-time

[0020] If a large-scale optimization model is directly constructed at a very fine time granularity, the solution process is complex and difficult to meet the real-time requirements of online dispatch; if only planning at a coarse granularity, it is also difficult to respond to business surges and temperature hotspots in a timely manner, which can easily lead to local overheating or frequent invalid start-stop. Some schemes propose multi-time scale dispatch, but mostly focus on the business layer or cooling layer, and do not integrate server health, temperature fluctuations, and other long-term factors with short-term load fluctuations into a coordinated and consistent dispatch framework.

[0021] 4. Energy consumption and life are not truly optimized together

[0022] Some schemes introduce start-stop cost, load migration cost and other penalty terms in the energy consumption optimization model, but often only as an auxiliary quantity in energy consumption optimization, and it is difficult to establish a clear quantitative relationship with server life attenuation, maintenance cost and other long-period goals. There is a lack of systematic methods to establish a connection between "health degree decline" and "life loss cost", and to jointly optimize them with IT energy consumption and cooling energy consumption, making it difficult to achieve true "energy consumption-life" collaborative optimization from the perspective of overall operation and maintenance cost.

[0023] Therefore, there is an urgent need for a new technical solution to introduce a sustainable server health degree index and temperature fluctuation memory mechanism on the basis of existing energy saving and temperature control, to unify IT energy consumption, cooling energy consumption and life loss cost in an optimization model through a multi-time scale, hierarchical and partitioned scheduling framework, to balance global planning and real-time response, and to achieve comprehensive optimization of data center energy consumption and server life. SUMMARY

[0024] The purpose of the present application is to overcome the shortcomings of the prior art, such as not explicitly modeling server health degree, not long-term memory of temperature fluctuations, and single scheduling time scale, energy consumption and life not collaborative optimization, and to propose a health degree-based thermal-aware data center energy consumption and life optimization method, system and storage medium.

[0025] The present application introduces a temperature fluctuation memory window to depict the thermal fatigue effect by constructing a sustainable server health degree scoring model, unifies IT energy consumption, cooling energy consumption and life loss cost converted from health degree decline into a multi-time period joint optimization model, and adopts a two-level time scale scheduling strategy to guide the server operating environment and load state to evolve towards "low energy consumption, low temperature impact, low failure rate" under the premise of ensuring business demand and temperature safety, and to achieve collaborative optimization of overall energy consumption and server life in the data center.

[0026] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0027] The health degree-based thermal-aware data center energy consumption and life optimization method is applied to a data center comprising multiple servers and a cooling system, and the method comprises at least the following steps:

[0028] Step 1. Health degree data acquisition and modeling

[0029] Periodically collect historical operating data and environmental data of each server in the data center, wherein the operating data at least includes processor utilization, cumulative running time, number of power-on and power-off, and fault records, and the environmental data at least includes server inlet temperature, cabinet temperature and room area temperature.

[0030] Based on the operation data and environment data, a server health score model is constructed to calculate a health score value H in the interval of 0~1 for each server, and the larger the health score value is, the better the long-term reliability is. The health score model includes a basic life factor and an operation stress factor, wherein the basic life factor is determined by static information such as server factory time, design life, and cumulative failure times; and the operation stress factor is calculated by dynamic information such as average utilization rate, peak utilization rate, start-stop frequency, average temperature, and temperature fluctuation index S_T.

[0031] Step 2. Temperature fluctuation memory construction and temperature fluctuation index generation

[0032] A temperature fluctuation memory window of a preset length is set for each server, and a server inlet temperature sequence is recorded in the window at a predetermined sampling period. The length of the temperature fluctuation memory window can be determined according to the business fluctuation period and the cooling system inertia, and the typical range is 30 minutes to 2 hours.

[0033] In the temperature fluctuation memory window, the maximum temperature change amplitude, the temperature change mean square value, and the number of times of crossing a set temperature difference threshold are calculated respectively, and the above indexes are normalized and combined into a temperature fluctuation index S_T in the interval of 0~1. The larger the temperature fluctuation index is, the more serious the recent temperature impact is. S_T is used as one of the input parameters of the health score model, and is used to update the health score value H. When S_T is greater than a preset threshold, the health decay coefficient is increased, so that the adverse effects of frequent temperature fluctuations on server life are explicitly reflected in the health evolution.

[0034] Step 3. Time discretization and classification of business requirements

[0035] A preset scheduling period is divided into multiple discrete time periods, the business request amount of each time period is predicted, and the business is classified according to the delay sensitivity and reliability requirements, which can be classified into real-time business, ordinary business, and postponable business, etc.

[0036] The equivalent computing resource amount required by each time period and each business level is estimated to obtain the working server number interval to be enabled for each time period and each business level, which provides a basis for subsequent health interval and business level binding.

[0037] Step 4. Joint optimization modeling of energy consumption and life

[0038] Under the multi-time period framework, the server start-stop state, load distribution ratio, and cooling system control quantity of each time period are used as optimization decision variables, and the weighted sum of IT energy consumption, cooling energy consumption, and life loss cost converted from the health score value H is used as the objective function, to construct a multi-time period energy consumption and life joint optimization model.

[0039] In this model, the following constraints are introduced:

[0040] Business demand satisfaction constraints for each time period ensure the quality of service requirements for each type of business;

[0041] Temperature safety constraints for each server inlet temperature do not exceed the temperature red line under the corresponding operating state;

[0042] Health threshold constraints for each server health score value H do not fall below the preset health threshold;

[0043] Cooling system control range constraints, including upper and lower limits of chilled water supply temperature, supply air temperature, fan speed, etc. control;

[0044] Constraints on the number of start-stop times, load migration times, and temperature variation amplitude to limit frequent start-stop and severe temperature fluctuations.

[0045] By introducing life loss cost in the objective function, and introducing health and temperature fluctuation related constraints in the constraints, the integration modeling and joint optimization of IT energy consumption, cooling energy consumption and life loss cost are realized.

[0046] Step 5. Server health interval division based on health and business binding

[0047] According to the health score value H of each server, the server is divided into health interval, and the server is divided into high health zone, medium health zone, low health zone and cooling recovery zone, etc. multiple health intervals, wherein:

[0048] Servers with health above the first threshold are classified as high health zone, and are preferentially loaded with real-time business;

[0049] Servers with health in the middle interval are classified as medium health zone, and carry ordinary business, and provide elastic expansion for real-time business when high health zone resources are insufficient;

[0050] Servers with lower health but not below the safety threshold are classified as low health zone, and carry delayed business or background tasks;

[0051] Servers with health below the minimum threshold are classified as cooling recovery zone, and only retain essential basic services or enter standby state, run at lower temperature and lower load, and slowly recover health.

[0052] Bind the business level with the health interval, and preferentially allocate business in the matching health interval in the subsequent scheduling process, to avoid long-term high-intensity business on servers with poor health from the source.

[0053] Step 6. Two-level time scale scheduling

[0054] In the first time scale (slow time scale), a plurality of discrete time periods are aggregated into a plurality of slow time scale time windows, in units of hours or days. Based on the joint optimization model, the current health score value H and the service prediction result, a server group level basic start-stop scheme, health zoning and the binding relationship between service levels and health intervals are generated; the number of servers required to be kept running in each slow time scale time window and the target interval of cooling system control quantity in each health zone are planned.

[0055] In the second time scale (fast time scale), each slow time scale time window is further divided into a plurality of fast time scale time periods, in units of minutes, for example, every 5 minutes a scheduling interval. In each fast time scale period, on the premise of the basic start-stop scheme, real-time service load, inlet temperature and temperature fluctuation index S_T are periodically obtained, and according to the service load deviation and temperature abnormality, only the start-stop state, load distribution and cooling system control quantity of part of the servers in the same health interval or adjacent health interval are locally adjusted under the preset constraint, and the cooling system control quantity is fine-tuned in a limited step, so as to obtain the final scheduling scheme meeting the real-time requirement.

[0056] During the adjustment process, the load migration range, the number of server start-stop times and the temperature change amplitude in each fast time scale period are limited to avoid large-scale migration introducing new temperature impact and unnecessary switching.

[0057] Step 7. Scheduling strategy execution and health closed loop update

[0058] According to the final scheduling scheme, the start-stop and load migration of the servers are controlled through the resource management platform, and the cooling control quantity such as chilled water supply temperature, air supply temperature and fan speed is adjusted through the cooling control system.

[0059] During the scheduling execution process, the latest operation data and temperature data are periodically collected, the temperature fluctuation memory window, the temperature fluctuation index S_T and the health score value H of each server are updated, and the updated health score value H is used for the next round of energy consumption and life joint optimization and two-level time scale scheduling, realizing the closed loop optimization control of energy consumption and life.

[0060] The application also provides a health-based thermally aware data center energy consumption and life optimization system for realizing the above method, comprising:

[0061] A health acquisition and evaluation module is used to periodically obtain server operation data and environment data from a monitoring system and an operation and maintenance system, and calculate the health score value H of each server based on a preset health model;

[0062] A temperature fluctuation analysis module is configured to maintain a temperature fluctuation memory window and generate a temperature fluctuation index S_T;

[0063] A service prediction and classification module is configured to predict service request volume in a future period of time, classify services into real-time services, ordinary services, and delayable services according to time delay sensitivity and reliability requirements, and generate resource demand intervals of each service level in each time period.

[0064] A two-level scheduling optimization module is configured to construct an energy consumption and life joint optimization model and perform slow time scale and fast time scale scheduling.

[0065] An execution control module is configured to control server start / stop and load migration according to the final scheduling scheme, adjust cooling system control variables, and feed back running data and temperature data to update the health score value H.

[0066] The modules are interconnected through a communication bus or network to cooperatively implement the above steps.

[0067] The application also provides a computer readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, performs each step of the above method.

[0068] Compared with the prior art, the application has the following beneficial effects:

[0069] 1. A sustainable server health index is introduced and incorporated into the scheduling optimization core state

[0070] The application integrates server long-term operation history, temperature environment, and fault records into a quantifiable health score, and uses the health score value as a constraint condition in a multi-time period joint optimization model and a basis for calculating life loss cost, so that the scheduling decision meets the current load and temperature requirements while considering the long-term reliability of the server.

[0071] 2. The thermal fatigue effect is depicted through a temperature fluctuation memory window

[0072] The application configures a temperature fluctuation memory window on each server, constructs a temperature fluctuation index S_T that comprehensively considers the maximum temperature change amplitude, temperature change mean square value, and temperature difference crossing frequency, and uses S_T as a key input of the health model. When the temperature fluctuation index is large, the health score decay rate is explicitly increased, thereby avoiding the problem of frequent temperature impact damage to hardware life that appears safe from the perspective of instantaneous temperature.

[0073] 3. Two-level time scale scheduling is adopted to balance global planning and real-time response

[0074] The application plans server group level start-stop, health zoning and service binding relationship on a slow time scale, and performs limited range local adjustment according to service burst and temperature fluctuation on a fast time scale, thereby effectively reducing the solving complexity of large-scale optimization of a single time scale, while retaining the fast response capability to service burst and hot area.

[0075] 4. Synergistic optimization of IT energy consumption, cooling energy consumption and life loss cost

[0076] The application considers IT energy consumption, cooling energy consumption and life loss converted from health degree reduction in the objective function, and flexibly balances between energy saving target and life protection target through adjustable weight coefficient, so as to realize long-term optimization from the perspective of overall operation and maintenance cost, rather than only reducing energy consumption index in the short term.

[0077] 5. Parameters and structure have engineering implementability and generalizability

[0078] The parameters such as data center scale, server power consumption range, temperature control range and sensor precision adopted by the application conform to current engineering practice, and the health degree interval, temperature red line, time scale and weight coefficient can be flexibly configured according to the scale and service characteristics of different data centers, so that the application is suitable for different construction conditions and operation and maintenance strategies of cloud data centers, edge data centers and industry private data centers, and has good engineering implementability and generalization value. BRIEF DESCRIPTION OF DRAWINGS

[0079] Figure 1 It is a schematic diagram of typical physical and deployment structure of the data center of the application.

[0080] Figure 2 It is a schematic diagram of health degree acquisition and evaluation process of the application.

[0081] Figure 3 It is a schematic diagram of temperature fluctuation memory window construction and temperature fluctuation index generation of the application.

[0082] Figure 4 It is a schematic diagram of service grading and server health interval binding relationship of the application.

[0083] Figure 5 It is a schematic diagram of overall process of two-level time scale scheduling of the application.

[0084] Figure 6 It is a schematic diagram of energy consumption and health degree comparison of the application and a traditional only energy consumption optimization scheme in a typical working day scenario. DETAILED DESCRIPTION

[0085] The technical solutions of the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the following embodiments are only used to illustrate the present application, but not to limit the present application; without departing from the spirit and essence of the present application, those skilled in the art can make various forms of equivalent replacement or improvement on the specific embodiments, which should be considered to fall within the protection scope of the present application.

[0086] Embodiment one: overall structure and module composition of the system

[0087] As shown in Figure 1 , the data center of the present embodiment includes several cabinets, cooling systems, power supply and distribution systems, network equipment and monitoring and management systems. The cabinet adopts a standard 42U cabinet and is arranged in a cold-hot aisle isolation manner, with the front side of each cabinet being a cold aisle and the rear side being a hot aisle.

[0088] At the software and control level, the following functional modules are deployed:

[0089] 1. Health degree acquisition and evaluation module

[0090] It is used to periodically obtain server running data and environment data from the monitoring system and operation and maintenance system, including but not limited to processor utilization, memory utilization, inlet temperature, start-stop state, cumulative running time, fault alarm record, etc., and calculate the health degree score value H of each server based on the preset health degree model.

[0091] 2. Temperature fluctuation analysis module

[0092] It is used to maintain a temperature fluctuation memory window for each server, record the inlet temperature sequence at a fixed time interval, calculate the maximum change amplitude of temperature, the mean square value of temperature change and the number of times of crossing the set temperature difference threshold, normalize the results to obtain the temperature fluctuation index S_T, and provide S_T to the health degree acquisition and evaluation module and the two-level scheduling optimization module.

[0093] 3. Business prediction and classification module

[0094] It is used to predict the business request volume in the future period of time, divide the business into real-time business, ordinary business and delayable business according to the delay sensitivity and reliability requirements of the business, and generate the resource demand interval of each business level in each time period.

[0095] 4. Two-level scheduling optimization module

[0096] It is used to construct an energy consumption and life joint optimization model, generate a server group level basic start-stop scheme and health zoning division at a slow time scale, and make local adjustment on the basic start-stop scheme according to the real-time business load, inlet temperature and temperature fluctuation index S_T at a fast time scale to generate a final scheduling scheme.

[0097] 5. Execution control module

[0098] For the final scheduling scheme, the resource management platform controls the start-stop state and load migration of the server, adjusts the chilled water supply temperature, air supply temperature and fan speed through the cooling control system, and feeds back the latest operation data and temperature data to the health degree acquisition and evaluation module and the temperature fluctuation analysis module, to realize closed-loop control.

[0099] The above modules can be deployed on the same physical server or different servers, or can be implemented in a distributed system manner.

[0100] Embodiment two: health degree modeling and temperature fluctuation memory

[0101] As shown in Figure 2 the health degree acquisition and evaluation process in this embodiment includes the following steps:

[0102] 1. Data acquisition

[0103] Get the CPU utilization rate, memory utilization rate, inlet temperature, start-stop state, cumulative running time, and recent fault records of each server from the computer room monitoring system at a fixed period (for example, every 5 minutes).

[0104] 2. Feature extraction

[0105] Statistical processing is performed on the collected raw data to obtain characteristic quantities representing long-term running stress, for example:

[0106] Average utilization rate and peak utilization rate within a period of time;

[0107] Average inlet temperature and maximum inlet temperature within a period of time;

[0108] Average start-stop frequency per unit time;

[0109] Recent fault frequency and fault level.

[0110] 3. Temperature fluctuation memory and temperature fluctuation index (such as Figure 3 )

[0111] Set a temperature fluctuation memory window for each server, and the window length can be 60 minutes. Record the inlet temperature sequence at 1 minute sampling within the memory window, and calculate:

[0112] Maximum temperature change amplitude;

[0113] Mean square value of temperature change;

[0114] Number of times the temperature crosses a set temperature difference (for example, 5℃).

[0115] The above indicators are normalized and combined into a temperature fluctuation indicator S_T in the interval 0-1. The larger S_T is, the more severe the temperature shock is.

[0116] 4. Health score calculation

[0117] The health score value H is obtained by combining the basic life factor and the operation stress factor. The basic life factor can be determined by the server manufacturing time, design life, cumulative failure times, etc.; the operation stress factor is obtained by weighted operation of the average utilization rate, peak utilization rate, average temperature, temperature fluctuation indicator, start-stop frequency, etc.

[0118] During operation, the health score value H gradually decays from 1 to 0 as time passes and stress accumulates. When the temperature fluctuation indicator S_T is large, the decay rate of the health score is increased accordingly, so as to reflect the adverse effect of strong temperature shock on life.

[0119] 5. Health zone division and threshold setting

[0120] The health score value is divided into multiple intervals, for example:

[0121] Servers with H greater than 0.8 are classified into the high health zone;

[0122] Servers with H between 0.5 and 0.8 are classified into the medium health zone;

[0123] Servers with H between 0.3 and 0.5 are classified into the low health zone;

[0124] Servers with H less than 0.3 are classified into the cooling recovery zone.

[0125] The thresholds of each interval can be adjusted according to the type of data center equipment and operation and maintenance strategy. Then, real-time business, ordinary business and postponable business are respectively bound to different health intervals for subsequent scheduling, as shown in Figure 4 .

[0126] Embodiment Three: Two-level time scale scheduling process

[0127] As shown in Figure 5 , this embodiment adopts a two-level time scale scheduling framework combining slow time scale and fast time scale.

[0128] 1. Slow time scale (hour level) scheduling

[0129] (1) Business prediction and demand aggregation

[0130] At the beginning of each daily or hourly scheduling period, a future several hours' traffic load prediction curve and resource requirement interval of each service level are generated according to historical traffic load and prediction model.

[0131] (2) Health zone resource planning

[0132] According to the current number of servers in each health zone and the distribution of their health degrees, the number of servers to be kept running in each hour is planned in high health zone, medium health zone, low health zone and cooling recovery zone.

[0133] (3) Basic start-stop scheme generation

[0134] Under the premise of ensuring business demand and health degree constraints, the basic start-stop scheme of each health zone in each hour is determined through a combined optimization model of energy consumption and life, that is, which servers are main working servers, which servers are in low-load operation, which servers enter standby or cooling recovery state, and the target interval of cooling system control quantity is also given.

[0135] 2. Fast time scale (minute level) scheduling

[0136] Each hour is divided into several fast time scale time periods, for example, every 5 minutes is a scheduling interval. In each fast time scale period, the following is performed:

[0137] (1) Real-time monitoring

[0138] The actual traffic request volume, server inlet temperature, temperature fluctuation index S_T and health degree update results in the past 5 minutes are obtained.

[0139] (2) Deviation analysis

[0140] The actual traffic load is compared with the predicted load under the slow time scale basic start-stop scheme to determine whether there is a significant deviation; according to the temperature and temperature fluctuation index, it is determined whether there are hot servers or servers with rapidly decreasing health degrees.

[0141] (3) Local scheduling adjustment

[0142] When the traffic load deviation or temperature anomaly exceeds the preset threshold, local adjustment is made to the load allocation and start-stop state within the limit. The adjustment is preferentially made within the same health zone or adjacent health zone to avoid new temperature impact caused by large-scale migration.

[0143] (4) Cooling system cooperative control

[0144] According to the latest load and temperature distribution, the chilled water supply temperature, supply air temperature and local cabinet fan speed are adjusted moderately. For example, the supply water temperature adjustment step is limited to no more than 1℃ per 10 minutes, and the cabinet fan speed is increased or decreased by 5% as a granularity.

[0145] 3. Implementation of objective function and constraint conditions

[0146] In slow time scale and fast time scale scheduling, the joint optimization objective includes at least IT energy consumption, cooling energy consumption and health degree degradation converted life loss, and the three parts are weighted and summed by adjustable weight coefficients. The constraint conditions include:

[0147] The service demand in all time periods must be met or controlled within the allowed service quality range;

[0148] All server inlet temperatures do not exceed the temperature red line of the corresponding operating state;

[0149] The health degree cannot be lower than the minimum safety threshold;

[0150] The number of cooling system control variables and server start-stop times does not exceed the maximum value allowed by the device.

[0151] Through the above two-level time scale scheduling process, the comprehensive optimization of energy consumption and life is realized under the premise of ensuring real-time and computability.

[0152] Example four: engineering parameter configuration example

[0153] This embodiment gives a set of specific parameters conforming to engineering practice to illustrate the implementability of the method of the present application in real data center scenarios.

[0154] 1. Room and cabinet scale

[0155] The room area is about 450 square meters, and cold and hot aisle isolation arrangement is adopted, with a cold aisle width of about 1.2 meters and a hot aisle width of about 1.0 meter;

[0156] The total number of cabinets is 32, arranged in two rows, with 16 cabinets in each row;

[0157] The cabinet is a 42U standard cabinet, of which not less than 36U is used for installing servers, and the remaining U is used for network equipment and power distribution units;

[0158] The IT load upper limit of each cabinet is designed to be 10 kilowatts, and the single cabinet IT load is controlled to be less than 8 kilowatts in actual operation, so as to reserve not less than 20% redundancy.

[0159] 2. Server type and power consumption range

[0160] There are two kinds of servers mainly deployed in data centers:

[0161] A-type servers (main computing type):

[0162] 2U rack-mounted dual-path servers, with processor thermal design power of about 140-205 watts, and maximum power supply capacity of no less than 750 watts;

[0163] Standby power consumption is about 90-130 watts, and the whole machine power consumption is about 230-280 watts, 360-420 watts and 500-600 watts at 30%, 60% and close to 100% load, respectively.

[0164] B-type servers (low-power type):

[0165] 1U rack-mounted dual-path servers, with maximum power supply capacity of about 550-800 watts;

[0166] Standby power consumption is about 60-90 watts, and the whole machine power consumption is about 150-190 watts, 250-300 watts and 340-400 watts at 30%, 60% and close to 100% load, respectively.

[0167] Single cabinet can mix and install A-type and B-type servers according to business needs, so that the single cabinet IT load can be long-term controlled in the safety range of 7-8 kilowatts.

[0168] 3. Sensor and monitoring device arrangement

[0169] At least 2 temperature sensors are arranged in front of and behind each cabinet to monitor the cold aisle inlet temperature and the hot aisle outlet temperature;

[0170] Server-level temperature monitoring probes are installed in some cabinets to record the server inlet temperature;

[0171] The temperature sensor accuracy is not less than ±0.5℃, and the sampling period is 1 minute;

[0172] Through the environmental monitoring system, temperature data, power consumption data and equipment state data are collected and stored uniformly, providing data sources for health degree model and temperature fluctuation memory.

[0173] 4. Health threshold and time scale configuration example

[0174] The high health threshold of health degree can be 0.8, the medium health threshold can be 0.5, and the low health threshold and cooling recovery threshold can be 0.3;

[0175] The temperature fluctuation memory window length is 60 minutes, and the temperature crossing threshold can be 5℃;

[0176] The slow timescale scheduling period is 1 hour, and the fast timescale scheduling interval is 5 minutes.

[0177] The cooling system water supply temperature is allowed to be adjusted within the range of 12 to 18℃, with a single adjustment step not exceeding 1℃.

[0178] 5. Application Effect Examples

[0179] In a typical workday scenario, this embodiment compares the method of the present invention with a traditional energy consumption optimization method. The results are as follows: Figure 6 As shown, this invention's method, while ensuring business service quality and temperature safety, can reduce IT energy consumption and cooling energy consumption on an annual scale, while slowing down the rate of decline in the health of critical servers, reducing lifespan loss caused by high-temperature shocks and frequent start-ups and shutdowns, and achieving lower overall operation and maintenance costs than traditional solutions.

[0180] Example 5: Optional Extensions and Variations

[0181] Without departing from the spirit of this invention, the invention can be extended and modified in various ways, including but not limited to:

[0182] 1. Health score models can employ various calculation methods, such as rule-based scoring methods, statistical learning-based regression models, and machine learning-based prediction models. The health score range can also be normalized within 0 to 100 or other intervals.

[0183] 2. The length of the temperature fluctuation memory window, the sampling period, and the temperature difference threshold can be adjusted according to the business fluctuation characteristics and cooling system inertia of different data centers, or variants such as sliding windows and adaptive windows can be used.

[0184] 3. Two-level time scale scheduling can be extended to multi-level time scale scheduling, such as introducing intermediate time levels between daily, hourly, and minute-level scheduling to meet the collaborative scheduling needs of hyperscale data centers or cross-regional data centers.

[0185] 4. In terms of business classification, in addition to real-time business, ordinary business and deferred business, it is also possible to add distinctions between types such as mission-critical business and batch processing business, and set differentiated health constraints and scheduling strategies for different business types.

[0186] 5. In terms of cooling system control, this invention can be applied not only to traditional air-cooled computer rooms, but also to data centers with liquid cooling, immersion cooling, or gas-liquid hybrid cooling. By adjusting the meaning and range of the cooling control quantities, the method of this invention can be extended to different cooling architectures.

[0187] 6. The method of the present application can also be combined with time-of-use pricing, renewable energy access, etc. strategies to further optimize the operation strategy of the data center in different electricity price periods under the premise of ensuring the health degree and temperature constraints.

[0188] Those skilled in the art can understand that various replacements, modifications or combinations of the above embodiments without changing the basic idea of the present application all fall within the protection scope of the present application.

Claims

1. A method for thermal-aware data center energy consumption and lifetime optimization based on health, characterized in that, The application comprises the following steps: Step 1: periodically collecting running data and environment data of each server, wherein the running data at least includes processor utilization, cumulative running time, number of power-on and power-off, and fault record, and the environment data at least includes server inlet temperature, cabinet temperature, and computer room area temperature; constructing a server health score model based on the running data and environment data, and calculating a health score value H located in a preset interval; Step 2: setting a temperature fluctuation memory window with a preset length for each server, recording a server inlet temperature sequence in the memory window at a predetermined sampling period, and generating a temperature fluctuation index S_T, wherein the S_T is used as an input parameter of the health score model to update the health score value H; Step 3: dividing a preset scheduling period into multiple discrete time periods, predicting the business request quantity of each time period, classifying the business according to the time delay sensitivity and reliability requirement, and obtaining the working server quantity interval required by each business level in each time period; Step 4: under the multi-time period framework, taking the server start-stop state, load distribution ratio, and cooling system control quantity of each time period as the optimization decision variables, and taking the weighted sum of IT energy consumption, cooling energy consumption, and life loss cost obtained by attenuation of the health score value H as the objective function, a multi-time period energy consumption and life joint optimization model is constructed; Step 5: based on the health score value H of each server, the server is divided into a health interval, and a binding relationship between the business level and the health interval is established, so that the business is preferentially distributed in the matched health interval; Step 6: two-level time scale scheduling is adopted: on the slow time scale, based on the joint optimization model, the current health score value H, and the business prediction result, a server group level basic start-stop scheme, health interval division, and binding relationship between the business level and the health interval are generated, and the server quantity interval required to be kept running in each slow time scale time window and the cooling system control quantity target interval of each health interval are planned; on the fast time scale, on the premise of the basic start-stop scheme, real-time business load, inlet temperature, and temperature fluctuation index S_T are periodically obtained, and only the start-stop state, load distribution, and cooling system control quantity of part of the servers in the same health interval or adjacent health interval are locally adjusted under the preset constraint according to the business load deviation and temperature abnormality, and the cooling system control quantity is fine-tuned with a limited step, so that the final scheduling scheme meeting the real-time requirement is obtained; Step 7: according to the final scheduling scheme, the start-stop and load migration of the server are controlled through the resource management platform, and the cooling control quantity such as chilled water supply temperature, air supply temperature, and fan speed is adjusted through the cooling control system; In the scheduling execution process, the latest running data and temperature data are periodically collected, the temperature fluctuation memory window, the temperature fluctuation index S_T, and the health score value H of each server are updated, and the updated health score value H is used for the next round of energy consumption and life joint optimization and two-level time scale scheduling, so that the closed-loop optimization control of energy consumption and life is realized.

2. The health-based thermal-aware data center energy consumption and lifetime optimization method of claim 1, wherein, The health score value H output by the health score model is a normalized value in the interval of 0-1, and the health score model includes a basic life factor and an operation stress factor, wherein the basic life factor is determined by the server factory time, design life and cumulative failure number, and the operation stress factor is calculated by the average utilization rate, peak utilization rate, start-stop frequency, average temperature and temperature fluctuation index S_T.

3. The health-based thermal-aware data center energy consumption and lifetime optimization method of claim 1, wherein, The length of the temperature fluctuation memory window is 30 minutes to 2 hours; the temperature fluctuation index S_T is calculated by the maximum temperature change amplitude, temperature change mean square value and the number of times that the temperature crosses the set temperature difference threshold, and is normalized to the interval of 0-1; when S_T is greater than a preset threshold, the health degradation coefficient is increased.

4. The health-based thermal-aware data center energy consumption and lifetime optimization method of claim 1, wherein, The constraint conditions of the joint optimization model include service demand satisfaction constraints in each time period, ensuring the quality of service of various services; temperature safety constraints, the server inlet temperature does not exceed the temperature red line under the corresponding operating state; health threshold constraints, the health score value H of each server is not less than a preset health threshold; the value range constraints of cooling system control quantities, including the upper and lower limits of the control quantities of chilled water supply temperature, air supply temperature and fan speed; start-stop frequency, load migration frequency and temperature change amplitude constraints, limiting frequent start-stop and severe temperature fluctuation.

5. The health-based thermal-aware data center energy consumption and lifetime optimization method of claim 1, wherein, The health interval includes a high health zone with a health score higher than a first threshold, a medium health zone with a health score in a medium interval, a low health zone with a lower health score but not lower than a safety threshold, and a cooling recovery zone with a health score lower than a minimum threshold; wherein the high health zone preferentially carries real-time services, the medium health zone carries ordinary services and provides elastic expansion for real-time services when the high health zone is insufficient, the low health zone carries postponable services or background tasks, and the cooling recovery zone only retains necessary basic services or enters standby, runs at a lower temperature and lower load, and slowly recovers the health score.

6. The health-based thermal-aware data center energy consumption and lifetime optimization method of claim 1, wherein, The slow time scale is a time window of hours or days, and the fast time scale is a minute-level scheduling interval.

7. A health-based thermal-aware data center energy consumption and lifetime optimization system, characterized in that, The optimization method is used to execute the optimization method of any one of claims 1-6, comprising a health score acquisition and evaluation module, a temperature fluctuation analysis module, a service prediction and classification module, a two-level scheduling optimization module, and an execution control module; The health score acquisition and evaluation module is used to collect the operation data and environmental data of each server, and construct a server health score model to obtain a health score value H; the temperature fluctuation analysis module is used to maintain a temperature fluctuation memory window and generate a temperature fluctuation index S_T; the service prediction and classification module is used to divide a preset scheduling period into multiple discrete time periods, predict the service request quantity in each time period, and classify the services according to the delay sensitivity and reliability requirements; the two-level scheduling optimization module is used to construct an energy consumption and life joint optimization model and execute slow time scale and fast time scale scheduling; the execution control module is used to control the server start-stop / load migration according to the final scheduling scheme and adjust the cooling system control quantity, and feedback the operation data and temperature data to update the health score value H.

8. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to execute the optimization method of any one of claims 1 to 6.