Cooperative scheduling method and device of data center power distribution system

By constructing a digital twin coupled with multiple physical fields and a multi-agent decision-making framework, the business continuity problem of data center power distribution systems under multiple concurrent failures is solved, and intelligent autonomous scheduling and zero-interruption business guarantee are achieved under extreme failure scenarios.

CN121769893APending Publication Date: 2026-03-31SUZHOU A RACK INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing data center power distribution systems cannot ensure business continuity under the constraints of cross-system resource conflicts when faced with multiple concurrent failures, especially in extreme failure scenarios where they cannot provide reliable zero-interruption guarantees.

Method used

A digital twin containing a multi-physics coupling model is constructed, and the influence chain mapping relationship between physical resources and logical IT load is established. The decision target weights are dynamically adjusted through a multi-agent decision framework to generate and select scheduling strategies that can maintain the performance of the risk load set above a preset safety threshold.

Benefits of technology

It enables full-range identification and dynamic global resource scheduling of the impact on business under concurrent failures of multiple physical resources, ensuring the business continuity and stable operation of the data center under extreme failure scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121769893A_ABST
    Figure CN121769893A_ABST
Patent Text Reader

Abstract

The invention discloses a collaborative scheduling method and device for a data center power distribution system, and relates to the field of data center power distribution scheduling, and the method comprises the steps: constructing a digital twin containing a multi-physics field coupling model, building a mapping relation, responding to a physical resource fault event, determining a risk load set, and carrying out the collaborative scheduling of the data center power distribution system. And establishing a multi-agent decision framework based on the digital twin, adjusting the weight to obtain a candidate scheduling strategy, simulating execution of the candidate scheduling strategy in the digital twin to screen a security strategy set, and converting the optimal security strategy into a control instruction to be issued and executed. According to the method, when a physical resource fault occurs in the data center power distribution system, the affected logic IT load can be accurately determined, the effective scheduling strategy is generated through multi-agent collaborative decision, and the security strategy which ensures that the load performance is above the preset security threshold is screened out and executed, so that the load performance is ensured to be more than the preset security threshold. And the fault-tolerant capability and the stability of the power distribution system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power distribution management technology, and in particular to a collaborative scheduling method and apparatus for a data center power distribution system. Background Technology

[0002] With the deepening of digital transformation, business continuity in data centers has become a core lifeline. Currently, high availability assurance for data center power distribution systems mainly relies on equipment redundancy and backup switching mechanisms (such as ATS and dual power supply). However, these solutions have inherent limitations: firstly, they can only cope with pre-set single device or link failures, and are powerless against multiple concurrent failures not covered by the plan, especially those spanning power distribution, cooling, and computing systems; secondly, their decision-making logic is static and local, and cannot dynamically perform global resource scheduling to prioritize the most critical business in conflict environments where system resources are limited due to failures.

[0003] While existing technologies have employed digital twins for system modeling and multi-agent reinforcement learning for resource optimization, their optimization objectives are primarily focused on economic indicators such as improving energy efficiency and reducing costs, or they treat system constraints as soft optimization conditions. These solutions lack effective mechanisms for dynamic, safe, and collaborative decision-making in extreme failure scenarios under the highest priority hard constraint of "zero business interruption." Therefore, existing technologies cannot provide reliable guarantees when facing complex failures that could lead to business disruption.

[0004] Therefore, there is an urgent need for a global autonomous method and system that can enable data center power distribution systems to self-sensitize, self-determine, and self-execute under extreme multi-node failure scenarios. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a collaborative scheduling method and apparatus for a data center power distribution system. This application can solve the technical problem that existing data center power distribution management systems cannot ensure zero interruption of the highest service level agreement (SLA) service under the constraint of cross-system resource conflicts when facing multiple concurrent failures.

[0006] Firstly, this application provides a collaborative scheduling method for a data center power distribution system, employing the following technical solution: A collaborative scheduling method for a data center power distribution system includes the following steps: Construct a digital twin containing a multiphysics coupling model and establish an influence chain mapping relationship between physical resources and logical IT load identifiers; In response to multiple detected failure events of the physical resources, based on the mapping relationship between the digital twin and the impact chain, the logical IT loads that have an abnormal impact associated with the failure events are identified and a risk load set is formed. A multi-agent decision-making framework is established based on the digital twin. The risk load set is input into the multi-agent decision-making framework, and the decision target weights of each agent in the multi-agent decision-making framework are dynamically adjusted to obtain at least one candidate scheduling strategy. Simulate the execution of all candidate scheduling strategies in the digital twin, and based on the simulation results, select a set of security strategies that can maintain the performance of the risk load set above a preset security threshold from the candidate scheduling strategies. The optimal security policy from the set of security policies is converted into control commands and issued for execution.

[0007] By adopting the above technical solutions, a digital twin containing a multi-physics coupling model is constructed and an influence chain mapping relationship is established, which can accurately reflect the relationship between physical resources and logical IT load. Through a closed-loop process of risk load location, multi-agent strategy generation, simulation screening, and instruction execution, loads affected by faults can be accurately identified. Dynamic global resource scheduling is performed in conflict environments with limited system resources, solving the problem that traditional power distribution scheduling cannot cope with concurrent faults of multiple physical resources. It realizes intelligent autonomous scheduling in data center power distribution system fault scenarios, taking into account both fault response speed and business continuity assurance.

[0008] Preferably, the construction of a digital twin including a multiphysics coupling model and the establishment of an influence chain mapping relationship between physical resources and logical IT load identifiers specifically includes the following steps: A digital twin is obtained by performing multi-physics coupling basic modeling. The multi-physics coupling model includes an electrical model, a thermodynamic model, and an IT business logic model. The IT business logic model is used to abstract virtual machines and containers running in the data center into logical IT loads with unique identifiers and record the physical resource information on which the logical IT loads depend for operation. In the digital twin, each logical IT load identifier is associated with the corresponding physical resource information to form a dynamic influence chain mapping relationship. The physical resource information includes the server rack where the server carrying the logical IT load is located, the power distribution unit port of the server rack, and the physical equipment for heat dissipation of the physical area where the server is located.

[0009] By adopting the above technical solutions, a multi-physics coupled digital twin integrating electrical, thermodynamic, and IT business logic is constructed to comprehensively and accurately model the data center power distribution system. The IT business logic model abstracts virtual machines and containers into logical IT loads and records the physical resource information they depend on, facilitating subsequent management and scheduling. A dynamic influence chain mapping relationship between logical IT loads and physical resources is established, providing a high-precision virtual simulation foundation for subsequent fault impact deduction and scheduling strategy pre-simulation, ensuring the accurate correlation between physical faults and business impacts, and avoiding blind fault location and resource matching.

[0010] Preferably, in response to the detected failure events of multiple physical resources, the step of determining the logical IT loads that have an abnormal impact associated with the failure events and forming a risk load set based on the mapping relationship between the digital twin and the impact chain specifically includes the following steps: Receive and aggregate multiple original alarm events that occur simultaneously or sequentially within a preset time window from the data center to generate a list of faulty physical resources; Based on each device identifier in the faulty physical resource list, find all associated primary logical IT load identifiers in the impact chain mapping relationship, and summarize them to obtain the primary associated load set; The list of physical resources containing concurrent failures and the set of primary associated loads are input into the digital twin for dynamic simulation to obtain the secondary logical IT loads whose performance is expected to drop below a preset survival threshold due to the concurrent failures, and the corresponding secondary logical IT load identifiers are summarized into a set of secondary impact loads. The primary associated load set and the secondary impact load set are merged to obtain the final risk load set.

[0011] By adopting the above technical solution, a list of faulty resources is generated by aggregating concurrent fault alarms, faulty equipment is accurately located, and directly related loads are quickly located by combining impact chain mapping. Then, indirectly affected loads are inferred through digital twins. Finally, directly related loads are merged to form a complete set of risk loads. This achieves full-range risk coverage of direct impact and indirect propagation under concurrent faults of multiple physical resources, constructs a comprehensive and complete risk view, avoids missing potential business interruption risk points, and provides an accurate basis for subsequent scheduling decisions.

[0012] Preferably, the dynamic deduction process specifically includes the following steps: In the electrical model of the digital twin, the normal power supply state and the fault power supply state after the concurrent fault is injected into the fault physical resource are simulated respectively. The server corresponding to the fault power distribution circuit that is newly added in the fault power supply state relative to the normal power supply state and whose electrical parameters exceed the preset safety limit is identified. Based on the electrical changes in the faulty power distribution circuit, the thermal power consumption changes of the corresponding supplied equipment are calculated, and the thermal power consumption changes are input into the thermodynamic model of the digital twin. In the thermodynamic model, temperature field simulation is performed based on the change in heat dissipation, and the server whose heat dissipation environment temperature exceeds the preset temperature threshold is predicted. Based on the influence chain mapping relationship, the logical IT load carried by the server identified in the electrical model and the thermodynamic model is determined, and the corresponding logical IT load identifiers are summarized into a direct physical risk load set. Based on a predefined business service dependency graph, starting with each logical IT load identifier in the direct physical risk load set, all directly and indirectly related logical IT load identifiers are traversed in reverse to form a business-related risk load set. The set of direct physical risk loads and the set of business-related risk loads are combined to obtain the set of secondary impact loads.

[0013] By adopting the above technical solutions, through a quantitative data chain from electrical parameters to thermal power consumption to temperature field, electrical faults are automatically and accurately transformed into specific risks of overheated servers and business loads, revealing cross-domain hidden faults; an automatic mapping and traversal rule for physical devices to business loads to dependent businesses is established, which can identify businesses indirectly damaged by dependencies without omission, and realize a panoramic insight from hardware faults to risks across the entire business chain.

[0014] Preferably, the step of establishing a multi-agent decision-making framework based on the digital twin, and inputting the risk load set into the multi-agent decision-making framework, specifically includes the following steps: A multi-agent decision-making framework is established based on the digital twin. The multi-agent decision-making framework includes management agents such as power distribution management agent, IT load management agent, and cooling management agent, as well as a business protection agent. The business protection agent is used to receive external input and coordinate other management agents. Each of the aforementioned management intelligent agents is configured so that each of the aforementioned management intelligent agents can obtain its corresponding real-time status data through the digital twin, wherein the real-time status data includes electrical data, computational data and thermodynamic data; The set of risky loads marked with business SLA levels is input into the business guardian agent. Each risky load in the set of risky loads has a risky load identifier and a business SLA level. The business SLA level is a predefined logical IT load business continuity assurance level.

[0015] By adopting the above technical solution and establishing a decision-making framework that includes business guardian agents and multiple management agents, each management agent can obtain real-time cross-domain data through digital twins. The risk load set marked with SLA level is input into the business guardian agent, thus constructing a command system with business SLA as the highest goal. This provides data support and architectural guarantee for differentiated weight adjustment and collaborative decision-making, and solves the problems of independent decision-making by multiple subsystems and disconnection from business needs in traditional scheduling. It helps to dynamically, securely and collaboratively schedule global resources in extreme failure scenarios.

[0016] Preferably, the dynamic adjustment of the decision objective weights of each agent in the multi-agent decision-making framework to obtain at least one candidate scheduling strategy, led by the business guardian agent, specifically includes the following steps: After receiving the set of risk loads, the business guardian intelligent agent calculates the business impact value of each risk load based on the preset SLA level weight, performance degradation percentage, and preset business importance coefficient. The performance degradation percentage is the degree of real-time performance degradation caused by the failure of the physical resources associated with the risk load. The business protection agent dynamically adjusts the objective function weights of the power distribution management agent, the IT load management agent, and the cooling management agent based on the business impact value; wherein, the risk load with a higher business impact value is assigned a higher weight to the relevant protection items in the objective function of each agent. Based on the weighted adjustments of the management agent, collaborative decision-making is performed to generate at least one candidate scheduling strategy that meets the business impact load requirements.

[0017] By adopting the above technical solution and introducing the quantitative indicator of business impact value, the priority, urgency and value of business are unified to measure the degree to which risk loads are affected by failures. In this way, the decision objectives of each agent can be dynamically reconstructed, making the decisions more in line with the protection needs of different risk loads. The business protection agent dynamically adjusts the decision objective weights of the power distribution management, IT load management and cooling management agents, forcing the independent decisions of these three agents to converge to a common goal, that is, to prioritize the continuity of the critical business under limited post-failure resources.

[0018] Preferably, simulating the execution of all the candidate scheduling strategies in the digital twin specifically includes the following steps: The real-time operating status of the data center power distribution system is synchronized to the digital twin for simulation environment initialization; The business SLA level of each risk load in the risk load set is used as an attribute label and associated with the corresponding business logic model node in the digital twin; For each candidate scheduling strategy, the multiphysics coupling model of the digital twin is driven to perform rapid pre-simulation to simulate the state of the power distribution system within a preset time period after the current candidate scheduling strategy is executed; During the rapid simulation, key performance indicators for each of the risk loads are continuously recorded. These key performance indicators include at least power supply stability indicators, IT load performance indicators, and thermal environment indicators.

[0019] By adopting the above technical solution, the digital twin simulation environment is initialized by synchronously and in real-time running status, and risk load SLA level labels are associated. The impact of different business continuity assurance levels can be reflected in the simulation. The multi-physics coupling model is driven to pre-simulate candidate strategies and record multi-dimensional key performance indicators, providing comprehensive and accurate simulation data support for subsequent security strategy selection. This facilitates the selection of security strategies that can maintain the performance of the risk load set above the preset security threshold, ensuring the objectivity and comprehensiveness of strategy effect evaluation, thereby ensuring the stable operation of the data center power distribution system under extreme failure scenarios.

[0020] Preferably, the rapid pre-simulation specifically includes the following steps: Based on the aforementioned impact chain mapping relationship, the critical physical device corresponding to each risk load is determined, and a pre-stored derating rule is matched for each critical physical device based on the associated service SLA level; wherein, the higher the service SLA level, the lower the actual maximum allowable current specified by the matched derating rule at the same temperature; Based on the updated load distribution of the candidate scheduling strategy, the system state is subjected to electrical model deduction to obtain the electrical state and equipment loss. The electrical state includes the initial current carrying parameters of the key physical equipment. Thermodynamic model deduction is performed based on the equipment loss to obtain the updated operating temperature of each key physical device, and the result is fed back to the electrical model. The corrected current-carrying parameters of each key physical device are then calculated in conjunction with the corresponding derating rules. Based on the modified current-carrying parameters, it is determined whether the preset convergence condition is met. If it is met, the system is determined to have reached a safe convergence steady state. If it is not met, an iterative strategy for rapid pre-simulation is determined according to the unmet convergence condition, until the safe convergence steady state is reached through iteration. Based on the final physical state under the secure convergence steady state, the IT business logic model is driven to simulate business performance and record the key performance indicator data.

[0021] By adopting the above technical solutions, the business level is directly transformed into the differentiated safety standards of the equipment in the simulation, which upgrades the safety verification from ensuring that the equipment does not fail to guarantee that the specified business is not interrupted. Through closed-loop iterative calculation of electrical state and temperature field, the transient and steady state of the system after the strategy is executed can be accurately simulated, and the risks of instantaneous overload and heat accumulation that are ignored by traditional static simulation can be discovered, thus realizing an order-of-magnitude improvement in reliability verification.

[0022] Preferably, the step of selecting a set of all security strategies capable of maintaining the performance of the risk load set above a preset security threshold from the candidate scheduling strategies based on the obtained simulation results specifically includes the following steps: Set a security threshold for the associated key performance indicator for each of the risk loads in the risk load set, the performance security threshold being associated with the service SLA level of the risk load; Based on the pre-simulation results corresponding to each candidate scheduling strategy, security verification is performed on each of the risk loads in the risk load set one by one; When the data of the key performance indicators of all the risk loads remain within their respective security thresholds for a preset period of time in the future, the corresponding candidate scheduling strategy is determined to be a security strategy. All the security policies mentioned are summarized to form a security policy set.

[0023] By adopting the above technical solution, differentiated safety thresholds related to business SLA levels are set for risky loads. The pre-simulation results of candidate strategies are verified load by load, and only strategies that meet the performance standards of all loads are retained to form a safety set. This achieves a strong binding between business priority, safety standards, and strategy selection, ensuring that the selected strategies meet the continuity requirements of different levels of business. It also prevents insufficient core business protection or resource waste caused by uniform thresholds, thereby ensuring the performance of the data center power distribution system for risky loads of different business SLA levels under extreme failure scenarios, realizing dynamic, safe, and collaborative decision-making, and avoiding business interruption.

[0024] Secondly, this application provides a collaborative scheduling device for a data center power distribution system, which adopts the following technical solution: A collaborative scheduling device for a data center power distribution system includes the following modules: The model building and mapping module is used to build a digital twin containing a multiphysics coupling model and establish an influence chain mapping relationship between physical resources and logical IT load identifiers. A multi-fault identification module is used to respond to multiple detected fault events of the physical resources, and to determine the logical IT loads that have abnormal impact associated with the fault events and form a risk load set based on the mapping relationship between the digital twin and the impact chain. The collaborative strategy generation module is used to establish a multi-agent decision-making framework based on the digital twin, input the risk load set into the multi-agent decision-making framework, dynamically adjust the decision target weights of each agent in the multi-agent decision-making framework, and obtain at least one candidate scheduling strategy. The policy security verification module is used to simulate the execution of all the candidate scheduling policies in the digital twin, and based on the simulation results, to select all the security policy sets that can maintain the performance of the risk load set above a preset security threshold from the candidate scheduling policies. The security instruction execution module is used to convert the optimal security policy in the security policy set into control instructions and issue them for execution.

[0025] By adopting the above technical solutions, modular design enables independent encapsulation and collaborative linkage of various functions. Data communication between modules such as model building, fault identification, policy generation, security verification, and instruction execution ensures the feasibility of the entire fault-tolerant scheduling scheme. At the same time, it improves the maintainability and scalability of the system and adapts to the power distribution and scheduling needs of data centers of different sizes.

[0026] In summary, this application includes at least the following beneficial effects: (1) This application achieves full-range identification of the impact of business under concurrent failure of multiple physical resources by multi-physics field coupling modeling and fault propagation simulation of digital twin, combined with load correlation analysis of impact chain mapping. It solves the defects of traditional scheduling that cannot cope with multiple concurrent failures and incomplete risk coverage, and significantly reduces the potential risk of business interruption. (2) This application enables multi-agent collaborative decision-making to prioritize core businesses with high SLA levels by quantitatively calculating the impact of business and dynamically adjusting the weights of business guardians, and dynamically adjusts the decision weights to achieve dynamic global resource scheduling of electrical, thermal and computing resources under multiple faults, ensuring the performance stability of core businesses under fault scenarios and improving the overall business continuity level. (3) This application deeply couples multi-agent decision-making with the physical simulation of digital twins, so that digital twins serve as a security verification layer before decision-making. Combined with the strategy screening of differentiated security thresholds, it can make dynamic, safe and collaborative decisions on extreme fault scenarios under the hard constraint of zero business interruption, and provide reliable guarantee for the stable operation of business under complex faults. Attached Figure Description

[0027] Figure 1 This is a flowchart of the collaborative scheduling method for the data center power distribution system in this application; Figure 2 This is an architecture diagram of the collaborative scheduling device for the data center power distribution system in this application. Detailed Implementation

[0028] This application provides a collaborative scheduling method and apparatus for a data center power distribution system. To make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below.

[0029] The following describes in further detail an embodiment of a collaborative scheduling method for a data center power distribution system according to the present application, with reference to the accompanying drawings.

[0030] This embodiment presents a collaborative scheduling method for a data center power distribution system, the process of which is as follows: Figure 1 As shown, the specific steps include the following: S1. Construct a digital twin containing a multiphysics coupling model and establish an influence chain mapping relationship between physical resources and logical IT load identifiers, specifically including the following steps: S11. Perform multi-physics coupling basic modeling to obtain a digital twin. The multi-physics coupling model includes an electrical model, a thermodynamic model, and an IT business logic model.

[0031] In one specific implementation method, the process of establishing a multiphysics coupling model is as follows: Establish an electrical model for the data center to simulate the distribution of current, voltage, and power in the power distribution system and its transient characteristics under fault conditions; Establish a thermodynamic model of the data center to simulate the airflow organization and temperature field distribution within the computer room; Establish an IT business logic model for the data center to abstract virtual machines and containers running within the data center into logical IT loads with unique identifiers, and record the physical resource information that the logical IT loads depend on for operation, including the resource configuration and performance metrics required for operation.

[0032] Specifically, the workload in this embodiment includes virtual machines and containerized workloads.

[0033] Virtual machine workloads abstract physical servers into multiple independent virtual machines using virtualization technologies (such as VMware and KVM). Each virtual machine is a logical workload that can be migrated between different physical servers. For example, a physical server can host three virtual machines, running a web service, a database service, and a caching service respectively, with each virtual machine being an independent logical IT workload.

[0034] Containerized workloads (Docker / Kubernetes Pods) are more lightweight logical workloads. Containers package an application and its dependencies (such as libraries, configuration files, and runtime environments) into standardized units. For example, an e-commerce platform's order processing service is packaged as a Docker container, which is a logical workload.

[0035] S12. In the digital twin, each logical IT load identifier is associated with its corresponding physical resource information to form a dynamic influence chain mapping relationship.

[0036] Physical resource information includes the server rack where the server carrying the logical IT load is located, the power distribution unit port (PDU port) of the rack, the UPS circuit that powers the PDU, and the physical equipment that dissipates heat from the physical area where the server is located.

[0037] In one specific implementation, by querying the configuration management database and the infrastructure management system, a dynamic association is established and maintained between logical IT load identifiers (such as virtual machine identifiers), the power distribution unit port identifiers of the cabinets supplying power to them, and the air conditioning equipment identifiers that dissipate heat from the physical area where they reside. This forms a mapping relationship that allows for the deduction of the impact path of physical resource failures on logical IT loads.

[0038] In one specific implementation, the influence chain mapping relationship is constructed into a dynamic association graph. The nodes in the dynamic association graph represent physical resources or logical IT loads, and the edges represent the dependencies between them. When a failure is detected in a node of a physical resource, all logical IT load nodes affected by it can be automatically deduced by traversing the association graph.

[0039] The digital twin implemented here can continuously and dynamically calibrate the electrical model, thermodynamic model, and IT business logic model by collecting real-time data from the physical equipment operating parameters, environmental sensor data, and IT load status data of the data center, so as to ensure the simulation accuracy of the digital twin.

[0040] S2. In response to multiple detected physical resource failure events, based on the mapping relationship between digital twins and impact chains, identify the logical IT loads that have an abnormal impact associated with the failure events and form a risk load set, specifically including the following steps: S21. Receive and aggregate multiple original alarm events that occur simultaneously or successively within a preset time window from the data center, and generate a list of faulty physical resources.

[0041] In a specific implementation, the original alarm event could be, for example, transformer A over-temperature alarm, distribution cabinet B bus voltage abnormality, air conditioner C compressor failure, etc.

[0042] The power distribution system preprocesses these events, such as deduplication and aggregation, and calls the digital twin to load the real-time operating parameters of each faulty physical resource, including current, voltage, temperature, and fault codes, forming a list of faulty physical resources to be analyzed, including the fault type.

[0043] S22. Based on each device identifier in the faulty physical resource list, find all associated primary logical IT load identifiers in the impact chain mapping relationship, and summarize them to obtain the primary associated load set.

[0044] The previously established influence chain mapping relationship is invoked to perform the first round of direct association query, which is a static, topology-based retrieval.

[0045] Specifically, based on each device identifier in the list of faulty physical resources, all logical IT load identifiers directly associated with it are searched in the mapping relationship as primary logical IT load identifiers.

[0046] In one specific implementation, for example, the input circuit identifier of power distribution cabinet B is mapped to the influence chain relationship, and the output is the identifier of server groups A1 and A2 that are directly powered by that circuit; Alternatively, input the identifier of air conditioner C to the influence chain mapping relationship, and output the identifier of the entire rack-Zone-1 that is cooled by air conditioner C.

[0047] S23. The list of faulty physical resources containing concurrent faults and the set of primary associated loads are used as initial conditions and input into the digital twin for dynamic simulation to simulate the chain reactions and coupling effects caused in multiple physical fields such as electricity, heat, and computing under multiple concurrent fault scenarios.

[0048] Because the impact of the fault can propagate, the results of the initial association in step S22 are incomplete. Therefore, the digital twin still needs to be extrapolated to identify servers outside the initial association load set. In a specific feasible implementation, the dynamic extrapolation process is as follows: S231. The simulation engine obtains input from the list of faulty physical resources and the initial associated load set, and sets them as the initial perturbation conditions for the digital twin: In the electrical model of the digital twin, physical equipment such as transformers and distribution cabinets involved in the fault physical resource list are included.

[0049] S232. The simulation process is carried out in a short time domain, for example, to simulate the physical time of the next 30 seconds to 5 minutes. The digital twin iterates according to a preset time step, such as 10 milliseconds.

[0050] Dynamic simulations include simulations based on electrical fields, thermodynamic fields, and IT business logic models. Electrical and thermodynamic physical simulations are used to predict secondary physical device anomalies caused by faults; while IT business logic simulations are used to predict upper-layer business service interruptions due to physical device failures, based on predefined dependencies between business components. The specific process is as follows: S233. Fault injection is used to simulate the electrical field: In the electrical model of the digital twin, the normal power supply state and the fault power supply state after the injection of concurrent faults are simulated respectively, and the new voltage / current distribution and equipment load rate under the fault power supply state are obtained.

[0051] Specifically, the simulation includes the following steps: Normal power supply state simulation: Simulate the normal power supply network state after removing all faulty devices, and record the current value, voltage value, or load rate of each distribution circuit as the baseline value; Concurrent power supply fault state simulation: Inject all concurrent fault states into the same electrical model to simulate the power supply network state after the fault, and record the current value, voltage value, or load rate of each distribution circuit as the fault value.

[0052] Scan all power distribution circuits, calculate the difference between the baseline value and the fault value for each circuit, and identify the server corresponding to a faulty power distribution circuit that is newly added under faulty power supply conditions compared to normal power supply conditions, and whose electrical parameters exceed preset safety limits. In a specific implementation, a faulty power distribution circuit is defined as a power distribution circuit whose fault value exceeds the circuit's maximum safe carrying capacity limit (e.g., 100% of the rated current), and the absolute value of the difference exceeds a preset sensitivity threshold (e.g., 5% of the rated current).

[0053] The physical state of these servers is marked as "power quality deterioration," which is the first direct impact of the electric field on the computing field.

[0054] Coupled deduction of thermal field and electric field: Based on the electrical changes of the faulty power distribution circuit obtained from the above electrical field deduction, calculate the corresponding changes in heat power consumption of the supplied equipment.

[0055] For each identified faulty power distribution circuit, the change in its heat dissipation is calculated by summing the products of the load increment of each device on the circuit, the rated power of the device, and its energy-to-heat conversion efficiency coefficient. The load increment is the difference between the electric field and the computational field coupling. For IT computing devices, such as servers and network switches, most of the electrical energy consumed is converted into heat energy, and its energy-to-heat conversion efficiency coefficient can be set to a first predetermined value, such as 0.95 to 1.0. For power distribution equipment, such as transformers and distribution units, the main function is the transmission and transformation of electrical energy. Most of the electrical energy is delivered to downstream loads, so its heat dissipation mainly comes from internal losses and is not converted into heat energy at the power distribution equipment. Therefore, its conversion efficiency coefficient is a second predetermined value determined based on the energy efficiency of the equipment. This value is significantly less than 1, such as 0.01 to 0.1. For example, for high-efficiency transformers, the conversion efficiency coefficient may be between 0.01 and 0.05, i.e., a loss rate of 1% to 5%.

[0056] The changes in heat dissipation of these loops and their physical location information (such as which computer room area they are located in) are used as boundary conditions for the new heat source and input into the thermodynamic model for temperature field simulation.

[0057] Coupled extrapolation from thermal field to computational field: In the thermodynamic model under thermal-electric coupling, temperature field simulation is performed based on the change in heat dissipation. Servers whose heat dissipation environment temperature exceeds the preset temperature threshold are predicted. That is, the air inlet temperature of all servers is monitored, and servers whose air inlet temperature is expected to exceed the temperature threshold for safe operation are marked as "facing overheating risk".

[0058] In one specific implementation, thermodynamic field simulation is also achieved by adjusting the cooling capacity parameters of the refrigeration equipment. Specifically, when a refrigeration equipment malfunctions as indicated in the fault physical resource list, its cooling capacity parameter is set to zero or to a failure value. Simultaneously, thermal power change data of each cabinet and server calculated from the electrical field simulation are received and used as input to the thermodynamic model for temperature field simulation.

[0059] The reason for considering the status of the cooling equipment here is that if only the change in heat power is considered, if the server group S stops and the heat generation decreases, the simulation may incorrectly predict whether the temperature in the current computer room area will drop or remain unchanged. If the failure of the cooling equipment is also considered, such as the air conditioner failing and the cooling capacity returning to zero, even if the heat generation of the server group S decreases, the heat generated by other equipment still running in the area will not be dissipated, and the simulation can correctly predict that the temperature in the area will rise sharply.

[0060] The computing field is integrated based on physical state marking, and all servers marked as "power quality deterioration" and "facing overheating risk" are integrated to form a comprehensive risk server set.

[0061] Coupled extrapolation from computing field to computing field: The content of IT business logic extrapolation is to simulate the performance changes of IT load using the results of electrical field and thermodynamic field as input. Based on the influence chain mapping relationship, the logical IT load carried by the servers in the comprehensive risk server set obtained from all previous extrapolations is determined, and the corresponding logical IT load identifiers are summarized into a direct physical risk load set.

[0062] Based on the analysis of the failure impact propagation based on the predefined business service dependency graph, starting from each logical IT load identifier in the direct physical risk load set, all logical IT load identifiers that directly or indirectly depend on the starting point are traversed in reverse to form the business-related risk load set. For example, when the simulation detects that the database service DB1 is interrupted due to a power outage on the server it is hosting, it automatically infers that all front-end application services that depend on DB1 will also be unavailable, even if the server hosting the front-end application service is healthy.

[0063] The specific process by which the impact of a business dependency failure propagates is as follows: For each business node in the business service dependency graph, if there is a direct connection between two business nodes, they are considered to have a direct dependency / association relationship.

[0064] Taking each load in the direct physical risk load set as the source, and combining it with the business service dependency graph for mapping, calculate the single-hop correlation coefficient between the current source business node and each business node directly related to it in the graph; The single-hop correlation coefficient between business nodes is determined based on the resource topology correlation factor and the physical resource risk factor, and converges through a preset convergence rule to iteratively determine all business nodes affected by its direct or indirect risk transmission. Among them, the resource topology association factor is calculated based on the influence chain mapping relationship and the proportion of underlying physical resources shared by the two business nodes; the physical resource risk factor is calculated based on the safety margin preset for each physical resource in the digital twin. The lower the safety margin, the higher the corresponding physical resource risk factor.

[0065] The convergence rule is, The single-hop correlation coefficient is compared with a preset first threshold to determine whether the direct dependency is sufficient to constitute significant risk transmission. For indirect dependencies, the cumulative transmission strength is obtained by calculating the product of the correlation coefficients of each edge on the path from the risk source node to the target node; ;in, The cumulative correlation coefficient is given by m, where m is the number of hops in the dependency path. This is the single-hop correlation coefficient.

[0066] The single-hop correlation coefficient is compared with a preset second threshold, while the number of path hops is limited to a preset hop count threshold. This serves as the basis for determining indirect correlation, thereby determining whether to include the target node in the business-related risk load set.

[0067] In a specific implementation, it is assumed that for any two directly related business nodes in the business service dependency graph, they are denoted as node P and node Q (P depends on the service of Q). Calculate the single-hop correlation coefficient between node P and node Q. ; in, The single-hop correlation coefficient between node P and node Q. Let be the resource dependency weight of node P on node Q. Let be the physical resource criticality weight of node Q.

[0068] Specifically, ; in, This refers to the amount of physical resources shared by nodes P and Q, such as power distribution circuits, cooling equipment, and server racks. This represents the total number of physical resources that node P depends on. The influence chain mapping relationship is directly extracted from the topology, eliminating the need for further calculations.

[0069] Physical resource criticality weight of node Q This is the inverse average of the safety margins of all physical resources that node Q depends on; Safety margin = (Safety limit - Normal operating value) / Safety limit; Both the safety limit and the normal operating value have been configured in the aforementioned step S233. The lower the safety margin, the higher the criticality weight, which means that the physical resources of the node are closer to the failure threshold and the risk transmission is stronger. Right now, ,in, For safety margin, n is the number of physical resources that node Q depends on.

[0070] In a specific implementation, a simple business chain in the business service dependency graph is A directly associated with B, B directly associated with C, A depends on B's service, and B depends on C's service. It is assumed that business node A has been identified as a direct physical risk load. Assume the convergence rule parameters are: First threshold (one-sided correlation threshold): T1=0.3; Second threshold (cumulative conduction strength threshold): T2=0.1; Maximum path hops: N_max=3; Starting with business node A, the correlation between A and B is compared with a first threshold. If the single-hop correlation between A and B is not greater than the first threshold, the correlation of the direct dependency edge has not reached the threshold for significant risk transmission. Therefore, node B is not included in the business-related risk load set in this round of analysis, nor does it continue to propagate to C. If the single-hop correlation between A and B is greater than the first threshold, node B is included in the business-related risk load set, and the analysis continues from B to downstream node C.

[0071] Calculate the single-hop correlation coefficient between B and its directly related node C. Then, combine this with the single-hop correlation coefficient between A and B to calculate the cumulative correlation coefficient from A to B to C. That is, the cumulative correlation coefficient = the single-hop coefficient between A and B × the single-hop coefficient between B and C. If the cumulative correlation coefficient is not greater than the second threshold, node C is not included in the business-related risk load set, and propagation to downstream nodes of C is terminated. If the cumulative correlation coefficient is greater than the second threshold, node C is included in the business-related risk load set. At the same time, considering the maximum hop count limit, it is determined whether to continue downstream analysis.

[0072] Ultimately, all nodes that directly and indirectly depend on this starting point are obtained, making the corresponding logical IT loads constitute a set of business-related risk loads.

[0073] S24. Collect all secondary logical IT loads formed by the direct physical risk load set and the business-related risk load set during the simulation process, and summarize the corresponding secondary logical IT load identifiers into a secondary impact load set. This set can accurately reflect those businesses that are not directly affected by the initial failure, but are indirectly damaged by the chain effect caused by the coupling of multiple failures.

[0074] S25. Summarize and merge the primary associated load set and the secondary impact load set to obtain the final risk load set. Mark the business SLA level and associated risk load identifier for each risk load. The business SLA level is a predefined logical IT load business continuity assurance level.

[0075] The above-described fault propagation analysis captures the comprehensive impact of faults on IT load, both direct and indirect, immediate and potential, after they propagate through multiple physical fields such as electricity and heat.

[0076] S3. Establish a multi-agent decision-making framework based on a digital twin, input the risk load set into the multi-agent decision-making framework, dynamically adjust the decision objective weights of each agent in the multi-agent decision-making framework, and obtain at least one candidate scheduling strategy. The specific steps include the following: S31. Establish a multi-agent decision-making framework based on digital twins. The multi-agent decision-making framework includes management agents such as power distribution management agent, IT load management agent, and cooling management agent, as well as business protection agents. The business protection agent is used to receive external input and coordinate other management agents.

[0077] The interface is used to enable each agent to obtain real-time status data from the digital twin.

[0078] S32. Configure interfaces for data interaction between each management agent and the digital twin, so that each management agent can obtain its corresponding real-time status data through the digital twin and send the action instructions generated by each agent's decision to the digital twin for simulated execution.

[0079] Real-time status data includes electrical data, computational data, and thermodynamic data. Specifically, the power distribution management agent acquires circuit load rate, UPS status, etc.; the IT load management agent acquires logical IT load resource utilization rate, etc.; and the cooling management agent acquires area temperature, air conditioning operating power, etc.

[0080] S33. Input the set of risk loads marked with the business SLA level into the business guardian agent. Each risk load in the risk load set has a risk load identifier and a business SLA level.

[0081] The business SLA level is a predefined business continuity assurance level for logical IT loads. That is, it is a priority label that is predefined for logical IT loads and quantifies their business continuity requirements. Its data comes from the service catalog or SLA agreement and may include one or more dimensions such as availability requirements, recovery time target (RTO), and business criticality.

[0082] In a specific feasible implementation, the business SLA level is classified according to the following criteria: SLA-A: SLA=100.00%, critical core business with zero interruption, requiring continuous and uninterrupted operation in fault scenarios, such as financial transactions, core payments, and user authentication services; SLA-B: SLA ≥ 99.99%, critical business operations with an annual downtime of approximately 0.88 hours, allowing for short-term fluctuations at the millisecond level but requiring rapid recovery, such as online transaction inquiries and real-time order processing services.

[0083] SLA-C: SLA≥99.90%, general business with an annual downtime of approximately 8.76 hours, acceptable for short-term interruptions but must be restored within a preset time, such as product browsing and non-real-time data analysis services; SLA-D: SLA≥99.00%, with approximately 87.6 hours of back-end service interruption per year, and can tolerate longer interruptions, such as data backup and log storage services.

[0084] S34. After receiving the set of risk loads, the business protection agent calculates the business impact value of each risk load based on the preset SLA level weight, performance degradation percentage, and preset business importance coefficient.

[0085] SLA level weights are fixed numerical weights that are pre-set in the business configuration database for different business SLA levels; they are values ​​between 0 and 1.

[0086] In one feasible implementation, the SLA rating weight can be directly derived from the business's Service Level Agreement (SLA) contract. This weight is directly assigned based on the availability requirements (such as 99.9%) defined in the contract.

[0087] For example, the table below shows the correspondence between SLA levels and weights: Table 1. Correspondence between SLA levels and weights

[0088] The percentage performance degradation is a dynamically calculated value that measures the real-time performance impact of a physical resource failure event associated with a risky load on a specific business function. Specifically, it is calculated by comparing key performance indicators of the logical IT load before and after the failure, and then comparing these indicators with normal baseline values ​​simulated by a digital twin.

[0089] For virtual machines / servers, key performance indicators include: CPU utilization, memory utilization, network I / O, and response time.

[0090] In a specific implementation, assuming that the average response time of a certain payment service is 100ms under normal conditions, and after a failure, the average response time of the current service rises to 500ms; then the performance degradation percentage = (500ms-100ms) / 100ms = 400%.

[0091] It should be noted that in this embodiment, since the performance degradation percentage exceeds 100%, it indicates a very serious performance degradation. Therefore, in practical applications, an upper limit (such as 200%) can be set to prevent a single indicator from excessively affecting the overall calculation.

[0092] The business importance coefficient is a preset, fixed coefficient. It is a fixed coefficient set in advance based on the revenue contribution or user impact of the business carried by the logical IT load. In other words, it is the strategic importance or value of the business to the enterprise, in addition to the business SLA level.

[0093] It is usually determined through joint consultation between the business and IT departments of an enterprise, and is a quantitative value converted from a qualitative assessment.

[0094] For example, in a specific implementation, the following business importance coefficient correspondence table is provided: Table 2. Business Importance Coefficient Correspondence Table

[0095] In one specific implementation method, the Business Impact Value (BID) is calculated using the following formula: BID = SLA level weight × performance degradation percentage × business importance coefficient.

[0096] Among them, SLA-A=100%, which has the highest business importance coefficient.

[0097] S35. The business protection agent dynamically adjusts the objective function weights of the power distribution management agent, the IT load management agent, and the cooling management agent based on the business impact value.

[0098] For example, if a payment transaction with a BID value of 95 and a log analysis transaction with a BID value of 20 are affected at the same time, the primary goal of the system is to ensure the stable operation of the payment transaction, and the secondary goal is to restore the log analysis transaction.

[0099] The business protection agent determines the global objective priority based on the aforementioned business impact values ​​and dynamically assigns weights to the objective functions of other management agents through a weight allocation function. Specifically, risk loads with higher business impact values ​​receive higher weights in the relevant protection items of each agent's objective function. For example, the weights of items related to ensuring power supply stability for high-business-impact loads are increased in the power distribution management agent's objective function; the weights of items related to ensuring computing resources for high-business-impact loads are increased in the IT load management agent's objective function; and the weights of items related to ensuring heat dissipation requirements for high-business-impact loads are increased in the cooling management agent's objective function.

[0100] In a specific implementation, the objective function weights of each executing agent are dynamically adjusted, including: The business protection intelligence has at least one set of weighted policy mapping functions pre-built. The function takes the business impact value as input. The higher the business impact value, the greater the weight coefficient assigned to the sub-objectives that are directly related to ensuring the continuity of the business impact load in the output weight vector.

[0101] The current risk load set and its business impact value are input into the selected weight strategy mapping function to calculate and output a set of weight vectors. Each weight vector corresponds to a management agent, including power distribution, IT load and cooling. Each element in the vector corresponds to the weight coefficient of a certain sub-objective item in the objective function of the agent. The calculated weight vectors are sent to the corresponding management agents respectively, and each management agent uses the received weight coefficients to reconstruct its objective function.

[0102] Taking the objective function of a power distribution management agent as an example, the objective function may include three sub-objectives: load balancing, power supply stability, and energy efficiency. For a payment transaction with a BID value of 95 and a log analysis transaction with a BID value of 20, calculate the weighting factor of their BID load. ; in, For BID, a high-risk load, For BID with a lower risk load, this embodiment ; Objective function of power distribution management agent ; in, For load balancing, For power supply stability, For energy efficiency.

[0103] when When the load is very high, it indicates that the core business is in jeopardy, and the weight of power supply stability should be greatly increased, even if the weight of load balancing can be temporarily sacrificed.

[0104] for The calculation is as follows: ;in The base value is k, where k is the amplification factor.

[0105] For example, ;at the same time, , Setting it to zero means that the power distribution agent is almost solely concerned with power supply stability in this decision.

[0106] The baseline value is obtained by analyzing the normal contribution ratio of each sub-objective to the overall health of the system under historical normal operation data; the amplification factor k represents the degree of aggression of the system in emergency fault conditions when it shifts from a conventional multi-objective optimization mode to a high-priority business assurance mode.

[0107] It can be directly correlated based on the SLA level. For example, when the highest SLA is 100.00%, k=0.95 is used for full protection; when the highest SLA is 99.9%, k=0.6 is used for moderate protection.

[0108] S36. Based on the weighted management agents, collaborative decisions are made. The weighted power distribution management, IT load management, and cooling management agents make interactive decisions based on the acquired real-time status data, and iteratively generate at least one candidate scheduling strategy that meets the load requirements of business impact.

[0109] In a specific feasible implementation, the IT load management agent proposes a migration scheduling scheme that is most conducive to ensuring high BID load, such as immediately migrating all virtual machines of core business to the currently most idle Server_X and Server_Y.

[0110] The power distribution management agent and the cooling management agent evaluate the migration scheduling scheme based on the real-time status of the digital twin and feed it back to the IT load management agent. The power distribution management agent calculates the changes in the load of the power supply circuit when the scheme is executed, and the cooling management agent calculates the new heat generated by the scheme and compares it with the current cooling capacity.

[0111] After receiving feedback, the IT load management agent makes corrections, such as migrating most of the virtual machines for core business to Server_X; moving virtual machines for secondary processing business out of Server_X to free up resources and reduce heat load; and moving another non-core business away from the target loop to alleviate power distribution pressure.

[0112] The power distribution and cooling agents evaluate the revised scheme again based on their own objective functions until the sub-schemes output by each agent can jointly satisfy all the physical constraints of the feedback. The sub-schemes of the agents that finally reach a consensus are combined to generate at least one candidate scheduling strategy.

[0113] S4. Simulate the execution of all candidate scheduling strategies in the digital twin. Based on the simulation results, select a set of security strategies that can maintain the performance of the risky load set above a preset security threshold. This includes the following steps: S41. Through the data bus, synchronize the real-time operating status of the data center power distribution system at the moment of decision-making to the digital twin for simulation environment initialization. The real-time operating status includes the fault status of electrical, thermodynamic, and IT load status, the load of each circuit, the area temperature, and the distribution of IT load.

[0114] In one specific implementation, the electrical status includes the switching status of all circuit breakers and ATS, and the real-time current, voltage, and power readings of each circuit. Thermodynamic conditions include temperature sensor readings for each zone, operating modes, set temperatures, and fan speeds for all precision air conditioners; IT load status includes the resource utilization (CPU, memory, I / O) and distribution of all physical servers and their virtual machines.

[0115] S42. The system reads the inherent business SLA level of each risk load in the risk load set from the Business Configuration Management Database (CMDB), and binds the business SLA level as an immutable attribute label to the corresponding business logic model node in the digital twin, so that the simulation engine can identify and track each key business entity.

[0116] S43. For each candidate scheduling strategy, drive the multi-physics coupling model of the digital twin to perform rapid pre-simulation, simulating the state of the power distribution system within a preset time period after the current candidate scheduling strategy is executed.

[0117] Specifically, the multiphysics coupling model is driven to advance the simulation with small time steps (such as 10-100 milliseconds), and the dynamic behavior of the system after executing the current candidate scheduling strategy is deduced within a preset time period according to the preset time step.

[0118] The rapid simulation includes an electro-thermal bidirectional coupling iterative process driven by the business SLA level to solve the steady state that the system may reach after the execution of candidate scheduling strategies. This process is performed on a set of key physical devices whose states are affected by the execution of candidate scheduling strategies.

[0119] S431. Based on the influence chain mapping relationship, determine the key physical equipment corresponding to each risk load, and match a pre-stored derating rule for each key physical equipment based on the associated business SLA level; wherein, the higher the business SLA level, the lower the actual maximum allowable current specified by the matched derating rule at the same temperature.

[0120] To achieve differentiated physical security assurance based on SLA levels, the system pre-stores multiple sets of equipment derating rules.

[0121] First derating rule set: Corresponds to standard SLA requirements (not greater than 99.9%). This rule set adopts the standard commercial derating curves published by the equipment manufacturer, aiming for higher resource utilization while ensuring equipment lifespan.

[0122] The second derating rule set corresponds to extremely high SLA requirements (not less than 99.99%). This rule set uses a reinforcement curve with a larger derating range recommended by the manufacturer for high-reliability applications. In other implementations, if no recommended reinforcement curve is available, an additional safety factor (e.g., 0.8) can be preset to conservatively adjust the standard commercial derating curve.

[0123] S432. Based on the candidate scheduling strategy, update the system state after the load distribution, perform electrical model simulation to obtain the electrical state and equipment losses. Equipment losses include power loss and heat generation. Specifically, the electrical state simulation is to obtain the initial temperature state of each key physical device through the real-time operating state obtained in step S41, thereby determining the initial current-carrying parameters, which include the operating current of the device.

[0124] S433. Based on the heat loss power of the equipment and the heat generation power of the IT equipment, a thermodynamic model is used to deduce the updated operating temperature of each key physical device and feed it back to the electrical model. Combined with the corresponding rules, the corrected current carrying parameters of each key physical device are calculated. The corrected current carrying parameters are the actual maximum allowable current of the device at the current updated operating temperature.

[0125] S434. Based on the corrected current-carrying parameters, determine whether the preset convergence conditions are met. If they are met, the system is determined to have reached a safe convergence steady state. If not, determine the rapid pre-simulation iterative strategy based on the unmet convergence conditions until the iteration reaches a safe convergence steady state.

[0126] In one specific feasible implementation, the convergence conditions include the electro-thermal self-consistency condition, the electrical safety condition, and the SLA safety condition, wherein... The electro-thermal self-consistency is achieved by re-executing steps S432-S433 based on the calculated actual maximum allowable current. The variation in the result is less than a preset threshold, such as a temperature change of less than 0.5℃ or a loss change of less than 1%.

[0127] Electrical safety conditions: The actual operating current of each critical physical device does not exceed its own actual maximum permissible current determined by the current temperature. SLA safety conditions: All key performance indicators meet the thresholds required by their associated SLA level.

[0128] If all convergence conditions are met, the system is determined to have reached a safe convergence steady state. If the electrical safety condition or the SLA safety condition is not met, the candidate scheduling strategy is immediately determined to have failed and the iteration is terminated. If only the electro-thermal self-consistency condition is not met, the current-carrying parameters are updated with the calculated actual maximum allowable current, and the process returns to S432 for the next iteration. If the iteration count exceeds the preset upper limit and the condition is still not met, the candidate scheduling strategy is determined to be unsafe.

[0129] S435. Based on the final physical state under safe convergence steady state, drive the IT business logic model to simulate business performance and record key performance indicator data.

[0130] S44. During the rapid rehearsal process (e.g., within the next 3 minutes), continuously record the key performance indicators of each risk load in the form of time series data. The key performance indicators include at least power supply stability indicators, IT load performance indicators, and thermal environment indicators.

[0131] Among them, the power supply stability index corresponds to the voltage fluctuation range and frequency deviation of the circuit; IT performance metrics such as CPU utilization, memory usage, and response time for this workload; Thermal environment indicators include the air intake temperature of the server where the load is located and the average temperature of the corresponding area.

[0132] S45. Set a security threshold for the associated key performance indicators for each risk load in the risk load set. The performance security threshold is related to the business SLA level of the risk load. The higher the business SLA level, the more stringent the corresponding security threshold.

[0133] In one specific implementation, the system pre-defines a mapping rule between SLAs and safety thresholds. Specifically: For services with an SLA of 99.99%, the response time safety threshold is set to ≤100 milliseconds, the power supply voltage deviation threshold is set to ±2%, and the air inlet temperature threshold is set to ≤23℃.

[0134] For services with an SLA of 99.90%, the response time safety threshold can be relaxed to ≤500 milliseconds, and the voltage deviation threshold is set to ±5%.

[0135] S46. Based on the pre-simulation results corresponding to each candidate scheduling strategy, perform security verification on each risk load in the risk load set one by one.

[0136] The verification rule is as follows: traverse each risky load under the candidate scheduling strategy, check all the recorded key performance indicators of the load, and determine the worst value on the entire simulation timeline to see if all of these worst values ​​do not exceed the safety threshold corresponding to the load's own SLA level.

[0137] When the key performance indicators of all risky loads remain within their respective security thresholds for a preset period of time, the corresponding candidate scheduling policy is determined to be a security policy, and the performance indicators of the corresponding key business set will not be triggered to interrupt.

[0138] If any critical performance metric of any risky load exceeds its safety threshold, the strategy is immediately rejected. For example, a strategy might perfectly solve the power supply problem, but if it causes a server serving a core business to overheat (e.g., exceed 23°C), it will be eliminated.

[0139] All security policies are aggregated to form a security policy set. If this set is empty, an alarm is triggered and fed back to the decision-making module, requesting the generation of new candidate policies.

[0140] S5. Convert the optimal security policy in the security policy set into control commands and issue them for execution.

[0141] In a specific feasible implementation, the criteria for optimal security strategy prioritize maximizing business value and minimizing system disturbance.

[0142] Calculate a global optimization score for each security policy in the set. This score is a function of several weighted factors, including a business assurance factor, a resource efficiency factor, and a system stability factor.

[0143] The business assurance factor represents the change in the sum of the BIDs of all risky loads after the strategy is executed. The strategy with the largest decrease in the sum of the BIDs is prioritized. In general, the business assurance factor has the highest weight. The resource efficiency factor represents the total resource scheduling cost involved in the strategy, such as the total number of virtual machines migrated, the total number of power supply switching times, and the estimated additional cooling power consumption. The strategy with the lower resource cost is prioritized. The system stability factor represents the load imbalance of the entire power distribution system or the average load rate of key equipment after the strategy is executed. The strategy that can bring more stable operation is prioritized.

[0144] All security policies are sorted in descending order based on their overall scores. The policy with the highest score is automatically selected as the optimal security policy. In case of a tie, the policy with the higher business protection factor is selected first.

[0145] The generated optimal security policy is a cross-system cooperative sequence of operations, such as closing circuit breaker A, migrating virtual machine B to host C, and adjusting the air conditioner D valve to 50%. The system will decompose the optimal security policy into independent, sequential atomic instructions.

[0146] A predefined device driver adapter layer is used to convert atomic instructions into control commands that conform to different subsystem protocols.

[0147] In one possible implementation, for example, the command "Close circuit breaker A" is converted into an MMS protocol command conforming to the IEC 61850 standard and sent to the intelligent power distribution system. Convert "Migrate Virtual Machine B" into a command that conforms to the VMware vSphere API or Kubernetes CSI specification and send it to the cloud management platform; The command "Adjust air conditioning D valve" is converted into a command conforming to the BACnet protocol and sent to the building automation system.

[0148] In automatic execution mode, before issuing instructions, the system will pop up a final confirmation prompt on the human-computer interaction interface, displaying a list of operations to be executed and their impact assessment, providing the operations and maintenance administrator with a final opportunity to intervene.

[0149] Instructions are issued strictly according to the timing and dependencies determined by the strategy deduction, and a timeout is set for each instruction.

[0150] If an instruction times out or returns a failure signal, the system will suspend subsequent instructions and trigger a rollback process or advanced alarm. After each instruction is successfully executed, the system will verify whether the state of the physical world is consistent with the expectation through the monitoring system, update the state of the digital twin accordingly, and then execute the next instruction.

[0151] Based on the same inventive concept described above, this application also discloses a collaborative scheduling device for a data center power distribution system, the structure of which is as follows: Figure 2 As shown, the device includes the following modules: The model building and mapping module is used to build a digital twin containing a multiphysics coupling model and establish an influence chain mapping relationship between physical resources and logical IT load identifiers. The multi-fault identification module is used to respond to multiple physical resource failure events detected by monitoring, and to identify logical IT loads that have abnormal impacts associated with the failure events and form a risk load set based on the mapping relationship between digital twins and impact chains. The collaborative strategy generation module is used to establish a multi-agent decision-making framework based on a digital twin. It inputs the risk load set into the multi-agent decision-making framework, dynamically adjusts the decision target weights of each agent in the multi-agent decision-making framework, and obtains at least one candidate scheduling strategy. The policy security verification module is used to simulate the execution of all candidate scheduling policies in the digital twin. Based on the simulation results, it selects all security policies that can maintain the performance of the risk load set above the preset security threshold from the candidate scheduling policies. The security instruction execution module is used to convert the best security policy in the security policy set into control instructions and issue them for execution.

[0152] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in the computer-readable storage medium, which includes, for example, various media capable of storing program code such as: USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk.

[0153] Those skilled in the art will understand that the step numbers of the above methods or processes are only used to distinguish different steps and do not constitute an absolute restriction on the execution order. Some steps may be executed simultaneously or in a different order than the numbers.

[0154] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for coordinated dispatch of a data center power distribution system, the method comprising: The method comprises the following steps: constructing a digital twin containing a multi-physical field coupling model, and establishing a mapping relationship between a physical resource and a logical IT load identifier; in response to a monitored failure event of a plurality of the physical resources, determining a logical IT load associated with the failure event according to the digital twin and the mapping relationship, and forming a risk load set; based on the digital twin, establishing a multi-agent decision framework, inputting the risk load set into the multi-agent decision framework, dynamically adjusting the decision target weight of each agent in the multi-agent decision framework, and obtaining at least one candidate scheduling strategy; in the digital twin, simulating the execution of all the candidate scheduling strategies, and according to the simulation results obtained, screening a safe strategy set from the candidate scheduling strategies, which can maintain the performance of the risk load set above a preset safety threshold; converting the optimal safe strategy in the safe strategy set into a control instruction and executing it.

2. The method of claim 1, wherein, The method comprises the following steps: performing multi-physical field coupling basic modeling to obtain a digital twin, wherein the multi-physical field coupling model comprises an electrical model, a thermodynamic model, and an IT business logic model, wherein the IT business logic model is used to abstract virtual machines and containers running in the data center into logical IT loads with unique identifiers, and record physical resource information on which the logical IT loads depend; in the digital twin, each logical IT load identifier is associated with corresponding physical resource information to form a dynamic influence chain mapping relationship.

3. The method of claim 1, wherein, The method comprises the following steps: receiving and aggregating a plurality of original alarm events occurring simultaneously or successively within a preset time window from the data center to generate a list of failed physical resources; according to each device identifier in the list of failed physical resources, finding all associated primary logical IT load identifiers in the influence chain mapping relationship, and obtaining a primary associated load set; inputting the list of failed physical resources containing concurrent failures and the primary associated load set into the digital twin for dynamic deduction to obtain secondary logical IT loads affected by the concurrent failures, and collecting corresponding secondary logical IT load identifiers into a secondary impact load set; merging the primary associated load set and the secondary impact load set to obtain a final risk load set.

4. The method of claim 3, wherein, The dynamic deduction process comprises the following steps: In the electrical model of the digital twin, the normal power supply state and the fault power supply state after injecting the concurrent fault in the fault physical resource set are simulated respectively, and a fault power distribution circuit corresponding to a server is identified, which is newly added in the fault power supply state relative to the normal power supply state and whose electrical parameter exceeds a preset safety limit value; Based on the electrical variation of the fault power distribution circuit, a thermal power consumption variation of the corresponding equipment is calculated, and the thermal power consumption variation is input into a thermodynamic model of the digital twin; In the thermodynamic model, a temperature field simulation is performed according to the thermal power consumption variation, and a server whose heat dissipation environment temperature exceeds a preset temperature threshold is predicted; According to the influence chain mapping relationship, the logical IT load carried by the server identified in the electrical model and the thermodynamic model is determined, and the corresponding logical IT load identifiers are summarized into a direct physical risk load set; Based on a predefined business service dependency graph, each logical IT load identifier in the direct physical risk load set is taken as a starting point, and all directly and indirectly associated logical IT load identifiers are traversed in reverse to form a business-related risk load set; The direct physical risk load set and the business-related risk load set are summarized to obtain a secondary influence load set.

5. The method of claim 2, wherein, The multi-agent decision framework is established based on the digital twin, and the risk load set is input into the multi-agent decision framework, which specifically includes the following steps: The multi-agent decision framework is established based on the digital twin, and the multi-agent decision framework includes management agents such as power distribution management agents, IT load management agents, and cooling management agents, and a business guardian agent, which is used to receive external input and coordinate other management agents; Each management agent is configured to obtain its corresponding real-time state data through the digital twin, including electrical data, computing data, and thermodynamic data; The risk load set marked with a business SLA level is input into the business guardian agent, each risk load in the risk load set corresponds to a risk load identifier and the business SLA level, and the business SLA level is a pre-defined logical IT load business continuity assurance level.

6. The method of claim 5, wherein, The decision target weight of each agent in the multi-agent decision framework is dynamically adjusted to obtain at least one candidate scheduling strategy, which is dominated by the business guardian agent, and specifically includes the following steps: After receiving the risk load set, the business guardian agent calculates the business impact degree value corresponding to each risk load according to the preset SLA level weight, the performance decline percentage, and the preset business importance coefficient of each risk load, and the performance decline percentage is the real-time performance decline degree caused by the failure of the physical resource associated with the risk load; The business daemon intelligent agent dynamically adjusts target function weights of the power distribution management intelligent agent, the IT load management intelligent agent, and the cooling management intelligent agent according to the business impact degree value; wherein the higher the business impact degree value of the risk load is, the higher the weight of the related guarantee item in the target function of each intelligent agent is allocated to be; Collaborative decision-making is performed according to the management intelligent agent after weight adjustment, and at least one candidate scheduling strategy meeting the business impact degree load demand is generated.

7. The method of claim 5, wherein, The simulation of the execution of all the candidate scheduling strategies in the digital twin specifically includes the following steps: The real-time running state of the data center power distribution system is synchronized to the digital twin for simulation environment initialization; The business SLA level of each risk load in the risk load set is taken as an attribute tag and is associated to the corresponding IT business logic model in the digital twin; For each candidate scheduling strategy, the multi-physical field coupling model of the digital twin is driven to perform fast rehearsal to simulate the state of the power distribution system in a future preset time after the current candidate scheduling strategy is executed; During the fast rehearsal, the key performance indicator data of each risk load is continuously recorded, and the key performance indicator at least includes a power supply stability indicator, an IT load performance indicator, and a thermal environment indicator.

8. The method of claim 7, wherein, The fast rehearsal specifically includes the following steps: Based on the impact chain mapping relationship, the key physical equipment corresponding to each risk load is determined, and a pre-stored de-rating rule is matched for each key physical equipment based on the associated business SLA level; wherein the higher the business SLA level is, the lower the actual maximum allowable current specified by the matched de-rating rule at the same temperature is; Based on the system state after the load distribution is updated according to the candidate scheduling strategy, electrical model deduction is performed to obtain electrical state and equipment loss, and the electrical state includes the initial current-carrying parameter of the key physical equipment; Based on the equipment loss, thermodynamic model deduction is performed to obtain the updated running temperature of each key physical equipment, and the updated running temperature is fed back to the electrical model to calculate the corrected current-carrying parameter of each key physical equipment in combination with the corresponding de-rating rule; Based on the corrected current-carrying parameter, it is determined whether a preset convergence condition is met, if the convergence condition is met, it is determined that the system reaches a safe convergence steady state, and if the convergence condition is not met, an iteration strategy of the fast rehearsal is determined according to the non-convergence condition until the iteration reaches the safe convergence steady state; Based on the final physical state under the safe convergence steady state, the IT business logic model is driven to simulate business performance, and the key performance indicator data is recorded.

9. The method of claim 7, wherein, Based on the obtained simulation results, a safe strategy set that can maintain the performance of the risk load set above a preset safety threshold is screened from the candidate scheduling strategies, and the safe strategy set specifically includes the following steps: A safety threshold of the key performance indicator associated to each risk load in the risk load set is set, and the performance safety threshold is associated to the business SLA level of the risk load. According to a rehearsal result corresponding to each candidate scheduling strategy, each risk load in the risk load set is checked for security one by one; When the key performance indicators of all the risk loads are within the corresponding security threshold for a future period of time, the corresponding candidate scheduling strategy is determined as a safe strategy; All the safe strategies are summarized to form a safe strategy set.

10. A co-scheduling device of a data center power distribution system, characterized in that, The method comprises the following modules: A model construction and mapping module is configured to construct a digital twin comprising a multi-physical field coupling model, and establish an influence chain mapping relationship between physical resources and logical IT load identifiers; A multi-fault identification module is configured to, in response to a monitored fault event of a plurality of physical resources, determine a logical IT load having an abnormal influence correlation with the fault event according to the digital twin and the influence chain mapping relationship, and form a risk load set; A collaborative strategy generation module is configured to establish a multi-agent decision-making framework based on the digital twin, input the risk load set into the multi-agent decision-making framework, dynamically adjust the decision-making target weight of each agent in the multi-agent decision-making framework, and obtain at least one candidate scheduling strategy; A strategy security verification module is configured to simulate execution of all the candidate scheduling strategies in the digital twin, and select a safe strategy set from the candidate scheduling strategies according to a simulation result obtained, wherein the safe strategy set is capable of maintaining the performance of the risk load set above a preset security threshold; A safe instruction execution module is configured to convert an optimal safe strategy in the safe strategy set into a control instruction and execute the control instruction.