PDU load distribution system based on reinforcement learning

The PDU load distribution system, which combines a multi-dimensional execution cost model with a digital twin environment, solves the problem of efficient scheduling of virtual machine migration and server startup and shutdown in the data center, achieves a balance between energy consumption and performance under different load scenarios, prioritizes the protection of key businesses, and improves the security and reliability of the data center.

CN120596258AActive Publication Date: 2025-09-05ANHUI WEIYUAN NEW ENERGY TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510689960.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing reinforcement learning cannot simultaneously quantify the instantaneous impact of virtual machine migration/server startup and shutdown on the power supply link and the combined cost of critical business security redundancy in data center PDU load distribution. As a result, there is a lack of an effective autonomous judgment mechanism when the load pattern drifts or the security level is increased, which can easily cause phase overload or current shock, and cannot meet the high reliability requirements of financial and government-level bearer environments.

Method used

By combining a multi-dimensional execution cost model with a digital twin environment, efficient scheduling of PDU loads, server migration, and start-up and shutdown is achieved. Offline pre-training and hierarchical control are used to balance energy consumption and performance in different load scenarios. Online learning optimization is performed under concept drift detection, and security costs and fault tolerance overheads are superimposed to prioritize the protection of critical business and sensitive data.

Benefits of technology

It achieves efficient, secure, and scalable management of data centers in multiple dimensions, avoids resource impact and performance jitter caused by large-scale operations, flexibly responds to environmental changes, and ensures the security and reliability of key businesses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596258A_ABST
    Figure CN120596258A_ABST
Patent Text Reader

Abstract

The invention discloses a PDU load distribution system based on reinforcement learning, relates to the technical field of data center resource management, and realizes efficient scheduling of PDU load, server migration and start-stop through combination of a multi-dimensional execution cost model and a digital twin environment. Energy consumption and performance can be balanced in different load scenes by using offline pre-training and hierarchical control, and continuous optimization is performed through online learning under concept drift detection; in addition, after the safety cost and the fault-tolerant overhead are overlaid, key services and sensitive data can be protected preferentially, safety risks and downtime losses are effectively reduced, and finally multi-dimensional collaborative efficient, safe and extensible data center management is achieved. Meanwhile, resource impact and performance jitter caused by large-scale operation are further avoided through batch migration and staged starting and stopping, flexible response can be achieved under multi-dimensional risks, and it is ensured that scheduling robustness and sustainability are kept under heterogeneous loads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data center resource management, and in particular to a PDU load distribution system based on reinforcement learning. Background Art

[0002] As energy-intensive businesses such as cloud computing, AI inference and training, and big data analytics rapidly converge in hyperscale data centers, computing loads are experiencing transient peaks, non-stationary waveforms, and a diversification of business types. Server cabinet power consumption often increases exponentially with second-by-second traffic surges. The power supply must coordinate PDU (Power Distribution Unit) phase balancing, backup power switching, and busbar redundancy within milliseconds. Simultaneously, hot and cold aisle temperature control, liquid cooling pump speed, and UPS energy storage scheduling must also be coordinated in real time. Traditional heuristic scheduling, based solely on CPU utilization or task queue depth, is unable to balance power peak suppression, energy utilization improvement, and the stability requirements of the long-tail load of AI training. In recent years, deep reinforcement learning (RL) has been used for resource scheduling, but most studies only use server power consumption or electricity costs as reward signals, ignoring physical and operational limitations such as phase line capacity differences in PDU topology, backup battery discharge depth, critical business security isolation, and multi-replica fault tolerance overhead. This results in slow policy convergence and easy breakthrough of power risk thresholds when implemented in real scenarios, making it difficult to meet the high reliability requirements of "zero failure and zero data loss" in financial and government-level hosting environments.

[0003] After searching, a Chinese invention patent with application publication number CN109324875A discloses a method for power consumption management and optimization of data center servers based on reinforcement learning. Reinforcement learning methods are used to solve the power consumption management and optimization problems of data centers. By continuously observing the load arrival, load distribution and power consumption information of the random system of the data center, decisions are made sequentially. That is, based on the state observed at each moment, an action is selected from the available action set to make a decision. The decision maker makes a new decision based on the newly observed state, and this process is repeated. The present invention can directly optimize the load distribution strategy of the data center online without any prior knowledge, thereby reducing the overall operating power consumption of the data center.

[0004] Existing scheduling frameworks are unable to simultaneously quantify the combined costs of "the instantaneous impact of virtual machine migration / server startup and shutdown on the power supply link" and "critical business security redundancy." This results in the RL agent lacking an effective autonomous judgment mechanism when faced with power supply fluctuations, load pattern drift, or security level increases. When multiple GPU training tasks in the same cabinet are migrated in batches to adjacent cabinets at night to save cooling energy, the algorithm does not consider the instantaneous margin and phase line imbalance of the target PDU circuit, often causing phase overload or current surges. If the core financial business virtual machine is migrated to a node with only a single backup and is in a maintenance window at this time, the fault tolerance level will drop sharply. Once the PDU overload triggers a circuit breaker or the host dual-machine hot standby fails, the result will be AI task interruption, transaction system SLA breach, and source station log data corruption, which will then lead to chain compensation and brand reputation loss.

[0005] Therefore, there is an urgent need for an intelligent scheduling method that can simultaneously embed power topology constraints, security isolation levels, and fault-tolerant redundancy overheads in a multi-dimensional execution cost model, and can adaptively retrain when load concept drift occurs, so as to achieve coordinated optimization of data center energy consumption, performance, and security and reliability. Summary of the Invention

[0006] (1) Technical problems solved

[0007] In response to the deficiencies of the existing technology, the present invention provides a PDU load distribution system based on reinforcement learning. By combining a multi-dimensional execution cost model with a digital twin environment, efficient scheduling of PDU loads, server migration, and start-up and shutdown is achieved. Offline pre-training and hierarchical control are used to balance energy consumption and performance in different load scenarios, and continuous optimization is achieved through online learning under concept drift detection. In addition, after superimposing security costs and fault-tolerant overheads, key businesses and sensitive data can be protected first, effectively reducing security risks and downtime losses, and ultimately achieving multi-dimensional collaborative, efficient, secure, and scalable data center management. At the same time, through batch migration and phased start-up and shutdown, resource impacts and performance jitters caused by large-scale operations can be further avoided, and flexible responses can be made under multi-dimensional risks; the technical problems recorded in the background technology are solved.

[0008] (2) Technical solution

[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: a reinforcement learning-based PDU load distribution system, including: when a data center load fluctuation is detected, calling a real-time monitoring component to obtain raw monitoring data, deeply quantifying the energy consumption and latency of virtual machine migration and server startup and shutdown through a multi-dimensional execution cost model, and outputting a dynamic cost sequence;

[0010] After receiving the dynamic cost sequence and raw monitoring data, multiple rounds of reinforcement learning offline training are conducted on the PDU power distribution topology, server cluster, and typical load scenarios in the digital twin environment to select pre-trained strategies for hierarchical scheduling.

[0011] After successfully loading the pre-trained policy, the upper-level reinforcement learning controller generates macro migration and start-stop instruction sets. The lower-level flexible scheduling module executes them in batches based on multi-dimensional execution costs and safety limits, and records specific scheduling feedback to the system log.

[0012] If the concept drift metric exceeds the set threshold or the policy benefit drops significantly, a small-scale online learning is performed based on the digital twin environment and the latest scheduling feedback, and a resource allocation plan is deployed in advance using the load forecasting model;

[0013] Security costs and fault-tolerance overheads are injected into the original multi-dimensional execution cost model, and reinforcement learning strategies are prioritized through security-related penalties or rewards during offline pre-training. During the hierarchical scheduling and online adaptation stages, redundancy, encryption, or batch migration of critical servers or sensitive businesses are prioritized.

[0014] Preferably, the multi-dimensional execution cost model sets independent dimensions for virtual machine migration, server start and stop, and PDU switching, and quantifies the time delay, energy consumption overhead, and business performance impact of each dimension in a weighted manner to dynamically measure the comprehensive cost of each scheduling action in the current state and write it into the monitoring database.

[0015] Preferably, the deployed monitoring component collects server utilization, PDU load, power startup time and network bandwidth in a fixed time window, and maps the collected results to the cost model in real time to dynamically update the cost data; and generates an available margin threshold based on the bandwidth safety factor and the temperature safety factor as the basis for subsequent batch scheduling judgment.

[0016] Preferably, the constructed digital twin environment fully maps the PDU power distribution topology, server hardware configuration, network bandwidth limitations and typical business load curves of the real data center, and uses performance incentive coefficients and cost penalty coefficients in offline simulations to guide the reinforcement learning agent to generate pre-training strategies.

[0017] Preferably, multiple rounds of reinforcement learning simulation training are performed on various load scenarios in the digital twin environment. During the offline training process, the cost sensitivity is adjusted to amplify the negative returns of high-cost actions, and the difference between the new and old strategies is regularized and constrained by the strategy stability coefficient, and pre-training strategies and performance evaluation results that take into account both energy efficiency and performance are output.

[0018] Preferably, after the upper-level reinforcement learning controller calls the pre-training strategy to generate an instruction set, it first filters out the cost peak instructions based on the real-time cost and priority label, and then hands them over to the lower-level module for execution in batches according to the scheduling amplification factor and flexibility control parameters; wherein, before generating the instruction set, instructions with costs exceeding the threshold are filtered out based on the multi-dimensional execution cost model, and execution priority is assigned to the remaining instructions.

[0019] Preferably, the lower-level scheduling module executes the instructions in the instruction set in batches according to the bandwidth margin and the server temperature margin, and after each batch is completed, returns the scheduling feedback including the actual execution time, power consumption changes and performance jitter together with the updated cost vector to the upper-level controller for use in the online adaptive stage.

[0020] Preferably, the difference between the actual system state and the expected state is monitored to calculate the environmental drift degree;

[0021] When the drift degree exceeds a preset threshold, a drift alarm is generated, and the triggering timestamp and main offset indicators are recorded to trigger subsequent online training or transfer learning actions.

[0022] Preferably, the online training process selects the most recent several scheduling cycles to form a micro-training data set, uses the learning rate to incrementally update the old strategy to obtain a new strategy, and redeploys it after completing the rapid regression test in the digital twin environment.

[0023] The micro-training dataset includes the current state, the actions generated by the upper-level reinforcement learning controller, the multi-dimensional execution costs, and the actual benefits.

[0024] Preferably, historical load sequences and recent monitoring data are called up, combined with long-short-term memory networks to predict the load for several future time periods, and when the mean absolute percentage error of the prediction deviation exceeds a set threshold, it is fed back to the time series prediction model to iteratively update the parameters, and accordingly pre-set the available computing and power resources.

[0025] Preferably, new security cost parameters and fault tolerance overhead parameters are added to the original multi-dimensional execution cost model to form a comprehensive cost model, in which the priority of key business migration, redundant copy synchronization and sensitive node power management is controlled by the security weight coefficient and the fault tolerance weight coefficient.

[0026] Preferably, during the offline simulation process, positive security rewards and negative fault-tolerant penalties are added to the reward function to guide the scheduling strategy to prioritize the secure isolation of sensitive services and high-availability redundancy of key nodes. If the host where the sensitive service is located has not completed encryption isolation or copy backup, the shutdown operation is delayed and the log is recorded in the execution feedback.

[0027] Preferably, after introducing additional offsets related to safety alarms and redundant status of key nodes, during online adaptation, the safety and fault tolerance weight coefficients are dynamically adjusted according to the frequency of safety alarms and the number of fault triggers to maintain the preset balance target between energy consumption, performance and reliability, and the actual execution results are fed back to the digital twin environment again.

[0028] (3) Beneficial effects

[0029] The present invention provides a PDU load distribution system based on reinforcement learning, which has the following beneficial effects:

[0030] This solution is linked through steps 1 to 5, combined with the multidimensional cost function C(a i ,s t ) and the digital twin environment Ω, enabling coordinated management of PDU load distribution, server startup and shutdown, and safety fault tolerance, achieving the following beneficial results:

[0031] Step 1: Quantify the time delay, energy consumption and performance jitter of virtual machine migration, server startup and shutdown operations into a multidimensional cost function C(a i ,s t ) and introduces real-time monitoring to accurately characterize the cost of each action. Step two embeds this multi-dimensional execution cost into the offline training process, using the digital twin environment Ω to simulate various load scenarios and iterate through multiple rounds of pre-training strategies to ensure that the reinforcement learning agent can master the ability to spontaneously avoid high-cost actions before actual deployment.

[0032] Step 3: Use a hierarchical control architecture: upper layer reinforcement learning controller π top Based on pre-trained strategies, macro-load instructions are generated. Lower-level modules then flexibly schedule and execute them in batches, avoiding system impacts caused by large-scale migrations or frequent starts and stops. This layered mechanism not only balances energy consumption and performance but also continuously collects scheduling feedback during execution, providing high-precision data for online adaptation in step four.

[0033] Step 4: When facing non-stationary loads, the concept drift metric is used to capture environmental changes, and small-scale online training is combined with rapid adjustment of strategies in the digital twin environment Ω, which is both agile and stable, allowing the system to maintain high efficiency in long-term operation. Step 5: Add a new safety cost function and fault-tolerance cost function In the original multidimensional cost function C(a i ,s t ) is combined to obtain the comprehensive cost function C'(a i ,s t), it can impose safety / fault-tolerance rewards or restrictions in offline training (step 2) and hierarchical scheduling (step 3), and dynamically increase weights as needed during online adaptation (step 4), achieving fine-grained control over key server redundancy and sensitive business isolation.

[0034] Through the unified state t 、Action a i With the cost function C′(a i ,s t ) are mutually coupled: early offline simulation lays the strategic foundation for later hierarchical scheduling, and hierarchical execution feedback feeds back to online adaptation. The addition of safety and fault-tolerant factors gives the system as a whole more resilience and expands the scope of strategic adaptation. It not only maintains the continuous advancement of energy consumption optimization, but also can flexibly respond to multi-dimensional risks (performance jitter, environmental drift, security threats) to achieve comprehensive management and control effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a structural diagram of the PDU load distribution system based on reinforcement learning in the present invention. DETAILED DESCRIPTION

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0037] See also Figure 1 The present invention provides a PDU load distribution system based on reinforcement learning, comprising:

[0038] Step 1: When load fluctuations are detected in the data center, the real-time monitoring component is called to collect in-depth data on server utilization, PDU load, power startup time, network bandwidth and other indicators, and the multi-dimensional cost function C(a i ,s t ) performs nonlinear calculations on action costs using the energy consumption, latency, and performance jitter weights in the model, dynamically updates the comprehensive cost data format, and enables instant evaluation and quantitative characterization of virtual machine migration or server startup and shutdown. The generated timing cost information is then provided as a key input for subsequent offline pre-training and online adaptation.

[0039] The step 1 includes the following:

[0040] Step 101: Establish a multi-dimensional execution cost model

[0041] Step 101 defines a unified multi-dimensional execution cost function for various execution actions in the data center (including virtual machine migration, server startup and shutdown, and PDU switching) based on historical operation and maintenance data and predictive analysis results.

[0042] Let a i represents the i-th type of action that can be performed in the data center environment, such as migrating a virtual machine to the target host or activating a dormant server; let s t A comprehensive description of the current system status, such as server utilization, PDU remaining power margin, network bandwidth load, etc.

[0043] Let E(a i ,s t ) means in state s t Next, perform action a i The expected additional energy consumption; let M(a i ,s t ) represents the performance migration overhead or service jitter degree that may be caused by the action; let L(a i ,s t ) represents the operation delay caused by the action (the time from the instruction to the effectiveness of the process). Based on the above three key cost factors, a multidimensional cost function C(a) can be constructed under dimensionless conditions. i ,s t ), as follows:

[0044]

[0045] Among them, ζ1, ζ2, and ζ3 are positive weighting coefficients, all ranging from 0 to 1, and their sum is 1. They are used to balance the relative importance of energy consumption, performance jitter, and latency.

[0046] φ1, φ2, φ3 are action sensitivity coefficients, which are greater than 0 and are dimensionally matched according to the average cost magnitude of different action categories. They are used to control the amplification or attenuation degree of each cost factor in the exponential or logarithmic transformation.

[0047] Multidimensional cost C(a i ,s t ) returns a higher value, indicating that the overall execution cost of the action in the current state is greater; by splitting the three key dimensions of energy consumption, performance, and latency, it is easier to identify high-risk or high-overhead actions and accurately characterize the action cost; the universality and portability of the parameter system: through the weighting coefficient ζ i and sensitivity coefficient φ i The adjustable properties can adapt to data centers of different sizes and business needs, and realize the multi-scenario reuse of a set of function structures.

[0048] Step 102: Real-time monitoring and dynamic updating

[0049] In establishing the multidimensional cost function C(a i ,s t ), step 102 further deploys real-time monitoring components in the data center to dynamically collect and update the key parameters defined in step 101;

[0050] Deploy monitoring probes to collect hardware-layer data such as server CPU utilization, network I / O usage, and power startup time;

[0051] Deploy a load analysis engine to track the real-time performance status and service delay of each virtual machine and convert it into migration overhead M(a i ,s t ) and operation delay L(a i ,s t ) required statistics, integrating the output of the power monitoring module to calculate and summarize the additional energy consumption generated when the action is executed (a i ,s t ) estimated value.

[0052] The monitored incremental energy consumption, performance jitter, and delay information are continuously input into the multi-dimensional cost function C(a) of step 101. i ,s t ), get the latest execution cost value; the latest multidimensional cost function C(a i ,s t ) and the current system state s t The unified storage forms a traceable and time-evolving cost time series data set for subsequent offline pre-training and online adaptive learning. Through real-time updates, the execution cost of each action under different load conditions can be dynamically characterized, avoiding decision bias caused by historical or static configuration alone. When the scale or hardware type of the data center expands, the new cost factors can be incorporated into the multidimensional cost function C(a) by adding or upgrading monitoring components. i ,s t ) system, with good flexibility and maintainability.

[0053] Step 2: When the execution cost data and real-time monitoring records are obtained, the PDU distribution topology, server cluster and typical load pattern are loaded into the digital twin environment Ω, and a reinforcement learning agent is used to perform offline training for multiple rounds of simulation interaction. The nonlinear objective function is used to balance energy consumption, performance and multidimensional cost function C(a i ,s t ) overhead, screen out pre-training strategies, complete strategy performance evaluation and solidification, and provide accurate and feasible macro-decision templates for subsequent hierarchical scheduling;

[0054] The second step includes the following:

[0055] Step 201: Build a digital twin environment and load the execution cost model

[0056] According to the system status s output in step 1 t 、Action a i and its corresponding multidimensional cost function C(a i ,s t ) and build a digital twin environment Ω that is highly consistent with the real data center in the virtual simulation platform;

[0057] Define the complete PDU power distribution topology, computing nodes (server clusters), and network link attributes within the digital twin environment Ω, and set up corresponding load injection modules to simulate workloads in various typical business scenarios (such as daytime peak loads, nighttime low loads, and highly fluctuating loads for AI training).

[0058] Combine the real-time monitoring data from step 1 with the historical load pattern to generate the initial state set of the digital twin environment Ω , where each state s t Including server utilization, PDU power supply margin, network occupancy, etc.;

[0059] Use a i Action Set Strictly reproduce the impact of the execution action in the digital twin environment Ω and call the multidimensional cost function C(a i ,s t ) to evaluate the comprehensive costs of energy consumption, delay, and performance jitter of each action in the simulation environment; build an offline data interaction interface so that the reinforcement learning agent can obtain the current state s at any time when interacting with the digital twin environment Ω t And based on action a i Calculate the corresponding comprehensive cost C(a i ,s t );

[0060] The PDU topology, server cluster and load characteristics are reproduced through the digital twin environment Ω, making the training and verification phases as close to the production environment as possible, reducing the trial and error risk of the real system. The multidimensional cost function C(a i ,s t ) evaluates the true cost of each action, ensuring that offline simulations accurately reflect the operational costs of subsequent deployments, providing a reliable basis for policy training. High-cost actions can be accurately identified in the digital twin environment, effectively suppressing frequent migrations or low-benefit switching behaviors during simulation training.

[0061] Step 202: Offline pre-training and strategy verification

[0062] Taking into account the positive incentives for improving system performance and the negative penalties for execution costs, the instantaneous benefit function R(s) of the reinforcement learning agent is defined. t ,a i ), as shown below:

[0063]

[0064] Where: α is the performance incentive coefficient, which is greater than 0 and is used for dimension matching and balancing convergence speed, which can amplify the current state s t The overall performance index of the system is Λ(s t ) of the income; Λ(s t ) can be obtained by comprehensive measurement of data center throughput, task completion rate, etc. The larger the value, the better the performance. The specific value can be obtained by weighting after linear normalization; ∈ is the cost penalty coefficient, which is greater than 0 and is used to suppress high-cost actions; δ is the cost sensitivity, which is greater than 0 and is used for dimensional matching and amplifying C(a) in logarithmic or exponential operations. i ,s t )’s impact.

[0065] In the digital twin environment Ω, from the state s t Starting, the agent tries different actions a i , and according to the above instant income R(s t ,a i ) performs multiple rounds of iterative updates on the strategy; by simulating real load fluctuations through the monitoring data collected in step 1, the intelligent agent can learn the optimal or suboptimal load distribution method in various situations such as high load and low load, network redundancy and network congestion; in the later stage of training convergence, by increasing the cost sensitivity δ, the impact of the execution cost can be further amplified, guiding the intelligent agent to reduce high-cost actions with limited performance benefits.

[0066] After the training is completed, the agent's cumulative benefits, average execution cost, load balancing and other indicators in each scenario are statistically analyzed to obtain several candidate strategies. The strategy with the highest comprehensive score in energy efficiency and SLA performance is selected as the pre-training strategy, and the corresponding performance evaluation result Θ is generated, including the average migration frequency, overall energy consumption and response time under the strategy. The obtained pre-training strategy and performance evaluation result Θ are output together.

[0067] SLA performance refers to the key quality indicators agreed upon in the Service Level Agreement (SLA) to measure whether the system or service meets the contract requirements, including availability, response latency, and throughput.

[0068] Large-scale migration or testing in a real data center environment will bring high risks and costs. Step 202 completes a large number of simulation iterations in the digital twin environment Ω, which can quickly find the best strategy and avoid affecting production business. t ,a i ) is embedded with a nonlinear amplification method, which enables the intelligent agent to consciously avoid operations that have limited contribution to the system benefits but are too costly to execute, thereby improving the stability and energy efficiency during subsequent deployment and effectively suppressing low-benefit and high-cost actions; once the simulation training converges, a pre-trained strategy that can be directly transplanted is generated, and the output strategy can be directly applied in the subsequent hierarchical control architecture, which can greatly shorten the online debugging cycle and obtain better scheduling efficiency in the initial stage.

[0069] Step 3: When the pre-trained strategy is deployed to the actual data center, the upper reinforcement learning controller uses the macro load distribution solution Π selected in step 2 to distribute the macro load to the actual data center. top Generate migration or start and stop instruction set {a i ′}, and pass it to the lower scheduling module to refer to the real-time monitoring information of step 1 and the multi-dimensional cost function C(a i ,s t Dynamic data is executed in batches, with segmented control and safety limit verification to smoothly implement virtual machine migration and server startup and shutdown, generating traceable execution feedback for subsequent online adaptive calls;

[0070] The step three includes the following:

[0071] Step 301: Upper layer reinforcement learning controller deployment and global load distribution

[0072] Load the pre-training strategy output in step 2 into the upper reinforcement learning controller π top , refers to the decision-making module deployed at the top level of the data center scheduling architecture and running the reinforcement learning strategy; the upper reinforcement learning controller Π top (s t ) can be used given the current system state s t When generating macro load distribution decisions, that is, giving high-level instruction sets {a i};

[0073] The performance evaluation result Θ generated in step 2 is placed on the upper reinforcement learning controller Π top A reference database is used to compare the actual effect and detect deviations during the strategy execution process. i ,s t) maintains linkage and inputs the latest real-time monitoring information (server utilization, PDU load, power startup time, etc.) into the upper reinforcement learning controller Π top In the state observation, the controller is ensured to be aware of the possible high-cost operations of the current system;

[0074] When the upper reinforcement learning controller Π top Calculate an action instruction set {a i}, based on the multidimensional cost function C(a i ,s t ) Re-filter command actions and automatically lower the priority of actions with excessively high costs or unclear benefits to avoid large-scale or frequent operations at the source;

[0075] The upper layer reinforcement learning controller π top The final set of action instructions retained in {a i ′} is defined as a global load distribution instruction, such as migrating several virtual machines to a specified host to enable / disable a specific server group, etc.; the corresponding action instruction set {a i '} together with its corresponding priority tag is output to the lower-level scheduling module in step 302 for implementing specific operations in batches or segments.

[0076] During the execution process, if the lower scheduling module (step 302) returns execution delay or partial operation failure information, the upper reinforcement learning controller Π top A rapid simulation evaluation will be conducted in conjunction with the digital twin environment Ω constructed in step two to determine whether a temporary strategy adjustment is needed. This intermediate feedback allows the controller to fine-tune the current strategy and also provides more realistic operating data for the subsequent step four: online adaptive learning and continuous iterative optimization.

[0077] When in use, the pre-trained strategy is directly used for scheduling in the real environment, which greatly shortens the online tuning time and makes the initial strategy have a good compromise between energy consumption and performance, which can ensure the global optimal tendency of macro decision-making; by real-time reference to the multi-dimensional cost function C(a i ,s t ), it can prioritize or delay those control instructions with excessive overhead when generating actions, reducing the execution burden of subsequent lower layers. Once the lower-layer execution results do not meet expectations, the strategy can be quickly tested and corrected in the digital twin environment Ω, forming a rapid test-adjustment closed loop. In traditional systems, problems such as excessive migration or excessive start-stopping are often only discovered during the execution phase. This solution uses the multidimensional cost function C(a) at the upper layer to i ,s t ) for pre-screening, which can significantly reduce execution oscillation.

[0078] Step 302: Lower-layer flexible scheduling and multi-stage execution

[0079] The lower-level scheduling module obtains the action instruction set output by the upper-level {a i ′}, and read its corresponding priority tag;

[0080] According to the dynamic monitoring data in step 1 (such as network bandwidth BW t 、Server Temperature Temp t etc.), determine whether the current system has the ability to execute multiple instructions {a i ′} simultaneous operation capability. If simultaneous execution is not possible, scheduling is performed according to priority or dependency order. For example, the system will first schedule each migration or start / stop instruction a i Estimate its resource requirements for network bandwidth and host cooling, and read the current available bandwidth BW in real time avail = Current total bandwidth BW cap -Used bandwidth BW t Temperature margin ΔT tolerated by the server avail (s) = Safety temperature upper limit Temp max -Current temperature t (s). The system then accumulates the bandwidth requirements of all instructions and adds them to BW avail ×μ bw (Bandwidth Safety Factor) comparison; At the same time, for each target server, the temperature change ΔT caused by the instruction i and ΔT avail (s)×μ tmp (Temperature Safety Factor). The system only considers parallel execution possible if both network and temperature conditions are met for all commands. Otherwise, the commands are split into smaller batches to ensure that each operation is executed within the bandwidth and temperature safety margins. This prevents performance jitter or service interruptions caused by bandwidth overload or data center overheating.

[0081] For a large number of virtual machines that need to be migrated or multiple servers that need to be enabled, adopt a batch and segmented execution strategy:

[0082]

[0083] Where: k represents the number of operations in the current batch, such as the number of virtual machines that need to be migrated;

[0084] ω and γ are the scheduling amplification factor and flexibility control parameter, respectively. Both are positive real numbers and are used to determine the specific execution scale of each batch. When γ>1, the execution scale of subsequent batches gradually increases with the increase of batch number, and vice versa.

[0085] Through the above batch control function, you can first perform a small number of migration or start-stop tests to verify whether their impact on server load and network I / O is within the simulation range of step 2, and then increase the volume in batches based on the execution results.

[0086] After each batch or segment is executed, monitor the current system status s t Whether it reaches or exceeds the pre-set safety limit (such as the maximum available margin of PDU load, the critical value of network occupancy), if it is close to or exceeded, the subsequent operation is suspended and the upper layer reinforcement learning controller π is sent to the upper layer. top Report;

[0087] The execution process data (including actual migration duration, server power-on time, success / failure logs) is synchronized to the digital twin environment Ω in step 2 and the next step of online adaptive learning for reference.

[0088] If the upper-level instruction includes shutting down multiple idle or low-load servers, the lower-level module will first check whether the current task migration is complete to ensure that power is not cut off before the migration is complete. When it is confirmed that the shutdown is possible, the servers are shut down in small batches to avoid a high current shock at one time. The lower-level module refers to the execution unit deployed in the data center control architecture and is specifically responsible for implementing the upper-level scheduling instructions.

[0089] Splitting large-scale migration or start-stop actions into multiple segments and gradually observing the system impact caused by each segment execution can effectively prevent resource bottlenecks and performance jitters; by continuously monitoring status t By comparing key indicators with safety thresholds, the operation scale can be flexibly controlled during execution, ensuring the protection of SLA and equipment health.

[0090] Upper-level reinforcement learning controller π top Using the pre-training strategy output from step 2, the multi-dimensional execution cost data C(a) of step 1 is integrated when generating the global load distribution instruction. i ,s t ), reducing high-cost actions at the source; the lower-level flexible scheduling then allocates the execution order based on the real-time monitoring safety limits and multi-segment batch strategies to ensure a smooth transition of the system during large-scale migration, start-stop, or PDU switching; the instructions generated by the upper-level controller are gradually implemented at the lower level, and the execution process data of the lower level is fed back to the upper level and the digital twin environment Ω, further providing real-scene feedback for the online adaptive training in step 4.

[0091] Therefore, the combination of pre-training strategy and batch / segment scheduling mechanism significantly improves the security and efficiency when implemented in a real data center environment, avoids the common defects of relying solely on macro-optimization while ignoring the impact of the execution layer, and improves the global optimization capability and execution stability.

[0092] Step 4: When the concept drift metric CD(Δ(s t When the threshold is exceeded or the policy benefit drops significantly, fast, small-scale online training is performed based on the actual scheduling feedback collected in step 3 and the digital twin environment Ω to incrementally fine-tune the reinforcement learning policy. Combined with the prediction model to predict short- and medium-term loads, virtual machine migration or backup power activation plans are pre-arranged. The updated policy is then deployed in the real environment to maintain global optimization of multi-dimensional cost and performance indicators under non-stationary loads.

[0093] The step 4 includes the following contents:

[0094] Step 401: Environmental drift monitoring and concept drift detection

[0095] After the lower-level flexible scheduling is completed in step 3, the actual system performance of each execution batch is compared. The expected state verified in the digital twin environment Ω in step 2 Collect the errors of key indicators (such as server utilization, PDU power supply margin, and network occupancy) after load distribution;

[0096] Let Δl(s τ ) represents the difference vector between the actual system state and the ideal or expected state at time τ. To capture the cumulative deviation and the rate of change of the deviation in the recent period, the sliding time window length can be defined as L. When evaluating the drift at the current time t, the evolution of the difference in the interval [tL, t] is comprehensively considered, and the following metric function is given:

[0097]

[0098] Where: Δ(s τ ) represents the difference between the actual state and the expected state of the system at time τ, in the form of a vector, which is used to measure the deviation of multi-dimensional load indicators (such as CPU utilization, PDU load, network occupancy, etc.) at that time;

[0099] W represents a configurable positive definite symmetric weighting matrix, which is used to assign differential weights between the various indicators in the vector. For example, CPU utilization deviation can be given a higher weight, while temperature or network indicators are given a lower weight;

[0100] [Δ(s τ )] T W[Δ(s τ )] can be regarded as a quadratic form under this weighting matrix, which is used to measure the comprehensive impact of multidimensional deviations.

[0101] It represents the first-order derivative of the deviation vector with respect to time τ, which is used to reflect the speed of change of the deviation with time; κ is the drift speed sensitivity coefficient, which is greater than 0 and is used to control the deviation of Δ(s τ ) the degree of concern about the rate of change over time;

[0102] β is the overall sensitivity coefficient of drift, which is greater than 0, and L is the length of the sliding time window;

[0103] If the obtained deviation CD(Δ t ) exceeds the offset threshold τ cd , it can be considered that a significant drift has occurred. When the drift monitoring module detects CD(Δ t )>τ cd , triggers subsequent online training or transfer learning actions.

[0104] If the deviation CD (Δ t ) is greater than the offset threshold τ cd A drift alarm is generated and details such as the triggering timestamp and main offset indicators are recorded. This drift alarm is simultaneously sent to step 402 to start a targeted strategy fine-tuning process and notify step 403 to adjust the strategy in advance in combination with the load forecasting mechanism.

[0105] When in use, by continuously monitoring the difference between the actual and expected states, it can quickly capture environmental changes, avoid using outdated strategies for a long time under non-stationary loads, and identify strategy mismatches in real time; only when the deviation CD (Δ t ) value exceeds the deviation threshold τ cd Subsequent learning is performed only when the training is complete, which can reduce unnecessary training overhead and maintain system stability.

[0106] Step 402: Small-scale online training and transfer learning

[0107] After receiving the drift alert, extract data from the real operation logs in the recent execution cycles and the execution feedback collected in step 3 to form a micro training dataset Micro training dataset Including current status Upper-level reinforcement learning controller π top The action generated i ′、Multi-dimensional execution cost and information such as actual benefits (SLA performance, energy consumption).

[0108] Importing the micro training dataset into the digital twin environment Ω Using the pre-training strategy output in step 2 as the starting point, perform small-scale online training or transfer learning:

[0109]

[0110] Where: old is the original policy function, such as the upper layer reinforcement learning controller Π top The current strategy used; η is the learning rate, which can be adjusted according to the drift amplitude and system stability requirements;

[0111] It is a loss function for reinforcement learning or deep learning, used to minimize the loss in a small training dataset. The estimation error and cost penalty on ; a composite loss function based on policy gradient is used to train the upper reinforcement learning controller:

[0112]

[0113] The first term calculates the negative logarithmic probability weighted reward for the state-action-reward triple (s, a, r) ​​in the dataset D, and realizes the balance between the performance incentive αΛ(s) and the cost penalty ∈(e δC(a,s) -1);

[0114] The second term is in the old strategy Π old The Kullback-Leibler divergence regularization is introduced between the new strategy π and the new strategy π, and the weight coefficient λ is used to smooth the strategy update to avoid decision instability caused by a one-time large adjustment, thereby ensuring that during the offline pre-training and online fine-tuning stages, the strategy can not only continuously improve the overall benefit of the system, but also maintain the continuity and robustness with the verified strategy; in this process, the execution cost function C(a i ,s t ) to evaluate the action cost and ensure that the strategy after transfer learning does not sacrifice sensitivity to high-cost actions;

[0115] The trained strategy Π new Perform multi-scenario rapid verification in the digital twin environment Ω (the number of simulations is less than the complete training in step 2). If the benefits are significantly improved and the system fluctuations are controllable, replace the trained strategy π in the upper controller. new If the new strategy is found to be unstable or limited by insufficient data sets during the verification process, the trained strategy Π is maintained. old And only incrementally update some actions.

[0116] After the concept drift alarm, small-scale online training can be carried out to promptly correct the policy parts that do not match the new environment, avoiding the high cost and long delay caused by full retraining; new Rapid verification can be performed to screen out high-risk solutions before they are applied to real systems, ensuring data center business continuity. To reflect the latest environment, efficient simulation fine-tuning is performed in the digital twin environment Ω, so that online training can quickly complete convergence without disturbing the production system, realizing the dual linkage of micro training data + digital twin.

[0117] Step 403: Combine the prediction model to conduct forward-looking strategic layout

[0118] While concept drift monitoring and small-scale online training are running in parallel, historical load series and recent monitoring data are called, and a forward-looking algorithm based on a time series forecasting model (such as a long short-term memory network LSTM) is applied to derive load trend estimates for the next T periods. Combined with the multidimensional cost function C(a i ,s t ) Conduct advance assessments of potential virtual machine migrations or server startups and shutdowns, prioritizing the preparation of sufficient available servers or load offline windows;

[0119] Estimate the predicted load trend Input the updated (or retained) strategy π in step 402 and simulate the scheduling decision for the next period of time in the digital twin environment Ω. If it is found that the upcoming high-load period may trigger a large number of migrations or device startups, batch preheating or resource provisioning is performed in the real system in advance to avoid concentrated operations when approaching high load, and realize forward-looking strategy layout. This forward-looking strategy layout is submitted to the upper-level reinforcement learning controller π together with the execution log. top , and observe whether it further alleviates system fluctuations in the next stage of concept drift detection;

[0120] If the forecast deviation is large, that is, the load trend estimation If the error index between the observed loads during the actual arrival period exceeds a preset threshold, the difference is fed back to the aforementioned time series prediction model, and the prediction parameters are iteratively updated to improve subsequent accuracy.

[0121] By proactively understanding load fluctuations, the system can utilize low-load periods to prepare resources or remove redundant equipment. This proactively avoids emergency dispatch during peak load periods, significantly reducing the risk of sudden migrations and startups and shutdowns during peak periods. When dealing with unsteady loads, the system can not only conduct passive online learning but also implement more proactive strategic planning with the support of predictive data.

[0122] The established online adaptive mechanism ensures rapid response to real-time environmental changes while proactively preventing future fluctuations, achieving a deep integration of monitoring, learning, prediction, and execution. This deep collaboration goes beyond simply stacking concept drift detection, online training, and load forecasting, and instead reinforces each other, maintaining optimal PDU load distribution and energy consumption optimization over the long term. This provides effective technical support for the sustainable operation of large-scale data centers in non-stationary load environments.

[0123] Step 5: When the data center needs to strengthen security protection or improve fault tolerance level, in the multidimensional cost function C(a i ,s t ) Add safety cost and fault tolerance overhead Then the comprehensive cost function C′(a i ,s t ), and prioritize the screening of reinforcement learning strategies through security-related penalties or rewards in offline pre-training, and prioritize redundancy, encryption, or batch migration of key servers or sensitive businesses in the hierarchical scheduling and online adaptation stages, ultimately achieving high-dimensional coordination of data security and redundant fault tolerance while meeting energy consumption and performance requirements.

[0124] The step five includes the following:

[0125] Step 501: Incorporate security and fault tolerance overhead into the multi-dimensional execution cost model

[0126] Inherit the multidimensional cost function C(a i ,s t ), further define the security cost function and fault-tolerance cost function Weigh the execution action a separately i In state s t The degree of impact on the system security strategy and fault tolerance strategy;

[0127] It can represent the additional load brought by encryption, authentication, and isolation operations; It can be used to describe the additional resources consumed by fault-tolerant redundant deployment (such as multi-replica synchronization and dual-machine hot standby for key nodes).

[0128] Incorporate security and fault tolerance costs into the original cost model and define the comprehensive cost function C′(a i ,s t ):

[0129]

[0130] Where: C(ai ,s t ) is the multi-dimensional cost of the original energy consumption, delay, and performance jitter (defined in step 1);

[0131] ρ1, ρ2 are the safety and fault tolerance weight coefficients, both ranging from 0 to 1, which determine the degree of priority amplification for safety and fault tolerance. or When it is large, the fault tolerance weight coefficient It will suppress the probability of choosing high-risk or low-tolerance actions in subsequent decisions;

[0132] Based on the monitoring components in step 1, add the collection of information such as security alerts, vulnerability detection, and key node availability, and update the fault tolerance weight coefficient in real time and fault-tolerance cost function The value of the comprehensive cost function C′(a i ,s t ) reflects the latest security and fault-tolerance status at all times; with the help of the original monitoring mechanism and the newly added security alarm interface, it can realize dynamic tracking of security vulnerabilities, key node redundancy and other status, and can perceive security and fault-tolerance requirements in real time.

[0133] Step 502: Introduce safety / fault tolerance constraints or rewards in offline pre-training and layered scheduling

[0134] In the digital twin environment Ω, the newly defined comprehensive cost function C′(a i ,s t ) for assessment;

[0135] Incorporate negative rewards or positive incentives related to security and fault tolerance into the objective function of reinforcement learning. For example, provide additional positive rewards for necessary security operations (such as high-priority encryption) and impose exponential penalties on unsafe or intolerant actions, so that the agent can learn to balance security and fault tolerance during the offline training phase.

[0136] Upper-level reinforcement learning controller π top When generating a macro load distribution plan, the security status of key servers and the completeness of redundant nodes are given priority: if an action is likely to lead to weak security or insufficient fault tolerance, then based on the comprehensive cost function C′(a i ,s t ) automatically increases its execution cost and reduces scheduling priority;

[0137] The lower-level flexible scheduling module also considers security fault tolerance indicators when performing batch migrations or starting and stopping devices. If the host hosting sensitive services has not yet completed encryption isolation or replica backup, shutdown is delayed and the log is recorded in the execution feedback, allowing for subsequent fine-tuning during online adaptation. This eliminates the need to rebuild independent security policy training. During hierarchical scheduling, more cautious batch or batch scheduling can be implemented for security-critical and fault-tolerant operations, avoiding potential risks to system security and ensuring the execution of high-priority security requirements.

[0138] Step 503: Online adaptive enhancement of security and fault tolerance

[0139] When performing concept drift detection in step 4, additional offsets related to security alerts and key node redundancy status are introduced. For example, when the overall system load is still normal but a high-density network attack or multi-node hardware failure occurs, the offset CD (Δ (s t The weight of the security dimension within )) is increased, so that subsequent online training focuses on security countermeasures.

[0140] If security alarms occur frequently or redundant nodes are insufficient, the fault tolerance overhead function If the value of the training set increases significantly, small-scale online training or transfer learning is enabled in step 4;

[0141] Using a small training dataset Collect recent real execution logs and incorporate them into safety fault tolerance indicators, increase the penalty coefficient for unsafe actions or increase the incentive value for redundant operations, and ensure that the intelligent agent quickly converges to a new strategy that balances safety and performance.

[0142] Combined with the time series prediction model in step 4, if a potential attack peak or extreme fluctuation in critical load is predicted in the future, more redundant servers are provisioned in advance or stricter authentication and isolation measures are enabled, and a temporary security silent window is set up to complete the backup if necessary; the actual execution results are fed back to the digital twin environment Ω again, reserving an interface for subsequent higher-level security or fault-tolerant expansion.

[0143] During use, when sudden changes in the data center environment are primarily due to security or fault tolerance factors, the concept drift detection mechanism will quickly capture the risk, triggering online updates, keeping the strategy and environment evolving in sync, and dynamically adapting to security / fault tolerance requirements. By proactively preventing attacks or failure peaks, critical data isolation or multi-copy deployment can be completed before the main load overheats, improving the overall resilience of the system. Step 501 couples the security and fault tolerance dimensions into the original multi-dimensional execution cost model, allowing all execution actions to quantify their security and fault tolerance costs within the same framework. Step 502 inserts security / fault tolerance-related reward or penalty mechanisms into the offline pre-training and hierarchical scheduling stages (corresponding to steps 2 and 3), allowing the intelligent agent to achieve energy consumption and performance optimization while also taking into account security compliance and fault tolerance stability. Step 503, for the online adaptation scenario (corresponding to step 4), incorporates security and fault tolerance index weights into concept drift detection and small-scale online training, enabling emergency response to high-risk security situations or sudden scenarios with insufficient redundancy, and proactive deployment in conjunction with load forecasting.

[0144] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0145] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0146] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is only for some logical functions. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0147] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0148] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A PDU load distribution system based on reinforcement learning, characterized by: include, When data center load fluctuations are detected, the real-time monitoring component is called to obtain raw monitoring data. A multi-dimensional execution cost model is used to deeply quantify the energy consumption and latency of virtual machine migration and server startup and shutdown, and a dynamic cost sequence is output. After receiving the dynamic cost sequence and raw monitoring data, multiple rounds of reinforcement learning offline training are conducted on the PDU power distribution topology, server cluster, and typical load scenarios in the digital twin environment to select pre-trained strategies for hierarchical scheduling. After successfully loading the pre-trained policy, the upper-level reinforcement learning controller generates macro migration and start-stop instruction sets. The lower-level flexible scheduling module executes them in batches based on multi-dimensional execution costs and safety limits, and records specific scheduling feedback to the system log. If the concept drift metric exceeds the set threshold or the policy benefit drops significantly, a small-scale online learning is performed based on the digital twin environment and the latest scheduling feedback, and a resource allocation plan is deployed in advance using the load forecasting model; Security costs and fault-tolerance overheads are injected into the original multi-dimensional execution cost model, and reinforcement learning strategies are prioritized through security-related penalties or rewards during offline pre-training. During the hierarchical scheduling and online adaptation stages, redundancy, encryption, or batch migration of critical servers or sensitive businesses are prioritized.

2. The PDU load distribution system based on reinforcement learning according to claim 1, characterized in that: The multi-dimensional execution cost model sets independent dimensions for virtual machine migration, server start and stop, and PDU switching, and quantifies the time delay, energy consumption overhead, and business performance impact of each dimension in a weighted manner. It is used to dynamically measure the comprehensive cost of each scheduling action in the current state and write it into the monitoring database.

3. The PDU load distribution system based on reinforcement learning according to claim 2, characterized in that: The deployed monitoring component collects server utilization, PDU load, power startup time and network bandwidth in a fixed time window, and maps the collected results to the cost model in real time to dynamically update the cost data; The available margin threshold is generated based on the bandwidth safety factor and the temperature safety factor, which serves as the basis for subsequent batch scheduling.

4. The PDU load distribution system based on reinforcement learning according to claim 3, characterized in that: The constructed digital twin environment fully maps the PDU power distribution topology, server hardware configuration, network bandwidth limitations and typical business load curves of a real data center, and uses performance incentive coefficients and cost penalty coefficients in offline simulations to guide the reinforcement learning agent to generate pre-training strategies.

5. The PDU load distribution system based on reinforcement learning according to claim 4, characterized in that: Multiple rounds of reinforcement learning simulation training are performed on various load scenarios in the digital twin environment. During offline training, cost sensitivity is adjusted to amplify the negative rewards of high-cost actions. The difference between the new and old strategies is constrained by the strategy stability coefficient, and pre-training strategies and performance evaluation results that take into account both energy efficiency and performance are output.

6. The PDU load distribution system based on reinforcement learning according to claim 5, characterized in that: After the upper-level reinforcement learning controller calls the pre-trained strategy to generate an instruction set, it first filters out the costly instructions based on real-time cost and priority labels, and then passes them to the lower-level module for batch execution according to the scheduling amplification factor and flexibility control parameters. Before generating the instruction set, instructions with costs exceeding a threshold are screened out based on the multi-dimensional execution cost model, and execution priorities are assigned to the remaining instructions.

7. The PDU load distribution system based on reinforcement learning according to claim 6, characterized in that: The lower-level scheduling module executes the instructions in the instruction set in batches according to the bandwidth margin and server temperature margin. After each batch is completed, the scheduling feedback including the actual execution time, power consumption changes and performance jitter is returned to the upper-level controller together with the updated cost vector for use in the online adaptation stage.

8. The PDU load distribution system based on reinforcement learning according to claim 7, characterized in that: Monitor the difference between the actual system state and the expected state to calculate the environmental drift; When the drift degree exceeds a preset threshold, a drift alarm is generated, and the triggering timestamp and main offset indicators are recorded to trigger subsequent online training or transfer learning actions.

9. The PDU load distribution system based on reinforcement learning according to claim 8, characterized in that: The online training process selects several recent scheduling cycles to form a micro-training dataset, uses the learning rate to incrementally update the old strategy to obtain a new strategy, and redeploys it after completing rapid regression testing in the digital twin environment; The micro-training dataset includes the current state, the actions generated by the upper-level reinforcement learning controller, the multi-dimensional execution costs, and the actual benefits.

10. The PDU load distribution system based on reinforcement learning according to claim 9, characterized in that: The system calls on historical load sequences and recent monitoring data, and combines them with the long short-term memory network to predict the load for several future time periods. When the mean absolute percentage error of the prediction deviation exceeds the set threshold, it feeds back to the time series prediction model to iteratively update the parameters, and accordingly pre-sets available computing and power resources.

11. The PDU load distribution system based on reinforcement learning according to claim 10, characterized in that: New security cost parameters and fault tolerance overhead parameters are added to the original multi-dimensional execution cost model to form a comprehensive cost model, in which the priority of key business migration, redundant copy synchronization and sensitive node power management is controlled by the security weight coefficient and the fault tolerance weight coefficient.

12. The PDU load distribution system based on reinforcement learning according to claim 11, characterized in that: During offline simulation, we add security positive rewards and fault-tolerant negative penalties to the reward function, guiding the scheduling strategy to prioritize the secure isolation of sensitive services and high availability redundancy of key nodes. If the host where the sensitive business is located has not completed encryption isolation or copy backup, the shutdown operation will be delayed and the log will be recorded in the execution feedback.

13. The PDU load distribution system based on reinforcement learning according to claim 12, characterized in that: After introducing additional offsets related to safety alarms and the redundant status of key nodes, during online adaptation, the safety and fault tolerance weight coefficients are dynamically adjusted according to the frequency of safety alarms and the number of fault triggers to maintain the preset balance target between energy consumption, performance and reliability, and the actual execution results are fed back to the digital twin environment again.

Citation Information

Patent Citations

  • Data center server power consumption management and optimization method based on reinforcement learning

    CN109324875A

  • Multi-microgrid cooperative scheduling method based on migration reinforcement learning and related device

    CN117526416A

  • Big data dynamic allocation and optimal scheduling method based on reinforcement learning

    CN119311407A

  • Multi-chiller scheduling using reinforcement learning with transfer learning for power consumption prediction

    EP3885850A1

  • Method for automatically regulating explicit congestion notification of data center network based on multi-agent reinforcement learning

    US20240080270A1

Cited By

  • Energy consumption prediction and scheduling control method based on machine learning

    CN120930892A

  • Data center resource dynamic scheduling method based on artificial intelligence

    CN121658186A