Distributed computing-based intelligent design method and system for electric power engineering cloud resources
By building a multi-dimensional resource modeling system and adaptive clustering algorithm, combined with multi-objective optimization and deep learning, elastic scheduling and rapid failure recovery of power engineering cloud resources are achieved, solving the problems of mismatch between resource supply and demand and insufficient response capabilities in the existing technology, and improving the real-time and reliability of the system.
Patent Information
- Application Number
- CN202510611620.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the existing power engineering cloud resource scheduling system faces the multi-dimensional dynamic characteristics of the power system, it is difficult to achieve elastic redistribution of resources, resulting in mismatch between resource supply and demand, insufficient energy waste and dynamic response capabilities, and cannot meet the real-time requirements of the power system.
Build a multi-dimensional resource modeling system, adopt adaptive dynamic clustering algorithm and quantum genetic optimization mechanism, combine multi-objective optimization models and deep learning prediction networks, and deploy anomaly perception and self-healing mechanisms to achieve elastic scheduling of resources and rapid failure recovery.
It improves resource matching accuracy, improves network jitter fault tolerance and fault identification speed, shortens service interruption time, optimizes resource utilization efficiency and system reliability, and meets the real-time requirements of the power system.
Smart Images

Figure CN120494574A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart grid technology, and in particular to a method and system for intelligent design of power engineering cloud resources based on distributed computing. Background Art
[0002] With the advancement of smart grid construction, power engineering cloud resource dispatching systems are required to handle a wide range of heterogeneous tasks, including real-time monitoring, data analysis, and control command issuance. Existing technologies generally employ static or semi-static resource allocation strategies, implementing resource allocation by pre-setting virtual machine specifications and binding them to physical nodes. However, power system operation is characterized by significant time-varying and uncertainties: fluctuations in renewable energy output cause sudden changes in computing load, failures trigger a surge in emergency tasks, and network security threats lead to dynamic changes in resource availability. Traditional static allocation models struggle to adapt to these dynamic demands in real time, leading to two significant issues: a severe mismatch between resource supply and demand; fixed resource pools cannot be elastically expanded, leading to task backlogs during peak load periods, resulting in excessive SCADA control command delays; and a large number of idle resources during low-load periods, resulting in energy waste. Insufficient dynamic response capabilities: When regional node failures caused by typhoons occur, existing technologies require manual intervention to reconfigure resources, resulting in recovery times of several minutes, failing to meet the mandatory requirement of 300ms or less for critical service recovery times as stipulated in the "Regulations on the Safety Protection of Power Monitoring Systems." Of particular note, resource dynamics in power engineering scenarios exhibit multi-dimensional coupling: computational load fluctuations are correlated with grid frequency regulation requirements, cyberattack risk levels vary in real time with security protection status, and equipment reliability is nonlinearly affected by ambient temperature and humidity conditions. Existing technologies lack mechanisms for collaboratively sensing and comprehensively analyzing these multi-dimensional dynamic factors, resulting in a significant disconnect between resource scheduling strategies and real-time operational status.
[0003] Therefore, there is an urgent need for an intelligent scheduling method that can deeply perceive the multi-dimensional dynamic characteristics of the power system and realize flexible resource reallocation based on this, so as to fundamentally solve the problems of inefficient resource utilization and deteriorated service quality caused by the static allocation mode. Summary of the Invention
[0004] The purpose of this invention is to solve the shortcomings of the existing technology and propose an intelligent design method for power engineering cloud resources based on distributed computing, including: S1: Build a multi-dimensional resource modeling system for power engineering projects, dividing computing nodes into three types of heterogeneous resource units: real-time processing cores, data analysis cores, and disaster recovery cores. Establish a dynamic attribute matrix that includes latency sensitivity, energy consumption coefficient, and risk assessment values. The real-time processing cores are equipped with hardware-accelerated instruction sets and support microsecond-level response. S2: An adaptive dynamic clustering algorithm is used to generate flexible resource clusters tailored to task requirements based on the spatial topological relationships of resource units and load variation trends. The clustering process incorporates a quantum genetic optimization mechanism to dynamically adjust the coupling between clusters and establish cross-cluster communication links based on credibility assessment. S3: Establish a resource allocation model based on multi-objective optimization to simultaneously optimize task completion time, energy consumption cost, and system reliability. Cross-node decision-making coordination is achieved through a distributed consensus algorithm. The reliability metric integrates hardware failure rate and network security parameters. S4: Design an intelligent scheduling engine tailored to the characteristics of power engineering projects, integrating a deep spatiotemporal prediction network to predict resource demand fluctuations in real time, dynamically adjust resource allocation strategies, and optimize task queues through a sliding time window mechanism; S5: Deploy an anomaly perception and self-healing mechanism, using multi-level heartbeat detection combined with an improved convolutional neural network anomaly recognition algorithm to achieve rapid isolation of faulty resources and backup and recovery based on hot migration technology.
[0005] This technical solution proposes a distributed computing-based intelligent design method for power engineering cloud resources, encompassing five core steps: constructing a dynamic attribute matrix containing three types of heterogeneous resource units; generating elastic resource clusters using a dynamic clustering algorithm integrated with quantum genetic optimization; establishing a multi-objective optimization resource allocation model; designing an intelligent scheduling engine integrating deep learning prediction; and deploying an anomaly self-healing mechanism based on multi-level detection. This approach addresses the technical shortcomings of traditional power engineering resource scheduling, such as the inability of static resource partitioning to adapt to dynamic load changes; the difficulty in balancing energy consumption and reliability due to single-objective optimization; and the reliance on manual intervention for fault recovery, which cannot meet the real-time requirements of power systems.
[0006] Preferably, the method for constructing the dynamic attribute matrix in S1 includes: S11: Delay Sensitivity It is calculated as the square of the ratio of the maximum allowed delay of the task to the actual transmission delay, and is dynamically calibrated according to the number of hops and bandwidth utilization of the transmission path; S12: Energy consumption coefficient The energy consumption per unit computing amount using dynamic voltage and frequency adjustment technology is calculated, and a correction factor for the effect of ambient temperature on chip leakage current is introduced; S13: Risk Assessment Value It is calculated based on the node failure history, current load rate and environmental parameters, including the temperature and humidity of the computer room, the remaining battery life of the UPS and the matching degree of the network attack signature library.
[0007] Adopting the above technical solution: This solution can solve the shortcomings of existing resource modeling in existing technologies by refining the dynamic attribute matrix construction method, including: delay evaluation only considers transmission delay and ignores the impact of network topology hop count; energy consumption calculation does not consider the nonlinear relationship between chip leakage current and ambient temperature.
[0008] Further preferably, the adaptive dynamic clustering algorithm in S2 includes: S21: Based on the initial clustering of the improved firefly algorithm, the Lévy flight mechanism is introduced to enhance the global search capability, and the number of population iterations is set to be proportional to the square root of the node size; S22: Inter-cluster communication establishes a topological connection based on credibility evaluation, with a credibility threshold of T_h = μ + 3σ, where μ is the mean of the historical communication success rate and σ is the standard deviation, and a weighted addition coefficient is set for the encrypted communication link; S23: The clustering stability coefficient is calculated by taking into account node mobility and load change rate. A sliding window mechanism is used to update the cluster structure, and the window size is adaptively adjusted according to the network jitter amplitude.
[0009] This technical solution, which provides an adaptive dynamic clustering algorithm encompassing an improved firefly algorithm, trustworthy topological connections, and a sliding window update mechanism, can address technical issues with traditional clustering algorithms in power engineering scenarios, including the local optimal solution caused by fixed step sizes and the difficulty of static topologies in adapting to the intermittent high jitter of power communication networks.
[0010] Further preferably, the multi-objective optimization model in S3 uses a decomposition evolutionary algorithm to process the coupling relationship between objectives, including: S31: Establish a three-dimensional coordinate space mapping of the Pareto front solution set, where each coordinate axis corresponds to task delay, energy consumption cost, and reliability loss respectively; S32: Design a solution set distribution evaluation mechanism based on KL divergence to ensure the diversity of optimization results in the solution space; S33: Monte Carlo simulation is introduced to predict the long-term effects of different resource allocation strategies, with the number of simulations proportional to the logarithm of the system size.
[0011] This technical solution proposes a specific implementation method for a multi-objective optimization model, employing a decomposition-based evolutionary algorithm to address the coupled relationship between task completion time, energy cost, and system reliability. This solution addresses two major bottlenecks of traditional multi-objective optimization methods: Inadequate objective conflict resolution: The weighted summation method cannot effectively balance the conflict between latency-sensitive tasks and high energy consumption requirements, resulting in a single-dimensional optimization result. Existing models (such as NSGA-II) focus solely on short-term optimization and ignore the long-term impact of resource allocation strategies on power equipment lifespan and grid stability.
[0012] Further preferably, the objective function of the resource allocation model in S3 is: , in, represents the expected completion time of the i-th type of task; represents the energy consumption cost of the jth computing node; is the system reliability index; it is calculated as the geometric mean of the reliability of each node; 、 as well as is the dynamic weight coefficient, satisfying .
[0013] This technical solution, which integrates the total task completion time, total energy cost, and system reliability loss into the objective function through dynamic weight coefficients, addresses technical flaws in existing resource allocation models, including the problem of fixed weights (α=β=γ=1 / 3) that are incapable of adapting to sudden failures or load surges in power projects. It also addresses the problem of one-sided reliability modeling, which only considers hardware failure rates and ignores the impact of cyberattacks on system reliability.
[0014] Further preferably, the load balancing calculation formula used by the intelligent scheduling engine in S4 is: , in, is the difference between the current load and the predicted load; Indicates the real-time computing capacity of the nth node; Calculate the average capacity for the system; is the load sensitivity coefficient, and its value is [0.5,1.2].
[0015] The above technical solution limits the load balancing formula used by the intelligent scheduling engine, which can solve two technical problems existing in traditional load balancing methods. These include: static capacity assessment, which only uses CPU utilization as a load indicator and does not consider the dedicated computing power of heterogeneous computing units (such as FPGA accelerators); prediction and real-time disconnection, and the fixed threshold trigger mechanism leads to delayed resource allocation, resulting in a load imbalance rate of 45% in wind farm power fluctuation scenarios.
[0016] Further preferably, the anomaly recognition algorithm in S5 adopts a fault detection model of an improved convolutional neural network: ; in, represents discrete wavelet transform feature extraction; is the convolution kernel weight of the lth layer; is the bias term; is the ReLU activation function; L is 5-7 layers.
[0017] This technical solution, which establishes an improved convolutional neural network fault detection model, addresses the shortcomings of traditional anomaly detection methods, including: Single feature extraction: Relying solely on time-domain signal analysis, the method misses transient faults by up to 34%. Insufficient model generalization: Fixed threshold detection results in a false alarm rate exceeding 40% in novel network attack scenarios, such as false data injection.
[0018] A system, applied to a distributed computing-based intelligent design method for power engineering cloud resources as described in any one of the above, comprising: The resource modeling module is equipped with a heterogeneous resource classifier and a dynamic attribute collection unit. The classifier uses a support vector machine to identify node types, and the collection unit integrates a temperature sensor and a network sniffer. Cluster building module, integrating quantum genetic optimizer and dynamic clustering controller. The optimizer is equipped with qubit rotation gate adjustment strategy, and the controller supports minimum spanning tree topology construction. The optimization decision-making module includes a multi-objective solution engine and a distributed negotiation agent. The solution engine has a built-in NSGA-II algorithm framework, and the negotiation agent uses the Raft consensus protocol. The intelligent scheduling module has a built-in deep prediction network and elastic scaling manager. The prediction network uses a hybrid architecture of LSTM and graph convolution, and the manager implements resource scaling at the virtual machine granularity. The fault-tolerant control module is deployed with multi-level detection probes and self-healing executors. The detection probes cover the hardware layer, virtualization layer, and application layer. The self-healing executor supports hot migration and snapshot rollback.
[0019] The intelligent design system proposed in this application, comprising five core modules, addresses architectural flaws in existing systems, including isolated module operations; the separation of resource modeling and scheduling decisions, which results in cluster reconstruction delays exceeding 500ms; incomplete detection coverage; and the inability of single-layer probes to identify the associated risks of hardware firmware vulnerabilities and virtualization escape attacks.
[0020] More preferably, the resource modeling module further comprises: A three-dimensional spatial encoder maps resource attributes to an orthogonal feature space and uses principal component analysis dimensionality reduction technology to eliminate attribute redundancy; Dynamic association analyzer, detects implicit dependencies between resources and builds dependency graphs based on Granger causality tests; The risk assessment unit integrates the Bayesian network prediction model, and the input layer includes voltage fluctuation rate and firewall log analysis results.
[0021] The aforementioned technical solution: The enhanced design of the resource modeling module provided by this solution can overcome the limitations of traditional modeling tools. The "curse of dimensionality": Direct input of high-dimensional attributes causes the convergence speed of the clustering algorithm to decrease by 83%. The "risk assessment lag": Static scoring based on the rule engine cannot reflect the APT attack penetration process in real time.
[0022] More preferably, the intelligent scheduling module further includes: The task priority classifier uses the improved ABC classification method to divide the emergency task levels and adds the power system transient stability coefficient as a classification dimension; The resource reservation manager establishes a dynamic reservation pool based on a sliding time window, and the pool capacity is adjusted at hourly granularity based on load forecast results; The energy efficiency optimizer implements dynamic voltage and frequency adjustment and task migration collaborative optimization, and the migration decision takes into account the cross-cabinet communication overhead.
[0023] The above technical solution addresses the problems of traditional solutions, such as crude task classification; prioritization based solely on service type, without considering the real-time stability of the power grid; and rigid reservation strategies; fixed reservation ratios, resulting in insufficient backup resources in extreme weather conditions and wasted resources in everyday scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 This is a flow chart of the intelligent design method for power engineering cloud resources based on distributed computing in this application; Figure 2 This is a block diagram of the power engineering cloud resource intelligent design system based on distributed computing for this application. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0026] See also Figure 1 Traditional power engineering resource scheduling has the following technical flaws: static resource partitioning cannot adapt to dynamic load changes; single-objective optimization makes it difficult to balance energy consumption and reliability, and fault recovery relies on manual intervention; and the response speed cannot meet the real-time requirements of the power system. Based on this, this application provides a distributed computing-based intelligent design method for power engineering cloud resources, including: S1: Build a multi-dimensional resource modeling system for power engineering projects, dividing computing nodes into three types of heterogeneous resource units: real-time processing cores, data analysis cores, and disaster recovery cores. Establish a dynamic attribute matrix that includes latency sensitivity, energy consumption coefficient, and risk assessment values. The real-time processing cores are equipped with hardware-accelerated instruction sets and support microsecond-level response. S2: An adaptive dynamic clustering algorithm is used to generate flexible resource clusters tailored to task requirements based on the spatial topological relationships of resource units and load variation trends. The clustering process incorporates a quantum genetic optimization mechanism to dynamically adjust the coupling between clusters and establish cross-cluster communication links based on credibility assessment. S3: Establish a resource allocation model based on multi-objective optimization to simultaneously optimize task completion time, energy consumption cost, and system reliability. Cross-node decision-making coordination is achieved through a distributed consensus algorithm. The reliability metric integrates hardware failure rate and network security parameters. S4: Design an intelligent scheduling engine tailored to the characteristics of power engineering projects, integrating a deep spatiotemporal prediction network to predict resource demand fluctuations in real time, dynamically adjust resource allocation strategies, and optimize task queues through a sliding time window mechanism; S5: Deploy an anomaly perception and self-healing mechanism, using multi-level heartbeat detection combined with an improved convolutional neural network anomaly recognition algorithm to achieve rapid isolation of faulty resources and backup and recovery based on hot migration technology.
[0027] It is worth mentioning that this application proposes an intelligent design method for power engineering cloud resources based on distributed computing, which includes five core steps: constructing a dynamic attribute matrix containing three types of heterogeneous resource units; using a dynamic clustering algorithm integrated with quantum genetic optimization to generate elastic resource clusters; establishing a multi-objective optimization resource allocation model; designing an intelligent scheduling engine integrated with deep learning prediction; and deploying an abnormal self-healing mechanism based on multi-level detection.
[0028] The technical effects of this embodiment include: being able to achieve multi-dimensional quantification of resource characteristics through a dynamic attribute matrix (latency sensitivity, energy consumption coefficient, risk assessment value), and improving resource matching accuracy compared to the traditional CPU / memory single indicator division method.
[0029] The clustering convergence speed of the quantum genetic algorithm is improved, and the credibility threshold design of the inter-cluster communication link increases the network jitter tolerance rate to 99.7%.
[0030] The anomaly detection algorithm achieves 200ms-level fault identification under GPU acceleration, and hot migration technology reduces service interruption time to less than 50ms.
[0031] Existing resource modeling has two major shortcomings: latency assessment only considers transmission delay and ignores the impact of network topology hop count; energy consumption calculation does not consider the nonlinear relationship between chip leakage current and ambient temperature. Based on this, the construction method of the dynamic attribute matrix described in S1 includes: S11: Delay Sensitivity It is calculated as the square of the ratio of the maximum allowed delay of the task to the actual transmission delay, and is dynamically calibrated according to the number of hops and bandwidth utilization of the transmission path; S12: Energy consumption coefficient The energy consumption per unit computing amount using dynamic voltage and frequency adjustment technology is calculated, and a correction factor for the effect of ambient temperature on chip leakage current is introduced; S13: Risk Assessment Value It is calculated based on the node failure history, current load rate and environmental parameters, including the temperature and humidity of the computer room, the remaining battery life of the UPS and the matching degree of the network attack signature library.
[0032] Adopting the above technical solution: This solution can solve the shortcomings of existing resource modeling in existing technologies by refining the dynamic attribute matrix construction method, including: delay evaluation only considers transmission delay and ignores the impact of network topology hop count; energy consumption calculation does not consider the nonlinear relationship between chip leakage current and ambient temperature.
[0033] It is worth mentioning that this embodiment refines the dynamic attribute matrix construction method, including: Delay sensitivity = (maximum allowed delay / actual transmission delay)² + hop count calibration item Energy consumption coefficient =DVFS basic energy consumption × temperature correction factor Risk Assessment Value =Failure rate × load rate × environmental risk factor.
[0034] After the hop count calibration item is introduced in the design of the above embodiment, the delay prediction error is reduced from 15.3% to 3.8% in the cross-regional data center scenario.
[0035] The temperature correction factor is constructed based on the Arrhenius equation, which can improve the accuracy of energy consumption calculation during the high temperature period in summer.
[0036] The environmental risk factor combines UPS battery life and network attack feature matching, successfully identifying 97.2% of APT attacks in power grid attack and defense drills.
[0037] For example, the traditional K-means clustering algorithm has two major problems in power engineering scenarios: the fixed step size leads to local optimal solutions; and the static topology is difficult to adapt to the intermittent high jitter of the power communication network. Based on this, the adaptive dynamic clustering algorithm described in S2 includes: S21: Based on the initial clustering of the improved firefly algorithm, the Lévy flight mechanism is introduced to enhance the global search capability, and the number of population iterations is set to be proportional to the square root of the node size; S22: Inter-cluster communication establishes a topological connection based on credibility evaluation, with a credibility threshold of T_h = μ + 3σ, where μ is the mean of the historical communication success rate and σ is the standard deviation, and a weighted addition coefficient is set for the encrypted communication link; S23: The clustering stability coefficient is calculated by taking into account node mobility and load change rate. A sliding window mechanism is used to update the cluster structure, and the window size is adaptively adjusted according to the network jitter amplitude.
[0038] It is worth mentioning that the adaptive dynamic clustering algorithm designed in this solution includes an improved firefly algorithm, a trustworthy topological connection, and a sliding window update mechanism. Levi's flight step length adjustment formula: ; Confidence threshold , is the historical average success rate, is the standard deviation.
[0039] Sliding window size .
[0040] The technical effects of this embodiment include: improving the global search capability, and the Levy flight mechanism increases the optimal solution coverage rate of the algorithm at a scale of 1000 nodes from 78.4% to 95.2%.
[0041] The communication reliability is guaranteed. The credibility threshold design based on the 3σ principle can maintain 92.3% effective communication even in a network environment with a 15% packet loss rate.
[0042] Dynamic adaptability is enhanced, and the sliding window mechanism reduces the cluster structure update delay to 50ms, compared to 200ms required by traditional periodic update solutions. The cluster reconfiguration success rate is increased to 99.1% in wind farm power fluctuation scenarios.
[0043] Traditional multi-objective optimization methods have the following technical problems: insufficient handling of objective conflicts; the weighted summation method cannot effectively balance the contradiction between delay-sensitive tasks and high energy consumption requirements, resulting in optimization results that are biased towards a single dimension; and the lack of long-term effect prediction; existing models only focus on short-term optimization and ignore the long-term impact of resource allocation strategies on power equipment life and grid stability. Based on this, the multi-objective optimization model described in S3 uses a decomposition evolutionary algorithm to handle the coupling relationship between objectives, including: S31: Establish a three-dimensional coordinate space mapping of the Pareto front solution set, where each coordinate axis corresponds to task delay, energy consumption cost, and reliability loss respectively; S32: Design a solution set distribution evaluation mechanism based on KL divergence to ensure the diversity of optimization results in the solution space; S33: Monte Carlo simulation is introduced to predict the long-term effects of different resource allocation strategies, with the number of simulations proportional to the logarithm of the system size.
[0044] It is worth mentioning that this embodiment proposes a specific implementation method of a multi-objective optimization model, which uses a decomposition evolutionary algorithm to handle the coupling relationship between task completion time, energy consumption cost, and system reliability. It specifically includes three core steps: Three-dimensional coordinate space mapping: Map the Pareto front solution set to the three-dimensional space consisting of task delay, energy consumption cost, and reliability loss, and define the optimization boundaries of each objective through the coordinate axis.
[0045] KL divergence evaluation mechanism: Kullback-Leibler divergence is used to measure the uniformity of the distribution of the solution set in space, ensuring that the optimization results cover a variety of trade-off scenarios.
[0046] Monte Carlo simulation prediction: The number of simulations is set based on the logarithmic relationship of the system scale to predict the impact of different resource allocation strategies on the long-term operation of the power grid, such as the effects of equipment aging and load accumulation.
[0047] The technical effects of this embodiment include: improving the accuracy of multi-dimensional optimization: in the IEEE 39-node test, the three-dimensional coordinate mapping method increased the Pareto solution set coverage from 72.5% to 93.8%, and the solution set diversity index HV value increased by 0.21.
[0048] Reliability is guaranteed. After combining Monte Carlo simulation with the equipment aging model, the error in predicting the system failure rate within 5 years is reduced to 3.2%, while the error of traditional methods is 18.7%.
[0049] Computational efficiency has been optimized, and the KL divergence evaluation mechanism has reduced the number of algorithm iterations by 42%. In a power grid scenario with more than 1,000 nodes, the optimization time is shortened to 8.3 minutes, while the original solution required 14.6 minutes.
[0050] Further preferably, the objective function of the resource allocation model in S3 is: , in, represents the expected completion time of the i-th type of task; represents the energy consumption cost of the jth computing node; is the system reliability index; it is calculated as the geometric mean of the reliability of each node; 、 as well as is the dynamic weight coefficient, satisfying .
[0051] This formula is a linear weighted sum expression for a multi-objective optimization problem, using dynamic weight coefficients. The three conflicting objectives (task completion time, energy cost, and system reliability) are integrated into a single optimization objective.
[0052] In the above formula: The total time to complete the task. It is defined as the difference between the maximum allowed delay of the i-th task and the expected delay under the actual resource allocation policy. The difference is designed to penalize task allocation schemes that exceed the SLA constraints.
[0053] is the total energy consumption cost. The calculation is based on the dynamic voltage frequency scaling (DVFS) technology, combined with the real-time data of the node temperature sensor. The specific formula is:
[0054] in is the static power consumption, is the calculation efficiency factor, is the ambient temperature. The square term of temperature reflects the strong nonlinear relationship between chip leakage current and temperature.
[0055] is the system reliability loss term. The geometric mean method is used to integrate the reliability of each node. The specific calculation formula is: , The design is more sensitive to low-reliability nodes than the geometric mean method, ensuring that the resource allocation scheme does not cause systemic collapse due to a single point of failure.
[0056] Dynamic weight design: The weight coefficients α, β, and γ are dynamically adjusted according to the grid operation status: α: During load peaks or transient instability, the α weight is increased to 0.6-0.8 to prioritize task timeliness.
[0057] β: During periods of low electricity prices or high penetration of renewable energy, the β weight increases to 0.4-0.5 to reduce operating costs.
[0058] γ: When a network attack is detected or the hardware failure rate exceeds the standard, the γ weight is automatically increased to enhance system reliability.
[0059] The design rules of the above weight coefficients are implemented based on a fuzzy logic controller, and the input variables include load fluctuation rate, electricity price signal and security situation score.
[0060] It's worth noting that the objective function of this embodiment integrates the total task completion time, total energy consumption cost, and system reliability loss through dynamic weight coefficients. This addresses technical flaws in existing resource allocation models, including the problem of fixed weights (α=β=γ=1 / 3) that cannot adapt to sudden failures or load surges in power projects. It also addresses the problem of one-sided reliability modeling, which only considers hardware failure rates and ignores the impact of cyberattacks on system reliability.
[0061] The load balancing calculation formula used by the intelligent scheduling engine in S4 is: , in, is the difference between the current load and the predicted load; Indicates the real-time computing capacity of the nth node; Calculate the average capacity for the system; is the load sensitivity coefficient, and its value is [0.5,1.2].
[0062] In the above formula: Sigmoid function term: Used to dynamically adjust load sensitivity, where Indicates the deviation of the current load from the predicted load.
[0063] is the load sensitivity coefficient, with a value range of [0.5, 1.2], and is dynamically configured according to the task type: Real-time monitoring tasks (k=1.2): Respond quickly to load fluctuations and avoid task accumulation.
[0064] Offline analysis tasks (k=0.5): Allow moderate load imbalance to save migration overhead.
[0065] The function output range is (0,1). , the load exceeds expectations, the output is close to 1, triggering resource expansion; when When the output approaches 0, the energy saving mode is enabled.
[0066] Standard deviation term: Compute node real-time capacity and the system mean The Euclidean distance of represents the dispersion of load distribution.
[0067] in The calculation formula is: , The weight , , The above formula reflects the contribution of CPU utilization (U), memory bandwidth (B), and IO throughput (IOPS) to the overall computing power.
[0068] The square design of the standard deviation term amplifies the negative impact of high-load nodes, forcing the scheduling algorithm to prioritize balancing hot spots.
[0069] It is worth mentioning that the above solution limits the load balancing formula used by the intelligent scheduling engine, which can solve two technical problems existing in traditional load balancing methods. These include: static capacity assessment, which only uses CPU utilization as a load indicator and does not consider the dedicated computing power of heterogeneous computing units (such as FPGA accelerators); prediction and real-time disconnection, and the fixed threshold trigger mechanism leads to delayed resource allocation, resulting in a load imbalance rate of 45% in wind farm power fluctuation scenarios.
[0070] The anomaly recognition algorithm described in S5 adopts an improved convolutional neural network fault detection model: ; in, represents discrete wavelet transform feature extraction; is the convolution kernel weight of the lth layer; is the bias term; is the ReLU activation function; L is 5-7 layers.
[0071] In the above formula: Wavelet transform layer: : Perform 3-layer discrete wavelet decomposition on the input signal x to extract time-frequency domain features: Layer 1: Uses Daubechies4 wavelet basis to capture high-frequency transient faults such as voltage sag and lightning overvoltage.
[0072] Layer 2: Use the Symlet5 wavelet basis to extract intermediate frequency oscillation components, such as subsynchronous oscillations and harmonic interference.
[0073] Layer 3: Apply the Haar wavelet basis to analyze low-frequency trend items, such as the progressive degradation caused by equipment aging.
[0074] Convolutional Neural Networks: Represents the weight of the convolution kernel in layer l, which is optimized through adversarial training: Adversarial sample generation: adding Gaussian noise to the original data and pulse interference (amplitude ±20%) to improve the robustness of the model.
[0075] Loss function design: cross entropy loss + feature similarity loss, which can achieve a balance between classification accuracy and waveform fidelity.
[0076] L: The network depth is set to 5-7 layers. The deep network (L=7) is used for complex fault diagnosis (such as multiple harmonic superposition), and the shallow network (L=5) meets the real-time requirements.
[0077] It is the ReLU activation function, which solves the gradient disappearance problem. The calculation formula is: .
[0078] The dynamic adjustment mechanism of the resource allocation model weight coefficients α, β, and γ in the above formula and the load balancing formula designed in this embodiment can form a closed-loop control: When α increases, the k value in the design of this embodiment is automatically adjusted higher to speed up the hotspot migration.
[0079] When γ increases, the weight of the standard deviation term is strengthened in the load balancing calculation to avoid node overload.
[0080] The output of this design Can be used as Input parameters: If an abnormality is detected , immediately triggering the increase of reliability weight γ and freezing the resource allocation of suspicious nodes.
[0081] It's worth noting that this solution, through the establishment of an improved convolutional neural network fault detection model, addresses the shortcomings of traditional anomaly detection methods, including: Single feature extraction: Relying solely on time-domain signal analysis, the missed detection rate for transient faults reaches 34%. Insufficient model generalization: Fixed threshold detection results in a false alarm rate exceeding 40% in new network attack scenarios (such as False Data Injection).
[0082] See also Figure 2 , for example, existing systems have architectural flaws including: Modules operate in isolation, resource modeling is separated from scheduling decisions, resulting in cluster reconstruction delays exceeding 500ms. Detection coverage is incomplete, and a single-layer probe cannot identify the associated risks of hardware firmware vulnerabilities and virtualization escape attacks. Based on this, the present application provides a system for use in a distributed computing-based intelligent design method for power engineering cloud resources as described in any of the above, including: The resource modeling module is equipped with a heterogeneous resource classifier and a dynamic attribute collection unit. The classifier uses a support vector machine to identify node types, and the collection unit integrates a temperature sensor and a network sniffer. Cluster building module, integrating quantum genetic optimizer and dynamic clustering controller. The optimizer is equipped with qubit rotation gate adjustment strategy, and the controller supports minimum spanning tree topology construction. The optimization decision-making module includes a multi-objective solution engine and a distributed negotiation agent. The solution engine has a built-in NSGA-II algorithm framework, and the negotiation agent uses the Raft consensus protocol. The intelligent scheduling module has a built-in deep prediction network and elastic scaling manager. The prediction network uses a hybrid architecture of LSTM and graph convolution, and the manager implements resource scaling at the virtual machine granularity. The fault-tolerant control module is deployed with multi-level detection probes and self-healing executors. The detection probes cover the hardware layer, virtualization layer, and application layer. The self-healing executor supports hot migration and snapshot rollback.
[0083] It is worth mentioning that the intelligent design system provided by this patent includes five core modules: Resource Modeling Module: Identifies node types based on a support vector machine (SVM) classifier and integrates temperature and humidity sensors and a network sniffer to collect dynamic attributes.
[0084] Cluster building module: The quantum genetic optimizer uses a qubit rotation gate strategy to adjust population genes, and a dynamic clustering controller generates a minimum spanning tree topology.
[0085] Optimization decision module: The NSGA-II algorithm framework generates Pareto solutions, and the Raft protocol implements distributed inter-node policy negotiation.
[0086] Smart Scheduling Module: The LSTM-GCN hybrid network predicts load fluctuations, and the elastic scaling manager adjusts resources at the virtual machine level.
[0087] Fault-Tolerance Control Module: Probes at the hardware layer (BMC chip), virtualization layer (Hypervisor log), and application layer (API call chain) provide full coverage detection.
[0088] It can achieve end-to-end optimization, and pipeline collaboration between modules can reduce resource allocation decision-making delays. It also provides full-stack protection, and multi-level probes successfully intercepted 96.5% of cross-layer attacks in power grid attack and defense drills. It also ensures elastic scalability, supports access to 1,000+ nodes in minutes, and achieved a linear growth slope of 0.98 in provincial power grid expansion tests.
[0089] For example, traditional modeling tools have the following limitations: Direct input of high-dimensional attributes causes the convergence speed of the clustering algorithm to decrease by 83%. Risk assessment is delayed, and static scoring based on the rule engine cannot reflect the APT attack penetration process in real time. Based on this, the resource modeling module also includes: A three-dimensional spatial encoder maps resource attributes to an orthogonal feature space and uses principal component analysis dimensionality reduction technology to eliminate attribute redundancy; Dynamic association analyzer, detects implicit dependencies between resources and builds dependency graphs based on Granger causality tests; The risk assessment unit integrates the Bayesian network prediction model, and the input layer includes voltage fluctuation rate and firewall log analysis results.
[0090] It is worth mentioning that the enhanced design of the resource modeling module provided in this embodiment includes: Three-dimensional Space Encoder: Reduces 12-dimensional resource attributes to a 3-dimensional orthogonal space through principal component analysis (PCA), eliminating attribute redundancy.
[0091] Dynamic Correlation Analyzer: Builds a resource dependency graph based on Granger causality tests to identify implicit relationships (such as the correlation between storage IO and computing tasks).
[0092] Risk Assessment Unit: The Bayesian network inputs voltage fluctuation rate, firewall logs, and industrial control protocol compliance data, and outputs a node risk score.
[0093] It can improve dimensionality reduction efficiency: PCA processing reduces the number of clustering algorithm iterations by 65%, and reduces modeling time from 58 seconds to 19 seconds at a 2000-node scale.
[0094] Relationship Mining: Granger causality test successfully identified 92.3% of implicit dependencies.
[0095] Dynamic Risk Warning: Bayesian Networks enable the detection window of new ransomware attacks to be advanced to 8 minutes after intrusion. For example, traditional scheduling schemes prioritize services only by type, without considering the real-time stability of the power grid. Fixed reservation ratios lead to insufficient backup resources in extreme weather conditions or waste of resources in daily scenarios. Based on this, the intelligent scheduling module further includes: The task priority classifier uses the improved ABC classification method to divide the emergency task levels and adds the power system transient stability coefficient as a classification dimension; The resource reservation manager establishes a dynamic reservation pool based on a sliding time window, and the pool capacity is adjusted at hourly granularity based on load forecast results; The energy efficiency optimizer implements dynamic voltage and frequency adjustment and task migration collaborative optimization, and the migration decision takes into account the cross-cabinet communication overhead.
[0096] Adopting the above technical solution: the task priority classifier designed in this embodiment improves the ABC classification method and adds the transient stability coefficient as an emergency task judgment dimension.
[0097] Resource Reservation Manager: Dynamically adjusts the reserved pool capacity using a sliding time window, synchronizing load forecast results with meteorological data at hourly granularity.
[0098] Energy Efficiency Optimizer: Collaboratively implements DVFS voltage regulation and virtual machine migration, and migration decisions take into account the power consumption and latency costs of optical modules across cabinets.
[0099] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for intelligent design of power engineering cloud resources based on distributed computing, characterized in that: include: S1: Build a multi-dimensional resource modeling system for power engineering projects, dividing computing nodes into three types of heterogeneous resource units: real-time processing cores, data analysis cores, and disaster recovery cores. Establish a dynamic attribute matrix that includes latency sensitivity, energy consumption coefficient, and risk assessment values. The real-time processing cores are equipped with hardware-accelerated instruction sets and support microsecond-level response. S2: An adaptive dynamic clustering algorithm is used to generate flexible resource clusters tailored to task requirements based on the spatial topological relationships of resource units and load variation trends. The clustering process incorporates a quantum genetic optimization mechanism to dynamically adjust the coupling between clusters and establish cross-cluster communication links based on credibility assessment. S3: Establish a resource allocation model based on multi-objective optimization to simultaneously optimize task completion time, energy consumption cost, and system reliability. Cross-node decision-making coordination is achieved through a distributed consensus algorithm. The reliability metric integrates hardware failure rate and network security parameters. S4: Design an intelligent scheduling engine tailored to the characteristics of power engineering projects, integrating a deep spatiotemporal prediction network to predict resource demand fluctuations in real time, dynamically adjust resource allocation strategies, and optimize task queues through a sliding time window mechanism; S5: Deploy an anomaly perception and self-healing mechanism, using multi-level heartbeat detection combined with an improved convolutional neural network anomaly recognition algorithm to achieve rapid isolation of faulty resources and backup and recovery based on hot migration technology.
2. The method for intelligent design of power engineering cloud resources based on distributed computing according to claim 1, characterized in that: The method for constructing the dynamic attribute matrix described in S1 includes: S11: Delay Sensitivity It is calculated as the square of the ratio of the maximum allowed delay of the task to the actual transmission delay, and is dynamically calibrated according to the number of hops and bandwidth utilization of the transmission path; S12: Energy consumption coefficient The energy consumption per unit computing amount using dynamic voltage and frequency adjustment technology is calculated, and a correction factor for the effect of ambient temperature on chip leakage current is introduced; S13: Risk Assessment Value It is calculated based on the node failure history, current load rate and environmental parameters, including the temperature and humidity of the computer room, the remaining battery life of the UPS and the matching degree of the network attack signature library.
3. The method for intelligent design of power engineering cloud resources based on distributed computing according to claim 1, characterized in that: The adaptive dynamic clustering algorithm described in S2 includes: S21: Based on the initial clustering of the improved firefly algorithm, the Lévy flight mechanism is introduced to enhance the global search capability, and the number of population iterations is set to be proportional to the square root of the node size; S22: Inter-cluster communication establishes a topological connection based on credibility evaluation, with a credibility threshold of T_h = μ + 3σ, where μ is the mean of the historical communication success rate and σ is the standard deviation, and a weighted addition coefficient is set for the encrypted communication link; S23: The clustering stability coefficient is calculated by taking into account node mobility and load change rate. A sliding window mechanism is used to update the cluster structure, and the window size is adaptively adjusted according to the network jitter amplitude.
4. The method for intelligent design of power engineering cloud resources based on distributed computing according to claim 1, characterized in that: The multi-objective optimization model described in S3 uses a decomposition evolutionary algorithm to handle the coupling relationship between objectives, including: S31: Establish a three-dimensional coordinate space mapping of the Pareto front solution set, where each coordinate axis corresponds to task delay, energy consumption cost, and reliability loss respectively; S32: Design a solution set distribution evaluation mechanism based on KL divergence to ensure the diversity of optimization results in the solution space; S33: Monte Carlo simulation is introduced to predict the long-term effects of different resource allocation strategies, with the number of simulations proportional to the logarithm of the system size.
5. The method for intelligent design of power engineering cloud resources based on distributed computing according to claim 1, characterized in that: The objective function of the resource allocation model in S3 is: , in, represents the expected completion time of the i-th type of task; represents the energy consumption cost of the jth computing node; is the system reliability index; it is calculated as the geometric mean of the reliability of each node; 、 as well as is the dynamic weight coefficient, satisfying .
6. The method for intelligent design of power engineering cloud resources based on distributed computing according to claim 1, characterized in that: The load balancing calculation formula used by the intelligent scheduling engine in S4 is: , in, is the difference between the current load and the predicted load; Indicates the real-time computing capacity of the nth node; Calculate the average capacity for the system; is the load sensitivity coefficient, and its value is [0.5,1.2].
7. The method for intelligent design of power engineering cloud resources based on distributed computing according to claim 1, characterized in that: The anomaly recognition algorithm described in S5 adopts an improved convolutional neural network fault detection model: ; in, represents discrete wavelet transform feature extraction; is the convolution kernel weight of the lth layer; is the bias term; is the ReLU activation function; L is 5-7 layers.
8. A system, applied to the distributed computing-based intelligent design method for power engineering cloud resources according to any one of claims 1 to 7, characterized in that: include: The resource modeling module is equipped with a heterogeneous resource classifier and a dynamic attribute collection unit. The classifier uses a support vector machine to identify node types, and the collection unit integrates a temperature sensor and a network sniffer. Cluster building module, integrating quantum genetic optimizer and dynamic clustering controller. The optimizer is equipped with qubit rotation gate adjustment strategy, and the controller supports minimum spanning tree topology construction. The optimization decision-making module includes a multi-objective solution engine and a distributed negotiation agent. The solution engine has a built-in NSGA-II algorithm framework, and the negotiation agent uses the Raft consensus protocol. The intelligent scheduling module has a built-in deep prediction network and elastic scaling manager. The prediction network uses a hybrid architecture of LSTM and graph convolution, and the manager implements resource scaling at the virtual machine granularity. The fault-tolerant control module is deployed with multi-level detection probes and self-healing executors. The detection probes cover the hardware layer, virtualization layer, and application layer. The self-healing executor supports hot migration and snapshot rollback.
9. A system according to claim 8, characterized in that: The resource modeling module also includes: A three-dimensional spatial encoder maps resource attributes to an orthogonal feature space and uses principal component analysis dimensionality reduction technology to eliminate attribute redundancy; Dynamic association analyzer, detects implicit dependencies between resources and builds dependency graphs based on Granger causality tests; The risk assessment unit integrates the Bayesian network prediction model, and the input layer includes voltage fluctuation rate and firewall log analysis results.
10. A system according to claim 8, characterized in that: The intelligent scheduling module further includes: The task priority classifier uses the improved ABC classification method to divide the emergency task levels and adds the power system transient stability coefficient as a classification dimension; The resource reservation manager establishes a dynamic reservation pool based on a sliding time window, and the pool capacity is adjusted at hourly granularity based on load forecast results; The energy efficiency optimizer implements dynamic voltage and frequency adjustment and task migration collaborative optimization, and the migration decision takes into account the cross-cabinet communication overhead.
Citation Information
Cited By
Resource management method and system for green cloud computing and storage medium
CN120832248A
Data center power distribution optimization method and system
CN120930089A
A method and system for data center power distribution optimization
CN120930089B
Data flow-energy flow coupling scheduling method and equipment for integrated energy system
CN121073080A
Computer system intelligent backup method based on thermal power plant
CN121166446A