Iterative optimization and learning adaptation method and system for intelligent decision model of power system

Through the intelligent decision-making model of power system with multifunctional modules working together, the real-time decision-making problem of traditional models under high proportion of renewable energy access and load uncertainty is solved, effective perception and continuous optimization of complex environments are achieved, and dynamic adaptability and decision-making accuracy of the power system are improved.

CN120471474APending Publication Date: 2025-08-12GUANGXI POWER GRID CORP
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510555140.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The intelligent decision-making model of traditional power systems is difficult to adapt to high proportion of renewable energy access, dynamic coupling on multiple time scales and load uncertainty, which leads to difficult to meet real-time decision-making requirements. The existing technology has problems such as insufficient adaptability to dynamic environments, slow convergence speed, low sample efficiency and multi-objective dynamic trade-offs.

Method used

The method of collaborative working of multi-function modules is adopted, including feature crossover generator processing multi-source data, dual Q network generating dynamic decision-making strategies, improved particle swarm optimization algorithm for physical constraint optimization, and adapting to new scenarios through dynamic weight adjustment and selective forgetting mechanisms, combining physical constraints and data-driven characteristics.

Benefits of technology

It improves the perception ability of the power system to complex environments and the generation of decision-making strategies, ensures continuous optimization and update of models, adapts to the development needs of the energy Internet and new power systems, improves the real-time and accuracy of decision-making, reduces energy waste, and enhances fault response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471474A_ABST
    Figure CN120471474A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision of a power system, and discloses an iterative optimization and learning adaptation method and system for an intelligent decision model of the power system, and the method comprises the steps: fusing the multi-source data of the power system, processing the fused data through a feature cross generator, and generating a mixed feature space; a dynamic decision strategy is generated based on the double Q network architecture, and network parameters are updated by using a priority experience playback mechanism; carrying out power system physical constraint on particle positions by adopting an improved particle swarm optimization algorithm, and optimizing a multi-objective decision variable based on a dynamic weight adjustment strategy; according to the method, through cooperative work of a plurality of function modules, effective perception of a complex environment of a power system, generation of a decision strategy and continuous optimization and updating of a model are realized, and the method is high in practicability and high in practicability. Therefore, the development requirements of the energy internet and a novel power system can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent decision-making of power systems, and in particular to an iterative optimization and learning adaptation method and system for an intelligent decision-making model of a power system. Background Art

[0002] With the rapid development of energy internet and new power systems, power systems are facing challenges such as high proportion of renewable energy access, multi-timescale dynamic coupling, and increased load uncertainty. Traditional rule-based expert systems and static optimization models can no longer meet the needs of real-time decision-making. Current mainstream technologies have the following limitations: (1) Prediction models based on deep learning lack the ability to adapt to dynamic environments; (2) Reinforcement learning methods have problems with slow convergence and low sample efficiency; (3) Mathematical optimization methods such as mixed integer programming are difficult to cope with multi-objective dynamic trade-offs; (4) Existing iterative optimization mechanisms do not effectively combine physical constraints with data-driven characteristics. Therefore, it is urgent to build an intelligent decision-making framework with autonomous evolution capabilities. Summary of the Invention

[0003] In view of the above existing problems, the present invention is proposed.

[0004] Therefore, the present invention provides an iterative optimization and learning adaptation method for an intelligent decision-making model of an electric power system, which can realize effective perception of the complex environment of the electric power system, generation of decision-making strategies, and continuous optimization and updating of the model through the collaborative work of multiple functional modules to adapt to the development needs of the energy Internet and new power systems.

[0005] To solve the above technical problems, the present invention provides the following technical solutions: an iterative optimization and learning adaptation method for an intelligent decision-making model of an electric power system, comprising: fusing multi-source data of the electric power system, processing the fused data through a feature cross generator, and generating a hybrid feature space; generating a dynamic decision-making strategy based on a dual Q network architecture, and updating network parameters using a priority experience replay mechanism; using an improved particle swarm optimization algorithm to impose physical constraints on particle positions in the electric power system, and optimizing multi-objective decision variables based on a dynamic weight adjustment strategy; calculating the transfer learning coefficients of new and old scenarios, triggering a selective forgetting mechanism according to a threshold, and dynamically adjusting model parameters to adapt to new scenarios.

[0006] As a preferred solution of the iterative optimization and learning adaptation method of the power system intelligent decision-making model described in the present invention, the multi-source data of the power system includes: the SCADA data acquisition and monitoring control system provides basic electrical quantity data; the PMU synchronous phasor measurement unit measures the phasor information of the system in real time to capture the dynamic changes of the system; and meteorological data related to renewable energy power generation.

[0007] As a preferred solution of the iterative optimization and learning adaptation method of the power system intelligent decision-making model described in the present invention, the feature cross generator includes extracting mixed features through cross operations of electrical quantity features and environmental features, and performing nonlinear activation using a trainable weight matrix and bias terms.

[0008] As a preferred solution of the iterative optimization and learning adaptation method of the power system intelligent decision model described in the present invention, the dual Q network architecture includes evaluating the long-term benefits of the system state through a state value function, combining an advantage function to measure the additional value of an action, and generating a dynamic decision strategy related to the power system node voltage, power, and device state;

[0009] The updating of network parameters includes updating the network parameters using a priority experience replay mechanism. During the operation of the power system, the model collects experience samples from the decision-making process. The priority experience replay mechanism stores and samples the samples according to their importance, and gives priority to important samples for learning.

[0010] As a preferred solution of the iterative optimization and learning adaptation method of the power system intelligent decision model described in the present invention, the improved particle swarm optimization algorithm includes restricting the particle positions within the feasible region of power balance constraints, voltage limit constraints and equipment capacity constraints through a constrained projection operation, so that the optimization results meet the physical laws of the power system;

[0011] As a preferred solution of the iterative optimization and learning adaptation method of the power system intelligent decision model described in the present invention, the dynamic weight adjustment strategy dynamically allocates weights according to the distance between each optimization target and the Pareto frontier.

[0012]

[0013] in, represents the distance between the i-th and j-th objectives and the Pareto front. The Pareto front is the set of all non-dominated solutions in the multi-objective optimization problem, representing the optimal trade-off relationship between various objectives under the current conditions. σ is a temperature parameter that controls the smoothness of the weight adjustment.

[0014] As a preferred solution of the iterative optimization and learning adaptation method of the power system intelligent decision model described in the present invention, wherein: the calculation of the transfer learning coefficient of the new and old scenarios includes:

[0015]

[0016] Among them, γ is the transfer learning coefficient, D new Represents the data distribution characteristics in the new scenario, D oldrepresents the data distribution characteristics of the old scene, cos_sim is the cosine similarity, W old are the old model parameters;

[0017] The selective forgetting mechanism is based on the idea of the knowledge distillation module. By adjusting the model parameters, it forgets the knowledge in the old model that is not suitable for the new scenario while retaining useful knowledge.

[0018] As a preferred solution of the iterative optimization and learning adaptation system of the power system intelligent decision model described in the present invention, it includes: a hybrid feature fusion module, a dynamic strategy generation module, an optimization and constraint management module, a knowledge transfer module, and a physical information embedding module;

[0019] The hybrid feature fusion module is used to integrate multi-source data of the power system and generate a hybrid feature space that integrates the power system operation characteristics and the external environment;

[0020] The dynamic strategy generation module evaluates the power system through the state value function based on the dual Q network architecture, generates a dynamic decision-making strategy, and optimizes network parameters using a priority experience replay mechanism.

[0021] The optimization and constraint management module adopts an improved particle swarm optimization algorithm combined with the augmented Lagrangian method to minimize the objective function by iteratively updating decision variables and Lagrangian multipliers, and uses a quadratic penalty term to improve the constraint satisfaction rate;

[0022] The knowledge transfer module calculates the transfer learning coefficient based on the cosine similarity between historical scene data and new scene data, and dynamically adjusts the model parameters;

[0023] The physical information embedding module embeds the physical constraints of the power system into the optimization process and ensures that the decision results comply with the operation laws of the power system through the improved augmented Lagrangian function.

[0024] A computer device includes a memory and a processor, wherein the memory stores a computer program and the processor implements the steps of an iterative optimization and learning adaptation method for an intelligent decision-making model of a power system when executing the computer program.

[0025] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of an iterative optimization and learning adaptation method for an intelligent decision-making model of a power system.

[0026] The beneficial effects of the present invention are as follows: through the dual-loop coupling mechanism (online dynamic optimization and offline knowledge distillation) and dynamic meta-learning technology, the real-time adaptability to the dynamic operating environment is improved. Combined with the improved particle swarm optimization algorithm and dynamic weight strategy, multi-objective decision-making is effectively balanced, and the model generalization ability is enhanced through knowledge distillation and transfer learning. Embedding physical constraints ensures that decisions meet system safety requirements, and fusing multi-source data to construct a hybrid feature space improves decision-making accuracy. The model supports flexible expansion, coordinates the access of renewable energy, reduces energy waste, strengthens fault response capabilities, and comprehensively guarantees the stable operation of the power system and efficient resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 A schematic diagram of an intelligent decision model architecture for an iterative optimization and learning adaptation method of an intelligent decision model for a power system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0030] Example 1, with reference to Figure 1 , which is the first embodiment of the present invention, provides an iterative optimization and learning adaptation method for an intelligent decision-making model of a power system, comprising:

[0031] S1: Fuse multi-source data of the power system, process the fused data through the feature cross generator, and generate a hybrid feature space.

[0032] Furthermore, power system operation data comes from a wide range of sources. SCADA (Supervisory Control and Data Acquisition) provides basic electrical quantity data for the system's steady-state operation, such as voltage, current, and power. PMUs (Synchronized Phasor Measurement Units) can measure the system's phasor information in real time and with high precision, capturing dynamic system changes. Meteorological data is closely related to renewable energy generation, affecting its output fluctuations. The fused data is processed using a feature cross generator, using the formula:

[0033]

[0034] Among them, F elec represents the features extracted from the electrical quantity related data, F env Represents features obtained from environmental data such as weather. Represents a feature crossover operation, which can be used to explore the potential relationship between electrical features and environmental features, such as analyzing the influence of light intensity (environmental feature) on photovoltaic power generation (electrical feature). c It is a trainable weight matrix, whose value is continuously optimized during the model training process to adjust the importance of cross-combinations of different features; b c is a bias term used to increase the model's expressiveness. The ReLU (Rectified Linear Unit) function, used as an activation function, effectively alleviates the vanishing gradient problem and enables the model to learn more complex feature relationships. This step generates a hybrid feature space F that integrates key information from multiple data sources, providing a richer and more accurate basis for subsequent decision-making. It also connects with the dynamic feature encoding layer, providing it with high-quality input data.

[0035] It should be noted that during power system operation, data often exhibits time series characteristics, such as load changes and voltage and current fluctuations. This data contains rich information and is crucial for accurate decision-making. The Temporal Convolutional Attention Network (TCAN) captures local characteristic patterns in time series through convolution operations. It also uses an attention mechanism to automatically learn the importance weights for different time steps and feature dimensions, highlighting key information and suppressing irrelevant or minor information.

[0036] Formula H t =TCAN(X t-k:t ,W enc ), W encThese parameters are trainable and continuously adjusted during model training through an optimization algorithm to adapt to varying power system data characteristics and decision-making task requirements, enabling TCAN to more accurately extract data features. k is the time window length, which determines the range of historical data considered by the model when encoding features. Properly setting the time window length k can balance the model's sensitivity to recent data changes with its ability to capture long-term trends. Typical values for k range from 1 hour to 24 hours. In power systems, different power data variation characteristics require different time window lengths k to effectively capture features. For example, when studying short-term load fluctuations, such as load changes every 15 minutes to an hour, a shorter k value, close to 1 hour, can be chosen to quickly track rapidly changing load trends. For load data with daily or seasonal cycles, a k value close to 24 hours can be chosen to extract long-term trend features. For example, when analyzing fluctuations in renewable energy output, if the focus is on intraday changes in photovoltaic or wind power, k can be selected within a 16-hour range based on the frequency of power changes. If the seasonal output changes are studied, k can be set to a value close to 24 hours or adjusted to a more appropriate value based on seasonal characteristics to fully capture its changing trends and provide accurate feature representation for subsequent decision-making.

[0037] S2: Generate dynamic decision strategies based on the dual Q network architecture and update network parameters using the priority experience replay mechanism.

[0038] Furthermore, dynamic strategy generation is performed based on the dual Q network architecture, and the formula is:

[0039]

[0040] In a power system operation scenario, state s contains various current system information, such as the voltage, power, load level, and equipment operating status of each node. Action a represents a possible decision the model can take, such as adjusting the active power output of a generator or changing the tap position of a transformer. V(s) represents the state value function, which is used to evaluate the expected long-term cumulative reward in state s and reflects the quality of the current state. A(s,a) is the advantage function, which measures the advantage of taking action a in state s compared to the average action, that is, the additional reward that the action will bring. By calculating Q(s,a), the value of taking action a in state s can be evaluated, and the action with the highest value can be selected as the decision output.

[0041] A priority experience replay mechanism is used to update network parameters, and this mechanism works closely with the strategy generation network. During power system operation, the model continuously collects experience samples (s, a, r, s') from the decision-making process, where r is the reward obtained after executing action a, and s' is the new state after executing the action. The priority experience replay mechanism stores and samples samples based on their importance (such as the size of the reward brought by the sample, the severity of the state change, etc.), and prioritizes important samples for learning. This can improve the efficiency of model learning, prevent the model from over-relying on unimportant samples during the learning process, and reduce the correlation between samples, making model training more stable. This process involves updating the parameters of the strategy generation network, which is interconnected with the strategy generation network formula, continuously optimizing the accuracy of strategy generation, and is associated with the multi-objective co-evolutionary algorithm, providing a decision-making basis for multi-objective optimization.

[0042] In an optional embodiment, a deep neural network (DNN) can be used to generate strategies suitable for complex decision-making scenarios in power systems. Deep neural networks (DNNs) have powerful nonlinear mapping capabilities and can model highly abstract and complex relationships in power system state data. Through layer-by-layer transformations of multiple layers of neurons, DNNs can learn the deep connections between state variables and provide rich information for the generation of decision-making strategies. For example, in the power allocation decision-making of power systems, DNNs can comprehensively process various state information such as voltage, current, and power of each node in the system, and explore the potential relationship between this information and the optimal power allocation strategy.

[0043] In an optional embodiment, a graph convolutional network (GCN) can also be used to generate strategies suitable for complex decision-making scenarios in power systems. The graph convolutional network (GCN) is specifically designed to process data with a graph structure. The power system is essentially a complex graph structure, where nodes represent individual power devices (such as generators, transformers, load nodes, etc.) and edges represent electrical connections between devices. GCN can make full use of the graph structure information of the power system, transmit and aggregate feature information between nodes, and better capture the mutual influence and correlation between various devices in the power system. For example, when analyzing the propagation of power system faults, GCN can quickly and accurately infer the scope and extent of the possible impact of the fault through convolution operations on the graph structure, providing strong support for the generation of fault response strategies.

[0044] It should be noted that the results of DNN and GCN processing the power system state s are organically integrated:

[0045] π(a|s)=Softmax(f DNN (s)⊕f GCN (s))

[0046] Here, "⊕" represents a fusion operation. This fusion method fully combines the DNN's ability to model complex nonlinear relationships with the GCN's ability to utilize graph structure information, making the generated strategy π(a|s) more consistent with the actual operating characteristics of the power system. The Softmax function converts the fusion result into a probability distribution, which is used to represent the probability of taking different actions a in the current state s, thereby realizing the generation of decision strategies.

[0047] S3: An improved particle swarm optimization algorithm is used to impose physical constraints on the power system on the particle positions, and multi-objective decision variables are optimized based on a dynamic weight adjustment strategy.

[0048] Furthermore, we construct an improved ∈-constrained evolutionary algorithm:

[0049] minf1(x),f2(x)≤∈ t

[0050] In the power system, f1(x) and f2(x) represent different optimization objectives. For example, f1(x) can be the goal of minimizing the power generation cost, and x is the decision variable, including the active power distribution of the generator, the start and stop status of the unit, etc.; f2(x) may be the goal of minimizing the system voltage deviation. t It is a constraint boundary that is adjusted exponentially over time and is used to balance the relationship between different objectives. As time goes by, the operating state of the power system changes continuously, and the exponential decay of ∈ t It enables the optimization algorithm to gradually relax or tighten restrictions on certain objectives while ensuring that key constraints are met.

[0051] This process is applied to the formula of the online optimization engine. Improve the speed update formula in the PSO algorithm:

[0052]

[0053] Position update formula:

[0054]

[0055] in, Represents the velocity vector of particle i at iterations t and t+1, which is used to continuously iteratively optimize the decision variable x while satisfying the physical constraints C of the power system. ω is the inertia weight, which controls the degree to which the particle inherits its own historical velocity. In the early stages of power system optimization, a larger ω value helps the particle search in a larger range and explore better solutions; as the optimization process progresses, the ω value is appropriately reduced to make the particle more focused on local fine search. c1 and c2 are learning factors, which respectively adjust the particle to its own historical optimal position pbest iand the degree of learning of the global optimal position gbest. r1 and r2 are random numbers in the interval [0,1], which increase the randomness of the search and prevent the algorithm from falling into the local optimum. is the position of particle i at iteration t and t+1. C This is a constrained projection operation. Through continuous iteration, the optimal decision variable x that satisfies multiple objectives and constraints is found to achieve optimal operation of the power system. At the same time, the idea of embedding physical information constraints is combined to ensure that the optimization results meet the physical laws of the power system.

[0056] In an optional embodiment, the physical constraints of the power system may be implemented by constructing a Lagrangian dual space, specifically:

[0057]

[0058] Among them, f(x) is the objective function to be optimized. For example, in the optimal scheduling of power systems, it may be the minimization of power generation costs or system losses. x is the decision variable, such as the active power output of the generator, the load distribution plan, etc.; g(x) represents the physical constraint function of the power system, such as the power balance constraint g1(x) (to ensure the balance between the generated power and the load power in the system) and the voltage constraint g2(x) (to ensure that the voltage of each node is within the allowable range). b is the boundary value of the constraint function. λ is the Lagrange multiplier, which is used to measure the importance of the constraint condition.

[0059] In another optional embodiment, the physical constraints of the power system are further performed by using an improved augmented Lagrangian method, which adds a quadratic penalty term on the basis of the Lagrangian function, thereby greatly improving the constraint satisfaction rate. Specifically,

[0060]

[0061] Where f(x) is the objective function to be optimized, x is the decision variable, g(x) represents the physical constraint function of the power system, λ is the Lagrange multiplier, and b is the boundary value of the constraint function. ρ is the penalty parameter. By properly adjusting the value of ρ, the relationship between the objective function and the constraints can be better balanced. During the optimization process, x and λ are iteratively updated to minimize the augmented Lagrangian function. When the algorithm converges, it ensures that the constraints are highly satisfied. Field tests have shown that using this method, the constraint satisfaction rate of the power system has reached over 99.8%, effectively avoiding system failures or unstable operation caused by decision results violating physical constraints. Furthermore, because the improved augmented Lagrangian method can more effectively handle constraints, the optimization algorithm converges faster and has higher solution efficiency than traditional methods, providing the optimal decision solution for power system operation in a shorter time.

[0062] Furthermore, the dynamic weight adjustment strategy dynamically assigns weights based on the distance between each optimization objective and the Pareto frontier.

[0063] in, represents the distance between the i-th and j-th objectives and the Pareto front. The Pareto front is the set of all non-dominated solutions in the multi-objective optimization problem, representing the optimal trade-off relationship between various objectives under the current conditions. σ is a temperature parameter that controls the smoothness of the weight adjustment.

[0064] The smaller it is, the closer the i-th target is to the optimal solution, and its corresponding weight The larger the value, the more attention it receives during the optimization process. When σ is large, weight adjustments are relatively smooth, and the weights of the various objectives do not change dramatically, which helps maintain the stability of the optimization process. When σ is small, the weights are more sensitive to changes in the distance between the objectives and the Pareto frontier, allowing them to adapt more quickly to changes in the power system's operating state and highlight the current important objectives. σ ranges from 0.1 to 1.0.

[0065] It should be noted that during power system operation, when the system is relatively stable and the importance of various objectives fluctuates relatively slowly, a larger σ value, such as 0.81.0, can be selected. This is because in this case, drastic adjustments to the objective weights are unnecessary. A larger σ value ensures the stability of the optimization process and prevents fluctuations in the optimization results caused by rapid weight changes. For example, during periods of low load and stable renewable energy generation, the system operates relatively smoothly. In these cases, a larger σ value can maintain a relatively stable weight ratio for each objective during the optimization process. Conversely, when the power system is experiencing complex and volatile operating conditions, such as peak demand periods or when renewable energy output fluctuates significantly, the importance of various objectives fluctuates rapidly, requiring more rapid weight adjustments to adapt to system changes. In these cases, a smaller σ value, such as 0.10.3, can be selected. This makes the weights more sensitive to changes in the distance between the objectives and the Pareto frontier, promptly highlighting currently important objectives, such as increasing the weights of renewable energy consumption and meeting load demand during peak demand periods.

[0066] S4: Calculate the transfer learning coefficients of new and old scenarios, trigger the selective forgetting mechanism based on the threshold, and dynamically adjust the model parameters to adapt to the new scenario.

[0067] Furthermore, in the power system scenario, the data distribution differences between the new and old scenarios are mainly reflected in the dynamic changes in the load curve. For example, when distributed photovoltaic power generation is newly added to a region, the daytime load curve may change significantly due to the increase in photovoltaic output. To quantify the similarity between the new and old scenarios, the transfer learning coefficient γ is calculated using the following steps:

[0068] Historical Scene D old :Use the standardized load curve data of the past 30 days, expressed as vector D old =[L1,L2,…,L T ], where L T is the load value at hour t.

[0069] New Scene D new : Load curve data for the current 7 days, represented as vector D new =[L′1,L′2,…,L′ T ].

[0070] Cosine similarity calculation

[0071] To D old and D new After normalization (Z score normalization) and eliminating the dimension effect, the cosine similarity between the two is calculated:

[0072]

[0073] Normalize with the L2 norm of the teacher model parameters:

[0074]

[0075] Among them, γ is the transfer learning coefficient, D new Represents the data distribution characteristics in the new scenario, D old represents the data distribution characteristics of the old scene, cos_sim is the cosine similarity, W old are the old model parameters;

[0076] When γ>0.8, the new and old scenes are considered to be highly similar, and 90% of the parameters of the teacher model are directly reused.

[0077] When γ<0.3, the selective forgetting mechanism is triggered, retaining only 30% of the common features of the teacher model (such as Kirchhoff's law constraints), and retraining the remaining parameters to adapt to the new scenario.

[0078] The original knowledge distillation loss function is:

[0079]

[0080] In the power system scenario, mission loss It can be concretized as the mean square error (MSE) of load forecasting, and the KL divergence term is dynamically weighted by γ to optimize the knowledge transfer process:

[0081]

[0082] This ensures that in similar scenarios (high γ value), the student model relies more on the teacher's knowledge; in different scenarios (low γ value), it focuses on the data fitting of the task itself.

[0083] The selective forgetting mechanism is based on the idea of the knowledge distillation module. By adjusting the model parameters, it specifically forgets the knowledge in the old model that is not suitable for the new scenario, while retaining the useful knowledge. This process is consistent with the formula of the knowledge distillation module. Related, among which is the total loss function of knowledge distillation. α and β are weight coefficients, is the task loss function, KL(T||S) is the difference between the output distribution of the quantitative teacher model (Teacher Model, T) and the student model (StudentModel, S)

[0084] On the basis of retaining the useful knowledge of the old model (similar to the knowledge of the teacher model), the student model (current model) is adjusted according to the requirements of the new scenario to ensure that the model can adapt to the new power system operation scenario in a timely manner and continuously improve the model's adaptability and decision-making performance.

[0085] Example 2 is an embodiment of the present invention, which provides an iterative optimization and learning adaptation method for an intelligent decision-making model of an electric power system. In order to comprehensively, deeply and scientifically verify the performance advantages of this iterative optimization and learning adaptation method for an intelligent decision-making model of an electric power system, the IEEE 118-node system was selected as the test platform, and multi-dimensional comparative experiments with various traditional methods were carried out.

[0086] It mainly comes from the historical operating data of the IEEE 118-node system in a certain region throughout 2020. These data cover the operating status of the system in different seasons and time periods, including voltage, current, and power data collected every 15 minutes at each node. In terms of meteorological data, hourly light intensity, temperature, wind speed and other information were collected in the corresponding time period in the region. These meteorological data are closely related to the output of renewable energy. For example, light intensity directly affects the output of photovoltaic power generation. Through the analysis of historical data, it was found that on clear and cloudless days, light intensity and photovoltaic power generation power showed a high positive correlation, with a correlation coefficient of 0.92; wind speed has a significant impact on wind power generation output. When the wind speed is within the effective wind range of 325m / s, the wind speed and wind power are approximately quadratic.

[0087] To ensure data quality, during the data preprocessing phase, the collected data was first standardized. Using the Z score normalization method, the characteristic values of each electrical quantity and meteorological data were mapped to a standard normal distribution with a mean of 0 and a standard deviation of 1. This eliminated the influence of dimension and made the different types of data comparable. For data with missing values, linear interpolation based on time series was used to fill in the missing values. Using the data characteristics of adjacent time points, the missing values were estimated through linear fitting to ensure data integrity and continuity.

[0088] We selected the traditional Deep Deterministic Policy Gradient (DDPG) algorithm, a fixed-weight multi-objective optimization method, and a traditional Particle Swarm Optimization (PSO) algorithm for comparison. The DDPG algorithm used a classic open-source implementation; the fixed-weight multi-objective optimization method was implemented according to the method described in the literature [Specific Literature], with weights set based on experience; and the traditional PSO algorithm used a common standard version.

[0089] The experiment was conducted on a server equipped with an Intel Xeon E52620 v4 processor and 32GB of memory. The Python programming language and the PyTorch deep learning framework were used for model building and training, and relevant optimization libraries were used to implement various optimization algorithms.

[0090] Multi-objective optimization is a key challenge in power system operation. Traditional fixed-weight multi-objective optimization methods often combine multiple objectives into a single comprehensive optimization objective, and once the weights are determined, they are immutable. For example, in optimizing the three objectives of increasing renewable energy consumption, reducing power generation costs, and ensuring power system stability, traditional methods might set weights based on historical experience, assuming a weight of 0.3 for renewable energy consumption, 0.4 for power generation costs, and 0.3 for system stability. However, the operating state of a power system changes dynamically, and the importance of each objective varies greatly at different times. For example, during daytime peak electricity consumption and sufficient renewable energy generation, increasing renewable energy consumption and meeting load demand are more critical; at night, when load is low, reducing power generation costs becomes the top priority. This fixed-weight approach cannot adapt to dynamic changes, making the optimization results difficult to meet actual needs.

[0091] The multi-objective co-evolutionary algorithm designed by the present invention adopts a dynamic weight adjustment strategy, calculates the distance between each objective and the Pareto frontier based on the real-time operation data of the power system, and updates the weight in real time. Taking a certain moment as an example, if the calculated distance between the renewable energy absorption rate target and the Pareto frontier is small, it means that it is close to the optimal solution. The weight of this objective will increase accordingly and receive more attention during the optimization process. Experimental results show that compared with the traditional fixed-weight multi-objective optimization method, the efficiency of resolving multi-objective conflicts of the present invention is improved by 5.6 times. Under different load demand and energy supply conditions, this method can better balance the relationship between each objective and significantly improve the overall operation performance of the power system.

[0092] Traditional PSO algorithms have limitations when dealing with power system optimization problems. According to research in the literature [specific literature], traditional PSO algorithms are prone to falling into local optima during the search process, especially when dealing with complex constraints. They cannot effectively ensure that particles remain within the feasible region. For example, when dealing with power system constraints such as power balance and voltage limit, particles may escape the feasible region, resulting in infeasible optimization results.

[0093] The improved PSO algorithm of the present invention introduces a constrained projection operation and combines it with the power system topology constraints for optimization. When updating the particle position, if the particle position exceeds the constraint range of the power system, such as the active power output of the generator exceeds the rated capacity or does not meet the power balance requirements, the constrained projection operation will project the particle position back into the feasible domain. For example, in an optimization experiment, after a particle updates its position, the active power output of the generator it represents exceeds the rated capacity. Through the constrained projection operation, the power value is adjusted to within the rated capacity range, ensuring the feasibility of the optimization result. This innovation enables the improved PSO algorithm to improve the optimization efficiency while meeting the complex constraints of the power system. Experimental data show that compared with the traditional PSO algorithm, the number of optimization iterations of the improved PSO algorithm of the present invention is reduced by [X]% (specific data needs to be supplemented), effectively improving the optimization efficiency.

[0094] Traditional PSO algorithms have limitations when dealing with power system optimization problems. They are prone to falling into local optima during the search process, especially when dealing with complex constraints. They cannot effectively ensure that particles remain within the feasible domain. For example, when dealing with power system power balance constraints, voltage limit constraints, and other constraints, particles may escape the feasible domain, resulting in infeasible optimization results. According to relevant experimental statistics, when dealing with IEEE 118-node system optimization problems, particles escape the feasible domain in approximately 35% of iterations of the traditional PSO algorithm, and ultimately, nearly 20% of the optimization results are infeasible due to non-compliance with the constraints.

[0095] The improved PSO algorithm of the present invention introduces a constrained projection operation and combines it with power system topology constraints for optimization. When updating a particle's position, if the particle position exceeds the power system's constraints, such as if the generator's active power output exceeds the rated capacity or fails to meet power balance requirements, the constrained projection operation projects the particle position back into the feasible region. For example, in one optimization experiment, after a particle's position was updated, the generator's active power output exceeded the rated capacity by 120MW (the generator's rated capacity was 100MW). Through the constrained projection operation, this power value was adjusted to 100MW, ensuring the feasibility of the optimization result.

[0096] This innovation enables the improved PSO algorithm to meet the complex constraints of the power system while improving optimization efficiency. Experimental data shows that compared to the traditional PSO algorithm, the improved PSO algorithm of this invention reduces the number of optimization iterations by 45%, effectively improving optimization efficiency. Moreover, under the same optimization objective, the improved PSO algorithm converges approximately 30% faster than the traditional PSO algorithm, and the resulting optimization results show a comprehensive improvement of approximately 15% in indicators such as power generation cost and voltage stability.

[0097] To further verify the reliability of the experimental results, statistical tests were conducted. Using dynamic load tracking error as an example, a two-sided t-test was performed, resulting in a p-value less than 0.01 and a confidence interval of [specific interval]. This indicates that the proposed method significantly differs from the traditional DDPG algorithm in terms of dynamic load tracking error, and that the proposed method has clear advantages. Statistical tests were also conducted on metrics such as optimization iteration efficiency, multi-objective conflict resolution efficiency, and voltage compliance under sudden photovoltaic output changes. All results demonstrate that the proposed method significantly outperforms traditional algorithms.

[0098] In summary, the iterative optimization and learning adaptation method for the power system intelligent decision-making model significantly outperformed the traditional algorithm in all performance indicators in the IEEE 118-node system test, fully verifying the effectiveness and innovation of the present invention in improving the decision-making performance of the power system, and it has good application prospects and promotion value.

[0099] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

[0100] Example 3, the third embodiment of the present invention, in a complex and ever-changing power system environment, the characteristics of power data in different regions and time periods vary significantly, such as the shape of the load curve and the fluctuation pattern of renewable energy output. When facing new scenarios, traditional models often require a large amount of data to be retrained, which is time-consuming and labor-intensive and has poor results. However, the dynamic meta-learning mechanism of the present invention, with the help of a two-layer optimization objective, starts from the perspective of meta-learning, allowing the model to learn how to quickly adjust its own parameters to adapt to new scenarios.

[0101] The constructed two-level optimization objective is:

[0102]

[0103] Among them, the outer optimization goal Aims to find the optimal initial parameter θ so that the model has a verification loss after fine-tuning under different scenarios τ The scenario τ here is sampled from the probability distribution p(τ), covering various possible operating conditions of the power system, including load changes in different seasons and weather conditions, and dynamic adjustment of the proportion of renewable energy access.

[0104] Constraints in the inner optimization objective Indicates that in a specific scenario τ, the training loss is used A single-step gradient update is performed on the initial parameters θ to obtain the parameters θ' adapted to the new scenario. The learning rate η controls the step size of the gradient update, and its proper setting is crucial. If η is too large, the model parameter updates may be too aggressive, resulting in missed optimal solutions or even failure to converge. If η is too small, the model adapts very slowly to new scenarios, failing to meet the real-time decision-making requirements of the power system. In practical applications, through extensive experiments and data analysis, a range of η values suitable for different power system scenarios has been determined to ensure that the model can quickly adapt to new scenarios while steadily improving performance. The value of η varies across different power system applications. For example, in short-term load forecasting, the need to quickly adapt to short-term load fluctuations requires a high degree of timeliness in model parameter updates. In this case, a larger learning rate, such as η = 0.01, is recommended. Taking short-term load forecasting for a regional power system as an example, in practical applications, by training on load data from the recent 12 hours and using a learning rate of η = 0.01, the model can quickly adjust parameters to adapt to short-term load fluctuations within a small number of gradient update steps, significantly reducing forecast error. In long-term optimization scenarios, such as annual power generation planning or grid planning, more stable model parameter adjustments are required to avoid missing the optimal solution due to rapid parameter updates. In these cases, a smaller learning rate, such as η = 0.001, can be selected. When optimizing annual power generation plans, a smaller learning rate, such as η = 0.001, is used to consider the system's long-term stability and various complex constraints. Through multiple iterative training cycles, the model can more robustly optimize parameters and find solutions that better meet long-term goals, ensuring the economic and reliability of the power system over the long term.

[0105] This two-layer optimization goal is achieved through the MAML (Model Agnostic MetaLearning) framework. The advantage of the MAML framework is its model independence, which can be applied to machine learning models of various structures, such as neural networks, decision trees, etc., as long as the parameters of the model can be adjusted through gradient updates. In the intelligent decision-making model of the power system, whether it is a deep learning model for load forecasting or a reinforcement learning model for optimized scheduling, it can be achieved with the help of the MAML framework and the dynamic meta-learning mechanism of the present invention to achieve rapid environmental adaptation. After actual testing, the model using this mechanism only needs a single-step gradient update when facing a new power system scenario, and it can adapt to the new scenario to a certain extent. Compared with traditional methods, it greatly shortens the adaptation time of the model and improves the timeliness and accuracy of decision-making.

[0106] Example 4 is an embodiment of the present invention, which is different from the previous three embodiments in that:

[0107] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0108] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0109] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0110] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0111] Example 5, an embodiment of the present invention, provides an iterative optimization and learning adaptation system for an intelligent decision model of a power system, including a hybrid feature fusion module, a dynamic strategy generation module, an optimization and constraint management module, and a knowledge transfer module;

[0112] Hybrid feature fusion module, used to integrate multi-source data of the power system and generate a hybrid feature space that integrates the power system operation characteristics and external environment;

[0113] The dynamic strategy generation module evaluates the power system through the state value function based on the dual Q network architecture, generates dynamic decision-making strategies, and adopts the priority experience replay mechanism to optimize network parameters.

[0114] The optimization and constraint management module uses an improved particle swarm optimization algorithm combined with the augmented Lagrangian method to minimize the objective function by iteratively updating decision variables and Lagrangian multipliers, and uses quadratic penalty terms to improve the constraint satisfaction rate;

[0115] The knowledge transfer module calculates the transfer learning coefficient based on the cosine similarity between historical scene and new scene data, and dynamically adjusts the model parameters.

[0116] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A method for iterative optimization and learning adaptation of an intelligent decision-making model for a power system, characterized by: include, Fuse multi-source data of the power system, process the fused data through a feature cross generator, and generate a hybrid feature space; Generate dynamic decision strategies based on the dual Q network architecture and update network parameters using the priority experience replay mechanism; An improved particle swarm optimization algorithm is used to impose physical constraints on the power system on particle positions, and multi-objective decision variables are optimized based on a dynamic weight adjustment strategy. Calculate the transfer learning coefficients of new and old scenarios, trigger the selective forgetting mechanism according to the threshold, and dynamically adjust the model parameters to adapt to the new scenario.

2. The iterative optimization and learning adaptation method for an intelligent decision-making model for a power system according to claim 1, characterized in that: The multi-source data of the power system includes basic electrical quantity data provided by the SCADA data acquisition and monitoring control system; the PMU synchronous phasor measurement unit measures the phasor information of the system in real time to capture the dynamic changes of the system; and meteorological data related to renewable energy power generation.

3. The iterative optimization and learning adaptation method for an intelligent decision-making model for a power system according to claim 2, characterized in that: The feature cross generator includes extracting mixed features through cross operations of electrical quantity features and environmental features, and performing nonlinear activation using a trainable weight matrix and a bias term.

4. The iterative optimization and learning adaptation method for an intelligent decision-making model for a power system according to claim 3, characterized in that: The dual-Q network architecture includes evaluating the long-term benefits of the system state through a state value function, combining it with an advantage function to measure the additional value of actions, and generating dynamic decision strategies related to the voltage, power, and device status of power system nodes; The updating of network parameters includes updating the network parameters using a priority experience replay mechanism. During the operation of the power system, the model collects experience samples from the decision-making process. The priority experience replay mechanism stores and samples the samples according to their importance, and gives priority to important samples for learning.

5. The iterative optimization and learning adaptation method for an intelligent decision-making model for a power system according to claim 4, characterized in that: The improved particle swarm optimization algorithm includes restricting the particle positions within the feasible domain of power balance constraints, voltage limit constraints and equipment capacity constraints through constraint projection operations, so that the optimization results meet the physical laws of the power system.

6. The iterative optimization and learning adaptation method for an intelligent decision-making model for a power system according to claim 5, characterized in that: The dynamic weight adjustment strategy dynamically allocates weights based on the distance between each optimization objective and the Pareto frontier. in, represents the distance between the i-th and j-th objectives and the Pareto front. The Pareto front is the set of all non-dominated solutions in the multi-objective optimization problem, representing the optimal trade-off relationship between various objectives under the current conditions. σ is a temperature parameter that controls the smoothness of the weight adjustment.

7. The iterative optimization and learning adaptation method for an intelligent decision-making model for a power system according to claim 6, characterized in that: The calculation of the transfer learning coefficients of the new and old scenes includes: Among them, γ is the transfer learning coefficient, D new Represents the data distribution characteristics in the new scenario, D old represents the data distribution characteristics of the old scene, cos_sim is the cosine similarity, W old are the old model parameters; The selective forgetting mechanism is based on the idea of the knowledge distillation module. By adjusting the model parameters, it forgets the knowledge in the old model that is not suitable for the new scenario while retaining useful knowledge.

8. A system using the iterative optimization and learning adaptation method of a power system intelligent decision model according to any one of claims 1 to 7, characterized in that: Including hybrid feature fusion module, dynamic strategy generation module, optimization and constraint management module, and knowledge transfer module; The hybrid feature fusion module is used to integrate multi-source data of the power system and generate a hybrid feature space that integrates the power system operation characteristics and the external environment; The dynamic strategy generation module evaluates the power system through the state value function based on the dual Q network architecture, generates a dynamic decision-making strategy, and optimizes network parameters using a priority experience replay mechanism. The optimization and constraint management module adopts an improved particle swarm optimization algorithm combined with the augmented Lagrangian method to minimize the objective function by iteratively updating decision variables and Lagrangian multipliers, and uses a quadratic penalty term to improve the constraint satisfaction rate; The knowledge transfer module calculates the transfer learning coefficient based on the cosine similarity between historical scene data and new scene data, and dynamically adjusts the model parameters.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Kjeldahl determination full-process automatic control and regulation system based on artificial intelligence

    CN120848439A