Power system dynamic adaptive decision generation method, system, device and medium
By building a power system decision-making method with dynamic network structure and real-time parameter updates, the uncertainty and dynamic adaptability problems of the power system under the grid of high-proportion new energy are solved, efficient and accurate adaptive decision-making are achieved, and the system's operating efficiency and stability are improved.
Patent Information
- Application Number
- CN202510672263.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-25
AI Technical Summary
When the existing power system decision-making methods face the grid connection of high proportion of new energy, there are problems such as increased uncertainty, insufficient dynamic adaptability and high computational complexity, especially the static network structure cannot respond to rapidly changing operating conditions in real time.
The dynamic adaptive decision generation method is adopted, and the initial dynamic network structure is constructed, and the structural dynamic optimization algorithm and parameter adaptive update method are used, and the variational inference algorithm and random sampling algorithm are combined to realize real-time probability inference and decision optimization, and the feedback mechanism is used to optimize the model.
It improves the adaptability and accuracy of the power system to the dynamic environment, enhances the scientificity and reliability of decision-making, improves the overall operating efficiency and stability of the system, and achieves a multi-objective balance.
Smart Images

Figure CN120377359A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system decision-making generation, and particularly to a method, system, device and medium for dynamically adaptive decision-making generation in a power system. Background Art
[0002] With the large-scale grid connection of new energy and the increasing complexity of power systems, traditional power system decision-making methods face the following challenges: enhanced uncertainty: the randomness of wind and solar power generation, load fluctuations, and equipment failures make it difficult to accurately predict the system state; insufficient dynamic adaptability: existing decision-making methods based on rules or static models (such as expert systems, linear programming) cannot respond in real time to rapidly changing operating conditions; high computational complexity: deep learning methods rely on massive amounts of data and lack interpretability, making it difficult to meet the real-time decision-making requirements.
[0003] Bayesian networks have been partially applied to power system risk assessment due to their probabilistic reasoning ability and uncertainty handling advantages, but existing studies mostly use static network structures and do not solve the problem of adaptive decision-making in dynamic environments. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method, system, device and medium for dynamically adaptive decision-making generation in a power system to solve the problem that existing studies mostly use static network structures and do not have adaptive decision-making in dynamic environments.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for dynamically adaptive decision-making generation in a power system, including:
[0008] Obtain parameter variables, perform data processing on the parameter variables, and construct an initial dynamic network structure according to the parameter variables after data processing;
[0009] Dynamically adjust the initial dynamic network structure using a structure dynamic optimization algorithm, and update the parameters in the initial dynamic network structure in real time using a parameter adaptive update method to obtain a first dynamic network model;
[0010] Perform real-time probabilistic inference on state variables using the first dynamic network model through a variational inference algorithm to obtain a real-time inference result, use the real-time inference result as an input, and combine it with an objective function to obtain a decision-making plan;
[0011] Based on the decision-making plan, generate a candidate decision set using a random sampling algorithm, and evaluate the candidate decision set to obtain an optimal decision;
[0012] According to the optimal decision, the first dynamic network model is optimized using a feedback mechanism to achieve dynamic adaptive decision generation.
[0013] As a preferred solution of a method for generating a dynamic adaptive decision in a power system according to the present invention, wherein: the data processing of the parameter variables includes:
[0014] Perform a first encoding on the parameter variables to obtain first parameter variables;
[0015] Construct an initial dynamic network structure through a time slicing mechanism to obtain a joint probability distribution.
[0016] The beneficial effect of this preferred technical solution is that by performing a first encoding on the parameter variables and constructing an initial dynamic network structure, parameter simplification and dynamic modeling can be achieved, improving the calculation efficiency and adaptability of the model.
[0017] As a preferred solution of a method for generating a dynamic adaptive decision in a power system according to the present invention, wherein: the dynamic adjustment of the initial dynamic network structure includes:
[0018] Obtain historical parameter variables, perform a time series analysis on the historical parameter variables using a time window, and perform a first calculation on the dynamic mutual information between the historical parameter variables;
[0019] According to the first calculation result, for the mutual information change exceeding the first threshold, trigger network structure adjustment.
[0020] The beneficial effect of this preferred technical solution is that by dynamically adjusting the network structure, the change of variable correlation can be captured in real time, enhancing the adaptability and accuracy of the model to the dynamic environment.
[0021] As a preferred solution of a method for generating a dynamic adaptive decision in a power system according to the present invention, wherein: the real-time update of the parameters in the initial dynamic network structure includes:
[0022] Update the conditional probability table of the initial dynamic network structure using an incremental learning algorithm;
[0023] Adopt a double-window mechanism to update the parameters of the initial dynamic network structure, evaluate the stability of the initial dynamic network structure, and adjust the initial dynamic network structure to obtain a first dynamic network model.
[0024] The beneficial effect of this preferred technical solution is that by real-time updating the parameters and evaluating the network stability, the dynamic response ability and decision reliability of the model are improved.
[0025] As a preferred solution of a method for generating a dynamic adaptive decision in a power system according to the present invention, wherein: the real-time probabilistic inference of the state variables includes:
[0026] Perform a second calculation on the posterior probability distribution of the state variables through a variational inference algorithm;
[0027] Use a first optimization algorithm to adjust the parameters of the variational distribution and minimize the difference between the variational distribution and the posterior probability distribution;
[0028] Use a second optimization algorithm to generate candidate decision-making schemes, evaluate the candidate decision-making schemes in combination with the objective function, and obtain the decision-making schemes.
[0029] The beneficial effects of this preferred technical solution are real-time probabilistic inference and optimized decision-making, improving the scientificity and reliability of decision-making, and enhancing the system's ability to cope with uncertainties.
[0030] As a preferred scheme of a method for generating a dynamic adaptive decision in a power system according to the present invention, wherein: evaluating the candidate decision set includes:
[0031] Evaluating the expected utility of the candidate decision set through a Bayesian network includes:
[0032] By traversing all the next states, calculating the product of the probability of each next state and the corresponding reward value, and accumulating the products.
[0033] As a preferred scheme of a method for generating a dynamic adaptive decision in a power system according to the present invention, wherein: optimizing the first dynamic network model by using a feedback mechanism includes:
[0034] Obtain decision-related data, and calculate the deviation between the decision execution result and the expected result according to the decision-related data;
[0035] According to the deviation, use a third optimization algorithm to adjust the parameters of the conditional probability table of the first dynamic network model;
[0036] Define a reward function, construct a state space and an action space, use a first reinforcement learning algorithm to update the weights of the objective function, obtain the optimal weights, and realize the optimization of dynamic adaptive decision generation.
[0037] In a second aspect, the present invention provides a power system dynamic adaptive decision generation system, including:
[0038] A data processing module, configured to obtain parameter variables, perform data processing on the parameter variables, and construct an initial dynamic network structure according to the parameter variables after data processing;
[0039] A model construction module, configured to dynamically adjust the initial dynamic network structure by using a structure dynamic optimization algorithm, and update the parameters in the initial dynamic network structure in real time by using a parameter adaptive update method to obtain a first dynamic network model;
[0040] A real-time inference module, which is used to perform real-time probabilistic inference on state variables through a variational inference algorithm by using the first dynamic network model, obtain a real-time inference result, and use the real-time inference result as an input, and combine it with an objective function to obtain a decision-making scheme;
[0041] A decision evaluation module, which is used to generate a candidate decision set based on the decision-making scheme by using a random sampling algorithm, and evaluate the candidate decision set to obtain an optimal decision;
[0042] A model optimization module, which is used to optimize the first dynamic network model by using a feedback mechanism according to the optimal decision, and realize dynamic adaptive decision generation.
[0043] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of the method for generating a dynamic adaptive decision for a power system are implemented.
[0044] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that when the computer program is executed by a processor, the steps of the method for generating a dynamic adaptive decision for a power system are implemented.
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows: through real-time dynamic modeling, intelligent decision optimization, and a closed-loop feedback mechanism, the present invention solves the problems of uncertainty, dynamic adaptability, and multi-objective balance of a power system under high-proportion new energy grid connection, and provides key technical support for building a new power system. Verified by actual cases, its technical indicators reach the international leading level, and it has significant economic, social, and environmental benefits. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic diagram of the overall process logic of a method for generating a dynamic adaptive decision for a power system provided by an embodiment of the present invention. Detailed Embodiments
[0048] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0049] Example 1, referring to Figure 1 , which is an embodiment of the present invention, provides a method for generating a dynamic adaptive decision for a power system, including:
[0050] S100: Obtain parameter variables, perform data processing on the parameter variables, and construct an initial dynamic network structure based on the parameter variables after data processing;
[0051] S200: Dynamically adjust the initial dynamic network structure using a structure dynamic optimization algorithm, and update the parameters in the initial dynamic network structure in real time using a parameter adaptive update method to obtain a first dynamic network model;
[0052] S300: Use the first dynamic network model to perform real-time probabilistic inference on state variables through a variational inference algorithm to obtain a real-time inference result, and use the real-time inference result as input, combined with an objective function, to obtain a decision-making plan;
[0053] S400: Based on the decision-making plan, generate a candidate decision set using a random sampling algorithm, and evaluate the candidate decision set to obtain an optimal decision;
[0054] S500: Optimize the first dynamic network model using a feedback mechanism according to the optimal decision to achieve dynamic adaptive decision generation.
[0055] It should be noted that through dynamic network structure adjustment, real-time parameter update, real-time probabilistic inference, optimized decision generation, and feedback mechanism optimization, the efficient, accurate, and adaptive decision-making of the power system is realized; the dynamic network structure and parameter update improve the response ability of the model to system changes; real-time probabilistic inference and optimized decision-making enhance the scientificity and reliability of decision-making; the feedback mechanism further optimizes the model performance and improves the overall operation efficiency and stability of the system.
[0056] In the embodiment of the present invention, the above step S100 includes the following sub-steps A1 - A2;
[0057] In A1: Perform a first encoding on the parameter variables to obtain first parameter variables;
[0058] In A2: Construct an initial dynamic network structure through a time slicing mechanism to obtain a joint probability distribution;
[0059] In an alternative embodiment, the first encoding may be equal-width interval discretization encoding. Equal-width interval discretization divides the value range of a continuous variable into several intervals with equal widths, and each interval corresponds to a discrete value. For example, for the continuous variable of power generation in the power generation state, first determine its value range. Assuming the power generation range is 0 - 1000 MW, if it is to be divided into 5 intervals, then the width of each interval is (1000 - 0) / 5 = 200 MW. These 5 intervals are [0, 200), [200, 400), [400, 600), [600, 800), [800, 1000], and each interval can be encoded as 0, 1, 2, 3, 4 in sequence.
[0060] In an alternative embodiment, the first encoding may be equal-frequency interval discretization encoding. Equal-frequency interval discretization divides the value range of a continuous variable into several intervals such that the number of data points in each interval is approximately equal. For example, for the continuous variable of load demand, assuming there are 1000 data points and it is to be divided into 4 intervals, then each interval should contain approximately 250 data points. By sorting the data and then dividing according to the number of data points, the corresponding intervals are obtained, and then each interval is encoded.
[0061] In an alternative embodiment, the first encoding may be clustering-based discretization encoding. The clustering-based discretization method uses a clustering algorithm (such as K-Means clustering) to divide the data points of a continuous variable into several clusters, and each cluster corresponds to a discrete value. For example, for some continuous features in network topology (such as the connection strength between nodes), the K-Means algorithm can be used to divide these data points into 3 clusters. The center of each cluster represents the typical value of the cluster, and then the data points belonging to the same cluster are encoded as the same discrete value.
[0062] In the embodiment of the present invention, the first encoding includes discretization encoding.
[0063] Specifically, a four-dimensional variable set V = {X G , X L , X T , X F} including power generation state X G , load demand X L , network topology X T , and equipment failure X F is established, and each variable uses discretization encoding.
[0064] Exemplarily, the power generation state X G can be divided into three states: normal power generation, early warning of insufficient power generation, and power generation equipment failure; the load demand X LIt can be divided into states such as normal load, high load warning, and load overload fault. By accurately describing and encoding these key variables, the operating state of the power system can be better reflected;
[0065] Specifically, the initial dynamic network structure is a dynamic Bayesian network. A T-order dynamic Bayesian network is constructed, and the joint probability distribution is expressed as:
[0066]
[0067] Among them, represents the set of parent nodes of X i at time t, which contains relevant variables of the previous T time slices, V t is the set of variables at time point t, V t-1 ,…,V t-T is the set of variables from time point t - 1 to t - T, represents the i-th variable at time point t;
[0068] The time slicing mechanism divides the time series data into multiple time slices, and the variable relationships within each time slice are described by conditional probabilities; the multi-order Markov model assumes that the system state at the current moment is only related to the states of the previous T moments. In this way, the temporal change law of the system state can be effectively captured.
[0069] For example, in a power system, the power generation state at the current moment may be affected by factors such as the states of power generation equipment and load demand in the previous few moments. The multi-order Markov model can more accurately describe this dependence relationship.
[0070] It should be noted that through discretized encoding and dynamic Bayesian network modeling, the state description of a complex power system can be efficiently simplified, while capturing the temporal change law of the system state, providing an accurate model basis for subsequent dynamic decision-making, and significantly improving the scientificity and real-time nature of decision-making.
[0071] In the embodiment of the present invention, the above step S200 includes the following sub-steps B1 - B2;
[0072] In B1: Obtain historical parameter variables, perform temporal analysis on the historical parameter variables using a time window, and perform a first calculation on the dynamic mutual information between the historical parameter variables;
[0073] In B2: According to the first calculation result, trigger network structure adjustment for the mutual information change exceeding the first threshold;
[0074] In an alternative embodiment, the first calculation may be a sliding-window based mutual information calculation. A sliding window is set, which slides along the time series. At each time point, the mutual information of the variables under time delay is calculated, and the average mutual information is obtained by taking the average of the mutual information calculated for all segments.
[0075] In an alternative embodiment, the first calculation may be a wavelet-transform based mutual information calculation. The time series data of the variables is wavelet-transformed to obtain wavelet coefficients under different frequency components. At each frequency component, the mutual information between the wavelet coefficients of the variables is calculated, and the total mutual information is obtained by summing up the mutual information for all frequency components.
[0076] In the embodiment of the present invention, the first calculation includes performing a time series analysis on historical data using a time window of size W, and calculating the dynamic mutual information between variables, expressed as:
[0077]
[0078] where τ is the number of time delay steps, is the value of variable x i at time point t, is the value of variable x j at time point t - τ, and P(·) is the marginal probability distribution.
[0079] For example, in a power system, when a fault occurs in a certain line, the correlation between relevant power generation, load, and network topology variables will change, and this change can be quickly detected through the calculation of dynamic mutual information.
[0080] Specifically, when it is detected that the mutual information change exceeds the first threshold, network structure adjustment is triggered.
[0081] If then add an edge. When the correlation between two variables exceeds a certain threshold, it indicates that there is a strong causal relationship between them, and a corresponding edge needs to be added to the network structure to represent this relationship.
[0082] For example, when it is found that the mutual information between the change in load demand and the change in power generation status exceeds the threshold, it indicates that the load demand has a significant impact on the power generation status, and an edge from the load demand to the power generation status needs to be added to the network. I th is the threshold of the mutual information of the newly added edge, with the unit of bit.
[0083] If for W consecutive windows then remove the edge; when the correlation between two variables continuously remains below a certain minimum value, it indicates that the causal relationship between them has become very weak or does not exist. At this time, the corresponding edge in the network structure can be removed to reduce the complexity and redundant information of the model. Imin is the threshold of edge mutual information deletion, with the unit of bit;
[0084] Among them, the first threshold ε = 0.1 bit, I th = 0.3 bit, I min = 0.05 bit;
[0085] It should be noted that by calculating the dynamic mutual information, the change of correlation between variables can be monitored in real time; the sliding window mechanism slides forward continuously, and the data in each window is analyzed, so as to timely detect the dynamic change of correlation between variables.
[0086] In the embodiment of the present invention, after the steps B1 - B2 are completed in the above step S200, the following steps B3 - B4 are further included;
[0087] In B3: Use the incremental learning algorithm to update the conditional probability table of the initial dynamic network structure;
[0088] In B4: Adopt the dual - window mechanism to update the parameters of the initial dynamic network structure, evaluate the stability of the initial dynamic network structure, and adjust the initial dynamic network structure to obtain the first dynamic network model.
[0089] In an alternative embodiment, the incremental learning algorithm can be the sliding window method, which stores the most recent data points by maintaining a fixed - size window; the window slides on the data stream, and only the data within the window is considered each time; when new data arrives, the old data is removed from the window and the new data is added to the window; the update of the conditional probability table is based on the data statistics within the window;
[0090] In an alternative embodiment, the incremental learning algorithm can be online Bayesian update, which adapts to new data by gradually updating the prior probability; each time new data is received, the prior probability is updated to the posterior probability and then used for the next update.
[0091] In the embodiment of the present invention, the incremental learning algorithm includes using the exponential weighted moving average method with a forgetting factor to update the conditional probability table, which is expressed as:
[0092]
[0093] Among them, α ∈ [0, 1] is the forgetting factor, which is dynamically adjusted according to the exponential decay strategy, N obs (X i , Pa(X i )) is the number of times the combination (X i , Pa(X i )) is observed within the window, N obs (Pa(X i )) is the number of times Pa(X iThe total number of observations, P old (X i |Pa(X i )) is the original conditional probability table;
[0094] It should be noted that through the incremental learning algorithm, the model can continuously update the conditional probability table according to new data without retraining the entire dataset, improving the update efficiency and real-time performance of the model.
[0095] Specifically, a dual-window mechanism is adopted. The short-term window (W1 = 10s) is used for real-time parameter update. The short-term window can quickly capture short-term changes in the system state and update the model's parameters in a timely manner to ensure the model's response ability to real-time data. For example, when an instantaneous fault occurs in the power system, the short-term window can quickly detect changes in the data and update the conditional probability table;
[0096] The long-term window (W2 = 1h) is used for structural stability assessment. The long-term window can analyze the change trend of the system state from a more macroscopic perspective and evaluate the stability of the network structure. If it is found that the relationship between certain variables has changed continuously within the long-term window, it may be necessary to adjust the network structure;
[0097] It should be noted that the dual-window mechanism combines short-term fast response and long-term stability assessment to improve the model's dynamic adaptability and decision-making reliability.
[0098] In the embodiment of the present invention, the above step S300 includes the following sub-steps C1 - C3;
[0099] In C1: The posterior probability distribution of the state variables is calculated for the second time through the variational inference algorithm;
[0100] In C2: The parameters of the variational distribution are adjusted using the first optimization algorithm to minimize the difference between the variational distribution and the posterior probability distribution;
[0101] In C3: A candidate decision-making scheme is generated using the second optimization algorithm, and the candidate decision-making scheme is evaluated in combination with the objective function to obtain the decision-making scheme.
[0102] In an alternative embodiment, the second calculation can be the Markov chain Monte Carlo method, which generates a series of samples by performing random walks in the parameter space, and the distribution of the samples gradually approaches the true posterior distribution;
[0103] In an alternative embodiment, the second calculation can be the particle filter method, which represents the posterior distribution by maintaining a set of particles, each particle representing a possible system state. As time goes by, the particles adapt to new observation data through prediction and update steps.
[0104] In the embodiments of the present invention, the second calculation includes calculating the posterior probability distribution of the key state based on the variational inference algorithm;
[0105] Specifically, it is expressed using Bayes' theorem as:
[0106]
[0107] Where Q is a latent variable. For example, in a power system, it can represent the fault probability, which reflects some state information in the system that is difficult to directly observe; in industrial production, it may represent the potential fault risk of equipment, etc.; E is the observed evidence. For example, in a power system, it is voltage, current data, etc., which can be obtained in real time through various sensors; in industrial production, it can be the operating parameters of equipment, production progress data, etc.
[0108] The variational inference algorithm is an approximate inference method. Its core idea is to approximate the true posterior probability distribution by finding an approximate distribution. First, it is necessary to select a suitable variational distribution according to the specific problem scenario and data characteristics. Common variational distributions include Gaussian distribution, Bernoulli distribution, etc. For example, when dealing with continuous data, the Gaussian distribution may be a better choice; while for discrete data, the Bernoulli distribution is more suitable.
[0109] The first optimization algorithm is used to adjust the parameters of the variational distribution to minimize the difference between the variational distribution and the true posterior probability distribution. Commonly used optimization algorithms include stochastic gradient descent method, variational autoencoder, etc.;
[0110] In the embodiments of the present invention, the first optimization algorithm includes the stochastic gradient descent method. By continuously iteratively updating the parameters of the variational distribution, the value of the objective function is gradually reduced, thereby finding the optimal variational distribution.
[0111] It should be noted that during real-time operation, the system will continuously obtain new observed evidence E and use these new evidences to update the parameters of the variational distribution, and then update the posterior probability distribution P(Q|E) of the key state; it can track the state changes of the system in real time and provide accurate probability information for subsequent decision-making.
[0112] In an optional embodiment, the objective function can be reliability, which is the ability of the system to complete the specified function under the specified conditions and within the specified time, including aspects such as the continuity of power supply, voltage and frequency stability. In a power system, reliability is a crucial indicator because power outages may have serious impacts on industrial production, commercial activities, and residents' lives. By incorporating reliability into the objective function, the decision-making generation method can pay more attention to improving the reliability of the power system, such as optimizing the power grid structure, increasing backup power supplies, etc.;
[0113] In an alternative embodiment, the objective function can be flexibility, which refers to the system's ability to quickly respond to load and generation changes; a power system with higher flexibility can better adapt to the intermittency and volatility of renewable energy and improve energy utilization efficiency; taking flexibility as part of the objective function can prompt the decision-making generation method to take measures to improve the flexibility of the power system, such as using flexible power generation equipment and optimizing energy storage configuration.
[0114] In the embodiments of the present invention, the objective function includes three main objectives: economy, security, and environmental friendliness.
[0115] Specifically, under certain constraint conditions, minimize the weighted sum of the objective functions, that is:
[0116]
[0117] where u represents the decision variable, which represents various decision-making schemes that can be taken, and w i is the objective weight, which is dynamically adjusted according to the real-time posterior probability. For example, when the calculated fault probability is relatively high in real time, increase the weight w2 of the security objective to ensure that the system can prioritize ensuring safe operation in high-risk situations. In a power system, if it is detected that the fault probability of a certain line increases, the monitoring intensity of this line will be increased, and at the same time, the power generation and transmission strategies will be adjusted to reduce the impact of the fault on the system; g i is the constraint function, and g i (u) ≤ 0 is the constraint condition, such as equipment capacity limitation, power balance constraint, production process requirements, etc.
[0118] Use various optimization algorithms to solve the multi-objective optimization problem, such as genetic algorithms, particle swarm algorithms, etc.
[0119] In the embodiments of the present invention, the second optimization algorithm includes a genetic algorithm, which continuously iteratively searches for the optimal solution by simulating the biological evolution process. In each generation, the algorithm generates a set of candidate solutions (i.e., decision-making schemes) and evaluates these solutions according to the objective function and constraint conditions; then, through operations such as selection, crossover, and mutation, generate the next generation of candidate solutions until the required optimal solution is found.
[0120] It should be noted that by dynamically adjusting the objective weights through the multi-objective optimization algorithm and solving the optimal decision-making scheme, it is possible to effectively balance the economy, security, and environmental friendliness of the power system, while meeting various constraint conditions and improving the overall performance of the system. Economy can reduce operating costs and increase benefits; security ensures reliable power supply and avoids major losses; environmental friendliness responds to environmental requirements and promotes sustainable development.
[0121] In the embodiments of the present invention, the above step S400 includes the following sub-steps D1 - D2;
[0122] In D1: Evaluating the expected utility of the candidate decision set through a Bayesian network includes:
[0123] In D2: It is obtained by traversing all the next states, calculating the product of the probability of each next state and the corresponding reward value, and accumulating the products.
[0124] In an alternative embodiment, the random sampling algorithm can be policy gradient search. The neural network simulation policy is adjusted online through policy gradient updates, and the policy is directly optimized through policy gradient to dynamically select the optimal action during the search process;
[0125] In an alternative embodiment, the random sampling algorithm can be to gradually approximate the target distribution by constructing a Markov chain. By performing random walks in the parameter space, a series of samples are generated, and the distribution of these samples gradually approaches the true posterior distribution;
[0126] In the embodiment of the present invention, the random sampling algorithm includes Monte Carlo Tree Search (MCTS) to generate a candidate decision set. By simulating a large number of possible decision paths, a search tree is gradually constructed, and a candidate decision set is generated according to the search results;
[0127] Specifically, in the selection stage, the algorithm selects a node from the current search tree for expansion according to a certain strategy; common selection strategies include the Upper Confidence Bound Trees (UCT) algorithm, etc. The UCT algorithm comprehensively considers the number of visits to the node and the average reward, and selects a node with high potential for expansion. For example, in power system decision-making, the UCT algorithm will select a scheme that is most likely to produce the optimal result for further exploration according to the historical performance and potential benefits of different power generation and transmission schemes;
[0128] In the expansion stage, the child nodes of the node are generated; each child node represents a possible decision scheme; for example, in industrial production scheduling, if the current node represents the current production task allocation situation, then the expanded child nodes may represent different new task allocation schemes;
[0129] The expanded node is randomly simulated to evaluate its performance. During the simulation process, the algorithm randomly selects a decision path starting from the current node until a termination state is reached. Then, the performance of the decision path is evaluated according to the reward value of the termination state. For example, in the power system, the simulation process may simulate the operation of different power generation and transmission schemes over a period of time, and calculate indicators such as power generation cost and voltage stability as the reward value;
[0130] Feed the simulation results back into the search tree to update the statistical information of the nodes. During the backtracking process, the algorithm starts from the terminal node and gradually updates the visit count and average reward of its parent nodes. In this way, the statistical information of the nodes in the search tree is continuously updated to reflect the performance of different decision-making schemes.
[0131] Evaluate the expected utility of each decision through a Bayesian network, expressed as:
[0132]
[0133] where s is the current state, s ′ is the next state, and R is the reward function; the reward function R(s ′ , a) is used to measure the gain or loss after taking decision a in state s and reaching state s ′ . In the power system, the reward function may comprehensively consider factors such as generation cost, voltage stability, and carbon emissions; in industrial production, the reward function may include aspects such as production cost, production efficiency, and product quality.
[0134] It should be noted that by calculating the expected utility of each decision, select the decision with the maximum expected utility as the final decision result, so as to achieve dynamic adaptive decision-making; continuously repeat the above processes of real-time probability inference, multi-objective decision optimization, and dynamic decision generation, and adjust the decision-making scheme in real time according to the real-time state of the system and the latest observation evidence to adapt to the complex and changeable environment.
[0135] In the embodiment of the present invention, the above step S500 includes the following sub-steps E1 - E3;
[0136] In E1: Obtain decision-related data, and calculate the deviation between the decision execution result and the expected result according to the decision-related data;
[0137] In E2: According to the deviation, use the third optimization algorithm to adjust the conditional probability table parameters of the first dynamic network model;
[0138] In E3: Define the reward function, construct the state space and action space, and use the first reinforcement learning algorithm to update the weights of the objective function to obtain the optimal weights, so as to achieve the optimization of dynamic adaptive decision-making generation.
[0139] Specifically, during the decision-making execution process, the system will collect various data related to the decision in real time, including input variables, output results, and intermediate state information, etc.; these data will serve as the basis for subsequent analysis; according to the collected data, calculate the deviation between the decision-making execution result and the expected result; the deviation index can be defined according to the specific application scenario. For example, in the power system, it can be the difference between the actual power generation and the predicted power generation, or the deviation between the actual voltage and the rated voltage, etc.; based on the calculated deviation index, use the third optimization algorithm to adjust the conditional probability table parameters.
[0140] Common third optimization algorithms include the gradient descent method, Newton's method, etc. By continuously iteratively adjusting the conditional probability table parameters, the deviation index is gradually reduced, thereby improving the inference accuracy.
[0141] In a load forecasting model of a power system, if there is a large deviation between the load value predicted according to the current conditional probability table parameters and the actual load value, the system will adjust the probability distribution of the variables related to the load in the conditional probability table according to this deviation, making the next prediction more accurate.
[0142] Specifically, in complex systems, decisions often need to consider multiple objectives simultaneously, such as economy, safety, environmental protection, etc.; in order to balance the relationship between these multiple objectives, a weight w is introduced i to represent the importance of each objective; the long-term feedback mechanism uses the method of reinforcement learning to update the weight w according to the overall performance of the system over a long period of time i so as to optimize the multi-objective preference;
[0143] The reward function is the core of reinforcement learning and is used to measure the performance of the system in each state. In the present invention, the reward function comprehensively considers the achievement of multiple objectives. For example, in the power system, the reward function can be defined according to indicators such as power generation cost, voltage stability, carbon emissions, etc. When the system achieves a better multi-objective balance under a certain decision, the value of the reward function is higher; on the contrary, the value of the reward function is lower.
[0144] The state space describes the state of the system at different times, including various observable variables and parameters; the action space represents the set of decisions that the system can take. For example, in the power system, the actions can be the start and stop of power generation equipment, power distribution, etc.
[0145] The first reinforcement learning algorithm can be Q learning, deep Q network (DQN), etc. According to the specific application scenario and problem complexity, select an appropriate reinforcement learning algorithm to update the weight w i ;
[0146] At each time step, the system selects an action (makes a decision) based on the current state and executes the action. Then, according to the reward value calculated by the reward function, the weights w are updated using a reinforcement learning algorithm. i Through continuous iteration, the weights w i will gradually converge to an optimal value, enabling the system to achieve an optimal balance of multiple objectives in the long-term operation.
[0147] In an industrial production scheduling system, the long-term feedback mechanism continuously adjusts the weights w of objectives such as economy, quality assurance, and environmental protection based on the comprehensive performance of multiple objectives such as production efficiency, product quality, and energy consumption. i to achieve the overall optimization of the production process.
[0148] It should be noted that by adjusting the conditional probability table parameters through deviation feedback and dynamically updating the multi-objective weights in combination with reinforcement learning, the decision-making accuracy and system performance are effectively improved, the multi-objective dynamic balance is achieved, and the economy, security, and environmental protection of the power system operation are enhanced.
[0149] The above is a schematic solution of a power system dynamic adaptive decision-making generation method according to this embodiment. It should be noted that the technical solution of the power system dynamic adaptive decision-making generation system belongs to the same concept as the technical solution of the above power system dynamic adaptive decision-making generation method. For the details not described in detail in the technical solution of the power system dynamic adaptive decision-making generation system in this embodiment, reference can be made to the description of the technical solution of the above power system dynamic adaptive decision-making generation method.
[0150] The power system dynamic adaptive decision-making generation system in this embodiment includes:
[0151] A data processing module, configured to obtain parameter variables, perform data processing on the parameter variables, and construct an initial dynamic network structure based on the parameter variables after data processing.
[0152] A model construction module, configured to dynamically adjust the initial dynamic network structure using a structure dynamic optimization algorithm and update the parameters in the initial dynamic network structure in real time using a parameter adaptive update method to obtain a first dynamic network model.
[0153] A real-time inference module, configured to perform real-time probability inference on state variables using the first dynamic network model through a variational inference algorithm to obtain a real-time inference result, and use the real-time inference result as an input and combine it with an objective function to obtain a decision-making plan.
[0154] A decision evaluation module, configured to generate a candidate decision set based on the decision-making plan using a random sampling algorithm and evaluate the candidate decision set to obtain an optimal decision.
[0155] A model optimization module, configured to optimize the first dynamic network model according to the optimal decision by using a feedback mechanism, so as to realize the generation of dynamic adaptive decisions.
[0156] This embodiment also provides a computer device, which is applicable to the situation of dynamic adaptive decision generation in a power system, including:
[0157] A memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement a method for generating dynamic adaptive decisions in a power system as proposed in the above embodiment.
[0158] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a method for generating dynamic adaptive decisions in a power system as proposed in the above embodiment.
[0159] The storage medium proposed in this embodiment and the method for generating dynamic adaptive decisions in a power system proposed in the above embodiment belong to the same inventive concept. For technical details not described in detail in this embodiment, reference can be made to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a floppy disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (FLASH), a hard disk, or an optical disc of a computer, etc., including several instructions for causing a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present invention.
[0161] Embodiment 2. Referring to Table 1, this embodiment is different from the first embodiment and provides a verification test for a method for generating dynamic adaptive decisions in a power system to verify and illustrate the technical effects adopted in this method.
[0162] A certain regional power grid includes: new energy: 500 MW of wind power and 300 MW of photovoltaic power; conventional power sources: 2 × 300 MW thermal power units; load characteristics: industrial and commercial loads account for 60%, and the maximum load is 1200 MW; goal: to achieve safe and economic dispatching under high-proportion new energy access;
[0163] Variable encoding includes: power generation status: wind power / solar power (normal / downturn / disconnected), thermal power (normal / maintenance); load status: normal (<80%), warning (80%-95%), overload (>95%); network status: bus voltage (normal / overlimit), line power flow (normal / overload); fault status:
[0164] equipment fault (0-1 encoding);
[0165] Deploy PMU (synchronized phasor measurement unit) devices to collect voltage, current, and power data at 100ms intervals, and obtain device status, power generation plans, etc. through the SCADA (supervisory control and data acquisition) system;
[0166] The mutual information threshold is determined using the 3σ principle: calculate the mutual information between variables for the historical one-year grid operation data (a total of 3,504,000 time points); statistically obtain the mutual information mean μ = 0.08bit and standard deviation σ = 0.04bit, and set:
[0167] ε = μ + 0.5σ = 0.10bit (the minimum change amount triggering structural adjustment)
[0168] I th = μ + 3σ = 0.20bit (the strict threshold for new edges, the original patent value of 0.3bit is for engineering margin)
[0169] I min = μ - 2σ = 0.00bit (set to 0.05bit in actual engineering to avoid misdeletion)
[0170] If on a certain summer day at 14:00, the load suddenly increases to 1150MW (exceeding the threshold of 95%); the power of the photovoltaic system drops suddenly by 200MW due to cloud cover;
[0171] It is monitored that the mutual information between the load and the thermal power output suddenly increases to 0.35bit (>0.3bit threshold), and a new causal edge of load → thermal power is added;
[0172] If the mutual information between the photovoltaic system and the bus voltage is detected to be <0.05bit for 3 consecutive long-term windows (3 hours), the photovoltaic → voltage edge is deleted;
[0173] Adopt the exponential decay model:
[0174] α(t) = α0 - (α0 - α min )π·2 -t / T
[0175] where: α0 = 0.95 (initial forgetting factor); α min = 0.5 (minimum forgetting factor); T = 100 (half-life period)
[0176] The model satisfies: at time t = 0, α = 0.95 (initial trust historical data);
[0177] When t → ∞, α → 0.5 (the weights of new and old data are equal);
[0178] Derivative Ensure monotonic decrease;
[0179] The information entropy H(α) = -αlnα - (1 - α)ln(1 - α) decreases with time, which conforms to the forgetting mechanism.
[0180] Short - term window (10s): When load overload is detected, the start - up probability of thermal power increases from 0.6 to 0.85; the probability of voltage violation increases from 0.05 to 0.12
[0181] Long - term window (1h): It is found that the load overload probability at 14:00 every Wednesday is stable at 0.35, and the prior probability of load status is adjusted; the forgetting factor α decays from 0.95 to 0.82 to strengthen the influence of new data;
[0182] Decision - making scenario: Current state: load 1150MW, PV output 100MW, wind power 300MW, thermal power 200MW;
[0183] Constraint conditions: The maximum output of thermal power is 600MW, and the line N - 1 security check is carried out;
[0184] Objective function: Economy: Generation cost (thermal power cost is 80 yuan / MWh); Safety: Voltage violation risk (weight increased to 0.4); Environmental protection: Carbon emission intensity (new energy proportion ≥ 40%)
[0185] Decision - making result: Start the second thermal power unit, and the output increases to 500MW; Put into the SVG reactive power compensation device to maintain voltage stability; Dispatching cost: 98,000 yuan / h → 86,000 yuan / h after optimization (a 12% reduction);
[0186] Short - term feedback: After implementing the decision, it is found that the voltage still violates the limit, and the thermal power output is corrected → voltage CPT parameters; After the measured data is updated, the prediction error of the voltage violation probability drops from 15% to 8%;
[0187] Long - term feedback: Statistics for a continuous week show that: The excessive safety weight leads to an increase in generation cost; Reinforcement learning is used to adjust the weight: the economic weight is increased from 0.3 to 0.45;
[0188] Table 1 Implementation effect
[0189] Indicator Traditional method This solution Improvement rate Decision response time 2.3s 0.15s 93% Power generation cost 125,000 yuan / h 102,000 yuan / h 18.4% Number of voltage violations 4.2 times / day 1.1 times / day 73.8% New energy consumption rate 85% 94% 10.6%
[0190] As shown in Table 1, the threshold ε = 0.1 bit successfully identifies 92.3% of the key structural changes in historical data (verified by the confusion matrix), and the false positive rate is controlled below 4.7%; the α decay strategy is verified by KL divergence, and the model deviation is reduced by 38% compared with the model with fixed α = 0.95 in the mutation data scenario; the short-term window W1 = 10 s can cover 95% of the grid dynamic response frequency (0.05 - 1 Hz) through Fourier transform analysis; the weight w i Verified by 5000 Monte Carlo simulations, the decision-making robustness is still maintained under ±15% parameter fluctuations.
[0191] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for generating a dynamic adaptive decision in a power system, characterized in that Including: Obtain parameter variables, perform data processing on the parameter variables, and construct an initial dynamic network structure based on the parameter variables after data processing; Dynamically adjust the initial dynamic network structure using a structure dynamic optimization algorithm, and update the parameters in the initial dynamic network structure in real time using a parameter adaptive update method to obtain a first dynamic network model; Use the first dynamic network model to perform real-time probability inference on state variables through a variational inference algorithm to obtain a real-time inference result, and use the real-time inference result as input, combined with an objective function, to obtain a decision-making plan; Based on the decision-making plan, generate a candidate decision set using a random sampling algorithm, and evaluate the candidate decision set to obtain an optimal decision; According to the optimal decision, use a feedback mechanism to optimize the first dynamic network model to achieve dynamic adaptive decision-making generation.
2. The dynamic adaptive decision generation method for a power system according to claim 1, wherein Performing data processing on the parameter variables includes: Perform a first encoding on the parameter variables to obtain first parameter variables; Construct an initial dynamic network structure through a time slicing mechanism to obtain a joint probability distribution.
3. A method for generating a dynamic adaptive decision in a power system according to claim 1 or 2, characterized in that, Dynamically adjusting the initial dynamic network structure includes: Obtain historical parameter variables, perform time series analysis on the historical parameter variables using a time window, and perform a first calculation on the dynamic mutual information between the historical parameter variables; According to the first calculation result, trigger network structure adjustment for the mutual information change exceeding the first threshold.
4. The dynamic adaptive decision-making generation method for a power system according to claim 3, wherein, Updating the parameters in the initial dynamic network structure in real time includes: Update the conditional probability table of the initial dynamic network structure using an incremental learning algorithm; Adopt a double window mechanism to update the parameters of the initial dynamic network structure, evaluate the stability of the initial dynamic network structure, and adjust the initial dynamic network structure to obtain a first dynamic network model.
5. The dynamic adaptive decision-making generation method for a power system according to claim 4, characterized in that, Performing real-time probability inference on state variables includes: Perform a second calculation on the posterior probability distribution of state variables through a variational inference algorithm; Use a first optimization algorithm to adjust the parameters of the variational distribution to minimize the difference between the variational distribution and the posterior probability distribution; Use a second optimization algorithm to generate a candidate decision-making plan, and evaluate the candidate decision-making plan in combination with the objective function to obtain a decision-making plan.
6. The dynamic adaptive decision generation method for a power system according to claim 5, wherein Evaluating the candidate decision set includes: Evaluating the expected utility of the candidate decision set through a Bayesian network includes: By traversing all next states, calculate the product of the probability of each next state and the corresponding reward value, and accumulate the products to obtain.
7. The dynamic adaptive decision generation method for a power system according to claim 6, wherein, Using a feedback mechanism to optimize the first dynamic network model includes: Obtain decision-related data, and calculate the deviation between the decision execution result and the expected result according to the decision-related data; According to the deviation, use a third optimization algorithm to adjust the parameters of the conditional probability table of the first dynamic network model; Define a reward function, construct a state space and an action space, and use a first reinforcement learning algorithm to update the weights of the objective function to obtain optimal weights, realizing the optimization of dynamic adaptive decision-making generation.
8. A power system dynamic adaptive decision-making generation system, which applies a power system dynamic adaptive decision-making generation method as described in any one of claims 1-7, characterized in that, Including: A data processing module for obtaining parameter variables, performing data processing on the parameter variables, and constructing an initial dynamic network structure based on the parameter variables after data processing; A model construction module, configured to dynamically adjust the initial dynamic network structure by using a structure dynamic optimization algorithm, and to update the parameters in the initial dynamic network structure in real time by using a parameter adaptive update method, so as to obtain a first dynamic network model; A real-time inference module, configured to perform real-time probabilistic inference on state variables by using the first dynamic network model through a variational inference algorithm, obtain a real-time inference result, use the real-time inference result as an input, and combine it with an objective function to obtain a decision-making scheme; A decision evaluation module, configured to generate a candidate decision set by using a random sampling algorithm based on the decision-making scheme, and evaluate the candidate decision set to obtain an optimal decision; A model optimization module, configured to optimize the first dynamic network model by using a feedback mechanism according to the optimal decision, so as to realize dynamic adaptive decision-making generation.
9. A computer device, characterized in that, Comprising: A memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of a method for generating a dynamic adaptive decision in a power system according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor, the steps of a method for generating a dynamic adaptive decision in a power system according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Wharf yard scheduling method, device, equipment, medium and product
CN122549885A