Multi-time interval weather prediction distribution parallel confidence strategy optimization generation control method
Patent Information
- Application Number
- CN202210967237.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-08-12
AI Technical Summary
[0002]现有新型电力系统自动发电控制存在未充分考虑环境因素的问题,这导致新型电力系统无法准确跟踪环境进行调节
[0105](1)在分布式系统中,各个模块相互独立,整个系统是多线平行的架构,不会因为其中一个模块出现问题而影响整体正常运行,有强鲁棒性;在新型电力系统中加入实时气象预测网络,使新型电力系统与环境交互充分交互,能让新型电力系统准确跟踪环境进行并进行智能发电控制调节。
Smart Images

Figure CN115238592B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power generation control in new power systems, and involves artificial intelligence, quantum technology and power generation control methods. It is applicable to power generation control in new power systems and integrated energy systems. Background Technology
[0002] Existing automatic generation control systems for new power systems have the problem of not fully considering environmental factors, which makes it impossible for new power systems to accurately track and adjust to the environment.
[0003] In addition, traditional policy optimization networks require a large amount of data for training. Due to the high dimensionality of the data, the training speed of the network is slow and it is prone to the curse of dimensionality.
[0004] Therefore, a parallel trust strategy optimization method for power generation control based on multi-time-spacing meteorological forecast distribution is proposed. This method can solve the problem that new power systems cannot accurately track the environment for regulation, while also accelerating the training speed of new power systems and eliminating the curse of dimensionality. Summary of the Invention
[0005] This invention proposes a multi-temporal meteorological forecast distributed parallel reliability strategy optimization power generation control method. This method combines multi-temporal meteorological forecasting, distributed parallelism, and reliability strategy optimization neural networks for power generation control in novel power systems. The steps in using the proposed multi-temporal meteorological distributed parallel reliability strategy optimization power generation control method are as follows:
[0006] Step (1): Define each controlled power generation area as an intelligent agent, and label each area as follows: , where i is the label of each power generation area; each power generation area does not interfere with each other, but is interconnected, and has strong robustness;
[0007] Step (2): Initialize the parameters of the stacked autoencoder neural network and the gated recurrent unit, collect the meteorological data of wind intensity and light intensity for the past three years, extract the meteorological feature dataset, and input the meteorological feature dataset into the stacked autoencoder neural network;
[0008] A stacked autoencoder neural network is composed of multiple autoencoder neural networks stacked together. Let x be the input meteorological feature data vector, where x is an n-dimensional vector. The hidden layer h of the autoencoder neural network AE1 (1) The hidden layers h of the autoencoder network AE2 are used as input to train the autoencoder network AE2. (2)As input to the autoencoder network AE3, and so on; after stacking layers, the feature dimensionality of the weather data decreases, which speeds up the training of the gated recurrent unit while preserving key information of the data; since each hidden layer has a different dimension, we have:
[0009] (1)
[0010] Among them, h (1) It is the hidden layer of the autoencoder neural network AE1, h (2) It is a hidden layer of the autoencoder neural network AE2, h (p -1) It is an autoencoder neural network (AE) p-1 The hidden layer, h (p) It is an autoencoder neural network (AE) p The hidden layer; W (1) It is the hidden layer h (1) The parameter matrix, W (2) It is the hidden layer h (2) The parameter matrix, W (p) It is the hidden layer h (p) The parameter matrix; b (1) It is the bias of the autoencoder neural network AE1, b (2) It is the bias of the autoencoder neural network AE2, b (p) It is an autoencoder neural network (AE) p The bias; It is the activation function; p is the number of layers in the stack of the stack autoencoder network; It is a normalized exponential function, used as a classifier;
[0011] In a stacked autoencoder neural network, if For m dimensions, For k dimensions, the process of stacking an autoencoder network AE1 to an autoencoder network AE2 involves training a... The network structure; first, train the network. To get The transformation, and then the network is trained. ,get The transformation is performed, and finally, the autoencoder network AE1 and the autoencoder network AE2 are stacked to obtain the network. ; via autoencoder network AE1 to AE p The layers are stacked, and the output vector is finally obtained through the softmax function. The stacked autoencoder neural network was trained to obtain network parameters with certain initial values and dimensionality-reduced meteorological features. As output;
[0012] Step (3): Let Let t be the output vector of the autoencoder neural network. The initial values of network parameters pre-trained from the stacked autoencoder neural network and meteorological characteristics are used. The input is fed into the gating loop unit, making the inputs for updating and resetting the gates in the gating loop unit as follows: The outputs of updating the gate and resetting the gate in the gated loop unit are as follows:
[0013] (2)
[0014] in, Update the gate output at time t. The gate output is reset at time t. Let (t-1) be the hidden state of the gated loop unit. Let be the input at time t, and [] denote that the two vectors are connected. To update the gate weight matrix, To reset the weight matrix of the gate, It is the sigmoid function;
[0015] The gated recurrent unit discards and memorizes the input information through two gates, thus obtaining the candidate hidden state value at time t. , for:
[0016] (3)
[0017] in This represents the tanh activation function. To hide state values The weight matrix; * denotes the matrix product;
[0018] The tanh activation function, after obtaining updated state information through the update gate, creates a vector of all possible values based on the input and calculates the candidate hidden state values. Then, the state at time t is calculated through the network. , for:
[0019] (4)
[0020] The reset gate determines the number of past states you want to remember; when When the value is 0, the state information at time (t-1) It will be forgotten, hidden. It will be reset to the information input at time t; the update gate determines the number of past states in the new state; when When the value is 1, it is in a hidden state. Update to the state at time t The gated loop unit updates and resets gates and filters information, retains important features through gate functions, captures dependencies through learning, and thus obtains weather forecast values.
[0021] Step (4): After the stacked autoencoder neural network and the gated recurrent unit are trained, the meteorological data to be predicted is input into the gated recurrent unit through the stacked autoencoder neural network, and the obtained meteorological prediction value is input into the new power system. In the new power system, each power generation area is equipped with three parallel trust strategy optimization networks: short time interval, medium time interval, and long time interval. The short time interval is one day, the medium time interval is fifteen days, and the long time interval is three months.
[0022] Step (5): Initialize the parameters of the parallel trust strategy optimization network in each region, set the strategy of the parallel trust strategy optimization network, and initialize the parallel expected value table in the parallel trust strategy optimization network. The initial expected value is 0.
[0023] Step (6): Set the number of iterations to X, set the initial value of the number of searches to a positive integer v, and initialize the number of searches for each intrinsic action to V=v;
[0024] Step (7): In the current state, the parallel trust policy optimization network in each agent selects an action based on the policy, obtains the corresponding reward value of the action in the current environment, and feeds the obtained reward value back to the parallel expected value table. Then the iteration number is incremented by one. If the current iteration number is equal to X, the iteration is completed and the trained parallel trust policy optimization network is obtained.
[0025] Step (8): Perform policy optimization and parallel value optimization in the parallel trust policy optimization network. The optimization method is as follows:
[0026] The core of the parallel trust optimization policy network is the actor-critic method; in the policy optimization of the parallel trust optimization policy network, the Markov decision process is a tuple. Where S is the wind intensity predicted by step (3), the light intensity predicted by step (3), and the frequency deviation. The state space is composed of the area control error (ACE) and tie-line power exchange performance index (CPS), and any state A represents the power change of different magnitudes. The action space formed by the quantum superposition state, where j is the action space of the quantum superposition state. dimensionality, arbitrary action P represents the state transition from any state s to state a via any action a. The transition probability distribution matrix; For the reward function; Initial state The probability distribution; Let be the discount factor; Represents a random policy In strategy The expected cumulative reward function for:
[0027] (5)
[0028] in This is the initial state. It is in state Next random strategy The chosen action Let be the discount factor at time t. Let t be the state at time t. Let t be the action at time t. In the state The reward value below, Indicating in strategy medium state Next action Perform sampling;
[0029] Introducing state-action value functions State value function Advantage function and probability distribution function ;
[0030] State-Action Value Function Seeking Strategy medium state Next action The resulting cumulative rewards, state-action value function for:
[0031] (6)
[0032] in The state at time (t+1); The state at time (t+l); The action at time (t+1); The action at time (t+l); Indicating in strategy medium state Next action Perform sampling; Let l be the discount factor at time (t+l); l is a positive integer. It is in state The reward value below;
[0033] State value function Seeking Strategy medium state The cumulative rewards below for Regarding the action The average value, state value function for:
[0034] (7)
[0035] in Indicates the strategy of seeking the answer. middle Regarding the action The average value;
[0036] Advantage function Find the advantage of taking any action a in any state s compared to the average situation, and define the advantage function. for:
[0037] (8)
[0038] It is a state For any state s, the action Let a be the state-action value function for any action a; It is a state Let be the state-value function for any state s;
[0039] probability distribution function Seeking Strategy any state The probability distribution and probability distribution function are as follows. for:
[0040] (9)
[0041] in The state at time t Let be the transition probability distribution matrix for any state s;
[0042] Parallel value optimization of the parallel trust optimization policy network is a function of the state-action value function. Perform parallel optimization; state-action value function Parallel optimization is towards The method introduces qubits and Grover's search method to accelerate network training and eliminate the curse of dimensionality.
[0043] Parallel value optimization will take action Quantization, and its application in states The search count V is updated instead of the state. The following method updates the action probabilities:
[0044] Assume the motion space has Each intrinsic action will Individual intrinsic movements The superposition is represented as ; to perform the action Quantization into j-dimensional quantum superposition state actions , Each of the qubits is and The superposition of two states; quantum superposition state actions Equivalent to j-dimensional quantum superposition state action The expression is:
[0045] (10)
[0046] in It is a j-dimensional quantum superposition state action Quantum actions after being observed; It is a quantum action Probability amplitude; Quantum Actions The modulus value satisfies ;
[0047] When quantum superposition state action After being observed, It will collapse into a quantum action quantum action Each qubit is a or Quantum Actions Each qubit is represented by a value indicating its expected value; these expected values differ depending on the state; these expected values are used to select actions in a given state, and the update rules for these expected values are as follows:
[0048] If in strategy In the middle, in the state Below, quantum superposition state action Collapse into quantum actions When cumulative rewards When adding, the action The qubits in each qubit are The expected value decreases, and the quantum bit is The expected value increases; the amount of increase and decrease in the expected value follows the strategy. Update the settings, and when selecting an action in the same state next time, set the qubits with positive expected values for each qubit of the action as... The remaining quantum positions are To obtain quantum actions; let the strategy The parameter vector is The specific action value is obtained by measuring the quantum action and then expressed through the parameter vector. Converted into power change As the output of Agent1 in this state; where For parameter vectors The parameters are given, where u is a positive integer;
[0049] Quantum action selection obtains the state at time t using the Grover search method. Observing quantum superposition actions The quantum action obtained by the collapse The rules for obtaining and updating the search count V are as follows:
[0050] First, all intrinsic actions are superimposed with equal weights by sequentially applying j Hadamard gates to j initial states, where each state is equal to a given value. The independent qubits yielded the following results:
[0051] (11)
[0052] Here, H is a Hadamard gate, which can convert the ground state... Convert to equal weight stacking ; It is the action of the quantum superposition state after initialization, by Actions with the same probability amplitude are superimposed together; This means applying j Hadamard gates sequentially to... One initialization intrinsic action; the two parts of the Grover iterative operator are: and They are respectively:
[0053] (12)
[0054] Where I is a unitary matrix of appropriate dimension; It is a quantum action The reverse quantum action; It is a quantum action The reverse quantum action; It is a quantum action The reverse quantum action; It is a quantum action The outer product; It is the action of the quantum superposition state after initialization. The outer product; and It is a quantum black box; Acting on intrinsic action hour, Change and The phase of states in the same direction changes by 180°. Acting on intrinsic action hour, Change and The phase of states in the same direction changes by 180°.
[0055] Let Grover's iterative process be denoted as... :
[0056] (13)
[0057] according to For each intrinsic action, iterate to obtain the closest iteration count with the fewest iterations, and update the iteration count V corresponding to each intrinsic action, while also updating the policy. middle;
[0058] From formula (9), we can obtain the state at time t. For any state At that time, the strategy is from Updated to Expected cumulative reward function for:
[0059] (14)
[0060] in For the updated strategy; In strategy State at time t Let be the transition probability distribution matrix for any state s; In strategy State at time t For any state s, the dominance function is used. In strategy The state at time t Let be the transition probability distribution matrix for any state s; In strategy medium state Next action Perform sampling; In strategy The sampled value a under any state s in the middle; In strategy any state The probability distribution below; The state at time t Take action below Advantages compared to the average level; It is a strategy The expected cumulative reward function is as follows;
[0061] In any state There is ,in For the updated strategy The state at time t Choosing the right action value increases the cumulative reward. Or maintain cumulative rewards when the expected advantage is zero. The strategy remains unchanged, but the cumulative rewards are continuously optimized through updates. ;
[0062] Because formula (14) requires the calculation of the updated strategy probability distribution This results in high complexity for formula (7), making it difficult to optimize. To reduce computational complexity, a substitution function is introduced. :
[0063] (15)
[0064] in It is a function that finds the largest independent variable in a given function; In strategy any state The probability distribution function under;
[0065] substitution function and The difference lies in the substitution function. Ignore changes in state access density caused by policy changes. use As access frequency rather than , Access frequency By utilizing the strategy The approximate acquisition when the strategy and When certain constraints are met, the substitution function It can replace the original expected cumulative reward function ;
[0066] In parameter vector In the update, the parameter vector is used. With any parameters The form of strategy Parameterization , In parameterization strategy Any action a in any state s; for any parameter When the strategy is not updated, the alternative function is used. and the original cumulative reward function They are completely equal, that is:
[0067] (16)
[0068] in It is a strategy After parameter vector Parameterized strategy;
[0069] When the substitution function and the original cumulative reward function have any parameter The derivative in the strategy The places are the same, that is, the strategy is from Updated to When there is a very small change, if the value of the substitution function Increase, cumulative rewards This will also increase, so we can use the substitution function as the optimization objective to improve the strategy, that is:
[0070] (17)
[0071] in It concerns arbitrary parameters The differential of the function;
[0072] Formulas (16) and (17) illustrate the strategy from Updated to A sufficiently small step size can increase cumulative rewards. ;definition Define the intermediate divergence variable as the strategy with the largest cumulative reward value among the old strategies. To increase cumulative rewards The lower bound is used to set a conservative iteration strategy. for:
[0073] (18)
[0074] in For new strategies; This is the current strategy; for and The maximum total variance divergence between them; In strategy The action chosen in any state s; In strategy The action chosen in any state s; yes and Total variance divergence between them; In strategy Any action a chosen in any state s; In strategy Any action a chosen in any state s;
[0075] For any random policy, let the intermediate entropy variable... ,in Let be the maximum absolute value of the function for choosing any action 'a' in any state 's'; using replace , replace ; Substitution function value and cumulative rewards The difference satisfies:
[0076] (19)
[0077] in It is a discount factor;
[0078] The maximum relative entropy is ;in In strategy The action chosen in any state s; In strategy The action chosen in any state s; yes and The relative entropy between them;
[0079] The relationship between total variance divergence and relative entropy satisfies ;in yes and Total variance divergence between them;
[0080] make Reconstraining relative entropy:
[0081] (20)
[0082] Where C is the penalty coefficient;
[0083] Under this constraint, for the constantly updated strategy ,exist ;
[0084] in This indicates the policy update process; It is a policy sequence of a parallel trust optimization policy network; It is the cumulative reward of each policy in the policy sequence of the parallel trust optimization policy network;
[0085] Considering parameterization strategy and parameter vector Remove the parameter vector Irrelevant items;
[0086] Expected cumulative reward function after parameter variable transformation for:
[0087] (twenty one)
[0088] Substitution function after parameter variable transformation for:
[0089] (twenty two)
[0090] Relative entropy after parameter variable transformation for:
[0091] (twenty three)
[0092] The constraints after the parameter variables are transformed are:
[0093] (twenty four)
[0094] in The variable is transformed to obtain the same value; This is the parameter vector that needs to be updated; It is a parameter vector Updated parameter vector; It is a strategy After parameter vector Parameterized strategy; It is a strategy After parameter vector Parameterized strategy; It is a strategy The expected cumulative reward function; It is a strategy The substitution function; yes and The relative entropy between them; It is the maximum value of the relative entropy after the parameter variables have been transformed;
[0095] The parameter vector of the parallel policy optimization network is obtained from formulas (21) to (24). The update process; through the parameter vector The update can optimize the selection weights of actions, thereby achieving the goal of optimizing parallel control;
[0096] To ensure cumulative rewards Increase, make Maximize; because C is the penalty coefficient, it will cause each time The value becomes very small, resulting in a very short step size for each update, which slows down the update speed. Therefore, the penalty term is changed into a constraint term:
[0097] (25)
[0098] in It is a constant;
[0099] Formula (14) is based on strategy Sampling is performed, due to the previous strategy. It is unknown and cannot be based on strategy. Therefore, importance sampling is used for the parameterized cumulative reward function. Rewrite; for the parameterized cumulative reward function Ignore and any parameters Irrelevant items, and use replace Finally, the update of the parallel trust optimization policy network becomes:
[0100] (26)
[0101] in It is about the parameterized strategy probability distribution and state-action value Perform sampling; In strategy medium state For any state s, the action Let a be the state-action value function for any action a;
[0102] Based on the set constraints, the parameter vector Perform an update, using the updated parameter vector to update the strategy. The policy update in the parallel policy optimization network is completed, and then the new policy is used to select the action in the current state, and the process is iterated step by step.
[0103] Step (9): After the iteration is completed, optimize the network according to the trained parallel trust strategy to regulate the power change in each power generation area of the new power system. This enables each region of the new power system to achieve the optimal tie-line power exchange performance index (CPS); each power generation region can achieve the optimal tie-line power exchange performance index (CPS) using the methods from steps (1) to (8); through training the network in each power generation region, regional cooperation is sought to achieve dynamic equilibrium, and finally the frequency deviation between each power generation region is reduced. As the power exchange performance index (CPS) approaches zero, the overall power exchange performance index (CPS) approaches 100%, and the entire new power system gradually reaches global optimization.
[0104] The present invention has the following advantages and effects compared with the prior art:
[0105] (1) In a distributed system, each module is independent of the others, and the whole system is a multi-line parallel architecture. It will not affect the normal operation of the whole system if one of the modules has a problem, and it has strong robustness. Adding a real-time weather forecast network to the new power system enables the new power system to fully interact with the environment, and allows the new power system to accurately track the environment and carry out intelligent power generation control and regulation.
[0106] (2) By adding a stacked autoencoder neural network, the feature dimension of weather data can be reduced and the training speed of the network can be accelerated.
[0107] (3) Compared with existing policy optimization networks, parallel trust policy optimization networks can reduce the dimension of the expectation value table and eliminate the curse of dimensionality. Attached Figure Description
[0108] Figure 1 This is a framework diagram of the parallel trust strategy optimization for meteorological forecast distribution in the method of this invention.
[0109] Figure 2 This is a control flowchart of the multi-time-distance meteorological forecast distribution parallel trust strategy optimization power generation method of the present invention.
[0110] Figure 3 This is a framework diagram of the stacked autoencoder network of the method of the present invention.
[0111] Figure 4 This is a framework diagram of the gated loop unit of the method of the present invention. Detailed Implementation
[0112] The present invention proposes a parallel reliability strategy optimization method for power generation control based on multi-temporal meteorological forecast distribution, which is described in detail below with reference to the accompanying drawings:
[0113] Figure 1 This is a framework diagram of the parallel trust strategy optimization for meteorological forecast distribution in the method of this invention.
[0114] First, each controlled power generation area is defined as a controlled intelligent agent, and each area is labeled as follows: ;
[0115] Then, initialize the parameters of the stacked autoencoder neural network and the gated recurrent unit in the meteorological prediction neural network, input the meteorological data from previous years, extract the meteorological features, and input the meteorological features into the stacked autoencoder neural network and the gated recurrent unit respectively;
[0116] Then, the initial parameter values obtained by training the stacked autoencoder neural network are used to train the gated recurrent unit. The trained gated recurrent unit is used to predict future meteorological data. If the prediction effect meets the target, the training ends.
[0117] Then, the meteorological data to be predicted is input into the gated recurrent unit through a stacked autoencoder neural network to obtain the prediction results, and the prediction results are input into the new power system;
[0118] Then, in each power generation area, three parallel trust strategies are set up to optimize the network: short time interval, medium time interval, and long time interval. The short time interval is one day, the medium time interval is fifteen days, and the long time interval is three months.
[0119] Then, initialize the system's parallel trust strategy optimization network parameters, set the parallel trust strategy optimization network strategy, initialize the parallel expected value table in the parallel trust strategy optimization network, set the initial expected value to 0, set the number of searches to V, and set the number of iterations to X;
[0120] Then, the parallel trust policy optimization network is pre-trained, and the initial values of the pre-trained network parameters are input into the parallel trust policy optimization network.
[0121] Then, in the current state, the parallel trust policy optimization network in each agent selects an action based on the policy, obtains the corresponding reward value for that action in the current environment, feeds the obtained reward value back to the parallel expectation value table, increments the iteration count by one, and determines whether the current iteration count is equal to X; if the iteration count is not equal to X, the parallel expectation value table is updated, the time difference error in the experience pool is updated, and the optimization policy is updated; if the iteration count is equal to X, the training of the parallel trust policy optimization network is complete.
[0122] Finally, the new power system is controlled by the optimized parallel trust strategy network, which regulates the power output of each agent to enable the new power system to achieve the optimal tie-line power exchange performance index (CPS).
[0123] Figure 2 This is a control flowchart of the multi-time-distance meteorological forecast distribution parallel trust strategy optimization power generation method of the present invention.
[0124] The diagram contains two parts: a weather forecasting neural network and a new power system. First, the weather forecasting neural network is operated by training a gated recurrent unit using past weather data. Then, the trained gated recurrent unit is used to predict future weather based on current weather data.
[0125] Then, taking Agent1 of the new power system as an example, Agent1 is equipped with three parallel trust strategy optimization networks for short-duration, medium-duration, and long-duration scenarios. Agent1 receives meteorological data output from the predictive neural network and uses the frequency deviation of Agent1... The regional control error ACE1 and tie-line power exchange performance index CPS are used to optimize the strategy selection actions in the network based on each parallel trust strategy.
[0126] Finally, update the parallel expected value table, update the time difference error in the experience pool, update the optimization strategy, and repeat this process to obtain the optimal tie-line power exchange performance index (CPS) for Agent1.
[0127] Except for Agent1, all other power generation areas can obtain the optimal tie-line power exchange performance index (CPS) for that area using this method.
[0128] Figure 3 This is a framework diagram of the stacked autoencoder network of the method of the present invention.
[0129] First, initialize the model parameters and input the past meteorological dataset into the autoencoder neural network; build the initial autoencoder neural network to compress the meteorological data from the original n dimensions to m dimensions;
[0130] Then, ignoring the input state y of the autoencoder neural network, and using the hidden layer h as the original novel, a new autoencoder is trained. Through layer stacking, the feature dimension of the data is reduced while retaining key information.
[0131] Then, the trained data is compared with the actual data, the loss function is calculated, and the system parameters are updated.
[0132] Finally, the initial values of the trained parameters are input into the autoencoder neural network.
[0133] Figure 4 This is a framework diagram of the gated loop unit of the method of the present invention.
[0134] First, initialize the model parameters and input the historical meteorological dataset into the gated recurrent unit;
[0135] Then, the network parameters are updated by resetting the short-term dependencies in the gate capture time series and updating the long-term dependencies in the gate capture time series.
[0136] Finally, the trained network will be used to predict current meteorological data.
Claims
1. A parallel trust strategy optimization method for power generation control based on multi-temporal meteorological forecast distribution, characterized in that, This method combines multi-timescale meteorological forecasting, distributed parallel processing, and a trust-based policy optimization neural network for power generation control in novel power systems. The steps in using the proposed multi-timescale meteorological forecasting, distributed parallel processing, and trust-based policy optimization power generation control method are as follows: Step (1): Define each controlled power generation area as an intelligent agent, and label each area as follows: , where i is the label of each power generation area; Step (2): Initialize the parameters of the stacked autoencoder neural network and the gated recurrent unit, collect the meteorological data of wind intensity and light intensity for the past three years, extract the meteorological feature dataset, and input the meteorological feature dataset into the stacked autoencoder neural network; A stacked autoencoder neural network is composed of multiple autoencoder neural networks stacked together. Let x be the input meteorological feature data vector, where x is an n-dimensional vector. The hidden layer h of the autoencoder neural network AE1 (1) The hidden layers h of the autoencoder network AE2 are used as input to train the autoencoder network AE2. (2) As input to the autoencoder network AE3, and so on; after stacking layers, the feature dimensionality of the weather data decreases, which speeds up the training of the gated recurrent unit while preserving key information of the data; since each hidden layer has a different dimension, we have: (1) Among them, h (1) It is the hidden layer of the autoencoder neural network AE1, h (2) It is a hidden layer of the autoencoder neural network AE2, h (p-1) It is an autoencoder neural network (AE) p-1 The hidden layer, h (p) It is an autoencoder neural network (AE) p The hidden layer; W (1) It is the hidden layer h (1) The parameter matrix, W (2) It is the hidden layer h (2) The parameter matrix, W (p) It is the hidden layer h (p) The parameter matrix; b (1) It is the bias of the autoencoder neural network AE1, b (2) It is the bias of the autoencoder neural network AE2, b (p) It is an autoencoder neural network (AE) p The bias; It is the activation function; p is the number of layers in the stack of the stack autoencoder network; It is a normalized exponential function, used as a classifier; In a stacked autoencoder neural network, if For m dimensions, For k dimensions, the process of stacking an autoencoder network AE1 to an autoencoder network AE2 involves training a... The network structure; first, train the network. To get The transformation, and then the network is trained. ,get The transformation is performed, and finally, the autoencoder network AE1 and the autoencoder network AE2 are stacked to obtain the network. ; via autoencoder network AE1 to AE p The layers are stacked, and the output vector is finally obtained through the softmax function. The stacked autoencoder neural network was trained to obtain network parameters with certain initial values and dimensionality-reduced meteorological features. As output; Step (3): Let Let t be the output vector of the autoencoder neural network. The initial values of network parameters pre-trained from the stacked autoencoder neural network and meteorological characteristics are used. The input is fed into the gating loop unit, making the inputs for updating and resetting the gates in the gating loop unit as follows: The outputs of updating the gate and resetting the gate in the gated loop unit are as follows: (2) in, Update the gate output at time t. The gate output is reset at time t. Let (t-1) be the hidden state of the gated loop unit. Let be the input at time t, and [] denote that the two vectors are connected. To update the gate weight matrix, To reset the weight matrix of the gate, It is the sigmoid function; The gated recurrent unit discards and memorizes the input information through two gates, thus obtaining the candidate hidden state value at time t. , for: (3) in This represents the tanh activation function. To hide state values The weight matrix; * denotes the matrix product; The tanh activation function, after obtaining updated state information through the update gate, creates a vector of all possible values based on the input and calculates the candidate hidden state values. Then, the state at time t is calculated through the network. , for: (4) The reset gate determines the number of past states you want to remember; when When the value is 0, the state information at time (t-1) It will be forgotten, hidden. It will be reset to the information input at time t; the update gate determines the number of past states in the new state; when When the value is 1, it is in a hidden state. Update to the state at time t The gated loop unit updates and resets gates and filters information, retains important features through gate functions, captures dependencies through learning, and thus obtains weather forecast values. Step (4): After the stacked autoencoder neural network and the gated recurrent unit are trained, the meteorological data to be predicted is input into the gated recurrent unit through the stacked autoencoder neural network, and the obtained meteorological prediction value is input into the new power system. In the new power system, each power generation area is equipped with three parallel trust strategy optimization networks: short time interval, medium time interval, and long time interval. The short time interval is one day, the medium time interval is fifteen days, and the long time interval is three months. Step (5): Initialize the parameters of the parallel trust strategy optimization network in each region, set the strategy of the parallel trust strategy optimization network, and initialize the parallel expected value table in the parallel trust strategy optimization network. The initial expected value is 0. Step (6): Set the number of iterations to X, set the initial value of the number of searches to a positive integer v, and initialize the number of searches for each intrinsic action to V=v; Step (7): In the current state, the parallel trust policy optimization network in each agent selects an action based on the policy, obtains the corresponding reward value of the action in the current environment, and feeds the obtained reward value back to the parallel expected value table. Then the iteration number is incremented by one. If the current iteration number is equal to X, the iteration is completed and the trained parallel trust policy optimization network is obtained. Step (8): Perform policy optimization and parallel value optimization in the parallel trust policy optimization network. The optimization method is as follows: The core of the parallel trust optimization policy network is the actor-critic method; in the policy optimization of the parallel trust optimization policy network, the Markov decision process is a tuple. Where S is the wind intensity predicted by step (3), the light intensity predicted by step (3), and the frequency deviation. The state space is composed of the area control error (ACE) and the tie-line power exchange performance index (CPS), and any state A represents the power change of different magnitudes. The action space formed by the quantum superposition state, where j is the action space of the quantum superposition state. dimensionality, arbitrary action P represents the state transition from any state s to state a via any action a. The transition probability distribution matrix; For the reward function; Initial state The probability distribution; Let be the discount factor; Represents a random policy In strategy The expected cumulative reward function for: (5) in This is the initial state. It is in state Next random strategy The chosen action Let be the discount factor at time t. Let t be the state at time t. Let t be the action at time t. In the state The reward value below, Indicating in strategy medium state Next action Perform sampling; Introducing state-action value functions State value function Advantage function and probability distribution function ; State-Action Value Function Seeking Strategy medium state Next action The resulting cumulative rewards, state-action value function for: (6) in The state at time (t+1); The state at time (t+l); The action at time (t+1); The action at time (t+l); Indicating in strategy medium state Next action Perform sampling; Let l be the discount factor at time (t+l); l is a positive integer. It is in state The reward value below; State value function Seeking Strategy medium state The cumulative rewards below for Regarding the action The average value, state value function for: (7) in Indicates the strategy of seeking the answer. middle Regarding the action The average value; Advantage function Find the advantage of taking any action a in any state s compared to the average situation, and define the advantage function. for: (8) It is a state For any state s, the action Let a be the state-action value function for any action a; It is a state Let be the state-value function for any state s; probability distribution function Seeking Strategy any state The probability distribution and probability distribution function are as follows. for: (9) in The state at time t Let be the transition probability distribution matrix for any state s; Parallel value optimization of the parallel trust optimization policy network is a function of the state-action value function. Perform parallel optimization; state-action value function Parallel optimization is towards The method introduces qubits and Grover's search method to accelerate network training and eliminate the curse of dimensionality. Parallel value optimization will take action Quantization, and its application in states The search count V is updated instead of the state. The following method updates the action probabilities: Assume the motion space has Each intrinsic action will Individual intrinsic movements The superposition is represented as ; to perform the action Quantization into j-dimensional quantum superposition state actions , Each of the qubits is and Superposition of two states; quantum superposition state actions Equivalent to j-dimensional quantum superposition state action The expression is: (10) in It is a j-dimensional quantum superposition state action Quantum actions after being observed; It is a quantum action Probability amplitude; Quantum Actions The modulus value satisfies ; When quantum superposition state action After being observed, It will collapse into a quantum action quantum action Each qubit is a or Quantum Actions Each qubit is represented by a value indicating its expected value; these expected values differ in different states; these expected values are used to select actions in a given state, and the update rules for these expected values are as follows: If in strategy In the middle, in the state Below, quantum superposition state action Collapse into quantum actions When cumulative rewards When adding, the action The qubits in each qubit are The expected value decreases, and the quantum bit is The expected value increases; the amount of increase and decrease in the expected value follows the strategy. Update the settings, and when selecting an action again in the same state, set the qubits with positive expected values for each qubit of the action as... The remaining quantum positions are , obtain quantum actions; let strategy The parameter vector is The specific action value is obtained by measuring the quantum action and then expressed through the parameter vector. Converted into power change As the output of Agent1 in this state; where parameter vector The parameters are as follows, where u is a positive integer; Quantum action selection obtains the state at time t using the Grover search method. Observing quantum superposition actions The quantum action obtained by the collapse The rules for obtaining and updating the search count V are as follows: First, all intrinsic actions are superimposed with equal weights by sequentially applying j Hadamard gates to j initial states, where each state is equal to a given value. The independent qubits yielded the following results: (11) Here, H is a Hadamard gate, which can convert the ground state... Convert to equal weight stacking ; It is the action of the quantum superposition state after initialization, by Actions with the same probability amplitude are superimposed together; This means applying j Hadamard gates sequentially to... One initialization intrinsic action; the two parts of the Grover iterative operator are: and They are respectively: (12) Where I is a unitary matrix of appropriate dimension; It is a quantum action The reverse quantum action; It is a quantum action The reverse quantum action; It is a quantum action The reverse quantum action; It is a quantum action The outer product; It is the action of the quantum superposition state after initialization. The outer product; and It is a quantum black box; Acting on intrinsic action hour, Change and The phase of states in the same direction changes by 180°. Acting on intrinsic action hour, Change and The phase of states in the same direction changes by 180°. Let Grover's iterative process be denoted as... : (13) according to For each intrinsic action, iterate to obtain the closest iteration count with the fewest iterations, and update the iteration count V corresponding to each intrinsic action, while also updating the policy. middle; From formula (9), we can obtain the state at time t. For any state At that time, the strategy is from Updated to Expected cumulative reward function for: (14) in For the updated strategy; In strategy State at time t Let be the transition probability distribution matrix for any state s; In strategy State at time t Let s be the dominance function for any state s; In strategy The state at time t Let be the transition probability distribution matrix for any state s; In strategy medium state Next action Perform sampling; In strategy The sampled value a under any state s in the middle; In strategy any state The probability distribution below; The state at time t Take action below Advantages compared to the average level; It is a strategy The expected cumulative reward function is as follows; In any state There is ,in For the updated strategy The state at time t Choosing the right action value will increase the cumulative reward. Or maintain cumulative rewards when the expected advantage is zero. The strategy remains unchanged, but the accumulated rewards are continuously updated and optimized. ; Because formula (14) requires the calculation of the updated strategy probability distribution This results in high complexity for formula (7), making it difficult to optimize. To reduce computational complexity, a substitution function is introduced. : (15) in It is a function that finds the largest independent variable in a given function; In strategy any state The probability distribution function under; substitution function and The difference lies in the substitution function. Ignore changes in state access density caused by policy changes. use As access frequency rather than , Access frequency By utilizing the strategy The approximate acquisition when the strategy and When certain constraints are met, the substitution function It can replace the original expected cumulative reward function ; In parameter vector In the update, the parameter vector is used. With any parameters The form of strategy Parameterization , In parameterization strategy Any action a in any state s; for any parameter When the strategy is not updated, the alternative function... and the original cumulative reward function They are completely equal, that is: (16) in It is a strategy After parameter vector Parameterized strategy; When the substitution function and the original cumulative reward function have any parameters The derivative in the strategy The places are the same, that is, the strategy is from Updated to When there is a very small change, if the value of the substitution function Increase, cumulative rewards This will also increase, so we can use the substitution function as the optimization objective to improve the strategy, that is: (17) in It concerns arbitrary parameters The differential of the function; Formulas (16) and (17) illustrate the strategy from Updated to A sufficiently small step size can increase cumulative rewards. ;definition Define the intermediate divergence variable as the strategy with the largest cumulative reward value among the old strategies. To increase cumulative rewards The lower bound is used to set a conservative iteration strategy. for: (18) in For the new strategy; This is the current strategy; for and The maximum total variance divergence between them; In strategy The action chosen in any state s; In strategy The action chosen in any state s; yes and Total variance divergence between them; In strategy Any action a chosen in any state s; In strategy Any action a chosen in any state s; For any random policy, let the intermediate entropy variable... ,in Let be the maximum absolute value of the function for choosing any action 'a' in any state 's'; using replace , replace ; Substitution function value and cumulative rewards The difference satisfies: (19) in It is a discount factor; The maximum relative entropy is ;in In strategy The action chosen in any state s; In strategy The action chosen in any state s; yes and The relative entropy between them; The relationship between total variance divergence and relative entropy satisfies ;in yes and Total variance divergence between them; make Reconstraining relative entropy: (20) Where C is the penalty coefficient; Under this constraint, for the constantly updated strategy ,exist ; in This indicates the policy update process; It is a policy sequence of a parallel trust optimization policy network; It is the cumulative reward of each policy in the policy sequence of the parallel trust optimization policy network; Considering parameterization strategy and parameter vector Remove the parameter vector Irrelevant items; Expected cumulative reward function after parameter variable transformation for: (21) Substitution function after parameter variable transformation for: (22) Relative entropy after parameter variable transformation for: (23) The constraints after the parameter variables are transformed are: (24) in The variable is transformed to obtain the same value; This is the parameter vector that needs to be updated; It is a parameter vector Updated parameter vector; It is a strategy After parameter vector Parameterized strategy; It is a strategy After parameter vector Parameterized strategy; It is a strategy The expected cumulative reward function; It is a strategy The substitution function; yes and The relative entropy between them; It is the maximum value of the relative entropy after the parameter variables have been transformed; The parameter vector of the parallel policy optimization network is obtained from formulas (21) to (24). The update process; through the parameter vector The update can optimize the selection weights of actions, thereby achieving the goal of optimizing parallel control; To ensure cumulative rewards Increase, make Maximize; because C is the penalty coefficient, it will cause each time The value becomes very small, resulting in a very short step size for each update, which slows down the update speed. Therefore, the penalty term is changed into a constraint term: (25) in It is a constant; Formula (14) is based on strategy Sampling is performed, due to the previous strategy. It is unknown and cannot be based on strategy. Therefore, importance sampling is used for the parameterized cumulative reward function. Rewrite; for the parameterized cumulative reward function Ignore and any parameters Irrelevant items, and use replace Finally, the update of the parallel trust optimization policy network becomes: (26) in It is about the parameterized strategy probability distribution and state-action value Perform sampling; In strategy medium state For any state s, the action Let a be the state-action value function for any action a; Based on the set constraints, the parameter vector Perform an update, using the updated parameter vector to update the strategy. The policy update in the parallel policy optimization network is completed, and then the new policy is used to select the action in the current state, and the process is iterated step by step. Step (9): After the iteration is completed, optimize the network according to the trained parallel trust strategy to regulate the power change in each power generation area of the new power system. This enables each region of the new power system to achieve the optimal tie-line power exchange performance index (CPS); each power generation region can achieve the optimal tie-line power exchange performance index (CPS) using the methods from steps (1) to (8); through training the network in each power generation region, regional cooperation is sought to achieve dynamic equilibrium, and finally the frequency deviation between each power generation region is reduced. As the power exchange performance index (CPS) approaches zero, the overall power exchange performance index (CPS) approaches 100%, and the entire new power system gradually reaches global optimization.