Integrated energy system control method for migration parallel deep reinforcement learning, medium and processor

Through transfer parallel deep reinforcement learning combined with quantum generative adversarial networks, Transformer and soft actor critics, the problem of insufficient control accuracy and flexibility in complex integrated energy systems is solved, and efficient control is achieved to quickly adapt to environmental changes.

CN120258435APending Publication Date: 2025-07-04GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510370029.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing intelligent power generation control method faces complex comprehensive energy systems, especially when the energy equipment capacity, coupling method, time constant and hill climb coefficient changes, the control accuracy is poor and the training time is long, resulting in supply and demand imbalance and system instability.

Method used

Transfer parallel deep reinforcement learning method is adopted, combining quantum generative adversarial network, Transformer and soft actor critics, through transfer learning and parallel system training models, quickly adjust control strategies, use quantum generative adversarial network and Transformer to predict, and soft actor critics output control instructions to ensure the stability of the system.

Benefits of technology

It improves the accuracy and flexibility of control in complex integrated energy systems, can quickly adapt to environmental changes, reduce training time, and ensure the stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258435A_ABST
    Figure CN120258435A_ABST
Patent Text Reader

Abstract

The invention provides an integrated energy system control method for migration parallel deep reinforcement learning, a medium and a processor, and the method comprises the steps: firstly, an integrated energy system central server collects the environment characteristics, frequency deviation and integrated energy system balance error of an integrated energy system; secondly, the central server of the integrated energy system constructs a parallel system for training according to the collected environmental characteristics, searches a model with the highest control accuracy on the integrated energy system, and updates parameters of the quantum generative adversarial network, the Transform and the soft actor commentator through transfer learning according to parameters of the model; the quantum generative adversarial network and the Transform are used for predicting the frequency deviation and the balance error of the integrated energy system, a soft actor commentator is used for outputting a control instruction, and the complex integrated energy system control problem that the coupling mode of the energy equipment of the integrated energy system changes or the parameters of the energy equipment changes can be efficiently solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the control field of power systems, multi-region integrated energy systems and integrated energy systems with environmental changes, and involves parallel systems, transfer learning, deep learning, quantum machine learning and deep reinforcement learning control methods, and is applicable to power generation control of integrated energy systems with environmental changes. Background Art

[0002] An integrated energy system refers to combining different types of energy systems through energy coupling and coordinated control to achieve efficient utilization and optimal allocation of energy. However, with the continuous access of a large number of renewable energy generating units, the existing intelligent power generation control methods have poor control accuracy for complex integrated energy systems. At the same time, the flexibility of intelligent power generation control methods is relatively low, and the control accuracy is poor when facing changes in the capacity of energy equipment, the coupling mode of energy equipment, the time constant of energy equipment, and the ramp coefficient of energy equipment in the integrated energy system.

[0003] In addition, when facing changes in the capacity of energy equipment, the coupling mode of energy equipment, the time constant of energy equipment, and the ramp coefficient of energy equipment in the integrated energy system environment, deep learning methods and deep reinforcement learning methods need to consume a large amount of training time, and the system may not be able to adjust in time, resulting in supply-demand imbalance and even system instability.

[0004] Therefore, a control method for an integrated energy system based on transfer parallel deep reinforcement learning is proposed, using a quantum generative adversarial network, a Transformer and a soft actor-critic method to solve the problem of poor control accuracy of intelligent power generation control methods in complex integrated energy systems. The quantum generative adversarial network has advantages in processing high-dimensional data and capturing complex patterns, and can accurately predict the frequency deviation signals of each region in the complex integrated energy system; the Transformer has high flexibility and excellent parallel computing power, and can efficiently predict the integrated energy system balance error signals of each region in the complex integrated energy system; the soft actor-critic has high learning ability and high adaptive adjustment ability, and can adaptively optimize the control strategy when changes occur in the capacity of energy equipment, the coupling mode of energy equipment, the time constant of energy equipment, and the ramp coefficient of energy equipment in the integrated energy system, reducing the frequency deviation and integrated energy system balance error of the integrated energy system. Using parallel systems and transfer learning, the training speed of the quantum generative adversarial network, the Transformer and the soft actor-critic method is accelerated, and the strategy is optimized and adjusted quickly when facing changes in the capacity of energy equipment, the coupling mode of energy equipment, the time constant of energy equipment, and the ramp coefficient of energy equipment in the system, ensuring the stable operation of the integrated energy system. Summary of the Invention

[0005] The present invention proposes a control method for an integrated energy system based on transfer parallel deep reinforcement learning, which combines transfer learning and parallel systems to train a model. The quantum generative adversarial network and Transformer are used to predict the system frequency deviation and the balance error of the integrated energy system respectively, and the soft actor-critic is used to control the integrated energy system. The central server of the integrated energy system collects the environmental characteristics, frequency deviation and balance error of the integrated energy system. The environmental characteristics include the capacity of energy equipment, the time constant of energy equipment, the ramp coefficient of energy equipment and the coupling mode of energy equipment. A parallel system is constructed according to the collected environmental characteristics. The parallel system contains parallel control systems. The environmental characteristics of each parallel control system are the same as those of the integrated energy system. Each parallel control system has a quantum generative adversarial network, a Transformer and a soft actor-critic with different weights and biases from those of the integrated energy system. To improve the model training speed of each parallel control system in the parallel system, a teacher model and a student model are included in each parallel control system. The teacher model consists of parallel prediction systems, and the student model consists of a student quantum generative adversarial network and a student Transformer in the parallel control system. The collected frequency deviation and balance error of the integrated energy system are divided into big data and small data in a ratio of 9:1. The big data is used to train the teacher model through parallel prediction systems. Every time, the weights and biases of the quantum generative adversarial network with the highest prediction accuracy for the frequency deviation of the integrated energy system and the weights and biases of the Transformer with the highest prediction accuracy for the balance error of the integrated energy system are selected in parallel prediction systems and copied to the student model of the parallel control system through transfer learning. The updated student model is used to predict the small data, and the prediction results of the student model for the small data are input into the soft actor-critic, and the soft actor-critic outputs control instructions. Every time, select The parameters of the controller of the parallel control system with the minimum frequency deviation of the integrated energy system and the balance error of the integrated energy system are copied to the controller of the integrated energy system through transfer learning; the parameters of the controller are the weights and biases of the quantum generative adversarial network, the weights and biases of the Transformer, and the weights and biases of the soft actor-critic; after the parameters of the controller of the integrated energy system are updated, the quantum generative adversarial network and the Transformer are used to predict the system frequency deviation and the balance error of the integrated energy system respectively, and the control instructions are output through the soft actor-critic; the control method of the integrated energy system based on transfer parallel deep reinforcement learning proposed by the present invention can efficiently adjust the parameters of the controller, adapt to the changes of the environment, and give accurate control instructions when the capacity of the energy equipment, the coupling mode of the energy equipment, the time constant of the energy equipment, and the ramp coefficient of the energy equipment in the integrated energy system change; the steps of the control method of the integrated energy system based on transfer parallel deep reinforcement learning in the using process are as follows:

[0006] Step (1): Establish the operation framework of the integrated energy system. In the integrated energy system The power supply and demand balance at time

[0007] (1)

[0008] Wherein, is the number of regions in the integrated energy system; is The total electric energy output by all power supply equipment in the th region of the integrated energy system at time is The total electric energy consumed by all electrical equipment in the th region of the integrated energy system at time

[0009] The power supply and demand balance in the th region of the integrated energy system at time is:

[0010] (2)

[0011] Wherein, is the electric energy transmitted through the tie line in the th region of the integrated energy system at time ;

[0012] The heat supply and demand balance in the integrated energy system at time is:

[0013] (3)

[0014] Wherein, for The first in the integrated energy system The total heat output of all heating equipment in a region; for The first in the integrated energy system The total heat energy consumed by all heat-using equipment in a region;

[0015] Integrated energy system Regions The heat supply and demand balance at the moment is:

[0016] (4)

[0017] in, The first Areas in The interconnecting lines at all times transmit heat energy;

[0018] Integrated energy system The gas supply and demand balance at the moment is:

[0019] (5)

[0020] in, for The first moment in the integrated energy system The total gas output of all gas supply equipment in a region; for The first in the integrated energy system The total gas consumption of all gas-using equipment in a region;

[0021] Integrated energy system Regions The gas supply and demand balance at the moment is:

[0022] (6)

[0023] in, The first Areas in The amount of gas transmitted by the interconnection line at each moment;

[0024] Integrated energy system Areas in Frequency deviation at time for:

[0025] (7)

[0026] in, The first The frequency of a region at moment;

[0027] In the integrated energy system, the th region at The integrated energy system balance error at moment is as follows:

[0028] (8)

[0029] Step (2): The central server of the integrated energy system collects the environmental characteristics of the integrated energy system, where the environmental characteristics are the capacity of the energy equipment, the time constant of the energy equipment, the ramp coefficient of the energy equipment, and the coupling method of the energy equipment; the central server of the integrated energy system uses transfer learning to establish a parallel system containing parallel control systems. Each parallel control system contains a student model, a teacher model, and a soft actor-critic; the student model contains 1 student quantum generative adversarial network and 1 student Transformer; the weights and biases of the student quantum generative adversarial networks in each parallel control system are different; the weights and biases of the student Transformers in each parallel control system are different; the weights and biases of the soft actor-critics in each parallel control system are different; the teacher model is constructed according to the structure of the student model parallel prediction systems. Each parallel prediction system contains 1 teacher quantum generative adversarial network with the same structure as the student quantum generative adversarial network and 1 teacher Transformer with the same structure as the student Transformer; the weights and biases of the teacher quantum generative adversarial networks in each parallel prediction system are different; the weights and biases of the teacher Transformers in each parallel prediction system are different;

[0030] At moment, each parallel control system central server collects the frequency deviation of each region of the virtual integrated energy system and the integrated energy system balance error of each region. Among them, the th parallel control system central server collects the frequency deviation of each region of the virtual integrated energy system at moment as follows:

[0031] (9)

[0032] Among them, , and are respectively the th parallel control system central server at Collect the frequency deviations of the 1st area, 2nd area and the th area of the virtual integrated energy system at the moment;

[0033] The th central server of the parallel control system collects the integrated energy system balance error of each area of the virtual integrated energy system at the moment as: :

[0034] (10)

[0035] where , and are respectively the integrated energy system balance errors of the 1st area, 2nd area and the th area of the virtual integrated energy system collected by the th central server of the parallel control system at the moment;

[0036] The th central server of the parallel control system divides the collected and into big data and small data according to the ratio of 9:1; the big data includes big data of frequency deviation and big data of integrated energy system balance error , and the small data includes small data of frequency deviation and small data of integrated energy system balance error ;

[0037] Step (3): Input the collected into the teacher model containing parallel prediction systems for training; the signal distribution input to the quantum generative adversarial network is ;

[0038] The quantum generator in the quantum generative adversarial network first randomly samples some random noise vectors from the noise distribution , then converts into a quantum state , then converts the quantum state into a generated state , and finally measures the generated state to obtain a generated data sample :

[0039] The quantum discriminator randomly samples some real data samples from the distribution of real samples , and combines and Input data samples that constitute a quantum discriminator ; The quantum discriminator takes the input data samples and converts them into a quantum state , and then converts them into a discrimination state ; The quantum discriminator evaluates the discrimination state and outputs a classification probability value , which includes the output probability value of the quantum discriminator for the input real data samples and the output probability value of the quantum discriminator for the input generated data samples ;

[0040] Calculate the loss function of the quantum generator :

[0041] (11)

[0042] wherein, is the expected value randomly sampled from ; ; is the logarithmic function;

[0043] Calculate the loss function of the quantum discriminator :

[0044] (12)

[0045] wherein, is the expected value randomly sampled from ; ;

[0046] By minimizing the loss function of the quantum generator, the probability that the quantum discriminator considers the generated data samples as real data samples is made as large as possible; by minimizing the loss function of the quantum discriminator, the ability of the quantum discriminator to distinguish between real data samples and generated data samples is improved;

[0047] Update the weights and biases of the quantum generator by the gradient descent method :

[0048] (13)

[0049] wherein, is the learning rate of the quantum generator; is the gradient of the loss function of the quantum generator with respect to ;

[0050] Update the weights and biases of the quantum discriminator by the gradient descent method :

[0051] (14)

[0052] Among them, is the learning rate of the quantum discriminator; is the gradient of the loss function of the quantum discriminator with respect to ;

[0053] Repeat step (3) until the training converges;

[0054] Step (4): Input the collected into the teacher model containing parallel prediction systems for training; The comprehensive energy system balance error vector input to the Transformer is ;

[0055] The self-attention mechanism calculates the correlation between the features of the input vector based on the attention features:

[0056] (15)

[0057] (16)

[0058] (17)

[0059] Among them, is the query matrix; is the key matrix; is the value matrix; , and are the weight matrices for training;

[0060] The self-attention mechanism calculates the attention scores and the output of the attention heads :

[0061] (18)

[0062] Among them, is the normalization function; is the transpose of the key matrix ; is the dimension of the key vector;

[0063] The multi-head attention mechanism concatenates the outputs of attention heads to obtain the output of the multi-head attention mechanism:

[0064] (19)

[0065] Among them, Represents a splicing operation; 、 and are the outputs of the first attention head, the second attention head, and the output of the attention head, respectively; The linear transformation matrix after splicing;

[0066] The feed-forward neural network consists of two linear transformation layers and an activation function. After passing through the feed-forward neural network, is non-linearly transformed:

[0067] (20)

[0068] where, is the activation function; is the weight of the first linear transformation layer in the feed-forward neural network; is the weight of the second linear transformation layer in the feed-forward neural network; is the bias of the first linear transformation layer in the feed-forward neural network; is the bias of the second linear transformation layer in the feed-forward neural network;

[0069] is the output of the self-attention mechanism or the output of the feed-forward neural network; After layer normalization and adding residual connections to the outputs of each multi-head attention mechanism and the output of the feed-forward neural network, the output of the encoder module is obtained :

[0070] (21)

[0071] where, is the layer normalization operation;

[0072] After linear transformation and activation function, the final output of the Transformer is obtained :

[0073] (22)

[0074] where, is the linear transformation;

[0075] Calculate the loss function of the Transformer :

[0076] (23)

[0077] where, is the input vector The number of elements in; is the predicted value of the th element in the input vector; is the true value of the th element in the input vector;

[0078] Calculate the gradient of the loss function using backpropagation and update the weights and biases of the Transformer :

[0079] (24)

[0080] where is the learning rate of the Transformer; is the gradient of the loss function of the Transformer with respect to ;

[0081] Repeat step (4) until the training converges;

[0082] Step (5): Every th parallel control system central server copies the weights and biases of the quantum generative adversarial network with the highest prediction accuracy in the parallel prediction systems in the teacher model to the student model together with the weights and biases of the Transformer through transfer learning; the student quantum generative adversarial network predicts the input to obtain the prediction result ; the student Transformer predicts the input to obtain the prediction result ; form the prediction results of the student model into a binary tuple and input it into the soft actor-critic for training; feedback the prediction accuracy of the student model to the teacher model through transfer learning and update the weights and biases of the quantum generative adversarial network and the Transformer in the teacher model;

[0083] Step (6): The soft actor-critic includes a policy network , a Q-value network , a Q-value network , a target Q-value network , a target Q-value network and an experience replay pool ;

[0084] The binary tuple input into the soft actor-critic is the current state ; according to the current policy , obtain the mean and standard deviation of the action distribution:

[0085] (25)

[0086] where is the mean of the action distribution; is the standard deviation of the action distribution;

[0087] Sample an action at the current state from the action distribution :

[0088] (26)

[0089] where represents the action distribution;

[0090] Execute the action , observe the state and reward at the next time step; store the transition samples of the state, action, reward, and the state at the next time step at the current state in the experience replay pool; randomly sample a batch of transition samples from the experience replay pool for training;

[0091] Sample an action for the state at the next time step according to the current policy of the policy network :

[0092] (27)

[0093] where is the state at the next time step all possible actions that can be obtained; represents at state under the current policy obtain probability;

[0094] Calculate the target Q value :

[0095] (28)

[0096] where is the discount factor; can take 1 and 2, and are the parameters of the target Q value network and respectively; is the target Q value network for the state at the next time step the action at the next time step The predicted value, and take the minimum value of the predicted values of the two target Q-value networks; is the entropy weight coefficient; represents at state under the current policy obtaining the probability of;

[0097] By minimizing the loss function of the Q-value network update the weights and biases of the Q-value network of and the weights and biases of the Q-value network of :

[0098] (29)

[0099] where is the expected value randomly sampled from ; the expected value of; is the predicted value of the Q-value network or for the current state and the action in the current state ;

[0100] By minimizing the loss function of the policy network update the weights and biases of the policy network :

[0101] (30)

[0102] where is the expected value randomly sampled from ; the expected value of; is the expected value randomly sampled from ; the expected value of; represents at state under the current policy obtaining the probability of;

[0103] Update the weights and biases of the target Q-value network of and the weights and biases of the target Q-value network of :

[0104] (31)

[0105] where is the update coefficient;

[0106] Repeat step (6) until the maximum number of training steps is reached;

[0107] Step (7): Every , select the parallel control system with the smallest frequency deviation of the integrated energy system and the smallest balance error of the integrated energy system among parallel control systems; through transfer learning, use the parameters of the controller in the parallel control system with the smallest frequency deviation of the integrated energy system and the smallest balance error of the integrated energy system to update the parameters of the controller in the integrated energy system; the parameters of the controller are the weights and biases of the quantum generative adversarial network, the weights and biases of the Transformer, and the weights and biases of the soft actor-critic;

[0108] After updating the parameters, the central server of the integrated energy system predicts the system frequency deviation through the quantum generative adversarial network, predicts the balance error of the integrated energy system through the Transformer, and outputs control instructions through the soft actor-critic; adjust the output power of the energy device;

[0109] Step (8): When the environmental characteristics of the integrated energy system change, repeat steps (2) to (7).

[0110] A computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned control method of the integrated energy system for transfer parallel deep reinforcement learning.

[0111] A processor, the processor is used to run a program, wherein when the program runs, it executes the above-mentioned control method of the integrated energy system for transfer parallel deep reinforcement learning. The present invention has the following advantages and effects compared with the prior art:

[0112] (1) The existing intelligent power generation control methods have limited processing capabilities when facing a more complex and non-linear power system caused by the further large-scale grid connection of renewable energy. The deep reinforcement learning soft actor-critic in the present invention uses a neural network as a function approximator to capture complex patterns and non-linear relationships in the system, and has stronger capabilities to handle complex environments and decision-making problems.

[0113] (2) The existing intelligent power generation control methods have limited adaptability. When the capacity of energy devices, the coupling method of energy devices, the time constant of energy devices, and the ramp coefficient of energy devices in the integrated energy system change, the control accuracy for complex integrated energy systems is poor. However, the present invention uses a deep reinforcement learning method to control the integrated energy system, which can perform online learning and update control strategies as the system environment changes, thus effectively coping with the dynamic changes in the coupling method, time constant, ramp coefficient, and capacity of energy devices in the integrated energy system.

[0114] (3) When dealing with time series signals, a single deep reinforcement learning method has limited feature extraction ability. However, the quantum generative adversarial network and Transformer adopted by the present invention have high feature representation ability. The quantum generative adversarial network can generate data distributions with rich features by combining the parallelism of quantum computing and the generative adversarial mechanism of the generative adversarial network; the Transformer can effectively capture the long-term dependencies and complex patterns in time series data through the self-attention mechanism. The information predicted by the quantum generative adversarial network and Transformer contains richer features. The present invention combines deep reinforcement learning, quantum generative adversarial network, and Transformer, enabling deep reinforcement learning to learn richer features in the predicted information, thereby improving the control accuracy of deep reinforcement learning in power generation control.

[0115] (4) The Soft Actor-Critic, quantum generative adversarial network, and Transformer have a large number of parameters and slow training speed, resulting in limited ability to adjust and update control strategies in response to the dynamic changes in the coupling method, time constant, ramp coefficient, and capacity of energy devices in the integrated energy system environment. However, the present invention combines the Soft Actor-Critic, quantum generative adversarial network, Transformer, parallel system, and transfer learning. The parallel system provides rich training data and diverse learning environments by simulating multiple possible system states and operation scenarios, and transfer learning uses the model parameters pre-trained in similar tasks to adapt to the new environment by fine-tuning the model parameters, thereby reducing the training time and accelerating the training speed of the Soft Actor-Critic, quantum generative adversarial network, and Transformer when the coupling method, time constant, ramp coefficient, and capacity of energy devices in the integrated energy system environment change, ensuring that the control strategy can be optimized and adjusted in a timely manner in the dynamic and complex integrated energy system environment, and guaranteeing the stability and efficient operation of the system. Description of the Drawings

[0116] Figure 1 is the control framework diagram of the integrated energy system of the method of the present invention.

[0117] Figure 2 It is the parallel system framework diagram of the method of the present invention.

[0118] Figure 3 It is the control flow diagram of the integrated energy system of the method of the present invention. Specific embodiments

[0119] A control method, medium and processor for an integrated energy system based on transfer parallel deep reinforcement learning proposed by the present invention will be described in detail with reference to the accompanying drawings as follows:

[0120] Figure 1 It is the control framework diagram of the integrated energy system of the method of the present invention. First, the central server of the integrated energy system collects the frequency deviation, the balance error of the integrated energy system and the environmental characteristics of the integrated energy system. The environmental characteristics are the capacity of the energy equipment of the integrated energy system, the coupling mode of the energy equipment, the time constant of the energy equipment and the ramp coefficient of the energy equipment. Then, a parallel system is constructed through transfer learning according to the collected environmental characteristics. The parallel system contains parallel control systems. Each parallel control system contains 1 student model and 1 teacher model. Every time, select the parallel control system with the smallest frequency deviation and balance error of the integrated energy system among the parallel control systems, and update the controller parameters of the integrated energy system according to the parameters of its controller. The parameters of the controller are the weights and biases of the quantum generative adversarial network, the weights and biases of the Transformer, and the weights and biases of the soft actor-critic. Finally, the central server of the integrated energy system predicts the system frequency deviation through the quantum generative adversarial network, predicts the balance error of the integrated energy system through the Transformer, inputs the frequency deviation prediction result and the balance error prediction result of the integrated energy system into the soft actor-critic, and the soft actor-critic outputs a control instruction to adjust the output power of the energy equipment.

[0121] Figure 2 It is the parallel system framework diagram of the method of the present invention. First, the central server of the parallel control system collects the frequency deviation of the system and the balance error of the integrated energy system. Then, the central server of the parallel control system divides the collected frequency deviation signal and balance error signal of the integrated energy system into big data and small data according to a ratio of 9:1. The central server of the parallel control system constructs a teacher model containing parallel prediction systems according to the structure of the student model. Input the frequency deviation big data in the big data into the teacher quantum generative adversarial network in the parallel prediction system for training. Input the balance error big data of the integrated energy system in the big data into the teacher Transformer in the parallel prediction system for training. Every time, select The parallel prediction system with the highest prediction accuracy in the parallel prediction systems updates the weights and biases of the student quantum generative adversarial network and the student Transformer respectively according to the weights and biases of its teacher quantum generative adversarial network and teacher Transformer. Then, the student model makes predictions on the input small data. The student quantum generative adversarial network is used to predict the small data of the frequency deviation of the input, and the student Transformer is used to predict the small data of the balance error of the integrated energy system of the input. The prediction results of the small data of the frequency deviation and the prediction results of the small data of the balance error of the integrated energy system are input into the soft actor-critic, and the soft actor-critic outputs control instructions to adjust the output power of the energy equipment. Finally, the accuracy rate of the prediction by the student model is fed back to the teacher model through transfer learning.

[0122] Figure 3 is the control flow chart of the integrated energy system of the method of the present invention. First, the central server of the integrated energy system collects the environmental characteristics, frequency deviation, and balance error of the integrated energy system. The environmental characteristics are the capacity of the energy equipment of the integrated energy system, the coupling method of the energy equipment, the time constant of the energy equipment, and the ramp coefficient of the energy equipment. Then, the central server of the integrated energy system constructs a parallel system including parallel control systems according to the collected environmental characteristics. The parallel control system includes a teacher model, a student model, and a soft actor-critic. The student model includes 1 student quantum generative adversarial network and 1 student Transformer; the teacher model in the parallel control system constructs parallel prediction systems according to the structure of the student model. Each parallel prediction system includes 1 teacher quantum generative adversarial network with the same structure as the student quantum generative adversarial network and 1 teacher Transformer with the same structure as the student Transformer. The weights and biases of the teacher quantum generative adversarial network in each parallel prediction system are different from each other. The weights and biases of the teacher Transformer in each parallel prediction system are different from each other. Then, the parallel control system collects the frequency deviation data and the balance error data of the system, and divides the collected data into big data and small data according to 9:1, and inputs the big data into the teacher model for training. Every time, In the parallel prediction system with the highest prediction accuracy, the weights and biases of the teacher quantum generative adversarial network and the teacher Transformer in the parallel prediction system are copied to the student model through transfer learning. In the student model, the student quantum generative adversarial network predicts the small data of the input frequency deviation, and the student Transformer predicts the small data of the balance error of the integrated energy system in the input. The prediction results obtained by the student model are input into the soft actor-critic for training, and the prediction accuracy is fed back to the teacher model through transfer learning to update the weights and biases of the teacher model. Every time, the controller parameters of the parallel control system with the smallest frequency deviation of the integrated energy system and the balance error of the integrated energy system in the parallel control systems are copied to the central server of the integrated energy system through transfer learning. The central server of the integrated energy system updates the parameters of the controller in the integrated energy system according to the controller parameters of the parallel control system with the smallest frequency deviation of the integrated energy system and the balance error of the integrated energy system. The parameters of the controller are the weights and biases of the quantum generative adversarial network, the weights and biases of the Transformer, and the weights and biases of the soft actor-critic. Finally, the central server of the integrated energy system predicts the system frequency deviation signal through the quantum generative adversarial network, predicts the balance error of the integrated energy system through the Transformer, and outputs control instructions through the soft actor-critic to adjust the output power of the energy equipment.

[0123] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A comprehensive energy system control method for migrating parallel deep reinforcement learning, characterized in that, The migration learning and parallel system are combined to train a model. The quantum generative adversarial network and Transformer are used to predict the system frequency deviation and the integrated energy system balance error respectively, and the soft actor-critic is used to control the integrated energy system. The steps in the process of a control method for an integrated energy system based on migration parallel deep reinforcement learning are as follows: Step (1): Establish an operation framework for the integrated energy system. The power supply-demand balance at the moment in the integrated energy system is as follows: (1) Among them, is the number of regions in the integrated energy system; is the total electric energy output by all power supply devices in the th region of the integrated energy system at time is the total electric energy consumed by all electrical equipment in the th region of the integrated energy system at time The power supply and demand balance at the th area moment in the integrated energy system is as follows: (2) Among them, is the electric energy transmitted through the tie line of the th area in the integrated energy system at time; In the integrated energy system The thermal energy supply-demand balance at time (3) Among them, is the total thermal energy output by all heating equipment in the th area of the integrated energy system at a certain moment; is the total thermal energy consumed by all heat-using equipment in the th area of the integrated energy system at a certain moment; The thermal energy supply-demand balance at the nth moment in the integrated energy system is as follows: (4) Among them, is the heat energy transmitted through the tie line of the th area in the integrated energy system at moment; In the integrated energy system The gas supply-demand balance at time (5) Among them, is the total gas volume output by all gas supply devices in the th area of the integrated energy system at a certain moment; is the total gas volume consumed by all gas-using devices in the th area of the integrated energy system at a certain moment; The gas volume supply-demand balance at the nth moment in the integrated energy system is as follows: (6) Among them, is the gas transmission volume of the tie line for the th area in the integrated energy system at time; The frequency deviation of the th area in the integrated energy system at moment is as follows: (7) Among them, is the frequency of the th area in the integrated energy system at moment; The th area in the integrated energy system at time has an integrated energy system balance error of as follows: (8) Step (2): The central server of the integrated energy system collects the environmental characteristics of the integrated energy system, where the environmental characteristics are the capacity of the energy equipment, the time constant of the energy equipment, the ramp coefficient of the energy equipment, and the coupling method of the energy equipment; the central server of the integrated energy system uses transfer learning to establish a parallel system containing parallel control systems, each parallel control system contains a student model, a teacher model, and a soft actor-critic; the student model contains 1 student quantum generative adversarial network and 1 student Transformer; the weights and biases of the student quantum generative adversarial networks in each parallel control system are different; the weights and biases of the student Transformers in each parallel control system are different; the weights and biases of the soft actor-critics in each parallel control system are different; the teacher model is constructed according to the structure of the student model parallel prediction systems, each parallel prediction system contains 1 teacher quantum generative adversarial network with the same structure as the student quantum generative adversarial network and 1 teacher Transformer with the same structure as the student Transformer; the weights and biases of the teacher quantum generative adversarial networks in each parallel prediction system are different; the weights and biases of the teacher Transformers in each parallel prediction system are different; Each parallel control system central server at collects the frequency deviation of each area of the virtual integrated energy system and the integrated energy system balance error of each area at the moment, where the th parallel control system central server at collects the frequency deviation of each area of the virtual integrated energy system at the moment is: (9) Among them, , and are respectively the frequency deviations of the 1st area, the 2nd area and the th area of the virtual integrated energy system collected by the th parallel control system central server at the th moment; The parallel control system central server collects the integrated energy system balance error of each region of the virtual integrated energy system at the moment, and the result is: ​ (10) Among them, , and are respectively the th parallel control system central servers collecting the integrated energy system balance errors of the 1st area, the 2nd area and the th area of the virtual integrated energy system at moment; The parallel control system central server divides the collected and into big data and small data in a ratio of 9:1; the big data includes big data of frequency deviation and big data of balance error of the integrated energy system , and the small data includes small data of frequency deviation and small data of balance error of the integrated energy system ; Step (3): Input the collected into the teacher model containing parallel prediction systems for training; the signal distribution input into the quantum generative adversarial network is ; The quantum generator in the quantum generative adversarial network first distributes the noise Randomly sample some random noise vectors in , and then Converted to quantum state , and then the quantum state Convert to generated state , and finally the generated state Take measurements and generate data samples : The quantum discriminator randomly samples some real data samples from the distribution of real samples and combines with to form the input data samples of the quantum discriminator ; The quantum discriminator converts the input data samples into a quantum state , and then converts into a discrimination state ; The quantum discriminator evaluates the discrimination state and outputs a classification probability value , , which includes the output probability value of the quantum discriminator for the input real data samples and the output probability value of the quantum discriminator for the input generated data samples ; Calculating the loss function of a quantum generator : (11) Among them, is sampled randomly from to obtain the expected value; is the logarithmic function; Calculate the loss function of the quantum discriminator : (12) wherein, is randomly sampled from to obtain the expected value; By minimizing the quantum generator loss function, the probability that the quantum discriminator considers the generated data sample as a real data sample is made as large as possible. By minimizing the loss function of the quantum discriminator, the ability of the quantum discriminator to distinguish between real data samples and generated data samples is improved. Update the weights and biases of the quantum generator by gradient descent : (13) Among them, is the learning rate of the quantum generator; is the gradient of the loss function of the quantum generator with respect to ; Update the weights and biases of the quantum discriminator by gradient descent : (14) Among them, is the learning rate of the quantum discriminator; is the gradient of the loss function of the quantum discriminator with respect to ; Repeat step (3) until the training converges. Step (4): Input the collected into the teacher model containing parallel prediction systems for training; the comprehensive energy system balance error vector input to the Transformer is ; The self-attention mechanism calculates the correlation between the features of the input vector based on the attention features: (15) (16) (17) Among them, is the query matrix; is the key matrix; is the value matrix; , and are the trained weight matrices; Self-attention mechanism calculates attention scores and the output of attention heads : (18) wherein, is a normalization function; is the key matrix transpose; is the dimension of the key vector; Concatenation of multi-head attention mechanism The outputs of attention heads are concatenated to obtain the output of the multi-head attention mechanism : (19) Among them, represents a splicing operation; , and are respectively the outputs of the first attention head, the outputs of the second attention head, and the outputs of the th attention head; The linear transformation matrix after splicing; The feedforward neural network consists of two linear transformation layers and an activation function. After passing through the feedforward neural network, a non-linear transformation is performed: (20) Among them, is the activation function; is the weight of the first linear transformation layer in the feedforward neural network; is the weight of the second linear transformation layer in the feedforward neural network; is the bias of the first linear transformation layer in the feedforward neural network; is the bias of the second linear transformation layer in the feedforward neural network; is the output of the self-attention mechanism or the output of the feed-forward neural network; after performing layer normalization on the output of each multi-head attention mechanism and the output of the feed-forward neural network and adding residual connections, the output of the encoder module is obtained : (21) Among them, is a layer normalization operation; After linear transformation and activation function, the final output of the Transformer is obtained : (22) Among them, is a linear transformation; Calculate the loss function of the Transformer : (23) Among them, is the number of elements in the input vector; is the predicted value of the th element in the input vector; is the true value of the th element in the input vector; Calculate the gradient of the loss function using backpropagation and update the weights and biases of the Transformer : (24) Among them, is the learning rate of the Transformer; is the gradient of the loss function of the Transformer with respect to ; Repeat step (4) until the training converges. Step (5): Every time, the th parallel control system central server copies the weights and biases of the quantum generative adversarial network with the highest prediction accuracy in the parallel prediction systems in the teacher model and the weights and biases of the Transformer to the student model through transfer learning; the student quantum generative adversarial network predicts the input to obtain the prediction result ; the student Transformer predicts the input to obtain the prediction result ; the prediction results of the student model are formed into a binary tuple and input into the soft actor-critic for training; the accuracy rate predicted by the student model is fed back to the teacher model through transfer learning to update the weights and biases of the quantum generative adversarial network and the weights and biases of the Transformer in the teacher model; Step (6): The soft actor-critic includes a policy network , a Q-value network , a Q-value network , a target Q-value network , a target Q-value network and an experience replay pool ; The tuple input to the soft actor-critic is the current state ; according to the current policy of the policy network , the mean and standard deviation of the action distribution are obtained: (25) Among them, is the mean of the action distribution; is the standard deviation of the action distribution; Sample an action in the current state from the action distribution : (26) Among them, represents the action distribution; Execute an action , and observe the state at the next time step and the reward ; transfer the state, action, reward, and the transition sample of the state at the next time step in the current state to the experience replay pool; randomly sample a batch of transition samples from the experience replay pool for training; Sample the action of the state at the next time step according to the current policy of the policy network : (27) Among them, is the state at the next time step all possible actions that may be obtained next; represents in the state under the current policy obtained probability; Calculate the target Q value : (28) Among them, is the discount factor; can take 1 and 2, and are the parameters of the target Q-value network and respectively; is the predicted value of the target Q-value network for the state at the next time step and the action at the next time step , and takes the minimum value of the predicted values of the two target Q-value networks; is the entropy weight coefficient; represents the probability of obtaining under the state according to the current policy ; By minimizing the loss function of the Q-value network Update the Q-value network weights and biases and the Q-value network weights and biases : (29) Among them, is sampled randomly from to obtain the expected value; is the Q-value network or for the current state and the action in the current state the predicted value; By minimizing the loss function of the policy network Update the weights and biases of the policy network : (30) wherein, is randomly sampled from to obtain the expected value of ; is randomly sampled from to obtain the expected value of ; represents the probability of obtaining under the current policy in the state ; Update the weights and biases of the target Q-value network and the weights and biases of the target Q-value network as well as the weights and biases of the target Q-value network : : (31) Among them, is the update coefficient; Repeat step (6) until the maximum number of training steps is reached. Step (7): Every time, select parallel control system with the smallest frequency deviation and integrated energy system balance error in the parallel control systems of the integrated energy system; through transfer learning, use the parameters of the controller in the parallel control system with the smallest frequency deviation and integrated energy system balance error of the integrated energy system to update the parameters of the controller in the integrated energy system; the parameters of the controller are the weights and biases of the quantum generative adversarial network, the weights and biases of the Transformer, and the weights and biases of the soft actor-critic. After updating the parameters, the central server of the integrated energy system predicts the system frequency deviation through the quantum generative adversarial network, predicts the integrated energy system balance error through the Transformer, and outputs control instructions through the soft actor-critic; adjust the output power of the energy device. Step (8): When the environmental characteristics of the integrated energy system change, repeat steps (2) to (7).

2. The method according to claim 1, wherein The central server of the integrated energy system collects the environmental characteristics, frequency deviation and integrated energy system balance error of the integrated energy system. The environmental characteristics include the capacity of the energy device, the time constant of the energy device, the ramp coefficient of the energy device and the coupling mode of the energy device; a parallel system is constructed according to the collected environmental characteristics. The parallel system includes parallel control systems. The environmental characteristics of each parallel control system are the same as those of the integrated energy system. Each parallel control system has a quantum generative adversarial network, a Transformer, and a soft actor-critic with different weights and biases from those of the integrated energy system. To improve the model training speed of each parallel control system in the parallel system, a teacher model and a student model are further included in each parallel control system. The teacher model is composed of parallel prediction systems. The student model is composed of a student quantum generative adversarial network and a student Transformer in the parallel control system. The collected frequency deviation and the balance error of the integrated energy system are divided into big data and small data in a ratio of 9:

1. The big data is used to train the teacher model through parallel prediction systems. Every time, the weights and biases of the quantum generative adversarial network with the highest prediction accuracy of the frequency deviation of the integrated energy system and the weights and biases of the Transformer with the highest prediction accuracy of the balance error of the integrated energy system are selected in parallel prediction systems and copied to the student model of the parallel control system through transfer learning. The updated student model is used to predict the small data, and the prediction results of the student model for the small data are input into the soft actor-critic, and the soft actor-critic outputs control instructions. Every time, the parameters of the controller of the parallel control system with the smallest frequency deviation of the integrated energy system and the balance error of the integrated energy system are selected from parallel control systems and copied to the controller of the integrated energy system through transfer learning. The parameters of the controller are the weights and biases of the quantum generative adversarial network, the weights and biases of the Transformer, and the weights and biases of the soft actor-critic. After the parameters of the controller of the integrated energy system are updated, the system frequency deviation and the balance error of the integrated energy system are predicted through the quantum generative adversarial network and the Transformer respectively, and control instructions are output through the soft actor-critic. A control method for an integrated energy system based on transfer parallel deep reinforcement learning can efficiently adjust the parameters of the controller, adapt to environmental changes, and give accurate control instructions when the capacity of the energy equipment, the coupling mode of the energy equipment, the time constant of the energy equipment, and the ramp coefficient of the energy equipment in the integrated energy system change.

3. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the control method for an integrated energy system based on migration parallel deep reinforcement learning as claimed in claim 1.

4. A processor, characterized in that, The processor is used to run the program, wherein when the program runs, it executes the control method for an integrated energy system based on migration parallel deep reinforcement learning as claimed in claim 1.