Data model dual-drive long-period electric power and electric quantity balance initial solution generation method and device
Through the dual-drive method of data model, combined with the self-attention mechanism, graph neural network and deep reinforcement learning model, the problem of long-term power balance analysis and calculation time in the power system is solved, and efficient power balance analysis and supply and demand balance guarantee are achieved.
Patent Information
- Application Number
- CN202411635790.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively solve the problem of long-term power balance in power systems under limited computing resources, especially when facing the randomness and volatility of new energy output, traditional methods have too long calculation time, making it difficult to ensure supply and demand balance.
Using the dual-driven method of data model, combining the multi-layer autoregressive model of the self-attention mechanism, the redundant network constraint reduction model of the graph neural network, and the deep reinforcement learning model of the deep deterministic strategy gradient algorithm, a mathematical optimization model of the long-term power balance problem of the power system is constructed. Through the interaction between the agent and the power system simulation environment, pre-decision and training are performed to generate initial feasible solutions to solve the optimal solution.
It effectively reduces the computational complexity and time, improves the efficiency of long-term power balance analysis of power systems, and can quickly obtain high-quality optimal solutions under limited computing resources, ensuring the supply and demand balance of power systems.
Smart Images

Figure CN120069353A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of power system planning and operation, and particularly to a method and device for generating an initial solution for long-term power and electricity balance driven by a dual data model. Background Art
[0002] With the continuous increase in the installed capacity and power generation ratio of new energy in the power system, the randomness and volatility of new energy output pose new challenges to the power supply and demand balance of the power system. This uncertainty may lead to insufficient power supply during certain periods, or force the system to abandon some new energy, resulting in waste of resources and economic losses. Therefore, how to effectively cope with the uncertainty of new energy output has become an urgent problem to be solved in the operation and planning of power systems.
[0003] In this context, it is particularly important to conduct accurate analysis of power and electricity balance on a medium- and long-term time scale. Through accurate balance analysis, potential supply-demand imbalance situations can be identified in advance, corresponding control measures can be formulated to ensure the safe and stable operation of the power system. At the same time, this analysis is of great significance for guiding the resource allocation scale of provincial or regional power systems and the assessment of annual power generation resource adequacy, helping to optimize resource allocation and improve system efficiency.
[0004] However, due to the introduction of resources with multiple types and multiple adjustment cycles such as annual-regulated hydropower and seasonal energy storage in the current power system, traditional daily or monthly rolling calculation methods can no longer fully utilize the adjustment capabilities of these resources. Therefore, there is an urgent need for a method for power and electricity balance analysis that can achieve full-cycle calculation. It should be noted that the mathematical models corresponding to such problems are usually complex large-scale mixed integer programming problems, with the phenomenon of dimensionality disaster, and the modeling and solution are extremely difficult, and it is impossible to guarantee obtaining a feasible solution under limited computing resources. If the above problems cannot be effectively solved, it will be difficult to accurately judge whether the power system resources can meet the supply-demand balance in the future time range. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a method and device for generating an initial solution for long-term power and electricity balance driven by a dual data model, and solve the problem of excessive calculation time existing in the existing long-term power and electricity balance analysis based on the artificial intelligence model method, providing a new solution for formulating long-term power and electricity balance solutions for power systems.
[0006] According to the first aspect of the embodiments of the present invention, a method for generating an initial solution for long-term power and electricity balance driven by a dual data model is provided, including: constructing a mathematical optimization model for the long-term power and electricity balance problem of the power system based on the obtained unit characteristics, energy storage device characteristics, and network structure characteristics of the power system; constructing a multi-layer autoregressive model based on the self-attention mechanism based on the obtained scenario load characteristics; constructing a redundant network constraint reduction model based on the graph neural network based on the topological structure and node characteristics of the power grid; constructing a deep reinforcement learning model based on the deep deterministic policy gradient algorithm according to the multi-layer autoregressive model and the redundant network constraint reduction model, where the deep reinforcement learning model includes a power system simulation environment and an agent, inputting the scenario load characteristics and the unit characteristics into the deep reinforcement learning model to pre-decide other unit start-stop variables; training the agent in the deep reinforcement learning model by interacting with the power system simulation environment; based on the pre-decision, fixing the start-stop states of some units through the multi-layer autoregressive model, removing some redundant constraints of the long-term power and electricity balance problem of the power system through the redundant network constraint reduction model, and inputting the output result of the agent as the initial feasible solution of the start-stop decision variables of all time periods of other unfixed units into the mathematical optimization model through the deep reinforcement learning model for further solution to obtain the optimal solution of the long-term power and electricity balance problem of the power system.
[0007] In one implementation, the improved self-attention neural network main body in the multi-layer autoregressive model consists of two parts: an encoding model and a decoding model. The encoding model and the decoding model are constructed by a self-attention module, a forward propagation module, and a cross-attention module. Among them, the output value of the β-th self-attention module is:
[0008] P z,β =L N {P x,β +D r [M HA (P x,β ,P x,β ,P x,β )]}
[0009] Where P z,β is the output value of the β-th self-attention module, P x,β is the global correlation load characteristic input by the β-th self-attention module, L N (·) is the layer normalization function, D r (·) is the neuron random removal function, and M HA (·) is the multi-head attention mechanism function.
[0010] In another implementation, the node feature extraction network composed of the fused edge feature map attention graph convolutional layer in the redundant network constraint reduction model is as follows:
[0011]
[0012] Where X l is the node feature vector of the l-th graph convolutional layer, and X l-1 is the node feature vector of the (l - 1)-th layer. E ..p l-1 is the p-channel vector of the edge feature of the (l - 1)-th layer, g l is the linear transformation function of the l-th graph convolutional layer, a l ..p is the attention parameter matrix of the p-channel of the edge feature of the l-th graph convolutional layer, f l (·) is the calculation function of the attention weight of each channel of the edge feature of the k-th graph convolutional layer. is the weight coefficient of the p-th channel of the node i, j line, and E ijp l-1 is the p-channel vector of the edge feature of the node i, j of the (l - 1)-th layer, E l is the edge feature vector of the l-th graph convolutional layer, and α l is the weight matrix after attention normalization of the l-th graph convolutional layer.
[0013] In another implementation, training the agent in the deep reinforcement learning model by interacting with the power system simulation environment specifically includes: when training the agent, first observing the power system simulation environment and obtaining the optimal action that the agent needs to execute in the current state according to the agent's own policy, and then interacting with the power system simulation environment to obtain samples and storing them in the agent's experience pool and regularly extracting samples for training; among them, the deep reinforcement learning model uses two independent networks to approximate the value function and the policy function, so that the agent selects actions based on the value function and the policy function and interacts with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, the network parameters are optimized by minimizing the loss function L(θ Q ):
[0014] L(θ Q ) = E[y t - Q(s t , a t |θ Q )] 2
[0015] Where y t - Q(s t , a t |θQ ) is the temporal difference error, y t is the target Q value;
[0016] The policy network improves actions based on gradient information, and the sampled policy gradient is shown as follows:
[0017]
[0018] wherein, is the gradient information.
[0019] According to the second aspect of the embodiments of the present invention, a data model dual-driven long-period power and electricity balance initial solution generation device is provided, including: a first construction module for constructing a mathematical optimization model of the long-period power and electricity balance problem of the power system based on the obtained unit characteristics, energy storage device characteristics, and network structure characteristics of the power system; a second construction module for constructing a multi-layer autoregressive model based on the self-attention mechanism based on the obtained scenario load characteristics; a third construction module for constructing a redundant network constraint reduction model based on the graph neural network based on the topological structure and node characteristics of the power grid; a fourth construction module for constructing a deep reinforcement learning model based on the deep deterministic policy gradient algorithm according to the multi-layer autoregressive model and the redundant network constraint reduction model, the deep reinforcement learning model includes a power system simulation environment and an agent, and inputs the scenario load characteristics and the unit characteristics into the deep reinforcement learning model to pre-decide other unit start-stop variables; a training module for training the agent in the deep reinforcement learning model by interacting with the power system simulation environment; a solving module for fixing the start-stop states of some units through the multi-layer autoregressive model based on the pre-decision, removing some redundant constraints of the long-period power and electricity balance problem of the power system through the redundant network constraint reduction model, and inputting the output result of the agent as the initial feasible solution of the start-stop decision variables of all time periods of other unfixed units into the mathematical optimization model through the deep reinforcement learning model for further solution to obtain the optimal solution of the long-period power and electricity balance problem of the power system.
[0020] In one implementation, the improved self-attention neural network main body in the multi-layer autoregressive model consists of two parts: an encoding model and a decoding model, and the encoding model and the decoding model are constructed by a self-attention module, a forward transfer module, and a cross-attention module. Among them, the output value of the β-th self-attention module is:
[0021] P z,β =L N {P x,β +D r [M HA (P x,β ,Px,β ,P x,β )]}
[0022] Among them, P z,β is the output value of the β-th self-attention module, and P x,β is the global relevance load feature input to the β-th self-attention module, L N (·) is the layer normalization function, D r (·) is the neuron random removal function, M HA (·) is the multi-head attention mechanism function.
[0023] In another implementation, the node feature extraction network composed of the fused edge feature map attention graph convolutional layer in the redundant network constraint reduction model is:
[0024]
[0025] Among them, X l is the node feature vector of the l-th graph convolutional layer, and X l-1 is the node feature vector of the (l - 1)-th layer, E ..p l-1 is the p-channel vector of the edge feature of the (l - 1)-th layer, g l is the linear transformation function of the l-th graph convolutional layer, a l ..p is the attention parameter matrix of the p-channel of the edge feature of the l-th graph convolutional layer, f l (·) is the calculation function of the attention weight of each channel of the edge feature of the k-th graph convolutional layer, is the weight coefficient of the p-th channel of the line between nodes i and j, E ijp l-1 is the p-channel vector of the edge feature between nodes i and j of the (l - 1)-th layer, E l is the edge feature vector of the l-th graph convolutional layer, α l is the weight matrix after attention normalization of the l-th graph convolutional layer.
[0026] In another implementation, the training module is specifically used for: when training the agent, first obtain the optimal action that the agent needs to execute in the current state by observing the power system simulation environment and according to the agent's own strategy, and then interact with the power system simulation environment to obtain samples and store them in the agent's experience pool and regularly extract samples for training; among them, the deep reinforcement learning model uses two independent networks to approximate the value function and the policy function, so that the agent selects actions based on the value function and the policy function and interacts with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, the network parameters are optimized by minimizing the loss function L(θ Q )
[0027] L(θ Q ) = E[y t - Q(s t , a t |θ Q )] 2
[0028] where y t - Q(s t , a t |θ Q ) is the temporal difference error, and y t is the target Q value;
[0029] The policy network improves actions based on gradient information, and the sampled policy gradient is shown as follows:
[0030]
[0031] where is the gradient information.
[0032] According to the third aspect of the embodiments of the present invention, an electronic device is provided, including a processor and a memory storing a program. Among them, the program includes instructions that, when executed by the processor, cause the processor to execute the steps performed by the method in the first aspect as described above.
[0033] According to the fourth aspect of the embodiments of the present invention, a computer storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the method in the first aspect as described above.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] (1) The present invention proposes a multi-layer autoregressive model based on the self-attention mechanism, which effectively reduces the scale of integer variables. This method utilizes the powerful feature extraction and sequence modeling capabilities of the Transformer model, reduces the complexity of the optimization problem, improves the computational efficiency, and provides a dimensionality reduction solution for the curse of dimensionality phenomenon in large-scale mixed integer programming.
[0036] (2) The present invention designs a redundant network constraint reduction model based on the graph neural network, which can intelligently identify and delete redundant network constraints in the optimization model. By introducing the graph neural network, the structural characteristics of the power system network topology are fully utilized, the number of constraint conditions is reduced, the computational complexity of the model is further reduced, and the solution speed is accelerated.
[0037] (3) The present invention constructs a deep reinforcement learning model based on the Deep Deterministic Policy Gradient algorithm to pre-determine the remaining decision variables in the long-term power and electricity balance problem. Through interaction training with the environment, the agent can learn the strategy to optimize the decision variables and input the obtained results as the initial feasible solution into the optimization model. This not only improves the quality of the initial solution, enhances the convergence of the solver, but also can quickly obtain a high-quality optimal solution under limited computing resources, and relevant research work can be carried out by engineering practitioners based on this. Brief Description of the Drawings
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0039] Figure 1 It is a flowchart of the steps of the data model double-driven long-term power and electricity balance initial solution generation method of the present invention;
[0040] Figure 2 It is a structural block diagram of the data model double-driven long-term power and electricity balance initial solution generation device of the present invention;
[0041] Figure 3 It is a schematic structural diagram of an electronic device of the present invention. Detailed Embodiments
[0042] In order to have a clearer understanding of the technical features, objectives, and effects of the embodiments of the present invention, the following will describe the specific embodiments of the embodiments of the present invention with reference to the accompanying drawings.
[0043] In this article, "exemplarily" means "serving as an instance, example, or illustration", and any illustration or embodiment described as "exemplarily" in this article should not be interpreted as a more preferred or more advantageous technical solution.
[0044] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention should fall within the scope of protection of the embodiments of the present invention.
[0045] The following further illustrates the specific implementation of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention.
[0046] SeeFigure 1 , the present invention provides a method for generating an initial solution of long-term power and electricity balance driven by a dual data model, which mainly includes the following steps:
[0047] Step S1: Based on the obtained unit characteristics, energy storage device characteristics, and network structure characteristics of the power system, construct a mathematical optimization model for the long-term power and electricity balance problem of the power system;
[0048] Step S2: Based on the obtained scenario load characteristics, construct a multi-layer autoregressive model based on the self-attention mechanism;
[0049] Step S3: Based on the topological structure and node characteristics of the power grid, construct a redundant network constraint reduction model based on the graph neural network;
[0050] Step S4: According to the multi-layer autoregressive model and the redundant network constraint reduction model, construct a deep reinforcement learning model based on the deep deterministic policy gradient algorithm. The deep reinforcement learning model includes a power system simulation environment and an agent. Input the scenario load characteristics and the unit characteristics into the deep reinforcement learning model to pre-determine the start-stop variables of other units;
[0051] Step S5: Train the agent in the deep reinforcement learning model by interacting with the power system simulation environment;
[0052] Step S6: Based on the pre-determination, fix the start-stop states of some units through the multi-layer autoregressive model, eliminate some redundant constraints of the long-term power and electricity balance problem of the power system through the redundant network constraint reduction model, and use the output result of the agent as the initial feasible solution of the start-stop decision variables of all periods of other unfixed units and input it into the mathematical optimization model through the deep reinforcement learning model for further solution to obtain the optimal solution of the long-term power and electricity balance problem of the power system.
[0053] It should be understood that when further solving, it can be solved by existing mathematical solvers Gurobi or Cplex.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] (1) The present invention proposes a multi-layer autoregressive model based on the self-attention mechanism, which effectively reduces the scale of integer variables. This method utilizes the powerful feature extraction and sequence modeling capabilities of the Transformer model, reduces the complexity of the optimization problem, improves the calculation efficiency, and provides a dimension reduction solution for the curse of dimensionality phenomenon in large-scale mixed integer programming.
[0056] (2) The present invention designs a redundant network constraint reduction model based on a graph neural network, which can intelligently identify and delete redundant network constraints in the optimization model. By introducing the graph neural network, the structural characteristics of the power system network topology are fully utilized, the number of constraint conditions is reduced, the computational complexity of the model is further reduced, and the solution speed is accelerated.
[0057] (3) The present invention constructs a deep reinforcement learning model based on the deep deterministic policy gradient algorithm to pre-decide the remaining decision variables in the long-term power and electricity balance problem. Through interactive training with the environment, the agent can learn the strategy of optimizing the decision variables and input the obtained result as the initial feasible solution into the optimization model. This not only improves the quality of the initial solution and enhances the convergence of the solver, but also can quickly obtain a high-quality optimal solution under limited computing resources, and relevant research work can be carried out by engineering practitioners accordingly.
[0058] Optionally, the improved self-attention neural network body in the multi-layer autoregressive model consists of two parts: an encoding model and a decoding model. At the same time, the position embedding module is removed. The encoding model and the decoding model are constructed by a self-attention module, a forward propagation module, and a cross-attention module. Among them, the output value of the β-th self-attention module is:
[0059] P z,β =L N {P x,β +D r [M HA (P x,β ,P x,β ,P x,β )]}
[0060] Among them, P z,β is the output value of the β-th self-attention module, P x,β is the global relevance load feature input by the β-th self-attention module, L N (·) is the layer normalization function, D r (·) is the neuron random removal function, and M HA (·) is the multi-head attention mechanism function.
[0061] Optionally, the node feature extraction network composed of the fused edge feature map attention graph convolutional layer in the redundant network constraint reduction model is:
[0062]
[0063] Among them, X l is the node feature vector of the l-th graph convolutional layer, X l-1 is the node feature vector of the l-1 layer, and E ..p l-1is the p-channel vector of the edge feature of layer l-1, g l is the linear transformation function of the l-th layer graph convolutional layer, a l ..p is the attention parameter matrix of the p-channel of the edge feature of the l-th layer graph convolutional layer, f l (·) is the calculation function of the attention weight of each channel of the edge feature of the k-th layer graph convolutional layer, is the weight coefficient of the p-th channel of the node i, j line, E ijp l-1 is the p-channel vector of the edge feature of the nodes i, j of layer l-1, E l is the edge feature vector of the l-th layer graph convolutional layer, α l is the attention-normalized weight matrix of the l-th layer graph convolutional layer.
[0064] Optionally, training the agent in the deep reinforcement learning model by interacting with the power system simulation environment specifically includes: when training the agent, first obtaining the optimal action that the agent needs to execute in the current state by observing the power system simulation environment and according to the agent's own policy, and then interacting with the power system simulation environment to obtain samples and storing them in the agent's experience pool and regularly extracting samples for training; wherein, the deep reinforcement learning model uses two independent networks to approximate the value function and the policy function, so that the agent selects actions based on the value function and the policy function and interacts with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, use the minimization of the loss function L(θ Q ) to optimize the network parameters:
[0065] L(θ Q ) = E[y t - Q(s t , a t |θ Q )] 2
[0066] wherein, y t - Q(s t , a t |θ Q ) is the temporal difference error, y t is the target Q value;
[0067] The policy network improves the action based on the gradient information, and the sampling policy gradient is shown as follows:
[0068]
[0069] wherein, is the gradient information.
[0070] Specifically, the solution of the present invention is further described according to the following examples:
[0071] The method for generating an initial solution for long-term power and electricity balance driven by a dual data model is implemented through the following steps:
[0072] (1) Modeling and model preprocessing of long-term power and electricity balance problems
[0073] The main purpose of long-term power and electricity balance analysis is to provide an assessment conclusion of the adequacy of power generation resources for provincial or regional power systems on an annual or multi-year time scale, and to guide the formulation of resource allocation planning schemes for power sources, grid frameworks, etc. of future power systems. With the widespread access of uncertain power sources such as new energy, the time scale of concern for the operating characteristics of power systems has gradually deepened to scales such as hourly and minute-level. Therefore, the analysis scale of long-term power and electricity balance problems needs to be extended to more refined time granularity. The mathematical essence of long-term power and electricity balance problems is a complex large-scale mixed integer programming problem. In the context where both the breadth and depth of problem analysis need to be extended, this problem gradually exposes solving difficulties such as the curse of dimensionality in computational analysis performance.
[0074] The objective function of the mathematical optimization model for long-term power and electricity balance problems is shown in Equation (1), where t / T is the index / set of the analysis period; N T is the total number of time periods; ρ / Ω A is the index / set of the region; σ load , σ gex and σ ne are the weight coefficients of load shedding risk, external power purchase penalty, and renewable energy curtailment risk, respectively; P Llossρ,t is the load shedding amount of region ρ at time period t; P Gexρ,t is the external power purchase amount of region ρ at time period t; i W,ρ / Ω W,ρ and i S,ρ / Ω S,ρ are the indexes / sets of wind turbine generators and photovoltaic generator sets, respectively; and are the maximum outputs of wind turbine generator i W,ρ and photovoltaic generator set i S,ρ at time period t, respectively; P and are the outputs of wind turbine generator i W,ρ and photovoltaic generator set i S,ρ at time period t, respectively; i ES,ρ / Ω ES,ρ is the index / set of the energy storage device in region ρ; and are the energy storage device i ES,ρOperating cost of unit discharge and charge power; and are the charge and discharge powers of energy storage device i ES,ρ at time t; i C,ρ / Ω C,ρ is the index / set of coal-fired units in area ρ; i G,ρ / Ω G,ρ is the index / set of gas-fired units in area ρ; is the power generation of coal-fired unit i C,ρ at time t; and are the Boolean variables representing the operation, startup, and shutdown states of coal-fired unit i C,ρ at time t; is the piecewise linearized function of the operating cost of coal-fired unit i C,ρ δ C,ρ is the index of the piecewise linearized function segment of the operating cost of coal-fired units; and are the costs of single startup and shutdown of coal-fired unit i C,ρ respectively; is the power generation of gas-fired unit i G,ρ at time t; and are the Boolean variables representing the operation, startup, and shutdown states of gas-fired unit i G,ρ at time t; is the piecewise linearized function of the operating cost of gas-fired unit i G,ρ δ G,ρ is the index of the piecewise linearized function segment of the operating cost of gas-fired units; and are the costs of single startup and shutdown of gas-fired unit i G,ρ i respectively; H,ρ / Ω H,ρ is the index / set of hydroelectric units; is the piecewise linearized function of the flow-output characteristic of hydroelectric unit i H,ρ ; is the power generation flow of hydroelectric unit i H,ρ at time t; δ H / Ω δH is the index / set of the piecewise linearized function segment of the flow-output characteristic of hydroelectric units; is the out-flow of hydroelectric unit i H,ρ at time t.
[0075]
[0076] The constraint conditions corresponding to the gas turbine units are shown in equations (2)-(12), where and are the times required for the startup and shutdown of gas turbine unit i G,ρ since the initial time; and are the minimum startup and shutdown times of gas turbine unit i G,ρ respectively; and are the times experienced by gas turbine unit i G,ρ since the initial time to the most recent startup and shutdown; and are the minimum and maximum technical outputs of gas turbine unit i G respectively; and are the maximum downward and upward ramp rates of gas turbine unit i G,ρ respectively; and are the maximum downward ramp rate in the shutdown state and the maximum upward ramp rate in the startup state of gas turbine unit i G,ρ respectively.
[0077]
[0078] The constraints corresponding to the coal-fired power units are shown in equations (13)-(22), where and are the times required for the startup and shutdown of coal-fired power unit i C,ρ since the initial time; and are the minimum startup and shutdown times of coal-fired power unit i C,ρ respectively; and are the times experienced by coal-fired power unit i C,ρ since the initial time to the most recent startup and shutdown; and are the maximum downward and upward ramp rates of coal-fired power unit i C,ρ respectively; and are the maximum downward ramp rate in the shutdown state and the maximum upward ramp rate in the startup state of coal-fired power unit i C,ρ respectively.
[0079]
[0080] The constraints corresponding to the energy storage devices are shown in equations (23)-(33), where and are the new energy storage devices i ES,ρBoolean variable characterizing the charging and discharging states during period t; and are respectively the minimum charging and discharging powers of the new energy storage device i ES,ρ ; and are respectively the maximum charging and discharging powers of the energy storage device i ES,ρ ; and are respectively the charging and discharging efficiencies of the new energy storage device i ES,ρ ; is the stored energy of the new energy storage device i ES,ρ during period t; and are respectively the lower and upper limits of the stored energy of the new energy storage device i ES,ρ ; i ESY,ρ / Ω ESY,ρ 、i ESM,ρ / Ω ESM,ρ 、i ESW,ρ / Ω ESW,ρ and i ESD,ρ / Ω ESD,ρ are respectively the device indexes / sets with annual, monthly, weekly, and daily energy storage cycles in region ρ; and are respectively the initial stored energies of the devices with annual, monthly, weekly, and daily energy storage cycles in region ρ; and are respectively the energy storage values of the devices with annual, monthly, weekly, and daily energy storage cycles in region ρ during period N T 、during period t MN 、during period t WN 、during period t DN ; t MN / Ω MN 、t WN / Ω WN and t DN / Ω DN are respectively the end - time indexes / sets of each month, week, and day.
[0081]
[0082] The constraints of the hydropower unit are shown in Eqs. (34) - (43), where and are respectively the first - order term and constant - term coefficients of the δ H,ρ th segment of the piece - wise linearization function of the flow - output characteristic of hydropower unit i H,ρ ; is the water storage volume of the reservoir corresponding to hydropower unit i H,ρ during period t; is the natural inflow of hydropower unit i H,ρ at time period t; Δt is the duration of a unit time period; and are respectively the lower and upper limits of the outflow of hydropower unit i H,ρ from the corresponding power station; and are respectively the lower and upper limits of the water storage in the reservoir corresponding to hydropower unit i H,ρ ; i HY,ρ / Ω HY,ρ 、i HM,ρ / Ω HM,ρ 、i HW,ρ / Ω HW,ρ and i HD,ρ / Ω HD,ρ are respectively the indexes / sets of hydropower units with annual, monthly, weekly, and daily cycles in region ρ; and are respectively the initial water storages of the power stations corresponding to hydropower units with annual, monthly, weekly, and daily cycles in region ρ; and are respectively the water storages of the power stations corresponding to hydropower units with annual, monthly, weekly, and daily cycles in region ρ within the N T time period, within the t MN time period, within the t WN time period, within the t DN time period.
[0083]
[0084]
[0085] The spinning reserve constraint is shown in Eqs. (44)-(53), where and are respectively the positive spinning reserve capacity and negative spinning reserve capacity of coal-fired unit i C,ρ at time period t; is the spinning reserve response time of coal-fired unit i C,ρ ; and are respectively the positive spinning reserve capacity and negative spinning reserve capacity of gas-fired unit i G,ρ at time period t; is the spinning reserve response time of gas-fired unit i G,ρ ; and are respectively the positive spinning reserve capacity and negative spinning reserve capacity of hydropower unit i H,ρ at time period t; P and are respectively the positive spinning reserve capacity and negative spinning reserve capacity of new energy storage device i ES,ρPositive and negative spinning reserve capacities during period t; ε L+ and ε L- are the positive and negative spinning reserve demand coefficients of the load respectively; ε R+ and ε R- are the positive and negative spinning reserve demand coefficients of the new energy respectively; P Lρ,t is the load demand of area ρ at time period t.
[0086]
[0087]
[0088] The constraints of the new energy units are shown in equations (54)-(55):
[0089]
[0090] The power balance constraints are shown in equations (56)-(57), where P Gexρ,t and P Lexρ,t are the electricity purchased from outside and the electricity output to the outside by area ρ at time period t respectively; P Gexmax and P Lexmax are the maximum values of the electricity purchased from outside and the electricity output to the outside by the system at a single moment respectively; Ω Line is the set of inter-area tie lines; is the transmission electricity of the tie line connecting area ρ and area at time period t.
[0091]
[0092] (2) Construction of a multi-layer autoregressive model based on the self-attention mechanism
[0093] Construct a deep neural network model based on the improved Transformer architecture, that is, construct a multi-layer autoregressive model based on the self-attention mechanism, which is used to reduce the multi-period integer variables of the unit commitment problem of power and electricity balance. Transformer is an autoregressive model that makes predictions one part at a time and then uses its own output so far to decide what to do next. Arrange the load demands of all areas in the system in sequence according to time series to form matrix P M , and use P M as the feature input to the neural network, as shown in equation (61), where N is the number of areas and T is the number of time periods within the analysis period:
[0094]
[0095] Construct the target vector Q of the output Transformer neural network, i.e., the self-attention neural network, from all the boolean variables of the unit start-stop states in the system M , which is used as the output value of the neural network, as shown in Equation (62):
[0096]
[0097] The improved Transformer neural network, i.e., the improved self-attention neural network, consists of two parts: an encoding model (Encoder) and a decoding model (Decoder). The encoding model and the decoding model are constructed by a self-attention module, i.e., the Self-Attention module, a feed-forward module, and a cross-attention module. Combining the input and output structural characteristics of the Transformer neural network, the method of constructing the load characteristics and the target output in different time periods enables the Self-Attention mechanism under the Transformer architecture, i.e., the self-attention mechanism, to be applied to the data structure of unit commitment. The multi-head attention mechanism is used to extract the correlation of load characteristics between different time periods, and the Self-Attention mechanism is calculated for each group of input features respectively to achieve the extraction of load characteristics in different subspaces, as shown in Equation (63):
[0098] P z,β = L N {P x,β + D r [M HA (P x,β , P x,β , P x,β )]} (63)
[0099] Among them, P z,β is the output value of the β-th self-attention module, P x,β is the global correlation load characteristic input to the β-th self-attention module, L N (·) is the layer normalization function, D r (·) is the neuron random removal function, i.e., the Dropout function, and M HA (·) is the multi-head attention mechanism function.
[0100] The fitting ability of the model is enhanced by setting the feed-forward module. First, the output of the previous module is connected to the feed-forward neural network layer for calculation. The feed-forward neural network consists of two fully connected networks. Secondly, it passes through Dropout and residual connection, and finally outputs through layer normalization, as shown in Equation (64), where F FN (·) is the feed-forward neural network function; x is the output of the previous module and also the input of this module; W 1 and W 2is the weight parameter of the feedforward neural network; b 1 and b 2 are the biases of the first and second layers of the feedforward neural network respectively; S(·) is the ReLU activation function, i.e., the rectified activation function.
[0101] F FN (x) = S(xW 1 +b 1 )W 2 +b 2 (64)
[0102] The calculation process of the forward propagation module is shown in Equation (65), where y is the output of the forward propagation module.
[0103] y = L N {x + D r [F FN (x)]} (65)
[0104] The Cross-Attention module, i.e., the cross-attention module, is in the model decoding part and is used to combine the global encoding features and decode the target sequence. The calculation process of this module is shown in Equation (66), where P z,α is the output of the α-th Cross-Attention module; P y,α is the output of the upper layer module and also the input of this module; P e is the output value of the encoding part.
[0105] P z,α = L N {P y,α + D r [M HA (P e W b , P e W b , P y,α )]} (66)
[0106] Establish a time-series coupling feature matrix for the features of the input load sequence, and connect this matrix to the improved Transformer neural network, i.e., the improved self-attention neural network, to pre-identify the unit start-stop variables. Then, at the final output layer of the network, map the pre-identification result to the (0,1) interval in the form of an output probability value, and determine the confidence state of the unit prediction value through a pre-set confidence threshold. The unit start-stop variables that meet the confidence level are regarded as redundant integer variables, and the values of this type of integer variable are fixed to realize the construction of the multi-period integer variable reduction model.
[0107] (3) Construction of the redundant network constraint reduction model based on the graph neural network
[0108] Input the topological graph and node feature data of the power grid into the node feature extraction network composed of the edge feature graph attention graph convolutional layer, i.e., the EGAT graph convolutional layer, to extract the node feature vector, as shown in Equation (67):
[0109]
[0110] where, X l is the node feature vector of the l-th layer graph convolutional layer, X l-1 is the node feature vector of the (l - 1)-th layer, E ..p l-1 is the p-channel vector of the edge feature of the (l - 1)-th layer, g l is the linear transformation function of the l-th layer graph convolutional layer, a l ..p is the attention parameter matrix of the p-channel of the edge feature of the l-th layer graph convolutional layer, f l (·) is the calculation function of the attention weight of each channel of the edge feature of the k-th layer graph convolutional layer, is the weight coefficient of the p-th channel of the line between nodes i and j, E ijp l-1 is the p-channel vector of the edge feature between nodes i and j of the (l - 1)-th layer, E l is the edge feature vector of the l-th layer graph convolutional layer, α l is the weight matrix after attention normalization of the l-th layer graph convolutional layer.
[0111] Concatenate the feature vectors of the two ends of each branch with the feature of the branch itself, and use it as the input to the branch feature extraction network composed of fully connected layers to generate the branch feature vector. Pass the branch feature vector through the fully connected layer and input it into the S-shaped curve function, i.e., the Sigmoid function, to calculate the probability that the branch safety constraint becomes an active constraint after the current iterative solution, and sample according to this probability. If the sampling result is 1, add the corresponding network constraint; if the sampling result is 0, consider it as a redundant network constraint and cut it.
[0112] (4) Construction of the deep reinforcement learning model based on the Deep Deterministic Policy Gradient algorithm (DDPG) and generation of the initial solution for the long-term power and energy balance
[0113] Construct a deep reinforcement learning model based on the Deep Deterministic Policy Gradient (DDPG) algorithm to quickly infer the initial solution of the long-term power and energy balance problem after reducing redundant integer variables and redundant network constraint variables. When training the agent, first observe the power system simulation environment and obtain the optimal action that the agent needs to execute in the current state according to the agent's own policy, and then interact with the power system simulation environment to obtain samples, store them in the agent's experience pool, and regularly extract samples for training. The deep reinforcement learning model based on the DDPG algorithm uses two independent networks to approximate the value function and the policy function, so that the agent can select actions based on the value function and the policy function and interact with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, use the minimization of the loss function L(θ Q ) to optimize the network parameters, as shown in Equation (68), where y t -Q(s t ,a t |θ Q ) is the temporal difference error, and y t is the target Q value.
[0114] L(θ Q ) = E[y t -Q(s t ,a t |θ Q )] 2 (68)
[0115] The policy network improves the action based on the gradient information. The sampled policy gradient is as shown in Equation (69), where is the gradient information.
[0116]
[0117] Interact with the power system simulation environment to complete the agent training process.
[0118] Based on the pre-decision, fix the start-stop states of some units through a multi-layer autoregressive model, eliminate some redundant constraints of the long-term power and energy balance problem of the power system through the redundant network constraint reduction model, and use the output result of the deep reinforcement learning model as the initial feasible solution of the start-stop decision variables of all time periods of other unfixed units to input into the mathematical optimization model for further solution to obtain the optimal solution of the long-term power and energy balance problem of the power system.
[0119] It should be understood that when further solving, it can be solved by existing mathematical solvers such as Gurobi or Cplex.
[0120] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0121] (1) The present invention proposes a multi-layer autoregressive model based on the self-attention mechanism, which effectively reduces the scale of integer variables. This method utilizes the powerful feature extraction and sequence modeling capabilities of the Transformer model, reduces the complexity of the optimization problem, improves the computational efficiency, and provides a dimension reduction solution for the curse of dimensionality in large-scale mixed-integer programming.
[0122] (2) The present invention designs a redundant network constraint reduction model based on a graph neural network, which can intelligently identify and delete redundant network constraints in the optimization model. By introducing the graph neural network, the structural characteristics of the power system network topology are fully utilized, the number of constraint conditions is reduced, the computational complexity of the model is further reduced, and the solution speed is accelerated.
[0123] (3) The present invention constructs a deep reinforcement learning model based on the deep deterministic policy gradient algorithm to pre-decide the remaining decision variables in the long-term power and electricity balance problem. Through interactive training with the environment, the agent can learn the strategy of optimizing decision variables and input the obtained results as the initial feasible solution into the optimization model. This not only improves the quality of the initial solution, enhances the convergence of the solver, but also can quickly obtain high-quality optimal solutions under limited computational resources, and relevant research work can be carried out by engineering practitioners based on this.
[0124] As another example, see Figure 2 , the present invention provides a data model double-driven long-term power and electricity balance initial solution generation device 200, including:
[0125] The first construction module 210 is used to construct a mathematical optimization model for the long-term power and electricity balance problem of the power system based on the obtained unit characteristics, energy storage device characteristics, and network structure characteristics of the power system;
[0126] The second construction module 220 is used to construct a multi-layer autoregressive model based on the self-attention mechanism based on the obtained scenario load characteristics;
[0127] The third construction module 230 is used to construct a redundant network constraint reduction model based on a graph neural network based on the topological structure and node characteristics of the power grid;
[0128] The fourth construction module 240 is used to construct a deep reinforcement learning model based on the deep deterministic policy gradient algorithm according to the multi-layer autoregressive model and the redundant network constraint reduction model. The deep reinforcement learning model includes a power system simulation environment and an agent, and inputs the scenario load characteristics and the unit characteristics into the deep reinforcement learning model to pre-decide other unit start-stop variables;
[0129] A training module 250, configured to train an agent in the deep reinforcement learning model by interacting with the power system simulation environment;
[0130] A solving module 260, configured to, based on the pre-decision, fix the start-stop states of some units through the multi-layer autoregressive model, remove some redundant constraints of the long-term power balance problem of the power system through the redundant network constraint reduction model, and use the output result of the agent of the deep reinforcement learning model as the initial feasible solution of all-time period start-stop decision variables of other unfixed units to input into the mathematical optimization model for further solution, so as to obtain the optimal solution of the long-term power balance problem of the power system.
[0131] Optionally, the improved self-attention neural network main body in the multi-layer autoregressive model is composed of an encoding model and a decoding model, and the encoding model and the decoding model are constructed by a self-attention module, a forward transmission module, and a cross-attention module. Among them, the output value of the β-th self-attention module is:
[0132] P z,β =L N {P x,β +D r [M HA (P x,β ,P x,β ,P x,β )]}
[0133] Among them, P z,β is the output value of the β-th self-attention module, P x,β is the global correlation load feature input by the β-th self-attention module, L N (·) is a layer normalization function, D r (·) is a neuron random removal function, M HA (·) is a multi-head attention mechanism function.
[0134] Optionally, the node feature extraction network composed of a fused edge feature map attention graph convolutional layer in the redundant network constraint reduction model is:
[0135]
[0136] Among them, X l is the node feature vector of the l-th graph convolutional layer, X l-1 is the node feature vector of the l-1 layer, E ..p l-1 is the l-1 layer edge feature p-channel vector, g l is the linear transformation function of the l-th graph convolutional layer, a l ..pis the attention parameter matrix for the p channels of the edge features of the l-th layer graph convolutional layer, f l (·) is the calculation function for the attention weights of each channel of the edge features of the k-th layer graph convolutional layer, is the weight coefficient for the p-th channel of the node i, j circuit, E ijp l-1 is the p-th channel vector of the edge features of nodes i and j in the l-1 layer, E l is the edge feature vector of the l-th layer graph convolutional layer, α l is the weight matrix after attention normalization of the l-th layer graph convolutional layer.
[0137] Optionally, the training module is specifically configured to: when training the agent, first obtain the optimal action that the agent needs to execute in the current state by observing the power system simulation environment and according to the strategy of the agent itself, and then interact with the power system simulation environment to obtain samples and store them in the experience pool of the agent and regularly extract samples for training; wherein, the deep reinforcement learning model uses two independent networks to approximate the value function and the policy function, so that the agent selects actions based on the value function and the policy function and interacts with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, the network parameters are optimized by minimizing the loss function L(θ Q ):
[0138] L(θ Q ) = E[y t - Q(s t , a t |θ Q )] 2
[0139] wherein, y t - Q(s t , a t |θ Q ) is the temporal difference error, and y t is the target Q value;
[0140] The policy network improves the action based on the gradient information, and the sampling policy gradient is shown as follows:
[0141]
[0142] wherein, ▽ a Q(s, a|θ Q ) is the gradient information.
[0143] It should be understood that the data model double-driven long-period power and electricity balance initial solution generation device in this embodiment is used to implement the corresponding methods in the foregoing multiple method embodiments and has the beneficial effects of the corresponding method embodiments.
[0144] As another example, refer to Figure 3 , an electronic device 300 is provided. Now, the structural block diagram of the electronic device 300 that can be used as the server or client of the present invention will be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0145] The electronic device 300 may include: a processor 302, a communications interface 304, a memory 306, and a communication bus 308.
[0146] The processor 302, the communications interface 304, and the memory 306 communicate with each other through the communication bus 308. The communications interface 304 is used to communicate with other electronic devices or servers.
[0147] The processor 302 is used to execute the program 310, and specifically can execute the relevant steps in the above method embodiments.
[0148] Specifically, the program 810 may include program code, and the program code includes computer operation instructions.
[0149] The processor 802 may be a processor CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0150] The memory 306 is used to store the program 310. The memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0151] When executed by the processor 302, the program 310 is used to enable the electronic device to execute the data model double-driven long-term power and electricity balance initial solution generation method, including constructing a mathematical optimization model for the long-term power and electricity balance problem of the power system based on the obtained unit characteristics, energy storage device characteristics, and network structure characteristics of the power system; constructing a multi-layer autoregressive model based on the self-attention mechanism based on the obtained scenario load characteristics; constructing a redundant network constraint reduction model based on the graph neural network based on the topological structure and node characteristics of the power grid; constructing a deep reinforcement learning model based on the deep deterministic policy gradient algorithm according to the multi-layer autoregressive model and the redundant network constraint reduction model, where the deep reinforcement learning model includes a power system simulation environment and an agent, inputting the scenario load characteristics and the unit characteristics into the deep reinforcement learning model to pre-decide the start-stop variables of other units; training the agent in the deep reinforcement learning model by interacting with the power system simulation environment; based on the pre-decision, fixing the start-stop states of some units through the multi-layer autoregressive model, removing some redundant constraints of the long-term power and electricity balance problem of the power system through the redundant network constraint reduction model, and inputting the output result of the agent as the initial feasible solution of the start-stop decision variables of all time periods of other unfixed units into the mathematical optimization model through the deep reinforcement learning model for further solution to obtain the optimal solution of the long-term power and electricity balance problem of the power system.
[0152] In addition, for the specific implementation of each step in the program 310, reference can be made to the corresponding steps and descriptions in the corresponding units in the foregoing method embodiments, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.
[0153] The exemplary embodiments of the present invention further provide a computer storage medium storing a computer program, wherein when the computer program is executed by a processor, the methods of the embodiments of the present invention are implemented. Reference can be made to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.
[0154] The method according to an embodiment of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code that is originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and will be stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0155] So far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing may be advantageous.
[0156] It should be understood that although this specification is described according to various embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0157] Finally, it should be noted that the above embodiments are only used to illustrate the embodiments of the present invention, rather than to limit the embodiments of the present invention. Those of ordinary skill in the relevant technical field can still make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. A data model dual-driven long-period power and quantity balance initial solution generation method, characterized in that: include: Based on the acquired characteristics of the power system units, energy storage equipment and network structure, a mathematical optimization model for the long-term power and electricity balance problem of the power system is constructed; Based on the acquired scene load characteristics, a multi-layer autoregressive model based on the self-attention mechanism is constructed; Based on the topological structure and node characteristics of the power grid, a redundant network constraint reduction model based on graph neural network is constructed; According to the multi-layer autoregressive model and the redundant network constraint reduction model, a deep reinforcement learning model based on a deep deterministic policy gradient algorithm is constructed, wherein the deep reinforcement learning model includes a power system simulation environment and an intelligent agent, and the scenario load characteristics and the unit characteristics are input into the deep reinforcement learning model to make a pre-decision on the start-stop variables of other units; Training an agent in the deep reinforcement learning model by interacting with the power system simulation environment; Based on the pre-decision, the start and stop states of some units are fixed through the multi-layer autoregressive model, and some redundant constraints of the long-term power and quantity balance problem of the power system are eliminated through the redundant network constraint reduction model. The output results of the intelligent agent are used as the initial feasible solution of the start and stop decision variables of all time periods of other non-fixed units through the deep reinforcement learning model, and are input into the mathematical optimization model for further solution to obtain the optimal solution to the long-term power and quantity balance problem of the power system.
2. The method according to claim 1, characterized in that The improved self-attention neural network in the multi-layer autoregressive model consists of two models: an encoding model and a decoding model. The encoding model and the decoding model are constructed by a self-attention module, a forward transfer module and a cross attention module. The output value of the β-th self-attention module is: P z,β =L N {P x,β +D r [M HA (P x,β ,P x,β ,P x,β )]} Among them, P z,β is the output value of the βth self-attention module, P x,β is the global correlation load feature input to the βth self-attention module, L N (·) is the layer normalization function, D r (·) is the neuron random removal function, M HA (·) is the multi-head attention mechanism function.
3. The method according to claim 2, characterized in that The node feature extraction network composed of the fused edge feature map attention map convolution layer in the redundant network constraint reduction model is: Among them, X l is the node feature vector of the lth graph convolutional layer, X l-1 is the node feature vector of layer l-1, E ..p l-1 is the p-channel vector of the edge feature at layer l-1, g l is the linear transformation function of the lth graph convolution layer, a l ..p is the attention parameter matrix of the edge feature p channel of the l-th layer graph convolution layer, f l (·) is the attention weight calculation function for each channel of the edge feature of the k-th graph convolution layer, is the weight coefficient of the line channel p of node i,j, E ijp l-1 is the edge feature p channel vector of node i,j in layer l-1, E l is the edge feature vector of the lth graph convolutional layer, α l is the normalized attention weight matrix of the l-th graph convolutional layer.
4. The method according to claim 1, characterized in that: The training of the agent in the deep reinforcement learning model by interacting with the power system simulation environment specifically includes: When training the intelligent agent, firstly, the optimal action that the intelligent agent needs to perform in the current state is obtained by observing the power system simulation environment and according to the strategy of the intelligent agent itself, and then the intelligent agent interacts with the power system simulation environment to obtain samples, stores them in the experience pool of the intelligent agent, and regularly extracts samples for training; The deep reinforcement learning model uses two independent networks to approximate the value function and the policy function, so that the agent selects actions based on the value function and the policy function and interacts with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, the loss function L(θ Q ) to optimize network parameters: L(θ Q )=E[y t -Q(s t ,a t |θ Q )] 2 Among them, y t -Q(s t ,a t |θ Q ) is the timing difference error, y t is the target Q value; The policy network improves actions based on gradient information, and the sampling policy gradient is shown as follows: in, is the gradient information.
5. A data model dual-driven long-period power and quantity balance initial solution generation device, characterized in that: include: The first construction module is used to construct a mathematical optimization model for the long-term power and quantity balance problem of the power system based on the acquired unit characteristics, energy storage equipment characteristics and network structure characteristics of the power system; The second building module is used to build a multi-layer autoregressive model based on the self-attention mechanism based on the acquired scene load characteristics; The third building module is used to build a redundant network constraint reduction model based on graph neural network based on the topological structure and node characteristics of the power grid; A fourth construction module is used to construct a deep reinforcement learning model based on a deep deterministic policy gradient algorithm according to the multi-layer autoregressive model and the redundant network constraint reduction model, wherein the deep reinforcement learning model includes a power system simulation environment and an intelligent agent, and the scenario load characteristics and the unit characteristics are input into the deep reinforcement learning model to make a pre-decision on the start-stop variables of other units; A training module, used for training an agent in the deep reinforcement learning model by interacting with the power system simulation environment; A solution module is used to fix the start and stop states of some units through the multi-layer autoregressive model based on the pre-decision, eliminate some redundant constraints of the long-term power and quantity balance problem of the power system through the redundant network constraint reduction model, and use the output results of the intelligent agent as the initial feasible solution of the start and stop decision variables of all time periods of other non-fixed units through the deep reinforcement learning model to input the mathematical optimization model for further solution, so as to obtain the optimal solution to the long-term power and quantity balance problem of the power system.
6. The device according to claim 5, characterized in that The improved self-attention neural network in the multi-layer autoregressive model consists of two models: an encoding model and a decoding model. The encoding model and the decoding model are constructed by a self-attention module, a forward transfer module and a cross attention module. The output value of the β-th self-attention module is: P z,β =L N {P x,β +D r [M HA (P x,β ,P x,β ,P x,β )]} Among them, P z,β is the output value of the βth self-attention module, P x,β is the global correlation load feature input to the βth self-attention module, L N (·) is the layer normalization function, D r (·) is the neuron random removal function, M HA (·) is the multi-head attention mechanism function.
7. The device according to claim 6, characterized in that The node feature extraction network composed of the fused edge feature map attention map convolution layer in the redundant network constraint reduction model is: Among them, X l is the node feature vector of the lth graph convolutional layer, X l-1 is the node feature vector of layer l-1, E ..p l-1 is the p-channel vector of the edge feature at layer l-1, g l is the linear transformation function of the lth graph convolution layer, a l ..p is the attention parameter matrix of the edge feature p channel of the l-th layer graph convolution layer, f l (·) is the attention weight calculation function for each channel of the edge feature of the k-th graph convolution layer, is the weight coefficient of the line channel p of node i,j, E ijp l-1 is the edge feature p channel vector of node i,j in layer l-1, E l is the edge feature vector of the lth graph convolutional layer, α l is the normalized attention weight matrix of the l-th graph convolutional layer.
8. The device according to claim 5, characterized in that The training module is specifically used for: When training the intelligent agent, firstly, the optimal action that the intelligent agent needs to perform in the current state is obtained by observing the power system simulation environment and according to the strategy of the intelligent agent itself, and then the intelligent agent interacts with the power system simulation environment to obtain samples, stores them in the experience pool of the intelligent agent, and regularly extracts samples for training; The deep reinforcement learning model uses two independent networks to approximate the value function and the policy function, so that the agent selects actions based on the value function and the policy function and interacts with the power system simulation environment. The two independent networks are the value network and the policy network. For the value network, the loss function L(θ Q ) to optimize network parameters: L(θ Q )=E[y t -Q(s t ,a t |θ Q )] 2 Among them, y t -Q(s t ,a t |θ Q ) is the timing difference error, y t is the target Q value; The policy network improves actions based on gradient information, and the sampling policy gradient is shown as follows: Among them, a Q(s,a|θ Q ) is the gradient information.
9. An electronic device, characterized in that: include: processor; A memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the steps of the method as claimed in any one of claims 1 to 4.
10. A computer storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.