Demand side source load storage system optimization method and device based on digital model enhancement

By establishing a digital current model in a multi-microgrid system and combining multi-agent deep reinforcement learning, the problem of modeling difficulties in traditional methods is solved, efficient optimization and real-time control of the multi-microgrid system is achieved, and the stability and response capabilities of the system are improved.

CN120454201APending Publication Date: 2025-08-08ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510651229.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

Traditional methods are difficult to accurately model trends and hot trends in multi-micro grid systems, resulting in poor optimization effects of source load storage systems on the demand side. Especially when the parameters and topological structure of electrical and thermal networks are difficult to determine, the existing technology cannot effectively improve the economic benefits and power supply reliability of the system.

Method used

By establishing a digital trend model, using a supervised learning mechanism to train it, combining a multi-agent deep reinforcement learning model, a target control framework is established, and a decision output network and a decision-making guidance network are used for optimization processing, real-time coordinated control of the multi-microgrid system is achieved.

Benefits of technology

It has improved the optimization effect of the demand-side source load storage system of the multi-micro grid system, and can operate stably and efficiently in a complex and changeable energy environment, adapt to new energy fluctuations, and improves the system's real-time response and expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120454201A_ABST
    Figure CN120454201A_ABST
Patent Text Reader

Abstract

The invention relates to a demand side source load storage system optimization method and device based on digital model enhancement. The method comprises the following steps: establishing a digital power flow model based on actual operation data of a demand side source load storage system; training the digital power flow model by adopting a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used for predicting the power grid power flow state, and the power grid power flow state comprises the voltage of each node in the multi-micro-grid system and the interaction power between the multi-micro-grid system and an external main grid; establishing a target control framework of the multi-microgrid system according to the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework comprises a to-be-optimized target function; and optimizing the objective function by adopting a decision output network and a decision guide network of the multi-micro-grid system. By adopting the method, a control framework independent of a physical model can be established, and the optimization effect of the demand side source load storage system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of multi-microgrids, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for optimizing a demand-side source-load-storage system based on digital model enhancement. Background Art

[0002] As renewable energy becomes increasingly important in the power system, microgrids, as their key application platform, can improve the economic benefits and power supply reliability of the power system through regional energy sharing and the construction of interconnected multi-microgrid systems.

[0003] Currently, technologies for refined management of multi-microgrids with demand-side source-load-storage systems primarily rely on traditional optimization methods. These methods construct mathematical models and employ online iterative optimization algorithms to achieve active power control of energy storage systems and dispatchable loads, as well as reactive power regulation of renewable energy inverters. Traditional methods rely on precise electrical and thermal network parameters and topologies for function calculations, but obtaining accurate data is extremely difficult in reality. Line parameters and topologies in electrical systems are also difficult to determine, and the complex response mechanisms and user distribution in thermal systems make accurate modeling a significant challenge.

[0004] Therefore, there is a problem in the related technology that the optimization effect of the demand-side source-load-storage system is poor. Summary of the Invention

[0005] Based on this, it is necessary to provide a demand-side source-load-storage system optimization method, device, computer equipment, computer-readable storage medium and computer program product based on digital model enhancement, which can improve the optimization effect of the demand-side source-load-storage system in response to the above technical problems.

[0006] In a first aspect, the present application provides a method for optimizing a demand-side source-load-storage system based on digital model enhancement, the method comprising:

[0007] A digital power flow model is established based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system comprising a plurality of microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner;

[0008] The digital power flow model is trained using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes a voltage of each node in the multi-microgrid system and an interaction power between the multi-microgrid system and an external main grid;

[0009] Establishing a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized;

[0010] The objective function is optimized using the decision output network and decision guidance network of the multi-microgrid system.

[0011] In one embodiment, establishing a digital power flow model based on actual operating data of the demand-side source-load-storage system includes:

[0012] Acquire actual operating data of the multi-microgrid system in a preset historical period as sample data; the sample data is a small sample type, and the sample data includes a known information sample and a predicted information sample;

[0013] Taking the known information sample as input and the predicted information sample as output, constructing a mapping relationship between the input and the output to obtain the digital power flow model;

[0014] Among them, the known information samples include load demand samples, renewable energy output samples and decision information samples of each of the nodes in the multi-microgrid system; the predicted information samples include voltage samples of each of the nodes in the multi-microgrid system, and interaction power samples between the multi-microgrid system and the external main grid.

[0015] In one embodiment, the adopting a supervised learning mechanism to train the digital power flow model to obtain a trained digital power flow model includes:

[0016] Based on the sparse variational Gaussian process, the mapping relationship between the input and output in the digital power flow model is optimized, and the posterior distribution of the predicted value in the digital power flow model is optimized by using the variational inference technology to obtain the trained digital power flow model.

[0017] In one embodiment, the optimizing the mapping relationship between input and output in the digital power flow model based on a sparse variational Gaussian process, and optimizing the posterior distribution of predicted values in the digital power flow model using a variational inference technique, include:

[0018] The variational inference technique is adopted to obtain an optimal variational posterior distribution through a first maximization of the evidence lower bound process, and to obtain optimized parameters for the sparse variational Gaussian process through a second maximization of the evidence lower bound process.

[0019] In one embodiment, the objective function to be optimized is expressed as:

[0020]

[0021] in, The electricity interaction cost between the multi-microgrid system and the external main grid; is the penalty term used to maintain a stable voltage; is the operating cost and fuel cost of the diesel generator of the g-th microgrid; G is the number of microgrids in the multi-microgrid system.

[0022] In one embodiment, the optimizing process of the objective function using the decision output network and the decision guidance network of the multi-microgrid system includes:

[0023] Based on the output decision of the decision output network guided by the decision guidance network, the parameters are updated in the direction of the optimal decision, and an update formula of the decision output network is obtained; the output decision of the decision output network includes the active variable and the reactive variable of the multi-microgrid system;

[0024] According to the update formula of the decision output network, the active variables and reactive variables of the multi-microgrid system are optimized, and the optimization processing of the objective function is performed.

[0025] In a second aspect, the present application further provides a demand-side source-load-storage system optimization device based on digital model enhancement, the device comprising:

[0026] A digital power flow model building module is used to establish a digital power flow model based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system including multiple microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner;

[0027] a model training module, configured to train the digital power flow model using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes the voltage of each node in the multi-microgrid system and the interaction power between the multi-microgrid system and an external main grid;

[0028] A control framework establishment module is used to establish a target control framework of the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized;

[0029] The optimization processing module is used to optimize the objective function using the decision output network and decision guidance network of the multi-microgrid system.

[0030] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.

[0031] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when executed by a processor.

[0032] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the above method when executed by a processor.

[0033] The above-mentioned demand-side source-load-storage system optimization method, device, computer equipment, computer-readable storage medium and computer program product enhanced by digital model establish a digital power flow model based on the actual operation data of the demand-side source-load-storage system. The demand-side source-load-storage system is a multi-microgrid system including multiple microgrids, each microgrid is equipped with a microgrid controller, and each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner. Then, a supervised learning mechanism is used to train the digital power flow model to obtain a trained digital power flow model. The trained digital power flow model is used to predict the power grid power flow state. The power grid power flow state includes the voltage of each node in the multi-microgrid system and the interaction power between the multi-microgrid system and the external main grid. According to the trained digital power flow model and the multi-agent deep reinforcement learning model, a target control framework of the multi-microgrid system is established. The target control framework includes the objective function to be optimized, and then the decision output network and decision guidance network of the multi-microgrid system are used to optimize the objective function. By using digital power flow models to realize the real calculation of power flow and thermal flow, and incorporating the trained digital power flow models into multi-agent deep reinforcement learning control to establish a control framework that does not rely on physical models, the optimization effect of the demand-side source-load-storage system is effectively improved, and the multi-microgrid demand response can be realized based on multi-agent deep reinforcement learning enhanced by digital models. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0035] Figure 1 1 is a flow chart of a method for optimizing a demand-side source-load-storage system based on digital model enhancement in one embodiment;

[0036] Figure 2 A schematic flow chart of the optimization process steps in one embodiment;

[0037] Figure 3 Schematic diagram of a flow chart of a demand-side source-load-storage system optimization method based on digital model enhancement in another embodiment;

[0038] Figure 4 This is a structural block diagram of a demand-side source-load-storage system optimization device based on digital model enhancement in one embodiment;

[0039] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0040] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0041] Renewable energy has become an integral part of the power system, effectively alleviating the dual challenges of energy shortages and environmental pollution. As a key platform for renewable energy applications, microgrids are increasingly prominent in modern power systems. However, due to the regional nature of power supply, microgrids' energy supply capacity is inherently limited. To address this shortcoming, regional energy sharing mechanisms between adjacent microgrids have successfully overcome the capacity bottleneck of a single microgrid. This allows for direct connection of isolated microgrids to the external grid, while also integrating multiple microgrids into interconnected multi-microgrid systems for unified access. Control strategies based on interconnected modes not only improve the economic benefits of demand response within the multi-microgrid system itself but also enhance the power supply reliability of the external grid. Research on coordinated control strategies for multi-microgrid systems has become a focus of both academia and industry.

[0042] Traditional technologies for refined microgrid management primarily rely on traditional optimization methods to achieve active power control of energy storage systems and dispatchable loads, as well as reactive power regulation of renewable energy inverters. This traditional approach typically follows a systematic process: First, researchers or engineers, based on in-depth theoretical analysis and practical experience, construct a detailed and accurate mathematical model for the multi-microgrid system. This model incorporates not only the charging and discharging characteristics of the energy storage system, the power regulation range of the dispatchable loads, but also the reactive power compensation capabilities of the renewable energy inverters. Once the model is constructed, the optimization process begins. This process often employs an online iterative optimization approach. Based on real-time power system data, the values of control variables are continuously adjusted to find a solution that maximizes economic benefits or optimizes other pre-defined objectives (such as reducing carbon emissions or improving energy utilization) while ensuring safe and stable system operation. Online iterative optimization methods include gradient descent, genetic algorithms, and particle swarm optimization.

[0043] Because traditional methods rely on power and heat flows for the calculation of objective functions, they require a precise understanding of the parameters and topology of the power and thermal networks. However, developing accurate physical models based on real-world electrical and thermal systems is a difficult task; for electrical systems, it is difficult for operators to determine the exact line parameters and topology, and accurately estimating electrical system characteristics is also challenging due to the need for a large number of time-stamped measurements or integrated data from phasor measurement units. Similarly, for thermal systems, accurate pipeline parameters and topology are crucial for calculating heat flow. Given the complex response mechanisms, high nonlinearity, and significant hysteresis in thermal systems, it is difficult to obtain accurate pipeline parameters. In addition, the wide distribution of users within the thermal network complicates the topology of secondary pipelines, posing a challenge to accurately establishing a real-world mathematical model of the thermal network.

[0044] In response to the problem that the above-mentioned traditional optimization methods rely on the clear parameters and structure of the electrical and thermal networks, making it difficult to accurately model the power flow and thermal flow to determine the real results, this application proposes a demand-side source-load-storage system optimization method based on digital model enhancement. By using a digital power flow model to realize the real calculation of power flow and thermal flow, the trained digital power flow model can be incorporated into the multi-agent deep reinforcement learning control setting, thereby establishing a control framework that does not rely on physical models.

[0045] In an exemplary embodiment, Figure 1As shown, a method for optimizing a demand-side source-load-storage system based on digital model enhancement is provided. This embodiment uses the method applied to a terminal as an example. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps 101 to 104. Among them:

[0046] Step 101: Establish a digital power flow model based on actual operation data of the demand-side source-load-storage system.

[0047] The demand-side source-load-storage system is a multi-microgrid system consisting of multiple microgrids. Each microgrid can be configured with a microgrid controller, which can autonomously coordinate and control the multi-microgrid system. By adopting a distributed deployment strategy, a microgrid controller is configured within each microgrid. Each microgrid controller has the ability to operate independently and autonomously coordinate and control the multi-microgrid system. Through the mutual cooperation of the microgrid controllers, the stable operation of the entire multi-microgrid system can be maintained.

[0048] In practical applications, unlike traditional methods, the demand-side source-load-storage system optimization method based on digital model enhancement in this embodiment does not require complex mathematical modeling work for multi-microgrid systems. Instead, a digital power flow model is established through the actual operation data of the multi-microgrid system to further obtain the final control strategy after the digital power flow model is established and the training of the control strategy is assisted.

[0049] Step 102: Train the digital power flow model using a supervised learning mechanism to obtain a trained digital power flow model.

[0050] As an example, the trained digital power flow model can be used to predict the power flow state of the power grid, which may include the voltage of each node in the multi-microgrid system and the interaction power between the multi-microgrid system and the external main grid.

[0051] In specific implementation, in order to address the key challenge of accurately simulating and predicting the power grid flow state faced in the research of multi-microgrid systems, the demand-side source-load-storage system optimization method based on digital model enhancement in this embodiment proposes the concept of an agent model (i.e., a digital power flow model). The agent model can be trained with the help of a supervised learning mechanism, thereby having the ability to accurately simulate and predict the power grid flow state.

[0052] For example, by providing the proxy model with the load demand of each node, the output of renewable energy, and the decision information of the controller, the proxy model can effectively estimate the power exchange between the multi-microgrid system and the external main grid, as well as the voltage information of each node in the multi-microgrid system based on the known information obtained.

[0053] Step 103: Establish a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized.

[0054] Specifically, by using a digital power flow model to realize the realistic calculation of power flow and thermal flow, the trained digital power flow model can be incorporated into the multi-agent deep reinforcement learning control setting based on the multi-agent deep reinforcement learning model to establish a control framework that does not rely on the physical model (i.e., the target control framework).

[0055] Step 104 : Optimizing the objective function using the decision output network and decision guidance network of the multi-microgrid system.

[0056] In one example, the reward function is equivalent to the objective function, which is the core optimization goal of the demand-side source-load-storage system optimization method enhanced based on the digital model in this embodiment. By treating each microgrid in the multi-microgrid system as an intelligent agent, a reinforcement learning method is adopted, and the actor network (i.e., the decision output network) and the critic network (i.e., the decision guidance network) based on the policy gradient algorithm are used for policy optimization and value function estimation. They can work together to improve the learning efficiency of each intelligent agent, thereby enabling the multi-microgrid demand response to be realized based on the multi-agent deep reinforcement learning enhanced by the digital model.

[0057] Compared to traditional methods, which have a strong dependence on specific interconnected multi-microgrid environments, the versatility of traditional optimization strategies is greatly limited when applied in different multi-microgrid environments. The technical solution of this embodiment demonstrates excellent generalization performance, giving the microgrid controller a unique advantage. Based on this feature, the microgrid controller can quickly respond to fluctuations in renewable energy to achieve real-time coordinated control of the multi-microgrid system, effectively ensuring the stable and efficient operation of the multi-microgrid system even in complex and changing energy environments, and facilitating the expansion of its application to new control scenarios.

[0058] In terms of the ability to cope with environmental changes, as the proportion of renewable energy in multi-microgrid systems gradually increases, its short-term volatility is extremely strong, requiring the system to have a control strategy capable of real-time decision-making to effectively cope with the fluctuations of renewable energy. Traditional methods, however, struggle to respond promptly to changing conditions and require continuous resolution of emerging issues, which can very likely cause delays in real-time decision-making. The technical solution of this embodiment, based on its excellent generalization, can meet the requirements of real-time control of multi-microgrid systems, providing a strong guarantee for the stable operation of the system.

[0059] In the above-mentioned demand-side source-load-storage system optimization method based on digital model enhancement, a digital power flow model is established based on the actual operating data of the demand-side source-load-storage system. The digital power flow model is then trained using a supervised learning mechanism to obtain a trained digital power flow model. Based on the trained digital power flow model and the multi-agent deep reinforcement learning model, a target control framework for the multi-microgrid system is established. The decision output network and decision guidance network of the multi-microgrid system are then used to optimize the objective function. By using the digital power flow model to realize the actual calculation of power flow and thermal power flow, and incorporating the trained digital power flow model into the multi-agent deep reinforcement learning control to establish a control framework that does not rely on the physical model, the optimization effect of the demand-side source-load-storage system is effectively improved, and the multi-microgrid demand response can be realized based on the multi-agent deep reinforcement learning enhanced by the digital model.

[0060] In an exemplary embodiment, establishing a digital power flow model based on actual operating data of the demand-side source-load-storage system may include the following steps:

[0061] Actual operating data of the multi-microgrid system in a preset historical period is obtained as sample data; the sample data belongs to a small sample type, and the sample data includes a known information sample and a predicted information sample; the known information sample is used as input and the predicted information sample is used as output, a mapping relationship between the input and the output is constructed, and the digital power flow model is obtained.

[0062] Among them, the known information samples include the load demand samples of each node in the multi-microgrid system, the renewable energy output samples, and the decision information samples of each microgrid controller; the predicted information samples include the voltage samples of each node in the multi-microgrid system, and the interaction power samples between the multi-microgrid system and the external main grid.

[0063] In practical applications, the proxy model constructs the mapping relationship between input and output variables by learning from historical datasets. Unlike deep learning technologies, which rely on massive amounts of data, historical data is relatively scarce in multi-microgrid system applications. The proxy model proposed in this embodiment has the unique advantage of small-sample learning. It can accurately construct complex mapping relationships between input and output variables under limited data resources.

[0064] In an alternative embodiment, assuming the input and output If the Gaussian process is followed, then we can conclude that:

[0065] (1)

[0066] In formula (1), represents Gaussian noise; represents the variance; I is the identity matrix with appropriate dimension; Represents a Gaussian process mapping between input and output:

[0067] (2)

[0068] in, , the mean and covariance matrices are respectively given by and If the prior number is small, we can assume =0, covariance matrix The elements in can be represented as:

[0069] (3)

[0070] in, Denotes the covariance. After specifying m and k, the prior distribution of y can be expressed as:

[0071] (4)

[0072] When a new input x appears * When y and The joint prior distribution of is:

[0073] (5)

[0074] The posterior distribution of the predicted values can be derived from this:

[0075] (6)

[0076] in, as well as They represent the mean and variance of the test data, which can be calculated according to the following formula:

[0077] (7)

[0078] (8)

[0079] When the relevant parameters are known , the mean and variance of the test data can be calculated according to formulas (7) and (8). The update is done by maximizing To achieve it.

[0080] In this embodiment, by obtaining the actual operating data of the multi-microgrid system in a preset historical period as sample data, and then taking the known information sample as input and the predicted information sample as output, a mapping relationship between the input and output is constructed to obtain a digital power flow model, which can achieve accurate modeling and prediction of the operating status of the multi-microgrid system.

[0081] In an exemplary embodiment, the method of training the digital power flow model using a supervised learning mechanism to obtain a trained digital power flow model may include the following steps:

[0082] Based on the sparse variational Gaussian process, the mapping relationship between the input and output in the digital power flow model is optimized, and the posterior distribution of the predicted value in the digital power flow model is optimized by using the variational inference technology to obtain the trained digital power flow model.

[0083] Optionally, the computational complexity of the calculation steps of the above formulas (1) to (8) is , N represents the size of the input or output data. When processing large amounts of data, the Gaussian process will face greater computational pressure. In order to reduce the computational complexity, a surrogate model based on sparse variational Gaussian process can be used. This model introduces M (M is much smaller than N) auxiliary inputs and the corresponding output By introducing M induction points to approximate the original N-dimensional Gaussian process regression, the computational complexity can be reduced to For example, the joint density of y, f, and u is:

[0084] (9)

[0085] in, as well as If you know , we can get the above joint density. Since u and f are generated by the same Gaussian process, there is the following relationship between u and f:

[0086] (10)

[0087] in, as well as Represent the mean and variance of f, respectively, which can be obtained by the following formula:

[0088] (11)

[0089] (12)

[0090] In one example, to obtain the true posterior To reduce computational complexity, variational inference techniques can be used. Variational inference works by minimizing the Kullback-Leibler divergence between the variational posterior q and the true posterior p to find the optimal variational posterior distribution. The Kullback-Leibler divergence measures the distance between two distributions; smaller values indicate closer distributions.

[0091] In this embodiment, the mapping relationship between input and output in the digital power flow model is optimized based on the sparse variational Gaussian process, and the posterior distribution of the predicted value in the digital power flow model is optimized using variational inference technology to obtain a trained digital power flow model, which can effectively improve the accuracy and computational efficiency of the model.

[0092] In an exemplary embodiment, the steps of optimizing the mapping relationship between input and output in the digital power flow model based on a sparse variational Gaussian process and optimizing the posterior distribution of predicted values in the digital power flow model using a variational inference technique may include the following steps:

[0093] The variational inference technique is adopted to obtain an optimal variational posterior distribution through a first maximization of the evidence lower bound process, and to obtain optimized parameters for the sparse variational Gaussian process through a second maximization of the evidence lower bound process.

[0094] In practical applications, the variational posterior can be found by maximizing the evidence lower bound (i.e., the first maximization of the evidence lower bound process):

[0095] (13)

[0096] The lower limit of evidence can be simplified as:

[0097] (14)

[0098] Therefore, the Gaussian marginal probability It can be expressed as:

[0099] (15)

[0100] in, , and It can be expressed as:

[0101] (16)

[0102] (17)

[0103] From the above content, we can see that the optimization parameters of sparse variational Gaussian process include , the optimization basis of these parameters is to maximize formula (14), that is, the second maximum evidence lower limit processing. After obtaining the trained parameters by maximizing the evidence lower limit, we can enter the testing phase. At this time, enter a new test point , according to the calculation of formula (15), the posterior distribution of the predicted value can be obtained:

[0104] (18)

[0105] According to the calculation results of formula (18), the voltage of each node in the multi-microgrid system and the interaction power between the multi-microgrid system and the external main grid can be obtained.

[0106] In this embodiment, by adopting variational inference technology, the optimal variational posterior distribution is obtained through the first maximization of the evidence lower bound processing, and the optimization parameters for the sparse variational Gaussian process are obtained through the second maximization of the evidence lower bound processing, which can achieve efficient probability modeling and sparsification processing.

[0107] In an exemplary embodiment, the objective function to be optimized is expressed as:

[0108]

[0109] in, The cost of electricity interaction between the multi-microgrid system and the external main grid; is the penalty term used to maintain a stable voltage; is the operating cost and fuel cost of the diesel generator of the g-th microgrid; G is the number of microgrids in the multi-microgrid system.

[0110] Specifically, it is assumed that there is power interaction between the multi-microgrid system and the external main grid, and the power interaction cost is recorded as ; The penalty term set to maintain stable voltage is , where the voltage penalty term of the g-th microgrid is defined as , the total number of nodes in the multi-microgrid system is I; the operating cost and fuel cost of the diesel generator in the g-th microgrid are expressed as The total number of microgrids in the multi-microgrid system is represented by G, and the interaction power between the multi-microgrid system and the external main grid is , the node voltage in the system is The fuel cost of DEG is calculated by the formula The calculation cost is , a, b, c are the coefficients in the diesel generator fuel cost equation, It is the start-stop loss of the diesel generator. represents the working state of the diesel generator, then the reward function (i.e. the objective function to be optimized) can be expressed as: .

[0111] In an exemplary embodiment, Figure 2 As shown, step 104, using the decision output network and decision guidance network of the multi-microgrid system, optimizes the objective function, including steps 201 to 202.

[0112] Step 201: Based on the output decision of the decision output network guided by the decision guidance network, the parameters are updated in the direction of the optimal decision to obtain an update formula for the decision output network; the output decision of the decision output network includes the active variables and reactive variables of the multi-microgrid system.

[0113] Step 202 : Optimizing the active variables and reactive variables of the multi-microgrid system according to the update formula of the decision output network, and performing optimization processing on the objective function.

[0114] In a specific implementation, the multi-agent deep reinforcement learning model can consist of two neural network-based parts, where the actor network (i.e., the decision output network) is responsible for outputting decisions that cover the active and reactive variables of the microgrid system. ,like represents the active power of the energy storage system at the i-th node of the g-th microgrid, Indicates the active power of the diesel generator, Represents the reactive output of photovoltaic, Indicates the reactive output of the fan. represents the reactive power of the energy storage system, Represents the reactive output of the diesel generator. By optimizing these variables, the reward function can be optimized, thereby achieving the goal of optimizing the operation of the microgrid.

[0115] The other component is the critic network (also known as the decision guidance network), whose primary function is to guide the actor network toward the optimal policy. Because the decisions output by the actor network rely on the guidance and improvement of the critic network, the design of the critic network plays a key role in the entire method.

[0116] In one example, the critic network consists of three main components: Multilayer Perceptron 1 (MLP1), the attention layer, and MLP2. These three components process the data stream in tandem, processing the raw input data in sequence. MLP2 ultimately outputs the processed results.

[0117] For example, suppose there is a multi-microgrid system consisting of G microgrids. represents the state of the g-th microgrid, Represents the action variable of the g-th microgrid. Taking the first microgrid as an example, the mathematical expression of MLP1 is:

[0118]

[0119] Among them, concat represents the concatenation operation, ReLu represents the rectified linear unit, is the input of the u-th hidden layer, and are the weight parameters and bias terms of the u-th hidden layer of the first agent, U is the number of hidden layers, The hidden layer is composed of U layers from as well as The features extracted from and are the weights and biases of the first agent’s output layer.

[0120] In another example, the mathematical expression of the attention mechanism layer is as follows. Taking the first agent as an example, the calculation process involves operations on data related to other microgrids.

[0121]

[0122] in, , represent and dimension. It is used to measure the attention paid by the first microgrid to the operating conditions of other microgrids when it is controlled by its controller. The contributions of other microgrid controllers are calculated by the weighted sum of their values:

[0123]

[0124] After repeating the above calculation S times, the results of S times are integrated to obtain the attention features [O1, O2, ..., OS], which can then be input to MLP2; MLP2 can extract features from the input e1 and the attention features [O1, O2, ..., OS] and map them to the action-value function of the first agent, thereby guiding the actor network to output decisions and update parameters in the direction of the optimal decision. The calculation method is as follows:

[0125]

[0126] in, is the input of the e-th hidden layer, 、 are the weight parameters and bias of the e-th hidden layer of the first microgrid controller, is the output of the E-th hidden layer, and are the weight parameters and bias terms of the output layer of the first microgrid controller MLP2.

[0127] The mathematical expression of the actor network structure is:

[0128]

[0129] in, The input of the pth hidden layer, It is the first microgrid controller The weight parameters and bias of the hidden layer, is the output of the Pth hidden layer, and are the weight parameters and bias terms for the output decision of the first microgrid controller.

[0130] The update formula of the actor network is:

[0131]

[0132] Where G represents the number of microgrids in the multi-microgrid system, is the temperature parameter, which is used to adjust the amplitude of the update parameters. The value of comes from the output of MLP2 in the above formula, which is the key to the critic network's ability to guide the actor network parameter update.

[0133] In this embodiment, by guiding the output decision of the decision output network based on the decision guidance network, the parameters are updated in the direction of the optimal decision, and the update formula of the decision output network is obtained. Then, according to the update formula of the decision output network, the active variables and reactive variables of the multi-microgrid system are optimized, and the optimization processing of the objective function is performed. It is possible to dynamically adjust the active and reactive variables of the multi-microgrid system, achieve efficient optimization of the objective function, and help improve the overall performance of the system.

[0134] In an exemplary embodiment, Figure 3 As shown, a flow chart of another method for optimizing a demand-side source-load-storage system based on digital model enhancement is provided. In this embodiment, the method includes the following steps:

[0135] In step 301, actual operating data of the multi-microgrid system over a preset historical period is obtained as sample data. This sample data is a small sample type and includes known information samples and predicted information samples. In step 302, the known information samples are used as input and the predicted information samples as output, and a mapping relationship between the input and output is established to obtain a digital power flow model. In step 303, the mapping relationship between the input and output of the digital power flow model is optimized using a sparse variational Gaussian process, and the posterior distribution of the predicted values in the digital power flow model is optimized using variational inference techniques, resulting in a trained digital power flow model. In step 304, a target control framework for the multi-microgrid system is established based on the trained digital power flow model and a multi-agent deep reinforcement learning model. In step 305, the decision guidance network guides the output decisions of the decision output network, and the parameters are updated towards the optimal decision, resulting in an update formula for the decision output network. In step 306, the active and reactive variables of the multi-microgrid system are optimized based on the update formula of the decision output network, executing the optimization process for the objective function.

[0136] It should be noted that the specific definitions of the above steps can be found in the above specific definitions of a demand-side source-load-storage system optimization method based on digital model enhancement, which will not be repeated here.

[0137] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0138] Based on the same inventive concept, embodiments of the present application also provide a digital model-enhanced demand-side source-load-storage system optimization device for implementing the aforementioned digital model-enhanced demand-side source-load-storage system optimization method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the digital model-enhanced demand-side source-load-storage system optimization device provided below can be found in the above-mentioned limitations of the digital model-enhanced demand-side source-load-storage system optimization method, and will not be repeated here.

[0139] In an exemplary embodiment, Figure 4 As shown, a demand-side source-load-storage system optimization device based on digital model enhancement is provided, comprising:

[0140] A digital power flow model building module 401 is configured to establish a digital power flow model based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system comprising a plurality of microgrids, each of which is equipped with a microgrid controller; and each microgrid controller autonomously coordinates and controls the multi-microgrid system.

[0141] A model training module 402 is configured to train the digital power flow model using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, including the voltage of each node in the multi-microgrid system and the interaction power between the multi-microgrid system and an external main grid;

[0142] A control framework establishment module 403 is configured to establish a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized;

[0143] The optimization processing module 404 is used to optimize the objective function using the decision output network and the decision guidance network of the multi-microgrid system.

[0144] In one embodiment, the digital power flow model construction module 401 is specifically used to obtain the actual operation data of the multi-microgrid system in a preset historical period as sample data; the sample data belongs to a small sample type, and the sample data includes a known information sample and a predicted information sample; the known information sample is used as input and the predicted information sample is used as output, and a mapping relationship between the input and the output is constructed to obtain the digital power flow model; wherein, the known information sample includes the load demand sample of each node in the multi-microgrid system, the renewable energy output sample and the decision information sample of each microgrid controller; the predicted information sample includes the voltage sample of each node in the multi-microgrid system, and the interaction power sample between the multi-microgrid system and the external main grid.

[0145] In one embodiment, the model training module 402 is specifically used to optimize the mapping relationship between the input and output in the digital power flow model based on a sparse variational Gaussian process, and to optimize the posterior distribution of the predicted value in the digital power flow model using variational inference technology to obtain the trained digital power flow model.

[0146] In one embodiment, the model training module 402 is further used to adopt the variational inference technology to obtain the optimal variational posterior distribution through a first maximization of the evidence lower bound processing, and to obtain the optimization parameters for the sparse variational Gaussian process through a second maximization of the evidence lower bound processing.

[0147] In one embodiment, the objective function to be optimized is expressed as:

[0148]

[0149] in, The electricity interaction cost between the multi-microgrid system and the external main grid; is the penalty term used to maintain a stable voltage; is the operating cost and fuel cost of the diesel generator of the g-th microgrid; G is the number of microgrids in the multi-microgrid system.

[0150] In one embodiment, the optimization processing module 404 is specifically used to guide the output decision of the decision output network based on the decision guidance network, update parameters in the direction of the optimal decision, and obtain an update formula of the decision output network; the output decision of the decision output network includes the active variables and reactive variables of the multi-microgrid system; according to the update formula of the decision output network, the active variables and reactive variables of the multi-microgrid system are optimized, and the optimization processing of the objective function is performed.

[0151] Each module in the aforementioned digital model-enhanced demand-side source-load-storage system optimization device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in the computer device's memory in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0152] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication, and the wireless communication can be implemented via Wi-Fi, a mobile cellular network, near-field communication (NFC), or other technologies. When executed by the processor, the computer program implements a demand-side source-load-storage system optimization method based on digital model enhancement.

[0153] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0154] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0155] A digital power flow model is established based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system comprising a plurality of microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner;

[0156] The digital power flow model is trained using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes a voltage of each node in the multi-microgrid system and an interaction power between the multi-microgrid system and an external main grid;

[0157] Establishing a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized;

[0158] The objective function is optimized using the decision output network and decision guidance network of the multi-microgrid system.

[0159] In one embodiment, when the processor executes the computer program, it also implements the steps of the demand-side source-load-storage system optimization method based on digital model enhancement in the other embodiments described above.

[0160] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0161] A digital power flow model is established based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system comprising a plurality of microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner;

[0162] The digital power flow model is trained using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes a voltage of each node in the multi-microgrid system and an interaction power between the multi-microgrid system and an external main grid;

[0163] Establishing a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized;

[0164] The objective function is optimized using the decision output network and decision guidance network of the multi-microgrid system.

[0165] In one embodiment, when the computer program is executed by the processor, it also implements the steps of the demand-side source-load-storage system optimization method based on digital model enhancement in the other embodiments described above.

[0166] In one embodiment, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the following steps:

[0167] A digital power flow model is established based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system comprising a plurality of microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner;

[0168] The digital power flow model is trained using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes a voltage of each node in the multi-microgrid system and an interaction power between the multi-microgrid system and an external main grid;

[0169] Establishing a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized;

[0170] The objective function is optimized using the decision output network and decision guidance network of the multi-microgrid system.

[0171] In one embodiment, when the computer program is executed by the processor, it also implements the steps of the demand-side source-load-storage system optimization method based on digital model enhancement in the other embodiments described above.

[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0173] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0174] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0175] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A demand-side source-load-storage system optimization method based on digital model enhancement, characterized in that: The method comprises: A digital power flow model is established based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system comprising a plurality of microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner; The digital power flow model is trained using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes a voltage of each node in the multi-microgrid system and an interaction power between the multi-microgrid system and an external main grid; Establishing a target control framework for the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized; The objective function is optimized using the decision output network and decision guidance network of the multi-microgrid system.

2. The method according to claim 1, characterized in that The digital power flow model is established based on the actual operation data of the demand-side source-load-storage system, including: Acquire actual operating data of the multi-microgrid system in a preset historical period as sample data; the sample data is a small sample type, and the sample data includes a known information sample and a predicted information sample; Taking the known information sample as input and the predicted information sample as output, constructing a mapping relationship between the input and the output to obtain the digital power flow model; Among them, the known information samples include load demand samples, renewable energy output samples and decision information samples of each of the nodes in the multi-microgrid system; the predicted information samples include voltage samples of each of the nodes in the multi-microgrid system, and interaction power samples between the multi-microgrid system and the external main grid.

3. The method according to claim 1, characterized in that The method of training the digital power flow model using a supervised learning mechanism to obtain a trained digital power flow model includes: Based on the sparse variational Gaussian process, the mapping relationship between the input and output in the digital power flow model is optimized, and the posterior distribution of the predicted value in the digital power flow model is optimized by using the variational inference technology to obtain the trained digital power flow model.

4. The method according to claim 3, characterized in that The method of optimizing the mapping relationship between input and output in the digital power flow model based on a sparse variational Gaussian process and optimizing the posterior distribution of predicted values in the digital power flow model using a variational inference technique includes: The variational inference technique is adopted to obtain an optimal variational posterior distribution through a first maximization of the evidence lower bound process, and to obtain optimized parameters for the sparse variational Gaussian process through a second maximization of the evidence lower bound process.

5. The method according to claim 1, wherein The objective function to be optimized is expressed as: in, The electricity interaction cost between the multi-microgrid system and the external main grid; is the penalty term used to maintain a stable voltage; is the operating cost and fuel cost of the diesel generator of the g-th microgrid; G is the number of microgrids in the multi-microgrid system.

6. The method according to any one of claims 1 to 5, characterized in that The method of optimizing the objective function by using the decision output network and the decision guidance network of the multi-microgrid system includes: Based on the output decision of the decision output network guided by the decision guidance network, the parameters are updated in the direction of the optimal decision, and an update formula of the decision output network is obtained; the output decision of the decision output network includes the active variable and the reactive variable of the multi-microgrid system; According to the update formula of the decision output network, the active variables and reactive variables of the multi-microgrid system are optimized, and the optimization processing of the objective function is performed.

7. A demand-side source-load-storage system optimization device based on digital model enhancement, characterized in that: The device comprises: A digital power flow model building module is used to establish a digital power flow model based on actual operating data of a demand-side source-load-storage system; the demand-side source-load-storage system is a multi-microgrid system including multiple microgrids, each of which is equipped with a microgrid controller; each microgrid controller implements coordinated control of the multi-microgrid system in an autonomous manner; a model training module, configured to train the digital power flow model using a supervised learning mechanism to obtain a trained digital power flow model; the trained digital power flow model is used to predict a power grid power flow state, wherein the power grid power flow state includes the voltage of each node in the multi-microgrid system and the interaction power between the multi-microgrid system and an external main grid; A control framework establishment module is used to establish a target control framework of the multi-microgrid system based on the trained digital power flow model and the multi-agent deep reinforcement learning model; the target control framework includes an objective function to be optimized; The optimization processing module is used to optimize the objective function using the decision output network and decision guidance network of the multi-microgrid system.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Natural gas pipeline network operation failure model construction method and operation failure determination method

    CN121435436A