Dynamic reconstruction method and system of distribution network based on electrical-thermal coordination based on DDQN-KRR
The output of the coupling unit is predicted through the DDQN-KRR model, and the problem of insufficient consideration of electrical-gas-thermal coordination in the existing distribution network reconstruction method is solved, and rapid and economical dynamic reconstruction is achieved, reducing operating costs and improving system flexibility and voltage levels.
Patent Information
- Application Number
- CN202211318009.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-10-26
AI Technical Summary
The existing distribution network reconstruction method fails to effectively consider the coordination of electric-gas-thermal multi-energy flows, which makes it difficult to meet the requirements in terms of economy and carbon emissions, and the dynamic reconstruction calculation time is long and the strategy is not time-consuming.
The dynamic reconstruction method of electrical thermal collaborative distribution network based on DDQN-KRR is adopted, and the optimization objective function and constraints are constructed, and the output of the coupled unit is predicted using the DDQN-KRR model to achieve a rapid dynamic reconstruction strategy.
The rapid, economical and effective distribution network reconstruction is achieved when the electric-gas-heat multi-energy flow coordination is considered, which reduces operating costs, improves the system flexibility and voltage level, and reduces the switching operation costs.
Smart Images

Figure CN115618538B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to distribution network reconstruction, and in particular to a distribution network dynamic reconstruction method and system based on DDQN-KRR electrical-thermal synergy. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Distribution networks using a single energy source are no longer able to meet economic and carbon emission requirements. Distribution network reconfiguration (DNR), which considers the coordinated flow of electricity, gas, and heat, has become an important means to improve the economic efficiency of distribution systems and accommodate new energy sources. With the development of machine learning theory, data-driven dynamic reconfiguration of distribution networks has, to some extent, addressed the problem that dynamic reconfiguration cannot meet the system's timeliness and economic requirements. Therefore, data-driven dynamic reconfiguration of distribution networks that considers integrated energy systems (IES) has become a research hotspot. Research on DNR is relatively comprehensive, and DNR can be divided into static and dynamic reconfiguration. While static reconfiguration is computationally efficient, it does not consider source and load variations over time, making it difficult to obtain an ideal reconfiguration strategy. Dynamic reconfiguration, because it requires time period division, takes a relatively long computational time, and the resulting reconfiguration strategy is often not timely.
[0004] The literature "Yang Ming, Zhai Hefeng, Ma Jiayi, et al. Dynamic reconstruction of three-phase asymmetric distribution network considering the imbalance constraint of distributed generation [J]. Proceedings of the CSEE, 2019, 39(12): 3486-3499" proposes a dynamic reconstruction method for distribution network considering the three-phase asymmetry of distributed generation (DG), which improves the safety and economy of the three-phase unbalanced network. The literature "Zhai Hefeng, Yang Ming, Zhao Ligang, et al. Robust dynamic reconstruction of distribution network to improve the acceptance capacity of distributed generation [J]. Automation of Electric Power Systems, 2019, 43(18): 35-42." Based on the two-stage robust optimization method, the dynamic reconstruction of the distribution network is used to improve the grid connection capability of uncontrollable distributed generation. References: "Yi Haichuan, Zhang Peter, Wang Haiying, et al. Dynamic Reconfiguration Method for Distribution Network to Improve DG Acceptance Capacity [J]. Power System Technology, 2016, 40(05): 1431-1436." In response to the problem of large-scale DG access to the distribution system, a dynamic reconstruction method for distribution network is proposed with the goal of improving DG acceptance capacity, reducing the risk of abandoning DG. "Li Yang, Wei Gang, Ma Yu, et al. Dynamic Reconfiguration of Active Distribution Network with Electric Vehicles and Distributed Generation [J]. Automation of Electric Power Systems, 2018, 42(05): 102-110." A dynamic reconstruction strategy based on interval number method with electric vehicles and distributed generation is proposed, which realizes the dynamic reconstruction of active distribution network under various uncertain factors. In recent years, the emergence of emerging algorithms has greatly reduced the time to obtain the reconstruction solution for dynamic reconstruction. The paper "MALEKSHAHS, RASOULI A, MALEKSHAH Y, et al. Reliability-driven distribution power network dynamic reconfiguration in presence of distributed generation by the deep reinforcement learning method[J]. Alexandria Engineering Journal, 2022, 61(8): 6541-6556." proposes a dynamic reconstruction method to improve the reliability of the distribution network in the presence of distributed generation through reinforcement learning algorithm.The literature "JI X, YIN Z, ZHANG Y, et al. Real-time autonomous dynamic reconfiguration based on deep learning algorithm for distribution network [J]. Electric Power Systems Research, 2021, 195 (03): 107132." constructed a real-time autonomous dynamic reconfiguration model of distribution network based on long short-term memory neural network (LSTM), which reduced the power loss and switching cost of distribution network. The literature "BUI V, SU W. Real-time operation of distribution network: A deep reinforcement learning-based reconfiguration approach [J]. Sustainable Energy Technologies and Assessments, 2022, 50 (01): 101841." proposed a dynamic network reconfiguration strategy based on the three-stage deep deterministic policy gradient algorithm (DDPG) to achieve rapid reconstruction and thus improve the stability and reliability of the distribution network. The above literatures have effectively promoted the development of the field of dynamic reconfiguration of distribution networks, but have not considered the impact of IES on the dynamic reconfiguration of distribution networks.
[0005] Comprehensively considering the deep integration of the electric-gas-heat IES and the traditional distribution system, and deeply exploring the flexibility resources of the IES system in the process of energy transmission and conversion, the effectiveness of the dynamic reconstruction of the distribution network can be effectively improved. The literature "Jin Xiaolong, Mu Yunfei, Jia Hongjie, et al. Optimal hybrid power flow calculation of regional integrated energy system considering distribution network reconstruction [J]. Automation of Electric Power Systems, 2017, 41(01): 18-24+56." proposed an optimal hybrid power flow algorithm for the electric-gas regional integrated energy system (ICES) considering distribution network reconstruction, optimized the ICES scheduling from multiple aspects of "source-grid-load", and reduced the ICES cost. The literature "Chen Zexing, Zhang Yongjun, Huang Yu, et al. Research on the optimization and reconstruction method of integrated energy distribution network based on conditional risk value [J]. Global Energy Internet, 2020, 3(06): 590-599." established an electricity-gas IES optimization operation model considering distribution network reconstruction, and used the conditional risk value theory to quantify the operation risk brought by IES uncertainty, thereby improving the operation capability of IES and making the distribution network reconstruction result more robust. The literature "Zhou Buxiang, Yao Xianyu, Zang Tianlei. Optimal reconstruction of integrated energy distribution network considering electricity-gas bidirectional coupling [J / OL]. Electrical Measurement and Instrumentation, 2021[2022-10-05]." introduced a power-to-gas device and constructed an electricity-gas bidirectional coupling IES model considering distribution network reconstruction. It was verified that distribution network reconstruction can reduce the operating cost of IES and improve the operating voltage level of the distribution system and the gas pressure of the gas distribution network. The paper "Li Peng, Wang Zixuan, Hou Lei, et al. Analysis of Optimal Operation of Regional Integrated Energy Systems Based on Repeated Game Theory [J]. Automation of Electric Power Systems, 2019, 43(14): 81-89" proposed an IES optimization method for power-gas regions based on repeated game theory, using distribution network reconstruction as one of the optimization methods to effectively reduce system operating costs. Although the above studies comprehensively consider IES and distribution network reconstruction, few studies have simultaneously considered the coordinated operation of power, heat, and gas networks. Moreover, existing studies mostly focus on static reconstruction of a single time section, without considering dynamic reconstruction on a continuous time scale. Summary of the Invention
[0006] In order to overcome the above-mentioned deficiencies of the prior art, the present invention provides a method and system for dynamic reconstruction of distribution network based on electrical and thermal synergy of DDQN-KRR.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: a method for dynamic reconstruction of a distribution network based on electrical-thermal coordination using DDQN-KRR, comprising:
[0008] Taking the minimization of the sum of the distribution network operation cost and the gas source cost of the gas distribution network as the optimization objective function, and the distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints as the constraints, a distribution network dynamic reconstruction model is constructed.
[0009] The dynamic reconstruction model of the distribution network is solved using DDQN-KRR. The output of the coupled units is predicted by the KRR model based on the node load and DG output. The reconstruction strategy is obtained by the DDQN model based on the node load, DG output and the coupled unit output data.
[0010] The second aspect of the present invention provides a distribution network dynamic reconstruction system based on DDQN-KRR electrical and thermal coordination, including:
[0011] Model building module: Taking the minimization of the sum of the distribution network operating cost and the gas source cost of the gas distribution network as the optimization objective function, and using the distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints as constraints, a distribution network dynamic reconstruction model is constructed;
[0012] Model solving module: Uses DDQN-KRR to solve the distribution network dynamic reconstruction model. The KRR model predicts the output of coupled units based on node loads and DG outputs. The DDQN model is used to obtain the reconstruction strategy based on the node loads, DG outputs and the output data of the coupled units.
[0013] One or more of the above technical solutions have the following beneficial effects:
[0014] (1) A dynamic reconstruction model of the electricity-gas-heat coordinated distribution network is proposed. Based on a continuous data set of the system's historical status, the DNR solution space is compressed through static reconstruction, and a database corresponding to the distribution network status is established. In this process, the switching action cost is considered according to the time series, and a dynamic reconstruction strategy is formulated.
[0015] (2) A dynamic reconstruction model based on DDQN is proposed, which transfers the time-consuming flow calculation of the traditional dynamic reconstruction model to the model training process, deeply mines the mapping relationship between the distribution network operation status and the reconstruction strategy in the database, and after the model training is completed, the reconstruction strategy can be quickly obtained directly according to the current state of the system.
[0016] (3) A coupled unit output prediction method based on KRR is proposed to explore the mapping rules between node load, DG output and coupled unit output. It can accurately deduce the nonlinear functional relationship between different data sets, thereby realizing the real-time acquisition of the optimal output of the coupled unit during online application.
[0017] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0019] Figure 1 Schematic diagram of the DDQN model improvement and encoding method in Example 1 of the present invention;
[0020] Figure 2 Schematic diagram of the model training and application process in Example 1 of the present invention;
[0021] Figure 3 Schematic diagram of KRR prediction results in Example 1 of the present invention;
[0022] Figure 4 The solution space changes after static reconstruction in the first embodiment of the present invention;
[0023] Figure 5 Schematic diagram comparing total line cost and voltage in four scenarios in Example 1 of the present invention;
[0024] Figure 6 This is a schematic diagram of the output of different power generation equipment in Example 1 of the present invention;
[0025] Figure 7 This is the E33-G20-H16 system diagram in Example 1 of the present invention. DETAILED DESCRIPTION
[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0027] It should be noted that the terms used herein are for describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present invention.
[0028] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0029] Example 1
[0030] This embodiment discloses a method for dynamic reconstruction of a distribution network based on electrical-thermal synergy using DDQN-KRR. First, a database containing the state of the distribution network with electrical-gas-thermal synergy and its static reconstruction results is established. A dynamic reconstruction strategy is formulated based on the time series and the cost of switching operations. Second, a kernel ridge regression method is proposed to predict the optimal output of coupled units. Finally, a two-layer deep Q network is used to mine the mapping relationship between the distribution network state and the reconstruction results to achieve rapid dynamic reconstruction.
[0031] In this embodiment, the method for dynamic reconfiguration of a distribution network based on electrical-thermal coordination of DDQN-KRR specifically includes:
[0032] Step 1: Taking the minimization of the sum of the distribution network operation cost and the gas source cost of the gas distribution network as the optimization objective function, and using the distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints as constraints, a distribution network dynamic reconstruction model is constructed;
[0033] Step 2: Use DDQN-KRR to solve the distribution network dynamic reconstruction model, where the KRR model is used to predict the output of the coupled units based on the node load and DG output; and the DDQN model is used to obtain the reconstruction strategy based on the node load, DG output and the coupled unit output data.
[0034] In step 1 of this embodiment, the optimization objective is to minimize the sum of distribution network line loss costs, switch operation costs, substation power purchase costs, and natural gas grid gas source costs. To improve the utilization rate of renewable energy power generation equipment, this embodiment does not consider its cost; since both coupled units use natural gas as fuel, this embodiment includes the fuel costs of both in the gas source cost. The objective function is:
[0035]
[0036] Where, and are the distribution network operation cost and gas source cost of the gas distribution network during period t respectively; C L 、 C sub 、C d and C gas are the distribution network line loss, distribution network power purchase, distribution network switch operation and gas source output cost coefficient respectively; B, Ω, K and S are the distribution network branch set, distribution network node set, gas distribution network set and gas source set respectively; δ b,t and is a 0-1 variable, when δ b,t =1 / 0 respectively indicate that the switch of branch b changes from open to closed and the switch does not move during period t; when When represents the switch of branch b from closed to open and no action during period t respectively; is the current of the bth branch of the distribution network during period t; is the active power output of the substation connected to node i during period t; R b is the resistance of the bth branch of the distribution network; is the output of the sth gas source in the kth gas network during time period t.
[0037] Constraints include distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints.
[0038] Specifically, 1. Distribution network constraints include: distribution network flow constraints; system security constraints; network radial constraints.
[0039]
[0040]
[0041]
[0042]
[0043] Where b(i,j) is the branch b with the first node being i and the last node being j; b,t and Q b,t are the active and reactive powers flowing through branch b during period t respectively; and are the injected active and reactive powers of node i in period t, respectively; is the reactive power output of the substation connected to node i during period t; and are the active and reactive outputs of the DG connected to node i during period t; and are the active and reactive power outputs of the CHP connected to node i during period t; and are the active and reactive outputs of the GT connected to node i during period t; and are the active and reactive loads connected to node i during period t, respectively.
[0044]
[0045]
[0046]
[0047] Where U i,t is the voltage of node i during period t; and are the lowest and highest allowed voltages of node i respectively; γ b,t is a 0-1 variable, γ b,t=1 / 0 respectively indicate that branch b is connected and disconnected during period t; is the maximum allowable current of branch b; P i DG,max 、P i GT,max and P i CHP,max are the maximum active powers of DG, GT and CHP connected to node i; θ DG ,θ GT and θ CHP They are the power factor angles of DG, GT and CHP respectively.
[0048]
[0049] μ(i,j)+μ(j,i)=γ b,t (10)
[0050]
[0051]
[0052]
[0053] Where a is the total number of system nodes; μ is the node parent-child association matrix, μ(i,j)=1 means node i is the parent node of node j; a(L τ (j)) is a 0-1 variable, a(L τ (j))=1 / 0 respectively indicate that the jth switch of loop τ is connected and disconnected.
[0054] 2. Natural gas network constraints include: flow constraints; pipeline flow constraints; gas source output constraints; and node pressure constraints. The natural gas network involved in this embodiment is a distribution network with a relatively small coverage area, so there is no need to consider boosting stations.
[0055]
[0056] Where, is the gas load connected to node m during period t; and are the gas consumption of CHP unit c and GT unit g during period t respectively.
[0057]
[0058]
[0059] Where G k,mn,t is the pipeline flow between nodes mn during period t; p k,m,t and p k,n,t are the pressures of nodes m and n during period t; Jk,mn is the parameter of the pipeline between nodes mn, which is related to the pipeline diameter, etc. and are the minimum and maximum values of the pipeline mn flow rate respectively.
[0060]
[0061] Where, and The minimum and maximum values of the gas source output.
[0062]
[0063] Where, and are the minimum and maximum node pressures, respectively.
[0064] 3. Heating network constraints include: thermal power balance constraints; CHP and heat exchange station (HES) temperature constraints; heating network temperature constraints.
[0065]
[0066] Where, π CHP and Π HES for CHP and HES collections respectively; is the thermal power of CHP unit c during period t; is the thermal power of the hth HES in period t; and are the temperatures of the outlet pipe and return pipe of CHP unit c during period t; and are the outlet and return pipe temperatures of the hth HES in period t, respectively; and are the water flows of CHP unit c and the hth HES during period t, respectively; is the specific heat capacity of water.
[0067]
[0068] Where, and are the minimum and maximum values of the outlet water temperature of CHP unit c, respectively; and are the minimum and maximum return water temperatures of the hth HES, respectively.
[0069]
[0070] Where, π p- and Π p+are the pipe sets with o as the end point and starting point respectively; m k,l,t is the water flow in pipe l during period t; and are the outlet and inlet temperatures of the drainage pipe l during period t; and are the outlet and inlet temperatures of the return pipe l during period t; and are the mixed temperatures of node o on the outlet pipe and return pipe during period t respectively.
[0071] 4. Electric-gas-thermal coupling constraints include: electric-gas-thermal coupling constraints of CHP units; electric-gas coupling constraints of GT units.
[0072]
[0073]
[0074] Where, and are the conversion efficiencies of CHP unit c and GT unit g respectively; Q GV Natural gas has a high calorific value.
[0075] In this embodiment, the nonlinear equations in the distribution network's power flow constraints are linearized using the Big-M and second-order cone methods. To simplify the distribution network reconstruction process, leveraging the "closed-loop construction, open-loop operation" characteristics of the distribution network, a tie switch and multiple sectional switches are organized into a basic loop matrix (BLM). The BLM is compressed through static reconstruction, and the compressed matrix is denoted as BLM*.
[0076] In order to solve the problem that nonlinear equations in the original tidal current are difficult to solve, the variable and By using the big-M method, the formula is converted into a linear form:
[0077]
[0078] Wherein, M is a relatively large integer.
[0079] The quadratic equation in the formula is relaxed to the standard second-order cone form using the second-order cone method:
[0080]
[0081] Where ||·||2 represents the 2-norm.
[0082] Leveraging the "closed-loop construction, open-loop operation" characteristics of distribution networks, a tie switch and multiple section switches in the distribution network are considered a basic loop. The switches within these basic loops are numbered and arranged row-by-row to form a basic loop matrix (BLM). During reconstruction, the switches are encoded according to the basic loops. Reconstruction is completed by simply disconnecting a switch from each row of the BLM. In the BLM, there are certain switch combinations that result in high system operating costs. Static reconstruction can filter out these switch combinations, eliminating inferior reconstruction solutions from the BLM and thus compressing the DNR solution space. The resulting compressed matrix is denoted as BLM*.
[0083] In view of the quadratic function expression of pipeline flow constraints in the natural gas network, this embodiment uses an incremental piecewise linearization method to process it, specifically:
[0084]
[0085] Where, Ω ω is a 0-1 variable, indicating the position on the ωth segment interval; φ m It is a 0-1 variable.
[0086] make and And introduce auxiliary variables and Then formula (15) can be transformed into:
[0087]
[0088] Where, is an interval variable; N is the number of segmented intervals; It is a 0-1 variable.
[0089] The DDQN model consists of two neural networks with identical structures, Qpre and Qtar. Qpre derives switching actions based on the current distribution network state and updates the network parameters after each training session. Qtar is used to calculate the target Q value, and after a specified number of training sessions, the network parameters of Qpre are copied to Qtar. This asynchronous network update method solves the overestimation problem in the original deep Q network (DQN) by decoupling the selection of switching actions from the calculation of the target Q value. During the above training process, the DDQN model continuously accesses data such as the distribution network state and switching actions. Therefore, a database capable of carrying this information is required for the model to use. This database includes parameters such as St describing the distribution network state, the switch action combination At, and the reward value Rt.
[0090] Assume that the data of period t is selected from the database as the current state, the current Q value is calculated based on the state St+1 and action At+1 of period t+1, and the current action is selected according to the greedy principle:
[0091]
[0092] Where θ t is the network parameter of Qpre; χ is a random number between 0 and 1; ε is the greedy probability, and ε∈[0,1]. Thus, the Q network has a certain probability of selecting a random action, which can increase the diversity of search actions and prevent the model from falling into local optimality.
[0093] According to the current action A(S t+1 θ t ) Calculate the current target Q value:
[0094]
[0095] Where γ is the attenuation factor and γ∈[0,1]; It is the network parameter of Qtar.
[0096] Then use Q pre (S t ,A t θ t )and Calculate the loss value:
[0097]
[0098] Where M is the number of batch samples during training.
[0099] For the KRR algorithm, KRR introduces the kernel method into ridge regression, enabling it to establish a nonlinear relationship between the distribution network state and the output of the coupled units, and thus deal with such nonlinear problems
[19] .
[0100] Assume the linear model is:
[0101] g(C)=C T X (31)
[0102] Where C T =(c1,c2...c n ) is the parameter matrix; T is the matrix transpose; X=(x1,x2…x n ) is the distribution network state matrix, including node loads and DG output.
[0103] Define G=(g1,g2,…,g n ) is the coupled unit output matrix, then the loss function is:
[0104]
[0105] Assume that the estimated value of the model is:
[0106] C1=argminloss KRR (C) (33)
[0107] In order to solve the matrix irreversibility problem caused by insufficient samples, the regularization method is introduced:
[0108] argmin[loss KRR (C)+υP(C)] (34)
[0109] Where υ is the hyperparameter matrix; P(C) is the regularization function.
[0110] Solve the above equations to get the estimated value:
[0111] C1=(X T X+υ -1 E) -1 X T G (35)
[0112] Where the estimated value C1 is the sum of a semi-positive definite matrix and a diagonal matrix, so C1 is positive definite and reversible; E is the identity matrix.
[0113] Introducing the kernel space, the optimal solution of the parameter matrix C1 is converted into the inner product form, and simplified according to the matrix inversion lemma:
[0114]
[0115] Where ξ is the weight parameter matrix.
[0116] In step 2 of this embodiment, based on the above, the DDQN model calculates and maximizes the Q value based on St, At, and Rt in the database to complete the training and optimization process. However, at this time, the model can only output a single dimension. In order to meet the requirement of the reconstruction model to output a single switch state containing multiple basic loops, it is necessary to increase the output dimension of the DDQN.
[0117] Among them, such as Figure 1 As shown in the figure, the improvements to the DDQN model are as follows: a fully connected neural network corresponding one-to-one to the elements in the BLM is added after the output of the DDQN model to achieve multi-dimensional output of the DDQN model; the one-hot principle is introduced to encode the model output results to ensure that only one switch is disconnected in each basic loop.
[0118] In this embodiment, the process of establishing the DDQN model database is as follows:
[0119] 1) Static Reconfiguration. The static reconfiguration model for the distribution network, which takes into account the synergy of electricity, gas, and heat, uses historical information such as node loads and DG output to perform static reconfiguration, thereby obtaining the distribution network reconfiguration solution and coupled unit output data. The static reconfiguration model does not consider the switching operation cost of the distribution network. The rest of the model is the same as the dynamic reconfiguration model, so its objective function is:
[0120]
[0121] 2) Establish a database. Store the distribution network status, switch actions, and return values in the database. The return value is calculated using formula (38).
[0122] Among them, the return value Rt is the basis for updating and calculating the Q value. Therefore, in the distribution network reconstruction problem, R t It can be set to the inverse of the objective function, that is:
[0123]
[0124] Where C t is the objective function of the reconstruction model.
[0125] like Figure 2 As shown in the figure, the dynamic reconstruction model of the distribution network taking into account the coordination of electricity, gas and heat adopts the "offline training, online application" method to obtain the reconstruction strategy.
[0126] The offline training steps are as follows:
[0127] 1) Generate a database. Generate a database through static reconstruction based on historical DNR data and distribution network structure parameters.
[0128] 2) Train the KRR model. Calculate and update the model parameters using the formula, and save the trained model for use in online applications.
[0129] 3) Initialize the training state. Clear the model input state, action, and reward variables to prepare for model training.
[0130] 4) Calculate Q pre Randomly select H data from the database, calculate the Qpre value through the formula and decide the action based on the greedy principle.
[0131] 5) Calculate Q tar According to the action obtained in step 4), Q is calculated by the formula tar value.
[0132] 6) Calculate the loss value and update Q pre The obtained Q pre With Q tarSubstitute the value into the formula, calculate the loss value, and use the stochastic gradient descent method to update Q pre Network parameters.
[0133] 7) Determine and update Q tar Network. Determine whether the specified number of training times E has been reached. If so, copy the network parameters of Qpre at this time to Q tar Otherwise, it will not be copied.
[0134] 8) Output the model. When the maximum number of training cycles is reached, save the DDQN model for use in online applications. Otherwise, return to step 3).
[0135] The online application steps are as follows:
[0136] 1) Predict the output of coupled units. The output of coupled units is predicted using the KRR model based on the node load and DG output.
[0137] 2) Obtain a dynamic reconstruction strategy. Based on the node load, DG output, and the coupled unit output data in step 1), the reconstruction strategy is quickly obtained through the DDQN model.
[0138] This embodiment takes a 33-node distribution network, a 20-node gas distribution network, a 16-node heat distribution network (E33-G20-H16) and an E78-G40-H32 system as examples to verify the effectiveness of the proposed model. MATLAB is used to call the CPLEX solver for static reconstruction to establish a database; the DDQN model is trained using PYCHARM software, with the attenuation factor set to 0.99, the learning rate set to 0.001, the database capacity set to 50,000, the number of training rounds set to 5,000, the number of batches set to 128, and the greedy probability set to 0.8; the KRR model is trained using JUPYTERNOTEBOOK software, with the kernel function set to Sigmoid and the learning rate set to 0.001. The computer system used for the simulation is Win10; the CPU is an Intel Core i5-11300H with a base frequency of 3.10GHz; and the memory is 16GB.
[0139] The system structure diagram of the E33-G20-H16 example is as follows Figure 7As shown. The reference for the parameter settings of the gas distribution network and the heat distribution network is "Chen Xi. Research on the Optimal Scheduling Method of the Integrated Energy System Considering the Dynamic Characteristics of the Network [D]. Jinan: Shandong University, 2020.". The distribution network load and DG output are from the literature "ZHANG L, WANG G, GIANNAKIS G B. Real-Time Power System State Estimation and Forecasting via Deep Unrolled Neural Networks [J]. IEEE Transactions on Signal Processing. 2019, 67(15): 4069-4077.", and the gas load and heat load data are from the literature "Zhang Yumin, Zhang Xuan, Ji Xingquan, et al. Transmission and Distribution Coordinated Unit Combination Considering the Dynamic Characteristics of Electricity, Gas and Heat IES [J / OL]. Proceedings of the CSEE, 2022 [2022-10-05].". There are four DGs in the grid, two photovoltaic and two wind turbines, with their data shown in Table C1. Two GTs are installed in the gas grid to achieve electricity-gas coupling; two CHPs are installed in the heat grid to achieve electricity-gas-heat coupling. The data of the coupled units are shown in Table C2.
[0140] Table C1E33-G20-H16 Distributed Power Supply Data
[0141]
[0142] Table C2E33-G20-H16 coupled unit data
[0143]
[0144] In order to solve the problem of difficulty in obtaining the output of coupled equipment GT and CHP units in online applications, the KRR method that can effectively handle nonlinear problems is proposed to predict the output of coupled units. In order to simulate the measurement error existing in real situations, a bimodal Gaussian noise with a mean of 0, covariance of 10-3I and 10-7I, and weights of 0.65 and 0.35 is introduced. The prediction results of each coupled unit are shown in the following figure. Figure 3 As shown, the true value is the coupled unit output data obtained by static reconstruction.
[0145] Depend on Figure 3 As can be seen, the predicted and true curves are not significantly different. Even with the complex nonlinear mapping relationships between node load, DG output, and the coupled GT and CHP outputs, the KRR model can still accurately predict the outputs of the coupled GT and CHP units. This is because the KRR proposed in this example uses a kernel method to propagate low-dimensional data to high dimensions and perform nonlinear fitting in this high-dimensional space, giving it a powerful ability to solve nonlinear problems.
[0146] To reduce the optimization range of the DDQN-KRR algorithm, this embodiment eliminates poor-quality solutions from the original BLM through static reconstruction, thereby compressing the DNR solution space. BLM and BLM* are shown in Table 1 and Figure 4, respectively.
[0147] Table 1 IEEE 33-bus system BLM and BLM*
[0148]
[0149]
[0150] From Table 1 and Figure 4 As can be seen, the number of solutions in loop I has been reduced from 9 to 4; the number of solutions in loop II has been reduced from 7 to 2; the number of solutions in loop III has been reduced from 14 to 3; the number of solutions in loop IV has been reduced from 21 to 3; and the number of solutions in loop V has been reduced from 10 to 4. This shows that static reconstruction can eliminate poor-quality reconfiguration solutions, such as switch combinations that result in high system operating costs and low user voltage levels in the BLM. This significantly compresses the DNR solution space, allowing DDQN-KRR to find the optimal solution in a smaller solution space, improving the efficiency of the model solution.
[0151] To analyze the impact of the IES system on distribution network reconfiguration, this embodiment sets the following four scenarios:
[0152] Scenario 1: The power grid operates alone;
[0153] Scenario 2: Electricity-gas coordinated operation;
[0154] Scenario 3: Electricity-heat coordinated operation;
[0155] Scenario 4: Coordinated operation of electricity, gas and heat.
[0156] The distribution network line loss, total operating cost and minimum node voltage of the four scenarios are as follows: Figure 5 and as shown in Table 2.
[0157] Table 2 System operation results for four scenarios
[0158]
[0159] Depend on Figure 5 (a) It can be seen that the total operating cost of scenario 4 is the lowest in the 24 time periods, which shows that the synergistic effect of electricity-gas-heat multi-energy flow can give the system higher flexibility, thereby better balancing the relationship between line loss cost, electricity purchase cost and gas source cost, and achieving the goal of reducing the total operating cost. Figure 5(b) As can be seen, the minimum node voltage in the distribution network with electricity-gas-heat IES coordination has been improved to varying degrees. Although the minimum node voltage in scenario three is higher than that in scenario one over the 24-hour period, it remains below 0.95 pu during the 8-hour to 22-hour period, failing to meet user voltage requirements. The minimum node voltages in scenarios two and four are both above 0.95 pu, and scenario four's voltage is higher than that of the other three scenarios. This demonstrates that electricity-gas-heat multi-energy flow coordination has a greater advantage in improving node voltage levels than electricity-gas or electricity-heat coordination.
[0160] Table 2 shows that, after considering multi-energy flow coupling, the grid connection of GT and CHP units shortens the distance of electricity transmission and significantly reduces distribution network line losses. Scenario 1 has the highest operating cost, at 9732.92 USD, while the other three scenarios that consider multi-energy coupling have lower operating costs. Scenario 2 considers only electricity-gas coupling. GT units offer greater flexibility than CHP units, enabling them to provide power according to distribution network demand. Therefore, the minimum node voltage in Scenario 2 is higher than in Scenario 3. Scenario 3 considers only electricity-heat coupling, but the heat load demand causes the CHP output to be non-zero. Due to the CHP unit's "heat-based electricity generation and co-production" nature, this energy is not connected to the grid and is "wasted," ultimately resulting in higher costs for Scenario 2 than for Scenario 3. Scenario 4 combines the advantages of Scenarios 2 and 3, reducing total costs by 5.06% while increasing the minimum node voltage to 0.9586 pu. This shows that electricity-gas-heat synergy further enhances the coupling advantages of multiple energy sources, reducing total system operating costs while increasing system voltage levels.
[0161] In order to verify the performance advantage of DDQN-KRR in DNR, it is compared with static reconstruction, MISOCP dynamic reconstruction in the literature "Li Chao, Miao Shihong, Sheng Wanxing, et al. Active distribution network optimization operation strategy considering dynamic network reconstruction [J]. Transactions of the Chinese Society of Electrical Engineering, 2019, 34(18): 3909-3919." and LSTM dynamic reconstruction in the literature "JI X, YIN Z, ZHANG Y, et al. Real-time autonomous dynamic reconfiguration based on deep learning algorithm for distribution network [J]. Electric Power Systems Research, 2021, 195(03): 107132.", and the non-reconstructed network is used as the control group. Among them, the output of the coupled units in the LSTM dynamic reconstruction is also predicted using the KRR method; non-reconstruction means keeping the interconnecting switches S33, S34, S35, S36 and S37 in the IEEE33 system always in the disconnected state. The system operation results of different algorithms are shown in Table 3.
[0162] Table 3 System operation results of four scenarios
[0163]
[0164]
[0165] Table 3 shows that the unreconfigured group relies solely on the synergy of electricity, gas, and heat to reduce total system operating costs. Consequently, its coupled unit output to substation power purchase ratio is as high as 1.0152:1, allowing the system's lowest node voltage to reach 0.95 pu. However, due to topological limitations, the higher coupled unit output fails to reduce distribution network operating costs, ultimately resulting in higher total operating costs for the unreconfigured group. This demonstrates the need to consider both the IES system and the DNR. Static reconfiguration effectively balances coupled unit output and substation power purchase, but frequent switching operations result in a total operating cost that is 0.151% higher than the unreconfigured group. MISOCP dynamic reconfiguration, which takes switching costs into account, has a lower total operating cost than static reconfiguration. However, because the coupled unit output to substation power purchase ratio is only 0.4692:1, it results in higher line losses. Both the LSTM and DDQN-KRR methods rely on KRR to obtain the output of coupled units. Therefore, the ratio of coupled unit output to substation power purchase for both methods is similar to that of static reconstruction. However, the total operating cost of DDQN-KRR is as low as 9240 USD, and the minimum node voltage is increased to 0.9586 pu. This shows that the method proposed in this embodiment can more deeply mine the characteristics of historical data, determine a better reconstruction strategy, and improve the voltage level of distribution network nodes.
[0166] To verify the computational efficiency advantage of the DDQN-KRR method, we compared it with static reconstruction, MISOCP dynamic reconstruction, and LSTM dynamic reconstruction. The comparison results of the four algorithms are shown in Table 4.
[0167] Table 4 Training and running time of four algorithms
[0168]
[0169] As shown in Table 4, the average time it takes for the DDQN-KRR method proposed in this embodiment to obtain a reconstruction strategy is only 0.0131 seconds, which is significantly lower than the 0.0657 seconds of the LSTM method, the 3.23 seconds of the static reconstruction method, and the 342.64 seconds of the MISOCP dynamic reconstruction method. This shows that the DDQN-KRR method proposed in this embodiment has the highest computational efficiency. Both the DDQN-KRR and LSTM methods can achieve rapid reconstruction in milliseconds. However, the LSTM method models each basic loop, and the more basic loops there are, the longer it takes to obtain a reconstruction solution. In contrast, the DDQN-KRR method proposed in this embodiment integrates all loops into a single model, allowing it to directly obtain a reconstruction strategy for the entire system, significantly reducing reconstruction time.
[0170] In addition to reducing the total operating cost of the system, the coordinated operation of electricity, gas and heat can also achieve "peak shaving and valley filling" to a certain extent. Figure 6 shown.
[0171] Depend on Figure 6 As can be seen, during the 8-15 hour period, DG output is high, while GT and CHP output is relatively low. This indicates that during electricity-gas-heat coordinated operation, the output of the coupled GT and CHP units can be reduced during peak DG output periods, thereby reducing gas consumption. This effectively converts electricity into natural gas and stores it, providing ample capacity to accommodate DG output. During the 2-7 hour period, the load is low, and the coupled units proactively reduce their output to facilitate the grid's absorption of DG. During the 16-24 hour period, DG output is low, while GT and CHP output is high. This indicates that electricity-gas-heat coordinated operation can convert natural gas and heat into electricity to meet load demand during periods of low DG output. Therefore, IES coordination can, to a certain extent, smooth the power fluctuations caused by DG grid integration, achieving a "peak shaving and valley filling" effect.
[0172] E78-G40-H32 Example: To further validate the effectiveness and practicality of the proposed DDQN-KRR model and method, we analyze the E78-G40-H32 system, which consists of a 78-node power distribution system, two 20-node gas distribution networks, and four 8-node heat distribution networks. The E78 parameters are shown in Table C3. The system contains eight DGs, as detailed in Table C4; four GTs, and four CHPs, as detailed in Table C5.
[0173] Table C3 Network parameters of a city's improved 78-node distribution system
[0174]
[0175]
[0176]
[0177] Table C4 E78-G40-H32 system distributed power generation equipment data
[0178]
[0179] Table C5 E78-G40-H32 system coupling unit data
[0180]
[0181] To verify the superior performance of the DDQN-KRR method proposed in this example in large-scale power distribution systems, we compared it with methods such as no reconfiguration, static reconfiguration, MISOCP dynamic reconfiguration, and LSTM dynamic reconfiguration. No reconfiguration refers to the fact that S78, S79, S80, S81, and S82 are always disconnected. The results of the five algorithms are shown in Table 5.
[0182] Table 5 Running results of five algorithms
[0183]
[0184] Table 5 shows that the unreconfigured group, due to the inability to change the topology, has a higher ratio of coupled unit output to substation power purchase and higher line losses, resulting in a total operating cost of 17,428.20 USD. The static reconfiguration, due to frequent switch operations, results in a 0.115% higher total operating cost than the unreconfigured group. MISOCP dynamic reconfiguration, which accounts for switch operations, reduces total costs by 0.610%. Line losses are minimized due to the large ratio of coupled unit output to substation power purchase. This is because, after accounting for switch operations, the system prioritizes increasing coupled unit output over changing the topology to reduce total operating costs. The coupled unit output to substation power purchase ratios for the LSTM and DDQN-KRR methods are essentially the same as for the static reconfiguration, resulting in similar line losses for the three methods. The DDQN-KRR method achieves a 0.756% reduction in total cost, outperforming the other four methods. This demonstrates the applicability of the DDQN-KRR method proposed in this example to large-scale IES systems. The lowest node voltage of all methods is greater than 0.95pu. This is because the large number of coupled units put into operation enables the IES system to fully coordinate and operate, which improves the node voltage mean.
[0185] To verify the computational efficiency advantage of the DDQN-KRR method proposed in this example in large-scale power distribution systems, it is compared with static reconstruction, MISOCP dynamic reconstruction, and LSTM dynamic reconstruction. The comparison results of the four algorithms are shown in Table 6.
[0186] Table 6 Training and running time of four algorithms
[0187]
[0188] Table 6 shows that the DDQN-KRR algorithm has a more significant advantage in computational efficiency. The average reconstruction times for static and MISOCP dynamic reconstructions significantly increased to 7.44s and 516.35s, respectively. However, the average reconstruction times for the DDQN-KRR and LSTM methods were 0.0193s and 0.0947s, respectively, remaining within milliseconds and achieving rapid reconstruction. In summary, the DDQN-KRR method maintains good practicality in large-scale IES systems.
[0189] This example addresses the problem of low computational efficiency in dynamic reconfiguration and proposes a distribution network dynamic reconfiguration model based on DDQN-KRR for electricity, gas, and heat IES coordination. This is verified by using the E33-G20-H16 and E78-G40-H32 examples, yielding the following:
[0190] 1) The KRR method proposed in this embodiment can accurately predict the output of the coupled equipment GT and the CHP unit, solving the problem of difficulty in obtaining the output of the coupled unit in online applications.
[0191] 2) Taking into account the DNR of the electricity-gas-heat IES synergy, the system can leverage the resource advantages of different energy types based on their characteristics and complement each other. This can, to a certain extent, smooth out the power fluctuations caused by DG grid connection, playing the role of "peak shaving and valley filling", thereby reducing the total operating cost of the system and improving the voltage level of users.
[0192] 3) Compared with other reconstruction algorithms, the DDQN-KRR algorithm proposed in this embodiment takes into account the economy and computational efficiency of the reconstruction strategy and realizes dynamic and fast reconstruction, and has good applicability in large-scale electric-gas-heat IES systems.
[0193] Example 2
[0194] The purpose of this embodiment is to provide a distribution network dynamic reconstruction system based on DDQN-KRR electrical and thermal coordination, including:
[0195] Model building module: Taking the minimization of the sum of the distribution network operating cost and the gas source cost of the gas distribution network as the optimization objective function, and using the distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints as constraints, a distribution network dynamic reconstruction model is constructed;
[0196] Model solving module: Uses DDQN-KRR to solve the distribution network dynamic reconstruction model. The KRR model predicts the output of coupled units based on node loads and DG outputs. The DDQN model is used to obtain the reconstruction strategy based on the node loads, DG outputs and the output data of the coupled units.
[0197] Example 3
[0198] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the program.
[0199] Example 4
[0200] The purpose of this embodiment is to provide a computer-readable storage medium.
[0201] A computer-readable storage medium stores a computer program, which, when executed by a processor, performs the steps of the above method.
[0202] The steps involved in the apparatuses of Examples 2, 3, and 4 above correspond to those of Method Example 1. For detailed implementations, please refer to the relevant description of Example 1. The term "computer-readable storage medium" should be understood to mean a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and causing the processor to perform any method of the present invention.
[0203] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.
[0204] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A dynamic distribution network reconstruction method based on DDQN-KRR electrical and thermal synergy, characterized by: include: Taking the minimization of the sum of the distribution network operation cost and the gas source cost of the gas distribution network as the optimization objective function, and the distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints as the constraints, a distribution network dynamic reconstruction model is constructed. The dynamic reconstruction model of the distribution network is solved using DDQN-KRR. The KRR model is used to predict the output of the coupled units based on the node load and DG output. The DDQN model is used to obtain the reconstruction strategy based on the node load, DG output and the coupled unit output data. When solving the distribution network dynamic reconstruction model using DDQN-KRR, the offline training of DDQN-KRR specifically includes: Step 1: Generate a database through static reconstruction based on historical DNR data and distribution network structure parameters; Step 2: Train the KRR model and save the trained KRR model; Step 3: Initialize the input state, action, and reward value of the reconstruction model; Step 4: Randomly select H data from the database and calculate the DDQN model Q pre value and make decisions based on the greedy principle; Step 5: Calculate the action obtained in step 4 in the DDQN model Q tar value; Step 6: Calculated based on steps 4 and 5 Q pre Value and Q tar Calculate the loss value and update it using stochastic gradient descent method Q pre Network parameters; Step 7: Determine whether the number of training times has been reached. If it has been reached, Q pre The network parameters are copied to Q tar Otherwise, it will not be copied; the trained DDQN model is obtained; The improvements to the DDQN model are as follows: a neural network corresponding to the BLM is added after the DDQN model output to achieve multi-dimensional output of the DDQN model, and the one-hot encoding principle is introduced to ensure that only one switch is disconnected in each basic loop.
2. The method for dynamic reconstruction of distribution network based on electrical-thermal synergy of DDQN-KRR according to claim 1, characterized in that: The objective function is: in, and are the distribution network operation cost and gas source cost of the gas distribution network during period t respectively; 、 、 and They are distribution network line loss, distribution network power purchase, distribution network switch action and gas source output cost coefficient; B, , K and S are the distribution network branch set, distribution network node set, gas distribution network set and gas source set respectively; and is a 0-1 variable; is the current of the bth branch of the distribution network during period t; is the active power output of the substation connected to node i during period t; is the resistance of the bth branch of the distribution network; is the output of the sth gas source in the kth gas network during time period t.
3. The method for dynamic reconstruction of distribution network based on electrical-thermal synergy of DDQN-KRR according to claim 1, characterized in that: The distribution network constraints include: distribution network flow constraints; system safety constraints; network radial constraints; the gas distribution network constraints include: flow constraints; pipeline flow constraints; gas source output constraints; node pressure constraints; the heat distribution network constraints include: thermal power balance constraints; CHP and heat exchange station temperature constraints; heat network temperature constraints; the electric-gas-heat coupling constraints include: CHP unit electric-gas-heat coupling constraints; GT unit electric-gas coupling constraints.
4. The method for dynamic reconstruction of distribution network based on electrical-thermal synergy of DDQN-KRR according to claim 2, characterized in that: It also includes linearizing the nonlinear equations in the flow constraints using big-M and second-order cone methods; using the incremental piecewise linearization method to process the quadratic functions in the pipeline flow constraints in the gas distribution network constraints; and simplifying the distribution network reconstruction process as follows: compiling a connecting switch and multiple section switches into a basic loop matrix, and compressing the BLM through static reconstruction.
5. The method for dynamic reconstruction of distribution network based on electrical-thermal synergy of DDQN-KRR according to claim 1, characterized in that: When using DDQN-KRR to solve the distribution network dynamic reconstruction model, the establishment of the database includes: Based on the distribution network static reconstruction model taking into account the synergy of electricity, gas and heat, static reconstruction is performed using node loads and DG output historical data to obtain the distribution network reconstruction solution and coupled unit output data; The distribution network status, switch actions and reward values are stored in a database; the reward value is the inverse of the objective function of the reconstruction model.
6. The distribution network dynamic reconstruction system based on DDQN-KRR electrical and thermal coordination is characterized by: include: Model building module: Taking the minimization of the sum of the distribution network operating cost and the gas source cost of the gas distribution network as the optimization objective function, and using the distribution network constraints, gas distribution network constraints, heat distribution network constraints, and electricity-gas-heat coupling constraints as constraints, a distribution network dynamic reconstruction model is constructed; Model solving module: DDQN-KRR is used to solve the distribution network dynamic reconstruction model. The KRR model is used to predict the output of coupled units based on node loads and DG outputs. The DDQN model is used to obtain a reconstruction strategy based on the node loads, DG outputs, and the output data of the coupled units. In solving the distribution network dynamic reconstruction model using DDQN-KRR, offline training of DDQN-KRR specifically includes the following: Step 1: Generate a database through static reconstruction based on historical DNR data and distribution network structure parameters; Step 2: Train the KRR model and save the trained KRR model; Step 3: Initialize the input state, action, and reward value of the reconstruction model; Step 4: Randomly select H data from the database and calculate the DDQN model Q pre value and make decisions based on the greedy principle; Step 5: Calculate the action obtained in step 4 in the DDQN model Q tar value; Step 6: Calculated based on steps 4 and 5 Q pre Value and Q tar Calculate the loss value and update it using stochastic gradient descent method Q pre Network parameters; Step 7: Determine whether the number of training times has been reached. If it has been reached, Q pre The network parameters are copied to Q tar Otherwise, it will not be copied; the trained DDQN model is obtained; The improvements to the DDQN model are as follows: a neural network corresponding to the BLM is added after the DDQN model output to achieve multi-dimensional output of the DDQN model, and the one-hot encoding principle is introduced to ensure that only one switch is disconnected in each basic loop.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for dynamic reconstruction of a distribution network based on electrical-thermal coordination of DDQN-KRR are implemented as described in any one of claims 1 to 5.
8. A processing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method for dynamic reconstruction of a distribution network based on electrical-thermal coordination of DDQN-KRR are implemented as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Micro-power-grid energy storage scheduling method and device based on deep Q-value network (DQN) reinforcement learning
CN109347149A
Distribution network real-time dynamic reconstruction method and system based on branch dual deep Q network
CN114282330A