Power distribution network dynamic planning investment decision-making method and system based on deep double-Q network
By constructing a state-space vector based on a deep dual-Q network and performing iterative training, the problems of dynamic uncertainty and computational complexity in distribution network investment decisions are solved, and real-time response and globally optimal investment strategy output are achieved.
Patent Information
- Application Number
- CN202511631671.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2025-12-05
AI Technical Summary
Existing power distribution network investment decision-making methods cannot respond to dynamic uncertainties in a timely manner, lack robustness, have high computational complexity, and are difficult to achieve global optimization.
We employ a deep dual-Q network-based approach, constructing a state space vector and performing iterative training. By combining experience replay and parameter updates, we output a periodically optimal investment strategy.
It enables real-time response to dynamic uncertainties such as load growth and new energy access, improves the accuracy and robustness of investment decisions, and ensures the reliability and sustainability of planning results.
Smart Images

Figure CN121073262A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power system optimization, more particularly, the present application relates to a power distribution network dynamic planning investment decision method and system based on deep double Q network. BACKGROUND
[0002] In the construction and operation of power distribution networks, investment decisions face constantly changing uncertainties such as load demand, energy structure, and environmental policies. In order to ensure the effectiveness and long-term benefits of investment decisions, construction units need to accurately plan the investment of power distribution networks and fully consider these dynamic changes. However, current power distribution network investment decision methods still have many shortcomings.
[0003] Traditional power distribution network investment decision methods are usually based on static mathematical models or heuristic rules for planning, and often cannot respond to dynamic uncertainties such as load growth and new energy access in a timely manner. These methods lack real-time response to actual operating data, and most methods require pre-set fixed scenarios, resulting in poor robustness of the planning results and difficulty in adapting to complex future changes. In addition, traditional methods also have the problem of insufficient quantification of long-term technical-economic comprehensive benefits, especially when facing complex technical constraints and economic optimization, it is difficult to achieve global optimization.
[0004] In the prior art, some methods make decisions by manually setting parameters or based on historical experience, but these methods rely on human experience and lack quantitative and scientific basis, and are prone to local optimization or unreasonable decisions. At the same time, although the multi-stage stochastic programming method can handle uncertainty, its computational complexity increases exponentially with the increase in the number of scenarios, especially when facing long-period planning, the calculation time is too long and the efficiency is low. SUMMARY
[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present application provide a power distribution network dynamic planning investment decision method based on deep double Q network, which constructs a state space vector by normalizing the historical operating data of the power distribution network, and iteratively trains it in combination with deep double Q network, to solve the problems of existing power distribution network investment decision methods that cannot respond to dynamic uncertainty, lack of robustness, and high computational complexity.
[0006] To achieve the above-mentioned purposes, the present application provides the following technical solutions: The power distribution network dynamic planning investment decision method based on deep double Q network comprises the following steps: collecting and normalizing the historical operating data of the regional power distribution network; constructing a state space vector based on the normalized data; inputting the vector into the deep double Q network for iterative training, and realizing model convergence through experience replay and parameter update; outputting a periodic optimal investment strategy sequence based on the converged model.
[0007] In a preferred embodiment, the state space vector is constructed based on the normalized data, specifically: a covariance matrix of the data is calculated; an eigenvalue diagonal matrix is obtained by performing eigenvalue decomposition on the covariance matrix; a contribution rate of each eigenvalue is calculated, and a data state space vector corresponding to an eigenvalue with a cumulative contribution rate not lower than a preset threshold is selected.
[0008] In a preferred embodiment, the iterative training includes filtering invalid power distribution network planning investment actions by an action mask module.
[0009] In a preferred embodiment, the action mask module specifically: investment actions violating constraints are screened by calculating physical constraints of the power distribution network in real time; and the Q value of the investment actions is set to be negative or zero.
[0010] In a preferred embodiment, the iterative training includes designing a composite reward function, which is composed of an economic cost term, a technical reward term, and a constraint penalty term.
[0011] In a preferred embodiment, the economic cost term is calculated based on the discounted value of the initial investment cost and the operation and maintenance cost of the equipment, and the discounted value is determined according to the investment cycle and a preset discount rate; the technical reward term includes a reliability improvement reward, a line loss reduction reward, and a voltage qualification reward, and the reliability improvement reward includes adjusting an adaptive coefficient according to the planning stage; and the constraint penalty term is determined according to whether the power supply reliability index and the node voltage deviation exceed a preset threshold.
[0012] In a preferred embodiment, the experience replay includes setting a sample priority, specifically: a transition sample consisting of a current state, a selected investment action, a corresponding composite reward, a next state after the action is performed, and a termination flag is stored in an experience replay pool; the priority of each transition sample in the experience replay pool is determined according to the absolute size of the time difference error-based error, and a minimum constant is introduced in the calculation; the transition sample newly added to the experience replay pool is assigned the highest priority in initialization; and the probability of a sample being extracted for training is determined according to the relative size of the priority of each sample in the replay pool.
[0013] In a preferred embodiment, the determination of the probability of a sample being extracted for training further includes calculating an importance sampling weight based on the sample extraction probability.
[0014] In a preferred embodiment, the iterative training further includes using a greedy strategy to select investment actions.
[0015] This invention provides a dynamic planning investment decision-making system for distribution networks based on deep dual-Q networks, comprising: a data acquisition module for collecting and normalizing historical operation data of the regional distribution network; a data processing module for constructing a state space vector based on the normalized data; a model training module for inputting the vector into the deep dual-Q network for iterative training, achieving model convergence through experience playback and parameter updates; and a strategy output module for outputting a periodic optimal investment strategy sequence based on the converged model.
[0016] The technical effects and advantages of the multi-source load response intelligent distribution network hierarchical collaborative control method of this invention are as follows: This invention constructs a state-space vector by normalizing historical data of the distribution network operation and uses a deep dual-Q network for iterative training, achieving real-time response and optimization in the distribution network investment decision-making process. This method effectively addresses dynamic uncertainties such as load growth and renewable energy integration, avoiding the limitations of traditional methods, such as inability to adapt to dynamic changes, poor robustness, and high computational complexity. This improves the accuracy and robustness of investment decisions, ensuring the reliability and sustainability of planning results. Attached Figure Description
[0017] Figure 1 A schematic diagram of the dynamic planning investment decision-making method for distribution networks based on deep dual-Q networks provided in an embodiment of the present invention; Figure 2 This is a block diagram of the dynamic planning investment decision-making system for distribution networks based on deep dual-Q networks, provided in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] Example 1, Figure 1 This invention presents a dynamic planning investment decision-making method for distribution networks based on deep dual-Q networks, comprising the following steps: Step S1: Obtain historical data of regional power distribution network operation, normalize the data, and construct an 8-dimensional state space vector containing technical indicators, equipment status, economic parameters, and environmental factors. Step S2: Based on the deep dual-Q network, create an online network and a target network for distribution network analysis, and initialize the target network parameters to the online network parameters; Step S3, define a discretized set of distribution network planning investment actions, based on -Greedy strategy selects investment actions; Step S4: Execute the action and calculate the technical indicators, and design a composite reward function that integrates technical indicator rewards and economic cost penalties; Step S5: Construct a distribution network planning quadruple set, store it in the experience playback pool, and set sample priority using a priority playback mechanism; Step S6: Sample from the experience replay pool, train the Q network, and output the optimal investment strategy sequence.
[0020] In this embodiment, step S1 includes the following steps: Step S1.1: Obtain historical data on the operation of the regional power distribution network, including: technical indicators, equipment status, economic parameters and environmental factors within the power system, and extract dynamic features and spatial correlation features through calculation and analysis; Step S1.2: Standardize and normalize each indicator to eliminate dimensional differences; Step S1.3: PCA is used to analyze and compress redundant dimensions, compressing the original multidimensional state vector into an 8-dimensional state vector.
[0021] The above steps are as follows: Technical indicators were collected through the regional power company's distribution network operation systems (such as GIS and EMS systems), power dispatch center systems, and smart power terminal equipment. These indicators included: system average outage duration (SAIDI), system average outage frequency (SAIFI), comprehensive line loss rate (η), and node voltage deviation (ΔV); equipment status: distribution transformer load, switchgear aging coefficient, and line insulation degradation index; economic parameters: cumulative investment, unit outage loss cost, and maintenance cost ratio; and environmental factors: distributed power generation output and load forecasting error rate. The collected data underwent preprocessing.
[0022] Time alignment: Ensure that the data has the same time resolution (15 minutes); Missing value handling: interpolating or deleting missing data; Outlier Handling: Identify, correct, or remove outlier data points. By mining the dynamic evolution patterns and spatial correlation characteristics of the above data, more representative multidimensional state vectors can be generated. Specifically, the time-series rate of change characteristics can be expanded to include SAIDI trend, line loss rate acceleration, and investment benefit decay; spatial correlation characteristics can be expanded to include voltage deviation range, load rate spatial variance, net load fluctuation, and renewable energy penetration rate. All raw and extended data indicators are normalized and standardized to eliminate dimensional differences. Normalization mainly employs Min-Max standardization of the data indicators, expressed as: , in For data to be normalized, and These represent the minimum and maximum values in the data sample, respectively. The data normalization method used includes, but is not limited to, Min-Max standardization, exponential decay normalization, logarithmic function compression, and other normalization methods known to those skilled in the art. For all normalized data indicators, a multi-dimensional standardized state vector is generated, expressed as: , Principal Component Analysis (PCA) is used to compress the multidimensional state vector. First, the covariance matrix is calculated, expressed as: , Then, eigenvalue decomposition is performed, expressed as: , in, The eigenvalue diagonal matrix (in descending order, λ1≥...≥λ) n , λ n (where is the eigenvalue of the covariance matrix). This is the eigenvector matrix.
[0023] Finally, the contribution rate is calculated to obtain the principal components with a cumulative contribution rate ≥ 95%. The expression for calculating the contribution rate is as follows: , The calculation yielded eight principal components: PC1 (including transformer load and line loss rate (η)), PC2 (including SAIDI and SAIFI reliability indicators), PC3 (including voltage deviation (ΔV) and distributed generation output (DG)), PC4 (equipment aging rate (τ)), and PC5 (cumulative investment amount (Inv)). A state vector was constructed. , This embodiment obtains relevant data indicators through distribution network operation-related systems (such as GIS systems and EMS systems), power dispatch center systems, and intelligent power terminal equipment. By fusing dynamic and spatial features, the state vector can more comprehensively depict the distribution network operation status, providing a more accurate decision-making basis for deep dual-Q networks (DDQN). Through PCA dimensionality reduction, the training efficiency of DDQN is significantly improved while ensuring decision accuracy, making it particularly suitable for real-time rolling planning of distribution networks with massive nodes.
[0024] In this embodiment, step S2 includes the following steps: Step S2.1, construct DDQN, including creating the distribution network analysis online network and target network. The network structure includes an input layer, a hidden layer, and an output layer. Step S2.2: Initialize the target network parameters by setting the initial parameters of the target network to the online network parameters.
[0025] Specifically: The mesh is designed based on the dimensions of the state vector generated in step S1, where: Input layer: Number of neurons equals the dimension of the state vector ; Hidden layers: two fully connected layers, each with 256 neurons, using the ReLU activation function; First layer: , in, ; Second layer: , in, ; Output layer: The number of neurons equals the dimension of the action space, and it outputs the Q-value of each action, expressed as: , in, ; The online mathematical expression is: , , s is the distribution network state vector, a is the investment action, θ is the trainable parameters of the online network, and f is the distribution network state vector. online is the forward computation function of the online network, W1 is the weight of the feature extraction layer, W2 is the weight of the action value output layer, W3 is the weight of the output layer, b1 is the activation threshold of each neuron, b2 is the basic offset of the action value, and b3 is the bias vector of the output layer. The mathematical expression for the target network is: , , This embodiment effectively solves the overestimation problem in deep learning training by designing a dual network structure of online grid and target grid, and by separating action selection and target value calculation. It is the core design for stable convergence of dynamic programming in distribution networks.
[0026] In this embodiment, step S3 includes the following steps: Step S3.1: Define the discrete action space, including line switching, automation upgrades, substation capacity expansion, and maintaining the status quo; Step S3.2, policy definition, in terms of probability Randomly select an action (exploration) with a probability of 1- Select the action with the highest current Q value (utilize); Step S3.3: Generate uniformly distributed random numbers p ~ U(0,1), and compare p with the current exploration rate. This determines the type of strategy to execute; Step S3.4: Dynamically adjust the exploration rate; Step S3.5: Immediately output the selection result, i.e., the investment action; Step S3.6: Introduce an action mask module in the output layer to automatically filter invalid actions that violate physical rules such as voltage constraints and capacity limits.
[0027] Specifically: By selecting relevant investment actions from the problem-solving approach of distribution network planning, a discrete action space is defined. The expression is: , in, For line switching, For automation upgrades, To increase the capacity of the substation To maintain the status quo.
[0028] Policy definition, in terms of probability Randomly select an action (random exploration), with a probability of 1- Select the action with the maximum current Q value (exploitation), the expression is: , Generate uniformly distributed random numbers p ~ U(0,1), and compare p with the current exploration rate. This determines the type of execution strategy. If p≤ Perform random exploration, which starts from the action space. A uniformly random action is selected from the middle. If p > Execute the exploitation strategy and input the state. To online network Calculate the Q-value for all actions, expressed as: , Choose the optimal action. Thus, the state of the distribution network can be obtained. The optimal investment action is taken at the right time. When the distribution network condition changes, the exploration rate is dynamically adjusted. This allows for a re-screening of investment actions. Exploration rate The initial value is set to 1.0 for completely random exploration. The exploration rate is then dynamically adjusted using the following expression: , in, The initial exploration rate (e.g., 1.0). Minimum exploration rate (e.g., 0.01). The decay rate is 0.001, and t is the number of training steps.
[0029] By conducting The `-greedy` strategy outputs the selection result, i.e., the investment action. After the selection result is output, invalid actions that violate physical rules such as voltage constraints and capacity limits still need to be filtered out. Specifically: At the output layer, an action mask module is introduced. The function of this module is to automatically filter out invalid actions that violate physical rules such as voltage constraints and capacity limits.
[0030] Once an investment action is selected, the action mask module verifies it based on the physical constraints of the distribution network. The specific steps are as follows: 1) Calculate physical constraints: Call the power flow calculation model of the distribution network, calculate the new distribution network state (including physical parameters such as voltage, current, and power) based on the selected investment action, and check whether it meets the technical requirements such as voltage constraints and capacity constraints.
[0031] 2) Filtering invalid actions: If an investment action violates physical constraints (such as voltage over-limit or insufficient capacity), the action is considered invalid.
[0032] 3) Set the Q value to a negative or zero: For invalid actions that violate constraints, the action mask module sets their Q value to a negative or zero, thereby reducing the probability of these actions being selected. The specific formula is: , or: , This masking mechanism ensures that the selected action always conforms to physical constraints, thereby improving the feasibility and effectiveness of decision-making.
[0033] This embodiment applies... The -greedy strategy employs a dynamic decision-making process that balances exploration and utilization in selecting investment actions. The model learns from historical best decisions while also exploring potential better solutions, ultimately outputting investment actions that are both technically constrained and economically sound.
[0034] In this embodiment, step S4 includes the following steps: Step S4.1: Update the distribution network topology and equipment parameters according to the selected investment action; Step S4.2: Invoke power flow calculation to verify technical constraints, including power supply reliability, line loss, and voltage. Step S4.3: Set up a composite reward function, including economic costs, technological rewards, and constraint penalties; Step S4.4: Calculate the economic cost as a negative reward; Step S4.5: Set technical reward items as positive rewards, including reliability improvement rewards, line loss reduction rewards, and voltage qualification rewards; Step S4.6: Set the technical constraint violation penalty items as negative rewards, including power supply reliability and voltage deviation; Based on the investment actions selected in step S3, update the distribution network topology and equipment parameters. Verify the technical constraints using power flow calculations: whether the voltage deviation meets |ΔV| ≤ 7%, and whether the power supply reliability SAIDI ≤ SAIDI. 目标 Line loss rate η≤η 允许 Among them, SAIDI 目标 For target SAIDI, η 允许 To allow for line loss rate, a compound reward function is set based on changes in technical parameters and the economic cost of investment actions, expressed as: , in, For economic costs, Rewards for improving technical indicators Penalties for violating technical constraints. The economic cost is calculated as a negative reward; the expression for the current investment cost is: , in, The initial investment cost of the equipment. The maintenance cost in year k is... The discount rate is T, and the investment period is T. Technical incentives are set as positive rewards, including reliability improvement rewards. Reduced line loss reward Voltage qualification reward The specific expression is: , in, Value per unit of time; , in, For adjustment coefficients, Let t be the line loss rate; , in, Basic rewards; , , As an adaptive coefficient, it can be set based on the planned economic cost and actual operational experience. The weights need to be adjusted according to the planning phase, as expressed in the following expression: , Setting penalties for violating technical constraints as negative rewards includes power supply reliability and voltage deviation, with the specific expression as follows: , in, It is an exponential function (1 if violated, 0 otherwise). It is a large constant (set based on experience); This embodiment introduces a composite reward function to perform time value discounting, discounting operation and maintenance costs to their current value. It quantifies technological benefits, converting non-economic indicators (such as SAIDI) into monetary value. By quantitatively balancing technological and economic objectives, it guides DDQN to learn the optimal investment strategy.
[0035] In this embodiment, step S5 includes the following steps: Step S5.1: Select the 8-dimensional state vector obtained in step 1, the action space defined in step 3.1, and construct transition samples based on the composite reward after the action is executed, the new state vector after the action is executed, and the termination flag. Step S5.2: Store the transferred samples in the experience playback pool and set the sampling priority.
[0036] Using the 8-dimensional state vector obtained in step 1, the action space defined in step 3.1, and the state vector based on the composite reward after action execution, the new state vector after action execution, and the termination flag, a transition sample is constructed, expressed as: , The transferred samples are stored in the empirical replay pool. The priority of each sample is based on the absolute value of the TD error, expressed as: , in, =10 -5 To prevent zero priority, set it initially. (e.g., 1.0) As a discount factor, For the current moment's composite reward, It is used in the target network to evaluate the next state when calculating the TD error. Virtual action variables at time; Samples with large TD errors are assigned higher sampling probabilities, and the sampling probability expression is: , in, Control priority (usually set to 0.6); To prevent high TD error samples from being oversampled, which could cause gradient updates to deviate from the true expectation, importance weights are introduced. Reweighting is performed to maintain unbiasedness. If an investment action sample is frequently sampled due to high TD error, its gradient update will be... Lower the weight to avoid excessively influencing the strategy. The importance sampling weight is calculated using the following expression: , To control variance and stability, the weights are normalized. The weight range is limited to [0,1] to prevent extreme samples (such as...). hour This leads to gradient explosion. The normalization expression is: , in, from (e.g., 0.4) linearly increases to 1.0, where N is the total number of samples in the empirical replay pool. Initially ( (Smaller) Reduce weight differences, approximate uniform sampling, and encourage extensive exploration. Later ( Approaching 1) complete compensation of bias, focusing on high-value samples, and improving convergence accuracy.
[0037] Set the termination flag to "terminated". The termination conditions can be set to the end of the planning period or the technical indicators being seriously exceeded.
[0038] This embodiment constructs transition samples for DDQN and performs priority sampling, enabling DDQN to learn more frequently from high-error samples, accelerating the utilization of key experiences, and reducing training time in distribution network planning. The quality of the transition samples, as direct parameters for DDQN learning, directly affects the strategy optimization effect. In distribution network planning, it is essential to ensure that the samples cover various typical scenarios (such as high load, high renewable energy penetration, etc.) to improve model robustness.
[0039] In this embodiment, step S6 includes the following steps: Step S6.1: Based on batch processing Mini-Batch training, randomly sample B samples from the experience replay pool; Step S6.2, Network parameter update, including online network parameter update and target network parameter update; Step S6.3: Recalculate the TD error for each sample in the sampling batch and update the priority; Step S6.4: By inputting the initial state, planning period, and trained Q-network, output the optimal investment strategy sequence, including: period, investment action, cost, technical indicators, and cumulative investment.
[0040] The specific steps described above are as follows: Through batch training (Mini-Batch), B samples (B=64~256) are randomly sampled from the experience replay pool, and the target Q-value is calculated as follows: , The expression for the loss function is: , in, For importance sampling weights, The target Q value is calculated using the target network. Let Q be the predicted Q value of the online network, and B be the batch size.
[0041] Gradient calculation: via online network only Backpropagation.
[0042] Network parameter updates include online network parameter updates and target network parameter updates. The online network dynamically evaluates the value of each investment action in the current state and updates it through gradient descent at each training step.
[0043] The parameter update for the training step is expressed as follows: , For loss function, Recommended learning rate range (5×10) -4 ~2×10 -3 ).
[0044] The target network provides a reliable target for the Q-value, with a delayed update parameter (configurable to update every 500-1000 steps). The delayed update expression is as follows: , This is the target network update coefficient, and the recommended value range is (0.01~0.05).
[0045] After each training iteration, the TD error for each sample in the sampling batch is recalculated, and the priority is updated. The expression is as follows: , Given an initial state, a planning period, and a trained Q-network, the system outputs an optimal investment strategy sequence, including: period, investment actions, cost, technical indicators, and cumulative investment. The expression for this investment strategy sequence is: , Initialization starts from the actual state s0 of the distribution network and obtains the current operating indicators.
[0046] Rolling planning, for each time period Select the optimal action, then call the power flow calculation to simulate the impact of the action and generate a new distribution network state.
[0047] Termination condition checks include: Budget exhausted, remaining funds < cheapest action cost; Technical requirements are met: SAIDI ≤ SAIDI target value and line loss rate ≤ line loss rate target value; Forced termination, reaching the maximum planning period T; Technical indicators are severely exceeded (e.g., SAIDI > 10 hours).
[0048] This embodiment achieves stable convergence of the deep dual-Q network through Mini-Batch training and a soft update mechanism for the target network. The final output strategy sequence includes the optimal investment actions, costs, technical indicators, and cumulative investment for each period, providing intelligent decision-making basis for distribution network planning.
[0049] Example 2, Figure 2 A dynamic programming investment decision-making system for distribution networks based on deep dual-Q networks is presented, including: The data acquisition module is used to collect and normalize historical operation data of the regional power distribution network; The data processing module is used to construct a state space vector based on the normalized data; The model training module is used to input the vector into a deep double-Q network for iterative training, and achieve model convergence through experience replay and parameter update. The strategy output module is used to output a periodic optimal investment strategy sequence based on the converged model.
[0050] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0051] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0052] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0053] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0054] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0055] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for power distribution network dynamic planning investment decision based on deep double Q network, characterized in that, The method comprises the following steps: Collect and normalize regional power distribution network operation history data; Construct a state space vector based on the normalized data; Input the vector into a deep double Q network, perform iterative training, and realize model convergence through experience replay and parameter update; Output a periodic optimal investment strategy sequence based on the converged model.
2. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 1, characterized in that, The state space vector is constructed based on the normalized data, specifically: Calculate the covariance matrix of the data; Perform eigenvalue decomposition on the covariance matrix to obtain an eigenvalue diagonal matrix; Calculate the contribution rate of each eigenvalue, and select the data corresponding to the eigenvalues with a cumulative contribution rate not lower than a preset threshold to construct the state space vector.
3. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 2, characterized in that, The iterative training includes filtering invalid power distribution network planning investment actions through an action mask module.
4. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 3, characterized in that, The action mask module specifically comprises: Filter investment actions that violate constraints by calculating the physical constraints of the power distribution network in real time; Set the Q value of the investment action to negative or zero.
5. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 4, characterized in that, The iterative training includes designing a composite reward function, which is composed of an economic cost term, a technical reward term, and a constraint penalty term.
6. The deep double Q network-based power distribution network dynamic planning investment decision-making method according to claim 5, wherein: The economic cost term is calculated based on the discounted value of the initial investment cost and operation and maintenance cost of the equipment, and the discounted value is determined according to the investment period and a preset discount rate; The technical reward term includes reliability improvement reward, line loss reduction reward, and voltage qualification reward, and the reliability improvement reward includes adjusting an adaptive coefficient according to the planning stage; The constraint penalty term is determined according to whether the power supply reliability index and node voltage deviation exceed a preset threshold.
7. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 6, characterized in that, The experience replay includes setting sample priorities, specifically: Store the transition sample consisting of the current state, the selected investment action, the corresponding composite reward, the next state after the action is executed, and the termination flag into the experience replay pool; Determine the priority of each transition sample in the experience replay pool based on the absolute size of the time difference error, and introduce a minimum constant in the calculation; Assign the highest priority to the transition sample newly added to the experience replay pool during initialization; Determine the probability of the sample being extracted for training according to the relative size of each sample priority in the replay pool.
8. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 7, characterized in that, After determining the probability of the sample being extracted for training, it further includes calculating the importance sampling weight based on the sample extraction probability.
9. The power grid dynamic planning investment decision-making method based on deep double Q network according to claim 8, characterized in that, The iterative training further comprises using a greedy policy to select the investment action.
10. A system using the power distribution network dynamic planning investment decision method based on deep double Q network according to any one of claims 1-9, characterized in that, It includes: A data collection module for collecting and normalizing regional power distribution network operation history data; A data processing module for constructing a state space vector based on the normalized data; A model training module for inputting the vector into a deep double Q network, performing iterative training, and realizing model convergence through experience replay and parameter update; A strategy output module for outputting a periodic optimal investment strategy sequence based on the converged model.
Citation Information
Patent Citations
Robot path planning method and system based on priority experience playback mechanism
CN115509233A
Health degree evaluation system and method for endocrine nursing
CN120089386A
Task scheduling method and device based on edge computing, equipment and storage medium
CN120448069A
Multi-time scale demand response optimization method and device based on dual Q network
CN120875427A