Power distribution network dispatching optimization method based on exploration-question iterative optimization

By constructing an expert mechanism model and a deep reinforcement learning model that include long-term and short-term constraints, and combining the cross-questioning mechanism of exploratory decision-making and mechanism model, the problems of insufficient scenario coverage and poor economic efficiency of traditional distribution network dispatching models are solved, and efficient and safe dispatching optimization is achieved.

CN121840642APending Publication Date: 2026-04-10山东华科信息技术有限公司 +6
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing power distribution network dispatching models rely on expert experience, resulting in insufficient scenario coverage and limited optimization for unknown operating conditions. Traditional optimization methods are computationally inefficient and economical, making it difficult to meet real-time and intelligent requirements.

Method used

An expert mechanism model containing long-term and short-term constraints is constructed, and a deep reinforcement learning model with exploratory neurons in the input layer is established. The exploration coefficient is adaptively adjusted in the simulation environment, and the strategy is optimized through the cross-questioning mechanism between the exploration decision and the mechanism model.

Benefits of technology

It significantly improves the proactive adaptability and optimization efficiency of the distribution network dispatching model in the face of unknown and complex operating conditions, and realizes the continuous evolution and economic improvement of dispatching strategies while ensuring safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121840642A_ABST
    Figure CN121840642A_ABST
Patent Text Reader

Abstract

The invention provides a power distribution network dispatching optimization method based on exploration-question iterative optimization, and belongs to the technical field of power distribution network dispatching. Comprising the following steps: S1, exploration model training based on condition confidence: constructing a power distribution network dispatching optimization mechanism model containing long-term and short-term constraints based on expert experience; constructing a deep reinforcement learning scheduling optimization model containing exploration neurons, adaptively adjusting exploration coefficients, and optimizing an exploration model; s2, power distribution network dispatching optimization based on exploration decision and mechanism model cross question: cross question decision based on condition similarity and safety distance, double-strategy verification and challenge judgment, and model aggregation optimization based on performance gap and safety distance. According to the method, the active adaptive capacity and optimization efficiency of the power distribution network dispatching model for unknown complex working conditions are remarkably improved, continuous evolution and self-improvement of the dispatching strategy are realized, and the safety and economy of power distribution network dispatching are effectively balanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application provides a power distribution network scheduling optimization method based on exploration-questioning iterative optimization, and belongs to the technical field of power distribution network scheduling. BACKGROUND

[0002] At present, power distribution network scheduling mainly relies on mathematical models constructed based on physical mechanisms and expert experience, and solves scheduling strategies by strictly modeling long-term safety constraints (such as energy storage power balance) and short-term safety constraints (such as line thermal stability) and using optimization algorithms. The mechanism model based on expert experience is widely used in traditional power grids, can provide a deterministic safety boundary for the system, is the basic guarantee for the current power distribution network scheduling control, and ensures the stable operation of the power grid under known working conditions.

[0003] However, the existing scheduling optimization method based on expert experience has a series of limitations when facing high uncertainty of sources and loads. First, expert experience tends to be conservative, resulting in a too large safety domain set by the constructed mechanism model, which seriously sacrifices the operation economy and limits the utilization efficiency of assets. Secondly, the calculation efficiency of the traditional optimization method is low in the face of a large number of variable operation scenarios, and it is difficult to meet the real-time requirement. More importantly, existing experience is difficult to cover all extreme or unknown sudden scenarios, resulting in that the model often cannot output the optimal strategy when facing "no experience reference" new working conditions due to lack of flexibility, and even may produce suboptimal solution or no solution due to insufficient scenario coverage, which is difficult to meet the growing intelligent demand of power distribution network. SUMMARY

[0004] The application provides a power distribution network scheduling optimization method based on exploration-questioning iterative optimization, which has solved the following technical problems:

[0005] 1. In order to solve the problems of insufficient scene coverage and unknown working condition optimization limitation caused by existing scheduling model relying on expert experience, the application provides a training method of exploration model based on condition confidence. First, based on expert experience, a scheduling mechanism model distinguishing long-term safety constraints and short-term hard constraints is constructed, and a deep reinforcement learning model with exploration neurons in the input layer is established. The mechanism verification is used to ensure the safety of the output in the early shaping period. Secondly, in the simulation environment, the historical operation conditions are classified by using the clustering algorithm, and the exploration coefficient is adaptively adjusted according to the similarity of the current condition and the typical scene and the historical performance. Finally, only the short-term hard constraints are locked in the exploration strategy verification, and the long-term constraints are ignored to expand the search space, and the model parameters are updated through simulation feedback. Through this method, the application effectively overcomes the defects of traditional expert experience model, such as too large safety domain setting and insufficient scene coverage, and uses condition confidence to guide the model to carry out directional and efficient exploration within the safety boundary, which significantly improves the active adaptation ability and optimization efficiency of the distribution network scheduling model facing unknown complex working conditions.

[0006] 2. In order to solve the problems of poor economy and lack of evolution ability caused by the conservative safety domain of traditional mechanism model, the application provides a distribution network scheduling optimization method based on exploration decision and mechanism model cross questioning. First, the same power grid condition is input into the basic model and the exploration model at the same time, the strategy deviation is calculated, if the deviation exceeds the threshold, the challenge score is calculated combining the condition similarity, historical challenge success rate and hard constraint safety distance, and it is decided whether to initiate the challenge. Secondly, after triggering the challenge, the double strategy is substituted into the high-precision simulation environment for full-cycle verification, if the exploration strategy is more economical than the basic strategy under the premise of meeting all long-term and short-term constraints, the challenge is successful. Finally, the aggregation weight is calculated based on the performance improvement amplitude and the hard constraint safety margin, and the mechanism model is updated with weighting. Through the closed-loop iteration mechanism, the application can break the shackles of the conservative traditional mechanism model while strictly guaranteeing the safety of system operation, realize the continuous evolution and self-improvement of the scheduling strategy, and effectively balance the safety and economy of the distribution network scheduling.

[0007] The application provides a distribution network scheduling optimization method based on exploration-questioning iterative optimization. First, an expert mechanism model containing long-term and short-term constraints is constructed, and a deep reinforcement learning model with exploration neurons in the input layer is established, and the mechanism verification is used to assist the initial shaping of the model. Secondly, in the simulation environment, the exploration coefficient is adaptively adjusted based on the condition clustering similarity, and only the short-term hard constraints are relaxed for efficient exploration training. Finally, a cross questioning mechanism of exploration decision and mechanism model is established, and simulation challenge is initiated for conditions with large strategy deviation, and if the exploration strategy is better in performance under the premise of meeting all constraints, the model parameters are updated based on the safety margin aggregation. Through this process, the application realizes the autonomous evolution of the scheduling model while strictly guaranteeing the safety, effectively solving the problem of low efficiency of the traditional model.

[0008] The specific technical solution is:

[0009] A power distribution network scheduling optimization method based on exploration-question iteration optimization, comprising the following steps:

[0010] S1, exploration model training based on condition confidence

[0011] First, based on expert experience, a scheduling mechanism model is constructed to distinguish long-term safety constraints and short-term hard constraints, and a deep reinforcement learning model with exploration neurons in the input layer is established. In the early shaping period, the mechanism verification is used to ensure the safety of the output.

[0012] Secondly, in the simulation environment, the historical operation conditions are classified by using clustering algorithm, and the exploration coefficient is adjusted adaptively according to the similarity of the current condition and the typical scene and the historical performance.

[0013] Finally, in the exploration strategy verification, only the short-term hard constraint is locked, and the long-term constraint is ignored to expand the search space, and the model parameters are updated through simulation feedback.

[0014] S2, power distribution network scheduling optimization based on exploration decision and mechanism model cross-question

[0015] First, the same power grid condition is input into the basic model and the exploration model at the same time, the strategy deviation is calculated, if the deviation exceeds the threshold, the challenge score is calculated according to the condition similarity, historical challenge success rate and hard constraint safety distance, and it is decided whether to initiate the challenge.

[0016] Secondly, after triggering the challenge, the double-strategy is substituted into the high-precision simulation environment for full-cycle verification, if the exploration strategy is more economical than the basic strategy under the premise of meeting all long-term and short-term constraints, the challenge is successful.

[0017] Finally, based on the performance improvement amplitude and the hard constraint safety margin, the aggregation weight is calculated, and the mechanism model is updated with weighted.

[0018] Further:

[0019] S1, exploration model training based on condition confidence, specifically including the following sub-steps:

[0020] S1.1 Constructing a power distribution network scheduling optimization mechanism model containing long-term and short-term constraints based on expert experience

[0021] For the distribution network source-grid-load-storage integration scenario, a scheduling optimization mechanism model is constructed based on expert experience. This model takes minimizing the overall system operating cost within the scheduling cycle as its objective function and strictly divides the constraints into short-term and long-term security constraints. Short-term security constraints refer to the physical hard limitations that must be met at any single time point, while long-term security constraints refer to limitations involving time coupling or cumulative effects over the entire cycle. The mathematical model is established as follows:

[0022] (1)

[0023] In the formula, scheduling period Total cost within; It is a collection of generator sets, nodes, lines, energy storage devices, and on-load tap-changing transformers; This is a time-period index. For the first The power generation cost coefficient of the Taiwanese generator unit; For the first Time period The active power output of the Taiwanese generator set; For the first The operation and maintenance costs of the unit; For the first The unit price of electricity purchased from the grid during the specified time period; For the first Power exchanged with the upstream power grid during specific time periods; This is the network loss penalty coefficient; For the first System network loss during certain periods; For the first Lifetime loss cost coefficient of an energy storage device; For the first Time period The charging and discharging power of an energy storage device;

[0024] In short-run constraints: For the first The lower and upper limits of the unit's output; For the first Time period The voltage amplitude at each node, For the first The lower and upper limits of the voltage safety threshold for each node; For the first Time period The current amplitude of the line, For the first The thermal stability limit of the line; For the first Time period The load power of each node; For the first Time period The state of charge of an energy storage device; For the first The self-discharge rate of the energy storage unit; The first The charging and discharging efficiency of an energy storage device; The first The charging and discharging power components of an energy storage device; For the first The lower and upper limits of the State of Charge (SOC) of an energy storage battery; For the first The initial and final SOC states of an energy storage system.

[0025] In long-term constraints: For the first Time period The tap position of the transformer; For the first The maximum number of permissible operations for a transformer throughout its entire lifecycle.

[0026] S1.2 Constructing a deep reinforcement learning scheduling optimization model with exploratory neurons

[0027] Based on historical operational data, a distribution network scheduling optimization model, or basic model, based on deep reinforcement learning is constructed. An independent exploration neuron unit is set in the input layer of the model, with its value corresponding to the exploration weight. In the early stages of model shaping, the value of the exploration neuron is forcibly set to 0. At this time, the output result must pass the safety verification of the S1.1 mechanism model; if the constraints are violated, rule correction is applied. After the basic model training converges, the network structure and parameters of this model are copied to generate the exploration model. The network forward propagation and exploration mechanism are defined as follows:

[0028] (2)

[0029] (3)

[0030] In the formula, For the first The scheduling action vector output by the time-segment model; The activation function for the output layer; The output layer weight matrix and bias vector; This is the feature extraction function for the hidden layer of a neural network. For the first The performance of the time-period distribution network scheduling optimization model is defined in S1.3; Network parameters of the basic model; For the first Random noise signals generated over a period of time; The weight values ​​of neurons are determined during the scheduling phase to explore their potential. The adaptive exploration coefficients are calculated based on the state during the exploration phase; Phase indicates the current training phase of the model, Shaping indicates the model shaping phase, and Exploring indicates the exploration phase.

[0031] S1.3 Adaptive adjustment of exploration coefficient

[0032] The exploratory model was trained in a distribution network dispatch optimization simulation environment. A clustering algorithm was used to analyze the historical distribution network operation status. Divided into A typical scenario, among which For the first The first characteristic of the state in a typical scenario Dimensional value. Define the exploration coefficient. Its value depends on the similarity between the current situation and the center of a typical scene, as well as the model's historical performance in the corresponding scene. Higher similarity and better historical performance result in a greater exploration weight. For the current... Feature extraction is performed on the operating status of the distribution network during different time periods, and the first feature of the status is... Dimensional numerical representation is .but The calculation is as follows:

[0033] (4)

[0034] (5)

[0035] (6)

[0036] In the formula, For the first Current status of the time period With the Cluster centers The weighted Euclidean distance; The dimension of the state vector; For the first Weights of state features; For the current state Dimensional value; For the first The first cluster center Dimensional value. To represent the scenario that most closely resembles the typical case obtained from clustering, For the first Time period The weighted Euclidean distance between the cluster centers. The baseline exploration coefficient; This is the distance attenuation parameter; As a performance incentive factor; To explore the model in the first Historical average reward value in similar scenarios; The baseline reward value for the basic model; To prevent small quantities with a denominator of zero.

[0037] S1.4 Exploration Model Optimization

[0038] When performing safety checks and gradient updates, the exploration strategy only considers the short-term safety constraints defined in S1.1; the loss function of the exploration model is expressed as:

[0039] (7)

[0040] (8)

[0041] In the formula, To explore the loss function of the model; To explore model parameters; Indicates the expected value; Economic rewards for environmental feedback; The penalty weight for violating short-term hard constraints; This is the function for penalizing exceeding the limit; For the first The first period The physical quantity values ​​corresponding to each short-term safety constraint. This represents the total number of short-term constraints in the model. These correspond to the upper and lower limits of the constraints in S1.1.

[0042] S2. Distribution network dispatch optimization based on cross-questioning of exploratory decision-making and mechanism models, specifically including the following sub-steps:

[0043] S2.1 Cross-questioning decision based on situation similarity and safety distance

[0044] The same power grid operation status Simultaneously, the inputs are fed into both the basic scheduling optimization model and the exploration model. The policy deviation between the two is calculated. When the deviation exceeds a preset threshold, the exploration model calculates a challenge score based on situation similarity, historical challenge success rate, and hard constraint safety distance, and decides whether to initiate a challenge, expressed as:

[0045] (9)

[0046] (10)

[0047] In the formula, This is a strategy normalization bias; For the action dimension; The basic model and the exploratory model are respectively in the first two stages. Time period Dimensional action output. To score points for the challenge; These are the weighting coefficients; The maximum cluster distance; For the first Success rate of historical challenges in similar scenarios; This represents the absolute margin of the current physical quantity from the nearest hard constraint boundary. This is the nominal value of the physical quantity. Only when... and If all values ​​exceed the threshold, a challenge is initiated.

[0048] S2.2 Dual-Strategy Verification and Challenge Judgment

[0049] After the challenge is initiated, the two strategies are substituted into a high-precision distribution network dispatch optimization simulation environment. If the objective function value of the exploration model strategy is better than that of the mechanistic model, provided that all constraints described in S1.1 are satisfied, then the challenge is considered successful. The determination logic is as follows:

[0050] (11)

[0051] In the formula, The sign indicating a successful challenge; The total cost objective function defined for S1.1; These are the exploration strategy and the mechanism model strategy vectors, respectively. The set of short-run constraint violations is defined if and only if all The value is 0 when the constraint is satisfied. For a set of long-term constraint violations, if and only if all The value is 0 when the constraint is met.

[0052] S2.3 Model Aggregation Optimization Based on Performance Gap and Safety Distance

[0053] After successfully completing the challenge, the performance gap between the exploratory model and the mechanistic model is calculated, along with the hard constraint safety distance of the exploratory model. Based on this, aggregation weights are set to update the mechanistic model. For safety reasons, a weight cap is set. The aggregation update formula is expressed as:

[0054] (12)

[0055] (13)

[0056] In the formula, Aggregate weights for the model; This represents the upper limit of the weight. This represents the performance improvement value. To explore the normalized minimum distance of the strategy to the nearest short-term hard constraint boundary in the simulation; Impact factor weights; This is the sensitivity coefficient. For the updated mechanistic model parameters; The parameters of the mechanistic model before the update; These are the parameters for the current exploration model.

[0057] The beneficial effects of the technical solution of this invention are as follows:

[0058] 1. This invention proposes a training method for an exploration model based on situation confidence. First, a scheduling mechanism model that distinguishes between long-term safety constraints and short-term hard constraints is constructed based on expert experience. A deep reinforcement learning model with exploration neurons in the input layer is established, and output safety is ensured through mechanism verification in the early stages of model shaping. Second, training is conducted in a simulation environment. Historical operating conditions are classified using a clustering algorithm, and the exploration coefficients are adaptively adjusted based on the similarity between the current condition and typical scenarios, as well as historical performance. Finally, during exploration strategy verification, only short-term hard constraints are locked, ignoring long-term constraints to expand the search space, and model parameters are updated through simulation feedback. This invention effectively overcomes the shortcomings of traditional expert experience models, such as excessively large safety domains and insufficient scenario coverage. By using situation confidence to guide the model to conduct targeted and efficient exploration within the safety boundary, it significantly improves the proactive adaptability and optimization efficiency of the distribution network scheduling model in the face of unknown and complex operating conditions.

[0059] 2. This invention proposes a distribution network scheduling optimization method based on cross-questioning of exploratory decision-making and mechanistic models. First, the same grid condition is simultaneously input into both the basic model and the exploratory model, and the strategy deviation is calculated. If the deviation exceeds a threshold, a challenge score is calculated based on condition similarity, historical challenge success rate, and hard constraint safety distance to determine whether to initiate a challenge. Second, after triggering the challenge, both strategies are substituted into a high-precision simulation environment for full-cycle verification. If the exploratory strategy is more economical than the basic strategy while satisfying all long-term and short-term constraints, the challenge is considered successful. Finally, aggregate weights are calculated based on performance improvement and hard constraint safety margin, and the mechanistic model is updated with weighted averages. This invention, through a closed-loop iterative mechanism, can break the conservatism of traditional mechanistic models while strictly ensuring system operational safety, achieving continuous evolution and self-improvement of scheduling strategies, and effectively balancing the safety and economy of distribution network scheduling. Attached Figure Description

[0060] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0061] The specific technical solution of the present invention will be described in conjunction with the accompanying drawings.

[0062] This invention proposes a distribution network scheduling optimization method based on exploration-challenge iterative optimization. The method includes training a deep reinforcement learning model with adaptive exploration coefficients to generate new strategies, constructing a cross-challenge mechanism between the new strategy and the underlying mechanism model, and aggregating and updating the underlying mechanism model after a successful challenge. This achieves safe iteration and continuous optimization of the scheduling strategy. The specific process is as follows: Figure 1 As shown.

[0063] S1 Exploratory Model Training Method Based on State Confidence

[0064] This invention constructs a deep reinforcement learning model containing exploratory neurons and trains the exploratory model in a simulation environment using adaptive exploratory coefficients based on situation clustering similarity and historical performance feedback. This enables autonomous strategy exploration under the premise of ensuring short-term hard constraints, effectively solving the problems of insufficient coverage of expert experience scheduling scenarios and low optimization efficiency.

[0065] S1.1 Constructing a distribution network dispatch optimization mechanism model based on expert experience, incorporating both long-term and short-term constraints.

[0066] For distribution network source-grid-load-storage integration scenarios, a scheduling optimization mechanism model is constructed based on expert experience. This model takes minimizing the overall system operating cost within the scheduling cycle as its objective function and strictly divides the constraints into short-term and long-term security constraints. Short-term security constraints refer to the physical hard limitations that must be met at any single time point, while long-term security constraints refer to limitations involving time coupling or cumulative effects over the entire cycle. For example, the following mathematical model can be established, expressed as:

[0067] (1)

[0068] In the formula, scheduling period Total cost within; It is a collection of generator sets, nodes, lines, energy storage devices, and on-load tap-changing transformers; This is a time-period index. For the first The power generation cost coefficient of the Taiwanese generator unit; For the first Time period The active power output of the Taiwanese generator set; For the first The operation and maintenance costs of the unit; For the first The unit price of electricity purchased from the grid during the specified time period; For the first Power exchanged with the upstream power grid during specific time periods; This is the network loss penalty coefficient; For the first System network loss during certain periods; For the first Lifetime loss cost coefficient of an energy storage device; For the first Time period The charging and discharging power of an energy storage device (charging is positive, discharging is negative).

[0069] In short-run constraints: For the first The lower and upper limits of the unit's output; For the first Time period The voltage amplitude at each node, For the first The lower and upper limits of the voltage safety threshold for each node; For the first Time period The current amplitude of the line, For the first The thermal stability limit of the line; For the first Time period The load power of each node; For the first Time period The state of charge of an energy storage device; For the first The self-discharge rate of the energy storage unit; The first The charging and discharging efficiency of an energy storage device; The first The charging and discharging power components of an energy storage device; For the first The lower and upper limits of the State of Charge (SOC) of an energy storage battery; For the first The initial and final SOC states of an energy storage system.

[0070] In long-term constraints: For the first Time period The tap position of the transformer; For the first The maximum number of permissible operations for a transformer throughout its entire lifecycle.

[0071] This invention does not limit the objective function, constraints, and network topology in the above model. Technical personnel can flexibly expand or adjust the model according to the actual form of the distribution network (such as radial, ring, or independent microgrids) and specific business needs. For example, for scenarios where high-proportion distributed photovoltaic access leads to fluctuations on both the source and load sides, constraints related to preventing node voltage overruns and backflow can be added; for low-carbon-oriented dispatching needs, carbon emission indicators can be incorporated into the objective function; for vehicle-grid interaction or demand response scenarios, the adjustable capacity constraints of the electric vehicle fleet and the interruption characteristic constraints of flexible loads can be further expanded.

[0072] S1.2 Constructing a deep reinforcement learning scheduling optimization model with exploratory neurons

[0073] Based on historical operational data, a distribution network scheduling optimization model (basic model) based on deep reinforcement learning is constructed. An independent exploratory neuron unit is set in the input layer of the model, with its value corresponding to the exploratory weight. In the early stages of model shaping, the value of the exploratory neuron is forcibly set to 0. At this time, the output result must pass the safety verification of the mechanism model described in S1.1; if the constraints are violated, rule correction is applied. After the basic model training converges, the network structure and parameters of this model are copied to generate the exploratory model. The network forward propagation and exploration mechanism are defined as follows:

[0074] (2)

[0075] (3)

[0076] In the formula, For the first The scheduling action vector output by the time-segment model; The activation function for the output layer; The output layer weight matrix and bias vector; This is the feature extraction function for the hidden layer of a neural network. For the first The performance of the time-period distribution network scheduling optimization model is defined in S1.3; Network parameters of the basic model; For the first Random noise signals generated over a period of time; The weight values ​​of neurons are determined during the scheduling phase to explore their potential. The adaptive exploration coefficients are calculated based on the state during the exploration phase; Phase indicates the current training phase of the model, Shaping indicates the model shaping phase, and Exploring indicates the exploration phase.

[0077] S1.3 Adaptive adjustment of exploration coefficient

[0078] The exploratory model was trained in a distribution network dispatch optimization simulation environment. A clustering algorithm was used to analyze the historical distribution network operation status. Divided into A typical scenario, among which For the first The first characteristic of the state in a typical scenario Dimensional value. Define the exploration coefficient. Its value depends on the similarity between the current situation and the center of a typical scene, as well as the model's historical performance in the corresponding scene. Higher similarity and better historical performance result in a greater exploration weight. For the current... Feature extraction is performed on the operating status of the distribution network during different time periods, and the first feature of the status is... Dimensional numerical representation is .but The calculation is as follows:

[0079] (4)

[0080] (5)

[0081] (6)

[0082] In the formula, For the first Current status of the time period With the Cluster centers The weighted Euclidean distance; The dimension of the state vector; For the first Weights of state features; For the current state Dimensional value; For the first The first cluster center Dimensional value. To represent the scenario that most closely resembles the typical case obtained from clustering, For the first Time period The weighted Euclidean distance between the cluster centers. The baseline exploration coefficient; This is the distance attenuation parameter; As a performance incentive factor; To explore the model in the first Historical average reward value in similar scenarios; The baseline reward value for the basic model; To prevent small quantities with a denominator of zero.

[0083] S1.4 Exploration Model Optimization

[0084] The exploration strategy, when performing safety checks and gradient updates, only considers short-term safety constraints (such as voltage, current, and power balance) defined in S1.1, temporarily ignoring long-term coupling constraints to expand the search space and find potential global optima. The loss function of the exploration model is expressed as:

[0085] (7)

[0086] (8)

[0087] In the formula, To explore the loss function of the model; To explore model parameters; Indicates the expected value; Economic rewards for environmental feedback; The penalty weight for violating short-term hard constraints; The penalty function is for exceeding the limit. The specific form of the penalty function varies depending on the constraints. In actual engineering construction, it can be expressed differently according to the physical characteristics of the controlled object (such as transformer, line, inverter, etc.). For example, it can be defined as a quantitative value reflecting the risk of insulation aging cost caused by short-term overload operation, depreciation acceleration cost caused by line heating, or potential fault shutdown. This invention does not strictly limit the specific mathematical structure of the penalty function, and it can be defined according to actual operation needs and equipment physical characteristics. For the first The first period The physical quantity values ​​corresponding to each short-term safety constraint. This represents the total number of short-term constraints in the model. This corresponds to the upper and lower bounds of the constraints in S1.1. This step ensures that the strategy generated by the exploration model is instantaneously feasible at the physical level.

[0088] S2 A Distribution Network Dispatch Optimization Method Based on Cross-Questioning of Exploratory Decision-Making and Mechanism Models

[0089] This invention constructs a cross-questioning mechanism between exploratory decision-making and mechanistic models. When the strategy deviation is large, simulation challenges are triggered based on situation similarity. After the challenge is successfully completed, the mechanistic model is updated using aggregated weights based on performance gap and safety margin. This achieves safe iterative optimization of the distribution network scheduling model and breaks through the conservative limitations of traditional mechanistic models.

[0090] S2.1 Cross-questioning decision based on situation similarity and safety distance

[0091] The same power grid operation status Simultaneously, the inputs are fed into both the basic scheduling optimization model and the exploration model. The policy deviation between the two is calculated. When the deviation exceeds a preset threshold, the exploration model calculates a challenge score based on situation similarity, historical challenge success rate, and hard constraint safety distance, and decides whether to initiate a challenge, expressed as:

[0092] (9)

[0093] (10)

[0094] In the formula, This is a strategy normalization bias; For the action dimension; The basic model and the exploratory model are respectively in the first two stages. Time period Dimensional action output. To score points for the challenge; These are the weighting coefficients; The maximum cluster distance; For the first Success rate of historical challenges in similar scenarios; This represents the absolute margin of the current physical quantity from the nearest hard constraint boundary. This is the nominal value of the physical quantity. Only when... and If all values ​​exceed the threshold, a challenge is initiated.

[0095] S2.2 Dual-Strategy Verification and Challenge Judgment

[0096] After the challenge is initiated, the two strategies are substituted into a high-precision distribution network dispatch optimization simulation environment. If the objective function value of the exploration model strategy is better than that of the mechanistic model, provided that all constraints described in S1.1 (including short-term hard constraints and long-term safety constraints) are satisfied, the challenge is considered successful. The determination logic is as follows:

[0097] (11)

[0098] In the formula, The sign indicating a successful challenge; The total cost objective function defined for S1.1; These are the exploration strategy and the mechanism model strategy vectors, respectively. The set of short-run constraint violations is defined if and only if all The value is 0 when the constraint is satisfied. For a set of long-term constraint violations, if and only if all The value is 0 when the constraints are met. This step ensures that the exploration strategy is not only more economical but also meets the safety requirements throughout the entire lifecycle.

[0099] S2.3 Model Aggregation Optimization Based on Performance Gap and Safety Distance

[0100] After successfully completing the challenge, the performance gap between the exploratory model and the mechanistic model is calculated, along with the hard constraint safety distance of the exploratory model. Based on this, aggregation weights are set to update the mechanistic model. For safety reasons, a weight cap is set. The aggregation update formula is expressed as:

[0101] (12)

[0102] (13)

[0103] In the formula, Aggregate weights for the model; This represents the upper limit of the weight. This represents the performance improvement value. To explore the normalized minimum distance of the strategy to the nearest short-term hard constraint boundary in the simulation; Impact factor weights; This is the sensitivity coefficient. For the updated mechanistic model parameters; The parameters of the mechanistic model before the update; These are the parameters for the current exploratory model. This step enables the gradual optimization of the mechanistic model while ensuring safety.

Claims

1. A distribution network dispatch optimization method based on exploration-questioning iterative optimization, characterized in that, Includes the following steps: S1. Training of the Exploratory Model Based on State Confidence First, based on expert experience, a scheduling mechanism model is constructed to distinguish between long-term safety constraints and short-term hard constraints. A deep reinforcement learning model with exploratory neurons in the input layer is also established to ensure output safety through mechanism verification in the early stage of shaping. Secondly, training is conducted in a simulation environment, and clustering algorithms are used to classify historical operating conditions. Based on the similarity between the current condition and typical scenarios and historical performance, the exploration coefficient is adaptively adjusted. Finally, when exploring strategy verification, only short-term hard constraints are locked and long-term constraints are ignored to expand the search space, and the model parameters are updated through simulation feedback. S2. Distribution network dispatch optimization based on cross-questioning of exploratory decision-making and mechanism models First, the same power grid condition is simultaneously input into the basic model and the exploration model to calculate the strategy deviation. If the deviation exceeds the threshold, the challenge score is calculated by combining the condition similarity, historical challenge success rate and hard constraint safety distance to determine whether to initiate a challenge. Secondly, after the challenge is triggered, the two strategies are substituted into a high-precision simulation environment for full-cycle verification. If the exploration strategy is more economical than the basic strategy under the premise of satisfying all long-term and short-term constraints, the challenge is judged to be successful. Finally, the aggregate weights are calculated based on the performance improvement and the safety margin of hard constraints, and the mechanistic model is updated in a weighted manner.

2. The distribution network dispatch optimization method based on exploration-questioning iterative optimization according to claim 1, characterized in that, S1. Training the exploration model based on situation confidence, specifically including the following sub-steps: S1.1 Constructing a distribution network dispatch optimization mechanism model based on expert experience, incorporating both long-term and short-term constraints. For the distribution network source-grid-load-storage integration scenario, a scheduling optimization mechanism model is constructed based on expert experience. This model takes minimizing the overall system operating cost within the scheduling cycle as its objective function and strictly divides the constraints into short-term and long-term security constraints. Short-term security constraints refer to the physical hard limitations that must be met at any single time point, while long-term security constraints refer to limitations involving time coupling or full-cycle accumulation. The following mathematical model is established, expressed as: (1) In the formula, scheduling period Total cost within; It is a collection of generator sets, nodes, lines, energy storage devices, and on-load tap-changing transformers; Indexed by time period; For the first The power generation cost coefficient of the Taiwanese generator unit; For the first Time period The active power output of the Taiwanese generator set; For the first The operation and maintenance costs of the unit; For the first The unit price of electricity purchased from the grid during the specified time period; For the first Power exchanged with the upstream power grid during specific time periods; This is the network loss penalty coefficient; For the first System network loss during certain periods; For the first Lifetime loss cost coefficient of an energy storage device; For the first Time period The charging and discharging power of an energy storage device; In short-run constraints: For the first The lower and upper limits of the unit's output; For the first Time period The voltage amplitude at each node, For the first The lower and upper limits of the voltage safety threshold for each node; For the first Time period The current amplitude of the line, For the first The thermal stability limit of the line; For the first Time period The load power of each node; For the first Time period The state of charge of an energy storage device; For the first The self-discharge rate of the energy storage unit; The first The charging and discharging efficiency of an energy storage device; The first The charging and discharging power components of the energy storage; For the first The lower and upper limits of the State of Charge (SOC) of an energy storage battery; For the first The initial and final SOC states of an energy storage device; In long-term constraints: For the first Time period The tap position of the transformer; For the first The maximum permissible number of operations for a transformer throughout its entire lifecycle; S1.2 Constructing a deep reinforcement learning scheduling optimization model with exploratory neurons Based on historical operational data, a distribution network scheduling optimization model based on deep reinforcement learning, i.e., the basic model, is constructed. An independent exploration neuron unit is set in the input layer of the model, with its value corresponding to the exploration weight. In the early stage of model shaping, the value of the exploration neuron is forcibly set to 0. At this time, the output result must pass the safety verification of the S1.1 mechanism model; if the constraint is violated, rule correction is adopted. After the basic model training converges, the network structure and parameters of the model are copied to generate the exploration model. The network forward propagation and exploration mechanism are defined as follows: (2) (3) In the formula, For the first The scheduling action vector output by the time-segment model; The activation function for the output layer; The output layer weight matrix and bias vector; This is the feature extraction function for the hidden layer of a neural network. For the first The performance of the time-period distribution network scheduling optimization model is defined in S1.3; Network parameters of the basic model; For the first Random noise signals generated over a period of time; The weight values ​​of neurons are determined during the scheduling phase to explore their potential. For the adaptive exploration coefficients calculated based on the state during the exploration phase; Phase indicates the current training phase of the model, Shaping indicates the model shaping phase, and Exploring indicates the exploration phase; S1.3 Adaptive adjustment of exploration coefficient The exploratory model was trained in a power distribution network dispatch optimization simulation environment; Using clustering algorithms to analyze the historical operation status of the power distribution network Divided into A typical scenario, among which For the first The first characteristic of the state in a typical scenario Dimensional value; defining the exploration coefficient Its value depends on the similarity between the current situation and the center of a typical scene, as well as the model's historical performance in the corresponding scene; the higher the similarity and the better the historical performance, the greater the exploration weight is assigned; for the current... Feature extraction is performed on the operating status of the distribution network during different time periods, and the first feature of the status is... Dimensional numerical representation is ;but The calculation is as follows: (4) (5) (6) In the formula, For the first Current status of the time period With the Cluster centers The weighted Euclidean distance; The dimension of the state vector; For the first Weights of state features; For the current state Dimensional value; For the first The first cluster center Dimensional value; To represent the scenario that most closely resembles the typical case obtained from clustering, For the first Time period The weighted Euclidean distance between the cluster centers; The baseline exploration coefficient; This is the distance attenuation parameter; As a performance incentive factor; To explore the model in the first Historical average reward value in similar scenarios; The baseline reward value for the basic model; To prevent small quantities with a denominator of zero; S1.4 Exploration Model Optimization When performing safety checks and gradient updates, the exploration strategy only considers the short-term safety constraints defined in S1.1; the loss function of the exploration model is expressed as: (7) (8) In the formula, To explore the loss function of the model; To explore model parameters; Indicates the expected value; Economic rewards for environmental feedback; The penalty weight for violating short-term hard constraints; This is the function for penalizing exceeding the limit; For the first The first period The physical quantity values ​​corresponding to each short-term safety constraint. This represents the total number of short-term constraints in the model. These correspond to the upper and lower limits of the constraints in S1.

1.

3. The distribution network dispatch optimization method based on exploration-questioning iterative optimization according to claim 2, characterized in that, S2. Distribution network dispatch optimization based on cross-questioning of exploratory decision-making and mechanism models, specifically including the following sub-steps: S2.1 Cross-questioning decision based on situation similarity and safety distance The same power grid operation status Simultaneously, the data is input into both the basic scheduling optimization model and the exploration model; the policy deviation between the two is calculated. When the deviation exceeds a preset threshold, the exploration model calculates the challenge score based on situation similarity, historical challenge success rate, and hard constraint safety distance, and decides whether to initiate a challenge, as shown below: (9) (10) In the formula, This is a strategy normalization bias; For the action dimension; The basic model and the exploratory model are respectively in the first two stages. Time period Dimensional action output; To score points for the challenge; These are the weighting coefficients; The maximum cluster distance; For the first Success rate of historical challenges in similar scenarios; This represents the absolute margin of the current physical quantity from the nearest hard constraint boundary. This is the nominal value of the physical quantity; only when and If all values ​​exceed the threshold, initiate a challenge; S2.2 Dual-Strategy Verification and Challenge Judgment After the challenge is initiated, the two strategies are substituted into the high-precision distribution network dispatch optimization simulation environment. If the objective function value of the exploration model strategy is better than that of the mechanistic model under the premise of satisfying all the constraints described in S1.1, the challenge is considered successful. The determination logic is as follows: (11) In the formula, The sign indicating a successful challenge; The total cost objective function defined for S1.1; These are the exploration strategy and the mechanism model strategy vectors, respectively. The set of short-run constraint violations is defined if and only if all The value is 0 when the constraint is satisfied. For a set of long-term constraint violations, if and only if all The value is 0 when the constraint is satisfied. S2.3 Model Aggregation Optimization Based on Performance Gap and Safety Distance After successfully completing the challenge, the performance gap between the exploration model and the mechanism model is calculated, as well as the hard constraint safety distance of the exploration model. Based on this, the aggregation weight is set, and the mechanism model is updated. For safety reasons, a weight cap is set. The aggregation update formula is expressed as: (12) (13) In the formula, Aggregate weights for the model; This represents the upper limit of the weight. This represents the performance improvement value. To explore the normalized minimum distance of the strategy to the nearest short-term hard constraint boundary in the simulation; Impact factor weights; This is the sensitivity coefficient; For the updated mechanistic model parameters; The parameters of the mechanistic model before the update; These are the parameters for the current exploration model.