Security reinforcement learning power distribution network optimization scheduling method based on historical data

Through a secure reinforcement learning method based on historical data, the slack indicators of distribution network intelligent agents are calculated and classified, and a suitable group of intelligent agents is selected for dispatching strategy output. This solves the problems of insecurity and low efficiency of reinforcement learning methods in the early stages of exploration, and realizes efficient and safe distribution network optimization dispatching.

CN120657853APending Publication Date: 2025-09-16STATE GRID JIANGSU ELECTRIC POWER CO LIANYUNGANG POWER SUPPLY CO +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510695596.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing reinforcement learning distribution network scheduling methods are unsafe and computationally inefficient in the initial exploration stage, with low learning efficiency, and are difficult to meet complexity and timeliness requirements.

Method used

By collecting and preprocessing the operating sequence data of distribution network intelligent agents, calculating and classifying the slack indicators, selecting the intelligent agent group that matches the current state pattern, and using the reinforcement learning optimization model based on historical data to output the scheduling strategy, the strategy is coordinated and scheduled by region and function.

Benefits of technology

It simplifies the decision space, improves computational and learning efficiency, reduces the complexity of reinforcement learning methods, and ensures the safe and stable operation of the distribution network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120657853A_ABST
    Figure CN120657853A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power distribution network scheduling methods, and particularly provides a safety reinforcement learning power distribution network optimal scheduling method based on historical data, which comprises the following steps of: acquiring operation time sequence data of a plurality of intelligent agents in a power distribution network in a current period of time, and preprocessing the operation time sequence data; calculating the relaxation index of each agent, and classifying the agents according to the similarity between the relaxation indexes of different agents; selecting one agent group with the relaxation index corresponding to the current state mode of the power distribution network system as a strategy agent group; inputting the operation time sequence data of the intelligent agents in the strategy intelligent agent group into a reinforcement learning optimization model based on historical data, and outputting a scheduling strategy corresponding to each intelligent agent; and carrying out coordinated scheduling on the power distribution network equipment according to the plurality of scheduling strategies. According to the safety reinforcement learning power distribution network optimization scheduling method based on historical data provided by the invention, the complexity of an intensity learning method of the power distribution network can be reduced, and the calculation efficiency and the learning efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network dispatching methods, and in particular to a distribution network optimization dispatching method based on secure reinforcement learning of historical data. Background Art

[0002] Optimal distribution network scheduling is a key component of power system operation and management, directly impacting the safe, stable operation and economic benefits of power systems. With the rapid development of smart grids, distribution network structures are becoming increasingly complex, and the uncertainty of distributed energy, renewable energy, and load demand continues to increase, posing significant challenges to traditional distribution network scheduling.

[0003] In recent years, reinforcement learning (RL) technology has been applied to distribution network scheduling due to its adaptability and model-independence. These methods learn through the interaction between an intelligent agent and its environment, eliminating the need for a precise system model and enabling autonomous discovery of optimization strategies through trial-and-error. However, existing RL-based methods for distribution network scheduling suffer from two major challenges. First, RL algorithms require numerous random attempts in the initial exploration phase, which can potentially place the distribution network in an unsafe state. Second, traditional RL methods suffer from high system complexity and computational complexity, resulting in low learning efficiency and slow convergence.

[0004] In view of the problems of high insecurity factors and low learning efficiency in the initial exploration of the strength learning method of the distribution network in related technologies, no effective solution has been proposed so far. Summary of the Invention

[0005] The invention provides a method for optimizing and dispatching a distribution network through a secure reinforcement learning method based on historical data, which at least solves the problems of high complexity, low computational efficiency, and low learning efficiency of the strength learning method of the distribution network.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] The present invention provides a safe reinforcement learning distribution network optimization and scheduling method based on historical data, including: collecting the operating time series data of multiple intelligent agents in the distribution network in the current period, and preprocessing them to obtain preprocessed operating time series data; calculating the relaxation index of each intelligent agent based on the preprocessed operating time series data, and classifying the intelligent agents according to the similarity between the relaxation indexes of different intelligent agents to obtain multiple intelligent agent groups; according to the current state mode of the distribution network system, selecting an intelligent agent group corresponding to the relaxation index and the state mode as the strategic intelligent agent group; inputting the operating time series data of the intelligent agents in the strategic intelligent agent group into a reinforcement learning optimization model based on historical data, and outputting the scheduling strategy corresponding to each intelligent agent through the reinforcement learning optimization model; dividing the multiple scheduling strategies based on the region and function of the distribution network, and coordinating and scheduling the distribution network equipment.

[0008] Preferably, the operating time series data of each intelligent agent in the distribution network in the current period is collected in real time and preprocessed, including: collecting the operating time series data of each intelligent agent in the distribution network in the current period from the distribution network system in real time; wherein the operating time series data include: voltage data, flow data, equipment operating status, dispatching instructions, load changes, and time series data of distributed energy output; identifying missing points and abnormal points in the operating time series data, and correcting the data of the missing points and abnormal points to obtain cleaned operating time series data; performing time series alignment and normalization transformation on the cleaned operating time series data to obtain standardized data; identifying key features that have a significant impact on the operating status and dispatching decisions of the distribution network, and selecting standardized data corresponding to the key features ranked high in importance as preprocessed operating time series data; wherein the key features include: node voltage, line flow, equipment status, power factor, load rate and voltage deviation.

[0009] Preferably, the relaxation index of each intelligent agent is calculated based on the preprocessed runtime data, including: extracting a control feature set that reflects the control capability of each area in the distribution network from the preprocessed runtime data; calculating the maximum controllable margin, minimum controllable margin and normalization coefficient of each intelligent agent in the current state based on the control feature set; calculating the difference between the maximum controllable margin and the minimum controllable margin of each intelligent agent in the current state, and taking the ratio of the difference to the normalization coefficient as the basic relaxation data of each intelligent agent; performing time series analysis and uncertainty assessment on the basic relaxation data to generate a corresponding relaxation index.

[0010] Preferably, a control feature set is extracted from the preprocessed operation sequence data, including: extracting control-related parameters related to the control of distribution network equipment from the preprocessed operation data; establishing a control feature set that reflects the control capabilities of each area of ​​the distribution network based on the control-related parameters and related factors; wherein the control-related parameters include: equipment adjustment range, control accuracy, response time; and related factors include: distribution network load characteristics, network topology, and operating status.

[0011] Preferably, the relaxation index of each intelligent agent is calculated based on the preprocessed runtime data, and the intelligent agents are classified according to the similarity between the relaxation indicators of different intelligent agents to obtain multiple intelligent agent groups, including: based on the similarity between the relaxation indicators of each intelligent agent, the intelligent agents with similar relaxation indicators are divided into intelligent agent groups of the same category to obtain multiple intelligent agent groups; the average relaxation of each intelligent agent group is sorted, and the state mode of the corresponding distribution network system is identified according to the average relaxation of each intelligent agent group, and the state mode includes: high-load weekdays, low-load weekends, and high renewable energy penetration.

[0012] Preferably, before the runtime data of the agents in the strategy agent group are input into the reinforcement learning optimization model based on historical data, the method includes: collecting the historical runtime data of each agent in each distribution network system, and preprocessing it to obtain the preprocessed historical runtime data; creating a learning framework for the distribution network based on the preprocessed historical runtime data; generating an action space and a complete reward function in the distribution network learning framework; establishing safety constraints based on the preprocessed historical runtime data, the action space and the complete reward function; based on the safety constraints combined with the comprehensive reward function, using the historical runtime data to pre-train the deep learning reinforcement model to obtain a reinforcement learning optimization model based on historical data.

[0013] Preferably, a distribution network learning framework is created based on the preprocessed historical runtime data, including: generating the state vector and control action vector of the intelligent agent based on the preprocessed historical runtime data; considering the impact of the control action on the state vector, establishing a state transfer equation that describes the change relationship between the state vector at the current moment and the state vector at the next moment; based on the state transfer equation, establishing an observation equation that describes the mapping relationship between the state vector and the observation data; each intelligent agent only obtains its own observation data, and based on the observation equation, updates the global state estimate of the distribution network through local observation data to obtain a learning framework for the distribution network.

[0014] Preferably, safety constraints are established based on preprocessed historical operating time data, action space and complete reward function, including: establishing hard constraints and soft constraints based on preprocessed historical operating time data; wherein, hard constraints include: voltage stability constraints, line thermal stability constraints, equipment operating capacity constraints and system topology constraints; soft constraints include: economic constraints, environmental constraints and equipment life constraints; direct constraints are imposed on the actions that may be performed by each intelligent agent in the distribution network according to the hard constraints; a corresponding penalty function is generated according to each type of soft constraint, and the complete reward function is corrected by the penalty function to obtain a comprehensive reward function; wherein, the comprehensive reward function is used to indirectly constrain the actions that may be performed by the intelligent agent.

[0015] Preferably, based on safety constraints combined with a comprehensive reward function, the deep learning reinforcement model is pre-trained using historical runtime data to obtain a reinforcement learning optimization model based on historical data, including: based on the preprocessed historical runtime data, effective decisions are screened to generate initial decisions; using the initial decisions, the deep reinforcement learning model is pre-trained to obtain a pre-trained deep reinforcement learning model; using safety constraints, the initial strategy is modified to obtain a modified strategy; using the modified strategy, the deep learning reinforcement model is reinforcement trained to obtain a reinforcement learning optimization model based on historical data.

[0016] Preferably, the plurality of dispatching strategies are divided based on the area and function of the distribution network, and the distribution network equipment is coordinated and dispatched, including: dividing the plurality of agents in the strategy agent group into multiple levels based on the area and function of the distribution network; wherein the multiple levels include: strategic layer, tactical layer and execution layer; using the dispatching strategy of the agent in the strategic layer to perform global coordination; using the dispatching strategy of the agent in the tactical layer to perform coordination and optimization within each area of ​​the distribution network; using the agent in the execution layer to obtain the dispatching strategy and perform the dispatching operation of each physical device; and through the collaboration of the agents in the strategic layer, tactical layer and execution layer, the distribution network equipment is coordinated and dispatched.

[0017] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0018] The present invention provides a method for optimizing and dispatching a distribution network based on a secure reinforcement learning based on historical data. By calculating the slack index of each agent in the distribution network, the agent's control flexibility can be quantified. The agent can be classified by the similarity between the slack indexes of different agents, and agents with similar control flexibility can be clustered into one category, thereby simplifying the decision space and balancing the requirements of control accuracy and computational efficiency. The present invention can select an agent group corresponding to the slack index and the state mode as the strategy agent group according to the current state mode of the distribution network system, so that the reinforcement learning optimization model based on historical data does not need to calculate the scheduling strategy corresponding to each runtime data, but only needs to call the scheduling strategy related to the current distribution network state mode, which not only improves the computational efficiency and learning efficiency, but also reduces the complexity of the intensity learning method of the distribution network. The present invention divides multiple scheduling strategies based on the region and function of the distribution network, and adopts a collaborative approach to coordinate and dispatch the distribution network equipment, which can reduce the complexity of the intensity learning method, thereby solving the problems of high complexity, low computational efficiency and low learning efficiency of the intensity learning method of the distribution network. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without inventive effort.

[0020] Figure 1 1 is a flow chart of a method for optimizing and dispatching a distribution network using secure reinforcement learning based on historical data according to an embodiment of the present invention;

[0021] Figure 2 is a schematic diagram of a process for obtaining a reinforcement learning optimization model based on historical data in an embodiment of the present invention;

[0022] Figure 3 is a schematic diagram of a flow chart for defining a distribution network learning framework in an embodiment of the present invention;

[0023] Figure 4 This is an embodiment of the present invention Figure 2 In this paper, based on safety constraints and comprehensive reward functions, the deep learning reinforcement model is trained using historical runtime data to obtain a flow chart of the reinforcement learning optimization model based on historical data. DETAILED DESCRIPTION

[0024] The following describes embodiments of the present invention in more detail with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0025] Traditional dispatching methods for distribution networks in the related art primarily include optimization methods based on mathematical programming and dispatching methods based on heuristic algorithms. Mathematical programming methods typically use techniques such as linear programming, nonlinear programming, or mixed integer programming to establish a distribution network model and obtain the optimal dispatching solution by solving the objective function. Heuristic algorithms, such as genetic algorithms and particle swarm optimization, use iterative searches to find a near-optimal solution. These methods can achieve good results in deterministic environments, but they struggle to effectively address the high uncertainty and complexity of distribution network environments.

[0026] Conventional reinforcement learning techniques, in the early stages of algorithm development, require extensive random trial-and-error operations, which can lead to frequent over-limit operation of distribution network systems and pose safety risks. Furthermore, reinforcement learning frameworks suffer from high model complexity and high computational resource consumption, resulting in low learning efficiency and slow convergence, making them difficult to meet the timeliness requirements of practical engineering applications.

[0027] like Figure 1 As shown, in order to reduce the complexity of the strength learning method of the distribution network and improve the computing efficiency and learning efficiency, the embodiment of the invention provides a distribution network optimization scheduling method based on historical data security reinforcement learning, including:

[0028] Step S1, collecting the running time data of multiple intelligent agents in the distribution network in the current period, and preprocessing them to obtain the preprocessed running time data;

[0029] Step S2, calculating the relaxation index of each agent based on the preprocessed runtime data, and classifying the agents according to the similarity between the relaxation indexes of different agents to obtain multiple agent groups;

[0030] Step S3, according to the current state mode of the distribution network system, selecting an agent group whose slack index corresponds to the state mode as the strategic agent group;

[0031] Step S4, inputting the running time data of the agents in the strategy agent group into a reinforcement learning optimization model based on historical data, and outputting a corresponding scheduling strategy through the reinforcement learning optimization model;

[0032] Step S5: Divide the multiple dispatching strategies based on the area and function of the distribution network, and coordinate and dispatch the distribution network equipment.

[0033] Among them, an intelligent agent refers to an entity with the ability to perceive, make decisions and act. It can perceive its environment, reason and make decisions based on its own goals and knowledge, and influence the environment by performing corresponding actions.

[0034] Intelligent agents in the distribution network can generally include: feeder terminal intelligent agents, distribution terminal intelligent agents, transformer intelligent agents, energy storage intelligent agents, renewable energy intelligent agents and load intelligent agents.

[0035] The embodiments of the present invention can quantify the control flexibility of intelligent agents by calculating the relaxation index of each intelligent agent in the distribution network, and classify the intelligent agents according to the similarity between the relaxation indexes of different intelligent agents. Intelligent agents with similar control flexibility can be clustered into one category, thereby simplifying the decision space and balancing the requirements of control accuracy and computational efficiency.

[0036] The present invention can select an agent group corresponding to the relaxation index and the state mode as the strategy agent group according to the current state mode of the distribution network system, so that the reinforcement learning optimization model based on historical data does not need to calculate the scheduling strategy corresponding to each running time series data, but only needs to call the scheduling strategy related to the current distribution network state mode. This not only improves the computing efficiency and learning efficiency, but also reduces the complexity of the strength learning method of the distribution network.

[0037] The present invention divides multiple dispatching strategies based on the area and function of the distribution network, and adopts a collaborative approach to coordinate the dispatching of distribution network equipment, which can reduce the complexity of the intensity learning method.

[0038] In a preferred but non-limiting embodiment of the present invention, step S1 includes: collecting operation sequence data of multiple intelligent agents in a current period of time from various systems of the power distribution network.

[0039] Among them, the various systems of the distribution network include: SCADA system, energy management system (EMS), distribution management system (DMS) and other systems.

[0040] The collected operating sequence data include: voltage data, flow data, equipment operating status, dispatch instructions, load changes, distributed energy output and other sequence data.

[0041] Furthermore, the preprocessing of the runtime data in step S1 includes:

[0042] Step S11, identifying missing points and abnormal points in the collected runtime data, and correcting the data of the missing points and abnormal points to obtain cleaned runtime data;

[0043] Step S12, performing time series alignment and normalization transformation on the cleaned runtime data to obtain normalized data;

[0044] Step S13: Identify key features that significantly affect the operating status and dispatching decisions of the distribution network, and select standardized data corresponding to the key features ranked high in importance as preprocessed operating sequence data; wherein the key features include: node voltage, line flow, equipment status, power factor, load rate and voltage deviation.

[0045] By cleaning and correcting the runtime data, the embodiments of the present invention can obtain complete and coherent runtime data, improve data quality, and help improve the accuracy of the reinforcement learning optimization model; by performing time series alignment and standardization transformation on the cleaned time series operation data, the correlation of the data can be enhanced, which is convenient for the reinforcement learning optimization model processing; by identifying the key features that have a significant impact on the distribution network operation status and scheduling decisions, and selecting the top-ranked key features, resources can be saved and the complexity of the model can be reduced.

[0046] In a preferred but non-limiting embodiment of the present invention, step S11 includes:

[0047] Step S111, using a sliding window algorithm to identify missing points in the runtime series data, and classifying them into two types according to the missing pattern: random missing and continuous missing;

[0048] Step S112: Using the moving Z-score and box plot method to identify possible outliers in the runtime series data, uniformly marking the identified missing points and outliers to generate a data set to be corrected;

[0049] Step S113: correct missing points and abnormal points to obtain cleaned runtime data.

[0050] Specifically, in step S113, different correction methods are used according to the missing type of the missing point, for example:

[0051] For the random missing type, the linear interpolation method was used to correct the data of the missing points;

[0052] For consecutive missing data, the historical similar day pattern substitution method is used to correct the data of the missing points;

[0053] In step S113, the data of the abnormal point is corrected using the power system operation law, such as using median filtering or correction based on a physical model.

[0054] In a preferred but non-limiting embodiment of the present invention, step S12 includes:

[0055] Step S121: Perform time series alignment on the cleaned data by using a unified timestamp and interpolation method to ensure that data from different sources are aligned in the time dimension.

[0056] In step S122 , the maximum-minimum standardization method or the Z-score standardization method is used to normalize all types of data to the same scale range, perform a standardization transformation, and avoid the influence of dimensional differences on subsequent analysis to obtain a standardized data set.

[0057] In a preferred but non-limiting embodiment of the present invention, step S13 includes:

[0058] Step S131: identifying key features that significantly impact the distribution network's operating status and dispatching decisions. Key features include: primary features such as node voltage, line flow, and equipment status, as well as derived features such as power factor, load factor, and voltage deviation.

[0059] Step S132: Using an autoencoder to perform nonlinear dimensionality reduction on the high-dimensional features in the initial feature set, mapping the high-dimensional features to a low-dimensional hidden layer relationship, and retaining the main features of the high-dimensional features in the hidden layer relationship to obtain the key features after dimensionality reduction;

[0060] Step S133, using the random forest feature importance analysis method to evaluate the importance of each key feature after dimensionality reduction to the distribution network operation status and dispatch decision;

[0061] Step S134 , selecting features ranked high in importance for combination, and determining an optimal feature subset through cross-validation, and using the standardized data corresponding to the optimal feature subset as pre-processed distribution network operation sequence data.

[0062] In a preferred but non-limiting embodiment of the present invention, calculating the relaxation index of each agent based on the preprocessed runtime data in step S2 includes:

[0063] Step S21, extracting a control feature set reflecting the control capability of each area in the distribution network from the preprocessed operation sequence data;

[0064] Step S22, extracting the maximum controllable margin, minimum controllable margin and normalization coefficient of each agent in the current state based on the control feature set;

[0065] Step S23: Calculate the difference between the maximum controllable margin and the minimum controllable margin of each agent in the current state, and take the ratio of the difference to the normalization coefficient as the basic relaxation data of each agent;

[0066] Step S24, performing time series analysis and uncertainty assessment on the basic slack data to generate a corresponding slack index;

[0067] Step S25 , based on the similarity between the relaxation indices of the various agents, the agents with similar relaxation indices are divided into agent groups of the same category to obtain a plurality of agent groups.

[0068] Specifically, step S21 includes:

[0069] Extracting control-related parameters related to distribution network equipment control from pre-processed operating data;

[0070] According to the control-related parameters and related factors, a control feature set reflecting the control capability of each area of ​​the distribution network is established.

[0071] Among them, the control-related parameters include: equipment adjustment range, control accuracy and response time; related factors include: distribution network load characteristics, network topology and operating status.

[0072] For example, for energy storage system management equipment, the load regulation range, control accuracy, and response time must be determined based on the load characteristics of the distribution network, and the load regulation range, control accuracy, and response time must be used as control characteristic parameters related to the load characteristics of the distribution network.

[0073] In a preferred but non-limiting embodiment of the present invention, in step S23, for different agents, the difference between the maximum adjustable margin and the minimum adjustable margin and the normalization coefficient are different, for example:

[0074] For the transformer control agent, the difference between the maximum adjustable margin and the minimum adjustable margin can be expressed as the transformer tap adjustment margin, and the normalization coefficient is the voltage deviation;

[0075] For the distributed power control agent, the difference between the maximum adjustable margin and the minimum adjustable margin can be expressed as the adjustable power range, and the normalization coefficient is the power demand.

[0076] Furthermore, step S24 includes:

[0077] Step S241 , performing time series analysis on the basic slack data, and taking the integral or weighted sum of the basic slack data in the future time window as the time-varying slack data;

[0078] Step S242 , defining uncertain factors, and calculating the expected value of the time-varying slack data under the influence of the uncertain factors as a slack index.

[0079] Among them, uncertain factors include: new energy output and load fluctuations.

[0080] When calculating the expected value of time-varying slack data under the influence of uncertain factors, first calculate the probability of the occurrence of the uncertain factors, then calculate the time-varying slack data value under the uncertain factors, and then integrate the product of the time-varying slack data value and the probability to obtain the expected value of the time-varying slack data under the influence of uncertain factors.

[0081] Furthermore, before step S25, the method includes:

[0082] Based on the relaxation index of each agent, the state similarity, spatial similarity and functional similarity between agents are calculated;

[0083] The state similarity, spatial similarity and functional similarity are weighted and summed to obtain the comprehensive similarity.

[0084] The calculation formula for the state similarity between agents i and j in state s is as follows:

[0085] SimL(i,j,s) = exp(-α·|Laxity(i,s) - Laxity(j,s)|²)

[0086] Where α is a tuning parameter used to control the similarity decay rate, which can be set according to the difference between the two agents; Laxity(i,s) represents the relaxation index of agent i in state s, Laxity(j,s) represents the relaxation index of agent j in state s; exp() represents the natural exponential function.

[0087] The calculation method of spatial similarity and functional similarity is the same as state similarity, except that the corresponding adjustment coefficients are different. Spatial similarity and functional similarity determine the adjustment coefficients between agents from the functional and spatial perspectives respectively.

[0088] Furthermore, if the comprehensive similarity is close to 1, the relaxation indices of the two agents are considered to be similar; if the comprehensive similarity is close to 0, the relaxation indices of the two agents are considered to be dissimilar.

[0089] Step S25 includes grouping the agents using a clustering algorithm based on the comprehensive similarity between the relaxation indicators of the agents, resulting in agent groups of different categories. The clustering algorithm can be spectral clustering or hierarchical clustering. During the clustering process, the optimal number of groups and grouping method can be adaptively determined, balancing decision simplification with control accuracy.

[0090] Based on the similarity between the relaxation indicators of various agents, agents with similar relaxation indicators are divided into agent groups of the same category to obtain multiple agent groups.

[0091] Furthermore, a relaxation threshold is set for each category of agent groups. If the relaxation in a certain agent group exceeds the threshold, a clustering algorithm is used to re-divide the agent group, thereby achieving dynamic clustering.

[0092] In a preferred but non-limiting embodiment of the present invention, step S3 comprises:

[0093] Take the average of the slackness indicators of all agents in each agent group as the average slackness of the agent group;

[0094] The average slack of multiple agent groups is sorted, and the corresponding state mode of the distribution network system is identified according to the average slack of the agent group. The state mode includes: high load weekdays, low load weekends, and high renewable energy penetration.

[0095] Specifically, for example: high average slack corresponds to low-load weekends; low average slack corresponds to high-load weekends; and medium average slack corresponds to high renewable energy penetration.

[0096] like Figure 2 As shown, in a preferred but non-limiting embodiment of the present invention, before step S4, the method includes:

[0097] Step S01: collecting historical running time series data of each intelligent agent in each distribution network system and preprocessing it to obtain preprocessed historical running time series data;

[0098] Step S02: creating a learning framework for the distribution network based on the pre-processed historical operation time series data;

[0099] Step S03: generating an action space and a complete reward function in a distribution network learning framework;

[0100] Step S04: establishing safety constraints based on the pre-processed historical runtime data, the action space, and the complete reward function;

[0101] Step S05: Based on the safety constraints and the comprehensive reward function, the deep learning reinforcement model is pre-trained using the historical operation data to obtain a reinforcement learning optimization model based on the historical data.

[0102] In step S01 , the collection and preprocessing methods of historical runtime data can refer to the collection and preprocessing methods in step S1 .

[0103] like Figure 3 As shown, further, step S02 includes:

[0104] Step S021, generating a state vector and a control action vector of the intelligent agent based on the preprocessed historical runtime data;

[0105] Step S022, considering the influence of the control action on the state vector, establishing a state transition equation that describes the change relationship between the state vector at the current moment and the state vector at the next moment;

[0106] Step S023: Based on the state transition equation, establish an observation equation that describes the mapping relationship between the state vector and the observation data;

[0107] In step S024, each intelligent agent only obtains its own observation data, and based on the observation equation, updates the global state estimation of the distribution network through local observation data, realizes the definition of the linear incomplete information game model, and obtains the learning framework of the distribution network.

[0108] The embodiment of the present invention formalizes the distribution network scheduling problem into a model suitable for reinforcement learning processing by introducing a linear incomplete information game framework, effectively solving the uncertainty and incomplete information problems in the distribution network environment and improving the model's fitting accuracy to the actual system.

[0109] Specifically, step S021 includes:

[0110] Based on the pre-processed historical runtime data, the state vector and control action vector of the intelligent agent are generated to obtain the state space and observation space that reflect the real-time state of the system;

[0111] The state space is defined as S = {s | s = [v, p, q, d, r, e]}, where v represents the node voltage vector, p represents the active power vector, q represents the reactive power vector, d represents the discrete device state vector (such as switch state and transformer gear), r represents the renewable energy output vector, and e represents the energy storage system state vector.

[0112] The observation space is defined as the set of local observation data available to each agent.

[0113] In step S022, the established state transition equation can be expressed as:

[0114] x(t+1) = Ax(t) + Bu(t) + w(t)

[0115] Where x(t) represents the state vector at time t; u(t) represents the control action vector, such as the transformer tap position and the controllable load adjustment; A and B are the state transfer matrix and the control matrix, respectively, which are determined through historical data or expert knowledge; w(t) is the system noise, which represents random factors such as load fluctuations and uncertainty in renewable energy output.

[0116] In step S023, the observation equation established can be expressed as:

[0117] yi(t) = Ci·x(t) + vi(t)

[0118] Among them, yi(t) represents the observation value of agent i at time t, vi(t) represents the observation noise of agent i at time t, and Ci represents the observation matrix.

[0119] In step S024, it is defined that each intelligent agent can only obtain its own observation data and cannot directly access the information structure of the global state vector; and a belief update mechanism is defined based on the observation equation, in which each intelligent agent updates the global state estimation of the distribution network using the local observation data obtained.

[0120] Specifically, in step S024, the global state estimation is achieved by using a Bayesian update mechanism combined with state estimation techniques such as Kalman filtering.

[0121] By combining the state transfer equation, observation equation, information structure and belief updating mechanism, the definition of linear incomplete information game model is obtained, thus obtaining the learning framework of distribution network.

[0122] The embodiments of the present invention clarify the information sharing boundaries between intelligent agents by designing the information structure, and combine it with the belief updating mechanism to enable the multi-agent system in the distribution network to effectively collaborate under the condition of incomplete information, reduce the impact of information uncertainty on decision-making quality, and improve the overall performance of the system.

[0123] In a preferred but non-limiting embodiment of the present invention, step S03 includes:

[0124] Step S031: In the distribution network learning framework, controllable devices and control parameters are identified, and an initial action set is established;

[0125] Step S032: discretize and constrain the initial action set to form a standardized action space;

[0126] Step S033: constructing reward evaluation indicators based on the action space; wherein the reward evaluation indicators include: loss, voltage and cost;

[0127] Step S034: normalize and weight the reward evaluation indicators to generate a complete reward function.

[0128] Specifically, in step S031, the distribution network learning framework model is used to identify controllable devices and control parameters in the system, including transformer tap adjustment, capacitor bank switching, controllable load regulation, and distributed energy output control. For each device type, its control characteristics, adjustment range, and operational constraints are analyzed, and an initial action set containing all possible actions is established.

[0129] Step S032 includes:

[0130] The initial set of actions is discretized and constrained. For example, the energy storage charge and discharge power range [-Pmax, Pmax] is discretized into multiple discrete values, such as {-Pmax, -0.8Pmax, ..., 0, ..., 0.8Pmax, Pmax}. Taking into account the physical characteristics and operational constraints of distribution network equipment, time constraints are introduced, such as the minimum interval for transformer tap adjustment and the rate limit for energy storage charge and discharge.

[0131] According to the distribution network dispatching requirements, the set of dispatching actions that the intelligent agent can perform is clarified, and the action space is defined. The action space includes discrete and continuous actions such as transformer tap adjustment, capacitor bank switching, controllable load regulation, and distributed energy output control.

[0132] For discrete actions (such as switch operation and transformer tap adjustment), the finite set Ad = {a1, a2,..., an} is defined. For continuous actions (such as power regulation), the bounded interval Ac = [amin, amax] is defined, and appropriate discretization is performed, such as discretizing the energy storage charge and discharge power interval into multiple discrete values. Simultaneously, considering the physical characteristics of the equipment and operational constraints, and introducing time constraints, the resulting normalized hybrid action space A = Ad × Ac is formed. This is the Cartesian product of the discrete action space and the continuous action space, forming a structured high-dimensional action space.

[0133] In step S033, based on the normalized action space, a multi-objective reward evaluation index is constructed, which mainly includes: grid loss item, voltage deviation item, regulation cost item, equipment life impact item and renewable energy consumption item.

[0134] Among them, the grid loss term calculates the active power loss of the system, which is expressed as R loss = -w loss ·∑(I i 2 ·R i ), where I i is the current in line i, R i is the line resistance, w loss is the weight coefficient.

[0135] The voltage deviation term evaluates the degree to which the node voltage deviates from the nominal value and is represented by R voltage = -w voltage ·∑(|V j - V nom | / V nom ) 2 , where V j is the voltage at node j, V nom is the rated voltage, wvoltage is the weight coefficient.

[0136] The control cost item considers the economic cost of executing the dispatch action, including equipment operating costs, energy storage charging and discharging costs, etc., and is expressed as R cost = -w cost ∑C(ak), where C(ak) is the cost function for performing action ak, and w cost is the weight coefficient.

[0137] The equipment life impact item evaluates the impact of scheduling actions on the service life of equipment, especially the wear of mechanical parts caused by frequent operations, expressed as R life = -w life ∑D(ak), where D(ak) is the effect of action ak on the device life, w life is the weight coefficient.

[0138] The renewable energy consumption term encourages the system to maximize the use of renewable energy, which is expressed as

[0139] R renewable = w renewable ∑P renused / ∑P renavailable

[0140] Step S034 includes:

[0141] Normalize the constructed reward evaluation indicators to solve the dimensional differences between different indicators;

[0142] Sum up all reward evaluation indicators to form the complete reward function as follows:

[0143] R = R loss + R voltage + R cost + R life + R renewable .

[0144] In a preferred but non-limiting embodiment of the present invention, step S04 includes:

[0145] Step S041: Establishing hard constraints and soft constraints based on the pre-processed historical operating sequence data; wherein the hard constraints include: voltage stability constraints, line thermal stability constraints, equipment operating capacity constraints, and system topology constraints; and the soft constraints include: economic constraints, environmental constraints, and equipment life constraints;

[0146] Step S042, directly constraining the actions that each intelligent agent in the distribution network may perform according to the hard constraints;

[0147] In step S043, a corresponding penalty function is generated according to each type of soft constraint, and the complete reward function is modified by the penalty function to obtain a comprehensive reward function, wherein the comprehensive reward function is used to indirectly constrain the actions that each intelligent agent may perform.

[0148] The safety constraints of the embodiment of the present invention integrate hard safety constraints such as voltage stability and line capacity and soft constraints such as economy into the learning process through action masks and reward correction methods, respectively, ensuring the safety of the reinforcement learning agent during the optimization scheduling process.

[0149] Specifically, in step S041, the voltage stability constraint requires that the voltage of each node be maintained within a safe range. For example, in conventional distribution networks, the node voltage deviation is usually required to not exceed ±7%, which is expressed as: Vmin ≤ Vi ≤ Vmax; where Vi is the voltage amplitude of node i, and Vmin and Vmax are the minimum and maximum allowable voltage values, respectively.

[0150] The line thermal stability constraint restricts the line power to not exceed the thermal stability limit to avoid line overheating damage, which is expressed as: |Sj| ≤ Sjmax, where Sj is the apparent power of line j and Sjmax is the maximum allowable load of the line.

[0151] Equipment capacity constraints ensure that the operating parameters of various equipment, such as transformers and capacitor banks, are within their rated ranges and are expressed as a set of inequalities. System topology constraints ensure that the distribution network maintains reasonable network connectivity and a radial structure.

[0152] Among soft constraints, economic constraints seek to minimize system operating costs, including grid loss costs and dispatch operation costs. Environmental constraints focus on carbon emissions and renewable energy consumption rates, promoting green and low-carbon operations. Equipment life constraints consider the impact of dispatch decisions on equipment aging, preventing premature wear caused by frequent operation.

[0153] Furthermore, in step S042, based on the power system stability theory, a mathematical model of safety indicators such as voltage stability margin and line load rate is constructed, and a safety threshold is set.

[0154] Specifically, the voltage stability margin model assesses system voltage stability based on the eigenvalues ​​of the power flow Jacobian matrix. The voltage stability margin (VSM) is defined as (λmin - λcritical) / λcritical, where λmin is the minimum eigenvalue of the power flow Jacobian matrix and λcritical is the eigenvalue corresponding to the critical stability point. When the VSM falls below a preset threshold, the system approaches voltage instability and corrective measures are required.

[0155] The line load rate model is used to assess the thermal stability of a line. The load rate metric, LLR (Line Loading Rate), is defined as |Sactual| / Srated, where Sactual is the actual line load and Srated is the rated capacity. To prevent prolonged line overload, different thresholds are set, such as LLRwarning = 0.85 and LLRemergency = 0.95, corresponding to warning and emergency states, respectively. Furthermore, a time-dependent overload tolerance model is established to consider the line's transient overload capacity, allowing for moderate overloads exceeding the rated capacity for short periods of time.

[0156] For transformer equipment, a transformer load capacity assessment model, including a thermal model, was constructed. Based on IEC standards, a dynamic prediction model for the transformer's top oil temperature and winding hotspot temperature was established to determine the transformer's real-time load capacity and prevent thermal damage.

[0157] In terms of system topology constraints, a distribution network topology verification model based on graph theory is constructed to ensure that the network maintains a reasonable radial structure and avoid protection coordination issues caused by ring network operation. Network loops are detected using a depth-first search algorithm, and when an unallowed loop structure is detected, a topology constraint violation signal is generated.

[0158] Constraint integration mechanism design: Design a constraint integration mechanism to integrate hard constraints into the reinforcement learning algorithm through action masks or safety checkers, thereby limiting the possibility of the agent choosing unsafe actions.

[0159] Specifically, hard constraints directly limit the agent's range of action through action masking to ensure that the system does not enter an unsafe state.

[0160] Specifically, in each decision step, a feasible action mask M(s) is dynamically generated according to the current system state and safety constraints, marking the actions that may lead to violation of hard constraints as unavailable.

[0161] For complex constraints that are computationally complex or difficult to express precisely, a Safety Checker module is designed. Positioned between the agent's decision output and its interaction with the environment, the Safety Checker performs real-time safety assessments on the agent's selected actions. The checker comprises two levels of assessment: rapid assessment, based on a simplified model, to flag actions that may violate constraints. The precise assessment, through detailed power flow calculations or state predictions, makes a final safety determination on flagged actions. If an action is deemed unsafe, the Safety Checker replaces it with the nearest safe alternative or triggers a safety fallback strategy.

[0162] To balance hard constraints and exploration efficiency, we introduce a Constrained Exploration Strategy (CES), which restricts exploration to a safe action subspace and ensures the safety of the exploration process through techniques such as directed exploration or Gaussian process constrained exploration.

[0163] Furthermore, in step S043, a corresponding penalty function pi(s,a,s') is designed for each type of soft constraint to quantify the degree to which the state transition (s,a,s') violates the soft constraint. For example, the penalty function for the economic soft constraint can be expressed as pcost(s,a,s') = wcost (C(s,a,s') - Cref) / Cref. Where C(s,a,s') is the operating cost caused by action a, Cref is the reference cost baseline, and wcost is the weight coefficient.

[0164] Based on each penalty function, a modified reward function R'(s,a,s') = R(s,a,s') - ∑pi(s,a,s') is constructed, where R(s,a,s') is the original reward function, reflecting the main goal of distribution network optimization, such as minimizing grid losses.

[0165] Through these penalty functions, we can obtain a complete reward correction rule. Combining the safe action filter and reward correction rule forms a comprehensive safety constraint processing mechanism.

[0166] To balance the importance of different soft constraints, an adaptive weight adjustment mechanism is employed to dynamically adjust the weights of each penalty function based on the system's operating stage, the severity of constraint violations, and historical trends. For example, when the system is under high load for an extended period, the weight of the equipment life constraint is increased, while the impact of the economic constraint is reduced to protect the equipment from excessive wear.

[0167] Design a reward structure based on a safety margin. When the system state is far from the constraint boundary, the penalty intensity is reduced to encourage the agent to actively optimize the primary objective; when it approaches the constraint boundary, the penalty intensity is increased to guide the agent to maintain a safety margin.

[0168] Safety fallback strategy: Design an emergency fallback mechanism. When the system approaches a dangerous state, the preset safety fallback strategy is automatically activated to ensure system safety.

[0169] This embodiment of the present invention combines a safe action filter with reward correction rules to construct comprehensive safety constraints. Hard constraints directly limit the agent's range of action through action masks, ensuring that the system does not enter an unsafe state. Soft constraints indirectly guide the agent to learn safer and more efficient strategies by modifying the reward function R'(s,a,s') = R(s,a,s') - ∑pi(s,a,s'). Furthermore, a safety fallback strategy is designed to automatically initiate pre-set safety measures when the system approaches a dangerous state. This creates a multi-layered safety assurance mechanism, resulting in a learning system with safety constraints.

[0170] Specifically, a system safety status assessment mechanism is established, defining a multi-level safety status indicator (SSI). The SSI value categorizes the system status into four levels: normal, caution, warning, and danger. When the system enters the warning or danger level, a safety fallback mechanism is triggered.

[0171] The safety fallback strategy is divided into three levels: mild intervention, moderate intervention, and forced fallback. Mild intervention is activated when the system enters the attention state. It mainly modifies the agent's action selection probability distribution to increase the probability of selecting safe actions while still allowing the agent to maintain a certain degree of decision-making autonomy.

[0172] Moderate intervention is activated when the system enters a warning state. It temporarily takes over some control and executes a pre-defined sequence of safe recovery actions while maintaining normal control of non-critical variables. For example, if voltage approaches a limit, the moderate intervention mechanism might take over the reactive power control device and execute an action sequence to increase voltage while allowing the agent to continue controlling active power-related devices.

[0173] Forced fallback is activated when the system enters a dangerous state, completely taking over system control and executing the most conservative safety recovery strategy to quickly return the system to a safe zone. These pre-set safety fallback strategies are designed based on expert knowledge and historical emergency response experience, and their effectiveness and robustness are verified through offline simulation.

[0174] To prevent system control instability caused by frequent fallback triggering, state hysteresis and time delay mechanisms are introduced. The system state must remain at a certain safety level for a certain period of time before the corresponding intervention measures are triggered or released, preventing frequent switching caused by instantaneous fluctuations.

[0175] like Figure 4 As shown, in a preferred but non-limiting embodiment of the present invention, step S05 includes:

[0176] Step S051, based on the pre-processed historical runtime data, screening effective decisions and generating initial decisions;

[0177] Step S052: pre-training the deep reinforcement learning model using the initial decision to obtain a pre-trained deep reinforcement learning model;

[0178] Step S053, using the security constraint, modifying the initial policy to obtain a modified policy;

[0179] Step S054: Use the modified strategy to perform reinforcement training on the deep learning reinforcement model to obtain a reinforcement learning optimization model based on historical data.

[0180] By extracting effective decision patterns from historical scheduling data as initial strategies, this embodiment of the present invention significantly accelerates the convergence of reinforcement learning models, addressing the inefficiency caused by random exploration in related reinforcement learning techniques. Furthermore, by introducing safety constraints, the security of the reinforcement learning agent during the optimization scheduling process is ensured.

[0181] Specifically, step S051 includes:

[0182] Step S0511, calculate the relaxation index of each agent, classify the agents according to the similarity between the relaxation indexes of different agents, and obtain multiple clustered agent groups.

[0183] Step S0512, based on the similar relaxation of each clustered agent group, sort the relaxation of multiple clustered agent groups, and identify the corresponding state mode of the distribution network system according to the relaxation level. The state mode includes: high-load weekdays, low-load weekends, and high renewable energy penetration.

[0184] Step S0513, based on the similar relaxation of each clustered agent group, the state pattern is matched with the agent group, and a plurality of state-action pairs are generated according to the possible actions of each agent in the agent group;

[0185] Step S0514, calculate the probability of each state-action pair and its effect on improving the performance of the distribution network, select high-frequency and effective state-action pairs as effective decisions, and combine the effective decisions to generate a decision rule library.

[0186] Step S0515: convert the rules in the decision rule base into the initial strategy of the intelligent agent, that is, the probability distribution of selecting action a in state s.

[0187] The method for calculating the relaxation index and classifying the agents in step S0511 is the same as that in step S2. The method for executing step S0512 is the same as that in step S3 and will not be repeated here.

[0188] Specifically, in step S0513, conditional probability analysis is used to calculate the conditional probability P(a|s) of executing action a under state mode s, identifying high-frequency and effective state-action pairs. Simultaneously, the effectiveness of each decision is evaluated, and by constructing a decision-outcome association matrix, effective decisions that lead to improved system performance are identified.

[0189] To capture the temporal nature of decision-making, sequential pattern mining techniques (such as the PrefixSpan algorithm) are applied to extract frequent decision sequence patterns from continuous time series and identify multi-step decision strategies. After evaluation and screening, the extracted decision patterns are organized into a decision rule base, expressed as follows: If the system state satisfies condition C, the probability of executing action A is P.

[0190] Further, after step S052, the method includes:

[0191] Using expert knowledge rules and prior probability distribution, the initial decision is modified to obtain a modified decision;

[0192] Use the modified policy to perform reinforcement training on the pre-trained deep reinforcement learning model;

[0193] Specifically, expert knowledge rules are rules set by power system operation and dispatch experts and expressed as "if-then" rules. For example, if the voltage at a certain node is lower than 0.95 scalar power, reactive power compensation should be increased in that area.

[0194] These rules are encoded into a computer-processable form using a rule editor and are categorized as hard rules and soft rules. Hard rules must be strictly adhered to and typically involve system safety constraints, such as "line load must not exceed 90% of rated capacity." Soft rules are empirical recommendations that can be flexibly applied in specific circumstances.

[0195] Prior probability distributions are a priori estimates of the probability of system state transitions or optimal action selections. Through interaction with domain experts, Bayesian prior distributions are established, such as the prior distribution Pprior(a|s) for optimal actions under different states. These prior distributions are combined with reinforcement learning algorithms through a Bayesian inference framework to influence the direction and speed of the agent's exploration.

[0196] At the same time, a knowledge verification and updating mechanism is designed. When new data obtained during system operation significantly conflicts with expert knowledge, confidence assessment and contradiction detection are used to identify potentially outdated or inapplicable knowledge and guide experts to update their knowledge.

[0197] In step S053, the agent’s strategy π(a|s) is modified using the safety constraint mechanism to π'(a|s) = π(a|s)·M(s,a) / ∑a'[π(a'|s)·M(s,a')], where M(s,a) is a binary mask function that takes the value of 0 for unsafe actions and the value of 1 for safe actions, thereby forming a safe action filter.

[0198] In step S0521, a suitable reinforcement learning algorithm framework is first selected, such as the Deep Q Network (DQN), Proximal Policy Optimization (PPO) or Soft Actor-Critic (SAC) algorithm, and customized optimization is performed based on the characteristics of the distribution network.

[0199] Specifically, the algorithm selection was based on the characteristics of the distribution network scheduling problem. Considering that distribution network scheduling involves both discrete and continuous actions, it needs to handle a mixed action space. The system state space is high-dimensional, requiring good generalization and sample efficiency. Furthermore, safety constraints and multi-objective optimization require the algorithm to be stable and adaptable. Taking these factors into consideration, algorithms based on the actor-critic architecture were selected as the foundational framework, such as proximal policy optimization (PPO), soft actor-critic (SAC), or trust region policy optimization (TRPO).

[0200] A specialized neural network structure was customized as a function approximator for the high-dimensional state space of distribution networks. A graph neural network (GNN) combined with an attention mechanism was used as the underlying architecture to effectively capture the node relationships within the distribution network topology. Each node (such as a bus or device) is represented as a node in the GNN, with characteristics including physical quantities such as voltage and power. Physical connections between nodes (such as lines) are represented as edges in the graph, with characteristics including impedance, power flow, and other information.

[0201] To handle mixed action spaces, a dual-head output network structure was designed: the discrete action head uses a Softmax layer to output the probability distribution of discrete actions; the continuous action head outputs the mean and standard deviation of a Gaussian distribution, representing the distribution of the continuous control variable. The two output heads share feature extraction layers but have independent final output layers.

[0202] Distribution network scheduling is highly temporal. Long short-term memory (LSTM) layers or transformer modules are integrated into the network structure to enable the model to memorize and utilize historical information. Furthermore, a multi-timescale learning framework is designed to simultaneously train value functions and policy networks for both short time steps (e.g., minutes) and long time steps (e.g., hours), capturing optimal policies at different time scales.

[0203] Using the constructed initial policy library, a deep reinforcement learning model is pre-trained using imitation learning. State-action pairs extracted from historical data are used as expert demonstrations, and a supervised learning method is used to train the policy network to minimize the discrepancy between predicted actions and historical actions. Generative adversarial imitation learning is used to improve the quality of imitation. A discriminator network is introduced to distinguish between agent-generated trajectories and historical expert trajectories. Through adversarial training, the agent learns long-term policy logic, resulting in a pre-trained deep reinforcement learning model.

[0204] Perform reinforcement learning training on the pre-trained model. In the initial stage, it mainly relies on historical data for guidance, and gradually increases the reinforcement learning weight as the training progresses.

[0205] The initial policy is leveraged through imitation learning or pre-training to accelerate the reinforcement learning exploration process, achieving adaptive exploration based on uncertainty. The entropy of the action distribution output by the policy network is used as a measure of uncertainty. In areas with good historical data coverage, the model outputs a low-entropy (high-certainty) distribution; in areas with sparse data, it outputs a high-entropy distribution. During training, the exploration rate ε is dynamically adjusted to maintain a positive correlation with policy uncertainty, encouraging more exploration in areas with sparse data.

[0206] A reward-based data filtering mechanism is introduced, using a weighted, revised initial policy for reinforcement learning. A quality score is assigned to each state-action pair in the historical data using an environment model or expert evaluation function. In imitation learning, the importance of samples is weighted according to the quality score, prioritizing high-quality decisions and minimizing the imitation of suboptimal ones.

[0207] Safety constraints are introduced during the training process, and the initial strategy is modified using weighted corrections using multiple safety constraint mechanisms. The modified strategy is used to further strengthen learning to ensure that the intelligent agent does not make dangerous decisions during the exploration process. At the same time, the reward function is used to guide the learning of safety strategies.

[0208] Specifically, we implement training based on a constrained Markov decision process, using a Lagrangian approach to transform constrained optimization into a saddle point problem. We employ a progressive constraint tightening strategy, while integrating a safety layer to ensure action safety. We continuously optimize the policy through interaction with the environment, generating an optimized policy network.

[0209] Furthermore, robust training techniques have been incorporated. Through domain randomization, environmental parameters (such as load forecast errors and line parameter uncertainties) are randomly perturbed during training, enabling the agent to learn robust strategies that are robust against worst-case scenarios. Furthermore, quantile regression techniques are employed to estimate the distribution of the value function, rather than just the expected value. This allows the decision-making process to account for tail risks and enhances safety in extreme situations.

[0210] Furthermore, the present invention employs dual-timescale updating and policy locking techniques to address the challenges of non-stationary learning in multi-agent environments. Dual-timescale updating enables different agents to update their policies at different frequencies, reducing oscillations caused by synchronized updates. Policy locking, on the other hand, fixes the policies of some agents during each training cycle, allowing other agents to learn in a relatively stable environment. The locked groups of agents are then rotated to gradually achieve global collaborative optimization.

[0211] In a preferred but non-limiting embodiment of the present invention, after obtaining the reinforcement learning optimization model based on historical data, the method further includes, step S055, performing experience replay and parameter tuning based on the optimization strategy network to form an improved model.

[0212] Specifically, experience replay and parameter tuning are implemented based on the optimized policy network. An experience replay mechanism is used to reduce sample correlation, storing interactively collected experiences in a replay buffer and randomly sampling them during training. Prioritized experience replay is implemented, assigning priorities to experiences based on the time-delay error. A target network is integrated to mitigate overestimation of the value function, slowly updating the target network through soft updates. Gradient optimization techniques, such as generalized advantage estimation, are applied to reduce variance, enabling fine-tuning of parameters and ultimately forming an improved model.

[0213] After obtaining the improved model, the method includes: step S056, performing security verification and performance evaluation on the improved model, and outputting an optimized scheduling strategy.

[0214] Specifically, the improved model undergoes comprehensive safety verification and performance evaluation. Multi-dimensional evaluation metrics, including power quality, economic efficiency, safety margin, and algorithm convergence, are designed, and model performance is tested on an independent validation dataset. The model's effectiveness and superiority are verified by comparison with traditional scheduling methods and standard reinforcement learning methods. Particular attention is paid to the model's performance under extreme operating conditions and emergencies to ensure system safety and reliability. Necessary adjustments and optimizations are made based on the evaluation results, ultimately outputting an optimized scheduling strategy for use in actual distribution network optimization scheduling operations.

[0215] Step S056 includes:

[0216] Step S0561: Collect interaction data from the training process and create an experience playback buffer;

[0217] Step S0562, performing priority evaluation on the data in the experience replay buffer and generating sampling weights;

[0218] Step S0563: sampling empirical data based on sampling weights to form a training batch;

[0219] Step S0564, using the training batch to update the policy network parameters to obtain a phased model;

[0220] Step S0565, verify the phased model and output the improved training model.

[0221] Specifically, in step S0561, during reinforcement learning training, the system continuously collects data generated by the agent's interaction with the environment, including information such as state, action, reward, and next state (s, a, r, s'). This data is organized in time series and stored in an experience replay buffer. The buffer is designed as a fixed-capacity circular queue. When the buffer is full, new experience replaces the oldest stored experience, ensuring data timeliness.

[0222] In step S0562, the data in the experience replay buffer is prioritized, and each experience is assigned a priority based on the TD error |r +γV(s') - V(s)| or other importance metrics. Experiences with high errors indicate that the current value function prediction is inaccurate or has sudden changes, and therefore contain more valuable learning information. Sampling probabilities are calculated based on the priorities, generating a normalized sampling weight distribution.

[0223] In step S0563, based on the generated sampling weights, weighted sampling is performed from the experience replay buffer, prioritizing high-weighted experiences. The sampling process employs randomness to ensure that even low-weighted experiences have a certain probability of being selected, avoiding completely ignoring certain state spaces. The sampled experiences are organized into training batches, with batch sizes typically ranging from 32 to 256 to balance computational efficiency and learning stability.

[0224] In step S0564, the sampled training batches are used to update the parameters of the policy network and value network via the backpropagation algorithm. For algorithms based on policy gradients, the policy gradient is calculated and gradient ascent is applied; for algorithms based on value functions, the TD error is calculated and the loss function is minimized. Techniques such as gradient clipping and batch normalization are used during the update process to improve training stability. After a certain number of update iterations, the model parameters for this phase are obtained.

[0225] In step S0565, the performance of the phased model is verified by executing a certain number of evaluation rounds in the verification environment and collecting performance metrics such as average reward, success rate, and number of security breaches. If performance significantly improves compared to the previous model, the current model parameters are saved. If performance stagnates or declines, it may trigger adjustments to the learning rate or exploration strategy. Based on the verification results, a decision is made on whether to continue training or adjust hyperparameters, ultimately outputting an improved trained model that meets the performance requirements.

[0226] After obtaining the improved training model, the model is evaluated and deployed.

[0227] Specifically, it includes: the first step, multi-dimensional evaluation index design: design a multi-dimensional evaluation index system including power quality, economy, safety margin, algorithm convergence, etc. to comprehensively evaluate the performance of the reinforcement learning model.

[0228] The second step is comparative verification testing: Using the verification set data, the performance differences between the reinforcement learning model and the traditional scheduling method are compared to verify the effectiveness and superiority of the model.

[0229] The third step is model deployment architecture design: design a model deployment architecture suitable for the actual distribution network environment, including components such as data interface, decision engine, safety checker, and human-computer interaction interface.

[0230] Step 4: Online learning and adaptation mechanism: Implement an online learning mechanism based on real-time data so that the model can continuously learn from new operating data and adapt to environmental changes.

[0231] Step 5. Abnormal situation handling mechanism: Design an abnormal situation handling mechanism for extreme conditions and emergencies to ensure that the system can still make safe and reliable scheduling decisions under abnormal conditions.

[0232] In a preferred but non-limiting embodiment of the present invention, step S5 comprises:

[0233] Step S51, dividing each agent in the strategic agent group into multiple levels based on the area and function of the distribution network, including: a strategic level, a tactical level, and an execution level;

[0234] Step S52, performing global coordination through the scheduling strategy of the intelligent agent at the strategic layer;

[0235] Step S53, performing coordination and optimization in each area of ​​the distribution network through the dispatching strategy of the intelligent agent at the tactical layer;

[0236] Step S54, executing the scheduling strategy of the agent in the execution layer to execute the scheduling operation of each physical device, thereby obtaining a multi-agent collaboration mechanism;

[0237] Step S55: Coordinate and dispatch the distribution network equipment through the collaboration of the intelligent agents in the strategic layer, tactical layer and execution layer.

[0238] Specifically, in step S51, it includes:

[0239] Step S511: establish a communication network among intelligent agents using the physical topology of the distribution network, determine which intelligent agents can communicate directly with each other, and determine the frequency and priority of the communication.

[0240] Step S512: Analyze the dependency relationship between the various intelligent agents and generate an interaction strength matrix.

[0241] This step constructs an agent dependency graph (ADG) to quantify the degree of mutual influence between the decisions of the agents. For example, the dependency strength between the transformer agent and its connected capacitor agent in the same substation will be high because their decisions directly affect each other.

[0242] Step S513: designing a hierarchical collaboration strategy based on the interaction intensity matrix, dividing each agent in the strategic agent group into a strategic layer, a tactical layer, and an execution layer.

[0243] Specifically, in step S511, an inter-agent communication network is constructed based on the physical topology of the distribution network. This network considers physical distance, communication bandwidth, and latency, establishing direct communication links between physically adjacent or functionally closely related agents. Simultaneously, a communication protocol and message format are designed, and the content, frequency, and priority of information exchange are determined, forming an initial communication structure that supports multi-agent collaboration.

[0244] In step S512, the Networked Markov Decision Process (NMDP) framework is used to describe the interaction between agents. The formal definition is: NMDP =<N, S, A, T, R, O,γ> , where N is the set of agents, S is the global state space, A = ×Ai is the joint action space, T is the state transition function, R is the reward function, O defines the observation function of each agent (in the case of partial observability), and γ is the discount factor.

[0245] Four main types of agent interactions are identified and modeled: direct electrical coupling, resource competition, goal coordination, and information sharing. Direct electrical coupling stems from physical network connections, resource competition occurs when multiple agents compete for limited resources, goal coordination occurs when agents adjust their behaviors to achieve a global goal, and information sharing involves the exchange of observations and decision intentions between agents.

[0246] Build an Agent Dependency Graph (ADG). The ADG is a weighted directed graph G = (V, E, W), where the vertex set V corresponds to the set of agents, the edge set E represents the dependencies between agents, and the weight set W quantifies the strength of the dependencies. Edge weights are calculated based on the interaction type and the laxity metric, reflecting the degree of mutual influence between agent decisions.

[0247] In step S513, a hierarchical collaboration strategy is designed based on the interaction intensity matrix. Agents are organized into a three-tiered structure based on interaction intensity: strategic, tactical, and execution. Horizontal collaboration mechanisms are designed between agents on the same tier, such as auction-based resource allocation and consensus-based decision-making. Vertical coordination mechanisms are designed between tiers, such as a hierarchical coordination framework based on model predictive control. Communication rules, decision-making sequences, and conflict resolution methods are developed during the collaboration process to form a complete set of collaboration rules.

[0248] Specifically, the strategic layer is at the top, the tactical layer is at the middle, and the execution layer is at the bottom.

[0249] The central agent with central coordination capabilities is used as the strategic layer agent;

[0250] Based on the feeder structure of the distribution network, the power supply range of the substation, and the location of the tie switch, the network is divided into multiple relatively independent power supply areas. Each area corresponds to a regional agent, which serves as a tactical layer agent.

[0251] Based on the control objectives and equipment types, different types of functional agents are set as execution layer agents, such as transformer agents, energy storage agents, renewable energy agents, and load agents.

[0252] After step S513 , the method further includes: step S514 , implementing an information sharing mechanism based on the collaboration rule set and establishing a collaborative decision-making framework.

[0253] Specifically, based on a set of collaborative rules, an information sharing mechanism is implemented among agents. A value-based information filtering strategy is designed to evaluate the value of information to decision-making: V(info) = E[U(π(b⊕info))] - E[U(π(b))], where b is the current belief state, b⊕info is the updated belief state after receiving the information, π is the decision-making strategy, U is the utility function, and E represents the expectation.

[0254] Prioritize the transmission of high-value information. Implement distributed state estimation, enabling each agent to form a more accurate estimate of the global state based on local observations and information shared by neighbors. Establish a common belief update mechanism to coordinate the decision-making processes of each agent and form a consistent collaborative decision-making framework.

[0255] Furthermore, step S55 includes:

[0256] Agents on the same layer coordinate and schedule using horizontal collaboration, which can be auction-based or consensus-based.

[0257] Agents at different levels coordinate and schedule using a vertical coordination approach: the upper level transmits goals and constraints to the lower level, and the lower level feeds back execution results and status information to the upper level. Specifically, a hierarchical coordination framework based on model predictive control (MPC) is adopted: the strategic level runs low-frequency, long-term MPC to generate global optimization trajectories and key node constraints; the tactical level runs medium-frequency, medium-term MPC to optimize regional resource allocation under global constraints; and the execution level runs high-frequency, short-term control algorithms, such as PID or local reinforcement learning, to achieve precise control of equipment.

[0258] Establish an information sharing mechanism and collaborative decision-making framework, and integrate the entire collaborative mechanism with specific control strategies.

[0259] Integrate the collaborative decision-making framework with the control strategies of each agent to form a closed-loop multi-agent collaborative system. Design interface standards to ensure that decision information can be converted into specific control signals. Implement multi-timescale coordination, integrating fast control actions and slow planning decisions into a unified decision-making architecture. Design exception handling mechanisms to maintain basic system functionality when communication is interrupted or agents fail. Through comprehensive integration, output a complete multi-agent collaborative mechanism that is adaptive, robust, and scalable.

[0260] Assume that a high-load area in a city's distribution network experiences a voltage drop. The coordination mechanism is implemented as follows:

[0261] The strategic layer central agent identifies the problem and determines the global goal of increasing the regional voltage;

[0262] Through vertical coordination, the objectives are passed to the tactical-level agents responsible for the area;

[0263] Tactical-level agents assess resources within their jurisdiction and negotiate with agents in neighboring regions to determine whether additional support is available through horizontal collaboration.

[0264] The tactical-level agent decides to adjust transformer taps and input capacitor compensation within the region;

[0265] Through vertical coordination, specific instructions are sent to the transformer agent and capacitor agent at the execution layer;

[0266] The execution layer agent performs specific operations to adjust the parameters of physical devices;

[0267] The execution results are fed back to the upper level through vertical coordination to evaluate the completion of goals and adjust subsequent decisions.

[0268] The multi-agent collaboration mechanism in the embodiment of the present invention has the following functions:

[0269] First, the multi-agent collaboration mechanism defines the organizational structure and decision-making scope of the agents in the system. This provides a framework for identifying and classifying decisions from historical data. For example, knowing which decisions belong to the strategic level, which to the tactical level, and which to the execution level helps correctly map historical decisions to the corresponding decision-making levels and agents.

[0270] Second, agent dependency analysis within the collaborative mechanism helps identify patterns in historical decisions. Using the interaction intensity matrix, the system can identify which decisions in historical data are interconnected rather than independent, thereby extracting complete decision sequences rather than isolated decision points.

[0271] Furthermore, the hierarchical collaborative structure enables the system to extract decision patterns at different time scales and spatial scopes from historical data. Long-term planning decisions at the strategic level, mid-term coordination decisions at the tactical level, and short-term control decisions at the execution level correspond to different time windows and data characteristics. This hierarchical extraction approach enables the system to comprehensively capture expert decision-making knowledge at multiple scales.

[0272] To verify the effectiveness of the proposed method, a practical application test was conducted in a smart distribution network demonstration area in a certain city. The demonstration area covers approximately 15 square kilometers and includes 25 10kV distribution lines, 120 distribution transformers, 15 distributed photovoltaic power plants (total capacity 3.5MW), three energy storage systems (total capacity 2MWh), and approximately 20% of the controllable load. The following verification steps were performed:

[0273] Step 1: Data collection and preprocessing

[0274] Two years of historical operational data from the demonstration zone were collected, including power parameters such as voltage, current, active power, and reactive power at a 5-minute resolution, as well as dispatcher operation records, meteorological data, and load forecast data. The collected data was cleaned and preprocessed, addressing approximately 3% of missing data and 2% of outliers. After standardization, a dataset containing 152 features was constructed.

[0275] Step 2: Model construction and training

[0276] Based on preprocessed data, a linear incomplete information game model was constructed with a state space dimension of 86 and an action space containing 42 discrete actions and 8 continuous action variables. A comprehensive reward function was designed that incorporates grid losses, voltage deviation, regulation costs, and renewable energy consumption rate.

[0277] 978 typical decision rules were extracted from historical data to construct an initial policy library. Regarding safety constraints, 12 hard constraints (such as node voltage limits and line capacity limits) and 8 soft constraints (such as economic indicators and equipment life indicators) were defined.

[0278] Through Laxity metric calculation, 120 basic agents were dynamically aggregated into groups of 15-25 agents, significantly reducing computational complexity. Training was performed using the PPO algorithm, combined with historical data pre-training and safety constraint guidance. The training process executed 5 million environmental interaction steps and converged in approximately 24 hours (using parallel computing on four GPUs).

[0279] Step 3: Performance evaluation results

[0280] On independent three-month test data, we conducted a comparative test with traditional scheduling methods and common reinforcement learning methods. The results are shown in Table 1:

[0281] Table 1 Performance comparison of different scheduling methods

[0282] Evaluation Metrics Traditional scheduling methods General reinforcement learning Method of the present invention Network loss rate (%) 4.82 4.15 3.68 Voltage qualification rate (%) 95.3 97.2 99.1 Number of safety constraint violations 0 18 0 Renewable energy consumption rate (%) 78.5 85.7 91.2 Average scheduling time (s) 62 0.8 1.2 Convergence training rounds - >8000 3200

[0283] Results show that compared to traditional dispatch methods, the proposed method reduces network loss by 23.7%, increases voltage compliance by 3.8 percentage points, increases renewable energy consumption by 12.7 percentage points, and increases dispatch speed by approximately 50 times. Compared to conventional reinforcement learning methods, the proposed method reduces network loss by 11.3%, increases voltage compliance by 1.9 percentage points, completely avoids safety constraint violations, increases renewable energy consumption by 5.5 percentage points, and accelerates training convergence by approximately 60%. Furthermore, the proposed method demonstrates strong robustness and adaptability in the face of abnormal situations such as sudden load changes and equipment failures.

[0284] This example demonstrates the effectiveness and superiority of a historical data-guided secure reinforcement learning approach for optimal distribution network scheduling. Through innovative designs such as historical data-guided exploration, multiple security constraint integration, and Laxity-based multi-agent aggregation, this approach successfully addresses the security risks and learning efficiency challenges of traditional reinforcement learning in distribution network scheduling, achieving secure and efficient distribution network optimization.

[0285] An embodiment of the present invention further provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to perform the method of the embodiment of the present invention.

[0286] The embodiments of the present invention further provide a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiments of the present invention.

[0287] An embodiment of the present invention further provides an electronic device comprising: at least one processor; and a memory communicatively coupled to the at least one processor. The memory stores a computer program executable by the at least one processor, wherein the computer program, when executed by the at least one processor, causes the electronic device to perform the method of an embodiment of the present invention.

[0288] The computer programs for implementing the methods of the embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0289] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0290] It should be noted that the term "including" and its variations used in the embodiments of the present invention are open inclusions, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that unless the context clearly indicates otherwise, they should be understood as "one or more".

[0291] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances shall be provided for users to choose to authorize or refuse.

[0292] The various steps described in the method implementation methods provided by the embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method implementation methods may include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this respect.

[0293] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments are referenced to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.

[0294] The above-described embodiments merely represent several implementation methods of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of protection. It should be noted that a person of ordinary skill in the art would be able to make various modifications and improvements without departing from the scope of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A distribution network optimization scheduling method based on secure reinforcement learning of historical data, characterized by: include: Collect the running time data of multiple intelligent agents in the distribution network for a period of time, and pre-process them to obtain the pre-processed running time data; Calculating the relaxation index of each agent based on the preprocessed runtime data, and classifying the agents according to the similarity between the relaxation indexes of different agents to obtain multiple agent groups; According to the current state mode of the distribution network system, an agent group corresponding to the relaxation index and the state mode is selected as the strategic agent group; Inputting the running time data of the agents in the strategic agent group into a reinforcement learning optimization model based on historical data, and outputting the scheduling strategy corresponding to each agent through the reinforcement learning optimization model; The plurality of dispatching strategies are divided based on the regions and functions of the distribution network, and the distribution network equipment is coordinated and dispatched.

2. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 1 is characterized in that: Real-time collection of the running time data of each intelligent agent in the distribution network during the current period and pre-processing, including: Collecting real-time operational time series data of each intelligent agent in the distribution network over a period of time from the distribution network system; wherein the operational time series data includes: voltage data, power flow data, equipment operating status, dispatch instructions, load changes, and distributed energy output time series data; Identifying missing points and abnormal points in the runtime series data, and correcting the data of the missing points and abnormal points to obtain cleaned runtime series data; Performing time series alignment and standardization transformation on the cleaned runtime data to obtain standardized data; Identify key features that have a significant impact on the operating status and dispatching decisions of the distribution network, and select standardized data corresponding to the key features ranked high in importance as preprocessed operating series data; among them, key features include: node voltage, line flow, equipment status, power factor, load rate and voltage deviation.

3. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 1 is characterized in that: Calculating the relaxation index of each agent based on the preprocessed runtime data includes: Extracting the control feature set reflecting the control capability of each area in the distribution network from the pre-processed operation time series data; Calculate the maximum controllable margin, minimum controllable margin and normalization coefficient of each agent in the current state based on the control feature set; Calculate the difference between the maximum and minimum controllable margins of each agent in its current state, and take the ratio of the difference to the normalization coefficient as the basic slack data of each agent; The basic slack data are subjected to time series analysis and uncertainty assessment to generate the corresponding slack index.

4. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 3 is characterized in that: Extract the regulatory feature set from the preprocessed runtime data, including: Extracting control-related parameters related to distribution network equipment control from pre-processed operating data; Based on the control-related parameters and related factors, a control feature set reflecting the control capability of each area of ​​the distribution network is established; Among them, the control-related parameters include: equipment adjustment range, control accuracy, and response time; the related factors include: distribution network load characteristics, network topology, and operating status.

5. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 1 is characterized in that: The relaxation index of each agent is calculated based on the preprocessed runtime data, and the agents are classified according to the similarity between the relaxation indexes of different agents to obtain multiple agent groups, including: Based on the similarity between the relaxation indicators of each agent, the agents with similar relaxation indicators are divided into agent groups of the same category to obtain multiple agent groups; The average slack of each agent group is ranked, and the corresponding state mode of the distribution network system is identified according to the average slack of each agent group. The state mode includes: high-load weekdays, low-load weekends, and high renewable energy penetration.

6. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 1 is characterized in that: Before inputting the runtime data of the agents in the policy agent group into the historical data-based reinforcement learning optimization model, the method includes: Collect the historical running time series data of each intelligent agent in each distribution network system, and pre-process it to obtain the pre-processed historical running time series data; Create a learning framework for the distribution network based on pre-processed historical runtime data; Generate action space and complete reward function in the distribution network learning framework; Establish safety constraints based on pre-processed historical runtime data, action space, and complete reward function; Based on safety constraints and comprehensive reward functions, the deep learning reinforcement model is pre-trained using historical runtime data to obtain a reinforcement learning optimization model based on historical data.

7. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 6 is characterized in that: Based on the pre-processed historical runtime data, a distribution network learning framework is created, including: Generate the agent's state vector and control action vector based on the preprocessed historical runtime data; Considering the impact of the control action on the state vector, a state transition equation is established to describe the change relationship between the state vector at the current moment and the state vector at the next moment; Based on the state transfer equation, an observation equation is established to describe the mapping relationship between the state vector and the observation data; Each intelligent agent only obtains its own observation data and updates the global state estimation of the distribution network through local observation data based on the observation equation, thus obtaining a learning framework for the distribution network.

8. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 7 is characterized in that: Based on the preprocessed historical runtime data, the action space, and the complete reward function, safety constraints are established, including: Establish hard and soft constraints based on pre-processed historical operating sequence data. Hard constraints include voltage stability constraints, line thermal stability constraints, equipment operating capacity constraints, and system topology constraints. Soft constraints include economic constraints, environmental constraints, and equipment life constraints. Directly constrain the possible actions of each intelligent agent in the distribution network based on hard constraints; A corresponding penalty function is generated according to each type of soft constraint, and the complete reward function is modified by the penalty function to obtain a comprehensive reward function; wherein, the comprehensive reward function is used to indirectly constrain the actions that the intelligent agent may perform.

9. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 6 is characterized in that: Based on safety constraints and a comprehensive reward function, we pre-train the deep learning reinforcement model using historical runtime data to obtain a reinforcement learning optimization model based on historical data, including: Based on the pre-processed historical runtime data, effective decisions are screened and initial decisions are generated; Using the initial decision, pre-training the deep reinforcement learning model to obtain a pre-trained deep reinforcement learning model; Use security constraints to modify the initial policy to obtain a modified policy; The modified strategy is used to perform reinforcement training on the deep learning reinforcement model to obtain a reinforcement learning optimization model based on historical data.

10. The method for optimizing and dispatching a distribution network based on security reinforcement learning based on historical data according to claim 1, characterized in that: The plurality of dispatching strategies are divided based on the regions and functions of the distribution network, and the distribution network equipment is coordinated and dispatched, including: The multiple agents in the strategic agent group are divided into multiple levels based on the regions and functions of the distribution network; wherein the multiple levels include: a strategic level, a tactical level and an execution level; Perform global coordination using the scheduling policies of the agents in the strategy layer; Use the dispatch strategies of the intelligent agents in the tactical layer to perform coordination and optimization within each area of ​​the distribution network; Use the intelligent agent in the execution layer to obtain the scheduling strategy and execute the scheduling operations of each physical device; The distribution network equipment is coordinated and dispatched through the collaboration of intelligent agents in the strategic layer, tactical layer and execution layer.

Citation Information

Cited By

  • Flexible resource scheduling method and system considering state time sequence coupling and medium

    CN120851562A

  • Considering the flexibility of resource scheduling methods, systems, and media that involve state-temporal coupling

    CN120851562B

  • Method and device for generating training sample based on power flow calculation model

    CN122020184A

  • A method for selecting a subset of streaming data under multi-system constraints

    CN122451405A