A power distribution network abnormal working condition identification and self-healing method based on meta-reinforcement learning and related device

By using a meta-reinforcement learning-based approach, combined with multi-source heterogeneous data preprocessing and spatiotemporal graph neural networks, the problems of low identification rate and slow self-healing response of atypical faults in rural power distribution networks were solved, achieving efficient fault identification and self-healing decision-making, and improving the dynamic perception and resilience of the power distribution network.

CN122118744APending Publication Date: 2026-05-29STATE GRID HUBEI ELECTRIC POWER RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID HUBEI ELECTRIC POWER RES INST
Filing Date
2026-01-17
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies in rural power distribution networks suffer from low atypical fault identification rates and slow self-healing responses due to sparse measurements, weak communication, and a scarcity of historical fault samples. This makes it difficult to meet the dynamic and rapid response requirements of a high proportion of distributed renewable energy sources.

Method used

By employing a meta-reinforcement learning-based approach, through multi-source heterogeneous data preprocessing, spatiotemporal graph neural networks, and hierarchical meta-reinforcement learning models, combined with generative adversarial networks and an edge-cloud collaborative computing architecture, we can achieve accurate fault identification and self-healing decision-making.

Benefits of technology

It improves the identification rate of atypical faults, shortens the response time, realizes "identification equals decision-making", and has rapid self-healing capability in small sample scenarios, thereby improving the resilience and power supply reliability of rural power distribution networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122118744A_ABST
    Figure CN122118744A_ABST
Patent Text Reader

Abstract

The application provides a power distribution network abnormal working condition identification and self-healing method based on meta-reinforcement learning and related devices. The method comprises: acquiring multi-source heterogeneous data of a power distribution network, wherein the multi-source heterogeneous data comprises electrical quantity data, distributed photovoltaic output data, energy storage system state data and non-electrical quantity environmental data; preprocessing the multi-source heterogeneous data to obtain preprocessed data; mapping the power distribution network topology to a space-time graph neural network structure, and constructing a hierarchical meta-reinforcement learning model based on the space-time graph neural network structure; inputting the preprocessed data into the constructed hierarchical meta-reinforcement learning model, and synchronously outputting an identification result of an abnormal working condition and a self-healing control strategy. The application realizes "identification is decision-making" by integrating fault identification and self-healing decision-making into a unified model, significantly shortens the response time from fault occurrence to power restoration, and effectively improves the self-healing efficiency and operation resilience of the power distribution network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of power system automation and smart grid technology, and in particular to a method and related apparatus for identifying and self-healing abnormal operating conditions in distribution networks based on meta-reinforcement learning. Background Technology

[0002] Currently, the penetration rate of distributed generation (DG) in rural power distribution networks has increased significantly. The randomness and volatility of its output, coupled with the inherent weaknesses of rural power grids, pose significant challenges to the safe and stable operation of these networks. Against this backdrop, developing rapid and intelligent fault diagnosis and self-healing control technologies has become an important way to enhance the resilience of rural power grids.

[0003] First, the inherent characteristics of rural power grids—weak, remote, and scattered—are at odds with the high proportion of renewable energy. Rural power grids typically have a single-radial structure, with long power supply radii, high line impedance, and weak support capacity. After distributed photovoltaic (PV) grids are connected, the randomness and intermittency of their output can easily lead to problems such as over-limit voltage at the end, reverse power flow, and three-phase imbalance. Second, peak PV output in rural areas usually occurs at midday, while peak electricity load occurs mostly in the evening and at night, resulting in a supply-demand mismatch. Finally, existing fault handling technologies are ill-suited to the complex fault patterns under "PV-storage synergy." In power grids containing a large number of power electronic devices, fault characteristics exhibit nonlinearity and weak current, causing traditional overcurrent-dependent protection methods to frequently fail, especially for identifying atypical faults such as high-resistance grounding.

[0004] Currently, significant progress has been made in research on self-healing control of distribution networks both domestically and internationally. Some studies have also attempted to introduce artificial intelligence and deep learning algorithms to improve the intelligence level of fault diagnosis. However, many studies have not fully considered the key constraints in the operation of rural distribution networks during modeling, generally exhibiting a tendency to "emphasize models while neglecting data." On the one hand, the characteristics of rural power grids—"sparse measurements, numerous measurement blind spots, and weak communication coverage"—make traditional methods relying on high-density measurement data difficult to implement effectively. On the other hand, AI model training heavily relies on massive amounts of fault data, but relevant training samples for rural power grids are scarce, resulting in poor model generalization ability and difficulty in identifying atypical faults such as high-resistance grounding and intermittent arcing. Existing research has also simplified the characterization of the dynamic state transition process of distribution networks, failing to fully reflect the dynamic and rapid response requirements of the power grid to faults under the background of high-proportion distributed generation (DG) access.

[0005] Therefore, in response to the problems of current research models relying on massive amounts of data, being unable to adapt to small sample scenarios, and having insufficient dynamic identification capabilities, there is an urgent need for an intelligent method that can overcome the constraints of "few measurement points" and "scarce samples," while accurately identifying atypical faults and quickly generating self-healing decisions, so as to meet the dual needs of rural power distribution networks for dynamic perception and resilience enhancement. Summary of the Invention

[0006] This invention provides a method and related device for identifying and self-healing abnormal operating conditions in distribution networks based on meta-reinforcement learning. It aims to solve the operational stability problems caused by "spatiotemporal mismatch of light and load" and weak power grid structure in the prior art, as well as the technical difficulties in identifying atypical faults, making slow self-healing decisions, and failing to effectively coordinate light and energy storage resources for fault recovery under conditions of sparse measurement, limited communication, and scarce fault samples.

[0007] To achieve the above objectives, this invention provides a method for identifying and self-healing abnormal operating conditions in power distribution networks based on meta-reinforcement learning, comprising the following steps:

[0008] Acquire multi-source heterogeneous data of the power distribution network, including electrical quantity data, distributed photovoltaic power output data, energy storage system status data, and environmental data of non-electrical quantities;

[0009] The multi-source heterogeneous data is preprocessed to obtain preprocessed data. The preprocessing includes data cleaning and data augmentation based on generative adversarial networks (GANs).

[0010] The distribution network topology is mapped to a spatiotemporal graph neural network structure, and a hierarchical meta-reinforcement learning model is constructed based on the spatiotemporal graph neural network structure.

[0011] The preprocessed data described in step [1] is input into the constructed hierarchical meta-reinforcement learning model, which synchronously outputs the identification results of abnormal operating conditions and self-healing control strategies.

[0012] Furthermore, the preprocessing of the multi-source heterogeneous data specifically includes:

[0013] The Kalman filter algorithm is used to smooth and remove noise from the data, and the isolated forest algorithm is used to detect and remove false outlier data points to complete the data cleaning.

[0014] Generative Adversarial Networks (GANs) are used to augment scarce fault samples. Through adversarial training between the generator network and the discriminator network, the data distribution of real fault samples is learned to generate high-fidelity synthetic fault samples.

[0015] Furthermore, the mapping of the distribution network topology to a spatiotemporal graph neural network structure specifically includes:

[0016] (1) The busbars, load nodes and distributed power grid connection points in the power distribution network are abstracted as nodes in a graph structure, and the closed states of lines, transformers and switches are abstracted as edges;

[0017] (2) Establishing a mathematical representation for defining the graph structure, wherein establishing the mathematical representation specifically includes:

[0018] Based on the dynamic topology and switching states of the distribution network, a sparse adjacency matrix is ​​generated. It is used to define the spatial connection relationship between nodes in the graph structure and to reflect the topological changes of the power grid caused by switching operations;

[0019] Constructing the node feature matrix The graph structure is used to characterize the dynamic operating state of each node, characterized in that each node feature vector in the node feature matrix integrates data of electrical quantity time series features, photovoltaic power output features, energy storage status features and equipment health index.

[0020] (3) Construct the spatiotemporal graph neural network structure, wherein the spatiotemporal graph neural network structure receives the adjacency matrix. and the node feature matrix As input, it is composed of a graph convolutional network (GCN) and a temporal convolutional module (TCN).

[0021] Furthermore, the spatiotemporal graph neural network structure is used to extract the dynamic evolution features of the fault in the spatial and temporal dimensions, wherein:

[0022] The graph convolutional network (GCN) is used to propagate and aggregate information along the power grid topology defined by the adjacency matrix at each time step. By aggregating the feature information of neighboring nodes, it updates the feature representation of the central node, thereby extracting the spatial propagation pattern of fault signals in the power grid. Its update process follows the following propagation rules:

[0023]

[0024] in, It is the first The node feature matrix of the layer In the adjacency matrix The matrix obtained by adding self-loops to the base matrix , yes diagonal matrix, It is the first Layer-trainable weight matrix, It is a non-linear activation function;

[0025] The Temporal Convolutional Module (TCN) is used to receive the graph embedding sequence output by the GCN at consecutive time steps, and capture the long-range temporal dependencies in the fault evolution process through dilated causal convolution to distinguish between transient disturbances and permanent faults.

[0026] Furthermore, the hierarchical meta-reinforcement learning model is a Markov decision process (MDP), specifically:

[0027] (1) The hierarchical control architecture includes a high-level meta controller for operating on a slower time scale and setting abstract sub-objectives based on the global power grid situation extracted by the spatiotemporal graph neural network; and a low-level controller for receiving the sub-objectives and generating specific, fine-grained action sequences to achieve the sub-objectives; wherein, the low-level controller uses the Double Q Network (DQN) algorithm to learn the optimal action value function and outputs the identification results and the self-healing control strategy synchronously through a parallel output layer with an identification head and a control head.

[0028] (2) An action space (A), wherein the action sequence generated by the lower-level controller includes a switch “on” or “off” operation for topology reconfiguration, and an energy storage system “charge”, “discharge” or “standby” command for photovoltaic-storage synergy.

[0029] (3) A multi-objective reward function (R) is used to guide the learning of the model. The reward function is designed as a linear combination of multiple objectives to balance recovery benefits, operating costs and operational safety.

[0030] (3) Meta-learning training mechanism, characterized by training through a hybrid training mechanism that combines offline meta-training and online rapid adaptation on a digital twin platform.

[0031] Furthermore, the abnormal operating conditions include voltage over-limit or power flow overload, as well as high-resistance grounding faults or intermittent arcing faults.

[0032] Furthermore, the method is deployed on an edge-cloud collaborative computing architecture, the deployment including:

[0033] At the edge computing device, a lightweight feature extraction engine is deployed to perform real-time data cleaning and feature extraction close to the data source, reducing reliance on communication bandwidth and executing pre-defined emergency control strategies in the event of communication interruption.

[0034] At the cloud layer, its high computing power is used to support the complex offline training and online updates of the hierarchical meta-reinforcement learning model, and a digital twin system is deployed to perform rapid simulation and security verification of the self-healing control strategy before it is issued.

[0035] At the cloud layer, a blockchain-based failure case library is established to ensure the authenticity and traceability of historical cases, and the model is enabled to continuously learn from new events through online incremental learning.

[0036] A computer device includes: a memory and a processor; the memory stores a computer program, and when the processor executes the computer program, it implements the method described above.

[0037] A computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the method described above.

[0038] In summary, the technical solution provided by this invention has the following beneficial effects:

[0039] 1. Overcoming the small sample dilemma: By introducing meta-reinforcement learning and data augmentation based on generative adversarial networks (GANs), the problem of scarce historical fault samples in rural power distribution networks is solved.

[0040] 2. Improved identification rate of atypical faults: By constructing a spatiotemporal graph neural network (GCN+TCN) and integrating multi-source heterogeneous information, it is possible to dynamically characterize the distribution network topology and accurately extract the spatiotemporal correlation features of faults.

[0041] 3. Enables rapid collaborative decision-making: By integrating fault identification and self-healing strategy generation into the same model, "identification equals decision-making," which greatly shortens response time. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0043] Figure 1 This is a schematic diagram of the overall process of the method of the present invention.

[0044] Figure 2 This is a schematic diagram of the edge-cloud collaborative computing system architecture deployed by the method of the present invention.

[0045] Figure 3 This is a schematic diagram of the process of mapping the physical topology of a power distribution network into a spatiotemporal neural network structure.

[0046] Figure 4 This is a structural block diagram of a hierarchical meta-reinforcement learning model.

[0047] Figure 5 This is a schematic diagram of the multi-source heterogeneous data preprocessing process.

[0048] Figure 6 This is a timing diagram showing the identification accuracy and self-healing action sequence when the present invention is applied to a typical high-resistance grounding fault scenario. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0050] Reference Figure 1 This invention provides a method for identifying and self-healing abnormal operating conditions in distribution networks based on meta-reinforcement learning. This method aims to address the problems of low identification rate and slow self-healing response in rural distribution networks due to sparse measurements, weak communication, and a scarcity of historical fault samples. The method mainly includes the following steps:

[0051] Step S101: Acquire multi-source heterogeneous data from the distribution network. The system collects data from various sensors and data platforms deployed in the distribution network. This data includes traditional electrical quantity measurements, distributed photovoltaic power output data, energy storage system status data, and non-electrical quantity data that can provide key environmental background and equipment status information.

[0052] In order to achieve accurate identification of abnormal operating conditions in the power distribution network, this invention emphasizes the comprehensive utilization of multi-source heterogeneous data, breaking the limitations of relying solely on traditional electrical quantity measurements.

[0053] The data collected in this invention are divided into three main categories:

[0054] Electrical quantity data of power grid operation: including time-series data such as voltage, current, and power collected by SCADA, PMU or high-frequency recording devices.

[0055] New energy and energy storage status data: This is a key input for realizing "photovoltaic-storage synergy", including the real-time output of distributed photovoltaic (PV) and the state of charge (SOC) and charge / discharge power of energy storage system (ESS).

[0056] Non-electrical environmental and equipment data: including meteorological data such as light intensity, lightning strike location information, and equipment ledgers and health indices including inverters and converters.

[0057] Step S102: Preprocess the multi-source heterogeneous data. Due to the uneven quality of the original data and the existence of sample imbalance, it is necessary to clean, denoise, and standardize the data, and use advanced data augmentation techniques, especially generative adversarial networks (GANs), to expand the scarce fault sample set, laying the foundation for the robust training of the subsequent model.

[0058] Data preprocessing workflow reference Figure 5First, Kalman filtering and isolated forest algorithms are used for data cleaning. Second, to address the scarcity of rural power grid samples, generative adversarial networks (GANs) are employed for data augmentation. Through adversarial training between the generator and discriminator networks within GANs, high-fidelity synthetic fault samples are generated.

[0059] Data preprocessing workflow: Refer to Figure 5 Before being input into the model, the raw multi-source data must undergo a rigorous preprocessing process to ensure the quality and applicability of the data.

[0060] Step S201: Data Cleaning. This step employs two algorithms working in tandem. First, a Kalman filter algorithm is used to process the time-series data (such as voltage and current) acquired by the sensors. Then, an Isolation Forest algorithm is used to detect and remove outliers.

[0061] Step S202: Data Augmentation. This is a key step designed by the present invention to address the core problem of "scarcity of training samples for historical fault events" in rural power distribution networks. The present invention employs a Generative Adversarial Network (GAN) to generate new, high-quality fault samples. A GAN consists of two competing neural networks:

[0062] The generator's task is to learn the data distribution of real fault samples. It takes a random noise vector as input and attempts to output a synthetic sample that is statistically indistinguishable from real fault data (such as fault waveforms).

[0063] Discriminator: It is a standard binary classifier whose task is to determine whether a given sample comes from a real fault dataset or is synthesized by a generator.

[0064] To address the scarcity of rural power grid samples, a generative adversarial network (GAN) is used for data augmentation. Through adversarial training between the generator and discriminator networks within the GAN, high-fidelity synthetic fault samples are generated.

[0065] Step S103: Map the distribution network topology to a spatiotemporal graph neural network structure and construct a hierarchical meta-reinforcement learning model. The physical connections of the distribution network are abstracted into a graph structure, and a spatiotemporal graph neural network (ST-GNN) is used to extract the dynamic evolution features of faults in the spatial and temporal dimensions. Based on this, a hierarchical meta-reinforcement learning (Meta-RL) model is constructed to handle complex decision-making problems.

[0066] To enable deep learning models to understand and process the topology and dynamic behavior of power distribution networks, this invention proposes an innovative spatiotemporal graph modeling method, which transforms the physical entities and operating states of the power grid into mathematical representations that the model can process.

[0067] Graph structure construction: Refer to Figure 3 This process maps the physical topology of the distribution network to a mathematical graph. .

[0068] The set of nodes V (Nodes): Each bus in the distribution network, including substation buses, load nodes, distributed generation grid connection points, etc., is abstracted as a node in the graph. .

[0069] The set of edges E (Edges): Physical devices connecting any two buses, such as transmission lines, transformers, switches, etc., are abstracted as an edge in the graph. The presence or absence of edges directly reflects the state of the device.

[0070] Adjacency matrix A: Generated based on the dynamic topology of the power grid and switching states, used to define the spatial connectivity between nodes. Since the power grid topology changes due to switching operations, the adjacency matrix... It is dynamic, and the model needs to be able to handle this dynamic topology.

[0071] Node feature construction:

[0072] Used to characterize the dynamic operating state of each node. In a preferred embodiment, each node in the graph... They are all associated with a feature vector The eigenvectors of all nodes constitute the node feature matrix. This feature vector is multi-dimensional and designed to comprehensively describe the state of the node and its connected devices. In one embodiment, the feature vector for each node... It is composed of electrical quantity time-series characteristics, photovoltaic power output characteristics, energy storage status characteristics, and equipment health index.

[0073] Spatiotemporal Graph Neural Network (ST-GNN) Architecture: To simultaneously process the spatial topology and temporal dynamics of the power grid, this invention employs a Spatiotemporal Graph Neural Network (ST-GNN) architecture. This architecture consists of two stacked core modules:

[0074] Graph Convolutional Networks (GCNs) extract spatial features:

[0075] The GCN module is responsible for propagating and aggregating information along the power grid topology at each time step. Its core idea is that the feature representation (or embedding) of each node is updated by aggregating the feature representations of its neighboring nodes. A typical propagation rule for a GCN layer can be expressed as:

[0076]

[0077] in, It is the first The node feature matrix of the layer It is the trainable weight matrix of this layer. In the adjacency matrix The addition of self-loops on top of the existing structure means that nodes will consider their own information while aggregating neighbor information. yes diagonal matrix, It is a symmetric normalization of the adjacency matrix, a standard operation used to prevent nodes with high degree from having an excessive impact on the aggregation process and to maintain numerical stability. It is a non-linear activation function.

[0078] By stacking multiple layers of GCN, a node can aggregate information from its multi-hop neighbors, thereby capturing the spatial patterns of fault signals propagating in the grid topology.

[0079] Temporal Convolutional Networks (TCNs) capture temporal patterns:

[0080] The GCN module at each time point Each outputs a graph embedding that captures the spatial information at that moment. Embedding graphs at consecutive time steps into a sequence. As input, it is fed into the TCN module. TCN uses dilated convolutions to capture long-range temporal dependencies in the fault evolution process.

[0081] GCN is used to extract the spatial propagation pattern of fault signals in the power grid, while TCN is used to capture the long-range temporal dependencies in the fault evolution process. By combining GCN and TCN, the ST-GNN model can generate a highly condensed feature representation that simultaneously encodes the spatial propagation path and temporal evolution law of the fault in the power grid, providing high-quality input for subsequent decision models.

[0082] Step S104: Synchronously output identification results and self-healing strategies through the model. Input the preprocessed data into the trained model. A core innovation of this model is that its output layer is designed to simultaneously generate two results: one is an accurate identification of the current power grid operating conditions (e.g., fault type, location), and the other is the optimal self-healing control strategy under the identification results (e.g., a series of switching actions).

[0083] Reference Figure 4 The entire self-healing control problem is formalized as a Markov decision process (MDP), which specifically includes:

[0084] State space (S): The state at any given time. This is the feature vector output by the aforementioned ST-GNN module, which integrates spatiotemporal information. This state vector comprehensively describes the current operating status of the power grid.

[0085] Action Space (A): The actions that an agent can perform are a discrete set, mainly including performing "open" or "close" operations on sectionalizing switches and tie switches in the distribution network.

[0086] Reward function (R): This is key to guiding the agent's learning. Carefully designed to quantify the execution of actions The state reached by the system afterwards The degree of superiority or inferiority reflects the core idea of ​​"maximizing the benefits of restoration".

[0087] Hierarchical reinforcement learning architecture:

[0088] To address the complex and multi-layered objectives in distribution network restoration decisions, this invention employs a hierarchical reinforcement learning architecture. This architecture decomposes the complex decision-making task into two levels:

[0089] Meta-Controller: This controller operates on a slower timescale and is responsible for setting abstract sub-goals and selecting a strategic task based on the overall state of the current power grid.

[0090] Low-level Controller: This controller receives sub-goals from the higher-level meta-controller and is responsible for generating specific, fine-grained action sequences to achieve those sub-goals.

[0091] Policy learning and synchronous output based on Double DQN:

[0092] The lower-level controller uses a deep Q-network (DQN) to learn the optimal action value function. This function represents the state. Next action The maximum expected return that can be obtained is determined by the algorithm. This invention uses a Double DQN architecture. Compared to standard DQN, Double DQN uses two independent neural networks (an online network for selecting actions and a target network for evaluating the value of the action) to decouple action selection and value evaluation, effectively alleviating the problem of overestimation of Q-values ​​that is common in standard DQN. The formula is as follows:

[0093]

[0094] in The goal value, It's the reward for the next moment. It is a discount factor. It is the state in the next moment. These are online network parameters. These are the target network parameters.

[0095] A key innovation of this invention lies in the design of the model's output layer. The final layer of the neural network is designed to have both a recognition head and a control head, enabling simultaneous output of recognition and decision-making:

[0096] Identification Head: The classification output layer, typically using the Softmax activation function. It outputs a probability distribution representing the current state. The probability of each of the predefined abnormal operating conditions (such as single-phase grounding, two-phase short circuit, high-resistance grounding, normal, etc.) is given, along with the confidence level of the fault location.

[0097] Control Head: The regression output layer, outputting the Q-value corresponding to each possible switching action. The agent executes the optimal self-healing strategy by selecting the action with the highest Q-value.

[0098] Multi-objective reward function design:

[0099] The design of the reward function is crucial to the success of reinforcement learning, as it directly determines the learning objective of the agent. The reward function designed in this invention... It is a linear combination of multiple objectives, designed to balance recovery benefits, operating costs, and operational safety:

[0100]

[0101] The components are defined as follows:

[0102] The load that has successfully had its power restored (unit: kW) can be weighted according to the importance of the load (such as hospitals and government agencies, which are high priority), directly corresponding to the goal of "maximizing the restoration benefits".

[0103] The number of switching operations performed within this time step. This negative reward is introduced to penalize unnecessary or excessively frequent switching operations, as these operations can shorten equipment lifespan and potentially cause transient shocks.

[0104] : Voltage over-limit penalty term. This is a step function or smoothing function that applies when the voltage of any node in the network exceeds the normal operating range (e.g., the nominal voltage). When the value is 0, this term is a large negative value; otherwise, it is 0, to ensure that the self-healing strategy does not sacrifice the voltage stability of the entire system in order to restore part of the load.

[0105] Line overload penalty. Similar to voltage penalty, this term is a large negative value when the current in any line exceeds its thermal stability limit, to ensure the safety of the recovery plan.

[0106] : These are the weighting coefficients for each item, which can be adjusted by operation experts according to the specific requirements of the power grid to balance the priorities among different objectives.

[0107] By optimizing this comprehensive reward function, the RL agent learns to restore as much load as possible while taking into account the economy of operation and the safety constraints of the system, generating a comprehensive, practical and safe self-healing solution.

[0108] Meta-learning for rapid adaptation: The "meta" learning in this invention is embodied in its training mechanism, enabling the model to quickly adapt to new and unseen fault scenarios. The training process consists of two phases:

[0109] Offline meta-training: On a cloud-based digital twin platform, the model is trained on a series of different "tasks." Each "task" corresponds to a specific fault scenario (such as different fault types, locations, impedances, etc.). The goal of the model is to learn a universal initialization parameter, from which it can be quickly fine-tuned to its optimal state on any new task with a small number of samples.

[0110] Rapid online adaptation: Once the model is deployed to a real-world system, if it encounters a completely new type of fault not included in the offline training, the system can quickly simulate the fault several times in a digital twin, generating a small number of samples. Leveraging its "learning ability" acquired during the meta-training phase, the model can rapidly adjust its internal parameters using only these few samples, adapting to the new scenario and making the correct decisions.

[0111] To ensure reproducibility, Table 1 provides a typical hyperparameter configuration for training the AI ​​model in a preferred embodiment. These parameters are selected based on publicly available research and best practices in the field, and have been tailored for distribution network application scenarios.

[0112] Table 1: Examples of AI Model Hyperparameter Configurations

[0113]

[0114] Reference Figure 2 The method of this invention is deployed in an edge-cloud collaborative computing architecture, the deployment including:

[0115] At the edge computing device, a lightweight feature extraction engine is deployed to perform real-time data cleaning and feature extraction close to the data source, reducing reliance on communication bandwidth and executing pre-defined emergency control strategies in the event of communication interruption.

[0116] At the cloud layer, its high computing power is used to support the complex offline training and online updates of the hierarchical meta-reinforcement learning model, and a digital twin system is deployed to perform rapid simulation and security verification of the self-healing control strategy before it is issued.

[0117] At the cloud layer, a blockchain-based failure case library is established to ensure the authenticity and traceability of historical cases, and the model is enabled to continuously learn from new events through online incremental learning.

[0118] The edge layer is deployed in field devices close to the data source, performing data cleaning, feature extraction, and emergency control. The cloud layer consists of high-performance servers that handle compute-intensive tasks, performing model training, digital twin security verification, blockchain fault case library management, and online incremental learning.

[0119] After obtaining the refined state representation extracted by ST-GNN, this invention employs an innovative hierarchical meta-reinforcement learning model to identify abnormal operating conditions and make self-healing decisions.

[0120] To illustrate the technical effects of the present invention more specifically, a specific embodiment is described below.

[0121] Scene setup:

[0122] This embodiment employs a modified IEEE 123-node distribution network test system. The system has been modified to include multiple distributed photovoltaic (PV) power generation units to simulate a typical rural distribution network with a high proportion of renewable energy.

[0123] Fault Scenario: A challenging abnormal operating condition is set up: an intermittent high-impedance ground fault (HIF) occurs on a branch feeder with high photovoltaic penetration. This type of fault has a small current and inconspicuous characteristics, making it difficult to detect with traditional protection systems. Furthermore, the random power output of the photovoltaic system further interferes with voltage and current signals, increasing the difficulty of identification.

[0124] Method execution steps:

[0125] 1. Data Acquisition and Preprocessing: The system collects real-time voltage, current, and power data from all nodes in the network, as well as meteorological data (light intensity) and health indices of relevant switching equipment in the area. After Kalman filtering and isolated forest cleaning, the data is sent to the data augmentation module.

[0126] 2. Spatiotemporal Feature Extraction: The real-time topology of the distribution network (all switch closed states) and the feature vectors of each node are input into the ST-GNN module. The GCN layer captures the spatial propagation pattern of the voltage drop caused by the fault in the network topology, while the TCN layer identifies the "intermittent" temporal characteristics of the fault current and the temporal pattern of photovoltaic power output fluctuations. The ST-GNN ultimately outputs a highly condensed state vector. .

[0127] 3. Hierarchical decision-making: state vector They are fed into a hierarchical meta-reinforcement learning model.

[0128] The high-level element controller received Subsequently, it quickly determined that the main problem was a fault that was difficult to locate and was accompanied by voltage fluctuations. Therefore, it set a sub-objective: "Isolate the unknown fault source and stabilize the branch voltage".

[0129] The lower-level controller receives this target and state vector. The two output heads of its internal Double DQN model work simultaneously: the identification head analyzes the spatiotemporal characteristics of the input and outputs "the probability of a high-impedance ground fault occurring near node X is 99%"; the control head calculates the Q value of all available switch actions.

[0130] 4. Strategy Generation and Execution: The model selects the action sequence with the highest Q value. The intent of this strategy is to first isolate the network segment containing the suspected fault point, and then restore power to the downstream healthy portion of that network segment via a tie switch from another healthy feeder.

[0131] 5. Safety Verification and Recovery: Before being deployed, this strategy was verified by a cloud-based digital twin system. Simulation showed that the operation would not cause any line overload or voltage exceedance. After successful verification, control commands were sent to the field FTU to execute the switching operation. The fault area was successfully isolated, and most of the affected users had their power restored in a very short time.

[0132] Reference Figure 6 The figure shows a performance comparison between the method of the present invention and the traditional baseline method in dealing with the above-mentioned fault scenarios.

[0133] Accuracy and speed of identification: As shown in the figure, traditional methods may fail to detect faults at all due to the weak HIF features, or issue vague alarms only after several minutes. However, the method of this invention, with its powerful spatiotemporal feature extraction capability, accurately identifies the fault type and location with a confidence level of over 98% within hundreds of milliseconds after the fault occurs.

[0134] Self-healing efficiency: Traditional methods, even when detecting a fault, rely on pre-defined, fixed reconstruction logic, and the recovery process can take several minutes or even require manual intervention. The method of this invention generates the optimal recovery strategy simultaneously with fault identification, completing the entire "perception-decision-execution" closed loop within seconds.

[0135] Recovery benefits: Figure 6 The load recovery curves show that the method of this invention restored 95% of the lost load, while traditional methods may result in a continuous power outage of the entire feeder due to the inability to locate the fault. This fully demonstrates the significant advantages of this invention in improving power supply reliability and maximizing recovery benefits.

[0136] Another aspect of the present invention provides a computer device, including: a memory and a processor; characterized in that the memory stores a computer program, and when the processor executes the computer program, it implements the above-described method for identifying and self-healing abnormal operating conditions of power distribution networks based on meta-reinforcement learning.

[0137] In another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the aforementioned method for identifying and self-healing abnormal operating conditions in a power distribution network based on meta-reinforcement learning.

[0138] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0139] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for identifying and self-healing abnormal operating conditions in a power distribution network based on meta-reinforcement learning, characterized in that, Includes the following steps: Acquire multi-source heterogeneous data of the power distribution network, including electrical quantity data, distributed photovoltaic power output data, energy storage system status data, and environmental data of non-electrical quantities; The multi-source heterogeneous data is preprocessed to obtain preprocessed data. The preprocessing includes data cleaning and data augmentation based on generative adversarial networks (GANs). The distribution network topology is mapped to a spatiotemporal graph neural network structure, and a hierarchical meta-reinforcement learning model is constructed based on the spatiotemporal graph neural network structure. The preprocessed data described in step [1] is input into the constructed hierarchical meta-reinforcement learning model, which synchronously outputs the identification results of abnormal operating conditions and self-healing control strategies.

2. The method according to claim 1, characterized in that, The preprocessing of the multi-source heterogeneous data specifically includes: The Kalman filter algorithm is used to smooth and remove noise from the data, and the isolated forest algorithm is used to detect and remove false outlier data points to complete the data cleaning. Generative Adversarial Networks (GANs) are used to augment scarce fault samples. Through adversarial training between the generator network and the discriminator network, the data distribution of real fault samples is learned to generate high-fidelity synthetic fault samples.

3. The method according to claim 1, characterized in that, The mapping of the distribution network topology to a spatiotemporal graph neural network structure specifically includes: (1) The busbars, load nodes and distributed power grid connection points in the power distribution network are abstracted as nodes in a graph structure, and the closed states of lines, transformers and switches are abstracted as edges; (2) Establishing a mathematical representation for defining the graph structure, wherein establishing the mathematical representation specifically includes: Based on the dynamic topology and switching states of the distribution network, a sparse adjacency matrix is ​​generated. It is used to define the spatial connection relationship between nodes in the graph structure and to reflect the topological changes of the power grid caused by switching operations; Constructing the node feature matrix The graph structure is used to characterize the dynamic operating state of each node, characterized in that each node feature vector in the node feature matrix integrates data of electrical quantity time series features, photovoltaic power output features, energy storage status features and equipment health index. (3) Construct the spatiotemporal graph neural network structure, which receives the adjacency matrix. and the node feature matrix As input, it is composed of a graph convolutional network (GCN) and a temporal convolutional module (TCN).

4. The method according to claim 3, characterized in that, The spatiotemporal graph neural network structure is used to extract the dynamic evolution features of faults in the spatial and temporal dimensions, wherein: The graph convolutional network (GCN) is used to propagate and aggregate information along the power grid topology defined by the adjacency matrix at each time step. By aggregating the feature information of neighboring nodes, it updates the feature representation of the central node, thereby extracting the spatial propagation pattern of fault signals in the power grid. Its update process follows the following propagation rules: ; in, It is the first The node feature matrix of the layer, In the adjacency matrix The matrix obtained by adding self-loops to the base matrix , yes diagonal matrix, It is the first Layer-trainable weight matrix, It is a non-linear activation function; The Temporal Convolutional Module (TCN) is used to receive the graph embedding sequence output by the GCN at consecutive time steps, and capture the long-range temporal dependencies in the fault evolution process through dilated causal convolution to distinguish between transient disturbances and permanent faults.

5. The method according to claim 1, characterized in that, The hierarchical meta-reinforcement learning model is a Markov decision process (MDP), specifically: (1) The hierarchical control architecture includes a high-level meta controller for operating on a slower time scale, setting abstract sub-objectives based on the global power grid situation extracted by the spatiotemporal graph neural network; A low-level controller is used to receive the sub-target and generate specific, fine-grained action sequences to achieve the sub-target; wherein, the low-level controller uses the Double Q Network (DQN) algorithm to learn the optimal action value function, and outputs the identification result and the self-healing control strategy synchronously through a parallel output layer with an identification head and a control head; (2) An action space (A), wherein the action sequence generated by the lower-level controller includes a switch "on" or "off" operation for topology reconfiguration, and an energy storage system "charge", "discharge", or "standby" command for photovoltaic-storage synergy; (3) A multi-objective reward function (R) is used to guide the learning of the model. The reward function is designed as a linear combination of multiple objectives to balance recovery benefits, operating costs and operational safety. (4) Meta-learning training mechanism, characterized by training through a hybrid training mechanism that combines offline meta-training and online rapid adaptation on a digital twin platform.

6. The method according to claim 1, characterized in that, The abnormal operating conditions include voltage over-limit or power flow overload, as well as high-resistance grounding faults or intermittent arcing faults.

7. The method according to claim 1, characterized in that, The method is deployed in an edge-cloud collaborative computing architecture, and the deployment includes: At the edge computing device, a lightweight feature extraction engine is deployed to perform real-time data cleaning and feature extraction close to the data source, reducing reliance on communication bandwidth and executing pre-defined emergency control strategies in the event of communication interruption. At the cloud layer, its high computing power is used to support the complex offline training and online updates of the hierarchical meta-reinforcement learning model, and a digital twin system is deployed to perform rapid simulation and security verification of the self-healing control strategy before it is issued. At the cloud layer, a blockchain-based failure case library is established to ensure the authenticity and traceability of historical cases, and the model is enabled to continuously learn from new events through online incremental learning.

8. A computer device, comprising: A memory and a processor; characterized in that the memory stores a computer program, and when the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.