Power distribution network fault processing equipment

Through the neural network integrated with Transformer architecture and multi-layer self-attention mechanism, the adaptability and interpretability problems of complex scenarios in power distribution network fault processing are solved, efficient and safe fault diagnosis and response are achieved, and the reliability and user satisfaction of the power grid are improved.

CN119398106BActive Publication Date: 2025-08-15STATE GRID SHANDONG ELECTRIC POWER CO PINGYI COUNTY POWER SUPPLY CO

Patent Information

Application Number
CN202411182820.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2025-08-15
Estimated Expiration
2044-08-27

AI Technical Summary

Technical Problem

The existing power distribution network fault handling technology lacks a deep understanding of complex scenarios and lacks adaptive learning capabilities, making it difficult to deal with new and composite failures, and has shortcomings in data security, privacy protection and system interpretability.

Method used

The neural network based on Transformer architecture is adopted, combining multi-layer self-attention mechanism and feedforward neural network, and integrating pre-training modules, data acquisition layer, edge computing layer, core processing layer, context learning module, inference processing module, output and interaction layer and continuous learning and improvement layer. Through self-supervised learning, federated learning and reinforcement learning, failure response strategies are optimized, the system is realized, and the transparency of human-computer collaboration and decision-making is enhanced.

Benefits of technology

Significantly improves the accuracy and response speed of fault prediction and diagnosis, reduces power outage time, improves system availability and user satisfaction, enhances data security and system adaptability, and supports the rapid handling of complex faults and transparency in decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119398106B_ABST
    Figure CN119398106B_ABST
Patent Text Reader

Abstract

The present invention relates to an electric power distribution network fault processing device applied to the field of power distribution technology, aiming to improve the fault prediction, diagnosis and processing capabilities of the power system. The device includes a pre-training module, a neural network architecture module, a data acquisition layer, an edge computing layer, a core processing layer, a context learning module, an inference processing module, an output and interaction layer, and a continuous learning and improvement layer. The core of the present invention lies in the use of large-scale pre-trained Transformer models, combined with power system expertise, to achieve in-depth understanding and analysis of complex power grid states. Through multi-source data fusion, edge computing and center collaborative processing, the device can monitor the power grid state in real time, accurately predict potential faults, and provide intelligent diagnosis and repair suggestions. Self-supervised learning and federated learning technologies are used for model pre-training and optimization; reinforcement learning is used to optimize fault response strategies; the model's interpretability is achieved, and decision-making transparency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a power distribution network fault processing device, in particular to a power distribution network fault processing device applied in the field of power distribution technology. Background Art

[0002] As the scale of power systems continues to expand and the complexity continues to increase, fault handling in power distribution networks faces increasing challenges. Traditional fault handling methods mainly rely on manual experience and simple data analysis, which are difficult to cope with the complex and changing power grid environment and the increasing number of fault types. In recent years, with the development of artificial intelligence technology, some fault handling methods based on big data and digital twin technology have been proposed, but there are still some limitations.

[0003] Chinese invention patent CN108320043B discloses a distribution network equipment status diagnosis and prediction method based on power big data. This method uses holographic time-scale measurement data to diagnose and predict equipment status. It can discover weak links in distribution network equipment, quickly locate faults and provide treatment suggestions. However, this method mainly relies on predefined feature quantities and rules, and lacks deep understanding of complex scenarios and adaptive learning capabilities.

[0004] Chinese invention patent CN116885858B proposes a distribution network fault handling method and system based on digital twin technology. The system uses a digital twin model to achieve real-time monitoring and fault simulation of the distribution network, improving the efficiency of fault detection and handling. However, this method still has limitations when dealing with unknown fault types and complex fault scenarios, and lacks a mechanism for continuous learning and knowledge accumulation.

[0005] The above design improves the efficiency and accuracy of fault handling through big data analysis and digital twin technology, but it still has certain limitations, such as lack of understanding of complex contexts, limited adaptive learning capabilities, difficulty in handling new and complex faults, and insufficient human-computer collaboration. In addition, existing technologies also have shortcomings in data security, privacy protection, and system explainability, all of which limit their effectiveness and credibility in practical applications. Summary of the Invention

[0006] The purpose of the present invention is to develop a power distribution network fault handling device with higher intelligence, stronger adaptability and better interpretability.

[0007] The present invention is achieved through the following technical solutions:

[0008] To solve the above problems, the present invention provides a power distribution network fault processing device, which includes a pre-training module, a neural network architecture module, a data acquisition layer, an edge computing layer, a core processing layer, a context learning module, an inference processing module, an output and interaction layer, and a continuous learning and improvement layer;

[0009] A pre-training module for unsupervised learning on large-scale power system historical data to learn the statistical laws and semantic relationships of power grid operation;

[0010] Neural network architecture module, based on the Transformer architecture, including multi-layer self-attention mechanism and feedforward neural network, with billions to hundreds of billions of trainable parameters;

[0011] Data collection layer, including smart sensor networks, SCADA systems, PMU devices, weather stations, and social media API interfaces;

[0012] Edge computing layer, including local pre-processing unit, edge AI inference engine and real-time monitoring panel;

[0013] The core processing layer includes a data fusion center, a main AI diagnostic engine (using the Transformer architecture), a predictive maintenance module, a digital twin platform, and a reinforcement learning optimizer;

[0014] A contextual learning module that can understand and utilize long sequences of grid status information and focus on relevant information through an attention mechanism;

[0015] The reasoning processing module receives grid status information as input prompts and gradually generates fault diagnosis and treatment suggestions;

[0016] Output and interaction layer, including fault report generator, visual decision support system, repair suggestion optimizer and safety assessment module;

[0017] The continuous learning and improvement layer includes adaptive learning modules, human-computer collaboration interfaces, and knowledge base update systems.

[0018] In the above-mentioned power distribution network fault handling equipment, the understanding and handling capabilities of complex power grid environments and diverse fault types are improved, and continuous learning and adaptive optimization of the system are achieved to cope with the ever-changing power grid environment and emerging fault types, enhance human-computer collaboration, improve the transparency and explainability of decision-making, and enhance the system's real-time response capabilities and large-scale data processing capabilities.

[0019] As a further improvement of this application, the pre-training module adopts a self-supervised learning method, including but not limited to mask state prediction and next state prediction tasks, to learn the inherent laws of the power grid system, and realizes distributed training of the model through federated learning technology, thereby improving the accuracy and generalization ability of the model while protecting data privacy.

[0020] As a further improvement of this application, the edge AI inference engine of the edge computing layer is deployed with a lightweight model that has undergone knowledge distillation for fast preliminary fault detection and localization processing. The lightweight model inherits the core knowledge of the large-scale pre-trained model.

[0021] As a further improvement to this application, the digital twin platform of the core processing layer can create a virtual model of the power grid, perform real-time simulation and hypothetical scenario analysis, and combine with the Transformer model to enhance contextual understanding and prediction capabilities.

[0022] As another improvement of this application, the reinforcement learning optimizer of the core processing layer can dynamically optimize fault response strategies and system parameter adjustments through interactive training with the simulated environment of the digital twin platform. Its decision-making process uses the attention mechanism of the Transformer model to weigh multiple factors.

[0023] As another improved supplement to this application, the repair suggestion optimizer of the output and interaction layer adopts a sequence-to-sequence model based on Transformer, which can generate the optimal repair order and plan according to the current resource status and priority.

[0024] As another improved supplement to this application, the adaptive learning module of the continuous learning and improvement layer can continuously update and optimize the parameters of the AI model through incremental learning and fine-tuning technology based on newly collected data and feedback information, so as to maintain the model's adaptability to the latest power grid conditions.

[0025] As another improvement of this application, it also includes a high-speed and reliable data transmission network based on 5G or dedicated optical fiber network to ensure real-time data exchange between various layers of the system and support joint training and reasoning of large-scale distributed models.

[0026] It also includes multi-level backup and fault-tolerant mechanisms, as well as strict data security protocols to ensure the high availability and data security of the system; at the same time, it integrates an interpretability module to provide understandable explanations for fault diagnosis results by analyzing the model's attention weights and activation values.

[0027] The power distribution network fault processing method includes the following steps:

[0028] S1: Pre-training the Transformer model on large-scale power system historical data to learn general knowledge about power grid operation;

[0029] S2: Fine-tune and distill knowledge of the pre-trained model using local data for a specific power grid environment;

[0030] S3: Continuously collects multi-source, multi-modal real-time grid status information through the data acquisition layer and converts it into a tokenized input sequence;

[0031] S4: Use lightweight models for data preprocessing and preliminary fault detection at the edge computing layer;

[0032] S5: In the core processing layer, all input data is integrated and processed, and the main AI diagnostic engine is run to perform in-depth fault analysis.

[0033] S6: The contextual learning module integrates long-sequence historical information to build a comprehensive understanding of the current grid status;

[0034] S7: Use digital twin platforms for fault simulation and impact assessment;

[0035] S8: Generate the optimal repair strategy by combining the reinforcement learning optimizer with the attention mechanism of the Transformer model;

[0036] S9: Generates fault reports, visual decision support information, and multiple possible treatment plans at the output and interaction layer;

[0037] S10: Allow experts to review and intervene in AI decisions through a human-machine collaborative interface, and provide explainable analysis of the decision-making basis;

[0038] S11: Collect feedback on actual processing results, use adaptive learning modules and knowledge base to update the system, and continuously optimize system performance.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. Improved fault prediction and diagnosis capabilities, significantly improving the accuracy of fault prediction and early warning time, greatly improving the diagnostic accuracy of complex and rare faults, and significantly shortening fault response and diagnosis time;

[0041] 2. Improved system reliability and efficiency, significantly reducing the average annual power outage time and the frequency of large-scale power outages, significantly shortening fault repair time, improving repair efficiency, and significantly increasing the overall system availability and fault recovery speed;

[0042] 3. Optimize intelligent decision support to improve decision speed and quality. Through explainable analysis, enhance experts' trust in AI decisions, optimize repair strategies, and improve system recovery speed in complex fault situations.

[0043] 4. Predictive maintenance and long-term optimization to extend the average life of equipment, reduce replacement and maintenance costs, reduce unplanned downtime, improve overall system availability, continuously learn and adapt to new fault types and grid changes, build a comprehensive grid fault knowledge base, and promote experience transfer;

[0044] 5. Economic benefits and cost savings: significantly reduce annual operation and maintenance costs, reduce economic losses caused by power outages, improve equipment utilization efficiency, reduce replacement and maintenance expenses, improve energy efficiency, and support clean energy integration;

[0045] 6. Enhanced security and privacy protection, reducing the risk of data leakage, improving data transmission security, protecting data privacy through federated learning and differential privacy technologies, improving the system's ability to resist network attacks, and enhancing overall security;

[0046] 7. Improve user experience and satisfaction, increase customer satisfaction, reduce complaints, increase transparency in the decision-making process, enhance user trust, and provide timely fault notifications and personalized fault reports;

[0047] 8. Environmental impact and sustainable development: reducing carbon emissions by optimizing grid operations, improving the efficiency of renewable energy utilization, promoting energy transformation, reducing electronic waste by extending equipment life, providing technical support for the development of smart grids, and adapting to future energy structure changes;

[0048] 9. System adaptability and scalability: quickly adapt to new equipment and grid topology changes, support the access of large-scale distributed energy and new loads, improve grid flexibility, and adapt to the complex needs of future power systems;

[0049] 10. Knowledge accumulation and technological innovation: promote the digitization and systematization of power system expertise, promote the in-depth application of artificial intelligence technology in the power industry, and provide innovative solutions for the modernization and intelligentization of power grids. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a timing diagram of the work flow of the power distribution network fault processing device of this application;

[0051] Figure 2 This is a timing diagram of the pre-training, edge computing, and data transmission process of the power distribution network fault handling equipment of this application;

[0052] Figure 3 This is a timing diagram of the core AI processing flow of the power distribution network fault processing equipment of this application;

[0053] Figure 4 A timing diagram of the continuous learning, security and explainability process of the power distribution network fault handling device of this application;

[0054] Figure 5 This is a timing diagram of the complete workflow of the power grid fault handling device of this application. DETAILED DESCRIPTION

[0055] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention is clearly and completely described below in conjunction with the drawings of the present invention. Based on the embodiments in this application, other similar embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of this application. In addition, the directional words mentioned in the following embodiments, such as "up", "down", "left", "right", etc., are only reference to the directions of the drawings. Therefore, the directional words used are used to illustrate rather than limit the invention.

[0056] The present invention will be further described below with reference to the accompanying drawings.

[0057] Example 1: A power distribution network fault processing device, such as Figure 1 As shown, this embodiment provides a power distribution network fault processing device based on a large language model. The device integrates multiple advanced artificial intelligence technology modules to provide an efficient and accurate power distribution network fault diagnosis and processing solution, including the following main modules:

[0058] Pre-training module: Implementation method: Use self-supervised learning algorithms such as masked language model (MLM) and next sentence prediction (NSP); training data: Include large-scale text data such as historical fault records, equipment operation logs, and maintenance reports; hardware requirements: Use distributed GPU clusters such as the NVIDIA DGX SuperPOD system; training cycle: Initial training is approximately 4-6 weeks, followed by regular incremental training.

[0059] Neural network architecture module: The core architecture is based on Transformer and adopts variants of the BERT or GPT series models; the scale includes 24 layers of Transformer blocks, each with 16 attention heads; the number of parameters is approximately 175 billion parameters; the optimization algorithm is the Adam optimizer, using a learning rate warmup and linear decay strategy.

[0060] Data acquisition layer: intelligent sensor network, deploying high-precision voltage, current, and temperature sensors with a sampling rate of ≥10kHz; SCADA system, integrating existing SCADA systems to monitor the status of substations and distribution lines in real time; PMU equipment, installing phasor measurement units at key nodes with a sampling rate of ≥60Hz; weather stations, every 10km 2 Set up a micro weather station to monitor temperature, humidity, wind speed, etc., and use social media API interfaces to capture relevant complaint information from various platforms in real time.

[0061] Edge computing layer: local preprocessing unit, using FPGA to accelerate data cleaning and feature extraction, edge AI inference engine, deployment of TensorRT-optimized lightweight Transformer model, real-time monitoring panel, responsive web interface developed with React.js, and support for mobile device access.

[0062] Core processing layer: The data fusion center uses Apache Flink for real-time streaming data processing, the main AI diagnostic engine deploys a complete Transformer model, runs on NVIDIA A100 GPUs, and has a predictive maintenance module that combines time series analysis and machine learning algorithms (such as LSTM and Prophet). It also has a digital twin platform, a high-fidelity power grid simulation environment built on the Unity 3D engine, and a reinforcement learning optimizer using the Proximal Policy Optimization (PPO) algorithm.

[0063] Contextual Learning Module: Long sequence processing, using the Longformer architecture, supports input of up to 4096 tokens, and an attention mechanism to implement multi-head cross-attention and capture associations at different time scales and spatial locations.

[0064] Inference processing module: Input processing, encoding grid status information into a specially formatted prompt, inference strategy, using kernel sampling to maintain output diversity and rationality, acceleration technology, using ONNX Runtime for inference acceleration.

[0065] Output and interaction layer: fault report generator, natural language generation system based on templates and dynamic content filling, visual decision support system, interactive data visualization interface developed using D3.js, repair suggestion optimizer, integrated linear programming solver such as Gurobi, safety assessment module, risk assessment model based on Bayesian network.

[0066] Continuous learning and improvement layer: Adaptive learning module, implementation of online learning algorithms such as Online Gradient Descent, human-computer collaborative interface, web application developed based on Vue.js, support for expert feedback and manual intervention, knowledge base update system, and use of graph databases (such as Neo4j) to store and update domain knowledge.

[0067] Working process:

[0068] The system is initialized, pre-trained model parameters are loaded, various data sources are connected, data flow is initialized, and edge computing units and core processing modules are started.

[0069] Continuous data collection and monitoring: the data collection layer continuously collects grid status data, and the edge computing layer performs real-time data preprocessing and preliminary analysis.

[0070] Fault detection and diagnosis: When an anomaly is detected, a deep diagnosis process is triggered. The core processing layer integrates multi-source data and inputs it into the main AI diagnosis engine. The context learning module provides historical information support.

[0071] Fault analysis and solution generation: the reasoning processing module generates preliminary diagnostic results, the digital twin platform simulates the impact of the fault, and the reinforcement learning optimizer generates repair strategy recommendations.

[0072] Decision support and execution, the output and interaction layer generates visual reports and suggestions, and the human-computer collaborative interface allows experts to review and adjust and execute approved repair plans.

[0073] Continuous learning and optimization, collecting fault handling results and expert feedback, adaptive learning module updates model parameters, and knowledge base update system integrates new domain knowledge.

[0074] Efficient and accurate fault diagnosis, leveraging the powerful understanding capabilities of large language models, shortens diagnosis time to, intelligent repair plan generation, considers multi-dimensional factors, improves the optimality of repair plans, reduces human decision-making errors, improves system reliability, predicts maintenance capabilities, accurately predicts potential faults, improves preventive maintenance efficiency, significantly reduces unplanned power outages, and improves user satisfaction, real-time response and edge intelligence, edge computing achieves millisecond-level fault detection response, reduces the load on the central system, and improves overall system efficiency, continuous learning and knowledge accumulation, system performance continues to improve over time, fault handling efficiency is improved, and a inheritable intelligent knowledge base is formed, reducing dependence on expert experience, comprehensive situational awareness, multi-source data fusion provides comprehensive grid status awareness, improves the diagnostic accuracy of complex faults, especially rare faults, flexible human-machine collaboration, combines AI intelligence and human expertise, improves decision-making accuracy, provides an intuitive visual interface, reduces the cognitive load of operators, improves safety and reliability, multiple backup and safety mechanisms ensure system stability, and reduces economic losses caused by faults.

[0075] By integrating AI technology and power system expertise, this embodiment provides a comprehensive, intelligent, and efficient power distribution network fault handling solution, significantly improving the reliability, security, and operational efficiency of the power grid.

[0076] Example 2: A power distribution network fault processing device, such as Figure 2 As shown, the pre-training module adopts a self-supervised learning method and uses Masked State Prediction (MSP) to randomly mask 15% of the grid state data in the input sequence and train the model to predict these masked values.

[0077] Specific steps:

[0078] The grid state data is encoded into a discrete token sequence, such as [voltage_normal, current_high, temperature_normal]. 15% of the tokens are randomly selected for masking, 80% are replaced with the [MASK] tag, 10% remain unchanged, and 10% are randomly replaced with other tokens. The model training goal is to minimize the cross-entropy loss and accurately predict the masked original tokens. Next State Prediction (NSP): Given the grid state at the current time t, predict the state at time t+1.

[0079] Specific steps:

[0080] Construct (St, St+1) pairs, where St is the state at time t and St+1 is the state at time t+1; 50% of the samples use the true next state, and 50% randomly select the state at other times; the model training goal is a binary classification task, judging whether St+1 is the true next state of St;

[0081] Federated learning implementation:

[0082] Architecture: Adopts a horizontal federated learning framework and uses the FedAvg algorithm; Participants: Regional power companies act as independent data holders; Central server: Deployed in the State Grid data center, responsible for model aggregation; Communication protocol: Uses homomorphic encryption to ensure the security of model update transmission;

[0083] Training process:

[0084] The central server initializes the global model and distributes it to each participant. Each participant uses local data to train the model and calculate gradient updates. Each participant sends the encrypted gradient updates to the central server. The central server aggregates all updates and updates the global model. The updated global model is distributed to each participant. Steps 2-5 are repeated until convergence or the predetermined number of rounds is reached.

[0085] Edge computing layer:

[0086] The knowledge distillation lightweight model includes: the teacher model, a complete Transformer model with approximately 175 billion parameters; the student model, a DistilBERT architecture, with approximately 66 million parameters and a size of approximately 250MB;

[0087] Distillation process:

[0088] We use the hidden state and attention output of the teacher model as soft labels; combine the hard labels of the MSP and NSP tasks to construct a multi-task learning objective; use a softened cross-entropy loss with a temperature of 2 for knowledge transfer; and employ a progressive distillation strategy to train the student model layer by layer.

[0089] Edge AI inference engine;

[0090] Hardware platform: NVIDIA Jetson AGX Orin module; Inference optimization: Model quantization and optimization using TensorRT to achieve INT8 inference; Deployment strategy: Deploy one edge unit at each major substation to cover the surrounding distribution network;

[0091] Localization processing:

[0092] Real-time data preprocessing: noise reduction, outlier detection, and feature extraction; fast fault detection: using lightweight models for preliminary diagnosis with a latency of <10ms; local cache: storing the last 24 hours of raw data and processing results;

[0093] High-speed and reliable data transmission network;

[0094] 5G network deployment: Network architecture: 5G private network in SA (Standalone) mode; frequency band: 3.5 GHz, bandwidth: 100 MHz; base station density: one gNB (5G base station) per 10 square kilometers; network slicing: URLLC (Ultra-Reliable Low-Latency Communication) slices;

[0095] Dedicated fiber optic network: The topology adopts a star-ring hybrid architecture; the transmission technology uses DWDM (Dense Wavelength Division Multiplexing); the bandwidth is 100Gbps for backbone links and 10Gbps for access links; the redundancy design uses 1+1 protection switching for critical paths;

[0096] Network Convergence and Management: The SDN controller, based on the OpenDaylight platform, enables unified management of network resources. Intelligent routing dynamically selects 5G or fiber transmission based on data type and network conditions. QoS guarantees configure differentiated quality of service policies for different types of data flows. Security mechanisms utilize IPSec VPN to ensure end-to-end data transmission security.

[0097] The working process is as follows:

[0098] Pre-training process;

[0099] Data preparation: Each participant cleans and standardizes local power grid historical data; constructs training samples for MSP and NSP tasks;

[0100] Initialization: The central server initializes the Transformer model parameters and encrypts and distributes the initial model to all participants;

[0101] Local training: Participants use local data to train the model; perform MSP and NSP tasks simultaneously and update model parameters;

[0102] Model aggregation: Participants encrypt local model updates and send them to a central server; the central server decrypts and aggregates all updates;

[0103] Global update: The central server updates the global model and distributes the new global model to all participants;

[0104] Iterative optimization: Repeat steps 3-5 until the model performance converges or reaches the preset rounds;

[0105] Model evaluation: Evaluate global model performance on an independent test set to ensure the model's generalization capabilities across various power grid scenarios.

[0106] Edge computing process;

[0107] Knowledge distillation: Use a pre-trained large model as the teacher model; train the student model of the DistilBERT architecture to transfer core knowledge;

[0108] Edge deployment: Deploy the distilled lightweight model to edge devices; use TensorRT to optimize the model and improve inference efficiency;

[0109] Real-time data processing: Edge devices continuously receive raw data from sensors and perform real-time data cleaning, normalization, and feature extraction.

[0110] Fast fault detection: Use lightweight models to perform preliminary analysis on processed data; detect potential anomalies or fault conditions;

[0111] Local decision-making and reporting: For small-scale faults that can be handled locally, direct handling suggestions are generated; for complex faults, preliminary analysis results are reported to the central system;

[0112] Data caching and synchronization: Locally store the raw data and processing results of the last 24 hours; regularly synchronize with the central system for subsequent in-depth analysis and model updates;

[0113] The data transmission process is as follows:

[0114] Data classification: Classify data according to data type (such as real-time monitoring data, fault alarms, model updates, etc.);

[0115] Transmission strategy selection: The SDN controller selects the optimal transmission path based on data priority and network conditions. 5G URLLC slices are prioritized for critical real-time data. Large-capacity non-real-time data (such as historical data synchronization) is transmitted over the fiber network.

[0116] Secure encryption: All data is end-to-end encrypted before transmission; dynamic key management is used to ensure long-term security;

[0117] Real-time monitoring and adjustment: Continuously monitor network performance indicators (such as latency and packet loss rate); dynamically adjust network resource allocation to ensure service quality for key businesses;

[0118] Fault self-healing: When a network anomaly is detected, the backup path is automatically activated; in the event of a fiber failure, the 5G backup link is quickly switched to.

[0119] Data synchronization and consistency assurance: Use distributed transaction mechanisms to ensure data consistency across the network; perform data verification regularly to handle possible transmission errors;

[0120] Improve model performance and generalization capabilities. Through federated learning, the model can learn more diverse power grid scenarios with improved accuracy. While protecting data privacy, it realizes cross-regional knowledge sharing, improves the model's ability to identify rare faults, and brings real-time response to edge intelligence. The reasoning delay of lightweight models on edge devices is <10ms, which is faster than centralized processing. Local processing reduces 80% of data upload requirements and significantly reduces network load. The high efficiency of knowledge distillation means that lightweight models only take up 3.8% of the storage space of the original model, but retain more than 90% of the performance. The power consumption of edge devices is reduced by 70%, which greatly reduces equipment maintenance costs. Highly reliable data transmission, 5G URLLC slicing ensures that the end-to-end latency of critical data is less than 5ms, with high reliability. The fiber optic network provides high bandwidth, supports large-scale data analysis and model training, and intelligent scheduling of network resources. SDN control improves network utilization and reduces network congestion. Dynamic routing strategies ensure the service quality of critical businesses. Even in the event of partial network failures, federated learning avoids direct sharing of original data and reduces the risk of data leakage. End-to-end encryption ensures the security of the data transmission process. The edge computing architecture enables the system to be easily expanded to more distribution network nodes, reducing deployment costs. Distributed training speeds up model updates and adapts to the rapidly changing needs of the power grid. Localized processing and intelligent data transmission strategies reduce overall system energy consumption. Improved predictive maintenance capabilities extend equipment life and reduce replacement and repair frequency.

[0121] By integrating pre-training technology, edge computing, and high-speed network transmission, this embodiment not only significantly improves the efficiency and accuracy of power distribution network fault handling, but also brings significant improvements in privacy protection, energy efficiency, and system scalability. This distributed intelligent architecture lays a solid foundation for larger-scale and more complex smart grid management in the future.

[0122] Example 3: A power distribution network fault processing device,

[0123] Action generation: The Actor network generates candidate actions based on state evaluation and uses the Boltzmann exploration strategy to select the final action;

[0124] Execution and feedback: Simulate the execution of selected actions in the digital twin environment; calculate reward values and evaluate the effects of actions;

[0125] Strategy update: Use the PPO algorithm to update the strategy network and value network; store high-value experience in the priority experience replay buffer;

[0126] Continuous optimization: Regularly use accumulated experience for offline batch training to adapt to new grid configurations and operating modes;

[0127] Fixed suggestion optimizer workflow;

[0128] Input processing: receiving fault diagnosis results, resource status and system priority information; encoding input information into structured text sequences;

[0129] Contextual understanding: The T5 encoder processes the input sequence to understand the current context and uses the self-attention mechanism to capture the relationship between various elements;

[0130] Solution generation: The T5 decoder generates a preliminary sequence of repair steps; beam search is used to generate multiple candidate solutions;

[0131] Solution optimization: Evaluate the feasibility and efficiency of each candidate solution; consider multi-objective optimization criteria and select the best solution;

[0132] Natural language output: Convert the optimized solution into a clear, structured natural language description; generate step-by-step repair instructions;

[0133] Feedback and learning: Record actual implementation results and expert feedback; use the collected data for incremental learning and model adjustments;

[0134] Improve the accuracy of fault prediction. The combination of digital twins and Transformer models improves the accuracy of fault prediction, extends the average advance warning time, provides sufficient time for preventive maintenance, accelerates the fault diagnosis process, and shortens the diagnosis time of complex faults through real-time simulation and scenario analysis. Multimodal data fusion improves the recognition rate of rare faults and optimizes fault response strategies. The strategies generated by the reinforcement learning optimizer reduce the average power outage time compared with traditional methods. Dynamic decision-making capabilities increase the system's recovery speed in complex fault situations and improve repair efficiency. The solutions generated by the repair suggestion optimizer are faster than those made manually while ensuring quality. Resource utilization efficiency is improved, significantly reducing unnecessary waste of manpower and material resources, enhancing system adaptability, and the continuous learning mechanism improves system performance to a certain extent every month, continuously adapting to new power The system can handle network conditions and extreme situations significantly, successfully dealing with 95% of complex fault scenarios and improving decision transparency. The digital twin platform provides an intuitive visual interface, making the decision basis clearer. The attention visualization of the Transformer model helps understand the AI decision-making process and improves the trust of experts. The modular design makes the system easy to expand and supports the development of future smart grids. It can seamlessly integrate new energy sources (such as renewable energy) and has strong adaptability. Through simulation training, the system can better handle high-risk operations, reduce accident rates, improve defense capabilities against network attacks, and improve the overall security of the power grid. Knowledge accumulation and inheritance, continuous learning and optimization of the system have formed a valuable knowledge base, reduced dependence on individual expert experience, and promoted the systematization and standardization of knowledge.

[0135] By integrating digital twin technology, reinforcement learning, and a Transformer-based sequence-to-sequence model, this embodiment not only significantly improves the efficiency and accuracy of power distribution network fault handling, but also brings revolutionary improvements in system adaptability, decision-making transparency, and long-term optimization. This intelligent fault handling system lays a solid technical foundation for building a more reliable and efficient modern smart grid.

[0136] Example 4: A power distribution network fault processing device, such as Figure 4 As shown, the adaptive learning module of the continuous learning and improvement layer uses the Elastic Weight Consolidation (EWC) technology as the core algorithm, uses the Progressive Neural Networks (PNN) for incremental learning strategy, and uses the Apache Flink real-time stream processing framework for data stream processing. Model version control uses MLflow for model version management and experiment tracking.

[0137] Data collection and preprocessing;

[0138] Data sources: Real-time monitoring data from SCADA systems and smart sensors; fault reports, detailed troubleshooting reports completed by engineers; user feedback, user complaints and feedback collected through the customer service system; equipment maintenance records, including detailed records of scheduled inspections and emergency repairs; data cleaning, using Apache Spark for large-scale data cleaning and preprocessing; feature engineering, automatic feature generation and selection, using AutoML technology;

[0139] Model update strategy;

[0140] Update trigger mechanism: scheduled update, regular update once a week; performance trigger, when the model performance drops beyond the preset threshold, trigger update; new mode detection, update immediately when a new failure mode is detected;

[0141] Update method: Online fine-tuning makes small-scale parameter adjustments to the existing model; offline retraining, and full retraining when major changes are detected; model evaluation uses sliding window cross-validation to ensure that the new model performs better than the old model;

[0142] Multi-level backup and fault tolerance mechanism;

[0143] The data backup system uses Oracle GoldenGate for real-time data replication, with daily full cold backups stored in an off-site tape library. The system also includes a geographically distributed primary data center and at least two off-site disaster recovery centers. The system also features automated fault detection and data recovery processes, with an RTO of less than 15 minutes and an RPO of less than 5 minutes.

[0144] Fault detection and self-healing; health checks, using Prometheus for real-time system monitoring; auto-scaling to automatically adjust computing resources based on load; service mesh, using Istio service mesh for fine-grained flow control and fault isolation; circuit breaker mechanism, using Hystrix for service-level circuit breaking and degradation;

[0145] The attention visualization of the interpretability module uses the Captum library to visualize the attention weights of the Transformer model. The interactive interface is based on an interactive visualization dashboard developed with D3.js. Multi-level analysis supports word-level, sentence-level, and global-level attention analysis.

[0146] Feature importance analysis uses SHAP (SHapley Additive exPlanations) values to calculate feature importance, uses waterfall charts and force-directed graphs to display the contribution of features to prediction results, and supports real-time calculation and display of feature importance changes;

[0147] Decision path tracing, based on the Local Interpretability Model for Decision Tree Approximation (LIME), generates decision tree visualizations and natural language explanations, allowing users to explore what-if scenarios and understand the impact of input changes on outcomes;

[0148] Case library and comparative analysis, building a knowledge graph of power system faults, supporting similar case retrieval, automatically generating a comparison report between the current fault and historical cases, integrating an expert feedback mechanism, and continuously improving and verifying the accuracy of explanations;

[0149] The working process is as follows:

[0150] Continuously learn and improve processes, data collection and preprocessing: collect new data from various data sources in real time; use Apache Spark for data cleaning and feature engineering; store processed data in distributed storage systems;

[0151] Performance monitoring and update triggering: Continuously monitor the performance indicators of the model in real-world applications; trigger the update process when performance degradation or new failure modes are detected;

[0152] Incremental learning: Use the EWC algorithm to incrementally update the model; add new features through progressive neural networks while retaining the original capabilities;

[0153] Model evaluation and deployment: Use sliding window cross-validation to evaluate the performance of the new model; if the new model is better than the old model, perform a grayscale release; use MLflow to manage the model version and deployment process;

[0154] Continuous optimization: Collect feedback from model applications and continuously optimize learning strategies; regularly conduct comprehensive model diagnosis and optimization;

[0155] Interpretability module workflow;

[0156] Fault diagnosis result generation: The AI model processes the input grid status data and outputs fault diagnosis results and recommended repair solutions;

[0157] Attention weight analysis: Use Captum to extract the model's attention weights; generate multi-level attention visualization charts;

[0158] Feature importance calculation: Apply the SHAP algorithm to calculate the importance of each input feature; generate feature importance ranking and contribution charts;

[0159] Decision Path Reconstruction: Use the LIME algorithm to build a local interpretability model; generate decision tree visualization and text description;

[0160] Case comparison and knowledge graph analysis: Retrieve similar cases from the fault knowledge graph; generate a comparison report between the current fault and historical cases;

[0161] Interactive interpretation: Displays various interpretable analysis results through an interactive interface, allowing users to explore different hypothetical analysis scenarios;

[0162] Through continuous learning, the model's recognition accuracy for new fault types improves by approximately 2% per month, and its ability to handle rare faults is significantly enhanced. It can quickly adapt to changes in grid topology and new equipment, shortening the adaptation cycle from monthly to weekly. The prediction accuracy of seasonal changes is improved, effectively reducing the occurrence of seasonal faults. Multi-level backup and fault-tolerant mechanisms improve system availability and fault response speed. The explainability module makes the AI decision-making process transparent, enhances user trust, improves the efficiency of root cause analysis of error diagnosis, and accelerates the system improvement process. The construction of knowledge graphs and case libraries shortens the learning curve for novice engineers. The digitization and systematization of expert experience reduces dependence on specific individuals. Adaptive learning reduces the need for manual tuning and reduces maintenance costs. The improvement of predictive maintenance capabilities extends the average life of equipment. Due to the improvement of system reliability and response speed, customer satisfaction increases, transparent fault explanations reduce user complaints, and customer service efficiency is improved. The explainability module provides strong support for regulatory review and accelerates the compliance review process.

[0163] Example 5: A power distribution network fault processing device, such as Figure 5 shown

[0164] S1: Large-scale pre-training:

[0165] Data preparation;

[0166] Data sources: historical fault records, equipment operation logs, power grid topology information, meteorological data, etc.; the data volume is approximately 10TB of structured and unstructured data; Apache Spark is used for preprocessing to perform large-scale data cleaning and standardization;

[0167] Model architecture;

[0168] Base model: uses a variant of the GPT-3 architecture with 175B parameters; optimization: uses mixed precision training and model parallelism;

[0169] training strategies;

[0170] Target tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP); Training cycle: 4 weeks, approximately 1 trillion tokens; Learning rate strategy: warmup + cosine decay; Regularization: weight decay and dropout;

[0171] S2: Model fine-tuning and knowledge distillation:

[0172] Fine-tuning process;

[0173] Local data: Historical operating data of a specific power grid, approximately 100GB; Fine-tuning strategy: Use the AdamW optimizer with a learning rate of 1e-5; Task adaptation: Add task-specific output layers (such as fault classification and severity prediction);

[0174] Knowledge distillation;

[0175] Teacher model: fine-tuned large model; Student model: 12-layer Transformer, 80M parameters; Distillation method: soft label distillation + intermediate layer feature matching; Temperature parameter: T = 2.0;

[0176] S3: Real-time data collection:

[0177] Data sources: SCADA systems: real-time measurements of voltage, current, power factor, etc.; PMU devices: high-precision phasor data; weather stations: environmental data such as temperature, humidity, and wind speed; social media: user-reported power outage information;

[0178] Data integration: Use Apache Kafka for real-time data stream processing; Data standardization: Unify the format and units of data from different sources; Time synchronization: Use PTP (Precision Time Protocol) to ensure the accuracy of data timestamps;

[0179] Serialization: Convert multimodal data into a unified token sequence; use a custom vocabulary that includes power system-specific terminology; maximum sequence length: 2048 tokens;

[0180] S4: Edge computing preprocessing:

[0181] Deployment location: each major substation and key distribution node;

[0182] Lightweight model: Architecture: MobileNetV3 + small Transformer (4 layers, 4 attention heads); Model size: <50MB, suitable for edge deployment; Quantization: INT8 quantization, using TensorRT optimization;

[0183] Preliminary fault detection: Anomaly detection uses the Local Outlier Factor (LOF) algorithm; rapid classification provides preliminary classification of common fault types; processing delay is <10ms;

[0184] S5: Core AI Diagnosis:

[0185] Data fusion: multimodal fusion using attention mechanism; implementing the cross-attention layer of Transformer;

[0186] In-depth analysis: Fault location based on grid topology analysis using graph neural networks; multi-label fault classification, supporting complex fault identification; severity assessment regression tasks, predicting the scope and duration of impact;

[0187] Diagnostic Engine: Core model is a Transformer model with 175B parameters; inference optimization uses the NVIDIA Triton inference server; processes 100 complex diagnostic requests per second;

[0188] S6: Contextual Learning:

[0189] Long sequence processing; Technology: Longformer architecture, supports up to 16K tokens; Sliding window: retains relevant historical information for the last 7 days;

[0190] Attention mechanism; global attention focuses on key devices and abnormal patterns; local attention captures local patterns in time series;

[0191] State representation; dynamically embed vector representations of grid states that are updated in real time; and construct dynamic graph structures that reflect grid topology and temporal evolution using spatiotemporal graphs.

[0192] S7: Digital Twin Simulation:

[0193] Virtual model, the software uses Siemens PSS / E to simulate the power system, with a simulation accuracy of 0.1% and support for millisecond time steps;

[0194] Real-time synchronization, method: using Kalman filter for state estimation; frequency: synchronizing with the actual grid state every 5 seconds;

[0195] Scenario simulation; Monte Carlo simulation: generates 10,000 possible fault evolution scenarios; parallel computing: uses CUDA acceleration to support real-time scenario evaluation;

[0196] S8: Reinforcement Learning Optimization:

[0197] Environmental model; state space: grid topology, equipment status, load distribution, etc.; action space: switch operation, load transfer, power generation adjustment, etc.; reward function: comprehensive consideration of power supply reliability, economy, and equipment life;

[0198] Algorithm selection; Core algorithm: Proximal Policy Optimization (PPO); Neural network: Actor-Critic architecture combined with Transformer encoder; Exploration strategy: Exploration using parameter space noise;

[0199] Training process; Simulation environment: Deep integration with the digital twin platform; Training scale: Parallel training on 64 A100 GPUs; Experience replay: Prioritize experience replay, with a capacity of 10M experiences;

[0200] S9: Output generation:

[0201] Fault report generation; Method: Use T5 model for text generation; Personalization: Generate reports with different levels of detail based on user roles (technician / management);

[0202] Visual decision support; Tools: Use D3.js to develop an interactive visualization interface; Content: Fault location heat map, impact range prediction, repair priority sorting, etc.;

[0203] Multiple solution generation; Strategy: Use beam search to generate multiple candidate solutions; Evaluation: Multi-objective optimization to balance repair time, resource consumption, and reliability;

[0204] S10: Human-Robot Collaboration

[0205] Expert review interface; Interface: Web-based interactive review platform; Functions: scheme comparison, parameter adjustment, manual intervention;

[0206] Interpretability analysis; Method: Use SHAP (SHapley Additive exPlanations) values to explain model decisions; Visualization: Attention weight heat map, decision tree visualization;

[0207] Knowledge graph integration; Construction: Neo4j-based power system knowledge graph; Application: Assist decision-making explanation and provide relevant historical cases;

[0208] S11: Continuous Optimization:

[0209] Feedback collection; Source: actual repair results, expert evaluation, system performance indicators; Storage: Use the time series database InfluxDB to store performance data;

[0210] Adaptive learning; Method: Online gradient descent, dynamic adjustment of model weights; Frequency: Incremental updates weekly, full evaluation monthly;

[0211] Knowledge base update; Process: Automatically extract new knowledge and update the knowledge graph after manual review; Verification: Use A / B testing to verify the effectiveness of new knowledge;

[0212] The working process is as follows:

[0213] System initialization: Load the pre-trained and fine-tuned large Transformer model; initialize the digital twin platform and synchronize the current power grid status; start all edge computing nodes and data acquisition systems;

[0214] Continuous monitoring and data collection: S3: Real-time collection of multi-source, multi-modal grid status data; S4: Edge nodes perform preliminary data processing and anomaly detection;

[0215] Fault detection and preliminary analysis: When an edge node detects a potential fault, it triggers a deep analysis process; S5: The core AI diagnostic engine integrates all available data for comprehensive analysis; S6: The contextual learning module integrates historical information to provide long-term trend insights;

[0216] In-depth diagnosis and scenario simulation: S7: The digital twin platform generates multiple possible fault evolution scenarios based on the current status; simulates the effects of different intervention measures, and assesses potential risks and impacts;

[0217] Repair strategy generation: S8: The reinforcement learning optimizer combines real-time status and simulation results to generate a repair strategy; considering multiple objectives (such as repair time, cost, and reliability), it generates a Pareto optimal solution set;

[0218] Results presentation and human-computer interaction: S9: Generates detailed fault reports, visual decision support information, and multiple possible solutions; S10: Presents results through a human-computer collaborative interface, allowing experts to review and adjust; provides explainable analysis of decision-making basis to help experts understand the AI's reasons for recommendations;

[0219] Execution and feedback: Execute the selected repair plan and monitor the execution effect in real time; collect actual processing results and expert feedback;

[0220] Continuous Learning and Optimization: S11: Update model parameters and knowledge base based on collected feedback; regularly evaluate overall system performance and make necessary adjustments and optimizations;

[0221] Comprehensive fault prevention and rapid response, precise fault diagnosis, optimized repair strategies, enhanced decision support, improved system reliability, predictive maintenance capabilities, reduced unplanned downtime, knowledge accumulation and inheritance, shortened training time for new engineers, rapid mastery of complex fault handling skills, forming a comprehensive power grid fault knowledge base, capturing 90% of historical fault modes, adaptability and scalability, the system can adapt to new grid equipment or topology changes within 24 hours, support seamless integration of new energy equipment, improve the management capabilities of renewable energy, reduce carbon emissions annually by optimizing grid operation and rapid fault repair, improve the grid connection efficiency of renewable energy, and increase the proportion of clean energy use.

[0222] By implementing this comprehensive power distribution network fault handling process, the system not only significantly improves the reliability and efficiency of the power grid, but also brings huge economic and environmental benefits.

[0223] In view of current actual needs, the protection scope of the above-mentioned implementation mode adopted in this application is not limited to this. Various changes made within the knowledge scope of technical personnel in this field without departing from the concept of this application still fall within the protection scope of the present invention.

Claims

1. A power distribution network fault handling device, characterized by: It includes pre-training module, neural network architecture module, data acquisition layer, edge computing layer, core processing layer, context learning module, reasoning processing module, output and interaction layer, and continuous learning and improvement layer; The pre-training module is used to perform unsupervised learning on large-scale power system historical data to learn the statistical laws and semantic relationships of power grid operation; The neural network architecture module is based on the Transformer architecture and includes a multi-layer self-attention mechanism and a feedforward neural network with billions to hundreds of billions of trainable parameters; The data acquisition layer includes smart sensor networks, SCADA systems, PMU devices, weather stations, and social media API interfaces; The edge computing layer includes a local pre-processing unit, an edge AI inference engine, and a real-time monitoring panel; The core processing layer includes a data fusion center, a main AI diagnostic engine, a predictive maintenance module, a digital twin platform, and a reinforcement learning optimizer; The contextual learning module is able to understand and utilize long sequences of grid status information and focus on relevant information through an attention mechanism; The reasoning processing module receives the grid status information as input prompts and gradually generates fault diagnosis and treatment suggestions; The output and interaction layer includes a fault report generator, a visual decision support system, a repair suggestion optimizer, and a safety assessment module; The continuous learning and improvement layer includes an adaptive learning module, a human-computer collaboration interface, and a knowledge base update system; The core processing layer's digital twin platform can create a virtual model of the power grid, conduct real-time simulations and what-if scenario analysis, and integrate with the Transformer model to enhance contextual understanding and prediction capabilities. The reinforcement learning optimizer of the core processing layer can dynamically optimize fault response strategies and system parameter adjustments through interactive training with the simulated environment of the digital twin platform. Its decision-making process uses the attention mechanism of the Transformer model to weigh multiple factors.

2. The power distribution network fault handling device according to claim 1, characterized in that: The pre-training module adopts a self-supervised learning method, including masked state prediction and next state prediction tasks, to learn the inherent laws of the power grid system, and realizes distributed training of the model through federated learning technology, thereby improving the accuracy and generalization ability of the model while protecting data privacy.

3. The power distribution network fault handling device according to claim 1, characterized in that: The edge AI inference engine of the edge computing layer is deployed with a lightweight model that has undergone knowledge distillation for fast preliminary fault detection and localization processing. The lightweight model inherits the core knowledge of the large-scale pre-trained model.

4. The power distribution network fault handling device according to claim 1, characterized in that: The repair suggestion optimizer of the output and interaction layer adopts a sequence-to-sequence model based on Transformer, which can generate the optimal repair order and plan according to the current resource status and priority.

5. The power distribution network fault handling device according to claim 1, characterized in that: The adaptive learning module of the continuous learning and improvement layer can continuously update and optimize the parameters of the AI model through incremental learning and fine-tuning technology based on newly collected data and feedback information, so as to keep the model adaptable to the latest power grid conditions.

6. The power distribution network fault handling device according to claim 1, characterized in that: It also includes a high-speed and reliable data transmission network based on 5G or dedicated fiber optic networks to ensure real-time data exchange between all layers of the system and support joint training and inference of large-scale distributed models.

7. The power distribution network fault handling device according to claim 1, characterized in that: It also includes multi-level backup and fault-tolerant mechanisms, as well as strict data security protocols to ensure the high availability and data security of the system; at the same time, it integrates an interpretability module to provide understandable explanations for fault diagnosis results by analyzing the model's attention weights and activation values.

8. The power distribution network fault handling device according to any one of claims 1 to 7, characterized in that: The power distribution network fault processing method includes the following steps: S1: Pre-training the Transformer model on large-scale power system historical data to learn general knowledge about power grid operation; S2: Fine-tune and distill knowledge of the pre-trained model using local data for a specific power grid environment; S3: Continuously collects multi-source, multi-modal real-time grid status information through the data acquisition layer and converts it into a tokenized input sequence; S4: Use lightweight models for data preprocessing and preliminary fault detection at the edge computing layer; S5: In the core processing layer, all input data is integrated and processed, and the main AI diagnostic engine is run to perform in-depth fault analysis. S6: The contextual learning module integrates long-sequence historical information to build a comprehensive understanding of the current grid status; S7: Use digital twin platforms for fault simulation and impact assessment; S8: Generate the optimal repair strategy by combining the reinforcement learning optimizer with the attention mechanism of the Transformer model; S9: Generates fault reports, visual decision support information, and multiple possible treatment plans at the output and interaction layer; S10: Allow experts to review and intervene in AI decisions through a human-machine collaborative interface, and provide explainable analysis of the decision-making basis; S11: Collect feedback on actual processing results, use adaptive learning modules and knowledge base to update the system, and continuously optimize system performance.

Citation Information

Patent Citations

  • A method for condition diagnosis and prediction of power distribution network equipment based on power big data

    CN108320043B

  • A method and system for handling distribution network faults based on digital twin technology

    CN116885858B

  • Network fault management and control method based on parallel network large model digital expert

    CN118214676A

  • Dynamic prediction and optimization control method for power grid line loss driven by deep learning

    CN118469352A

Cited By

  • Electrical equipment fault diagnosis early warning system and method based on federated learning and hybrid intelligence

    CN122087675A