A power distribution network fault identification method and system based on data analysis

The defense system built using digital twin models and immune learning algorithms solves the problem of identifying small-sample faults in power distribution networks, achieving high-precision and highly generalizable fault identification, and improving the safety and operation and maintenance efficiency of power distribution networks.

CN122333773APending Publication Date: 2026-07-03STATE GRID JIBEI ELECTRIC POWER COMPANY LIMITED CHENGDE POWER SUPPLY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIBEI ELECTRIC POWER COMPANY LIMITED CHENGDE POWER SUPPLY
Filing Date
2026-04-08
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing power distribution network fault identification technologies rely on a balanced and complete fault sample library, which faces the problem of small sample size, making training difficult. They also lack the ability to generalize the identification of unknown faults. The fault diagnosis process is disconnected from the physical evolution mechanism, resulting in low identification accuracy and high false negative rate.

Method used

An immune defense system based on a digital twin model is constructed. By generating a self-sample set and training a mature detector with an adaptive immune learning algorithm, and combining a virtual-real fusion evolutionary inference mechanism, real-time identification of abnormal signals and tracing of their physical causes can be achieved.

Benefits of technology

It improves the accuracy and generalization ability of distribution network fault identification, reduces the false negative rate, realizes the accurate identification of unknown faults and the prediction of fault development paths, and enhances the predictability and scientific nature of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333773A_ABST
    Figure CN122333773A_ABST
Patent Text Reader

Abstract

This invention relates to the field of power system automation and fault prediction technology, and discloses a data analysis-based method and system for distribution network fault identification. The method includes: constructing a digital twin model of the distribution network to generate a self-sample set of normal operating conditions; training a set of mature detectors using an adaptive immune learning algorithm based on the self-sample set to define non-self regions; real-time acquisition and preprocessing of field operating data, inputting it into the detector set for anomaly identification; mapping the identified non-self signals back to the digital twin model, and determining the physical causes and development paths of faults through antigen epitope correlation analysis and evolutionary deduction. The system includes a digital twin simulation module, an immune learning processing module, a real-time acquisition and preprocessing module, a state matching monitoring module, and a fault deduction and correlation analysis module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system automation and fault prediction technology, specifically relating to a data analysis-based method and system for identifying distribution network faults. Background Technology

[0002] As a crucial link in the power system for transmitting electricity to end users, the safety and stability of the distribution network directly affect the continuous guarantee of socio-economic activities and the quality of life for the people. With the in-depth advancement of the energy internet construction, the topology of the distribution network is becoming increasingly complex and highly dynamic, placing higher demands on the accurate perception of system status and the timely detection of abnormal risks. Establishing an efficient operation and maintenance system has become a major task in the development of smart grids.

[0003] Data-driven power distribution network fault identification technology monitors electrical quantities such as voltage and current during line operation in real time, and uses mathematical modeling and pattern recognition to capture and classify abnormal states. This technology aims to replace traditional manual inspection methods by deeply mining massive amounts of operational data, quickly locating fault sources and providing early warning support in complex power grid environments, thus ensuring the continuity and reliability of the power supply system.

[0004] Existing technologies still face many problems in the field of distribution network fault identification: traditional deep learning models are highly dependent on a balanced and complete fault sample library, while rare fault data in actual distribution network operation is extremely scarce. This small sample problem makes it difficult for the model to be fully trained, which restricts the identification accuracy; existing identification mechanisms mostly focus on matching known fault types and lack the ability to generalize identification to deviations from normal states, which easily leads to missed detection when facing new or unknown faults generated in power grid operation; the fault diagnosis process is disconnected from the actual physical evolution mechanism, and there is a lack of virtual-real fusion methods to conduct in-depth correlation analysis of abnormal features, making it difficult to adaptively extrapolate the development path of faults. These problems result in the low efficiency of the distribution network security defense system in dealing with sudden and hidden faults. Summary of the Invention

[0005] The purpose of this invention is to provide a data analysis-based method and system for fault identification in distribution networks, which can solve the problems mentioned in the background. Addressing the structural pain points of distribution networks, such as difficulty in fault identification under small sample conditions, high false negative rates for unknown faults, and the disconnect between fault diagnosis and physical evolution mechanisms, this invention constructs a defense system with adaptive evolutionary prediction capabilities by deeply integrating artificial immune system mechanisms and digital twin simulation technology.

[0006] To achieve the above objectives, this invention proposes a data analysis-based method for fault identification in power distribution networks, comprising the following steps: Step 1: Construct a digital twin model of the distribution network and perform multi-condition simulations using the digital twin model to generate a self-sample set representing the normal operating state of the distribution network; Step 2: Based on the self-sample set, use the adaptive immune learning algorithm to train and generate a mature detector set, where the mature detector set is used to define the non-self region in the distribution network operation state space; Step 3: Acquire real-time field operation data of the distribution network, and after preprocessing, input it into a set of mature detectors for real-time matching and anomaly identification; Step 4: For the identified non-self-abnormal signals, map them back to the digital twin model for antigen epitope association analysis, and determine the physical cause and development path of the fault based on the evolutionary inference mechanism of virtual-real fusion.

[0007] Preferably, step 2 specifically includes the following steps: 21. A predetermined number of immature detectors are randomly generated within the feature space of the distribution network operation status. Each immature detector is represented by a hypersphere composed of a center vector and a preset coverage radius. 22. Compare the tolerance of the immature detector with the self-sample set, and calculate the Euclidean distance between the center vector of the immature detector and each data point in the self-sample set; 23. When the distance between an immature detector and any self-sample is less than a preset tolerance threshold, the detector is determined to be self-activated and eliminated. 24. For immature detectors that are not eliminated within the preset tolerance period, convert them into mature detectors and add them to the mature detector set.

[0008] Preferably, step 1, which involves constructing a digital twin model and generating a self-sample set, specifically includes: acquiring the physical topology parameters, line impedance parameters, transformer ratio information, and load characteristic distribution of the distribution network; constructing a quasi-steady-state simulation engine that is highly mapped to the physical entity in the digital space; setting the typical logic of load fluctuation step size, distributed power generation output randomness, and switching operation based on the historical operating range of the distribution network, and performing continuous power flow calculation under fault-free constraints; extracting the voltage vector, branch current vector, active power, and reactive power of each node as original features, and generating a self-sample set distributed on a specific manifold through dimensionality reduction processing.

[0009] Preferably, the adaptive immune learning algorithm further includes detector affinity evolution logic: calculating the overlap between detectors in the mature detector set; for detectors with an overlap exceeding a predetermined proportion, performing local search and radius compression through a cloning selection mechanism to ensure that the coverage rate of the mature detector set to the non-self space reaches a specific gradient, while minimizing detection holes.

[0010] Preferably, step 3, which involves acquiring and preprocessing the real-time operational data of the distribution network, specifically includes: acquiring current and voltage signals in real time through the distribution automation terminal and synchronous phasor measurement unit; performing sliding window sampling on the original signals and applying Fourier transform to extract each harmonic component and fundamental wave characteristics; using normalization processing to map electrical parameters of different dimensions to a selected value range to eliminate system deviations caused by differences in equipment sampling; and determining the spatial position of the preprocessed real-time data vector with each hypersphere in the mature detector set.

[0011] Preferably, step 4, the antigen epitope association analysis, specifically includes: extracting the key deviation direction of the identified non-self abnormal signal in the feature space and defining it as the antigen epitope feature vector; mapping the antigen epitope feature vector to the physical parameter layer of the digital twin model through a sensitivity matrix; iteratively adjusting physical parameters such as line insulation level, grounding resistance, and abnormal fluctuation of branch load in the digital twin model until the cosine similarity between the feature vector output by the simulation and the antigen epitope feature vector reaches a preset consistency threshold; and determining the combination of physical parameters in the digital twin model at this time as the physical cause of the fault.

[0012] Preferably, the evolutionary deduction mechanism based on virtual-real fusion in step 4 specifically includes: starting an accelerated simulation mode in the digital twin model when the physical cause of the fault is determined as the initial disturbance; predicting the spread trend of the fault current in the spatial topology based on the preset relay protection action logic and the automatic reclosing strategy of the distribution network; calculating the voltage sag depth and duration of the nodes around the fault point, deducing the probability distribution of the fault evolving from transient to permanent, and generating targeted operation and maintenance isolation suggestions.

[0013] Preferably, the update mechanism of the mature detector set includes: periodically acquiring newly added normal operation data of the distribution network and using it as an increment to perform secondary tolerance verification on the existing mature detector set; if a mature detector collides with the increment, the detector is immediately removed or adjusted to ensure the adaptive capability of the fault identification system to changes in the distribution network topology.

[0014] Preferably, a data analysis-based distribution network fault identification system implementing the above method includes: a digital twin simulation module for receiving distribution network physical parameters and generating a self-sample set; an immune learning processing module connected to the digital twin simulation module, configured to execute a negative selection algorithm and clonal selection evolution to generate and maintain a mature detector set; a real-time acquisition and preprocessing module for acquiring electrical operating parameters from the distribution network site and converting them into standard feature vectors; a state matching monitoring module, whose input is connected to the real-time acquisition and preprocessing module and whose reference is connected to the immune learning processing module, for determining whether the current operating state falls into a non-self region; and a fault inference and correlation analysis module connected to the state matching monitoring module and the digital twin simulation module, for performing antigen epitope mapping and physical evolution simulation after an anomaly is identified.

[0015] Preferably, the digital twin simulation module has a built-in high-performance matrix operation engine, which supports the reconstruction of the power flow section of the entire network topology within seconds, and can synchronize the topology switch status in real time according to the operation instructions of the power grid dispatching system.

[0016] Preferably, the state matching monitoring module adopts a multi-threaded parallel computing architecture, which supports the distribution of a large number of mature detectors on multiple computing units to achieve parallel retrieval and distance determination of high-frequency real-time data, ensuring that the real-time performance of fault identification meets the requirements of safe and stable operation of the power system.

[0017] Preferably, the fault deduction and correlation analysis module is equipped with an expert knowledge base matching unit, which can compare the parameter inversion results of the digital twin model with typical fault modes such as single-phase grounding, phase-to-phase short circuit, and resonant overvoltage, and output a structured fault diagnosis report.

[0018] Compared with existing technologies, the beneficial effects of this invention are as follows: By constructing an immune defense system based on digital twins, this invention cleverly solves the sample bottleneck and generalization ability problems in distribution network fault identification. Utilizing digital twin technology to generate large-scale self-samples of normal states in virtual space solves the problem of small-sample identification due to the extreme scarcity of fault samples in actual operation, making model training no longer dependent on difficult-to-obtain real fault data. Drawing on the non-self-recognition mechanism of artificial immune systems, the system only needs to learn normal states to define the entire anomaly space, which endows the system with extremely strong generalization ability, enabling it to identify any unknown faults or novel anomalies that deviate from the normal baseline, reducing the false negative rate in distribution network security defense. By mapping the identified anomaly features back to the digital twin model for evolutionary deduction, the barrier between traditional data analysis and physical mechanisms is broken down, realizing a leap in fault diagnosis from pure data classification to physical cause tracing and development path prediction, improving the predictability and scientific nature of distribution network operation and maintenance.

[0019] Furthermore, the self-sample set generated by the digital twin simulation module has extremely high coverage and completeness. By introducing various uncertainties such as load fluctuations and power output randomness during the simulation process, the self-samples can accurately characterize the normal operation boundaries of the distribution network under various complex operating conditions. This provides a solid and dynamic data foundation for the training of the immune detector, ensuring that the detector can accurately distinguish between normal grid fluctuations and real faults, thus reducing the false alarm rate of the system.

[0020] Furthermore, the adaptive negative selection algorithm employed in the immune learning processing module can automatically construct a detection grid without blind spots in the feature space through random distribution and tolerance screening. This detector structure based on hypersphere representation not only has low computational overhead but is also easy to expand and prune spatially. Affinity evolution achieved through clonal selection mechanism allows the detector ensemble to adaptively adjust its coverage density according to the complexity of the non-self space, improving the processing efficiency of large-scale high-dimensional data while ensuring detection accuracy.

[0021] Furthermore, the real-time matching logic implemented by the state matching monitoring module can quickly determine the spatial location of the operational data at each moment. Since the mature detector set has pre-excluded all self-regions, any sample point falling within the detector coverage area is considered a potential risk. This detection logic does not rely on specific fault feature matching, but rather on a process of elimination to monitor all anomalies comprehensively. This has significant technical advantages for identifying highly concealed early and subtle faults.

[0022] Furthermore, the antigen epitope mapping mechanism introduced in the fault simulation and correlation analysis module transforms the originally abstract mathematical anomaly characteristics into specific physical parameter changes. Through iterative parameter optimization in the digital twin model, the system can explain the physical nature of the current anomalous signals. This virtual-real fusion diagnostic approach provides maintenance personnel with intuitive and interpretable decision-making support. The evolutionary simulation function supports advanced simulation of future fault propagation trends, which buys valuable processing time for risk isolation and load transfer in the distribution network, reducing the scope of power outages and economic losses.

[0023] Furthermore, the system's adaptive update mechanism ensures that the fault identification logic keeps pace with the topology evolution of the physical distribution network. When the distribution network undergoes expansion, line upgrades, or significant adjustments to its operation, the digital twin model can quickly synchronize with the physical changes and update its own sample set, thereby automatically adjusting the detector set through incremental immune learning. This closed-loop self-evolutionary feature solves the problem of significant performance degradation in traditional fault identification schemes after changes in the power grid structure, ensuring the stability and reliability of the system during long-term operation.

[0024] Furthermore, by employing a multi-threaded parallel computing architecture and a high-performance matrix operation engine, this invention demonstrates excellent concurrency performance when processing massive electrical data. The system can complete the entire process from feature extraction and spatial matching to evolutionary deduction within a short sampling period, meeting the stringent requirements of modern smart distribution networks for sub-second fault early warning and response, and providing core technical support for building a highly reliable and resilient distribution network operation and maintenance system.

[0025] Furthermore, by introducing an expert knowledge base matching unit, the system can deeply integrate the correlation analysis capabilities of deep learning with traditional power system professional experience. Based on the physical causes output by the digital twin model, the system combines expert experience to perform multi-dimensional verification of the fault nature, improving the identification accuracy in complex cross-fault scenarios and ensuring the authority and practical value of the output early warning information. In summary, this invention, through the synergistic effect of digital twins and an artificial immune system, achieves high-precision, high-generalization, and deep-correlation fault identification in distribution networks under small sample constraints, which has significant technological advancements for ensuring the safe operation of power systems. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the overall technical solution architecture of the present invention; Figure 2 This is a schematic diagram of the core principle framework of the mature detector generation and evolution based on the adaptive immune learning algorithm in this invention; Figure 3 This is a flowchart illustrating the logical process of generating a self-sample set based on a digital twin model in this invention. Figure 4 This is a flowchart illustrating the logical process of real-time preprocessing of field operation data and spatial matching of detectors in this invention. Figure 5 This is a schematic diagram illustrating the physical cause tracing and evolution of faults based on antigen epitope association analysis and virtual-real fusion mechanism in this invention. Detailed Implementation

[0027] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0028] Example 1: This example provides a data analysis-based method for identifying faults in a power distribution network. Please refer to the appendix. Figure 1 To be continued Figure 5 This method is based on the complex operating environment of modern smart power distribution networks. It aims to build a dynamic defense system that can be self-sensing, self-identifying, and has the ability to trace back physical mechanisms by deeply coupling digital twin technology with artificial immune mechanisms.

[0029] In step 1, a digital twin model of the distribution network is constructed. This process is not simply geometric modeling, but rather a deep integration of multi-dimensional information about the physical entities. Physical topology parameters of the distribution network are obtained through distribution automation systems and geographic information systems. These parameters cover the connection relationships of all switching nodes, the branch distribution from substation outlets to end load points, and specific cable physical specifications. Detailed line impedance parameters are entered, including DC resistivity under different temperature environments, AC resistance increments due to the surface-to-surface effect, and susceptance and reactance parameters determined by the geometric mean distance between conductors. For transformers, the system needs to obtain their rated turns ratio, no-load and short-circuit loss parameters, and the logical constraints of tap changer adjustments. At the load end, a load characteristic distribution model is established using long-term monitoring data, including the fluctuation patterns of active and reactive power demand under different time periods and weather conditions.

[0030] Based on the aforementioned physical data, a quasi-steady-state simulation engine is constructed in the digital space. This engine is capable of performing high-precision power flow calculations. According to the historical operating range of the distribution network, load fluctuation steps are set, typically with a granular time unit of 15 minutes or finer, to simulate the transition between peak and off-peak electricity consumption. For distributed generation, the uncertainty of photovoltaic and wind power output is simulated by setting random offsets in the power prediction curve. Typical logic of switchgear switching operations also needs to be simulated, such as the closing of tie switches and the opening of sectionalizing switches. Continuous large-scale power flow calculations are performed under conditions ensuring no short-circuit, grounding, or open-circuit faults. Through these calculations, the voltage vector of each node is extracted, i.e., the effective value of the voltage and its phase angle relative to a reference base. Simultaneously, the current vector, active power flow direction, and reactive power distribution of each branch are extracted. These data constitute the original characteristics of the distribution network under healthy operating conditions.

[0031] To eliminate redundancy and retain key state features, the extracted high-dimensional original features are dimensionality-reduced. A nonlinear manifold learning logic is employed to project massive electrical features into a low-dimensional manifold space, forming a self-sample set that accurately characterizes the normal operating boundaries of the distribution network. These self-sample points exhibit a specific clustered distribution in the multi-dimensional feature space, representing the state trajectory of the distribution network under all legal operating conditions.

[0032] In step 2, based on the self-sample set generated above, an adaptive immune learning algorithm is used to train and generate a set of mature detectors. This process simulates the recognition logic of a biological immune system for foreign objects, defining normal operating states as self-identical and potential faults and anomalies as non-self-identical.

[0033] In step 21, a predetermined number of immature detectors are randomly generated within the feature space of the distribution network's operating status. Each immature detector is defined as a hypersphere structure, characterized by a center vector and a preset coverage radius. The center vector represents the detector's absolute position in the feature space, while the radius defines the detector's monitoring range.

[0034] In step 22, the generated immature detector is compared with the self-sample set for tolerance. To determine whether the detector will falsely trigger the normal state, the spatial distance between the center vector of the immature detector and each data point in the self-sample set needs to be calculated. The component values ​​of the detector center vector in each dimension are obtained, and the difference is calculated with the component values ​​of the corresponding self-sample points in the same dimension. The quadratic operation is performed on all the differences, and the quadratic operation results of all dimensions are summed. Finally, the square root of the sum is taken to obtain the scalar distance parameter reflecting the geometric displacement between the two.

[0035] In step 23, a self-activation determination is performed. The scalar distance parameter calculated above is compared with the system's preset tolerance threshold. If the scalar distance parameter is less than the preset tolerance threshold, it means that the detector's monitoring range covers the normal operating area, i.e., self-activation has occurred. To prevent the system from generating false alarms under normal operating conditions, the detector must be eliminated at this time and not allowed to proceed to the subsequent maturation stage.

[0036] In step 24, for immature detectors that have undergone multiple self-alignments within a preset tolerance period without being eliminated, it is determined that they are located in a non-self region in the feature space, i.e., a region where faults may occur. At this point, the system converts them into mature detectors and writes their relevant parameters into the database of the mature detector set.

[0037] To further optimize detector distribution efficiency, this embodiment also introduces detector affinity evolution logic. The system periodically calculates the overlap between the hyperspheres of each detector in the mature detector set. Specifically, it calculates the distance between the center vectors of two detectors and compares it with the sum of the radii of the two detectors. If the center distance is less than the sum of the radii, and the volume of the overlapping part exceeds a predetermined proportion, redundant coverage is determined to exist. At this time, a cloning selection mechanism is used to perform local search and radius compression on the detectors in the overlapping area. This dynamic adjustment ensures that the mature detector set can achieve maximum coverage of the non-self operating space with minimal computational redundancy, thereby establishing a tight monitoring grid without blind spots.

[0038] In step 3, real-time on-site operational data of the distribution network is acquired and preprocessed. High-frequency current and voltage raw signals are collected in real time using equipment installed at feeder automation switches, distribution transformer monitoring terminals, and synchronous phasor measurement units. These raw signals contain rich information reflecting the state of the power grid.

[0039] The preprocessing process first involves sliding window sampling of the original signal to ensure timely analysis. Fourier transform logic is then executed to decompose the acquired time-domain sequence. Signal features are correlated with basis functions of different frequencies to extract the effective value of the fundamental wave, its initial phase, and the amplitude distribution of each harmonic component. These components are crucial for identifying nonlinear load fluctuations or arc grounding faults. Normalization is then performed. Due to the significant differences in the dimensions and value ranges of physical quantities such as current, voltage, and power, the system employs feature scaling logic to map all electrical parameters to a selected value range, such as a closed interval between zero and one. This step effectively eliminates systematic biases caused by differences in sensor sampling accuracy and varying equipment ranges.

[0040] The preprocessed real-time data is encapsulated into a standardized real-time data vector and input into a set of mature detectors. The system then uses this vector to determine its spatial position relative to the hypersphere of each mature detector in the set. The core logic of this determination remains based on the aforementioned distance calculation, specifically whether the current real-time state point falls within the coverage radius of any detector. Once the real-time vector falls within the hypersphere of a detector, the system immediately triggers an anomaly detection signal.

[0041] In step 4, antigen epitope association analysis and evolutionary deduction are performed on the identified non-self abnormal signals. This step is the core of the present invention for achieving virtual-real fusion diagnosis.

[0042] First, the key deviation direction of the identified non-self-abnormal signal is extracted in the feature space. This deviation direction is defined as the epitope feature vector, which describes the deviation characteristics of the current abnormal state relative to the normal state manifold, such as a sudden drop in voltage amplitude or a surge in specific harmonic content. Next, this epitope feature vector is mapped to the physical parameter layer of the digital twin model using a sensitivity matrix. The sensitivity matrix characterizes the transfer function relationship between small changes in physical parameters and changes in electrical signal characteristics.

[0043] In the digital twin model, iterative optimization logic is initiated. The system automatically attempts to adjust the combination of physical parameters in the model, including but not limited to the insulation level of the line, the magnitude of the grounding resistance, and the abnormal fluctuation amplitude of the load on specific branches. For each parameter adjustment, the digital twin engine reruns the power flow simulation and outputs the corresponding simulation feature vector. The system calculates the degree of convergence between the simulation feature vector and the real-time acquired epitope feature vector. The specific degree of convergence is determined by cosine similarity logic: the components of the two vectors in corresponding dimensions are multiplied and summed, and the result is divided by the product of the magnitudes of the two vectors. When this ratio reaches a preset consistency threshold, the current simulated physical state is considered to have successfully reproduced the fault behavior in the field. This combination of physical parameters is identified as the physical cause of the fault, realizing reverse tracing from data anomalies to the physical essence.

[0044] After identifying the physical cause of the fault, a virtual-real integrated evolutionary simulation mechanism is initiated. Using this physical cause as the initial disturbance, an accelerated simulation mode is activated in the digital twin model. This mode considers the preset relay protection operation logic of the distribution network, such as the operating threshold, time-step configuration, and automatic reclosing logic strategy for instantaneous overcurrent protection. The system predicts the spread trend of the fault current in the spatial topology through simulation, determining which adjacent branches may be affected.

[0045] The system calculates the voltage sag depth and duration at nodes surrounding the fault point to assess the impact on sensitive power users. Through multi-condition simulations, it extrapolates the probability distribution of the current fault evolving from a transient ground fault into a permanent one. For example, if the physical cause indicates a carbonization path in the insulator, the model predicts the time chain for further expansion of the discharge channel under sustained voltage stress. Based on these extrapolation results, the system generates targeted operational isolation recommendations, such as remotely disconnecting specific sectionalizing switches or deploying energy storage systems to support the voltage levels of critical loads.

[0046] To ensure the accuracy of the system during long-term operation, this embodiment also includes an update mechanism for the mature detector set. The topology of the distribution network may slowly evolve due to line reconstruction or load relocation. The system periodically acquires newly added normal operation data of the distribution network and treats it as an increment. These increments are then subjected to a secondary tolerance check against the existing mature detector set. If a mature detector is found to overlap with the new normal data, i.e., a collision has occurred, the system will automatically remove the detector or reduce its radius. This ensures that the fault identification logic can adapt to the dynamic changes in the physical structure of the distribution network, avoiding the problem of traditional fixed criteria becoming ineffective due to environmental changes.

[0047] Example 2: Based on Example 1, this example focuses on active distribution network scenarios with a high proportion of distributed power sources. Specific technical optimizations were made to the generation of the self-sample set and the evolution logic of the detector. Because distributed power sources such as distributed photovoltaics and small-scale wind power have significant intermittent and random power output, traditional steady-state self-regions are insufficient to fully cover their fluctuation range.

[0048] In the digital twin construction phase of step 1, this embodiment further refines the dynamic characteristics of the power supply side. The digital twin model not only includes static power injection but also integrates the dynamic response model of the inverter controller. When generating its own sample set, Monte Carlo sampling logic is introduced to perform power flow simulations under a vast array of meteorological scenarios. For example, it simulates the sudden drop in photovoltaic output caused by cloud cover and the active power fluctuations caused by wind speed fluctuations.

[0049] To handle such highly complex normal fluctuations, in step 2, the self-sample set exhibits a non-convex multi-center distribution. Correspondingly, when generating immature detectors, the radius of the initially generated hypersphere is no longer a globally uniform constant, but is adaptively adjusted based on the density of self-points at its location in the feature space. In regions with dense self-point distribution, i.e., under normal load center conditions, the initial radius of the detector is set smaller to achieve high-precision state division; while in edge regions with sparse self-point distribution, i.e., under extreme output conditions, the radius of the detector is appropriately increased.

[0050] In the tolerance comparison process, this embodiment introduces a local density weighted algorithm. When determining the distance between the center vector of the immature detector and its own sample set, it does not only rely on the shortest distance, but also calculates the statistical distribution characteristics of the own samples within a certain neighborhood around the center vector. If the standard deviation of the distribution of the own samples at the detector's location exceeds a preset threshold, it indicates that the region belongs to the high-fluctuation normal zone, and the system will automatically increase the tolerance threshold value to reduce the risk of false identification caused by power fluctuations.

[0051] Furthermore, the epitope correlation analysis in this embodiment adds a deeper analysis of harmonic domain characteristics. For specific frequency components that may be generated by inverter faults, an electromagnetic transient simulation plugin is introduced into the digital twin model. Upon detecting an anomaly, the system adjusts impedance and load parameters, and also adjusts the inverter's control loop parameters, such as proportional-integral gain parameters or phase-locked loop time constants, to determine whether the fault is caused by a primary physical electrical circuit or a virtual fault caused by the instability of the power electronic equipment control. This differentiated diagnostic logic enables the present invention to accurately distinguish between physical line faults and control system anomalies, improving the operation and maintenance efficiency of active distribution networks.

[0052] Example 3: This example demonstrates the specific configuration and optimization logic of the present invention at the hardware implementation level, especially the engineering implementation scheme for meeting the needs of large-scale real-time data processing.

[0053] A data analysis-based power distribution network fault identification system that implements the above method adopts a layered distributed hardware architecture. The core of the system includes: a digital twin simulation module, an immune learning processing module, a real-time acquisition and preprocessing module, a state matching monitoring module, and a fault inference and correlation analysis module.

[0054] The digital twin simulation module incorporates a high-performance matrix operation engine, composed of multiple parallel floating-point units. This design supports the reconstruction of power flow sections for thousands of nodes across the entire network within seconds. To ensure synchronization between the simulation and the physical entity, the module is equipped with a high-speed data interface that directly connects to the power grid dispatching system's operational command database. This allows for real-time acquisition of switch opening and closing status, protection setting adjustment commands, and transformer tap change information, ensuring that the digital twin is always a shadow mapping of the physical entity.

[0055] The immune learning processing module is connected to the digital twin simulation module via an internal high-speed bus. Its core task is to execute large-scale rejection selection algorithms. To handle a massive number of mature detectors, the module employs a multi-threaded parallel computing architecture. Each computation thread is responsible for maintaining a specific subdomain of the feature space and performing clonal selection and radius evolution operations within it. This divide-and-conquer strategy reduces the computational complexity of the detector generation process.

[0056] Real-time data acquisition and preprocessing modules are deployed at key nodes in the power distribution network. Each acquisition unit is equipped with a high-precision analog-to-digital converter, with a sampling frequency of 256 points or higher per cycle. The preprocessing unit uses a field-programmable logic array (FPGA) to perform real-time Fourier transforms and vector normalization. The processed feature vectors are then transmitted to the central server's status matching and monitoring module via an encrypted dedicated power wireless network.

[0057] The state matching monitoring module employs a specially designed hardware acceleration array. This array can perform parallel searches of received real-time feature vectors with tens of thousands of mature detector parameters stored in memory. The decision-making logic is embedded in the hardware instruction set, which uses hardware multipliers and adders to quickly calculate the distance between vectors and compare it with a radius threshold stored in a lookup table. This implementation ensures that the decision-making latency for high-frequency real-time data is controlled within milliseconds, meeting the real-time requirements of power systems for fault identification.

[0058] The fault simulation and correlation analysis module is equipped with an expert knowledge base matching unit. This unit stores a database of typical fault cases accumulated over decades in the distribution network, covering various modes such as single-phase grounding, phase-to-phase short circuits, ferroresonance, and voltage transformer fuse blowout. After the digital twin model completes the physical cause tracing, the fault simulation and correlation analysis module will perform pattern matching between the inverted physical parameter combination and the case database, outputting a structured fault diagnosis report, which includes the fault type, estimated fault location coordinates, expected evolution consequences, and specific operation and maintenance operation suggestions.

[0059] Throughout the system's operation, a closed-loop feedback chain is formed between the modules. For example, when the state matching monitoring module identifies a novel, previously unseen anomaly pattern, this pattern is stored as an antigen sample. After the fault deduction and correlation analysis module provides an accurate diagnosis, the novel anomaly is assigned a specific label and fed back to the immune learning processing module to optimize the distribution strategy of subsequent detectors. This self-learning capability allows the system to continuously evolve, with its recognition accuracy and diagnostic depth continuously improving over time.

[0060] In the specific execution details of step 2, regarding the number of detectors generated, this embodiment employs a dynamic adjustment mechanism based on information entropy. The system monitors the coverage density of its own sample set in the feature space in real time. In areas with disordered feature point distribution and high information entropy, the system automatically increases the generation probability density of the initial detectors to ensure accurate characterization of complex boundary zones. This non-uniform generation strategy, compared to uniform generation across the entire space, saves computational resources and reduces the probability of detecting holes.

[0061] In step 4, the antigen epitope mapping, this embodiment details the calculation logic of the sensitivity matrix. The system applies minute perturbations to each key physical parameter in the digital twin model, such as the ground conductance of the k-th line segment. By comparing the difference in the output feature vector before and after the perturbation, the sensitivity coefficient of that parameter to each feature dimension is obtained. All sensitivity coefficients are arranged in matrix form. During source tracing and identification, the inverse or pseudo-inverse operation of this matrix is ​​used to quickly convert the observed feature deviation direction into the adjustment direction of the physical parameter. This sensitivity-guided search algorithm significantly shortens the parameter inversion convergence time of the digital twin model, typically requiring only three to five iterations to pinpoint the true physical cause of the fault in a broad parameter space.

[0062] For the fault evolution simulation, the system employs a probabilistic and statistical-based physical simulation. The digital twin engine simulates not only single deterministic processes but also the influence of random environmental factors. For example, when simulating the development of a flashover fault, the model retrieves real-time humidity data from the local weather station to analyze the impact of humidity changes on the recovery of air insulation strength. If it predicts that humidity will continue to rise within the next hour, the system automatically increases the risk weight of the fault evolving from instantaneous to permanent and pushes a high-priority warning to the mobile terminals of maintenance personnel in advance.

[0063] The system's adaptive update mechanism also reflects its intelligent characteristics. Whenever a large-scale topology reconfiguration occurs in the distribution network, such as adding a feeder or putting a new substation into operation, the system automatically enters a 24-hour self-learning period. During this period, the digital twin simulation module generates new self-samples at a high frequency, while the immune learning processing module executes incremental learning logic in the background, smoothly migrating and revising the existing detector set. This process is fully automated, requiring no manual intervention, ensuring the consistency between the fault identification logic and the physical power grid.

[0064] In terms of software architecture, this system adopts a microservice architecture. Each functional module can be upgraded and maintained independently. For example, when a more advanced signal processing algorithm emerges, only the logical image of the real-time acquisition and preprocessing module needs to be updated, without adjusting the core database of the entire system. The microservices communicate asynchronously through message queues, buffering the data surges generated when large-scale disturbances occur in the power grid and ensuring the stability of the system under high-pressure loads.

[0065] The system described in this embodiment has demonstrated extremely high practical value in actual deployment. By transforming the originally abstract electrical feature identification into a vivid physical evolution simulation, maintenance personnel can intuitively view a three-dimensional simulation diagram of the fault on the line topology in the command center. The diagram uses color depth to indicate the voltage risk level of the affected nodes and dynamic arrows to show the possible flow direction of the fault current. This visualization method, combined with in-depth logical analysis, enables the fault handling of the distribution network to shift from blind experience-based judgment to scientific and precise policy implementation.

[0066] Furthermore, the system is also equipped with a self-healing strategy evaluation function. During fault simulation, if it is determined that the fault point has been successfully isolated, the system will automatically try various load transfer schemes in the digital twin model. By calculating the voltage drop after the transfer and whether the line power flow exceeds the limit, the optimal power restoration scheme for the non-faulty area is selected. After these schemes have undergone safety verification, they can be directly sent to the distribution automation master station for execution, reducing the impact of the power outage.

[0067] In summary, this invention introduces the non-self-recognition intelligence of biological immunology into fault identification in power systems through an innovative fusion mechanism. In the absence of massive fault samples, high-quality self-samples generated by digital twins construct an elimination-based defense barrier for the system. This architecture, which does not rely on learning specific fault modes, gives the system a natural advantage in identifying unknown risks. Simultaneously, through the inversion and deduction of virtual and real-world data, it bridges the gap between data science and physical mechanisms, achieving in-depth and visualized fault diagnosis.

[0068] The method and system proposed in this invention have technical advantages in addressing single-phase grounding fault identification, early warning of hidden defects, and complex operating condition assessment after large-scale distributed power source integration in distribution networks. Its adaptive learning feature ensures that the system can evolve with the development of the power grid, reducing long-term maintenance costs. This has significant application value for building highly reliable and intelligent modern distribution networks and improving the safe operation and security of the power grid.

[0069] In practical implementation, those skilled in the art can flexibly adjust the detector's dimensions and sampling period according to the voltage level, coverage area, and specific communication link bandwidth of the power distribution network. For example, for microgrids in urban core areas, denser sampling and higher-dimensional feature spaces can be used to obtain the ultimate recognition accuracy; while for overhead line networks in remote areas, the feature dimensions can be appropriately simplified, and edge computing nodes can be used to complete the initial recognition and judgment to adapt to limited communication conditions. These adaptive adjustments in practical applications all fall within the protection scope of this invention.

[0070] To further ensure the completeness of the detector set, this embodiment also implements hole-filling logic. In the generated mature detector set, the system automatically searches for feature space regions that are not covered by any detector hypersphere or occupied by self-sample points. These regions are called detection holes, representing blind spots in system monitoring. For these regions, the system forcibly generates detectors with specific parameters to fill them, ensuring that the entire non-self-sample space is under the system's close monitoring. This proactive filling of detection holes enhances the system's ability to cope with extremely rare and severe failure conditions.

[0071] In the real-time acquisition and preprocessing module, to address the non-stationary characteristics of the sampled signal, this embodiment also employs wavelet decomposition as a supplement to Fourier transform. By decomposing the signal at different time and frequency scales, transient abrupt changes occurring within extremely short periods can be captured, such as inrush currents caused by lightning strikes or instantaneous overvoltages caused by switching operations. These microscopic features, after being extracted, are combined as additional components into the feature vector, improving the system's ability to distinguish between transient and permanent faults.

[0072] In the fault simulation and correlation analysis module, the accelerated simulation mode also considers the dynamic response characteristics of the load. For example, for motor-type loads, the simulation model includes the decay process of its feedback current. When simulating fault propagation, this dynamic simulation can more accurately determine whether the relay protection device will malfunction due to transient current. This in-depth simulation based on physical details provides maintenance personnel with valuable suggestions for optimizing protection settings.

[0073] Furthermore, this system supports multi-source data fusion analysis. In addition to electrical quantity information, it can also access data from line ambient temperature sensors, cable joint temperature measuring devices, and partial discharge monitoring data. These non-electrical quantity data are transformed into auxiliary feature dimensions, which, together with voltage and current features, constitute a multi-dimensional sensing vector. During epitope mapping, if a deviation in the electrical signal is found to be highly correlated with temperature anomalies or a surge in partial discharge levels, the system will automatically increase the weight of insulation aging in the physical causes, outputting a more comprehensive and accurate diagnostic conclusion.

[0074] Through the aforementioned series of detailed technical implementation methods, this invention constructs a complete closed-loop process from physical parameter input, virtual sample generation, immune defense establishment, real-time matching and identification to physical tracing and deduction. Each step is implemented through pure technical logic, without relying on subjective experience, and completely excludes the interference of mathematical formulas in the logical description, ensuring the openness, transparency, and feasibility of the technical solution.

[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A data analysis based power distribution network fault identification method, characterized in that, Includes the following steps: Step 1. Construct a digital twin model of the distribution network. The digital twin model is established by integrating the physical topology parameters, line impedance parameters, and transformer ratio parameters of the distribution network. Based on the digital twin model, continuous power flow calculation and simulation under multiple operating conditions are performed under preset load fluctuation step size and distributed power generation output randomness constraints. The voltage vector, branch current vector, active power, and reactive power of each node of the distribution network are extracted as electrical features. After nonlinear manifold learning dimensionality reduction processing, a self-sample set representing the normal operating state of the distribution network is generated. Step 2. Based on the self-sample set, a set of mature detectors is generated by training an adaptive immune learning algorithm. The adaptive immune learning algorithm randomly generates hyperspherical structure detectors defined by a center vector and a coverage radius in the feature space of the distribution network operation state, and performs self-tolerance screening and affinity evolution logic to exclude detectors that have spatial overlap with the self-sample set, so that the set of mature detectors can define the non-self region in the distribution network operation state space. Step 3. Acquire real-time field operation data of the power distribution network through the power distribution automation terminal, perform sliding window sampling on the field operation data, and apply Fourier transform to extract each harmonic component and fundamental wave characteristics. After normalization processing, form a real-time data vector and input it into the mature detector set. Perform real-time matching and anomaly identification by determining the spatial position relationship between the real-time data vector and the hypersphere structure detector in the feature space. Step 4. For the identified non-self-abnormal signals, extract their key deviation direction in the feature space as antigen epitope feature vectors, and map them back to the physical parameter layer of the digital twin model. Based on the evolutionary deduction mechanism of virtual-real fusion, the line insulation level, grounding resistance and branch load fluctuation parameters are iteratively adjusted in the digital twin model to reproduce the antigen epitope feature vectors, thereby determining the physical cause of the fault and the subsequent development path of the fault current in the spatial topology.

2. The data analysis-based distribution network fault identification method according to claim 1, characterized in that, The specific process of training and generating a mature detector set using the adaptive immune learning algorithm in step 2 includes: Within the characteristic space of the distribution network operation status, a controlled number of immature detectors are randomly generated, and each immature detector is represented by a hypersphere composed of a center vector and a preset coverage radius. The immature detector is subjected to tolerance comparison with the self-sample set to obtain the component values ​​of the center vector of the immature detector in each dimension, and the difference is calculated with the component values ​​of the self-sample points in the same dimension. The quadratic operation is performed on all the obtained differences, the quadratic operation results of all dimensions are accumulated and summed, and the arithmetic square root is taken on the accumulated sum to obtain the scalar distance parameter reflecting the geometric displacement between the two. When the scalar distance parameter is less than the preset tolerance threshold, the immature detector is determined to be self-activated and eliminated; for immature detectors that are not eliminated within the preset tolerance period, they are transformed into mature detectors and added to the set of mature detectors.

3. The data analysis-based distribution network fault identification method according to claim 2, characterized in that, The adaptive immune learning algorithm also includes detector affinity evolution logic: The overlap between mature detectors in the set of mature detectors is periodically calculated. The overlap is determined by calculating the distance between the center vectors of two mature detectors and comparing it with the sum of the radii of the two mature detectors. For mature detectors with an overlap exceeding a predetermined ratio, a cloning selection mechanism is used for local search and radius compression to ensure that the coverage of the mature detector set in the non-self space reaches a gradient, while reducing detection holes.

4. The data analysis-based distribution network fault identification method according to claim 1, characterized in that, The process of constructing a digital twin model and generating a self-sample set in step 1 specifically includes: acquiring the physical topology parameters, line impedance parameters, transformer ratio information and load characteristic distribution of the distribution network, and constructing a quasi-steady-state simulation engine that is highly mapped to the physical entity in the digital space. Based on the historical operating range of the distribution network, the typical logic of load fluctuation step size, distributed power output randomness and switching operation is set, and continuous power flow calculation is performed under fault-free constraint conditions. The voltage vector, branch current vector, active power, and reactive power of each node are extracted as original features, and a self-sample set distributed on the manifold is generated through dimensionality reduction.

5. The data analysis-based distribution network fault identification method according to claim 1, characterized in that, The process of acquiring and preprocessing the field operation data of the distribution network in real time in step 3 specifically includes: Current and voltage signals are acquired in real time through distribution automation terminals and synchronous phasor measurement units; the original signals are sampled by sliding window, and Fourier transform is applied to extract each harmonic component and fundamental wave characteristics; Normalization is used to map electrical parameters of different dimensions to a selected range of values ​​to eliminate system bias caused by differences in equipment sampling; the preprocessed real-time data vector is then used to determine the spatial position of each hypersphere in the set of mature detectors.

6. The data analysis-based distribution network fault identification method according to claim 1, characterized in that, The specific process of performing antigen epitope association analysis in step 4 includes: The key deviation directions of the identified non-self abnormal signals in the feature space are extracted and defined as antigen epitope feature vectors. The antigen epitope feature vector is mapped to the physical parameter layer of the digital twin model using a sensitivity matrix; In the digital twin model, the physical parameters of line insulation level, grounding resistance, and abnormal fluctuations in branch load are iteratively adjusted, and the power flow simulation is rerun to output the simulation feature vector; Calculate the cosine similarity between the simulation feature vector and the antigen epitope feature vector. When the cosine similarity reaches a preset consistency threshold, determine the combination of physical parameters in the digital twin model at this time as the physical cause of the current failure.

7. The data analysis-based distribution network fault identification method according to claim 6, characterized in that, The specific process of the evolutionary deduction mechanism based on virtual-real fusion in step 4 includes: starting the accelerated simulation mode in the digital twin model with the determined physical cause of the fault as the initial disturbance; predicting the spread trend of the fault current in the spatial topology based on the preset relay protection action logic and the automatic reclosing strategy of the distribution network; calculating the voltage sag depth and duration of the nodes around the fault point, deducing the probability distribution of the fault evolving from transient to permanent, and generating operation and maintenance isolation suggestions.

8. The data analysis-based distribution network fault identification method according to claim 1, characterized in that, It also includes an update mechanism for the mature detector set: periodically acquire new normal operation data of the distribution network and use it as an increment to perform a secondary tolerance check on the existing mature detector set; if the mature detector collides with the increment, the mature detector is removed or adjusted.

9. A data analysis-based distribution network fault identification system implementing the method of any one of claims 1 to 8, characterized in that, include: The digital twin simulation module is used to receive physical parameters of the distribution network and perform high-precision power flow calculations to generate its own sample set; An immune learning processing module, connected to the digital twin simulation module, is configured to execute an adaptive immune learning algorithm to generate and maintain a set of mature detectors through autotolerance screening and clonal selection evolution. The real-time acquisition and preprocessing module is used to acquire electrical operating parameters from the power distribution network site and perform sliding window sampling, Fourier transform and normalization processing to convert them into standard feature vectors. The state matching and monitoring module has its input end connected to the real-time acquisition and preprocessing module and its reference end connected to the immune learning and processing module. It is used to determine whether the current running state falls into a non-self region. The fault simulation and correlation analysis module is connected to the state matching monitoring module and the digital twin simulation module, and is used to perform antigen epitope mapping and physical evolution simulation after an anomaly is identified.

10. A data analysis-based power distribution network fault identification system according to claim 9, characterized in that, The fault inference and correlation analysis module is equipped with an expert knowledge base matching unit, which is used to compare the parameter inversion results of the digital twin model with the stored typical fault modes and output a structured fault diagnosis report. The state matching and monitoring module adopts a multi-threaded parallel computing architecture, which supports the distribution of mature detectors on multiple computing units to realize parallel retrieval and distance determination of real-time data.