Lightning stroke fault identification method and system for power distribution network containing distributed power supply
By simulating lightning faults in the distribution network, extracting multiple fault features and identifying them using a decision tree model, the problems of insufficient lightning identification accuracy and real-time performance in the existing technology are solved, and high-precision lightning fault identification is achieved.
Patent Information
- Application Number
- CN202510894864.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies for lightning strike identification suffer from complex environmental interference, insufficient real-time performance, and low classification accuracy, making it difficult to effectively monitor and identify lightning strike faults.
Using methods based on machine learning and deep learning, a distribution network model containing distributed power sources is built to simulate lightning faults, extract various fault features in the time domain, frequency domain, and time-frequency domain, and use a decision tree model for fault identification.
It improves the accuracy and real-time performance of lightning fault identification, adapts to the distribution network fault identification needs after a high proportion of distributed power sources are connected, and has important practical significance.
Smart Images

Figure CN120610109A_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of power grid signal processing, and in particular to a method and system for identifying lightning faults in a distribution network containing distributed power sources. Background Art
[0002] Lightning strikes are a common natural phenomenon, accompanied by intense electromagnetic radiation, sonic shocks, and instantaneous high-energy release. Their unpredictability and destructive nature pose a serious threat to human society. Natural disasters and industrial accidents caused by lightning strikes are commonplace each year. Lightning strikes not only endanger lives but can also cause secondary hazards such as fires, explosions, equipment damage, and communication disruptions. They pose a significant threat to key sectors such as power systems, aerospace, and communications networks. In power systems, in particular, transient overvoltages caused by lightning strikes can damage the insulation of transmission lines and substation equipment, and even cause widespread power outages. Furthermore, as a key characteristic of weather systems such as thunderstorms and severe convection, monitoring and analyzing lightning strikes can help improve the accuracy of short-term weather forecasts and provide a scientific basis for disaster prevention and mitigation.
[0003] Traditional lightning identification technology is centered around electromagnetic signal detection and optical observation, and mainly includes the following methods: (1) Ground-based lightning monitoring networks, such as the National Lightning Detection Network (NLDN) of the United States, which detects low-frequency and very high-frequency (VLF, VHF) electromagnetic radiation signals released by lightning based on the principle of electromagnetic wave propagation, and calculates the location and characteristics of lightning strikes through multi-station positioning methods; (2) Lightning locators, which use the time difference or direction of electromagnetic waves to perform single-point or multi-point positioning, and achieve real-time monitoring of lightning strikes within a region; (3) Optical detection and satellite observation, which analyzes the flash characteristics generated by lightning strikes. Representative equipment includes the Geostationary Lightning Imager (GLM), which can provide optical characteristics and spatiotemporal distribution information of lightning activities. These traditional methods have achieved important results in lightning detection and location, and have been widely used in meteorological monitoring and power system protection. However, these technologies also have limitations. For one thing, electromagnetic wave signals are susceptible to interference from complex environments, which can reduce detection accuracy. Furthermore, traditional technologies are inadequate for lightning strike classification (e.g., positive and negative polarity lightning strikes, cloud-to-ground lightning strikes, and cloud-to-cloud lightning strikes) and for multi-source information fusion, making it difficult to fully characterize lightning strike characteristics. Furthermore, due to the complexity of lightning signal processing algorithms, the real-time performance of the systems often fails to meet practical application requirements.
[0004] In recent years, with the development of artificial intelligence and big data technologies, lightning strike identification technology has gradually moved towards intelligentization. Machine learning and deep learning-based methods have begun to be applied to lightning strike identification tasks. By extracting features and recognizing patterns from massive amounts of lightning strike data, these methods have demonstrated significant advantages in improving identification accuracy, optimizing real-time performance, and classifying lightning strike types.
[0005] In summary, traditional lightning strike identification technology lays the foundation for lightning monitoring, but it still has shortcomings in coping with complex environmental interference, real-time performance, and classification accuracy. Intelligent methods, by introducing artificial intelligence and big data technologies, provide new solutions for lightning strike identification. To address the challenges currently faced in research, such as the need for multi-source data fusion, the balance between real-time performance and high accuracy, and model adaptability, future research will focus on developing more efficient, accurate, and intelligent lightning strike identification technologies to provide stronger support for lightning strike monitoring and protection. Summary of the Invention
[0006] In order to solve the technical problems existing in the prior art, the present invention provides a method and system for identifying lightning faults in a distribution network containing distributed power sources with high identification accuracy.
[0007] In order to solve the above technical problems, the technical solution proposed by the present invention is: A method for identifying lightning faults in a distribution network containing distributed power sources comprises the following steps: S1. Build a distribution network model with a high proportion of distributed generation (DGs). Simulate short-circuit and lightning faults under different fault types, fault locations, and transition resistance conditions, generating a raw data set containing lightning current and short-circuit current signals. For lightning faults, a standard lightning current waveform is simulated using a double exponential function and injected into the line end or tower grounding device, while also accounting for line coupling effects. S2. Perform multi-dimensional feature analysis on the original data set generated in step S1, extract time domain features, frequency domain features, and time-frequency domain features to form a fusion feature set, and label different fault types and locations; S3. Divide the fused feature set obtained in step S2 into a training set and a test set in proportion, construct a decision tree model, select the optimal splitting feature by information gain or Gini index, and iterate the training until the preset accuracy or convergence condition is met; the pruning strategy of the decision tree adopts pre-pruning or post-pruning to avoid overfitting, and the classification accuracy of the model on the test set is evaluated by the confusion matrix.
[0008] Preferably, the time domain features include peak-to-peak value, standard deviation and kurtosis; the frequency domain features include key frequency; and the time-frequency domain features include the average frequency of the intrinsic mode function component.
[0009] Preferably, the time-frequency domain features are extracted using a variational mode decomposition method. Specifically, the variational mode decomposition method decomposes the time domain signal into n IMF components, each IMF component corresponds to a specific frequency bandwidth, and the average frequency of each IMF component is taken to measure the index of the main contribution frequency of the IMF component in the frequency domain, reflecting the distribution of different frequency components in the signal and its evolution in the time domain.
[0010] Preferably, the specific steps of the variational mode decomposition method include: 1) Sampling time domain signals; 2) Determine the parameters of the variational mode decomposition method; 3) The sampled time domain signal is used as input and the variational mode decomposition method is applied to decompose it. Through an iterative process, the time domain signal is decomposed into multiple modal components, each of which represents the vibration mode of the signal at different frequencies and times. 4) Set the number of decomposed modal functions to n, then calculate the instantaneous frequency of each IMF component through Hilbert transform, and take the average frequency of each IMF component as the fault time-frequency feature.
[0011] Preferably, in step 2), the parameters decomposed by the variational mode decomposition method include a signal sampling rate, a number of decomposition layers, and a regularization parameter.
[0012] Preferably, the specific process of step S1 is: A distribution network model was built in Matlab / Simulink. First, a radial network was constructed using the three-phase power supply, transformer, line, and load modules in the SimPowerSystems library. A loop code script was written to control the three-phase fault detector to obtain batch current data by changing the fault type, fault location, and transition resistance value. At the same time, a lightning strike module was built. For lightning impulses, a customized lightning current waveform was defined, using a double exponential function to simulate the standard lightning current. The current was injected into the line end or tower grounding device through a current source, and the line coupling effect was considered. Batch lightning current data was obtained through a loop code, and a total of multiple fault currents and lightning fault currents were obtained.
[0013] Preferably, in step S3, pre-pruning is performed: during the process of building the tree, it is decided whether to stop splitting; post-pruning is performed: a complete tree is first built, and then some unimportant branches are pruned by evaluating the performance of the tree.
[0014] The present invention also discloses a computer program product, comprising a computer program, which executes the steps of the above method when executed by a processor.
[0015] The present invention further discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described above are executed.
[0016] The present invention also discloses a lightning fault identification system for a distribution network containing a distributed power supply, comprising a memory and a processor connected to each other, wherein a computer program is stored on the memory, and when the computer program is run by the processor, the steps of the above method are executed.
[0017] Compared with the prior art, the advantages of the present invention are: The present invention's lightning fault identification method for distribution networks containing distributed energy sources extracts multiple fault characteristic indicators in the time, frequency, and time-frequency domains when short-circuit faults and lightning faults occur in distribution networks with a high proportion of distributed power sources. It then uses this extensive multidimensional fault characteristic data to train a decision-tree-based fault identification model, leveraging artificial intelligence to improve fault identification accuracy. This invention adapts to the development trend of new energy and has important practical implications for fault identification in future distribution networks with a high proportion of distributed power sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a topological diagram of a lightning fault in the distribution network of the present invention.
[0019] Figure 2 This is the distribution network access fault diagram in the present invention.
[0020] Figure 3 This is a flow chart of an embodiment of the method for identifying lightning faults in a distribution network according to the present invention.
[0021] Figure 4 Generate a graph for the decision tree child nodes in the present invention.
[0022] Figure 5 This is the confusion matrix result diagram in the present invention. DETAILED DESCRIPTION
[0023] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0024] like Figure 3 As shown, the method for identifying lightning faults in a distribution network including distributed power sources provided by an embodiment of the present invention includes the following steps: Step S1: Build a distribution network model with a high proportion of distributed power sources, simulate short-circuit faults and lightning faults under different fault types, fault locations, and transition resistance conditions, and generate an original data set containing lightning current signals and short-circuit current signals; like Figure 1-Figure 2As shown, a distribution network model was built in Matlab / Simulink. First, a radial network was constructed using the three-phase power supply, transformer, line, and load modules from the SimPowerSystems library. Line parameters, such as resistance, inductance, and capacitance, were set according to actual cable specifications. A looping code script was written to control the three-phase fault detector to acquire batch current data by varying the fault type, location, and transition resistance. A lightning strike module was also built. For lightning impulses, a custom lightning current waveform was created, using a double exponential function to simulate a standard lightning current. This current was injected into the line end or tower grounding device via a current source, taking into account line coupling effects. The looping code was used to acquire batch lightning current data, totaling 3,000 fault currents and lightning fault currents. The current data was then divided into training and test sets in a 7:3 ratio.
[0025] Step S2: Perform multi-dimensional feature analysis on the original data set generated in step S1, extract time domain features, frequency domain features, and time-frequency domain features, form a fusion feature set, and label different fault types and locations; specifically: (1) Time domain characteristics Based on the above steps, the time domain waveform is used to calculate the three time domain characteristic indicators: peak-to-peak value, standard deviation, and kurtosis.
[0026] 1. Peak-to-peak value: Peak-to-peak value refers to the difference between the maximum and minimum values of a signal within a complete cycle. This metric is often used to describe the dynamic range of a signal and its amplitude. For non-periodic signals, peak-to-peak value refers to the difference between the maximum and minimum values within the entire signal range.
[0027] 2. Standard Deviation: Standard deviation is a statistical indicator used to measure the dispersion or volatility of values in a set of data. When the standard deviation is larger, it means that the volatility and dispersion of the data are higher; when the standard deviation is smaller, it means that the volatility of the data is lower. The formula for calculating standard deviation is as follows:
[0028] in, N is the number of observations in the dataset, It is observations, is the average value.
[0029] 3. Kurtosis: Kurtosis is used to measure the sharpness or flatness of the data distribution. The calculation formula of kurtosis can be implemented using various statistical software or programming languages.
[0030] (2) Frequency domain characteristics Fast Fourier Transform (FFT) can be used to convert fault characteristics in the time domain to the frequency domain for analysis. It is an effective method for converting time domain signals into frequency domain signals. The following are the basic steps for converting time domain signals into frequency domain signals using Fast Fourier Transform (FFT): 1. Sampling time domain signals; 2. Determine the sampling rate, which is the number of samples collected per second. The sampling rate determines the frequency resolution of the signal; 3. Choose an appropriate FFT length. The FFT length is usually a power of 2. 4. Take the sampled time domain signal as input and apply the FFT algorithm to convert the time domain signal into a frequency domain signal, where the frequency domain signal represents the amplitude and phase of the signal at different frequencies; 5. Obtain the frequency domain results, perform spectrum analysis, and calculate the center of gravity frequency as a frequency domain characteristic indicator. The center of gravity frequency is used to describe the center position of the spectrum distribution. It is the weighted average frequency of each frequency component in the spectrum, reflecting the contribution and distribution of each frequency component in the spectrum. The calculation formula is as follows:
[0031] in, It is k The frequency value of the frequency component, It is k The amplitude (energy or power) of the frequency components, N is the number of frequency components.
[0032] (3) Time-frequency domain characteristics The variational mode decomposition (VMD) method is used to convert the time-domain fault characteristics to the time-frequency domain. VMD is a signal processing method used to convert time-domain signals into the time-frequency domain. Unlike traditional Fourier transform methods, VMD is an adaptive decomposition technique that can adapt to the characteristics of nonlinear and nonstationary signals. The following are the basic steps for converting time-domain signals into the time-frequency domain using VMD: 1. Sampling time domain signals; 2. Determine the parameters of VMD decomposition, including signal sampling rate, number of decomposition layers, regularization parameters, etc.; 3. The sampled time domain signal is used as input and decomposed using the VMD algorithm. The VMD algorithm decomposes the time domain signal into multiple modal components through an iterative process. Each modal component represents the vibration mode of the signal at different frequencies and times. 4. Set the number of decomposed modal functions to 5, then calculate the instantaneous frequency of each IMF component through Hilbert transform, and take the average frequency of each IMF component as the fault time-frequency feature.
[0033] The final multi-dimensional fault characteristics are shown in Table 1: Table 1 Description of multi-dimensional fault characteristics
[0034] Step S3: Divide the fusion feature set obtained in step S2 into a training set and a test set in proportion, build a decision tree model, select the optimal split feature by information gain or Gini index, and iterate the training until the preset accuracy or convergence condition is met. Among them, the pruning strategy of the decision tree adopts pre-pruning or post-pruning to avoid overfitting, and evaluates the classification accuracy of the model on the test set through the confusion matrix, such as Figure 5 shown.
[0035] like Figure 4 As shown, the basic principle of decision tree 1. The structure of the decision tree A decision tree consists of a set of nodes and edges. The root node represents the entire dataset. Internal nodes represent the conditions for evaluating a particular feature. Leaf nodes represent the final prediction, typically a class label (for classification problems) or a numerical value (for regression problems). Edges represent the results of the conditional evaluation (e.g., the range of values for a feature).
[0036] 2. Tree construction process The decision tree construction process can be divided into the following steps: (2.1) Dataset division The dataset is divided into two or more subsets based on the value of a certain feature, and each subset is divided again until the stopping condition is met (such as the dataset is pure, the tree depth reaches a predetermined value, etc.). Each time the partition is made, the optimal feature is selected for the partition.
[0037] (2.2) How to choose the optimal features (splitting criteria) One of the most important aspects of decision trees is how to select the best features for data partitioning. Common selection criteria include information gain, Gini index, and mean squared error (MSE).
[0038] 3. Information Gain Information gain is a commonly used splitting criterion in decision trees, especially in algorithms such as ID3 and C4.5. Information gain measures how much the uncertainty of information is reduced after selecting a certain feature to split the data.
[0039] 3.1 Entropy Entropy is a measure of information uncertainty in information theory. For classification problems, suppose the categories of the data set are C 1 ,C 2 ,…,C k Its entropy is defined as:
[0040] in, Indicates that the data set belongs to the category The larger the entropy value, the higher the uncertainty of the data set.
[0041] 3.2 Information Gain In selecting features A To divide the data set D When It is defined as the difference in entropy before and after partitioning:
[0042] in, Values(A) It is a feature A All values of It is in the characteristics A The value is v The data subset at time is a subset The entropy of is a subset The number of samples, It is a dataset D The total number of samples.
[0043] The larger the information gain, the better the A After the data set is divided, the more the uncertainty is reduced, so the feature with the largest information gain is selected for division.
[0044] 4. Gini Index The Gini Index is another common decision tree splitting criterion, particularly in the CART (Classification and Regression Trees) algorithm. The Gini Index measures the impurity of a dataset, with smaller values indicating higher purity.
[0045] For classification problems, the Gini index is defined as:
[0046] in, Is the category in the data set The smaller the Gini index, the higher the purity of the data set.
[0047] 4.1 Minimization of the Gini Index In the process of building a decision tree, the feature with the smallest Gini index is selected for data partitioning. For a certain feature AAA, its Gini index calculation formula is:
[0048] in, It is the data subset when feature A takes value v, is a subset The Gini index.
[0049] 5. Splitting criteria for regression trees For regression problems, the goal of a decision tree is to predict a continuous value rather than a class label. A common splitting criterion is the mean squared error. (MSE) For a dataset D , and its mean square error is defined as:
[0050] in, It is a sample i The true value of It is a sample i The predicted value of (usually the mean of the data set). At each split, the feature that reduces the mean squared error the most is selected.
[0051] 6. Stop Condition The decision tree construction process is recursive until some stopping condition is met: All samples belong to the same category or have the same regression value; Reach the preset tree depth; The number of samples in the subset is less than a certain threshold; The information gain or Gini index after partitioning is reduced to less than a certain threshold.
[0052] 7. Pruning Decision trees can suffer from overfitting, especially when the tree is too deep. To avoid overfitting, the tree is usually pruned. There are two ways to prune the tree: Pre-pruning: During the tree construction process, decide whether to stop splitting.
[0053] Post-pruning: First build a complete tree, then prune some unimportant branches by evaluating the performance of the tree.
[0054] (3) Fault identification model training and processing process The processed data features are used for model training and fault identification. The specific steps are as follows: 1. Divide 200 9-dimensional data samples into training set and test set in a ratio of 7:3; 2. Build a decision tree model; 3. Input the training data into the decision tree and calculate the training loss; 4. Repeat step 3 until the accuracy and loss are met or the number of iterations is reached to complete the model training; 5. Input the test set data into the trained model to evaluate the diagnostic effect of the model. The present invention relates to the field of distribution network fault protection. When a short-circuit fault or lightning strike occurs in a distribution network with a high proportion of distributed power sources, the method extracts nine fault characteristic indicators in the time, frequency, and time-frequency domains. The method then uses this extensive nine-dimensional fault characteristic data to train a decision-tree-based fault identification model, leveraging artificial intelligence to improve fault identification accuracy. This method adapts to the development trend of new energy and has important practical significance for future fault identification in distribution networks with a high proportion of distributed power sources.
[0055] The present invention also discloses a computer program product, comprising a computer program, which executes the steps of the above method when executed by a processor.
[0056] The present invention further discloses a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described above are executed.
[0057] The present invention also discloses a lightning fault identification system for a distribution network containing a distributed power supply, comprising a memory and a processor connected to each other, wherein a computer program is stored on the memory, and when the computer program is run by the processor, the steps of the above method are executed.
[0058] The products, media and systems of the present invention correspond to the above-mentioned methods and also have the advantages of the above-mentioned methods.
[0059] The present invention can implement all or part of the process steps in the above-described method embodiments through hardware associated with computer program instructions. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of the above-described method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable storage media include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Memory is used to store computer programs and / or modules. The processor implements various functions by running or executing the computer programs and / or modules stored in the memory and accessing data stored in the memory. The memory may include a high-speed random access memory and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0060] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for identifying lightning faults in a distribution network containing distributed power sources, characterized in that: Including steps: S1. Build a distribution network model with a high proportion of distributed generation (DGs). Simulate short-circuit and lightning faults under different fault types, fault locations, and transition resistance conditions, generating a raw data set containing lightning current and short-circuit current signals. For lightning faults, a standard lightning current waveform is simulated using a double exponential function and injected into the line end or tower grounding device, while also accounting for line coupling effects. S2. Perform multi-dimensional feature analysis on the original data set generated in step S1, extract time domain features, frequency domain features, and time-frequency domain features to form a fusion feature set, and label different fault types and locations; S3. Divide the fused feature set obtained in step S2 into a training set and a test set in proportion, construct a decision tree model, select the optimal splitting feature by information gain or Gini index, and iterate the training until the preset accuracy or convergence condition is met; the pruning strategy of the decision tree adopts pre-pruning or post-pruning to avoid overfitting, and the classification accuracy of the model on the test set is evaluated by the confusion matrix.
2. The method for identifying lightning faults in a distribution network containing distributed power sources according to claim 1, wherein: The time domain features include peak-to-peak value, standard deviation and kurtosis; the frequency domain features include key frequency; and the time-frequency domain features include the average frequency of the intrinsic mode function component.
3. The method for identifying lightning faults in a distribution network containing distributed power sources according to claim 2, wherein: The time-frequency domain features are extracted using a variational mode decomposition method. Specifically, the variational mode decomposition method decomposes the time domain signal into n IMF components, each of which corresponds to a specific frequency bandwidth. The average frequency of each IMF component is taken to measure the index of the main contribution frequency of the IMF component in the frequency domain, reflecting the distribution of different frequency components in the signal and its evolution in the time domain.
4. The method for identifying lightning faults in a distribution network containing distributed power sources according to claim 3, wherein: The specific steps of the variational mode decomposition method include: 1) Sampling time domain signals; 2) Determine the parameters of the variational mode decomposition method; 3) The sampled time domain signal is used as input and the variational mode decomposition method is applied to decompose it. Through an iterative process, the time domain signal is decomposed into multiple modal components, each of which represents the vibration mode of the signal at different frequencies and times. 4) Set the number of decomposed modal functions to n, then calculate the instantaneous frequency of each IMF component through Hilbert transform, and take the average frequency of each IMF component as the fault time-frequency feature.
5. The method for identifying lightning faults in a distribution network containing distributed power sources according to claim 4, characterized in that: In step 2), the parameters decomposed by the variational mode decomposition method include the signal sampling rate, the number of decomposition layers and the regularization parameter.
6. The method for identifying lightning faults in a distribution network containing distributed power sources according to any one of claims 1 to 5, characterized in that: The specific process of step S1 is: A distribution network model was built in Matlab / Simulink. First, a radial network was constructed using the three-phase power supply, transformer, line, and load modules in the SimPowerSystems library. A loop code script was written to control the three-phase fault detector to obtain batch current data by changing the fault type, fault location, and transition resistance value. At the same time, a lightning strike module was built. For lightning impulses, a customized lightning current waveform was defined, using a double exponential function to simulate the standard lightning current. The current was injected into the line end or tower grounding device through a current source, and the line coupling effect was considered. Batch lightning current data was obtained through a loop code, and a total of multiple fault currents and lightning fault currents were obtained.
7. The method for identifying lightning faults in a distribution network containing distributed power sources according to any one of claims 1 to 5, characterized in that: In step S3, pre-pruning: during the tree construction process, decide whether to stop splitting; post-pruning: first build a complete tree, and then prune some unimportant branches by evaluating the performance of the tree.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 7.
10. A lightning fault identification system for a distribution network including a distributed power supply, comprising a memory and a processor connected to each other, wherein a computer program is stored in the memory, characterized in that: When the computer program is executed by a processor, the computer program performs the steps of the method according to any one of claims 1 to 7.