Electric leakage fault positioning and monitoring method and system for distributed photovoltaic access transformer area
By applying random forest algorithms and improved particle swarm algorithms in the distributed photovoltaic access station area, denoising and parameter optimization of leakage fault signals is solved, and the problem of difficult positioning of leakage faults in the low-voltage distribution station area is achieved, online high-precision positioning is improved, and positioning efficiency and safety are improved.
Patent Information
- Application Number
- CN202510333135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-24
AI Technical Summary
In the large-scale distributed photovoltaic access low-voltage distribution station area, leakage faults are difficult to quickly and accurately locate, resulting in long power outages, large safety hazards, and increased electrical fire risk. In addition, traditional methods are inefficient and low sensitivity, making online positioning impossible.
The random forest algorithm is used to combine the improved mixed threshold function and the dual strategy improved particle swarm algorithm to denoising the leakage fault signal and optimize the parameter, build a feature data set and train it to achieve online high-precision positioning of leakage fault points.
It realizes online high-precision positioning of leakage faults in the low-voltage distribution station area, shortens the fault power outage time, reduces safety hazards and electrical fire risks, and improves positioning efficiency and accuracy.
Smart Images

Figure CN120195579A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fault location monitoring, and particularly relates to a method and system for leakage fault location monitoring of a distribution photovoltaic access substation area. Background Art
[0002] With the large-scale distributed photovoltaic access to the substation area, the power flow distribution of the low-voltage distribution substation area becomes complex, bringing higher challenges to the leakage fault location. On the one hand, the lines in the low-voltage substation area are relatively scattered, with a wide coverage area and a harsh environment, and there are many potential leakage hazards. Leakage faults are serious safety hazards, and at the same time, they will cause an increase in the line loss of the substation area. On the other hand, due to the relatively complex leakage fault mechanism, there are many influencing factors, and the leakage signal is weak and the fault characteristics are not obvious. Therefore, leakage faults are relatively hidden faults, and it is difficult to accurately find the leakage fault point, further increasing the difficulty of fault location in the substation area. In the traditional leakage fault detection method, generally, power supply personnel conduct inspections along the line. It is very easy to be negligent during the process of finding the leakage fault, and the fault location cannot be checked in time, resulting in the failure to find the leakage fault point, causing a long power outage time for users, and posing a threat to personal safety. Moreover, the investigation has a large blindness, low detection sensitivity, low efficiency, and causes waste of manpower and material resources. At present, the form of a clamp meter is used in the market to find leakage faults, or power outage detection is required. There is still no method that can realize the online location of leakage faults in the low-voltage distribution substation area.
[0003] The current technology has the following defects, including the following: 1) Traditional fault location methods usually require a large amount of manual inspections, which consume a lot of time, delay the handling of faults and the restoration of power supply. This not only increases the complexity of the work but also reduces the efficiency; 2) The randomness and instantaneousness of leakage faults cause interference from other signals during signal acquisition. If the leakage is directly selected for the signal containing noise, it is very likely to cause misjudgment of leakage using incorrect information. 3) The traditional particle swarm algorithm is prone to falling into a local optimal solution and lacks a mechanism to jump out of this solution. The entire population may gather to this non-global optimal solution, and there are also problems of low convergence speed and accuracy, slow convergence speed, and the inability to find the optimal solution.
[0004] Therefore, the present invention proposes a method and system for leakage fault location monitoring of a distribution photovoltaic access substation area. Summary of the Invention
[0005] To solve the problem of leakage faults in low-voltage distribution substations with large-scale distributed photovoltaic access, the most important thing is to quickly and accurately locate the leakage faults online in sections, shorten the fault power outage time, eliminate the hidden dangers of leakage faults, improve personal safety, avoid the occurrence of electrical fires, and ensure the stability of the low-voltage substation power grid. The present invention provides a method and system for monitoring and locating leakage faults in a substation with distributed photovoltaic access to solve the technical problem of being unable to quickly and accurately locate leakage faults in a low-voltage distribution substation.
[0006] The present invention adopts the following technical solutions.
[0007] The present invention provides a method for monitoring and locating leakage faults in a substation with distributed photovoltaic access, including the following steps: S1: Collect the residual current data of the substation according to the operating conditions of each sample, set the normal and leakage fault labels, record the fault leakage current signal, and then perform preprocessing. After denoising the preprocessed fault leakage current signal through an improved hybrid threshold function, construct a feature dataset and divide it into a training set, a validation set, and a test set; S2: Sample and construct a training sample set from the training set obtained in step S1, and use the classification and regression tree algorithm to divide the sample set based on the Gini index to construct a decision tree. All the constructed decision trees together form a random forest model; S3: Use the improved particle swarm optimization algorithm with an improved dual strategy to iteratively optimize the relevant parameters in the random forest for the random forest performance model constructed in step S2, and use the validation set in step S1 for verification to complete the parameter tuning of the random forest performance model; S4: Input the test set into the random forest model, obtain the monitoring state with the largest proportion of test set results based on the dynamic weight allocation algorithm and the majority voting method based on the Gini index, and identify the corresponding fault state label information. Based on the label information, perform binary processing on the location of the leakage fault using the dichotomy method, output the predicted leakage fault point, and compare it with the actual fault point of the test set to realize the verification of the random forest model; S5: Real-time obtain the new residual current data of the substation collected according to the operating conditions of each sample, import it into the random forest model to obtain the leakage fault point corresponding to the fault leakage current, and complete the monitoring and location of the leakage fault in the substation.
[0008] Preferably, the preprocessing in step S1 further includes: Normalize the obtained fault leakage current signal data, and the calculation formula is as follows: (1) In the formula: is the data sample at a certain moment; is the largest sample data among all samples; is the minimum sample data among all samples.
[0009] Preferably, the denoising process of the preprocessed fault leakage current signal by improving the hybrid threshold function in step S1 further includes: a. For the part where the absolute value of the wavelet coefficient is greater than the threshold, perform a shrinking process from the sign function to the threshold point to retain the effective information; b. Introduce a lower threshold, perform a high-order function process on the part of the wavelet coefficient between the two thresholds, and extract the remaining effective information; c. Set the wavelet coefficients less than the lower threshold to zero; thus, an improved hybrid threshold function is obtained, and the calculation formula is: (2) In the formula: is the wavelet coefficient after being processed by the improved hybrid threshold function; is the wavelet coefficient without being processed by the improved hybrid threshold function; is the adaptive parameter; is the sign function; is the upper threshold; is the lower threshold; is the adjustment parameter; and , .
[0010] Preferably, the construction of the decision tree by splitting the sample set depending on the Gini index in step S2 further includes: The splitting criterion of the decision tree is measured by the Gini impurity. The classification and regression tree decision tree algorithm is used to split the sample set depending on the Gini index. The Gini coefficient measures the degree of improvement in the purity of the sub-datasets after division under a certain value of the feature . The Gini index of the decision tree , and the calculation formula is: (3) In the formula: represents the Gini coefficient of the data training set under a certain feature value of the feature ; is the constant coefficient; and represent the data training sets divided according to the feature The proportion of the number of samples in the sub - datasets D1 and D2 after the value division of Gini(D1) and Gini(D2) respectively represent the Gini indices of the sub - datasets D1 and D2.
[0011] Preferably, the iterative optimization of the relevant parameters in the random forest by using the particle swarm algorithm improved with a dual strategy in step S3 further includes: The random forest performance model introduces the encircling and hunting strategies of the whale optimization algorithm into the particle swarm algorithm, and adds the Levy flight strategy to the particle swarm algorithm and introduces it into the model, and proposes a particle swarm algorithm improved with a dual strategy. The improved update formula is as follows: (4) In the formula: is the d - dimensional position component of the i - th particle monomer; is the d - dimensional velocity component of the i - th particle monomer; is the compression ratio; is the contraction factor, used to suppress and control the speed magnitude; is the position where the whale is at the k - th iteration; is the probability selection factor; is the non - linear convergence factor; is the Levy flight search strategy; is the self - learning factor; is the social learning factor; is a random number between [0, 1]; is the position of the particle at the t - th iteration; is the position corresponding to the global optimal fitness value at the t - th generation; is the position corresponding to the optimal fitness value of the i - th particle monomer in the particle swarm.
[0012] Preferably, the calculation formula of the probability selection factor is as follows: (5) (6) The non - linear convergence factor has the following calculation formula: (7) In the formula: M is the population size of the particle swarm; n is the set parameter; G is the dimension of the search space; The subscript i represents the i - th particle in the population; is the particle fitness value; is the fitness of the optimal individual; is the average fitness of the population; is the base of the natural logarithm; is the number of iterations, is the maximum number of iterations.
[0013] Preferably, the Lévy flight search strategy has the following calculation formula: (8) The compression ratio has the following calculation formula: (9) In the formula: is to control the search direction; is a random number, and its value range is (0, 1); is a constant; is the gamma function.
[0014] Preferably, the dynamic weight allocation algorithm based on the Gini index in step S4 further includes: According to the output result of the Gini index of formula (3), a dynamic weight allocation algorithm based on the Gini index is proposed, and its dynamic weight calculation method is as follows: (10) In the formula: is the adjustment coefficient; is the decision tree Gini index.
[0015] 1. Preferably, the majority voting method in step S4 further includes: Using the majority voting method to select the prediction result. When the number of votes obtained by a prediction label is greater than half of the total number of votes of the decision trees, the final prediction is this prediction label; otherwise, this prediction label will be rejected. The specific formula is as follows: (11) In the formula: is the number of categories; is the marked category; is the number of categories; is the number of decision trees; is the dynamic weight; is the adjustment coefficient; Taking the monitoring state with the largest result proportion as the final result of the leakage monitoring method based on the random forest algorithm , and then obtaining the corresponding category.
[0016] The present invention also provides a leakage fault location and monitoring system for a distribution photovoltaic access substation area, which operates the foregoing leakage fault location and monitoring method for a distribution photovoltaic access substation area, including: A data acquisition and processing module, which is used to obtain the actual operating conditions of each sample, collect the residual current data of the substation area, set the normal and leakage fault labels, and record the fault leakage current signal; A data denoising processing module, which is used to perform denoising processing on the preprocessed fault leakage current signal through an improved hybrid threshold function, construct a feature data set based on the denoised fault leakage current signal, and divide it into a training set, a validation set, and a test set; A random forest construction module, which is used to jointly form a random forest model with all decision trees constructed by relying on the Gini index to divide the sample set using the classification and regression tree algorithm; A model parameter tuning module, which is used to iteratively optimize the relevant parameters in the random forest through a particle swarm algorithm improved by adopting a dual strategy, and complete the parameter tuning of the random forest performance model; A fault label discrimination module, which is used to take the monitoring state with the largest result proportion of the input test set as the final result of the leakage monitoring method based on the random forest algorithm through a dynamic weight allocation algorithm based on the Gini index, and discriminate the leakage fault state label information according to the final result; The fault location verification module is used to perform binary processing on the location of the leakage fault based on the leakage fault status tag information by the binary method, output the predicted leakage fault point, and compare it with the actual fault point in the test set to verify the random forest model; The fault location determination module is used to obtain the new remaining current data of each sample operating condition in real time, import it into the random forest model, and obtain the leakage fault point corresponding to the fault leakage current, so as to complete the leakage fault location monitoring of the distribution area.
[0017] Compared with the prior art, considering that the leakage signal in the low-voltage distribution area with distributed photovoltaic access is vulnerable to multiple factors and it is difficult to achieve accurate online leakage fault location, the present invention applies the random forest algorithm to the field of leakage fault location in the low-voltage distribution area, and finally realizes high-precision online leakage fault location in the low-voltage distribution area. The beneficial effects include: 1) By improving the hybrid threshold function, noise can be accurately separated, more useful information can be retained, fault information that has no impact on the positioning result can be eliminated, constant deviation can be effectively eliminated, the hybrid denoising effect can be achieved, and 、 adjustment parameters and adaptive parameters can more effectively separate signals and noise, better screen out useful information, reduce data redundancy, flexibly select between threshold functions, and finally achieve a better denoising effect.
[0018] 2) The present invention improves the calculation formula of the Gini index by adding a constant coefficient which can smooth the change of the Gini index, make the Gini index have better stability on data sets of different sizes, make the Gini index more flexible, and improve the generalization performance of the model on different data sets.
[0019] 3) The present invention proposes a particle swarm optimization algorithm improved by a dual strategy, introducing a probability selection factor 、specific gravity compression and a non-linear convergence factor and n setting parameters to improve the particle swarm optimization algorithm, improve the local search ability of the particle swarm optimization algorithm, improve the convergence ability of the algorithm, and have the ability to jump out of the local optimum twice; ensure fast and accurate convergence speed, reduce the probability of the algorithm falling into the local optimum, reduce unnecessary iterations, promote the algorithm to converge to the global optimum solution, and improve the rapidity and accuracy of the algorithm.
[0020] 4) The improved Gini index dynamic weighted decision algorithm adds an adjustment coefficient to avoid the deviation caused by the large number of feature values, alleviate the problem of unclear classification caused by the close probability values, enhance the classification accuracy, and improve the accuracy of classification. Description of the Drawings
[0021] Figure 1 The flow chart of an online leakage fault location method based on the random forest algorithm disclosed in an embodiment of the present invention; Figure 2 The signal denoising flow chart disclosed in an embodiment of the present invention; Figure 3 The flow chart of establishing a random forest disclosed in an embodiment of the present invention; Figure 4 The parameter tuning flow chart of the random forest performance model in the present invention. Specific embodiments
[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0023] To solve the problems existing in the online monitoring of leakage fault location in low-voltage distribution areas at the present stage, the present invention optimizes from aspects of data preprocessing, leakage monitoring and online monitoring of leakage fault location in low-voltage distribution areas. Firstly, the normalization algorithm is optimized and a denoising algorithm of a mixed sign function is proposed; secondly, the surrounding and hunting strategies of the whale optimization algorithm are introduced into the particle swarm optimization algorithm, and then the Levy flight strategy is added to the particle swarm optimization algorithm, and a particle swarm optimization algorithm with dual strategies is proposed; then the random forest model is evaluated, and according to the evaluation results, the dichotomy method is used to realize the verification of the leakage fault location, and new data is imported into the random forest model in real time to obtain the leakage fault point corresponding to the fault leakage current, and finally the leakage position and leakage line are found in a timely and accurate manner, ensuring the accurate location of the leakage fault in the low-voltage distribution area.
[0024] As Figures 1 to 4 shown, Embodiment 1 of the present invention provides a leakage fault location monitoring method for a distributed photovoltaic access area, including the following steps: S1: Collect the residual current data of the area according to the operating conditions of each sample, set the normal and leakage fault labels, record the fault leakage current signal, and then perform preprocessing. After denoising the preprocessed fault leakage current signal through an improved mixed threshold function, a feature dataset is constructed and divided into a training set, a validation set and a test set; Step S1 specifically includes: S1.1: Collect the residual current data of the area according to the actual operating conditions of each sample, set the normal and leakage fault labels, represented by 0 and 1 respectively, and record the fault leakage current signal; S1.2: Based on the fault leakage current signal collected and recorded in step S1.1, then periodically segment the fault leakage signal data according to the sliding time window; S1.3: Perform normalization preprocessing on the segmented data obtained in step S1.2; To improve the generalization ability, accuracy, and prediction fitness of leakage monitoring, ensure the reliable classification of fault leakage current signals, and prevent the data from being submerged during training due to the difference in the order of magnitude in the data, before the fault leakage current signal data vector is used as input data, it is necessary to perform normalization preprocessing. The calculation formula is as follows: (1) In the formula: is the data sample at a certain moment; is the largest sample data among all samples; is the smallest sample data among all samples.
[0025] S1.4: Denoise the fault leakage current signal obtained after preprocessing in step S1.3.
[0026] Since most of the working environments for leakage fault monitoring are outdoors, if the signal containing noise is directly selected for leakage, it is very likely to cause misjudgment using incorrect information. Therefore, it is necessary to perform noise reduction processing before extracting the feature data of the leakage selection criterion, ensure that noise does not enter again while extracting useful information, reduce redundant information, eliminate the problem of constant deviation, achieve the effect of hybrid denoising, and increase 、 adjustment parameters and adaptive parameters to better screen out useful information, reduce data redundancy, realize flexible selection between threshold functions, and finally achieve a better denoising effect.
[0027] In the embodiment of the present invention, corresponding to Figure 1 in the "Fault Characteristic Analysis and Feature Extraction" part, the fault leakage current signal after preprocessing is denoised by improving the hybrid threshold function. This algorithm can perform fine analysis on the leakage signal. According to the different time-frequency characteristics of the useful signal and noise at different decomposition scales, the useful signal and noise can be separated and extracted. At the same time, it is ensured that noise does not enter again while extracting effective information.
[0028] As Figure 2 shown, to realize flexible selection between threshold functions, it is necessary to further denoise the fault leakage current signal obtained after preprocessing.
[0029] Further preferably, step S1.4 includes: a. For the part where the absolute value of the wavelet coefficient is greater than the threshold, shrink it from the sign function to the threshold point to retain most of the useful information; b. Introduce a lower threshold, perform a high-order function processing on the part of the wavelet coefficient between the two thresholds, and extract the remaining small amount of useful information; c. Set the wavelet coefficients less than the lower threshold to zero; thus, an improved hybrid threshold function is obtained, and the calculation formula is: (2) In the formula: is the wavelet coefficient after being processed by the improved hybrid threshold function; is the wavelet coefficient of the unimproved hybrid threshold function; is the adaptive parameter; is the sign function; are the upper and lower limits of the threshold; is the adjustment parameter; and , .
[0030] S1.5: Based on the above steps S1.1 to S1.4, construct a feature dataset using the denoised fault leakage current signal, and divide the feature dataset into a training set , a validation set and a test set in proportion.
[0031] In the embodiment of the present invention, when dividing the feature dataset proportionally, a ratio of 7:1:2 is usually adopted to divide the feature dataset into a training set, a validation set, and a test set.
[0032] S2: Sample and construct a training sample set from the training set obtained in step S1, and use the classification and regression tree algorithm to rely on the Gini index to divide the sample set to construct a decision tree. All the constructed decision trees together form a random forest model; as Figure 3 shown, step S2 specifically includes: S2.1: From the training dataset , adopt the Bootstrap sampling method to perform resampling with replacement to obtain a dataset with the same number as the original dataset, and repeat the above operation times to construct training samples , where each dataset includes N samples and m feature variables for constructing a decision tree; S2.2: For each randomly selected feature and data sample, construct a decision tree using the feature variable and the data sample. The construction process of the decision tree includes selecting split nodes, split features, split thresholds, etc.; S2.3: The splitting criterion of the decision tree is usually measured using the Gini impurity. The present invention adopts the Classification and Regression Tree (CART) decision tree algorithm. The decision tree relies on the Gini index to partition the sample set. The Gini coefficient measures the degree of improvement in the purity of the sub-datasets after partitioning under a certain value of the feature Then, the CART decision tree minimizes the Gini index The calculation formula is: (3) In the formula: represents the Gini coefficient of the data training set under a certain feature value of the feature ; is a constant coefficient; and represent the proportion of the number of samples in the sub-datasets D1 and D2 after partitioning according to the value of the feature ; Gini(D1) and Gini(D2) represent the Gini indices of the sub-datasets D1 and D2 respectively.
[0033] S2.4: Repeat step S2.1 multiple times to generate all decision trees constructed based on steps S2.2 to S2.3 for all training samples. All the decision trees together form a random forest model, and the construction of the random forest performance model is completed.
[0034] S3: Use the improved particle swarm optimization algorithm with an improved dual strategy to iteratively optimize the relevant parameters in the random forest for the random forest performance model constructed in step S2, and use the validation set in step S1 for verification to complete the parameter tuning of the random forest performance model.
[0035] After establishing the random forest model, the optimal parameters are still unknown. The setting of the parameters not only affects the prediction accuracy of the model, but also affects the training effect of the model, and has a great impact on the online leakage fault location and monitoring of the entire model. Therefore, it is necessary to optimize the parameters. Among them, the number of decision trees n_estimators, the maximum number of features max_features considered by the optimal model, and the maximum depth of the decision tree max_depth are important parameters of the random forest model.
[0036] As Figure 4 shown, step S3 specifically includes: S3.1: Initialize the parameters to be optimized in the random forest model, including the number of decision trees n_estimators, the maximum number of features max_features considered by the optimal model, and the maximum depth of the decision tree max_depth. Each particle in the particle swarm represents a set of parameters; S3.2: Continuously iterate and select the optimal parameters through the particle swarm algorithm improved by the dual strategy in Equation (4).
[0037] To achieve the best classification effect of the random forest, the present invention introduces the encircling and hunting strategies of the whale optimization algorithm into the particle swarm algorithm, adds the Levy flight strategy to the particle swarm algorithm and introduces it into the random forest model, and proposes a particle swarm algorithm improved based on the dual strategy. The probability selection factor , specific gravity compression and non-linear convergence factor and the n setting parameters are used to improve the particle swarm algorithm. The improved update formula is as follows: (4) In the formula: is the dimensional position component of the particle monomer; is the dimensional velocity component of the particle monomer; is the compression specific gravity; The contraction factor is used to suppress and control the speed magnitude; is the position where the whale is at the th iteration; is the probability selection factor; is the non-linear convergence factor; is the Lecy flight search strategy; is the self-learning factor; is the social learning factor; is a random number between [0, 1]; is the position of the particle at the t-th iteration; is the position corresponding to the global optimal fitness value of the t-th generation; is the position corresponding to the optimal fitness value of a single particle in the particle swarm.
[0038] The said probability selection factor has the following calculation formula: (5) (6) The said non-linear convergence factor has the following calculation formula: (7) In the formula: M is the population size of the particle swarm; n is the set parameter; G is the dimension of the search space; The subscript i represents the i-th particle in the population; is the particle fitness value; is the fitness of the optimal individual; is the average fitness of the population; is the base of the natural logarithm; is the number of iterations, is the maximum number of iterations.
[0039] The said Levy flight search strategy has the following calculation formula: (8) The said compression ratio has the following calculation formula: (9) In the formula: is to control the search direction; is a random number, and its value range is (0, 1); is a constant; is the gamma function.
[0040] S3.3: Judge whether the maximum number of iterations is reached, S3.4: If the maximum number of iterations is reached, stop the iteration and output the optimal parameters; otherwise, return to S3.2.
[0041] S3.5: Introduce the particle swarm optimization algorithm improved by the dual strategy into the model to iteratively optimize the parameters in the random forest, and finally achieve a better classification effect. After the parameters are determined by the particle swarm optimization algorithm, the performance of the model will be tested on the validation set to determine the effectiveness of these parameters. The validation set helps evaluate the performance of each parameter combination and ensures that the selected optimized particle swarm model has good generalization ability on unknown data.
[0042] The model parameters after optimizing the random forest by the particle swarm optimization algorithm improved by the dual strategy are shown in Table 1.
[0043] Table 1 Optimal parameter settings Parameter name Parameter setting n_estimators 50 max_features 5 max_depth 15 S4: Input the test set into the random forest model, and obtain the monitoring state with the largest proportion of test set results and identify the corresponding fault state label information based on the dynamic weight allocation algorithm and majority voting method based on the Gini index. According to the label information, perform binary processing on the location of the leakage fault based on the dichotomy method, output the predicted leakage fault point, and compare it with the actual fault point in the test set to realize the verification of the random forest model; Step S4 specifically includes: Take the test feature data set as the input of the random forest, and each decision tree in the random forest will output a working state discrimination result. According to the Gini index output result of the above formula (3), better state discrimination performance is obtained, and a dynamic weight allocation algorithm based on the Gini index is proposed, which increases the adjustment coefficient to avoid the deviation caused by the large number of feature values. The specific calculation method of its dynamic weight is as follows: (10) In the formula: is the adjustment coefficient; is the Gini index of the decision tree.
[0044] The present invention uses the majority voting method to select the prediction result. When a prediction label obtains more than half of the decision tree votes, the final prediction is this prediction label; otherwise, this prediction label will be rejected. The specific formula is as follows: (11) In the formula: is the number of categories; is the marked category; is the number of categories; is the number of decision trees; is the dynamic weight; is the adjustment coefficient; Take the monitoring state with the largest result proportion as the final result of the leakage monitoring method based on the random forest algorithm , and then obtain the corresponding categories (normal state label: "0", leakage fault state label: "1") S5: Real-time obtain the operating conditions of each sample, collect new residual current data of the substation area, import it into the random forest model in step S4 to obtain the leakage fault point corresponding to the fault leakage current, and complete the leakage fault location monitoring of the substation area.
[0045] S6: After completing steps S1 to S5, confirm the final fault point, and the work process ends.
[0046] Embodiment 2 of the present invention provides a leakage fault location monitoring system for a distributed photovoltaic access substation area, which operates a leakage fault location monitoring method for a distributed photovoltaic access substation area as described in Embodiment 1, including: The data acquisition and processing module is used to obtain the actual operating conditions of each sample, collect the residual current data of the substation area, set normal and leakage fault labels, and record the fault leakage current signal; The data denoising processing module is used to denoise the preprocessed fault leakage current signal through an improved hybrid threshold function, construct a feature data set based on the denoised fault leakage current signal, and divide it into a training set, a validation set, and a test set; The random forest construction module is used to jointly form a random forest model by all decision trees constructed by relying on the Gini index to divide the sample set using the classification and regression tree algorithm; The model parameter tuning module is used to iteratively optimize the relevant parameters in the random forest through a particle swarm algorithm improved by a dual strategy to complete the parameter tuning of the random forest performance model; The fault label discrimination module is used to take the monitoring state with the largest result proportion of the input test set as the final result of the leakage monitoring method based on the random forest algorithm through the dynamic weight allocation algorithm based on the Gini index and using the majority voting method, and discriminate the leakage fault state label information according to the final result; The fault location verification module is used to perform binary processing on the location of the leakage fault based on the leakage fault state label information using the dichotomy method, output the predicted leakage fault point, and compare it with the actual fault point of the test set to realize the verification of the random forest model; The fault location determination module is used to real-time obtain the operating conditions of each sample, collect new residual current data of the substation area, import it into the random forest model to obtain the leakage fault point corresponding to the fault leakage current, and complete the leakage fault location monitoring of the substation area.
[0047] Compared with the prior art, considering that the leakage signals in low-voltage distribution substations with distributed photovoltaic access are vulnerable to multiple factors and it is difficult to achieve accurate online positioning of leakage faults, the present invention applies the random forest algorithm to the field of leakage fault positioning in low-voltage distribution substations, and finally realizes high-precision online positioning of leakage faults in low-voltage distribution substations. The beneficial effects of the present invention include: 1) By improving the hybrid threshold function, noise can be accurately separated, more useful information can be retained, fault information that has no effect on the positioning result can be eliminated, constant deviation can be effectively eliminated, the hybrid denoising effect can be achieved, and the and adjustment parameters and adaptive parameters can more effectively separate signals and noise, better screen out useful information, reduce data redundancy, flexibly select between threshold functions, and finally achieve a better denoising effect.
[0048] 2) The calculation formula of the Gini index is improved, and a constant coefficient is added to smooth the change of the Gini index, making the Gini index have better stability on datasets of different sizes, making the Gini index more flexible, and improving the generalization performance of the model on different datasets.
[0049] 3) The present invention proposes a particle swarm optimization algorithm improved by a dual strategy, introducing a probability selection factor , specific gravity compression and a non-linear convergence factor and n setting parameters to improve the particle swarm optimization algorithm, improving the local search ability of the particle swarm optimization algorithm, enhancing the convergence ability of the algorithm, and having the ability to jump out of local optima twice; ensuring a fast and accurate convergence speed, reducing the probability of the algorithm falling into local optima, reducing unnecessary iterations, promoting the algorithm to converge to the global optimum solution, and improving the rapidity and accuracy of the algorithm.
[0050] 4) The improved Gini index dynamic weighted decision algorithm adds an adjustment coefficient to avoid deviation caused by many feature values, alleviate the problem of unclear classification caused by close probability values, enhance classification accuracy, and improve the accuracy of classification.
[0051] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A method for locating and monitoring leakage faults in distributed photovoltaic access areas, characterized in that: The following steps are involved: S1: Collect the residual current data of the substation according to the operating conditions of each sample, set the normal and leakage fault labels, record the fault leakage current signal and perform preprocessing, denoise the preprocessed fault leakage current signal by improving the mixed threshold function, and then construct a feature data set and divide it into training set, verification set and test set; S2: Sampling and constructing a training sample set from the training set obtained in step S1, using the classification and regression tree algorithm to segment the sample set based on the Gini index to construct a decision tree, and all the constructed decision trees together constitute a random forest model; S3: The random forest performance model constructed in step S2 is iteratively optimized using an improved dual strategy improved particle swarm algorithm to optimize relevant parameters in the random forest and verified using the validation set in step S1 to complete parameter tuning of the random forest performance model; S4: Input the test set into the random forest model, obtain the monitoring state with the largest proportion of the test set results based on the dynamic weight allocation algorithm of the Gini index and the majority voting method, and identify the corresponding fault state label information. Based on the label information, the location of the leakage fault is binary processed based on the binary method to output the predicted leakage fault point and compare it with the actual fault point of the test set to realize the verification of the random forest model; S5: Obtain the new residual current data of the substation under each sample operating condition in real time and import it into the random forest model to obtain the leakage fault point corresponding to the fault leakage current, and complete the leakage fault location monitoring of the substation.
2. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 1, characterized in that: The pre-processing in step S1 further comprises: The obtained fault leakage current signal data is normalized, and the calculation formula is as follows: (1) Where: is a data sample at a certain moment; is the largest sample data among all samples; is the minimum sample data among all samples.
3. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 2, characterized in that: The step S1 of performing denoising on the pre-processed fault leakage current signal by improving the mixed threshold function further comprises: a. For the part where the absolute value of the wavelet coefficient is greater than the threshold, the sign function is used to shrink it toward the threshold point to retain the valid information; b. Introduce a lower threshold, process the wavelet coefficients between the two thresholds with a higher-order function, and extract the remaining valid information; c. Set the wavelet coefficients that are less than the lower threshold to zero; and then obtain the improved mixed threshold function, the calculation formula is: (2) Where: To improve the wavelet coefficients after mixed threshold function processing; The wavelet coefficients processed by the unmodified hybrid threshold function; is the adaptive parameter; is a symbolic function; is the upper threshold value; is the lower threshold; To adjust the parameters; and , .
4. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 1, characterized in that: The step S2 of segmenting the sample set based on the Gini index to construct a decision tree further includes: The splitting criterion of the decision tree is measured by Gini impurity. The classification regression tree decision tree algorithm is used to split the sample set based on the Gini index. The Gini coefficient measures the impurity of the feature. Under a certain value of , the purity of the sub-dataset after division is improved, and the decision tree Gini index , the calculation formula is: (3) Where: Indicated in the feature Data training set under a certain eigenvalue The Gini coefficient of is a constant coefficient; and According to the characteristics The percentage of samples in sub-datasets D1 and D2 after the value division; Gini(D1) and Gini(D2) represent the Gini index of sub-datasets D1 and D2 respectively.
5. The method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 1, characterized in that: The step S3 of iteratively optimizing the relevant parameters in the random forest using the particle swarm algorithm improved by the dual strategy further includes: The random forest performance model introduces the encirclement and hunting strategies of the whale optimization algorithm into the particle swarm algorithm, and adds the Levy flight strategy to the particle swarm algorithm to introduce it into the model, and proposes a particle swarm algorithm based on a dual strategy improvement, and its improved update formula is as follows: (4) Where: For the The particle monomer dimensional position component; For the The particle monomer dimensional velocity component; is the compression specific gravity; The contraction factor is used to suppress the speed; For the The whale is already in position at the first iteration; Select factors for probability; is the nonlinear convergence factor; Search strategy for Lecy flight; For the self-learning factor; is the social learning factor; is a random number between [0,1]; is the position of the particle at the tth iteration; is the position corresponding to the global optimal fitness value of the tth generation; It is the corresponding position of the optimal fitness value of the particle monomer in the particle swarm.
6. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 5, characterized in that: The probability selection factor The calculation formula is as follows: (5) (6) The nonlinear convergence factor The calculation formula is as follows: (7) Where: M is the population size of the particle swarm; n is the setting parameter; G is the dimension of the search space; The subscript i represents the i-th particle in the population; For particles Fitness value; is the fitness of the optimal individual; is the average fitness of the population; is the base of natural logarithms; is the number of iterations, is the maximum number of iterations.
7. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 1, characterized in that: The Lecy Fly Search Strategy The calculation formula is as follows: (8) The compression weight The calculation formula is as follows: (9) Where: To control the search direction; is a random number, and its value range is (0, 1); is a constant; is the gamma function.
8. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 1, characterized in that: The dynamic weight allocation algorithm based on the Gini index in step S4 further includes: According to the Gini index output result of formula (3), a dynamic weight allocation algorithm based on the Gini index is proposed. The dynamic weight calculation method is as follows: (10) Where: is the adjustment factor; is the Gini index of the decision tree.
9. A method for locating and monitoring leakage faults in distributed photovoltaic access areas according to claim 1, characterized in that: The majority voting method in step S4 further includes: The prediction results are selected using the majority voting method. When the number of votes for a prediction label is greater than half of the total number of votes in the decision tree, the final prediction is the prediction label. Otherwise, the prediction label will be rejected. The specific formula is as follows: (11) Where: is the number of categories; is the tag category; is the number of categories; is the number of decision trees; is the dynamic weight; is the adjustment factor; The monitoring state with the largest proportion of results is taken as the final result of the leakage monitoring method based on the random forest algorithm , and then get the corresponding category.
10. A distributed photovoltaic access area leakage fault location monitoring system, running a distributed photovoltaic access area leakage fault location monitoring method as claimed in any one of claims 1 to 9, characterized in that: include: The data acquisition and processing module is used to obtain the actual operating conditions of each sample, collect the residual current data of the substation, set the normal and leakage fault labels and record the fault leakage current signal; A data denoising processing module is used to denoise the preprocessed fault leakage current signal by improving the mixed threshold function, and to construct a feature data set based on the denoised fault leakage current signal and divide it into a training set, a validation set and a test set; The random forest construction module is used to combine all decision trees constructed by using the classification and regression tree algorithm to segment the sample set based on the Gini index to form a random forest model; The model parameter tuning module is used to iteratively optimize the relevant parameters in the random forest by using the particle swarm algorithm improved by the dual strategy, and complete the parameter tuning of the random forest performance model; A fault label identification module is used to use a dynamic weight allocation algorithm based on the Gini index and a majority voting method to take the monitoring state with the largest proportion of the input test set results as the final result of the leakage monitoring method based on the random forest algorithm, and to identify the leakage fault state label information according to the final result; The fault location verification module is used to perform binary processing on the location of the leakage fault based on the binary method according to the leakage fault state label information, output the predicted leakage fault point and compare it with the actual fault point of the test set to realize the verification of the random forest model; The fault location determination module is used to obtain the new residual current data of the substation under each sample operating condition in real time and import it into the random forest model to obtain the leakage fault point corresponding to the fault leakage current, so as to complete the leakage fault location monitoring of the substation.
Citation Information
Cited By
Dynamic risk assessment and suppression method for leakage current of photovoltaic access transformer area based on neural network and CVaR prediction algorithm
CN120952552A