Low-voltage transformer area electric leakage risk intelligent identification method and system based on machine learning
By constructing a multidimensional leakage current feature and graph structure model using machine learning methods, and combining G-GTBoost and Res-CNN-LSTM models, the problem of insufficient identification capability of traditional leakage current protection methods in low-voltage distribution areas is solved, and efficient and accurate leakage current risk assessment and classification are achieved.
Patent Information
- Application Number
- CN202511034036.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional leakage current protection methods struggle to identify complex leakage characteristics in low-voltage distribution areas, lack early warning capabilities, have high false alarm and missed alarm rates, cannot achieve trend prediction and risk classification, and cannot integrate multi-dimensional data for refined assessment.
A machine learning-based intelligent identification method for leakage risk in low-voltage distribution areas is adopted. Through multi-dimensional leakage feature construction and distribution area information aggregation analysis, combined with graph structure modeling, graph message passing mechanism and G-GTBoost risk scoring model, leakage risk is dynamically assessed. An improved Res-CNN-LSTM model is used for intelligent identification of leakage type and confidence determination.
It enables accurate identification and classification of leakage risks in low-voltage distribution areas, reduces false alarm and false alarm rates, improves identification accuracy and response speed, supports dynamic threshold adjustment and auxiliary decision-making, and enhances the reliability and efficiency of leakage detection.
Smart Images

Figure CN120929953A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-voltage distribution area leakage risk identification technology, and in particular to a machine learning-based intelligent identification method and system for low-voltage distribution area leakage risk. Background Technology
[0002] Low-voltage distribution networks, as a crucial link connecting the power system and end users, are widely used in residential, commercial, agricultural, and industrial sectors. However, with the increasing complexity of distribution network structures, the diversification of user-side electrical equipment, the frequent occurrence of aging lines, and the harsh environment for cable laying, leakage faults in low-voltage distribution areas exhibit characteristics such as high incidence, concealment, and diversity, seriously threatening the stable operation of the power supply system and the safety of people and property.
[0003] Traditional leakage current protection methods mainly rely on residual current operated devices (RCDs), which trip the circuit breaker by detecting whether the current amplitude exceeds a set threshold. This method has the following shortcomings: it has a weak ability to identify complex leakage current characteristics and it is difficult to distinguish between normal leakage current and dangerous leakage current; it cannot achieve trend prediction and risk classification, and it only operates when the leakage current exceeds the threshold, lacking early warning capability; the false alarm and missed alarm rates are high, especially in power supply areas with large load changes and complex grounding methods.
[0004] With the development of smart grid technology, machine learning methods are being applied more and more widely in power systems, showing great potential, especially in anomaly detection, pattern recognition and intelligent diagnosis. However, there is currently a lack of intelligent identification methods that can integrate multi-dimensional data such as leakage current waveform characteristics, historical trend information, transformer topology and electrical parameters for real-time perception and refined assessment of leakage risk in low-voltage distribution systems. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a machine learning-based intelligent identification method for leakage current risks in low-voltage distribution areas, comprising the following steps:
[0006] Step 1: Construction of multidimensional leakage current characteristics and aggregation and analysis of transformer area information;
[0007] Step 2: Dynamic assessment of leakage risk and adaptive adjustment of threshold;
[0008] Step 3: Intelligent identification and confidence level determination mechanism for leakage current type.
[0009] As a further supplement to this technical solution, the construction of multi-dimensional leakage current characteristics and the aggregation and analysis of transformer area information include the following steps:
[0010] Step 1) Data acquisition;
[0011] Deploy multi-point sensing devices to synchronously collect the residual current i of each branch in the transformer area.r (t), Three-phase voltage u abc (t), Load current i abc (t), temperature and humidity (T, H); enable time synchronization mechanism to ensure data is alignable and time-stamped;
[0012] Step 2) Data preprocessing;
[0013] The waveform signal is filtered to remove spikes and environmental noise; data standardization is performed using the Z-score standardization method.
[0014]
[0015] Step 3) Feature extraction to construct feature vector F1;
[0016] Representative feature vectors are constructed from the original acquired signals. A feature system is built through three dimensions: time domain features, frequency domain features, and topological and environmental features, and then fused to form the feature vectors.
[0017] As a further supplement to this technical solution, the dynamic assessment of leakage risk and adaptive threshold adjustment includes the following steps:
[0018] Step 1): Graph structure modeling and input representation;
[0019] Step 2): Diagram of message transmission mechanism and risk scoring process;
[0020] Step 3): Construction of the G-GTBoost risk scoring model.
[0021] As a further supplement to this technical solution, the graph structure modeling and input representation includes the following steps:
[0022] Step 1: Define the graph structure;
[0023] The entire transformer area is modeled as a directed graph G = (V, E), where V = {v1, v2, ..., v...}. n} represents each monitoring node in the distribution area, such as transformer outgoing lines, branch lines, and user meter boxes; Indicates electrical connection relationships; each edge e ij =(v i ,v j ) indicates the direction of electrical influence from node i to j;
[0024] Step 2: Node feature representation;
[0025] Each node v i The collected multidimensional features are represented as vectors:
[0026] x i =[PF iTHD i ,I rms,i ,ΔI i ,T i H i ,L i ,ρ i ,...]∈R d
[0027] It includes residual current characteristics within the time window, frequency domain disturbance characteristics, environmental parameters, and location factors;
[0028] All nodes constitute the input matrix:
[0029] X = [x1, x2, ..., x n ] T ∈R n×d .
[0030] As a further supplement to this technical solution, the graph message passing mechanism and risk scoring process include the following steps:
[0031] Step (1) Define the adjacency matrix;
[0032] Let A∈R n×n Let A be the adjacency matrix of the graph, representing the topological connections between nodes. ij =1 indicates that v exists. j →v i The information flow supports weighted edges, and voltage level difference and cable length can be added as edge weights;
[0033] Step (2) Feature enhancement aggregation;
[0034] Perform a GCN (Graph Convolutional Network) aggregation operation once for the input features of each node:
[0035]
[0036] Where N(i) represents the set of neighbors of node i; σ represents a nonlinear activation function, such as ReLU; W1∈R d×d′ The weight matrix represents the graph neural network; This represents the feature representation after aggregating neighbor information.
[0037] As a further supplement to this technical solution, the construction of the G-GTBoost risk scoring model includes the following steps:
[0038] Step 1: Input the basic features;
[0039] For any node v in the transformer area graph G=(V,E) i For each ∈V, its enhanced feature vectors are obtained through graph convolution:
[0040]
[0041] in, It is a feature representation that integrates local and neighbor information, and it is the input basis of GBDT;
[0042] Step 2: Gradient boosting tree regression mechanism;
[0043] Gradient boosting trees are an ensemble model that is trained iteratively, with one decision tree f trained in each round. t This is used to approximate the residual loss from the previous time, which is the difference between the actual risk value and the current model prediction.
[0044] The output of the model in round t is:
[0045]
[0046] in, f is the risk score prediction for node i in the t-th round of the model; η is the learning rate; f t The output of the t-th regression tree; initial prediction That is, the average risk value;
[0047] Step 3: Definition of loss function and derivation of residuals;
[0048] Let the real label be This represents the risk score calculated empirically, and the loss function used is the squared error loss:
[0049]
[0050] True risk value Configuration: If a tripping / grounding event occurs in the transformer area during a certain period, the value is assigned to 1.0; if an equipment malfunction occurs but the circuit does not trip, the value is set to 0.6 to 0.8; in normal conditions, the value is assigned to 0.0; when direct labels are lacking, they can be calculated through expert rules.
[0051] The target residual to be fitted in round t is:
[0052]
[0053] This is also known as "the error between the predicted value and the actual value of the previous model," and this error is determined by the new tree f. t Go and learn:
[0054]
[0055] Each subtree f t The learning objective is to approximate the residuals as closely as possible, thereby gradually approaching the true risk score.
[0056] Step 4: Overall model output and interpretability;
[0057] The final output of the model in round T is:
[0058]
[0059] Among them, f t (·) represents the regression function formed by several splitting conditions and leaf nodes in the tree model; ω t The weights are calculated as a whole; the output of each leaf node represents the risk increment for that category.
[0060] As a further supplement to this technical solution, the intelligent identification and confidence level determination mechanism for leakage current type includes the following steps:
[0061] Step 1: Model structure design;
[0062] Step 2: Convolutional Feature Extraction Module;
[0063] Step 3: Residual connection;
[0064] Step 4: LSTM timing modeling module;
[0065] Step 5: Classification and confidence output mechanism.
[0066] As a further supplement to this technical solution, the model structure design includes feature input and encoding, padding the temporal feature vectors into a fixed-dimensional tensor:
[0067] X∈R T×N
[0068] Where T is the number of time steps; N is the feature dimension of each step; X = [x1,...,x] T ] T Each line represents a time-time characteristic;
[0069] The convolutional feature extraction module uses a two-level one-dimensional convolutional structure to extract local spatial features of the signal;
[0070] The first-level convolution operation is as follows:
[0071] Z (1) =σ1(Conv1D1(X)+b1)
[0072] The second-level convolution operation is as follows:
[0073] Z (2) =σ2(Conv1D2(Z) (1) )+b2)
[0074] Where σ1 and σ2 are activation functions; Conv1D kOne-dimensional convolution operation with a kernel length of k; Z (2) It includes deep local features after convolution;
[0075] The residual connection directly connects the input X to Z across layers. (2) Then, feature fusion is performed to form enhanced features:
[0076] Z res =Z (2) +Proj(X)
[0077] Here, Proj(·) represents the projection transformation, ensuring dimensionality matching; the residual structure ensures that shallow information is not lost in deep networks, which is beneficial for stable training and preservation of multimodal features; it also provides dual-path features: Z (2) For nonlinear enhancement, X provides compensation for the original structure;
[0078] The LSTM temporal modeling module includes the residual-enhanced tensor Z. res The input is fed into stacked two-layer LSTM units to model the dynamic evolution and state transition patterns between feature sequences:
[0079] First layer LSTM:
[0080]
[0081] Second-layer LSTM:
[0082]
[0083] Finally, the global state representation is obtained:
[0084]
[0085] The classification and confidence output mechanism includes:
[0086] Classification prediction: Mapping the LSTM output to the leakage current type space through a fully connected layer:
[0087]
[0088] Use Softmax to calculate the probability distribution for each type:
[0089]
[0090] Classification results:
[0091]
[0092] Confidence output: Definition:
[0093]
[0094] Specifically, when γ≥0.85, the system has a high degree of confidence and will act automatically; when 0.6≤γ<0.85, the system has a moderate degree of confidence and will trigger an early warning; when γ<0.6, the system makes a fuzzy judgment and will recommend that the system use an expert system to assist in decision-making.
[0095] A machine learning-based intelligent identification system for leakage current risks in low-voltage distribution areas is provided, which uses the aforementioned machine learning-based intelligent identification method for leakage current risks in low-voltage distribution areas to identify leakage current risks.
[0096] A computer-readable storage medium, wherein computer program instructions are stored on the computer-readable storage medium, the computer instructions causing the computer to perform the method described above.
[0097] Its beneficial effect lies in the fact that by introducing the graph-enhanced gradient boosting tree model (G-GTBoost) and the improved residual convolutional-temporal neural network model (Res-CNN-LSTM), combined with multimodal feature fusion and dynamic threshold adjustment mechanism, significant performance improvement is achieved in leakage risk identification and classification. Attached Figure Description
[0098] Figure 1 This is a schematic diagram of the overall method flow of the present invention;
[0099] Figure 2 This is a schematic diagram of the specific process of the intelligent identification and confidence determination mechanism for leakage current type in step 3 of the present invention. Detailed Implementation
[0100] To facilitate a clearer understanding of this technical solution for those skilled in the art, the following will be described in conjunction with the appendix. Figure 1-2 The technical solution of the present invention is described in detail below:
[0101] A machine learning-based intelligent identification method for leakage current risk in low-voltage distribution areas includes the following steps:
[0102] Step 1: Construction of multidimensional leakage current characteristics and aggregation and analysis of transformer area information;
[0103] This step aims to perform high-precision modeling of leakage current behavior at different nodes within a low-voltage distribution area. By deploying various types of sensors (including residual current detection devices, voltage monitoring modules, and environmental sensing units) within the distribution area, key electrical parameters and operational status data are collected. Time-domain, frequency-domain, and statistical analyses are performed on the raw data to extract multi-scale, multi-angle leakage current characteristics, such as spikes, disturbances, and fluctuation frequencies. Based on this, structural information such as the distribution topology, power supply path within the distribution area, and load characteristics are incorporated into the feature set to form a unified multimodal feature vector. This vector comprehensively reflects the spatial structure, electrical response characteristics, and environmental correlation of leakage current behavior, providing rich data support for subsequent modeling.
[0104] Step 2: Dynamic assessment of leakage risk and adaptive adjustment of threshold;
[0105] Based on the high-dimensional feature vector obtained in the first step, this step constructs a machine learning risk assessment model to quantitatively score the leakage risk under the current operating conditions. Unlike traditional alarm mechanisms based on fixed values, this invention introduces a structure-enhanced gradient boosting tree model, which learns from historical operating samples and fault records to accurately identify complex leakage trends. Furthermore, considering the changing operating characteristics of the power distribution system under the influence of seasons, load fluctuations, and meteorological factors, this step designs a dynamically adjustable discrimination threshold mechanism. By introducing factors such as temperature and humidity, current disturbance intensity, and harmonic distortion rate, the risk judgment boundary is adjusted in real time to reduce the probability of false alarms and missed alarms, thereby improving the robustness and reliability of the risk assessment.
[0106] Step 3: Intelligent identification and confidence level determination mechanism for leakage current type;
[0107] When the system detects a leakage risk, it needs to identify the specific leakage type to support more efficient operation and maintenance decisions. This step introduces a deep neural network model, fusing residual connection mechanisms with time-series modeling capabilities to achieve automatic classification and identification of different leakage types. The main body of the model adopts an improved CNN-LSTM structure, introducing residual structures into the convolutional layers to enhance the network's ability to capture high-frequency details, and extracting the evolution trend of leakage features in the time series in the recurrent layers. Finally, the leakage type label is output through a fully connected layer, supplemented by a Softmax confidence output to determine the reliability of the classification result. When the confidence of the identification result is insufficient, the system can trigger a "fuzzy judgment" processing mechanism, combining historical operating conditions, environmental trends, and other information for auxiliary judgment, enhancing fault tolerance under boundary conditions.
[0108] For step 1, the construction of multi-dimensional leakage current characteristics and the aggregation and analysis of transformer area information:
[0109] This step aims to comprehensively model the leakage current characteristics in low-voltage distribution areas. By deploying monitoring terminals at key nodes, high-temporal-resolution multi-source electrical signals are collected, and auxiliary information such as distribution area topology, equipment type, and environmental factors are integrated to construct a high-dimensional feature vector with discriminative capabilities. This provides the training and input foundation for subsequent machine learning recognition models. The specific process is as follows:
[0110] 1) Data collection;
[0111] Deploy multi-point sensing devices to synchronously collect the residual current i of each branch in the transformer area. r (t), Three-phase voltage u abc (t), Load current i abc(t), temperature and humidity (T, H), etc.; enable time synchronization mechanism to ensure that data has alignability and time labeling.
[0112] 2) Data preprocessing
[0113] Use filters (such as wavelet denoising and Butterworth filtering) to remove spikes and environmental noise from the waveform signal; data standardization is performed using the Z-score standardization method.
[0114]
[0115] 3) Feature extraction (constructing feature vector F1);
[0116] To achieve efficient identification and classification of leakage current risks within the transformer substation area, this step focuses on constructing representative feature vectors from the original acquired signals. Considering that leakage current behavior has significant characteristics in both time-domain waveform morphology and frequency-domain energy distribution, and is also constrained by the combined influence of transformer substation structure, electrical topology, and environmental factors, a feature system is constructed from the following three dimensions:
[0117] 1. Time-domain characteristics: Time-domain characteristics mainly reflect the trend, amplitude fluctuation and statistical characteristics of the signal in the time dimension, and are suitable for identifying sudden, intermittent or continuous leakage current behavior.
[0118] Peak factor (PF): Describes the ratio of the amplitude of a signal spike to its effective amplitude, and is suitable for identifying arc-type leakage current characteristics. It is defined as:
[0119]
[0120] Among them, i r (t) represents the residual current signal, and N represents the number of sampling points. If PF > 5, it usually indicates the presence of spike interference or intermittent breakdown.
[0121] Effective value I rms : Measures the energy intensity of electric current; it is a fundamental energy characterization index. Defined as:
[0122]
[0123] Skewness and kurtosis: These reflect the symmetry and sharpness of the waveform, respectively, and are common indicators for identifying non-Gaussian waveforms.
[0124]
[0125] Where μ and σ are the mean and standard deviation, respectively; a skewness of <0 indicates that the waveform is biased to the left, and a kurtosis of >3 indicates the presence of a spike anomaly.
[0126] 2. Frequency domain characteristics: Frequency domain characteristics are mainly used to detect the differences in energy distribution of leakage current behavior at different frequencies, and are especially suitable for identifying complex modes such as harmonic disturbances and breakdown oscillations.
[0127] Total Harmonic Distortion (THD): Measures the proportion of non-fundamental energy in the leakage current waveform and is a key indicator for judging "power quality type" leakage. Defined as:
[0128]
[0129] Among them, I n I1: Amplitude of the nth harmonic component; I2: Fundamental component; High THD values are often associated with leakage events caused by ground oscillations or external interference.
[0130] Spectral kurtosis: A metric describing the degree of energy concentration in a spectrum, suitable for identifying abrupt changes in broadband and narrowband signals.
[0131]
[0132] Measuring the shift trend in energy distribution can serve as a basis for identifying leakage currents caused by equipment aging or uneven insulation.
[0133] 3. Topological and environmental characteristics:
[0134] To improve the adaptability of the identification model to different grounding methods, wiring methods, temperature and humidity changes, the following auxiliary features are constructed.
[0135] Density of branch roads in the area
[0136]
[0137] Where, N branch Number of branch roads in the district, L total Total power supply distance; this metric measures the impact of a "mesh" or "chain" structure on the signal transmission path.
[0138] Node position weight factor
[0139]
[0140] Among them, L i Let be the line length from node i to the transformer; the farther the node, the greater the attenuation of the leakage signal, so spatial compensation weights need to be given in the model.
[0141] Environmental factors T and H: Leakage behavior is significantly affected by temperature and humidity, especially during humid seasons or at high temperatures, the probability of leakage increases; this invention collects on-site temperature and humidity data and introduces feature vectors to enable the model to dynamically perceive climatic factors.
[0142] 4. Feature Vector Construction: All the above statistical indicators, physical quantities, and environmental parameters are integrated to form the final feature input vector.
[0143] F1 = [PF, I rms ,skew,kurt,THD,γ f ,f c ,ρ,α i [,T,H]
[0144] This feature vector serves as input to subsequent machine learning models, enabling quantitative modeling and dynamic assessment of leakage risk in transformer substations.
[0145] Next is step 2: Dynamic assessment of leakage current risk and adaptive adjustment of threshold:
[0146] To more comprehensively and accurately assess the risk level of leakage current behavior in low-voltage distribution areas, this step introduces a graph neural enhanced gradient tree model (G-GTBoost) to model and dynamically score the leakage current risk of each node within the distribution area under spatiotemporal disturbances. Unlike the traditional GBDT model, the G-GTBoost model not only considers the characteristics of a single node but also integrates the correlation information between adjacent nodes in the distribution area topology, achieving regional awareness and collaborative risk assessment of leakage current behavior. This enhances the model's accuracy and upstream / downstream adaptability. The specific process is as follows:
[0147] 1) Graph structure modeling and input representation
[0148] Graph structure definition:
[0149] The entire transformer area is modeled as a directed graph G = (V, E), where V = {v1, v2, ..., v...}. n} represents each monitoring node in the distribution area, such as transformer outgoing lines, branch lines, user meter boxes, etc. Indicates electrical connection relationships; each edge e ij =(v i ,v j ) indicates the direction of electrical influence from node i to j.
[0150] Node feature representation:
[0151] Each node v i The collected multidimensional features are represented as vectors:
[0152] x i =[PF i THD i ,I rms,i ,ΔI i ,T i H i ,L i ,ρi ,...]∈R d
[0153] It includes residual current characteristics within the time window, frequency domain disturbance characteristics, environmental parameters, location factors, etc.
[0154] All nodes constitute the input matrix:
[0155] X = [x1, x2, ..., x n ] T ∈R n×d
[0156] 2) Diagram message transmission mechanism and risk scoring process;
[0157] Adjacency matrix definition:
[0158] Let A∈R n×n Let A be the adjacency matrix of the graph, representing the topological connections between nodes. ij =1 indicates that v exists. j →v i The information flow supports weighted edges, and voltage level difference, cable length, etc. can be added as edge weights.
[0159] Feature enhancement aggregation:
[0160] Perform a GCN (Graph Convolutional Network) aggregation operation once for the input features of each node:
[0161]
[0162] Where N(i) represents the set of neighbors of node i; σ represents a nonlinear activation function, such as ReLU; W1∈R d×d′ The weight matrix represents the graph neural network; This represents the feature representation after aggregating neighbor information.
[0163] 3) Construction of the G-GTBoost risk scoring model;
[0164] The G-GTBoost model combines the spatial structure modeling capability of graph neural networks with the nonlinear regression fitting capability of gradient boosting trees (GBDT) to achieve a quantitative score of leakage risk for each monitoring node in the transformer substation.
[0165] Input feature basics:
[0166] For any node v in the transformer area graph G=(V,E) i For each ∈V, its enhanced feature vectors are obtained through graph convolution:
[0167]
[0168] in, It is a feature representation that integrates local and neighbor information, and it is the input basis of GBDT.
[0169] Gradient boosting tree regression mechanism:
[0170] Gradient boosting trees are an ensemble model that is trained iteratively, with one decision tree f trained in each round. t It is used to approximate the residual loss of the previous time, which is the difference between the true risk value and the current model prediction value.
[0171] The output of the model in round t is:
[0172]
[0173] in, f is the risk score prediction for node i in the t-th round of the model; η is the learning rate; f t The output of the t-th regression tree; initial prediction That is, the average risk value.
[0174] Loss function definition and residual derivation:
[0175] Let the real label be This represents the risk score calculated empirically, and the loss function used is the squared error loss:
[0176]
[0177] True risk value Configuration: If a tripping / grounding event occurs in the transformer area during a certain period, the value is assigned to 1.0; if an equipment malfunction occurs but the circuit does not trip, the value is set to 0.6 to 0.8; in normal conditions, the value is assigned to 0.0; when there is a lack of direct labels, it can be calculated through expert rules.
[0178] The target residual to be fitted in round t is:
[0179]
[0180] This is also known as "the error between the predicted value and the actual value of the previous model," and this error is determined by the new tree f. t Go and learn:
[0181]
[0182] Each subtree f t The learning objective is to approximate the residuals as closely as possible, thereby gradually approaching the true risk score.
[0183] Overall model output and interpretability
[0184] The final output of the model in round T is:
[0185]
[0186] Among them, f t (·) represents the regression function formed by several splitting conditions and leaf nodes in the tree model; ω t The weights are integrated; the output of each leaf node represents the risk increment under that category; the model structure has strong interpretability (the frequency, path distribution, and importance weight of each feature in the tree can be derived).
[0187] Finally, step 3 involves an intelligent leakage current type identification and confidence level determination mechanism:
[0188] This step aims to further intelligently classify identified abnormal leakage behaviors to distinguish between different leakage types such as arcing, insulation aging, and moisture breakdown. A confidence index is output for the judgment results to assist in the graded response strategy. To improve the model's ability to discriminate complex waveforms, superimposed disturbances, and multimodal leakage signals, this invention proposes a residual-connected multilayer CNN-LSTM network structure (Res-CNN-LSTM). Based on the original convolutional + temporal model, a cross-layer feature skipping mechanism is introduced to enhance the model's expressive power and training stability.
[0189] 1) Model structure design;
[0190] Feature input and encoding:
[0191] The time-series feature vector (such as the residual current waveform) is padded to a fixed-dimensional tensor:
[0192] X∈R T×N
[0193] Where T is the number of time steps; N is the feature dimension of each step; X = [x1,...,x] T ] T Each line represents a time-time characteristic;
[0194] 2) Convolutional Feature Extraction Module:
[0195] Functions of convolutional structures:
[0196] A two-level one-dimensional convolutional structure is used to extract the local spatial features of the signal.
[0197] The first-level convolution operation is as follows:
[0198] Z (1) =σ1(Conv1D1(X)+b1)
[0199] The second-level convolution operation is as follows:
[0200] Z (2) =σ2(Conv1D2(Z) (1) )+b2)
[0201] Where σ1 and σ2 are activation functions; Conv1D k One-dimensional convolution operation with a kernel length of k; Z (2) It contains deep local features after convolution.
[0202] 3) Residual connection:
[0203] To prevent gradient vanishing or degradation issues during training of deep neural networks, "residual connections" are added between convolutional and temporal models. These connections directly jump from shallow features to deep outputs, enabling the network to learn. The input X is then directly connected across layers to Z. (2) Then, feature fusion is performed to form enhanced features:
[0204] Z res =Z (2) +Proj(X)
[0205] Here, Proj(·) represents the projection transformation, ensuring dimensionality matching; the residual structure ensures that shallow information is not lost in deep networks, which is beneficial for stable training and preservation of multimodal features; it also provides dual-path features: Z (2) For nonlinear enhancement, X provides compensation for the original structure.
[0206] 4) LSTM timing modeling module:
[0207] The residual-enhanced tensor Z res The input is fed into stacked two-layer LSTM units to model the dynamic evolution and state transition patterns between feature sequences;
[0208] First layer LSTM:
[0209]
[0210] Second-layer LSTM:
[0211]
[0212] Finally, the global state representation is obtained:
[0213]
[0214] 5) Classification and confidence output mechanism:
[0215] Classification prediction:
[0216] The LSTM output is mapped to the leakage current type space through a fully connected layer:
[0217] z i =W i T ·hout +b i i = 1, ..., C
[0218] Use Softmax to calculate the probability distribution for each type:
[0219]
[0220] Classification results:
[0221]
[0222] Confidence output:
[0223] definition:
[0224]
[0225] Specifically, when γ≥0.85, the system has a high degree of confidence and will act automatically; when 0.6≤γ<0.85, the system has a moderate degree of confidence and will trigger an early warning; when γ<0.6, the system makes a fuzzy judgment and will recommend that the system use an expert system to assist in decision-making.
[0226] By introducing a graph-enhanced gradient boosting tree model (G-GTBoost) and an improved residual convolutional-temporal neural network model (Res-CNN-LSTM), combined with multimodal feature fusion and dynamic threshold adjustment mechanisms, significant performance improvements were achieved in leakage current risk identification and classification. In actual test data from the State Grid Electric Power Research Institute, the leakage current type identification accuracy reached 98.7% (a 26.9% improvement over the traditional SVM method), with an average confidence level of 0.93. Furthermore, a re-triggering rate of only 6.5% is required to ensure reliability under extreme operating conditions. This solves the problems of high false alarm rate and insufficient dynamic response in traditional methods, providing a precise and efficient leakage current detection solution for the safe operation and maintenance of low-voltage distribution transformer areas.
[0227] In typical transformer substations, the risk scoring accuracy is improved to 96.8%, which can effectively identify leakage hazards under fuzzy boundary conditions. The dynamic adjustment mechanism makes the judgment more consistent with the actual electrical state of the current transformer substation, reducing FPR to 6.7% and FNR to 4.5%. The average identification response time is shortened to 1.8 seconds, supporting online deployment and rapid response. Through comparative analysis of indicators, it can be seen that the dynamic leakage risk assessment mechanism in this invention has significant advantages over the traditional fixed threshold judgment method.
[0228] Performance indicators SVM GTBoost This patented method (Res-CNN-LSTM) Type recognition accuracy 71.8% 84.3% 98.7% Average confidence level 0.72 0.83 0.93 Re-trigger rate 21.0% 12.0% 6.5%
[0229]
[0230] A machine learning-based intelligent identification system for leakage current risks in low-voltage distribution areas is provided, which uses the aforementioned machine learning-based intelligent identification method for leakage current risks in low-voltage distribution areas to identify leakage current risks.
[0231] A computer-readable storage medium, wherein computer program instructions are stored on the computer-readable storage medium, the computer instructions causing the computer to perform the method described above.
[0232] The above technical solutions only embody the preferred technical solutions of the present invention. Any modifications that may be made by those skilled in the art to certain parts thereof embody the principles of the present invention and fall within the protection scope of the present invention.
Claims
1. A machine learning-based intelligent identification method for leakage current risk in low-voltage distribution areas, characterized in that, Includes the following steps: Step 1: Construction of multidimensional leakage current characteristics and aggregation and analysis of transformer area information; Step 2: Dynamic assessment of leakage risk and adaptive adjustment of threshold; Step 3: Intelligent identification and confidence level determination mechanism for leakage current type.
2. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 1, characterized in that, The construction of multidimensional leakage current features and the aggregation and analysis of transformer area information include the following steps: Step 1) Data acquisition; Deploy multi-point sensing devices to synchronously collect the residual current i of each branch in the transformer area. r (t), Three-phase voltage u abc (t), Load current i abc (t), temperature and humidity (T, H); enable time synchronization mechanism to ensure data is alignable and time-stamped; Step 2) Data preprocessing; The waveform signal is filtered to remove spikes and ambient noise; data standardization is performed using the Z-score standardization method. Step 3) Feature extraction to construct feature vector F1; Representative feature vectors are constructed from the original acquired signals. A feature system is built through three dimensions: time domain features, frequency domain features, and topological and environmental features, and then fused to form the feature vectors.
3. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 1, characterized in that, The dynamic assessment of leakage risk and adaptive adjustment of thresholds include the following steps: Step 1): Graph structure modeling and input representation; Step 2): Diagram of message transmission mechanism and risk scoring process; Step 3): Construction of the G-GTBoost risk scoring model.
4. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 3, characterized in that, The graph structure modeling and input representation includes the following steps: Step 1: Define the graph structure; The entire transformer area is modeled as a directed graph G = (V, E), where V = {v1, v2, ..., v...}. n } represents each monitoring node in the distribution area, such as transformer outgoing lines, branch lines, and user meter boxes; Indicates electrical connection relationships; each edge e ij =(v i ,v j ) indicates the direction of electrical influence from node i to j; Step 2: Node feature representation; Each node v i The collected multidimensional features are represented as vectors: x i =[PF i ,THD i ,I rms,i ,ΔI i ,T i ,H i ,L i ,ρ i ,...]∈R d It includes residual current characteristics within the time window, frequency domain disturbance characteristics, environmental parameters, and location factors; All nodes constitute the input matrix: X=[x1,x2,...,x n ] T ∈R n×d 。 5. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 4, characterized in that, The graph message transmission mechanism and risk scoring process include the following steps: Step (1) Define the adjacency matrix; Let A∈R n×n Let A be the adjacency matrix of the graph, representing the topological connections between nodes. ij =1 indicates that v exists. j →v i The information flow supports weighted edges, and voltage level difference and cable length can be added as edge weights; Step (2) Feature enhancement aggregation; Perform a GCN (Graph Convolutional Network) aggregation operation once for the input features of each node: Where N(i) represents the set of neighbors of node i; σ represents a nonlinear activation function, such as ReLU; W1∈R d×d' The weight matrix represents the graph neural network; This represents the feature representation after aggregating neighbor information.
6. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 5, characterized in that, The construction of the G-GTBoost risk scoring model includes the following steps: Step 1: Input the basic features; For any node v in the transformer area graph G=(V,E) i For each ∈V, its enhanced feature vectors are obtained through graph convolution: in, It is a feature representation that integrates local and neighbor information, and it is the input basis of GBDT; Step 2: Gradient boosting tree regression mechanism; Gradient boosting trees are an ensemble model that is trained iteratively, with one decision tree f trained in each round. t This is used to approximate the residual loss from the previous time, which is the difference between the actual risk value and the current model prediction. The output of the model in round t is: in, f is the risk score prediction for node i in the t-th round of the model; η is the learning rate; f t The output of the t-th regression tree; initial prediction That is, the average risk value; Step 3: Definition of loss function and derivation of residuals; Let the real label be This represents the risk score calculated empirically, and the loss function used is the squared error loss: True risk value Configuration: If a tripping / grounding event occurs in the transformer area during a certain period, the value is assigned to 1.0; if an equipment malfunction occurs but the circuit does not trip, the value is set to 0.6 to 0.8; in normal conditions, the value is assigned to 0.0; when direct labels are lacking, they can be calculated through expert rules. The target residual to be fitted in round t is: This is also known as "the error between the predicted value and the actual value of the previous model," and this error is determined by the new tree f. t Go and learn: Each subtree f t The learning objective is to approximate the residuals as closely as possible, thereby gradually approaching the true risk score. Step 4: Overall model output and interpretability; The final output of the model in round T is: Among them, f t (·) represents the regression function formed by several splitting conditions and leaf nodes in the tree model; ω t The weights are calculated as a whole; the output of each leaf node represents the risk increment for that category.
7. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 1, characterized in that, The intelligent identification and confidence level determination mechanism for leakage current type includes the following steps: Step 1: Model structure design; Step 2: Convolutional Feature Extraction Module; Step 3: Residual connection; Step 4: LSTM timing modeling module; Step 5: Classification and confidence output mechanism.
8. The intelligent identification method for low-voltage distribution area leakage risk based on machine learning according to claim 7, characterized in that, The model structure design includes feature input and encoding, padding the temporal feature vectors into a fixed-dimensional tensor: X∈R T×N Where T is the number of time steps; N is the feature dimension of each step; X = [x1,...,x] T ] T Each line represents a time-time characteristic; The convolutional feature extraction module uses a two-level one-dimensional convolutional structure to extract local spatial features of the signal; The first-level convolution operation is as follows: Z (1) =σ1(Conv1D1(X)+b1) The second-level convolution operation is as follows: Z (2) =σ2(Conv1D2(Z (1) )+b2) Where σ1 and σ2 are activation functions; Conv1D k One-dimensional convolution operation with a kernel length of k; Z (2) Includes deep local features after convolution; The residual connection directly connects the input X to Z across layers. (2) Then, feature fusion is performed to form enhanced features: From res =Z (2) +Proj(X) Here, Proj(·) represents the projection transformation, ensuring dimensionality matching; the residual structure ensures that shallow information is not lost in deep networks, which is beneficial for stable training and preservation of multimodal features; it also provides dual-path features: Z (2) For nonlinear enhancement, X provides compensation for the original structure; The LSTM temporal modeling module includes the residual-enhanced tensor Z. res The input is fed into stacked two-layer LSTM units to model the dynamic evolution and state transition patterns between feature sequences: First layer LSTM: Second-layer LSTM: Finally, the global state representation is obtained: The classification and confidence output mechanism includes: Classification prediction: Mapping the LSTM output to the leakage current type space through a fully connected layer: z i =W i T ·h out +b i ,i=1,...,C Use Softmax to calculate the probability distribution for each type: Classification results: Confidence output: Definition: Specifically, when γ≥0.85, the system has a high degree of confidence and will act automatically; when 0.6≤γ<0.85, the system has a moderate degree of confidence and will trigger an early warning; when γ<0.6, the system makes a fuzzy judgment and will recommend that the system use an expert system to assist in decision-making.
9. A machine learning-based intelligent identification system for leakage current risk in low-voltage distribution areas, wherein the machine learning-based intelligent identification method for leakage current risk in low-voltage distribution areas described in any one of claims 1-8 is used to identify leakage current risks.
10. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer program instructions that cause the computer to perform the method as described in any one of claims 1 to 8.
Citation Information
Cited By
Residual current-based street lamp electric leakage risk assessment method and device, and medium
CN121412731A
Method, device and medium for residual current-based street light leakage risk assessment
CN121412731B
Tundish erosion prediction method based on hierarchical hybrid expert framework
CN121561653A