Methods and systems for condition monitoring and fault early warning of distribution transformers, high and low voltage switchgear, and ring main units
By expanding the sample library through fault mechanism modeling and digital twin technology, and combining meta-learning and domain adaptive transfer learning, a small-sample enhancement and cross-domain transfer learning system is constructed. This solves the problems of scarce fault samples and heterogeneous differences in equipment in the power distribution equipment monitoring system, realizes real-time and accurate fault early warning of power distribution equipment, and improves the stability and reliability of the power supply system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA JIAOTONG UNIVERSITY
- Filing Date
- 2026-04-20
- Publication Date
- 2026-05-26
AI Technical Summary
The existing power distribution equipment monitoring system suffers from a lack of fault samples, heterogeneous equipment, and insufficient ability to identify multi-device cascading and complex faults. This leads to false alarms, missed alarms, and delayed early warnings. Furthermore, the system suffers from high risks of data transmission redundancy and privacy leaks, making it difficult to achieve real-time and accurate fault early warnings.
We expand the sample library by adopting fault mechanism modeling and digital twin technology, combine meta-learning and domain adaptive transfer learning to build a small sample enhancement and cross-domain transfer learning system, integrate static and dynamic parameters of equipment to build an adaptive dynamic threshold model, optimize data interaction through edge-cloud collaborative computing architecture, and combine federated learning and multimodal fusion algorithms to build a unified fusion framework for heterogeneous data.
It effectively solves the problem of insufficient identification accuracy caused by heterogeneous differences in equipment, reduces the risk of data transmission redundancy and privacy leakage, improves the ability to analyze the evolution law of complex faults and the timeliness of early warning of cascading faults, ensures unified status management of old and niche equipment and conventional power distribution equipment, reduces false alarms and missed alarms, and improves the operational stability and reliability of the power supply system.
Smart Images

Figure CN122092501A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of condition monitoring and fault early warning methods, and particularly relates to a condition monitoring and fault early warning method and system for distribution network transformers, high and low voltage switchgear, and ring main units. Background Technology
[0002] In power distribution networks, transformers, high and low voltage switchgear, and ring main units are core power distribution equipment. Their operational stability directly determines the safety and reliability of power supply. Traditional condition monitoring and fault early warning methods rely heavily on sufficient labeled samples and homogeneous equipment models. However, in actual field operations, there are common problems such as scarce fault samples, significant heterogeneous differences between old and niche equipment and mainstream mass-produced equipment, and difficulty in accurately analyzing the evolution patterns of multi-equipment cascading composite faults. Relying solely on a single fault identification model cannot meet the needs of small-sample learning and cross-equipment knowledge transfer. It also has weak ability to identify the temporal correlation of composite faults and the coexistence of multiple stages of states, which can easily lead to false alarms, missed alarms, and delayed early warnings.
[0003] Existing power distribution equipment monitoring systems mostly adopt centralized computing architectures. The large number of parameters in deep learning models makes them difficult to adapt to edge terminal deployments. Uploading all data results in transmission redundancy and poses a risk of data privacy leakage. At the same time, there is a lack of a unified fusion framework for multi-source heterogeneous data such as PRPD maps, dissolved gas in oil, and temperature resistance time series. Fault warning thresholds are mostly set in a fixed manner, which cannot be dynamically adapted to equipment operating conditions, service life, and environmental conditions. Edge-cloud collaboration, federated learning, adaptive threshold optimization, and other technologies have not been effectively integrated, making it difficult to achieve real-time, accurate, and adaptive fault warnings and full life cycle status management for various types of equipment in the power distribution network. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for condition monitoring and fault early warning of distribution network transformer-high and low voltage switchgear-ring mains cabinet, the method comprising: Based on fault mechanism modeling and digital twin technology, the sample library is expanded, and meta-learning and domain adaptive transfer learning are combined to construct a small sample enhancement and cross-domain transfer learning system. The evolution path of complex faults was analyzed and a fault chain state transition matrix was constructed. A fault chain temporal correlation model was established using a deep learning model with temporal attention mechanism and a multi-label classification algorithm. By integrating static and dynamic parameters of the equipment to construct a full-dimensional profile of the equipment, a dynamic threshold model is built based on a weighted regression algorithm, and a reinforcement learning mechanism is introduced to iteratively optimize the threshold calculation factor to construct an adaptive dynamic threshold adjustment algorithm. We build an edge and cloud collaborative computing architecture by modifying algorithms to be lightweight, optimize the data interaction mechanism, and combine federated learning and multimodal fusion algorithms to construct a unified fusion framework for heterogeneous data. A simulation test platform was built to conduct offline verification, typical lines were selected for pilot applications and online data collection and analysis, a closed loop of operation and maintenance feedback was established, and a full-process verification and iterative optimization mechanism was constructed.
[0005] Furthermore, embodiments of the present invention also provide a condition monitoring and fault early warning system for distribution network transformers, high and low voltage switchgear, and ring main units, comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the above-described method for monitoring and warning of faults in distribution transformers-high and low voltage switchgear-ring mains cabinets by executing the machine-executable instructions.
[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, the processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-mentioned method for monitoring and warning of faults in distribution network transformers-high and low voltage switchgear-ring network cabinets.
[0007] Based on the above, by unifying and dynamically adapting multi-source heterogeneous data, the problem of insufficient identification accuracy caused by the scarcity of fault samples and the heterogeneity of equipment in traditional condition monitoring methods can be effectively solved. By combining edge-cloud collaborative architecture and lightweight model design, the risk of data transmission redundancy and privacy leakage can be reduced, and unified condition management of old and niche equipment and conventional power distribution equipment can be achieved. This significantly improves the ability to analyze the evolution law of complex faults and the timeliness of early warning of cascading faults.
[0008] By relying on adaptive threshold dynamic adjustment and federated learning distributed training mode, the lag and limitations of traditional fixed threshold are eliminated. It maintains stable fault identification and early warning performance in scenarios of small sample learning and cross-device knowledge transfer, continuously optimizes the status management capability of distribution network equipment throughout its entire life cycle, reduces false alarms and missed alarms, improves the stability and reliability of power supply system operation, and provides efficient and feasible technical support for real-time and accurate fault early warning of distribution network equipment. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of the execution flow of the status monitoring and fault early warning method for distribution network transformers, high and low voltage switchgear, and ring main units provided in this embodiment of the invention.
[0010] Figure 2 This is a schematic diagram of exemplary hardware and software components of the distribution network transformer-high and low voltage switchgear-ring network cabinet status monitoring and fault early warning system provided in an embodiment of the present invention. Detailed Implementation
[0011] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a method for monitoring and warning of faults in a distribution network transformer-high and low voltage switchgear-ring mains unit according to an embodiment of the present invention. The following is a detailed description of this method.
[0012] Step S110: Expand the sample library based on fault mechanism modeling and digital twin technology, and construct a small sample enhancement and cross-domain transfer learning system by combining meta-learning and domain adaptive transfer learning; Based on the fault mechanism of distribution network equipment and digital twin technology, the sample library is expanded to construct a small-sample enhancement and cross-domain transfer learning system. First, theories from multiple disciplines such as electrical engineering, thermodynamics, and chemistry are integrated to analyze the physical evolution of fault occurrence and development, clarifying the correlation mechanism of characteristic parameters such as PRPD spectrum phase distribution, dissolved gas component ratio in oil, and temperature resistance time-series changes, and establishing a mathematical model of the fault mechanism. Second, based on equipment design drawings, material parameters, and historical operating data, a high-fidelity digital twin is built using multiphysics coupling simulation tools such as ANSYS. Fault excitation factors such as voltage fluctuations, mechanical wear, and environmental corrosion are injected to simulate the equipment operating state under multiple fault types, multiple evolution stages, and multiple operating condition combinations, generating highly realistic synthetic samples. Subsequently, meta-learning technology is used to process small-sample scenarios, and the MAML framework empowers the model to quickly adapt to new equipment / new scenarios; a cross-domain transfer module is constructed by integrating domain adaptive algorithms to minimize the feature distribution differences between heterogeneous equipment in the source and target domains. Finally, a structured sample library is constructed by integrating synthetic samples with real samples, forming a complete technical system for sample expansion, small sample adaptation, and cross-domain transfer, ensuring the model's fault identification performance in scenarios with scarce samples and heterogeneous equipment.
[0013] Step S111: Analyze the physical evolution law of distribution network equipment faults, clarify the change mechanism of characteristic parameters including PRPD spectrum characteristics, proportion of dissolved gas components in oil, and temperature resistance time series changes, construct a mathematical model of fault mechanism by combining multidisciplinary theories, and quantify the mapping relationship between fault severity and characteristic parameters; This study analyzes the physical evolution of faults in distribution network equipment and constructs a mathematical model of the fault mechanism to quantify the mapping relationship between fault severity and characteristic parameters. For core equipment such as transformers, high and low voltage switchgear, and ring main units, the evolution paths of different fault modes are systematically analyzed: In insulation degradation faults, the discharge phase distribution of the PRPD spectrum gradually diffuses to all phases, and the average discharge quantity increases exponentially with the deepening of degradation; under local overheating faults, the proportion of methane and ethylene components in the dissolved gases in the oil increases significantly, and the gas production rate is positively correlated with temperature; poor contact faults cause abrupt changes in resistance time-series data, and the slope of the abrupt change is negatively correlated with contact pressure. Combining theories from multiple disciplines such as electrical engineering, materials science, and chemical kinetics, a multivariate mathematical model is constructed, with fault severity (minor, moderate, severe, fatal) as the dependent variable and PRPD spectrum characteristic parameters, gas component ratios, and temperature-resistance time-series characteristics as independent variables. The model parameters are fitted using the least squares method.
[0014] Step S112: Based on equipment design, material and operation data, a high-fidelity digital twin is built using multi-physics coupling simulation technology, fault excitation factors are injected, the equipment operation status under multiple fault types, multiple stages and multiple working conditions is simulated, and multiple types of high-fidelity synthetic samples are generated. Based on equipment design parameters, material properties, and historical operating data, a high-fidelity digital twin is built using multiphysics coupling simulation technology to generate various types of highly realistic synthetic samples. First, basic parameters such as the equipment's 3D design drawings, winding material thermal conductivity, and contact resistance are collected. Combined with historical load curves and environmental temperature and humidity data, an electromagnetic, temperature, and fluid multiphysics coupling model is constructed in the COMSOL Multiphysics simulation platform to recreate the equipment's actual operating state. Second, fault excitation factors are injected: for insulation degradation faults, a gradient decay of the dielectric constant of the insulation medium is set; for SF6 leakage faults, a slow decrease in gas pressure is simulated; and for poor contact faults, the contact resistance is gradually increased. The simulation covers multiple operating conditions including 20%-100% load rate, -40℃-60℃ ambient temperature, and 5%-95%RH ambient humidity, generating various types of samples such as PRPD time series, dissolved gas component data in oil, and temperature / resistance time series curves. Each fault type covers four stages: initial, development, deterioration, and outbreak, with no fewer than 500 composite samples generated for each fault type to ensure the phased nature and diversity of the samples.
[0015] Step S113: Use statistical testing methods to verify the consistency of feature distribution between the synthesized sample and the real sample, optimize the parameters of the digital twin model to ensure the credibility of the sample, classify and label the sample according to the three dimensions of equipment type, failure mode, and operating condition parameters, and integrate various types of samples to build a structured sample library with full state coverage; Statistical testing methods were employed to verify the consistency of feature distributions between the synthetic and real samples, optimize the parameters of the digital twin model, and construct a structured sample library. First, the Kolmogorov-Smirnov test was used to verify the consistency of distributions of numerical features (such as gas component ratios and mean temperature), while the chi-square test was used to verify the distribution differences of discrete features such as the phase distribution of the PRPD spectrum. The significance level was set at 0.05, and a p-value > 0.05 was considered consistent. If distribution deviations existed, the boundary conditions of the digital twin model (such as heat dissipation coefficient and dielectric loss factor) or the intensity of fault excitation factors were adjusted until more than 95% of the feature dimensions met the distribution consistency requirements. Second, the samples were classified and labeled according to the three dimensions of "equipment type - fault mode - operating parameters": equipment type was labeled as transformer, high and low voltage switchgear, and ring main unit; fault mode was labeled as insulation degradation, poor contact, etc.; and operating parameters were labeled with specific load rates and ambient temperature and humidity ranges. Finally, real samples (including historical fault data and field test data) and optimized synthetic samples are integrated, and a structured sample library is built using a MySQL database, which supports fast retrieval in three dimensions.
[0016] Step S114: Divide the three-dimensional meta-task set based on the sample library, select a lightweight deep learning network and combine it with the MAML framework to build a meta-learning model, optimize the initialization parameters through an alternating training strategy, and give the model the ability to quickly adapt to the fault features of new devices / new scenarios. Based on a sample library, a three-dimensional meta-task set was divided. A lightweight deep learning network was selected and combined with the MAML framework to construct a meta-learning model. Initialization parameters were optimized through an alternating training strategy. First, feature engineering was performed on the sample library to extract and standardize key features of various data types. The meta-training / validation / test sets were divided in an 8:1:1 ratio, with the test set including samples from new devices / operating conditions that were not used in training. Second, a three-dimensional meta-task set was constructed according to device type, fault mode, and operating condition parameters. Each task was configured with 5-10 support sets per fault and 2-3 times the query set, covering more than 90% of operating scenarios and 10%-15% of edge conditions. Subsequently, a multi-branch feature extractor was built using MobileNetV3, which was fused to generate a 128-dimensional unified feature vector. This extractor was integrated with the MAML framework to construct a meta-learning model. A cross-entropy and contrastive loss fusion function was defined, and inner / outer loop hyperparameters were set and optimized in stages. Finally, an alternating training strategy of inner loop adaptation and outer loop update is adopted to optimize the model initialization parameters. The performance is evaluated using a meta-validation set every 100 meta-batch training, and the learning rate is dynamically adjusted until the model converges (validation set accuracy ≥ 90%), ensuring that the model can quickly adapt to new devices / new scenarios with 5-10 small sample support sets.
[0017] Step S1141: Perform feature engineering on the structured sample library, extract features from PRPD spectrum, dissolved gas in oil, and temperature / resistance time series data, and divide the meta-training / validation / test sets into 8:1:1 after Z-score standardization. The test set includes samples from new equipment / new operating conditions. Feature engineering was performed on the structured sample library to complete data standardization and dataset partitioning. For PRPD spectral samples, 15-20 key features, such as discharge phase distribution, average discharge quantity, and discharge frequency density, were extracted after grayscale conversion, normalization, and phase calibration. For dissolved gas samples in oil, feature ratios such as hydrogen / methane and acetylene / total hydrocarbons were calculated, and a component feature set was formed by combining them with gas production rate. For temperature / resistance time series samples, time-domain features such as mean, variance, and abrupt change slope, as well as frequency-domain features such as dominant frequency and harmonic proportion, were extracted. Z-score standardization was performed on all extracted features, calculated using the formula z=(x−μ) / σ, where μ is the mean of the training set features and σ is the standard deviation of the training set features, eliminating dimensional differences. The meta-training set, meta-validation set, and meta-test set were randomly partitioned in an 8:1:1 ratio. Stratified sampling was used during the partitioning process to ensure that each set contained the same proportion of equipment types, failure modes, and operating parameters. The meta-test set specifically includes samples of new equipment models (such as niche brand ring main units) and new operating conditions (such as high-altitude environments) that were not used in the training, in order to verify the generalization ability of the model and ensure that the test set samples do not overlap with the training set samples.
[0018] Step S1142: Construct a three-dimensional meta-task set according to equipment type, fault mode, and operating condition parameters. Each task is configured with 5-10 support sets per fault and 2-3 times the number of query sets. The task set covers more than 90% of the operating scenarios and includes 10%-15% of edge operating conditions. A meta-task set is constructed based on three dimensions: equipment type, fault mode, and operating condition parameters, ensuring task coverage and the proportion of edge operating conditions. The equipment type dimension includes three core distribution network equipment types: transformers, high and low voltage switchgear, and ring main units. The fault mode dimension covers typical fault types such as insulation degradation, poor contact, SF6 leakage, and related faults. The operating condition parameter dimension covers 20%-100% load rate (divided in 20% intervals), -40℃-60℃ ambient temperature (divided in 20℃ intervals), and 5%-95%RH ambient humidity (divided in 20% intervals), forming a multi-dimensional combination of operating conditions. Each meta-task is configured with samples according to the following rules: the support set randomly selects 5-10 samples for each fault type to ensure that the samples cover different evolution stages of the fault; the query set is 2-3 times the size of the support set, also sampled uniformly according to stages. During the construction of the meta-task set, statistical data on the actual operating conditions of distribution network equipment is used to ensure that over 90% of common operating condition combinations have corresponding meta-tasks. Simultaneously, extreme operating conditions are manually selected to construct edge-condition meta-tasks, with their proportion controlled at 10%-15%, improving the model's adaptability to complex scenarios. All meta-tasks are stored in JSON format, including task ID, three-dimensional dimension labels, and support set / query set sample indexes, facilitating rapid retrieval during model training.
[0019] Step S1143: Select the MobileNetV3 lightweight network to build a multi-branch feature extractor, extract features of different types of fault data respectively, and fuse them into a 128-dimensional unified feature vector; A lightweight MobileNetV3 network was selected to build a multi-branch feature extractor to achieve feature extraction and fusion of different types of fault data. This feature extractor includes three parallel branches: a convolutional branch for two-dimensional data such as PRPD maps, using depthwise separable convolutional layers (3×3 kernel size) of MobileNetV3, extracting spatial features through 4 convolutional layers + 2 pooling layers, and outputting a 64-dimensional feature vector; a lightweight LSTM branch for time-series temperature / resistance data, setting the number of hidden layer units to 64, extracting time-dependent features through 2 layers of bidirectional LSTM, and outputting a 64-dimensional feature vector; and a fully connected branch for numerical data such as dissolved gases in oil, using 2 fully connected layers (128 and 64 neurons respectively), combining BatchNorm layers and the ReLU activation function to extract statistical features, and outputting a 64-dimensional feature vector. In the feature fusion layer, the output feature vectors of the three branches are first concatenated into a 192-dimensional vector. Then, the weights of each branch's features are calculated using an attention mechanism (the initial weights are all 0.33, dynamically adjusted during training). After weighted summation, the vectors are mapped to a unified 128-dimensional feature vector through a fully connected layer. This structure ensures effective feature extraction while reducing computational complexity through lightweight network design, meeting edge deployment requirements, and keeping feature extraction time within 50ms.
[0020] Step S1144: Integrate a lightweight feature extractor with the MAML framework to build a meta-learning model, define cross-entropy and contrastive loss, and set hyperparameters such as inner / outer loop learning rate and meta-batch size; An end-to-end meta-learning model is constructed by integrating a lightweight feature extractor with the MAML framework, defining a fusion loss function and setting hyperparameters. First, the 128-dimensional output of the feature extractor is adapted to the inner loop input interface of the MAML framework, constructing a closed-loop architecture of "feature extraction - inner loop adaptation - outer loop update". All trainable parameters of the feature extractor (convolutional kernels, fully connected layer weights, LSTM hidden layer parameters) are associated with the parameter update interface of the MAML outer loop. Second, the loss function is defined: cross-entropy loss is used to measure fault classification error, and contrastive loss is used to optimize the feature space distribution. The two are fused into a total loss with an initial weight of α=0.7, and the gradient of the contrastive loss is clipped (norm ≤ 5.0). The hyperparameter settings are as follows: initial learning rate of 0.03 for the inner loop, 8 training epochs, and SGD optimizer (momentum 0.9); initial learning rate of 0.003 for the outer loop, meta-batch size of 24, and AdamW optimizer (weight decay factor 1e-5); other hyperparameters include an initial uniform distribution of attention weights in the feature fusion layer, a dropout rate of 0.2 in the fully connected layer, and a total of 10,000 training epochs. At the code level, conditional statements are used to limit the range of hyperparameters, ensuring that the inner loop learning rate is between 0.01 and 0.05, the outer loop learning rate is between 0.001 and 0.005, and the meta-batch size is between 16 and 32.
[0021] Step S11441: Standardize and adapt the interface of the lightweight feature extractor to the MAML framework, build an MAML module with inner and outer loop sub-modules based on the deep learning framework, build a closed-loop architecture of feature extraction, inner loop adaptation and outer loop update through hierarchical direct connection and reverse parameter association, and plan the model data flow. This paper standardizes the interface of a lightweight feature extractor with the MAML framework, constructs a MAML module with inner and outer loop submodules based on the PyTorch framework, and plans the data flow. First, the output interface of the feature extractor is standardized, converting the 128-dimensional feature vector into a tensor format compatible with the MAML framework, ensuring that the data dimensions are consistent with the input requirements of the inner loop. Second, the core MAML module is constructed: the inner loop submodule receives the feature vector and support set labels, updates temporary parameters through the SGD optimizer, iterates 8 times per round, and outputs temporary model parameters adapted to the current task; the outer loop submodule receives the loss value of the query set, updates the global initialization parameters through backpropagation using the AdamW optimizer, and achieves cross-task knowledge transfer. In terms of hierarchical connections, the output layer of the feature extractor and the input layer of the MAML inner loop are directly connected via tensors, and the parameter update interface of the outer loop is back-associated with all trainable parameters of the feature extractor through PyTorch's autograd mechanism. The data flow is planned as follows: original samples → feature extractor → 128-dimensional feature vector → branching to support set / query set → support set feature input inner loop generates temporary parameters → query set features calculate loss through temporary parameters → loss feedback outer loop updates global parameters, ensuring no format loss or dimension mismatch in data transmission.
[0022] Step S11442: Calculate the cross-entropy loss to measure the basic classification error, design the contrastive loss to bring the features of similar samples closer together and push away the features of dissimilar samples, fuse the cross-entropy loss and the contrastive loss according to the preset weights to obtain the total loss, and prune the gradient of the contrastive loss. Cross-entropy loss and contrastive loss are calculated and fused, and the gradient of the contrastive loss is pruned to avoid gradient explosion. During the calculation of cross-entropy loss, the predicted probability distribution of fault categories output by the model is matched with the one-hot encoding of the true labels. The classification error is quantified using a formula to ensure the model has basic fault recognition capabilities. In the design of contrastive loss, pairs of similar samples (different samples of the same fault type) and pairs of dissimilar samples (samples of different fault types) are randomly selected from the support set and query set. The Euclidean distance of sample features is calculated to narrow the feature distance of similar samples and widen the feature distance of dissimilar samples, with a marginal value m set to 0.5. The fusion loss uses a weighted summation method: total loss = α × cross-entropy loss + (1-α) × contrastive loss. α is initially set to 0.7, and can be dynamically reduced to 0.5 in small sample scenarios to balance classification accuracy and feature discriminative power. To avoid instability in the inner loop adaptation due to excessively large gradients in the contrast loss, a gradient pruning technique is adopted, setting an upper limit of 5.0 for the gradient norm. When the gradient norm of the contrast loss exceeds this upper limit, the gradient value is scaled proportionally to ensure that the gradient propagates within a reasonable range and to guarantee the stability of the inner loop parameter updates.
[0023] Step S11443: Select the inner loop hyperparameters, outer loop hyperparameters, and other hyperparameters and complete the initialization. The inner loop hyperparameters include the learning rate, number of training epochs, and optimizer. The outer loop hyperparameters include the learning rate, meta-batch size, and optimizer. The other hyperparameters include the attention weights of the feature fusion layer, the dropout rate of the fully connected layer, and the total number of training epochs. Add boundary constraints at the code level to limit the hyperparameters from exceeding the preset reasonable range. Select and initialize the inner loop hyperparameters, outer loop hyperparameters, and other hyperparameters, adding boundary constraints to ensure their rationality. Inner loop hyperparameters: initial learning rate is set to 0.03, with a range of 0.01-0.05 (step size 0.005), training epochs are set to 8 (range 5-10 epochs), the optimizer is SGD, and the momentum parameter is fixed at 0.9; Outer loop hyperparameters: initial learning rate is set to 0.003 (range 0.001-0.005), initial meta-batch size is set to 24 (range 16-32), the optimizer is AdamW, and the weight decay coefficient is set to 1e-5; Other hyperparameters: the attention weights of the feature fusion layer are initially evenly distributed across branches (0.33 for convolutional branches, 0.33 for LSTM branches, and 0.33 for fully connected branches), the dropout rate of the fully connected layer is set to 0.2, and the total number of training epochs is set to 10,000. At the Python code level, boundary constraints are added through if-else conditional statements: when the learning rate of the inner loop is lower than 0.01, it is forcibly reset to 0.01; when it is higher than 0.05, it is reset to 0.05. When the meta-batch size exceeds the range of 16-32, it is automatically adjusted to the nearest boundary value to ensure that the hyperparameters do not exceed the preset reasonable range during training, thus avoiding model training crashes or performance degradation.
[0024] Step S11444: Hyperparameter tuning is carried out in three stages: pre-tuning, fine-tuning, and final tuning. In the pre-tuning stage, the inner / outer loop learning rate is fixed, different meta-batch sizes are tested, and the lowest value of the validation set loss is selected and locked. In the fine-tuning stage, the meta-batch size is fixed, and the combination of the inner loop learning rate and the outer loop learning rate is searched in a grid and the optimal combination is selected. In the final tuning stage, the learning rate is locked, and the marginal value of the comparison loss and the loss fusion weight are fine-tuned to ensure that the query set accuracy reaches the preset standard in small sample scenarios. The optimal hyperparameters are saved as a configuration file for model reuse. Hyperparameter tuning is conducted in three stages: pre-tuning, fine-tuning, and final tuning, saving the optimal parameters for model reuse. Pre-tuning stage (first 2000 epochs): The inner loop learning rate is fixed at 0.03 and the outer loop learning rate at 0.003. Meta-batch sizes of 16, 24, and 32 are tested respectively. The validation set loss is recorded every 500 epochs, and the meta-batch size corresponding to the lowest loss value is locked (e.g., 24). Fine-tuning stage (2000-8000 epochs): The meta-batch size is fixed at 24. A grid search method is used to traverse 25 combinations of inner loop learning rates (0.01 / 0.02 / 0.03 / 0.04 / 0.05) and outer loop learning rates (0.001 / 0.002 / 0.003 / 0.004 / 0.005). The validation set accuracy is calculated every 500 epochs, and the combination with the highest accuracy is selected (e.g., inner loop 0.03, outer loop 0.003). Final optimization phase (8000-10000 rounds): Lock the core learning rate, fine-tune the marginal values of the contrastive loss (0.4 / 0.5 / 0.6) and the fusion weight α (0.6 / 0.7 / 0.8) to ensure query set accuracy ≥88% in small sample scenarios. Finally, save the optimal hyperparameters (including learning rate, meta-batch size, marginal values, α, etc.) as a JSON configuration file. The file contains information such as hyperparameter names, optimal values, and value ranges, which can be directly read and loaded during model training without repeated optimization.
[0025] Step S11445: Compile the model based on the deep learning framework, specify the custom MAML meta-optimizer and fusion loss function, and set accuracy, precision, and recall as evaluation metrics. Randomly select meta-task samples to input into the untrained model, verify the output results of each module, select a single meta-task to complete one round of inner / outer loop training, track the gradient values of each layer of the feature extractor to verify that the parameters can be effectively updated; use the initialized model to predict the meta-test set, and record the initial accuracy as the training baseline. The model was compiled using the PyTorch framework, and initialization verification was performed to ensure normal module functionality. First, model compilation parameters were configured in the PyTorch environment, specifying a custom MAML meta-optimizer (integrating inner / outer loop optimization logic), a cross-entropy + contrastive fusion loss function, and setting evaluation metrics as accuracy, precision, and recall. Precision and recall were used to evaluate the reliability of fault identification. Second, samples from 10 randomly selected meta-tasks (covering three types of equipment, five types of faults, and ten operating conditions) were input into the untrained model to verify the output of each module: the feature extractor output must be a 128-dimensional tensor, the loss value must be within the range of 0-5, and the updated temporary parameters of the inner loop must differ from the initial parameters. Then, a typical meta-task (such as transformer insulation degradation fault, 20% load rate + 25℃ environment) was selected to complete one round of inner / outer loop training. The gradient values of each layer of the feature extractor were tracked using TensorBoard, ensuring that the gradients were non-zero and ≥1e-6, verifying that the parameters could be updated effectively. Finally, use the initialized model to predict the meta-test set and record the initial accuracy (≥50%) as the training baseline. If the accuracy is not met, check the module connection logic and hyperparameter initialization values until the baseline meets the requirements.
[0026] Step S11446: Set a learning rate decay rule. When the accuracy of the validation set does not improve for a preset number of consecutive meta-batches, reduce the learning rate of the inner / outer loop. If there is still no improvement, trigger early stopping. Design an adaptive rule for the meta-batch size to adjust the meta-batch size according to the GPU memory usage. Design a dynamic adjustment rule for the loss weight to adjust the loss fusion weight according to the false positive rate and false negative rate of small sample adaptation to balance feature discrimination and classification accuracy.
[0027] A hyperparameter dynamic adaptation mechanism is designed to balance model training stability and performance. The learning rate decay mechanism monitors the accuracy of the meta-validation set. If the accuracy of five consecutive meta-batches shows no improvement (improvement < 0.5%), the learning rate of the inner / outer loop is reduced by 20%. If there is no improvement for three consecutive meta-batches, an early stopping mechanism is triggered to save the current optimal model parameters. The meta-batch size adaptive mechanism monitors memory usage in real-time via PyTorch's GPU memory monitoring interface. When the usage is > 85%, the meta-batch size is reduced to 16; when the usage is < 40%, it is increased to 32 to fully utilize hardware resources. The loss weight dynamic adjustment mechanism is used during the small sample adaptation phase. By statistically analyzing the model's false positive and false negative rates, if the false positive rate is > 5%, the fusion weight α is reduced to 0.6 to enhance the contrastive loss's optimization of feature discrimination; if the false negative rate is > 5%, α is increased to 0.8 to enhance the cross-entropy loss's guarantee of classification accuracy, ensuring a dynamic balance between feature discrimination and classification accuracy.
[0028] Step S1145: Optimize the initialization parameters using an inner loop adaptation and outer loop update strategy, periodically verify and adjust the learning rate until the model converges and the generalization error is ≤8%; A strategy combining inner-loop adaptation and outer-loop update is employed to optimize model initialization parameters, simultaneously adjusting the learning rate and determining convergence. During implementation, for each meta-task, the inner loop phase uses support set samples to temporarily update model parameters, generating local parameters adapted to the current task. The number of inner loop iterations for each meta-task is executed according to a preset value. In the outer loop phase, the loss value calculated based on query set samples is backpropagated to update the model's global initialization parameters, achieving cross-task knowledge aggregation. The model is validated at fixed meta-batch intervals (e.g., every 100 meta-batches), calculating the generalization error on the meta-validation set. If the error does not meet requirements, the inner / outer loop learning rate is adjusted by a preset percentage (e.g., decreasing by 20% each time). The inner-loop adaptation, outer-loop update, and learning rate adjustment process is repeated, continuously monitoring changes in generalization error until the model's generalization error on the meta-validation set drops below a preset threshold, and the error shows no significant fluctuations across multiple consecutive meta-batches, at which point the model is considered converged.
[0029] Step S1146: Take 5-10 samples from the new device / scene as the support set, perform 1-3 rounds of internal loop fine-tuning, verify the adaptation effect, and ensure that the accuracy reaches more than 85% of the traditional model with 5 samples and converges 10 times faster. Perform small-sample adaptation and performance validation for new devices / scenes. During implementation, randomly select 5-10 samples from the new device / scene samples to form a support set, covering the core fault types of the scenario. Input the support set into the trained meta-learning model and perform 1-3 rounds of inner loop fine-tuning, updating only the local adaptation parameters of the model without changing the global initialization parameters. After fine-tuning, select query set samples for the corresponding scenario and input them into the model for inference to verify the adaptation effect. Adaptation effect validation is performed according to preset standards, focusing on evaluating the model's classification accuracy and convergence speed under small-sample conditions. Ensure that the model's classification performance meets the preset reference standard with only 5 support set samples, and that the number of iterations required for convergence meets the set requirements. If the standard is not met, increase the number of support set samples or extend the inner loop fine-tuning rounds, and re-execute the adaptation and validation process.
[0030] Step S1147: After pruning and quantization, the model volume is compressed to 20%-30%, and a dynamic update and performance monitoring mechanism for meta-tasks is established. When the adaptation accuracy is lower than 80%, model retraining is triggered.
[0031] The model undergoes lightweighting, and a dynamic update and performance monitoring mechanism for meta-tasks is established. During implementation, structured pruning techniques are used to remove redundant convolutional kernels and fully connected layer neurons, compressing the model structure according to a preset ratio. Model parameters are quantized, converting floating-point parameters to low-precision integer parameters, compressing the model size to 20%-30% of its original size while maintaining model functionality. A dynamic update mechanism for meta-tasks is established; when the number of new samples reaches a set threshold or a new operating scenario emerges, the meta-task set is re-divided and added to the training library. A model performance monitoring module is built to periodically collect the model's adaptation accuracy data in practical applications. When the adaptation accuracy is detected to be below 80%, the model retraining process is triggered. During retraining, the updated meta-task set is loaded, the original training framework and hyperparameter configuration are used, and the training process is re-executed to optimize model performance.
[0032] Step S115: Identify the heterogeneous devices in the source and target domains, construct a transfer learning model using a domain adaptation algorithm, achieve knowledge transfer by minimizing the difference in feature distribution, and optimize the model's adaptability to heterogeneous devices by combining pseudo-label training with real samples from the target domain and a feature adaptation module. A transfer learning model is constructed to improve heterogeneous device adaptability through domain adaptation, pseudo-label training, and feature adaptation optimization. In implementation, the criteria for classifying heterogeneous devices in the source and target domains are first clarified, and their heterogeneous dimensions are identified and a correlation list is established. A three-layer model architecture, including shared feature extraction, domain classification, and fault classification, is built based on a lightweight feature extractor and integrated with a domain adaptation module. A gradient inversion layer is added to optimize feature learning. An appropriate feature distribution metric is selected, and a fusion loss function is constructed to minimize the difference in feature distribution between the two domains through backpropagation. Pseudo-labels for the target domain are generated using a pre-trained model in the source domain. High-confidence samples are selected and mixed with real samples to construct a training set, which is then iteratively trained with samples from the source domain. A feature adaptation module is added to the model to dynamically adjust feature weights and dimensions. Model training is performed in three stages, with a multi-index evaluation system. Performance is evaluated after each training round, and the model is optimized through parameter adjustment and sample selection until it meets the heterogeneous device adaptability requirements.
[0033] Step S1151: Define the source and target domains, sort out heterogeneous dimensions, and establish a list to clarify the feature directions for migration adaptation; The process involved defining the source and target domains, identifying heterogeneous dimensions, and clarifying the direction of transfer learning features. During implementation, the distribution of sample data from distribution network equipment was analyzed. Source domain equipment was defined based on sample sufficiency and fault type coverage. Target domain equipment was defined based on sample scarcity and differences from the source domain in terms of model, structure, operating environment, and data acquisition. The heterogeneity dimensions between the source and target domains were systematically analyzed, categorized into three dimensions: equipment structure, operating conditions, and data acquisition, clearly defining the specific differences under each dimension. Based on the analysis results, a "Domain-Different Dimension-Feature Impact" list was established, detailing the impact and degree of different heterogeneous dimensions on fault features. By analyzing the list, the key feature directions requiring adaptation during transfer learning were identified, clarifying the adaptation priority and adjustment focus for each feature direction, providing a basis for subsequent model construction.
[0034] Step S1152: Based on the lightweight feature extractor, integrate the domain adaptation module, split the model into three layers: shared feature extraction, domain classification, and fault classification, and add a gradient inversion layer to connect the shared feature extraction layer and the domain classification layer; A transfer learning model architecture with a domain adaptation module is built based on a lightweight feature extractor. During implementation, the network structure and parameters of the pre-built lightweight feature extractor are reused as the shared feature extraction layer of the model, responsible for extracting common fault features from samples in both the source and target domains. The domain adaptation core module is integrated after the shared feature extraction layer, splitting the model into a three-layer structure: a shared feature extraction layer, a domain classification layer, and a fault classification layer. The domain classification layer receives the feature vector output from the shared feature extraction layer and is used to distinguish the domain to which a sample belongs; the fault classification layer receives the same feature output and is responsible for classifying fault type and severity. A gradient inversion layer is added between the shared feature extraction layer and the domain classification layer. This layer inverts the gradient sign of the domain classification layer during gradient backpropagation, forcing the shared feature extraction layer to learn domain-independent common fault features, thus achieving domain adaptive learning. Interface adaptation and parameter association for each layer are completed to ensure smooth data flow in the model.
[0035] Step S1153: Select an appropriate feature distribution measurement method, construct a fault classification and domain adaptation fusion loss function, and optimize the feature distributions of the two domains through backpropagation; The process involves selecting a feature distribution measurement method, constructing a fusion loss function, and optimizing backpropagation. During implementation, fault feature types are first categorized and analyzed to clarify the data attributes and distribution characteristics of each feature type. Suitable distribution measurement methods are then selected for different feature types. Cross-entropy loss is used as the basic fault classification loss, combined with domain adaptation losses for different feature types, and a fusion total loss function is constructed with preset weights. The loss function is then regularized. In the forward propagation phase, features from source and target domain samples are extracted by a shared feature extraction layer and then input into the fault classification layer and domain classification layer respectively to calculate the corresponding losses. The total loss is then summed. In the gradient calculation and propagation phase, the gradient of the total loss with respect to each parameter is calculated based on the chain rule. The gradient of the domain classification layer is reversed by a gradient inversion layer and then backpropagated. In the parameter update phase, an adaptive optimizer is used to update the model parameters, making the feature distributions of the two domains converge. During training, the loss weights are dynamically adjusted. Relevant indicators are monitored after each training round, and the convergence state is determined according to preset conditions. If the classification accuracy significantly decreases, the weights are adjusted and retraining is performed.
[0036] Step S11531: Classify and sort out the fault characteristics of distribution network equipment, clarify the data attributes and distribution characteristics of each type of characteristic, divide the fault characteristics into three categories: numerical, high-dimensional vector, and time series, analyze the distribution pattern of each type of characteristic in the source domain and target domain, and clarify the distribution difference of different types of characteristics between the two domains. The fault characteristics of distribution network equipment were categorized and analyzed to clarify their data attributes, distribution patterns, and differences in distribution between the source and target domains. During implementation, the fault characteristics of distribution network equipment were classified into three categories: numerical, high-dimensional vector, and time-series. Numerical characteristics include gas component ratios and average temperatures, clearly defined as continuous data attributes. High-dimensional vector characteristics include PRPD spectrum extraction features and multimodal fusion features, defining their high-dimensional sparse data attributes. Time-series characteristics include resistance change sequences and partial discharge signal time-series features, determining their time-dependent data attributes. The distribution patterns of each type of feature in the source and target domains were analyzed separately, including whether they conform to specific distribution patterns, the degree of dispersion, and peak positions. A comparative analysis of the distribution differences of each type of feature between the two domains was conducted to clarify the specific manifestations of mean shift in numerical features, spatial distribution dispersion of high-dimensional vector features, and differences in trend consistency of time-series features, providing a basis for the selection of subsequent distribution measurement methods.
[0037] Step S11532: Numerical features use KL divergence and JS divergence to measure distribution similarity; high-dimensional vector features use MMD and Wasserstein distance to measure distribution distance; and time series features use dynamic time warping and time series distribution similarity to comprehensively measure distribution differences. For different types of fault features, appropriate distribution measurement methods are selected. In implementation, for numerical fault features, KL divergence and JS divergence are used as distribution similarity measures. The difference in probability density functions between the source and target domains is calculated to quantify the similarity between the two domains. JS divergence is used to avoid the potential infinity problem of KL divergence. For high-dimensional vector fault features, MMD (Maximum Mean Difference) and Wasserstein distance are used as distribution distance measures. MMD projects high-dimensional features onto a regenerating kernel Hilbert space through kernel function mapping and calculates the mean difference between the two domains. Wasserstein distance, based on optimal transport theory, is suitable for scenarios with less distribution overlap in high-dimensional spaces. For time-series fault features, Dynamic Time Warping (DTW) and Temporal Distribution Similarity (TDS) are used to comprehensively measure distribution differences. DTW calculates the elastic matching distance by aligning key nodes of the time-series features in both domains, while TDS combines the statistics of the time-series features for overall difference assessment, ensuring the accuracy of distribution difference measurement for various features.
[0038] Step S11533: Using cross-entropy loss as the basic fault classification loss, calculate based on the true labels of source domain samples and high-confidence pseudo labels of target domain samples, integrate domain adaptation losses of different feature types using a weighted summation method, set weights according to feature importance, and fuse fault classification loss and domain adaptation loss according to preset weights to obtain the total loss function. The weights of the two types of losses in the total loss function can be dynamically adjusted according to the training effect. A fusion loss function for fault classification and domain adaptation is constructed. In implementation, cross-entropy loss is used as the basic fault classification loss. For source domain samples, the loss is calculated using true labels, and for target domain samples, the loss is calculated using high-confidence pseudo-labels, measuring the basic error of the model's fault identification. Weights are assigned to the domain adaptation losses of different feature types according to feature importance. A weighted summation method is used to integrate the domain adaptation losses of numerical, high-dimensional vector, and temporal features to form the overall domain adaptation loss. The fault classification loss and domain adaptation loss are fused according to preset initial weights to construct the total loss function, expressed as: Total Loss = Fault Classification Loss × λ1 + Domain Adaptation Loss × λ2, where λ1 is the classification loss weight and λ2 is the domain adaptation loss weight. During training, the values of λ1 and λ2 are dynamically adjusted based on changes in model classification performance and domain adaptation effect to ensure a balance between the two, guaranteeing both model classification ability and promoting alignment of feature distributions in the two domains.
[0039] Step S11534: Introduce an L2 regularization term in the domain adaptation loss to limit the complexity of the model parameters, add a smoothing constraint term to the domain adaptation loss of time-series features to reduce noise interference, set a boundary threshold for the loss value, and fix its contribution ratio when the domain adaptation loss is lower than the threshold. Regularization is applied to the fusion loss function to optimize the stability of loss calculation. In implementation, an L2 regularization term is introduced into the domain adaptation loss. By penalizing the trainable parameters of the model, the numerical scale of the parameters is limited, reducing model complexity and avoiding overfitting caused by over-optimization of domain distribution differences. To address potential noise interference in temporal features, a smoothing constraint term is added to the corresponding domain adaptation loss. By smoothing the loss of adjacent time steps of the temporal features, the impact of noise on the distribution measurement results is reduced, making the loss calculation more focused on the core distribution trend. A boundary threshold is set for the domain adaptation loss. When the domain adaptation loss falls below this threshold, its contribution ratio in the total loss is fixed and no longer decreases during training. This prevents the loss of learned fault knowledge in the source domain due to over-domain adaptation, ensuring the model's core classification ability.
[0040] Step S11535: In the forward propagation stage, after the source domain and target domain samples extract features through the shared feature extraction layer, they are respectively input into the fault classification layer to calculate the classification loss, and the input domain classification layer to calculate the domain adaptation loss and summarize the total loss. In the gradient calculation and propagation stage, the gradient of the total loss with respect to each trainable parameter of the model is calculated based on the chain rule. The gradient of the domain classification layer is reversed by the gradient inversion layer and then propagated back to the shared feature extraction layer. In the parameter update stage, the adaptive optimizer is used to update the model parameters according to the gradient direction and magnitude, so that the feature distribution of the two domains gradually converges in the shared feature space. The design employs a backpropagation optimization process to simultaneously minimize classification loss and domain adaptation loss. During implementation, in the forward propagation phase, after the source and target domain samples are input into the model, common fault features are extracted by the shared feature extraction layer. These feature vectors are then input into the fault classification layer and the domain classification layer, respectively. The fault classification layer outputs the classification probability, which is compared with the label to calculate the classification loss. The domain classification layer, combined with the gradient reversal layer, outputs the domain determination result and calculates the domain adaptation loss. The two losses are summed to obtain the total loss. In the gradient calculation and propagation phase, the model layers are traversed using the chain rule to calculate the gradient of the total loss with respect to all trainable parameters of the shared feature extraction layer, the domain classification layer, and the fault classification layer. The gradient of the domain classification layer is reversed by the gradient reversal layer before backpropagation to the shared feature extraction layer, forcing it to learn domain-independent features. In the parameter update phase, an adaptive optimizer such as AdamW is used to adjust the model parameters based on the gradient direction and magnitude. Each iteration simultaneously optimizes both types of losses, gradually converging the source and target domain feature distributions within the shared feature space.
[0041] Step S11536: In the early stage of training, increase the weight of fault classification loss and decrease the weight of domain adaptation loss to ensure that the model can master the source domain fault classification ability. In the middle stage of training, gradually increase the weight of domain adaptation loss and decrease the weight of classification loss to focus on promoting the alignment of feature distributions in the two domains. In the later stage of training, dynamically adjust the weights according to the difference in feature distributions in the two domains to accelerate distribution convergence. The loss weights are dynamically adjusted to optimize the convergence efficiency of feature distributions between the two domains. In implementation, the model training process is divided into three stages: early, middle, and late training. In the early training stage (first 30% of iterations), a higher fault classification loss weight (λ1) and a lower domain adaptation loss weight (λ2) are set to prioritize strengthening the model's learning of source domain fault features and laying the foundation for fault classification. In the middle training stage (30%-70% of iterations), the domain adaptation loss weight (λ2) is gradually increased while the fault classification loss weight (λ1) is decreased, focusing on aligning the feature distributions of the source and target domains and enhancing knowledge transfer. In the late training stage (last 30% of iterations), the difference between the feature distributions of the two domains is calculated in real time. If the difference is close to a preset threshold, the current weight ratio is maintained; if the difference is still large, the domain adaptation loss weight (λ2) is temporarily increased to accelerate distribution convergence. During the weight adjustment process, the weight values at each stage and the corresponding model performance changes are recorded to ensure the rationality of the adjustment strategy.
[0042] Step S11537: After each training round, calculate the distribution metric, total loss, and fault classification accuracy of various features in the two domains, plot the indicator change curve, and set the convergence condition. When the decrease of the distribution metric of the features in the two domains is less than the preset threshold for several consecutive rounds, and the fault classification accuracy is stable in the target range, it is determined that the optimization has converged and the parameter update is stopped. If the classification accuracy decreases significantly during the training process, reduce the domain adaptation loss weight and re-iterate the training.
[0043] Monitor the optimization process and determine the model convergence status. During implementation, after each round of transfer training, calculate the distribution metrics (such as MMD value, KL divergence value), total loss value, and fault classification accuracy of various features in the source and target domains. Plot the metric change curves according to the iteration rounds to intuitively track the optimization progress. Set convergence criteria: when the decrease in the distribution metrics of the two domains is less than a preset threshold (such as 5%) for 10 consecutive rounds, and the fault classification accuracy is stable within the target range, the model optimization is considered to have converged, and parameter updates are stopped. If a significant decrease in fault classification accuracy is found during training (such as a decrease of more than 10%), reduce the domain adaptation loss weight (λ2), backtrack to the parameter state of the previous round, and restart iterative training to avoid classification performance degradation due to excessive domain adaptation and ensure that the model retains good classification ability while aligning the distribution.
[0044] Step S1154: Generate target domain pseudo-labels using the source domain pre-trained model, select high-confidence samples and mix them with real samples to construct a target domain training set, mix them with source domain samples for iterative training and update the pseudo-labels; Generating target domain pseudo-labels and conducting hybrid training. During implementation, the pre-trained model from the source domain is used to predict the fault type and severity of unlabeled samples in the target domain, outputting the predicted probability for each sample and generating initial pseudo-labels. A confidence threshold is set, and high-confidence pseudo-label samples that reach the threshold are selected, while low-confidence samples are removed to reduce noise interference. The selected high-confidence pseudo-label samples are mixed with a small number of real-label samples from the target domain, constructing a target domain training set at a preset ratio (e.g., 8:2). This target domain training set is then mixed with the source domain training set at a preset ratio (e.g., 1:4) and input into the model for iterative training. After each round of training, the optimized model is used to re-predict unlabeled samples in the target domain, updating the pseudo-labels to improve label reliability. Simultaneously, new high-confidence samples are added to the training set to continuously strengthen the model's learning of target domain features.
[0045] Step S1155: Add an attention mechanism feature adaptation module between the shared feature extraction layer and the fault classification layer to dynamically adjust feature weights, fine-tune dimensions, pre-train initial weights and update them synchronously. A device-specific feature adaptation module is added to the model to optimize feature representation. During implementation, this module is inserted between the shared feature extraction layer and the fault classification layer. This module employs an attention mechanism architecture and includes a feature weight allocation unit and a dimension adaptation unit. The feature weight allocation unit adjusts the weights of different fault feature dimensions based on the target domain device type. For example, it increases the weights of SF6 gas density and partial discharge signal-related features for ring main units, and strengthens the weights of dissolved gas and temperature-related features for transformers. The dimension adaptation unit fine-tunes the dimension of the feature vector output from the shared feature extraction layer through a fully connected network to match the dimensional distribution of the target domain device fault features. Based on the feature distribution of real samples in the target domain, the initial weights of the feature adaptation module are pre-trained to initially adapt to the target domain features. During transfer training, the parameters of this module are updated synchronously with the parameters of other layers in the model, continuously optimizing feature representation and making the output features more closely match the fault identification requirements of the target domain device.
[0046] Step S1156: The training is carried out in three stages: source domain pre-training, domain adaptation training, and target domain fine-tuning. The training is advanced according to the feature distribution difference value. During the target domain fine-tuning, some parameters are frozen and related modules are updated. A multi-stage transfer training and iterative optimization approach was implemented. The transfer training was divided into three stages: source domain pre-training, domain adaptation training, and target domain fine-tuning. In the source domain pre-training stage, only source domain samples were used to train the shared feature extraction layer and fault classification layer of the model, without activating the domain classification layer and gradient inversion layer, enabling the model to possess basic fault recognition capabilities. In the domain adaptation training stage, mixed source domain samples, target domain pseudo-label samples, and real samples were input, activating the domain classification layer and gradient inversion layer to minimize the difference in feature distribution between the two domains. The difference in feature distribution between the two domains was calculated every 50 training epochs; when the difference value decreased below a preset threshold, the model entered the next stage. In the target domain fine-tuning stage, only real samples and high-confidence pseudo-label samples from the target domain were used for training. The parameters of the first half of the shared feature extraction layer were frozen, and only the parameters of the feature adaptation module, fault classification layer, and the second half of the shared feature extraction layer were updated, further improving the model's adaptation accuracy in the target domain. After each stage, the model performance was validated using the corresponding dataset to ensure that the stage objectives were achieved.
[0047] Step S1157: Develop an evaluation system that includes recognition accuracy, false alarm rate, etc. Evaluate after each round of training. If the target is not met, adjust the parameters or screen samples. Repeat the process until the performance meets the requirements.
[0048] Establish a transfer learning effect evaluation system and implement model parameter calibration. During implementation, formulate a target domain model adaptation effect evaluation index system, with core indicators including fault identification accuracy, false positive rate, false negative rate, and source domain-target domain feature distribution difference value. After each round of transfer training, evaluate model performance using an independent target domain test set: if the fault identification accuracy is lower than a preset threshold, adjust the domain adaptation loss weights or increase the confidence threshold for pseudo-label screening, re-screen samples, and retrain; if the feature distribution difference value does not meet the standard, adjust the domain adaptation algorithm parameters (such as MMD kernel function type, gradient reversal coefficient), and re-conduct domain adaptation training. Record each evaluation result, parameter adjustment content, and corresponding model performance changes, repeating the "evaluation-adjustment-training" process until the model's fault identification performance in the target domain meets the preset requirements for engineering applications, completing parameter calibration.
[0049] Step S116: Continuously add new real samples and update the digital twin model parameters. Optimize the meta-learning task set and model parameters based on the new samples, establish a transfer learning effect evaluation mechanism, and trigger parameter calibration to form a dynamic closed loop of sample expansion, model training, transfer adaptation, effect evaluation, and parameter optimization.
[0050] A dynamic closed loop is constructed, encompassing sample expansion, model training, transfer learning adaptation, performance evaluation, and parameter optimization. During implementation, newly added real-world fault samples and normal samples generated during the operation of distribution network equipment are continuously collected and added to the sample library. Based on the characteristic distribution of the new samples, the relevant parameters of the digital twin model are adjusted, and synthetic samples are regenerated to enrich the coverage of the sample library. The meta-learning task set is optimized based on the new samples, supplementing meta-tasks corresponding to new equipment and operating conditions, while simultaneously updating the training data of the transfer learning model. A long-term evaluation mechanism for transfer learning effectiveness is established, regularly collecting fault identification data from actual model operation to evaluate the adaptation effect. When the evaluation results show a decline in model performance or the emergence of new heterogeneous equipment adaptation requirements, a parameter calibration process is triggered to readjust the relevant parameters of the meta-learning model and the transfer learning model, iteratively optimizing model performance and forming a continuously iterative, dynamically optimized closed-loop system to ensure that the model always adapts to changes in the operating status of distribution network equipment.
[0051] Step S120: Locate the complex fault evolution path and construct the fault chain state transition matrix. Use a deep learning model with temporal attention mechanism and a multi-label classification algorithm to establish a fault chain temporal correlation model. By analyzing the evolution paths of complex faults and constructing a state transition matrix, a fault chain temporal correlation model is established by combining a temporal attention mechanism and a multi-label classification algorithm. In implementation, typical complex faults of three types of distribution network equipment are systematically analyzed to clarify the chain relationship between primary and derivative faults, decomposed into a fourth-order evolutionary framework, and label critical points. A fault state set including normal states is defined, and an N×N dimensional state transition matrix is constructed and corrected based on historical data and simulation statistics of state transition probabilities. A GRU or bidirectional LSTM is used to build a temporal feature extraction layer, and an additive attention mechanism is designed to capture key temporal information and fuse it to generate a global feature vector. A multi-label output layer is designed, using a loss function that combines binary cross-entropy and Focal Loss, and dynamically setting the classification threshold based on the validation set. The state transition matrix is embedded into the model, and residual connections and gating units are added to form an end-to-end architecture. A preprocessed complex fault temporal dataset is constructed, trained in stages, and hyperparameters are optimized through grid search. Multi-metric evaluation of model performance is conducted, robustness testing is performed, and dynamic adjustments are made to ensure that the model can accurately identify fault states and evolution paths.
[0052] Step S121: Summarize the main fault and derivative fault chain relationships of typical composite faults of three types of equipment and form a list. Deconstruct the composite fault into a four-order evolution framework and clarify the triggering conditions, duration characteristics and transition critical points of each stage. The process involves identifying the chain relationships of complex faults, disassembling the evolutionary framework, and clarifying key parameters. During implementation, the focus is on three core equipment types: transformers, high- and low-voltage switchgear, and ring main units. The combination of typical faults such as insulation degradation, poor contact, and gas leakage is systematically analyzed to clarify the logical chain of "main fault triggering derivative faults," resulting in a structured list of complex faults. This list includes information such as fault combination types and chain triggering sequences. Each type of complex fault is broken down into a four-stage evolutionary framework: "initial stage - development stage - deterioration stage - outbreak stage." Combining fault mechanisms and operational data, the triggering conditions for each stage are clarified, quantified based on equipment operating parameters and characteristic parameter thresholds. The duration characteristics of each stage are defined, and the duration range is statistically analyzed according to fault type. The transition critical points between stages are marked, judged based on sudden changes in characteristic parameters and operating condition change thresholds. This ensures that each link in the evolutionary framework has a clear quantitative or logical basis, providing a foundation for subsequent state definition and model training.
[0053] Step S122: Define the set of fault states including normal states and the characteristic judgment criteria, and construct and correct the N×N-dimensional state transition matrix based on historical data and simulation statistical working conditions adapted to the state transition probability. Define a set of fault states, statistically analyze state transition probabilities, and construct a modified state transition matrix. In implementation, all evolution stages of a complex fault and the normal operating state of the equipment are considered as state nodes, forming a complete set of states. For each state node, corresponding fault characteristic judgment criteria are established, covering indicators such as the numerical range and changing trends of key characteristic parameters. Based on historical equipment fault records and digital twin simulation data, the transition frequency between any two states is statistically analyzed for each state node. The statistics on transition frequencies need to cover typical operating conditions such as different load rates and environmental temperature and humidity, and corresponding transition probability data are generated for each operating condition. An N×N dimensional state transition matrix is constructed (N is the total number of states). Matrix elements are explicitly defined as the probability value of transitioning from one state node to another, and the sum of elements in each row satisfies the normalization requirement. For low-probability transition paths that do not conform to the actual fault evolution logic (such as a direct jump from a normal state to the fault outbreak stage), constraints are set to correct their probabilities to 0, ensuring that the matrix conforms to the actual physical laws of fault evolution.
[0054] Step S123: Construct a temporal feature extraction layer using GRU or bidirectional LSTM, design an additive attention mechanism to calculate time step weights, and fuse them to obtain a global feature vector containing key temporal information; A temporal feature extraction layer is constructed, and a temporal attention mechanism is designed to fuse global feature vectors. In implementation, a GRU or bidirectional LSTM is selected as the base network, and a 3-5 layer stacked temporal feature extraction layer is built. This layer receives multi-dimensional temporal feature data, including PRPD map time series, dissolved gas component time series data in oil, temperature / resistance time series curves, etc. Temporal dependent features are extracted from the data through iterative computation of the base network, outputting the hidden layer feature sequence. An additive attention mechanism is designed, mapping the hidden layer features to query vector Q, key vector K, and value vector V through fully connected layers. A linear transformation with tanh activation is used as the score function to calculate the attention weight of each time step feature. The weight calculation results highlight the proportion of key nodes in fault evolution (such as feature mutation time steps). The attention weights are weighted and summed with the hidden layer features of the corresponding time steps to generate a global feature vector that fuses key temporal information, weakening the interference of irrelevant temporal noise on the model.
[0055] Step S124: Design a multi-label output layer, using binary cross-entropy and FocalLoss, and dynamically set the classification threshold for each state based on the validation set; A multi-label output layer and fusion loss function are designed, and the classification threshold is dynamically set and optimized. In implementation, the number of neurons in the output layer is set according to the total number of fault states, and a Sigmoid activation function is used to achieve non-mutually exclusive multi-state judgment. For each fault state, a binary cross-entropy loss is calculated, and FocalLoss is introduced to correct class imbalance. Weighted fusion is then used to obtain the total loss. Independent validation sets covering all states, all equipment, and all operating conditions are defined, and stratified sampling is performed according to fault states to enhance representativeness. For each state, a preset threshold range is traversed, and the initial threshold is selected based on the Youden index. The initial threshold is applied to the validation set. If the F1 score of a certain state does not meet the standard, the FocalLoss parameter is adjusted, and retraining and threshold selection are performed. This iteration is repeated until all states meet the requirements. A threshold update cycle is set, and the threshold is dynamically updated during training. If the change amplitude is less than a preset value for two consecutive cycles, the threshold is fixed. During inference, the judgment result is verified in conjunction with the fault chain state transition matrix. If the evolutionary logic is violated, the threshold is fine-tuned and the judgment is re-evaluated.
[0056] Step S1241: Based on the total number of fault states, set the number of output layer neurons to correspond one-to-one with the number of states. The output layer receives the feature vector of the preceding module and completes the dimension mapping through the fully connected layer. Sigmoid is selected as the activation function so that the output value of each neuron represents the probability of the existence of the corresponding fault state. It supports non-mutually exclusive independent judgment of multiple fault states and adapts to the multi-stage coexistence features of composite faults. The architecture design of the multi-label output layer was completed. During implementation, the number of neurons in the output layer was determined based on the total number of fault states (including normal states and each fault evolution stage), ensuring a one-to-one correspondence between the number of neurons and the number of fault states. The output layer receives the global feature vector output from the preceding module and maps it to a dimension matching the number of neurons through a fully connected layer, achieving adaptation between the feature dimension and the output dimension. The Sigmoid function was chosen as the activation function for the output layer. This function maps the output value of each neuron to the 0-1 interval, and this output value directly represents the probability of the corresponding fault state. This architecture design supports non-mutually exclusive independent determination of multiple fault states, meaning the model can output the probability of multiple fault states at the same time, perfectly adapting to the characteristics of multi-stage coexistence and continuous evolution of complex faults, ensuring the model can accurately capture the complex state combinations of complex faults.
[0057] Step S1242: For each fault state, calculate the binary cross-entropy loss state by state based on the predicted probability and the true label to reflect the basic error of single-state classification; introduce FocalLoss, adjust the loss weight of samples in different states by balancing coefficient, amplify the loss contribution of hard-to-classify samples and weaken the interference of easy-to-classify samples by focusing parameter, and weight and fuse the binary cross-entropy loss and FocalLoss, and take the average of the fused loss of each state as the total loss. A fusion loss function combining binary cross-entropy and FocalLoss is constructed. During implementation, for each fault state, the existence probability output by the model is compared state-by-state with the true label of that state to calculate the binary cross-entropy (BCE) loss, which directly reflects the classification basis error of a single fault state. To address the imbalance in the number of samples across different fault states, FocalLoss is introduced. A balance coefficient is set to adjust the loss weights of samples from different states, assigning higher weights to fault states with fewer samples. A focusing parameter is used to amplify the loss contribution of hard-to-classify samples (such as samples from the fault outbreak phase) while weakening the loss interference of easily-classified samples (such as samples from normal states). The binary cross-entropy loss and FocalLoss are weighted and fused according to preset weights. The average of the fused losses for all fault states is then used as the total loss for model training, balancing the model's basic classification accuracy with its ability to adapt to imbalanced samples.
[0058] Step S1243: Divide the data into independent validation sets covering all fault states, equipment types and operating conditions, perform the same preprocessing process as the training set, perform stratified sampling according to fault states, increase the proportion of fault states with fewer samples in the validation set, and avoid the distortion of threshold setting caused by sample distribution bias. The validation set is preprocessed and its representativeness is enhanced to provide reliable data support for threshold setting. During implementation, a separate validation set is created, covering all fault states, various equipment types, and different operating conditions to ensure the comprehensiveness of the validation data. The validation set undergoes the same preprocessing procedures as the training set, including time-series data length alignment, missing value imputation, and standardization, ensuring consistency in data format and distribution. A stratified sampling method based on fault state is used to construct the validation set. For fault states with fewer samples, their sampling proportion in the validation set is appropriately increased to improve the representativeness of these less-sampled fault states and avoid distortion in classification threshold setting due to sample distribution bias. After the validation set is constructed, data integrity and consistency are verified, and outlier data is removed to ensure that the validation set objectively reflects the model's performance in different scenarios, providing a reliable basis for subsequent threshold selection.
[0059] Step S1244: For each fault state, traverse the preset threshold range, calculate the sensitivity and specificity under each threshold, and select the threshold corresponding to the maximum value through the Youden index as the initial classification threshold for that state. Take into account the sample distribution and recognition difficulty differences of different fault states, and independently complete the threshold selection for each state. Based on the principle of maximizing the Youden index, the classification threshold for each fault state is dynamically set. During implementation, for each fault state, a threshold traversal interval of 0.05-0.95 is set, and candidate thresholds are selected sequentially in steps of 0.01. For each candidate threshold, the sensitivity (recall) and specificity of that fault state in the validation set are calculated. Sensitivity characterizes the model's coverage ability to identify that state, and specificity characterizes the model's ability to distinguish that state from other states. The Youden index corresponding to each candidate threshold is calculated using the formula "Youden index = sensitivity + specificity - 1". The candidate threshold corresponding to the maximum Youden index is selected as the initial classification threshold for that fault state. Since the sample distribution and recognition difficulty differ among different fault states, the above threshold selection process is executed independently for each fault state to avoid judgment bias caused by using a single threshold to adapt to all states.
[0060] Step S1245: Apply the initial threshold to the validation set, calculate the accuracy, recall and F1 score for each state. If the F1 score of a certain state does not meet the preset requirements, adjust the FocalLoss balance coefficient or focusing parameter corresponding to that state, retrain the model and then filter the threshold again. Repeat the parameter adjustment, model retraining and threshold filtering process until the F1 score of all states meets the requirements. Verify the effectiveness of the thresholds and perform iterative adjustments to ensure that the recognition performance of each state meets the standards. During implementation, the initial classification thresholds for each fault state are applied to the validation set, and the accuracy, recall, and F1 score for each state are calculated. The F1 score serves as the core evaluation metric, comprehensively reflecting the accuracy and recall of the classification. If the F1 score for a fault state does not meet the preset requirements, the FocalLoss balance coefficient or focusing parameter corresponding to that state is adjusted to enhance the loss weight or focusing strength for samples of that state. After adjustment, the model is retrained, and predictions are performed on the validation set again. The classification threshold for that state is then re-selected according to the method in step S1244. This closed-loop process of "parameter adjustment - model retraining - threshold selection" is repeated until the F1 scores for all fault states meet the preset standards, ensuring that the classification threshold for each state is adapted to its sample characteristics and recognition difficulty.
[0061] Step S1246: Set the threshold update cycle. In each cycle, the current model is used to re-predict the validation set, calculate the Youden index for each state and update the threshold. If the threshold change is less than the preset value for two consecutive cycles, the threshold for that state is fixed. If the model overfits, the threshold of the previous cycle is backtracked. The classification threshold is dynamically updated during model training to improve threshold adaptability. In implementation, a threshold update cycle is set, based on the meta-batch or iteration rounds of model training (e.g., every 50 iterations constitute one update cycle). After each update cycle, the model in the current training phase is used to re-predict the validation set, obtaining the latest prediction results for each fault state. Based on the new results, the sensitivity, specificity, and Youden index of each state are recalculated, a new optimal threshold is selected, and the classification threshold for that state is updated. If the threshold change for a fault state is less than a preset value (e.g., 3%) within two consecutive update cycles, it indicates that the threshold has stabilized, and the classification threshold for that state is fixed. If overfitting occurs during model training (e.g., increased validation set loss), the threshold is reverted to the previous update cycle to avoid overfitting the model and causing a decrease in generalization ability.
[0062] Step S1247: During inference, the predicted probability of each state is compared with the corresponding optimal threshold. If the probability is greater than or equal to the threshold, the state is determined to exist and the fault state combination is output. The judgment result is verified by combining the fault chain state transition matrix. If the fault evolution logic is violated, the corresponding state threshold is finely adjusted and the judgment is re-determined.
[0063] During the inference phase, optimal thresholds are applied and the rationality of the results is verified. In implementation, the probability of each fault state output by the model is compared one by one with its corresponding optimal classification threshold. If the probability of a state's existence is greater than or equal to its classification threshold, then the fault state is determined to exist, and all combinations of fault states determined to exist are finally output. The rationality of the determination results is verified using the constructed fault chain state transition matrix. The core of the verification is whether the determination results conform to the fault evolution logic, such as whether there are cases that violate the evolution path, such as "determining the existence of a development stage without an initial stage." If the verification finds that the determination results violate the evolution logic, the classification thresholds of the corresponding contradictory states are fine-tuned (the fine-tuning range is controlled within ±0.05), and the determination process is re-executed to ensure that the final output combinations of fault states conform to the actual fault evolution law.
[0064] Step S125: Embed the state transition matrix, add residual connections and gating units to form an end-to-end model with temporal extraction, attention fusion, matrix embedding, feature interaction, and multi-label output; An end-to-end fault chain temporal correlation architecture is constructed by integrating the state transition matrix with a deep learning model. In implementation, the fault chain state transition matrix is embedded as prior knowledge into the fully connected layer of the model. Matrix multiplication is used to combine the global feature vector fused with attention with the state transition matrix, generating feature vectors containing state transition tendencies, thus strengthening the model's learning of fault stage transition logic. A residual connection is added between the basic temporal feature extraction layer and the attention layer to alleviate the gradient vanishing problem in deep network training and ensure effective transfer of temporal features. A gating unit (such as a sigmoid gate) is added between the attention layer and the multi-label classification layer to control the fusion ratio of state transition features and original temporal features, improving the model's generalization ability. The modules are integrated according to the flow of "original temporal data → basic temporal feature extraction → temporal attention weighted fusion → state transition matrix embedding → cross-layer feature interaction → multi-label classification output" to form an end-to-end fault chain temporal correlation model, ensuring smooth data flow and functional synergy among modules.
[0065] Step S126: Construct and preprocess the composite fault time series dataset, train the model in stages, and optimize the hyperparameters through grid search; A composite fault time-series dataset was constructed and preprocessed, and the model was trained and hyperparameters optimized in stages. During implementation, real composite fault time-series data and digital twin simulation data of distribution network equipment were collected, and the data was organized into a standardized dataset according to the format of "time step-feature dimension-fault state label". The dataset was preprocessed: interpolation was used to fill in missing values in the time-series data, data length alignment was performed, Z-score standardization was used to eliminate dimensional differences, and the training, validation, and test sets were divided in a 7:2:1 ratio. A staged training strategy was adopted: in the first stage, the attention layer parameters were frozen, and only the basic time-series feature extraction layer and multi-label classification layer were trained, enabling the model to grasp the mapping relationship between basic time-series features and fault states; in the second stage, all parameters were unfrozen, and the attention layer and state transition matrix embedding layer were jointly trained to optimize the model's time-series correlation ability. A grid search method was used to optimize the core hyperparameters, including the number of basic network layers, the number of hidden layer units, the number of attention heads, the learning rate, and the number of iterations. The multi-label F1 score on the validation set was used as the core evaluation metric to select the optimal hyperparameter combination.
[0066] Step S127: Evaluate the recognition accuracy and path tracing accuracy using multiple indicators, qualitatively verify the ability to capture transition nodes, conduct robustness tests, and dynamically adjust the model.
[0067] Model validation and robustness testing were conducted, and model performance was dynamically adjusted and optimized. During implementation, a multi-dimensional evaluation index system was established. Quantitative evaluation indicators included multi-label classification accuracy, F1 score, and Hamming loss, used to measure the model's accuracy in identifying various fault states. The accuracy of evolution path tracing was measured by calculating the dynamic time warping (DTW) distance between the predicted state sequence and the actual state sequence. Typical complex fault cases were selected for qualitative validation. The probability change curves and attention weight distribution of each stage output by the model were visualized to verify the model's ability to capture stage transition nodes. Robustness testing was conducted: random noise of varying intensities was added to the test data to simulate sensor data fluctuations; sudden changes in operating conditions (such as a sudden increase in load rate or a sudden change in environmental temperature and humidity) were simulated to evaluate the model's stability under complex scenarios. If the model performance degraded beyond a preset threshold during testing, the attention weight calculation logic or state transition matrix constraints were adjusted, and the model was retrained to improve generalization ability.
[0068] Step S130: Integrate static and dynamic parameters of the device to construct a full-dimensional profile of the device, build a dynamic threshold model based on a weighted regression algorithm, introduce a reinforcement learning mechanism to iteratively optimize the threshold calculation factor, and construct an adaptive dynamic threshold adjustment algorithm; A comprehensive device profile is constructed by integrating static and dynamic parameters. An adaptive dynamic threshold adjustment algorithm is built based on weighted regression and reinforcement learning. During implementation, the static and dynamic parameter systems of the device are first reviewed, parameter standards are unified, and standardization is completed. A five-layer architecture for the comprehensive device profile is constructed, and feature vectors are generated by fusing static and dynamic features. Core input features are determined through feature selection. A multivariate weighted linear regression algorithm is used to train dynamic threshold models for different device types, outputting initial dynamic thresholds. The threshold optimizer is defined as a reinforcement learning agent, constructing a state space, action space, and multi-objective reward function. A simulation training environment is built based on historical data and fault cases, iteratively optimizing threshold calculation factors and forming a factor library. Threshold adjustment trigger conditions are set, and adjustments are executed according to the process of "real-time monitoring - condition judgment - factor matching - threshold calculation - smoothing verification - threshold update." A multi-index evaluation system is established, multi-scenario testing is conducted, and the algorithm is continuously optimized to ensure that the dynamic threshold can adapt to changes in device state.
[0069] Step S131: Organize the static and dynamic parameters of the equipment to form two types of parameter lists, unify the parameter naming, units and coding rules, clarify the sampling frequency and time granularity standards of dynamic parameters, normalize the static parameters, standardize the dynamic parameters using Z-score, fill in the missing parameters and identify and correct outliers using the 3σ principle; The equipment parameter system was streamlined and standardized to provide standardized data for profile construction and model training. During implementation, static parameters were compiled, covering basic attributes (model, manufacturer, years of operation, etc.), structural parameters (number of winding turns, contact material, etc.), and installation parameters (installation location, altitude, etc.), forming a static parameter list with unified parameter naming rules, units of measurement, and coding standards. Dynamic parameters were also compiled, including real-time operating parameters (load rate, voltage / current, etc.), fault characteristic parameters (PRPD spectrum features, gas composition, etc.), and environmental dynamic parameters (real-time temperature, humidity, wind speed, etc.), clarifying the sampling frequency of dynamic parameters and establishing a unified time granularity standard. Static parameters were normalized to eliminate dimensional differences, and dynamic parameters were standardized using Z-scores. Missing parameters were supplemented; static parameters were obtained by consulting manufacturer manuals and maintenance records, while dynamic parameters were supplemented using interpolation. Outliers were identified and corrected using the 3σ principle to ensure all parameter data were standardized, accurate, and free of redundancy.
[0070] Step S132: Construct a profile dimension framework including a basic attribute layer, structural feature layer, operating status layer, fault feature layer, and environment adaptation layer. Use the device's unique identifier as an index to associate standardized static parameters and time-series dynamic parameters to form structured profile data. Use time-series splicing and attention-weighted fusion to fuse multi-source dynamic parameters. Splice the fused dynamic features with static features and generate a full-dimensional feature vector of the device through a fully connected layer to establish a profile update mechanism. A comprehensive device profile is constructed and feature fusion is performed to generate standardized feature vectors. During implementation, a five-layer profile framework is built: "Basic Attribute Layer - Structural Feature Layer - Operating Status Layer - Fault Feature Layer - Environment Adaptation Layer." Using the device's unique identifier as an index, standardized static parameters and time-series dynamic parameters are linked to form structured profile data of "Device ID - Static Attribute - Time-series Dynamic Features." For multi-source dynamic parameters, a combination of time-series concatenation and attention-weighted fusion is used. Core fault feature parameters are assigned higher initial weights based on the importance of fault features, strengthening key dynamic information. The fused dynamic features are concatenated with static features and input into a fully connected layer to map to a unified dimension, generating a comprehensive device feature vector. A profile update mechanism is established: static parameters are updated as needed based on equipment modifications, maintenance changes, etc.; dynamic parameters are updated in real-time according to a set sampling frequency; and the integrity of the profile data is verified daily to ensure that the profile accurately reflects the device status in real time.
[0071] Step S133: Select features by ranking the importance of Pearson correlation coefficient and random forest features, determine the input feature set of the model, select multivariate weighted linear regression as the base model, use historical fault critical feature values as labels, input features and assign initial threshold calculation factors based on importance, output the initial dynamic threshold of each fault feature, and train the regression model according to equipment type. A dynamic threshold model is constructed based on a weighted regression algorithm to output the initial dynamic thresholds for each fault feature. During implementation, Pearson correlation coefficient is used to analyze the linear correlation between features and fault states, and a random forest algorithm is combined to evaluate feature importance. The results of both methods are used to select core features strongly correlated with fault states, eliminate redundant features, and determine the model input feature set. Multivariate weighted linear regression is selected as the basic model, with historical fault critical feature values used as model training labels, and the selected full-dimensional features used as model input. Initial threshold calculation factors are assigned to different input features based on their importance; features with higher importance have larger corresponding calculation factor weights. The dataset is divided according to equipment types such as transformers, high and low voltage switchgear, and ring main units, and regression models are trained independently to adapt the models to the fault characteristic patterns of different equipment. Each model outputs the initial dynamic thresholds for each fault feature corresponding to the equipment.
[0072] Step S134: Define the threshold optimizer as a reinforcement learning agent, take the device operation scenario as the environment, construct a state space that includes the device's full-dimensional feature vector, threshold calculation factor, fault identification index, real-time operating conditions, and action space for threshold calculation factor adjustment operations, design a multi-objective weighted reward function that integrates fault identification recall rate, false alarm rate, and threshold fluctuation amplitude, and build a simulation training environment based on historical device data and fault cases. A reinforcement learning mechanism was designed and a training environment was built to support the optimization of threshold calculation factors. In implementation, the "threshold optimizer" was defined as a reinforcement learning agent, and the equipment operation scenario (including a full-dimensional profile of the equipment, historical fault records, and real-time operation data) was defined as the training environment. A state space was constructed, containing key information such as the current full-dimensional feature vector of the equipment, the current threshold calculation factor, fault identification accuracy / false alarm rate, and real-time operating parameters. An action space was also constructed, containing adjustment operations for the threshold calculation factor, including factor increase / decrease and reset, with limited action granularity to ensure executable operations. A multi-objective weighted reward function was designed, integrating three core indicators: fault identification recall rate, false alarm rate, and threshold fluctuation range. Adjustment strategies with excessive false negative rates were given negative rewards, guiding the agent to learn the optimal adjustment strategy. A simulation training environment was built based on historical equipment operation data and typical fault cases to simulate the equipment operation status and fault evolution process under different operating conditions, generating sufficient training samples.
[0073] Step S135: Use the initial threshold calculation factor of the weighted regression model as the initial action of the agent, set the iteration rounds and exploration rate, load the profile data into the training environment, the agent obtains the state from the environment and selects factors to adjust the action, the environment calculates the reward value after executing the action, updates the action strategy through the reinforcement learning algorithm, iterates until the reward value converges, extracts the optimal threshold calculation factor and saves it according to the equipment working condition type to form a factor library. Iterative optimization training of threshold calculation factors is conducted to form an optimal factor library. During implementation, the initial threshold calculation factors output by the weighted regression model are used as the initial actions of the reinforcement learning agent. The number of iterative training rounds and the exploration rate are set (using an ε-greedy strategy). Full-dimensional device profile data is loaded into the simulation training environment. The agent obtains current state information from the environment (including device features, current factors, and identification indicators), selects an adjustment action for the threshold calculation factors according to a preset strategy, and calculates the fault identification indicator corresponding to the new threshold after the environment executes the action. The reward value is calculated based on the indicator. The Q-value function is updated using the Q-learning algorithm to optimize the agent's action selection strategy, enabling the agent to gradually learn the factor adjustment method that maximizes the reward value. This iterative process is repeated until the reward value converges. The threshold calculation factors corresponding to the optimal action in the converged state are extracted and categorized and saved according to device operating conditions (such as different load rates and ambient temperature and humidity) to form a threshold calculation factor library.
[0074] Step S136: Set threshold adjustment trigger conditions, including sudden changes in equipment operating conditions, excessive false alarms or missed alarms in fault identification, equipment profile updates, and timed triggers. When an adjustment is triggered, extract the current equipment profile features, match the optimal calculation factor in the factor library, input the weighted regression model to calculate the new threshold, and if the difference between the new and old thresholds reaches the preset standard, use exponential moving average to smooth the new threshold. Set the upper and lower limits of the threshold through the equipment's rated parameters or historical fault thresholds, and execute the process of real-time monitoring, condition judgment, factor matching, threshold calculation, smoothing verification, and threshold update. The system integrates an adaptive dynamic threshold adjustment algorithm and executes the adjustment process. During implementation, threshold adjustment trigger conditions are set, including sudden changes in equipment operating conditions (such as load rate changes exceeding preset values), excessive false / missed alarm rates in fault identification, updates to the equipment's full-dimensional profile, and timed triggers (such as once daily). When the trigger conditions are met, the system extracts the full-dimensional profile features of the current equipment and matches the corresponding optimal threshold calculation factor from the factor library based on the real-time operating conditions. The optimal calculation factor is input into a weighted regression model to calculate a new dynamic threshold for fault features. If the difference between the new threshold and the currently used old threshold is ≥5%, the new threshold is smoothed using an exponential moving average method. The smoothing formula is "new threshold = 0.7 × new calculated value + 0.3 × original value" to avoid threshold abrupt changes. Simultaneously, upper and lower limits for the threshold are set using the equipment's rated parameters and historical fault thresholds to prevent the threshold from exceeding a reasonable range. The entire process of "real-time monitoring - condition judgment - factor matching - threshold calculation - smoothing verification - threshold update" is executed to ensure that the dynamic threshold can adapt to changes in equipment status in real time.
[0075] Step S137: Establish an evaluation index system with fault identification accuracy, false alarm rate, false negative rate, and response speed as the core; conduct multi-scenario tests with multiple equipment types, multiple operating conditions, and multiple fault types; compare the identification effects of dynamic thresholds and fixed thresholds; regularly collect algorithm running records and fault identification results; retrain the regression model and reinforcement learning agent; update the factor library; supplement sample data for new operating conditions and new fault types; and optimize the reward function weights and action space.
[0076] Verify algorithm performance and conduct continuous optimization to improve its adaptability. During implementation, establish an evaluation index system, with core indicators including fault identification accuracy, false alarm rate, false negative rate, and adjustment response speed, to comprehensively measure the algorithm's identification effect and real-time performance. Conduct multi-scenario tests on different equipment types (transformers, high and low voltage switchgear, ring main units), different operating conditions (high / low load rate, extreme temperature and humidity), and different fault types (insulation degradation, poor contact, etc.) to compare the fault identification effects of adaptive dynamic thresholds and traditional fixed thresholds. Regularly collect algorithm operation records (including threshold adjustment records and identification results) and fault data, retrain the weighted regression model and reinforcement learning agent based on the new data, and update the threshold calculation factor library. For newly emerging operating conditions and new fault types, supplement sample data, optimize the weights of the reinforcement learning reward function and action space, adjust the feature weights of the weighted regression model, and continuously improve the algorithm's generalization ability and adaptability.
[0077] Step S140: Build an edge and cloud collaborative computing architecture through lightweight algorithm transformation, optimize the data interaction mechanism, and construct a unified fusion framework for heterogeneous data by combining federated learning and multimodal fusion algorithms; By streamlining algorithms, building an edge-cloud collaborative architecture, and optimizing data interaction mechanisms, a unified fusion framework for heterogeneous data is constructed by integrating federated learning and multimodal fusion algorithms. During implementation, the core fault identification algorithm undergoes layered lightweighting processing, including structured pruning, parameter quantization, and knowledge distillation, and is then deployed to the edge and cloud after functional decomposition. Next, the responsibilities of edge and cloud nodes are defined, and edge nodes are deployed according to the distribution network topology, building a collaborative architecture through low-latency links. A layered incremental data transmission strategy is designed, and lightweight protocols and encrypted verification mechanisms are used to optimize interaction. A horizontal federated learning framework is built to protect data privacy, and a three-level multimodal fusion mechanism is established to adapt to heterogeneous data. Finally, all modules are integrated, and standardized adaptation interfaces are designed to be compatible with existing systems, forming an end-to-end framework. The framework performance is verified through multi-scenario testing, and module parameters and rule bases are regularly optimized, with nodes expanded to adapt to new devices and fault modes.
[0078] Step S141: Perform hierarchical lightweight processing on the algorithm, use structured pruning to remove redundant components, quantize the model parameters, retain the core fault identification capability through knowledge distillation, decompose the model into functions, deploy the real-time inference module at the edge, and deploy the complex training and parameter optimization module in the cloud, forming a functional division of edge inference and cloud training. The core algorithms underwent a layered lightweight transformation, and functional decomposition and edge-cloud deployment were completed. During implementation, for core algorithms such as meta-learning, transfer learning, and fault chain temporal correlation, structured pruning techniques were used to remove redundant convolutional kernels and fully connected layer neurons, with the pruning degree dynamically adjusted according to the computing power of edge devices. 32-bit floating-point parameters were quantized to 8-bit integers or 16-bit half-precision, compressing the model size within a preset precision loss range. Knowledge distillation was used to transfer the core recognition capabilities of complex models to lightweight student models, retaining the core functions of fault classification and evolution path recognition. The model was functionally decomposed: low-latency modules such as real-time feature extraction, fault state inference, and threshold judgment were deployed at the edge; computationally intensive modules such as global model training, parameter optimization, and complex fault source analysis were deployed in the cloud, clearly defining the functional division of "edge inference responding to real-time needs, and cloud training supporting global optimization," ensuring that the lightweight model is adapted to edge hardware resources.
[0079] Step S142: Divide the responsibilities of edge nodes and cloud nodes. Edge nodes are responsible for real-time data collection, preprocessing, lightweight model inference, and local abnormal data caching. Cloud nodes are responsible for global model training, parameter aggregation and updating, full data storage, and complex fault tracing and analysis. Deploy edge nodes according to the distribution network topology and connect to the cloud through low-latency communication links to build a distributed collaborative management platform to realize resource scheduling, status monitoring, fault node switching, and load balancing. Establish an edge-cloud collaborative computing architecture, clearly defining node responsibilities and collaboration mechanisms. During implementation, the core responsibilities of each node are defined: edge nodes are responsible for real-time acquisition of multimodal data from equipment, noise reduction and standardization preprocessing, lightweight model inference operations, and local abnormal data caching (avoiding full uploads); cloud nodes are responsible for centralized global model training, parameter aggregation and updates for each edge node, long-term storage of all data, and complex fault tracing analysis across edge nodes. Edge nodes are deployed according to the distribution network topology, with edge gateways prioritized in densely populated areas such as substations and switchgear clusters. Low-latency communication links are established using 5G industrial modules or industrial Ethernet to ensure that data transmission latency between the edge and cloud meets preset requirements. A distributed collaborative management platform is built, integrating resource scheduling, node status monitoring, fault switching, and load balancing modules. This platform monitors the computing power utilization of edge nodes in real time, automatically offloading tasks to overloaded nodes and quickly switching faulty nodes to backup nodes, ensuring stable architecture operation.
[0080] Step S143: Design a layered data transmission strategy, where edge nodes only upload critical information, raw data is cached locally, incremental transmission is used to reduce duplicate transmissions, the data transmission format is optimized, a communication channel is built based on a lightweight communication protocol and an adaptive transmission frequency is set; establish a data transmission verification and encryption mechanism to ensure data integrity and transmission security. Optimize the data interaction mechanism to reduce transmission redundancy and ensure data security. During implementation, a layered data transmission strategy is designed: edge nodes only upload abnormal data fragments, model inference results, 128-dimensional key feature vectors, and model parameter gradients; raw monitoring data is locally cached for a preset duration to avoid bandwidth consumption from full data transmission; an incremental transmission mode is adopted, synchronizing only newly added or changed data fragments and eliminating duplicate transmissions. The data transmission format is optimized by serializing structured data such as feature vectors and model parameters into Protocol Buffers format to compress transmission volume; a data transmission channel is built based on the MQTT lightweight communication protocol, with an adaptive transmission frequency set: transmission is performed at a preset low frequency under normal operating conditions, automatically switching to a high-frequency transmission mode in case of abnormal operating conditions or equipment failure. A data transmission guarantee mechanism is established: data integrity is verified through CRC32 cyclic redundancy check, triggering retransmission upon detection of packet loss or errors; AES symmetric encryption algorithm is used to encrypt transmitted data to prevent data leakage during transmission and ensure secure and reliable data interaction.
[0081] Step S144: Construct a horizontal federated learning framework. Each edge node acts as an independent local training end to train a lightweight model based on local data, and only uploads the model parameter gradients. The cloud uses a parameter aggregation algorithm to perform weighted aggregation of the parameter gradients of each edge node, generates global optimization parameters, and then sends them to the edge nodes to update the local model. An abnormal parameter detection mechanism and differential privacy technology are set up to balance privacy protection and model accuracy. An integrated federated learning mechanism enables collaborative model optimization while protecting data privacy. In implementation, a horizontal federated learning framework is constructed, with each edge node acting as an independent local training endpoint and the cloud server as the parameter aggregation endpoint. Each edge node independently trains a lightweight fault identification model based on locally collected device data. During training, only the model parameter gradients are calculated; the original device data is not uploaded, avoiding the privacy risks associated with cross-regional data sharing. Edge nodes upload parameter gradients to the cloud at preset intervals. The cloud uses the FedAvg algorithm to weighted aggregate the parameter gradients of all edge nodes. The aggregation weights are dynamically allocated based on the data volume and local model training performance of each edge node, generating global optimization parameters. The cloud then sends the global parameters back to each edge node to update local model parameters, achieving cross-node knowledge transfer. An abnormal parameter detection mechanism is implemented, using threshold judgment and cluster analysis to identify abnormal gradients uploaded by malicious nodes and eliminate invalid interference. Small Gaussian noise is added before parameter gradient upload, and differential privacy technology is used to balance data privacy protection and model training accuracy.
[0082] Step S145: Construct a multimodal feature fusion framework, design a three-level fusion mechanism for different types of heterogeneous data, including feature layer attention-weighted fusion, decision layer weighted voting fusion, and data layer missing data imputation fusion, establish a modality adaptation rule base, and optimize the fusion strategy for different device types and fault types; A multimodal heterogeneous data fusion module is designed to adapt to the complementary fusion of different data types. During implementation, a three-level multimodal feature fusion framework is constructed: In the feature layer fusion stage, for numerical (gas composition, temperature), spectral (PRPD) data, and time-series (resistance change, partial discharge signal) data, the weights of each modal feature are calculated using an attention mechanism, assigning higher weights to core fault features. Different modal feature vectors are mapped to a unified dimension and then weighted for fusion. In the decision layer fusion stage, a weighted voting method is used to integrate the inference results of each single-modal model, with voting weights dynamically adjusted based on the fault identification accuracy of each modal model. In the data layer fusion stage, a generative adversarial network (GAN) is used to fill in missing modal data, improving fusion robustness. A modal adaptation rule base is established, with pre-set fusion strategies and weight configurations for equipment types such as transformers, high and low voltage switchgear, and ring main units, as well as fault types such as insulation degradation and poor contact, ensuring that the complementarity of heterogeneous data is fully utilized in different scenarios, improving the completeness of fault features and the reliability of identification.
[0083] Step S146: Integrate lightweight algorithms, collaborative architecture, optimized data interaction, federated learning, and multimodal fusion modules to form an end-to-end unified heterogeneous data fusion framework. After edge nodes complete data collection, preprocessing, and local inference, they upload key information. The cloud aggregates parameters, trains and optimizes the multimodal fusion model, and performs complex fault analysis through federated learning. The optimized model parameters and fusion strategy are then sent back to the edge nodes. The framework adaptability interface is designed to support data access from different brands of equipment and different types of sensors, and to be compatible with existing power distribution network monitoring systems. The integration and implementation of a unified heterogeneous data fusion framework was completed, ensuring its compatibility and practicality. During implementation, the lightweight algorithm module, edge-cloud collaborative architecture, optimized data interaction mechanism, federated learning module, and multimodal fusion module were integrated into a single end-to-end unified heterogeneous data fusion framework. The framework's operation flow is as follows: edge nodes collect multimodal device data in real time, preprocess it, and input it into the lightweight model to complete local inference, uploading only key information (abnormal results, feature vectors, parameter gradients) to the cloud; the cloud aggregates parameters from each edge node through federated learning, trains and optimizes the multimodal fusion model, and simultaneously performs complex fault analysis based on global data; the optimized model parameters and fusion strategy are then back-distributed to each edge node to iteratively improve edge inference accuracy. A standardized adaptation interface was designed to support communication protocol access from different brands of equipment, be compatible with data formats from different types of sensors, and allow for framework deployment without modifying existing power distribution network monitoring systems, reducing engineering application costs.
[0084] Step S147: Establish a multi-dimensional evaluation index system, select different distribution network scenarios to conduct tests, compare the relevant performance before and after framework deployment, regularly collect framework operation data, analyze and optimize the bottlenecks of each module, and expand the modal adaptation rule base and federated learning nodes for new equipment types and fault modes.
[0085] Conduct framework performance verification and continuous iterative optimization to improve the framework's generalization ability. During implementation, establish a multi-dimensional evaluation index system, with core indicators including data transmission latency, edge node CPU / memory resource utilization, fault identification accuracy, data privacy protection strength, and heterogeneous data adaptation capability. Select different regions such as urban distribution networks and mountainous distribution networks, different equipment types such as transformers and high and low voltage switchgear, and different operating scenarios such as normal and extreme operating conditions to conduct multi-scenario tests, comparing key indicators such as fault identification efficiency, data transmission volume, and operation and maintenance costs before and after framework deployment. Regularly collect framework operation data to analyze bottlenecks in each module. For example, optimize communication protocol parameters for excessively high transmission latency, adjust multimodal weight configuration for insufficient fusion accuracy; expand the policy entries of the modal adaptation rule base for new equipment types (such as new intelligent switchgear) and fault modes (such as new insulation aging faults), and add federated learning edge nodes for corresponding regions to continuously improve the framework's scenario adaptation capability and generalization performance.
[0086] Step S150: Build a simulation test platform to conduct offline verification, select typical lines for pilot application and online data collection and analysis, establish an operation and maintenance feedback loop, and build a full-process verification and iterative optimization mechanism.
[0087] A full-process verification and iterative optimization mechanism is established to continuously improve the framework through offline verification, pilot applications, and operation and maintenance feedback. During implementation, a distribution network equipment simulation test platform is built to simulate equipment operation under different fault types and operating condition combinations. Test data is collected to perform offline verification of the heterogeneous data unified fusion framework, testing the framework's performance indicators such as fault identification accuracy and response speed in a controlled environment. Typical distribution network lines are selected as pilot areas, and edge nodes and cloud-based collaborative systems are deployed to conduct online data collection and fault identification applications, recording the framework's operating status and actual operation and maintenance data in real time. An operation and maintenance feedback closed-loop mechanism is established to regularly collect framework application issues reported by operation and maintenance personnel (such as false alarms / missed alarms, interface compatibility issues). Combining online data collection and offline test results, the root causes of problems are analyzed, and framework parameters (such as multimodal fusion weights, federated learning aggregation strategies) are adjusted, and module functions (such as lightweighting and data interaction frequency) are optimized, forming a full-process mechanism of "offline verification - pilot application - feedback analysis - iterative optimization" to ensure the framework adapts to actual operation and maintenance needs.
[0088] Based on the same inventive concept, please refer to Figure 2 This paper shows a schematic block diagram of the structure of a distribution network transformer-high and low voltage switchgear-ring switchgear state monitoring fault early warning system 100 provided in this application embodiment for executing the above-described distribution network transformer-high and low voltage switchgear-ring switchgear state monitoring fault early warning method. The distribution network transformer-high and low voltage switchgear-ring switchgear state monitoring fault early warning system 100 may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0089] In this embodiment, both the machine-readable storage medium 120 and the processor 130 are located in the distribution transformer-high and low voltage switchgear-ring network cabinet status monitoring and fault early warning system 100 and are separately configured. However, it should be understood that the machine-readable storage medium 120 may also be independent of the distribution transformer-high and low voltage switchgear-ring network cabinet status monitoring and fault early warning system 100 and may be accessed by the processor 130 through a bus interface. Alternatively, the machine-readable storage medium 120 may also be integrated into the processor 130 and may communicate and interact with external systems through the communication unit 110.
[0090] The processor 130 is the control center of the distribution network transformer-high and low voltage switchgear-ring mains condition monitoring and fault early warning system 100. It connects to various parts of the system via various interfaces and lines. By running or executing software programs and / or modules stored in the machine-readable storage medium 120, and by calling data stored in the machine-readable storage medium 120, it performs various functions and processes data within the system, thereby providing overall monitoring of the distribution network transformer-high and low voltage switchgear-ring mains condition monitoring and fault early warning system 100. Optionally, the processor 130 may include one or more processing cores; for example, the processor 130 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor. The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to realize the status monitoring and fault early warning method for distribution transformer-high and low voltage switchgear-ring network cabinet provided in the aforementioned method embodiment.
[0091] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
Claims
1. A state monitoring and fault early warning method for a distribution network transformer-high and low voltage cabinet-ring network cabinet, characterized in that: Includes the following steps: Based on fault mechanism modeling and digital twin technology, the sample library is expanded, and meta-learning and domain adaptive transfer learning are combined to construct a small sample enhancement and cross-domain transfer learning system. The evolution path of complex faults was analyzed and a fault chain state transition matrix was constructed. A fault chain temporal correlation model was established using a deep learning model with temporal attention mechanism and a multi-label classification algorithm. By integrating static and dynamic parameters of the equipment to construct a full-dimensional profile of the equipment, a dynamic threshold model is built based on a weighted regression algorithm, and a reinforcement learning mechanism is introduced to iteratively optimize the threshold calculation factor to construct an adaptive dynamic threshold adjustment algorithm. We build an edge and cloud collaborative computing architecture by modifying algorithms to be lightweight, optimize the data interaction mechanism, and combine federated learning and multimodal fusion algorithms to construct a unified fusion framework for heterogeneous data. A simulation test platform was built to conduct offline verification, typical lines were selected for pilot applications and online data collection and analysis, a closed loop of operation and maintenance feedback was established, and a full-process verification and iterative optimization mechanism was constructed.
2. The method according to claim 1, characterized in that: The method of expanding the sample library based on fault mechanism modeling and digital twin technology, combined with meta-learning and domain adaptive transfer learning, constructs a small-sample enhancement and cross-domain transfer learning system, including: The physical evolution law of power distribution equipment faults is analyzed, the variation mechanism of characteristic parameters including PRPD spectrum characteristics, proportion of dissolved gas components in oil, and time-series changes in temperature resistance is clarified, and a mathematical model of fault mechanism is constructed by combining multidisciplinary theories to quantify the mapping relationship between fault severity and characteristic parameters. Based on equipment design, material and operation data, a high-fidelity digital twin is built using multi-physics coupling simulation technology, fault excitation factors are injected, and the equipment operation status under multiple fault types, multiple stages and multiple working conditions is simulated to generate multiple types of highly realistic synthetic samples. Statistical testing methods are used to verify the consistency of feature distribution between synthetic samples and real samples. The parameters of the digital twin model are optimized to ensure the credibility of the samples. Samples are classified and labeled according to three dimensions: equipment type, failure mode, and operating condition parameters. Various types of samples are integrated to build a structured sample library with full state coverage. Based on the sample library, a three-dimensional meta-task set is divided. A lightweight deep learning network is selected and combined with the MAML framework to build a meta-learning model. The initialization parameters are optimized through an alternating training strategy, giving the model the ability to quickly adapt to the fault characteristics of new devices / new scenarios. Identify heterogeneous devices in the source and target domains, construct a transfer learning model using a domain adaptation algorithm, achieve knowledge transfer by minimizing feature distribution differences, and optimize the model's adaptability to heterogeneous devices by combining pseudo-label training with real samples from the target domain and a feature adaptation module. Continuously add new real samples and update the parameters of the digital twin model. Optimize the meta-learning task set and model parameters based on the new samples, establish a transfer learning effect evaluation mechanism, and trigger parameter calibration to form a dynamic closed loop of sample expansion, model training, transfer adaptation, effect evaluation, and parameter optimization.
3. The method according to claim 2, characterized in that: The process involves dividing a three-dimensional meta-task set based on a sample library, selecting a lightweight deep learning network, and constructing a meta-learning model using the MAML framework. An alternating training strategy is employed to optimize initialization parameters, enabling the model to quickly adapt to fault characteristics of new devices / scenes. This includes: Feature engineering was performed on the structured sample library to extract features from PRPD spectra, dissolved gases in oil, and temperature / resistance time series data. After Z-score standardization, the meta-training / validation / test sets were divided into 8:1:1 sets, and the test set included samples from new equipment / new operating conditions. A three-dimensional meta-task set is constructed based on equipment type, fault mode, and operating condition parameters. Each task is configured with 5-10 support sets per fault and 2-3 times the number of query sets. The task set covers more than 90% of operating scenarios and includes 10%-15% of edge operating conditions. A multi-branch feature extractor was built using the MobileNetV3 lightweight network to extract features from different types of fault data and then merge them into a 128-dimensional unified feature vector. A lightweight feature extractor is integrated with the MAML framework to build a meta-learning model, defining cross-entropy and contrastive loss, and setting hyperparameters such as inner / outer loop learning rate and meta-batch size; The initialization parameters were optimized using an inner loop adaptation and an outer loop update strategy. The learning rate was periodically verified and adjusted until the model converged and the generalization error was ≤8%. For new devices / scenes, take 5-10 samples as the support set, fine-tune in 1-3 rounds of internal loop, verify the adaptation effect, and ensure that the accuracy reaches more than 85% of the traditional model with 5 samples and converges 10 times faster. After pruning and quantization, the model size is compressed to 20%-30%. A dynamic update and performance monitoring mechanism for meta-tasks is established, and model retraining is triggered when the adaptation accuracy is lower than 80%.
4. The method according to claim 3, characterized in that: The integrated lightweight feature extractor and MAML framework construct a meta-learning model, defining cross-entropy and contrastive loss, and setting hyperparameters such as inner / outer loop learning rate and meta-batch size, including: The interface of the lightweight feature extractor is standardized and adapted to the MAML framework. Based on the deep learning framework, a MAML module with inner and outer loop sub-modules is built. The closed-loop architecture of feature extraction, inner loop adaptation and outer loop update is built through hierarchical direct connection and reverse parameter association, and the model data flow is planned. The cross-entropy loss is calculated to measure the basic classification error. The contrastive loss is designed to bring the features of similar samples closer together and push the features of dissimilar samples apart. The cross-entropy loss and the contrastive loss are fused according to the preset weights to obtain the total loss. The gradient of the contrastive loss is then clipped. Select and initialize the inner loop hyperparameters, outer loop hyperparameters, and other hyperparameters. The inner loop hyperparameters include the learning rate, number of training epochs, and optimizer. The outer loop hyperparameters include the learning rate, meta-batch size, and optimizer. The other hyperparameters include the attention weights of the feature fusion layer, the dropout rate of the fully connected layer, and the total number of training epochs. Add boundary constraints at the code level to limit the hyperparameters from exceeding the preset reasonable range. Hyperparameter tuning is carried out in three stages: pre-tuning, fine-tuning, and final tuning. In the pre-tuning stage, the inner / outer loop learning rate is fixed, different meta-batch sizes are tested, and the lowest value of the validation set loss is selected and locked. In the fine-tuning stage, the meta-batch size is fixed, and the combination of the inner and outer loop learning rates is searched by grid and the optimal combination is selected. In the final tuning stage, the learning rate is locked, and the marginal value of the comparison loss and the loss fusion weight are fine-tuned to ensure that the query set accuracy reaches the preset standard in small sample scenarios. The optimal hyperparameters are saved as a configuration file for model reuse. The model is compiled based on a deep learning framework. A custom MAML meta-optimizer and fusion loss function are specified, and accuracy, precision, and recall are set as evaluation metrics. Meta-task samples are randomly selected and input into the untrained model to verify the output results of each module. A single meta-task is selected to complete one round of inner / outer loop training. The gradient values of each layer of the feature extractor are tracked to verify that the parameters can be updated effectively. The initialized model is used to predict the meta-test set, and the initial accuracy is recorded as the training baseline. A learning rate decay rule is set up so that when the accuracy of the validation set does not improve for a preset number of consecutive meta-batches, the learning rate of the inner / outer loop is reduced. If there is still no improvement, early stopping is triggered. An adaptive rule for the meta-batch size is designed to adjust the meta-batch size according to the GPU memory usage. A dynamic adjustment rule for the loss weight is designed to adjust the loss fusion weight according to the false positive rate and false negative rate of small sample adaptation, so as to balance feature discrimination and classification accuracy.
5. The method according to claim 2, characterized in that: The aforementioned heterogeneous devices with clearly defined source and target domains employ a domain adaptation algorithm to construct a transfer learning model. Knowledge transfer is achieved by minimizing the difference in feature distribution. Combined with pseudo-label training using real samples from the target domain and a feature adaptation module, the model's adaptability to heterogeneous devices is optimized, including: Define the source and target domains, identify heterogeneous dimensions, and establish a list to clarify the characteristic directions of migration adaptation; Based on a lightweight feature extractor integrated with a domain adaptive module, the model is split into three layers: shared feature extraction, domain classification, and fault classification. A gradient inversion layer is added to connect the shared feature extraction layer and the domain classification layer. Select an appropriate feature distribution measurement method, construct a fault classification and domain adaptation fusion loss function, and optimize it through backpropagation to make the feature distributions of the two domains converge. Use a source domain pre-trained model to generate target domain pseudo-labels, select high-confidence samples and mix them with real samples to construct a target domain training set, mix them with source domain samples for iterative training and update pseudo-labels; An attention mechanism feature adaptation module is added between the shared feature extraction layer and the fault classification layer to dynamically adjust feature weights, fine-tune dimensions, pre-train initial weights and update them synchronously. The training process consists of three stages: source domain pre-training, domain adaptation training, and target domain fine-tuning. The training is advanced according to the difference values of feature distribution. During the target domain fine-tuning, some parameters are frozen and related modules are updated. Establish an evaluation system that includes recognition accuracy and false alarm rate. Evaluate after each round of training. If the target is not met, adjust the parameters or screen samples. Repeat the process until the performance meets the requirements.
6. The method for condition monitoring and fault early warning of distribution network transformer-high and low voltage switchgear-ring mains cabinet according to claim 5, characterized in that: The selection of an appropriate feature distribution metric, the construction of a fault classification and domain adaptation fusion loss function, and the optimization through backpropagation to make the feature distributions of the two domains converge include: Classify and sort the fault characteristics of distribution network equipment, clarify the data attributes and distribution characteristics of each type of characteristic, divide the fault characteristics into three categories: numerical, high-dimensional vector, and time series, analyze the distribution pattern of each type of characteristic in the source domain and target domain, and clarify the different distribution differences of different types of characteristics between the two domains. Numerical features use KL divergence and JS divergence to measure distribution similarity; high-dimensional vector features use MMD and Wasserstein distance to measure distribution distance; and time-series features use dynamic time warping and time-series distribution similarity to comprehensively measure distribution differences. Cross-entropy loss is used as the basic fault classification loss. Based on the true labels of the source domain samples and the high-confidence pseudo labels of the target domain samples, the domain adaptation loss of different feature types is integrated by weighted summation. The weights are set according to the importance of the features. The fault classification loss and the domain adaptation loss are fused according to the preset weights to obtain the total loss function. The weights of the two types of loss in the total loss function can be dynamically adjusted according to the training effect. An L2 regularization term is introduced into the domain adaptation loss to limit the complexity of the model parameters. A smoothing constraint term is added to the domain adaptation loss of time-series features to reduce noise interference. A boundary threshold for the loss value is set, and its contribution ratio is fixed when the domain adaptation loss is lower than the threshold. In the forward propagation phase, after the source and target domain samples have their features extracted by the shared feature extraction layer, they are respectively input into the fault classification layer to calculate the classification loss, and the input domain classification layer to calculate the domain adaptation loss and summarize the total loss. In the gradient calculation and propagation phase, the gradient of the total loss with respect to each trainable parameter of the model is calculated based on the chain rule. The gradient of the domain classification layer is reversed by the gradient reversal layer and then propagated back to the shared feature extraction layer. In the parameter update phase, the adaptive optimizer is used to update the model parameters according to the gradient direction and magnitude, so that the feature distribution of the two domains gradually converges in the shared feature space. In the early stage of training, increase the weight of fault classification loss and decrease the weight of domain adaptation loss to ensure that the model can master the source domain fault classification ability. In the middle stage of training, gradually increase the weight of domain adaptation loss and decrease the weight of classification loss to focus on promoting the alignment of feature distributions in the two domains. In the later stage of training, dynamically adjust the weights according to the difference in feature distributions between the two domains to accelerate distribution convergence. After each training round, the distribution metrics of various features in the two domains, the total loss value, and the fault classification accuracy are calculated. The index change curves are plotted, and convergence conditions are set. When the decrease of the distribution metrics of the features in the two domains is less than the preset threshold for several consecutive rounds, and the fault classification accuracy is stable in the target range, it is determined that the optimization has converged and the parameter update is stopped. If the classification accuracy decreases significantly during the training process, the domain adaptation loss weight is reduced and the training is iterated again.
7. The method for condition monitoring and fault early warning of distribution network transformer-high and low voltage switchgear-ring mains cabinet according to claim 1, characterized in that: The process of identifying complex fault evolution paths and constructing a fault chain state transition matrix involves using a deep learning model with temporal attention mechanisms and a multi-label classification algorithm to establish a fault chain temporal correlation model, including: The chain relationships between the main faults and derivative faults of typical composite faults of three types of equipment were sorted out and listed. The composite faults were deconstructed into a four-order evolution framework, and the triggering conditions, duration characteristics and transition critical points of each stage were clarified. Define a set of fault states including normal states and a characteristic judgment standard. Based on historical data and simulation statistical operating conditions, adapt the state transition probabilities and construct and correct an N×N dimensional state transition matrix. A temporal feature extraction layer is built using GRU or bidirectional LSTM, and an additive attention mechanism is designed to calculate the weights of time steps. The resulting global feature vector containing key temporal information is obtained by fusing the data. The design employs a multi-label output layer, using binary cross-entropy and FocalLoss, and dynamically sets the classification threshold for each state based on the validation set. By embedding a state transition matrix and adding residual connections and gating units, an end-to-end model is formed that integrates temporal extraction, attention fusion, matrix embedding, feature interaction, and multi-label output. Construct and preprocess a composite fault time series dataset, train the model in stages, and optimize the hyperparameters through grid search; The model is evaluated using multiple indicators to assess recognition accuracy and path tracing accuracy, qualitatively verified to capture the transformation node capability, and robustness testing is conducted to dynamically adjust the model.
8. The method for condition monitoring and fault early warning of distribution network transformer-high and low voltage switchgear-ring mains cabinet according to claim 7, characterized in that: The multi-label output layer design employs binary cross-entropy and FocalLoss, dynamically setting classification thresholds for each state based on the validation set, including: Based on the total number of fault states, the number of neurons in the output layer is set to correspond one-to-one with the number of states. The output layer receives the feature vectors of the preceding modules and completes the dimension mapping through the fully connected layer. Sigmoid is selected as the activation function so that the output value of each neuron represents the probability of the existence of the corresponding fault state. It supports non-mutually exclusive independent judgment of multiple fault states and adapts to the multi-stage coexistence features of composite faults. For each fault state, a binary cross-entropy loss is calculated state by state based on the predicted probability and the true label to reflect the basic error of single-state classification. Focal Loss is introduced, and the loss weight of samples in different states is adjusted by the balancing coefficient. The loss contribution of hard-to-classify samples is amplified and the interference of easy-to-classify samples is weakened by the focusing parameter. The binary cross-entropy loss and Focal Loss are weighted and fused, and the average of the fused loss of each state is taken as the total loss. Divide into independent validation sets covering all fault states, equipment types, and operating conditions, perform the same preprocessing process as the training set, and perform stratified sampling according to fault states to increase the proportion of fault states with fewer samples in the validation set and avoid the distortion of threshold setting caused by sample distribution bias. For each fault state, the preset threshold range is traversed, and the sensitivity and specificity under each threshold are calculated. The threshold corresponding to the maximum value is selected by Youden index as the initial classification threshold for that state. The threshold selection for each state is completed independently according to the sample distribution and recognition difficulty differences of different fault states. The initial threshold is applied to the validation set to calculate the precision, recall, and F1 score for each state. If the F1 score for a certain state does not meet the preset requirements, the FocalLoss balance coefficient or focusing parameter corresponding to that state is adjusted, the model is retrained, and the threshold is selected again. The parameter adjustment, model retraining, and threshold selection process is repeated until the F1 score for all states meets the requirements. Set the threshold update cycle. In each cycle, the current model is used to re-predict the validation set, calculate the Youden index of each state and update the threshold. If the threshold change is less than the preset value for two consecutive cycles, the threshold of that state is fixed. If the model is overfitting, backtrack to the threshold of the previous cycle. During inference, the predicted probability of each state is compared with the corresponding optimal threshold. If the probability is greater than or equal to the threshold, the state is determined to exist and the fault state combination is output. The judgment result is verified by combining the fault chain state transition matrix. If the fault evolution logic is violated, the corresponding state threshold is finely adjusted and the judgment is re-determined.
9. The method for condition monitoring and fault early warning of distribution network transformer-high and low voltage switchgear-ring mains cabinet according to claim 1, characterized in that: The integrated static and dynamic parameters of the device construct a comprehensive profile of the device. A dynamic threshold model is built based on a weighted regression algorithm. A reinforcement learning mechanism is introduced to iteratively optimize the threshold calculation factor, and an adaptive dynamic threshold adjustment algorithm is constructed, including: The equipment's static and dynamic parameters were sorted out to form two types of parameter lists. The parameter naming, units and coding rules were standardized. The sampling frequency and time granularity standards for dynamic parameters were clarified. Static parameters were normalized and dynamic parameters were standardized using Z-score. Missing parameters were filled in and outliers were identified and corrected using the 3σ principle. A profile dimension framework is constructed, which includes a basic attribute layer, a structural feature layer, an operating status layer, a fault feature layer, and an environment adaptation layer. The standardized static parameters and time-series dynamic parameters are associated with the unique identifier of the device to form structured profile data. The multi-source dynamic parameters are fused by time-series splicing and attention weighting. The fused dynamic features are spliced with static features and generated by mapping through a fully connected layer to generate a full-dimensional feature vector of the device, and a profile update mechanism is established. Features were selected by ranking features by Pearson correlation coefficient and random forest feature importance to determine the input feature set of the model. Multivariate weighted linear regression was selected as the base model. Historical fault critical feature values were used as labels. Input features and assign initial threshold calculation factors based on importance. Output the initial dynamic threshold of each fault feature. Regression models were trained separately according to equipment type. The threshold optimizer is defined as a reinforcement learning agent, the equipment operation scenario is used as the environment, a state space is constructed including the equipment's full-dimensional feature vector, threshold calculation factor, fault identification index, real-time operating condition, and action space for threshold calculation factor adjustment operation. A multi-objective weighted reward function that integrates fault identification recall rate, false alarm rate, and threshold fluctuation amplitude is designed, and a simulation training environment is built based on historical equipment data and fault cases. The initial threshold calculation factor of the weighted regression model is used as the initial action of the agent. The iteration rounds and exploration rate are set, the profile data is loaded into the training environment, the agent obtains the state from the environment and selects factors to adjust the action, the environment calculates the reward value after executing the action, the action strategy is updated through reinforcement learning algorithm, iterates until the reward value converges, extracts the optimal threshold calculation factor and saves it according to the equipment working condition type to form a factor library. Set threshold adjustment trigger conditions, including sudden changes in equipment operating conditions, excessive false alarm or missed alarm rate in fault identification, equipment profile update, and timed trigger. When the adjustment is triggered, extract the current equipment profile features, match the optimal calculation factor in the factor library, input the weighted regression model to calculate the new threshold. If the difference between the new and old thresholds reaches the preset standard, use exponential moving average to smooth the new threshold, and set the upper and lower limits of the threshold through the equipment rated parameters or historical fault critical values. The process is executed according to the real-time monitoring, condition judgment, factor matching, threshold calculation, smoothing verification, and threshold update process. Establish an evaluation index system with fault identification accuracy, false alarm rate, false negative rate, and adjustment response speed as the core. Conduct multi-scenario tests with multiple equipment types, multiple operating conditions, and multiple fault types. Compare the identification effects of dynamic thresholds and fixed thresholds. Regularly collect algorithm running records and fault identification results. Retrain the regression model and reinforcement learning agent, update the factor library, supplement sample data for new operating conditions and new fault types, and optimize the reward function weights and action space.
10. A condition monitoring and fault early warning system for distribution network transformers, high and low voltage switchgear, and ring main units, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the condition monitoring and fault early warning method for distribution transformer-high and low voltage switchgear-ring mains cabinet as described in any one of claims 1 to 9 by executing the machine-executable instructions.