Early warning method and system for mechanical fault of transformer
By processing and evaluating multimodal data from transformer condition monitoring, the problem of low diagnostic confidence in single-sensor-single-parameter mode was solved, enabling early trend perception and high-reliability early warning of transformer mechanical faults.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 广州南网科研技术有限责任公司
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-15
AI Technical Summary
Existing transformer mechanical condition monitoring mainly adopts a single sensor-single parameter monitoring mode, which leads to low confidence in diagnostic conclusions, inability to effectively locate faults, and reduced reliability of transformer mechanical fault early warning.
By acquiring multiple training mechanical state data, a mechanical feature set is formed after preprocessing. The target feature extractor and evaluation network are used to extract and evaluate the multimodal fault feature data of the transformer state. The hierarchical early warning is then carried out by combining a multi-head cross-attention fusion module, an interpretable diagnostic network and a time series prediction model.
It enables early trend detection of transformer mechanical faults, reduces the probability of missed and false alarms, and improves the reliability of fault early warning.
Smart Images

Figure CN122045949A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of transformer monitoring technology, and in particular to a method and system for early warning of mechanical faults in transformers. Background Technology
[0002] Transformers are core equipment in power systems, and their safe and stable operation directly affects the reliability of the power grid and the quality of power supply. The core components inside a transformer, such as windings, core, and clamps, are subjected to electrodynamic and thermal stresses over long periods, and may also suffer hidden damage during transportation and installation. These factors can easily lead to mechanical faults such as component loosening, deformation, displacement, or insulation wear. These faults are often difficult to detect in their early stages; if not detected and addressed promptly, they can easily escalate into catastrophic accidents such as insulation breakdown or equipment burnout during transient events like short-circuit current surges, causing enormous economic losses and adverse social impacts. Therefore, reliable monitoring and early warning of transformer mechanical conditions are crucial.
[0003] Currently, existing transformer mechanical condition monitoring mainly adopts a "single sensor-single parameter" monitoring mode. For example, it relies solely on vibration acceleration sensors to monitor tank vibration or solely on acoustic sensors to capture signals. However, the internal structure of transformers is complex, and changes in a single physical quantity are often ambiguous. Different faults may exhibit similar signal characteristics. This information silo phenomenon leads to low confidence in diagnostic conclusions and makes it impossible to effectively locate fault positions, thus reducing the reliability of transformer mechanical fault early warning. Summary of the Invention
[0004] This invention provides a method and system for early warning of mechanical faults in transformers, which solves the problem that existing transformer mechanical condition monitoring mainly adopts a "single sensor-single parameter" monitoring mode, such as relying solely on vibration acceleration sensors to monitor tank vibration or using only acoustic sensors to capture signals. However, the internal structure of transformers is complex, and changes in a single physical quantity are often ambiguous. Different faults may exhibit similar signal characteristics. This information silo phenomenon leads to low confidence in diagnostic conclusions and makes it impossible to effectively locate the fault location, thus reducing the reliability of early warning of mechanical faults in transformers.
[0005] The first aspect of this invention provides a method for early warning of mechanical faults in transformers, comprising: Multiple training machine state data are acquired, and each training machine state data is preprocessed to obtain a corresponding machine feature set; The preset transformer state detection model is trained using the mechanical feature set to obtain the corresponding target transformer state detection model, wherein the target transformer state detection model includes a target feature extractor and a target evaluation network; The mechanical state data of the transformer under test is acquired, and the target feature extractor is used to extract features from the mechanical state data to obtain the corresponding multimodal fault feature data. The target evaluation network is used to evaluate the state of the multimodal fault feature data to obtain the corresponding evaluation results, and graded early warning is given based on the evaluation results.
[0006] Optionally, the step of training a preset transformer state detection model using the mechanical feature set to obtain a corresponding target transformer state detection model includes: The mechanical feature set is used to perform adversarial training on the feature extractor in the preset transformer state detection model to obtain the corresponding target feature extractor; The mechanical feature set is input into the target feature extractor to obtain the corresponding training modality feature dataset; The training modal feature dataset is input into the evaluation network of the transformer state detection model to obtain the corresponding training evaluation data; Based on a preset loss function, the training loss function value of the mechanical feature set is calculated according to the training evaluation data; When the training loss function value is greater than or equal to a preset loss threshold, the network parameters of the evaluation network are adjusted until the training loss function value is less than the loss threshold. When the training loss function value is less than the loss threshold, a corresponding target evaluation network is generated.
[0007] Optionally, the step of using the mechanical feature set to perform adversarial training on the feature extractor in the preset transformer state detection model to obtain the corresponding target feature extractor includes: The mechanical feature set is input into a preset generator to obtain the corresponding noise enhancement dataset; The mechanical feature set and the noise enhancement dataset are respectively input into the feature extractor in the preset transformer state detection model to obtain the noise feature dataset and the training state feature dataset; The noise feature dataset and the training state feature dataset are input into a preset domain discriminator to obtain the corresponding domain discrimination score data; Based on the preset adversarial loss function, the corresponding adversarial loss function value is determined according to the domain discrimination score data, the noise enhancement dataset, and the mechanical feature set; When the adversarial loss function value is greater than or equal to the preset adversarial loss threshold, the model parameters of the feature extractor and the domain discriminator are adjusted until the adversarial loss function value is less than the adversarial loss threshold. When the adversarial loss function is less than the adversarial loss threshold, a corresponding target feature extractor is generated.
[0008] Optionally, the target feature extractor includes a first feature extraction branch and a second feature extraction branch, the mechanical state data includes one-dimensional state data and two-dimensional state data, and the step of extracting features from the mechanical state data using the target feature extractor to obtain corresponding multimodal fault feature data includes: Temporal features are extracted from the one-dimensional state data through the first feature extraction branch to obtain the corresponding temporal feature data. Spatial feature extraction is performed on the two-dimensional state data through the second feature extraction branch to obtain the corresponding spatial feature data; The temporal feature data and the spatial feature data are used as multimodal fault feature data.
[0009] Optionally, the target evaluation network includes a multi-head cross-attention fusion module, an interpretable diagnostic network, and a temporal prediction model. The step of performing state evaluation on the multimodal fault feature data through the target evaluation network to obtain the corresponding evaluation result includes: The multi-head cross-attention fusion module performs feature fusion on the multimodal fault feature data to obtain the corresponding attention fusion feature vector; The attention fusion feature vector is used to perform fault detection through the interpretable diagnostic network to obtain the corresponding fault type confidence and visual heatmap. The attention fusion feature vector, the fault type confidence, and the visual heatmap are used to construct a corresponding multidimensional temporal feature sequence; The state prediction of the multidimensional time-series feature sequence is performed using the time-series prediction model to obtain the corresponding evaluation results.
[0010] Optionally, the interpretable diagnostic network includes a convolutional feature extraction layer, a fault localization branch, and a fault classification branch. The step of performing fault detection on the attention fusion feature vector through the interpretable diagnostic network to obtain the corresponding fault type confidence and visual heatmap includes: The attention fusion feature vector is extracted through the convolutional feature extraction layer to obtain the corresponding attention fusion feature map and the pre-convolutional feature map; The attention fusion feature map is used to diagnose faults through the fault classification branch to obtain the corresponding fault category confidence. The fault classification branch includes a fully connected stacked module and a Softmax activation function connected in sequence. The fault location is performed on the pre-convolutional feature map using the fault location branch and the fault category confidence, resulting in a corresponding visual heatmap.
[0011] A second aspect of the present invention provides a transformer mechanical fault early warning system, comprising: The preprocessing module is used to acquire multiple training machine state data, preprocess each training machine state data, and obtain the corresponding machine feature set. The training module is used to train the preset transformer state detection model using the mechanical feature set to obtain the corresponding target transformer state detection model, wherein the target transformer state detection model includes a target feature extractor and a target evaluation network; The extraction module is used to acquire the mechanical state data of the transformer under test, and to extract features from the mechanical state data through the target feature extractor to obtain the corresponding multimodal fault feature data. The evaluation module is used to perform state evaluation on the multimodal fault feature data through the target evaluation network, obtain the corresponding evaluation results, and perform graded early warning based on the evaluation results.
[0012] A third aspect of the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the transformer mechanical fault early warning method described above.
[0013] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the transformer mechanical fault early warning method as described above.
[0014] The fifth aspect of the present invention provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer performs the transformer mechanical fault early warning method as described above.
[0015] As can be seen from the above technical solutions, the present invention has the following advantages: By acquiring multiple training mechanical state data, preprocessing each training mechanical state data to obtain corresponding mechanical feature sets, and using these mechanical feature sets to train a pre-set transformer state detection model, a corresponding target transformer state detection model is obtained. The target transformer state detection model includes a target feature extractor and a target evaluation network. It acquires the mechanical state data of the transformer under test, extracts features from the mechanical state data using the target feature extractor to obtain corresponding multimodal fault feature data, and evaluates the state of the multimodal fault feature data using the target evaluation network to obtain corresponding evaluation results. Based on the evaluation results, a graded early warning is issued. This overcomes the technical problem that existing transformer mechanical state monitoring mainly adopts a "single sensor-single parameter" monitoring mode, resulting in low confidence of diagnostic conclusions and inability to effectively locate fault positions, thus reducing the reliability of transformer mechanical fault early warning. Compared with traditional transformer mechanical state monitoring methods, this invention uses a target evaluation network to evaluate the state of multimodal fault feature data and execute graded early warnings. It upgrades from passive fixed threshold alarms to proactive early warnings based on state time-series evolution, achieving early fault trend perception, effectively reducing the probability of missed and false alarms, and improving the reliability of transformer mechanical fault early warning. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the steps of a transformer mechanical fault early warning method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating the steps of a transformer mechanical fault early warning method provided in Embodiment 2 of the present invention; Figure 3 This is a structural block diagram of a transformer mechanical fault early warning system provided in Embodiment 3 of the present invention; Figure 4 This is a structural block diagram of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0018] This invention provides a method and system for early warning of mechanical faults in transformers, which addresses the problem that existing transformer mechanical condition monitoring mainly adopts a "single sensor-single parameter" monitoring mode, such as relying solely on vibration acceleration sensors to monitor tank vibration or using only acoustic sensors to capture signals. However, the internal structure of transformers is complex, and changes in a single physical quantity are often ambiguous. Different faults may exhibit similar signal characteristics. This information silo phenomenon leads to low confidence in diagnostic conclusions and makes it impossible to effectively locate the fault position, thus reducing the reliability of early warning of mechanical faults in transformers.
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. It should be noted that in the optional embodiments of the present invention, the object information and other related data involved require the permission or consent of the object when the embodiments of the present invention are applied to specific products or technologies, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of the present invention involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of the relevant departments, and in compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented with the authorization and consent of the object.
[0020] Please see Figure 1 , Figure 1 The flowchart illustrates the steps of a transformer mechanical fault early warning method provided in Embodiment 1 of the present invention.
[0021] This invention provides a method for early warning of mechanical faults in transformers, comprising: Step 101: Obtain multiple training machine state data, preprocess each training machine state data, and obtain the corresponding machine feature set.
[0022] Training mechanical condition data refers to the multi-source raw monitoring data collected for training the transformer condition detection model, which reflects different mechanical operating states (normal state and various fault states) of the transformer. It is the basic data for the model to learn fault characteristics and covers different types of monitoring data such as vibration, acoustic, infrared and electrical.
[0023] Preprocessing refers to a series of data cleaning and normalization operations performed on training machine state data. The purpose is to eliminate invalid interference information in the data, unify data standards, improve data quality, and provide a reliable data foundation for subsequent feature extraction and model training. Common operations include detrending, filtering and denoising, normalization, and spatiotemporal alignment.
[0024] The mechanical feature set refers to the dataset containing various effective mechanical operating features of a transformer, which is formed by integrating the training mechanical state data after preprocessing. This dataset extracts and retains key information that reflects the mechanical state of the transformer.
[0025] In this embodiment of the invention, multi-source raw monitoring data covering vibration, acoustic, infrared, and electrical types are collected from the transformer operation monitoring system as training mechanical state data. Preprocessing operations such as detrending, filtering and denoising, and data normalization are performed on each training mechanical state data in sequence to eliminate environmental interference, system errors, and dimensional differences in the data. Then, the preprocessed data with different dimensions and different acquisition frequencies are matched with timestamps and mapped to spatial dimensions through data spatiotemporal alignment, and finally integrated to form a mechanical feature set containing various effective state features.
[0026] It should be noted that the detrending term is used to eliminate trend biases in the training mechanical state data caused by factors such as system drift and slow environmental changes, restoring the true fluctuation characteristics of the data itself. Filtering and denoising refers to eliminating invalid signals such as background noise, electromagnetic interference, and environmental interference in the training mechanical state data using specific filtering algorithms, retaining the valid signals that reflect the transformer's mechanical state. Data normalization refers to mapping training mechanical state data with different dimensions and numerical ranges to the same numerical interval, eliminating differences in dimensions and numerical ranges between data, and improving the model's training efficiency and convergence speed. Spatiotemporal alignment refers to achieving synchronous matching in the time dimension based on timestamps, while simultaneously combining the transformer's physical structure to achieve dimensional mapping between monitoring data and equipment spatial location, ensuring consistency of different types of data in both time and space dimensions.
[0027] Step 102: Train the preset transformer state detection model using the mechanical feature set to obtain the corresponding target transformer state detection model, wherein the target transformer state detection model includes a target feature extractor and a target evaluation network.
[0028] The pre-built transformer condition detection model refers to a pre-built basic deep learning model framework for detecting mechanical fault conditions in transformers, which includes an untrained feature extractor and evaluation network.
[0029] The target transformer condition detection model refers to the final model that has the ability to extract mechanical fault features, assess condition, and provide early warning after being trained with a mechanical feature set. It consists of a target feature extractor and a target assessment network that have been trained and adapted.
[0030] In this embodiment of the invention, adversarial training is performed on the feature extractor in the preset transformer condition detection model based on the mechanical feature set (i.e., the feature extractor learns to remove noise interference and extract the essential features of the fault). After training, a suitable target feature extractor is obtained. The mechanical feature set is input into the trained target feature extractor to extract the training modal feature dataset that reflects the mechanical state of the transformer. The training modal feature dataset is input into the evaluation network of the transformer condition detection model for training. The loss function value during the training process is calculated according to the preset loss function. The network parameters of the evaluation network are continuously adjusted until the loss function value is lower than the preset loss threshold, thus obtaining the trained target evaluation network.
[0031] Step 103: Obtain the mechanical state data of the transformer under test, and extract features from the mechanical state data using a target feature extractor to obtain the corresponding multimodal fault feature data.
[0032] The transformer under test refers to a transformer device that is in actual operation and requires mechanical fault detection, condition assessment, and early warning.
[0033] Mechanical condition data refers to multi-source heterogeneous monitoring data collected in real time from the transformer under test, which reflects its current mechanical operating status. It includes two categories: one-dimensional condition data and two-dimensional condition data.
[0034] Multimodal fault feature data refers to a fault feature dataset extracted and integrated from mechanical state data of different modes by a target feature extractor. It integrates temporal and spatial features and can comprehensively reflect the mechanical fault attributes of transformers.
[0035] In this embodiment of the invention, mechanical state data, including one-dimensional and two-dimensional state data, is collected in real time from a sensor array deployed on the transformer under test and its key auxiliary components. The mechanical state data undergoes preprocessing operations consistent with those performed during model training, including detrending, filtering and denoising, normalization, and spatiotemporal alignment. A target feature extractor is then used to extract features from the preprocessed mechanical state data to obtain corresponding multimodal fault feature data.
[0036] Step 104: Perform state assessment on multimodal fault characteristic data through the target assessment network to obtain the corresponding assessment results, and conduct graded early warning based on the assessment results.
[0037] Tiered early warning refers to a mechanism that issues corresponding early warnings for different degrees of mechanical state anomalies in transformers based on assessment results and according to a preset multi-level early warning triggering logic, thereby achieving a gradient early warning prompt from early trend anomalies to fault confirmation.
[0038] The target assessment network refers to a network module composed of a multi-head cross-attention fusion module, an interpretable diagnostic network, and a time-series prediction model. It can perform feature fusion, fault detection, state prediction, and early warning judgment on multimodal fault feature data, and is the core module for realizing transformer state assessment and graded early warning.
[0039] In this embodiment of the invention, multimodal fault feature data is input into a target evaluation network for state evaluation (first, the multi-head cross-attention fusion module in the target evaluation network dynamically weights and fuses the multimodal fault feature data, automatically calculates the correlation weights between each modality feature, and generates an attention fusion feature vector. Then, the attention fusion feature vector is input into an interpretable diagnostic network to complete fault detection, outputs fault type confidence, and generates a visual heatmap reflecting the physical location of the fault. A multidimensional time-series feature sequence is constructed using the attention fusion feature vector, fault type confidence, and visual heatmap. The multidimensional time-series feature sequence is input into a time-series prediction model to complete the trend prediction of the transformer's mechanical state and the calculation of the health index, obtaining an evaluation result that includes state trend, health index, fault location, and type). Based on the evaluation result, the corresponding graded early warning operation is judged and executed according to the preset three-level early warning triggering logic.
[0040] In this embodiment of the invention, multiple training mechanical state data are acquired, and each training mechanical state data is preprocessed to obtain a corresponding mechanical feature set. The mechanical feature set is then used to train a preset transformer state detection model to obtain a corresponding target transformer state detection model. The target transformer state detection model includes a target feature extractor and a target evaluation network. Mechanical state data of the transformer under test is acquired, and the target feature extractor extracts features from the mechanical state data to obtain corresponding multimodal fault feature data. The target evaluation network evaluates the multimodal fault feature data to obtain corresponding evaluation results, and graded early warnings are issued based on the evaluation results. This overcomes the technical problem that existing transformer mechanical state monitoring mainly adopts a "single sensor-single parameter" monitoring mode, resulting in low confidence of diagnostic conclusions and inability to effectively locate fault positions, thus reducing the reliability of transformer mechanical fault early warning. Compared with traditional transformer mechanical condition monitoring methods, this invention uses a target evaluation network to perform state evaluation on multimodal fault characteristic data and execute hierarchical early warning. It upgrades from passive fixed threshold alarms to proactive early warning based on state time sequence evolution, realizes early fault trend perception, effectively reduces the probability of missed and false alarms, and improves the reliability of transformer mechanical fault early warning.
[0041] Please see Figure 2 , Figure 2 This is a flowchart illustrating the steps of a transformer mechanical fault early warning method provided in Embodiment 2 of the present invention.
[0042] This invention provides a method for early warning of mechanical faults in transformers, comprising: Step 201: Obtain multiple training machine state data, preprocess each training machine state data, and obtain the corresponding machine feature set.
[0043] In this embodiment of the invention, multi-source raw monitoring data covering vibration, acoustic, infrared, and electrical types are collected from the full-condition operation scenario of the transformer as training mechanical state data. This data includes monitoring information under normal transformer operation and various mechanical fault conditions. The training mechanical state data is preprocessed (i.e., detrending is performed sequentially (to eliminate deviations caused by system drift), filtering and noise reduction (to remove electromagnetic and environmental noise interference), and data normalization (to unify the numerical range of data with different dimensions). Then, the spatiotemporal alignment of the multi-source data is completed based on the timestamps of each sensor (to eliminate errors caused by asynchronous acquisition)). Finally, the data is integrated to form a mechanical feature set that retains the key characteristics of the transformer's mechanical state.
[0044] Step 202: Use the mechanical feature set to perform adversarial training on the feature extractor in the preset transformer state detection model to obtain the corresponding target feature extractor.
[0045] Further, step 202 includes the following sub-steps: S11. Input the mechanical feature set into the preset generator to obtain the corresponding noise enhancement dataset.
[0046] Noise-enhanced datasets refer to datasets obtained by superimposing target domain noise features onto a mechanical feature set after it is input into a generator. These datasets retain the core features of the transformer's mechanical state and also possess noise distribution characteristics consistent with those of actual substation sites. They are used for adversarial training to improve the model's anti-interference capabilities.
[0047] A generator is a neural network module that learns and models the noise distribution patterns of the target domain at a substation in advance, can simulate and generate noise signals that fit the actual operating scenario, and integrates them into clean data from the source domain.
[0048] In this embodiment of the invention, the mechanical feature set is input as clean source domain data into a preset generator to obtain the corresponding noise enhancement dataset. The generator has pre-learned the distribution patterns and characteristics of various noises in the target domain of the substation site. Based on these patterns, it can simulate and fuse the noise style of the input mechanical feature set. While retaining the core features of the transformer state in the mechanical feature set, it superimposes electromagnetic, mechanical, and environmental noise signals consistent with the actual operating scenario.
[0049] S12. Input the mechanical feature set and the noise enhancement dataset into the feature extractor in the preset transformer state detection model to obtain the noise feature dataset and the training state feature dataset.
[0050] The noise feature dataset refers to the dataset obtained by inputting the noise-enhanced dataset into the feature extractor and then extracting features that fuse the noise features of the target domain with the mechanical state features of the transformer, reflecting the state characteristics of the transformer with on-site noise.
[0051] The training state feature dataset refers to the core feature dataset that can truly represent the various mechanical operating states of a transformer after the mechanical feature set is input into the feature extractor and the feature is extracted, without interference from on-site noise.
[0052] In this embodiment of the invention, the mechanical feature set, which is clean data in the source domain, and the noise-enhanced dataset, which is generated by superimposing noise in the target domain, are simultaneously input into the feature extractor in the preset transformer state detection model (the two branches of the feature extractor will perform targeted feature extraction on the one-dimensional time-series signal and the two-dimensional spatial signal in the two types of data respectively, while retaining the core features of the transformer mechanical state in the data and stripping away invalid and redundant information), to obtain the noise feature dataset and the training state feature dataset.
[0053] S13. Input the noise feature dataset and the training state feature dataset into the preset domain discriminator to obtain the corresponding domain discrimination score data.
[0054] Domain discrimination score data refers to the set of discrimination probability values corresponding to various types of sample data output by the domain discriminator after it performs domain source discrimination on the input noise feature dataset and training state feature dataset. It quantitatively reflects the domain discriminator's judgment result on the domain source of each data sample.
[0055] In this embodiment of the invention, the noise feature dataset and the training state feature dataset are input into a preset domain discriminator (the domain discriminator, after initialization and training, has the ability to distinguish the source of the data domain, analyzes and judges the feature distribution of the two types of input datasets, calculates and outputs the domain discrimination probability value corresponding to each sample data, integrates these probability values to form the corresponding domain discrimination score data, thereby quantitatively representing the domain discriminator's judgment result on the source of each dataset), and the corresponding domain discrimination score data is obtained.
[0056] S14. Based on the preset adversarial loss function, determine the corresponding adversarial loss function value according to the domain discrimination score data, the noise enhancement dataset, and the mechanical feature set.
[0057] The adversarial loss function refers to the loss calculation function preset for adversarial training of the transformer state detection model. It is constructed based on the Wasserstein distance and incorporates a gradient penalty strategy. It is used to quantify the discrimination error of the domain discriminator and the noise simulation bias of the generator during adversarial training.
[0058] In this embodiment of the invention, the domain discrimination score data, the noise enhancement dataset, and the mechanical feature set are input into a preset adversarial loss function to obtain the corresponding adversarial loss function value.
[0059] It should be noted that the Wasserstein distance, a metric used to measure the similarity between two probability distributions, is used in this adversarial loss function to calculate the difference in feature distributions between the source domain (mechanical feature set) and the target domain (noise-enhanced dataset). The gradient penalty strategy, introduced into the adversarial loss function, is a constraint that adds a penalty term to the gradients of the feature extractor and the domain discriminator, preventing gradient vanishing or exploding problems during adversarial training and ensuring the stability of the training process.
[0060] It should be noted that the specific expression for the adversarial loss function is:
[0061] in, To counteract the loss function value, The loss value of the domain discriminator. The loss value of the feature extractor. The loss value is the gradient penalty term. The adaptive weighting coefficients for the gradient penalty term. E For mathematical expectation, x s These are samples from the source domain (samples from the mechanical feature set).P s For source domain data distribution, The source domain samples (mechanical feature set) follow the source domain data distribution. The target domain samples (noise-enhanced dataset) follow the target domain data distribution. x t The target domain samples (noise-enhanced samples from the noise-enhanced dataset). P t For the distribution of data in the target domain, To transform source domain samples into target domain samples, F (.) represents a feature extractor used to extract fault features from a sample. D (.) represents the domain discriminator, and its output is the domain discrimination score of the sample. These are random interpolated samples between the source domain samples and the noise-enhanced samples. The gradient of the interpolated sample is output by the domain discriminator.
[0062] S15. When the adversarial loss function value is greater than or equal to the preset adversarial loss threshold, adjust the model parameters of the feature extractor and the domain discriminator until the adversarial loss function value is less than the adversarial loss threshold.
[0063] The adversarial loss threshold refers to a pre-set critical value for the adversarial loss function, which is a quantitative standard for judging whether adversarial training has converged.
[0064] Model parameters refer to the trainable network parameters in the feature extractor and domain discriminator, mainly including the weights and biases of each network layer. They are the core elements that determine the model's feature extraction and domain discrimination capabilities.
[0065] In this embodiment of the invention, when the adversarial loss function value is greater than or equal to the preset adversarial loss threshold, it indicates that the feature extractor has not yet learned the domain-invariant fault essence features and the domain discriminator can still effectively distinguish the source domain and target domain feature data. Then, the gradient descent method is used to adjust the model parameters of the feature extractor and the domain discriminator respectively, and the process jumps to the step of inputting the mechanical feature set into the preset generator to obtain the corresponding noise enhancement dataset, until the adversarial loss function value is less than the adversarial loss threshold.
[0066] It should be noted that gradient descent is a classic machine learning parameter optimization algorithm. It calculates the gradient of the loss function with respect to the model parameters and iteratively adjusts the parameters along the negative gradient direction to minimize the value of the loss function.
[0067] S16. When the adversarial loss function is less than the adversarial loss threshold, the corresponding target feature extractor is generated.
[0068] The target feature extractor refers to the feature extraction model obtained after convergence through adversarial training.
[0069] In this embodiment of the invention, when the adversarial loss function value is less than the preset adversarial loss threshold, it indicates that the domain discriminator can no longer effectively distinguish the source domain and target domain feature data output by the feature extractor. The feature extractor has successfully learned the domain-invariant fault features that are separated from the noise interference at the substation site and are only related to the mechanical state of the transformer. At this time, the feature extractor that has completed adversarial training and parameter optimization is solidified and saved to generate the corresponding target feature extractor.
[0070] Step 203: Input the mechanical feature set into the target feature extractor to obtain the corresponding training modality feature dataset.
[0071] The training modal feature dataset refers to the dataset obtained by inputting the mechanical feature set into the target feature extractor and then extracting and integrating the features. It integrates the temporal and spatial features of the transformer's mechanical state and can represent various mechanical operating states of the transformer in multiple dimensions. It is used to evaluate the training of the network.
[0072] In this embodiment of the invention, the mechanical feature set is input into the target feature extractor to obtain the corresponding training modal feature dataset. The target feature extractor uses two branches of a two-stream network architecture to perform targeted feature extraction on the one-dimensional temporal signal and two-dimensional spatial signal in the mechanical feature set, respectively. This accurately extracts effective features that characterize the mechanical state and fault attributes of the transformer, removes redundant and invalid information, and then integrates the temporal and spatial features extracted by the two branches to obtain the corresponding training modal feature dataset.
[0073] Step 204: Input the training modal feature dataset into the evaluation network of the transformer state detection model to obtain the corresponding training evaluation data.
[0074] Training evaluation data refers to the comprehensive data set output by the evaluation network after the training modality feature dataset is input into the evaluation network. This data set includes fault type confidence, fault location virtual information, state trend prediction value, comprehensive health index, etc., and reflects the current training effect of the evaluation network.
[0075] In this embodiment of the invention, the training modal feature dataset is input into the untrained evaluation network of the transformer state detection model to obtain the corresponding training evaluation data.
[0076] Step 205: Based on the preset loss function, calculate the training loss function value of the mechanical feature set according to the training evaluation data.
[0077] In this embodiment of the invention, training evaluation data and mechanical feature set are input into a preset loss function to obtain the corresponding training loss function value.
[0078] It should be noted that the loss function is as follows:
[0079] in, To train the loss function values, The first adaptive weighting coefficient, For fault classification loss values, The second adaptive weighting coefficient, To determine the loss value for fault location, The third adaptive weighting coefficient, This is the time series prediction loss value. The number of training samples. C This represents the total number of categories of mechanical faults in the transformer. i For sample index, c For fault category index, For the sample i Fault Category c The true label, To evaluate the network's performance on samples i Fault category c The prediction confidence level for, for, H The high value of the heat map for fault location W The width of the fault location heatmap, For pixels ( h , w The actual fault location label value at () indicates that the fault core area is 1 and the non-fault area is 0. To evaluate the heatmap generated by the network at the pixel level ( h , w The predicted intensity value at () h The height coordinates of the pixel. w The width coordinates of the pixel. The number of feature dimensions in the time-series feature sequence. j Indexed by feature dimensions, For the first j The actual value of the dimensional feature. To evaluate the network's performance on the first j The predicted value of the dimensional feature.
[0080] Step 206: When the training loss function value is greater than or equal to the preset loss threshold, adjust the network parameters of the evaluation network until the training loss function value is less than the loss threshold.
[0081] Network parameters refer to all trainable core parameters in the evaluation network, mainly including the connection weights and bias values of each module and layer of the network. They are key elements that determine the network's ability to process features and evaluate its state.
[0082] The loss threshold refers to a pre-set critical standard for the training loss function value. It is a quantitative indicator for determining whether the evaluation network has completed effective training. When the training loss function value is lower than this threshold, the performance of the evaluation network is considered to meet the requirements of practical applications.
[0083] In this embodiment of the invention, when the training loss function value is greater than or equal to a preset loss threshold, it indicates that the current fault diagnosis, localization, and state prediction performance of the evaluation network has not met the preset training standard. At this time, gradient descent combined with backpropagation is used to perform gradient calculations layer by layer on all trainable network parameters, such as the network weights and biases of the multi-head cross-attention fusion module, the interpretable diagnostic network, and the temporal prediction model, based on the training loss function value. The parameters of the corresponding modules are adjusted differentially according to the errors of different tasks, such as fault classification, localization, and temporal prediction. Step 204 is executed for adjustment until the training loss function value is less than the loss threshold.
[0084] It should be noted that the backpropagation mechanism is the process of passing the error value of the training loss function backward from the output layer of the evaluation network to the input layer. It can calculate the contribution of each network parameter to the overall error layer by layer, providing a gradient basis for precise adjustment of the parameters.
[0085] Step 207: When the training loss function value is less than the loss threshold, the corresponding target evaluation network is generated. The target transformer state detection model includes a target feature extractor and a target evaluation network.
[0086] In this embodiment of the invention, when the training loss function value is less than the loss threshold, it indicates that the evaluation network has reached the preset performance standard in multi-task training of fault classification, fault location, and time series prediction, and the evaluation accuracy of the transformer mechanical state meets the requirements of practical applications. The evaluation network after parameter iterative optimization is then fixed with parameters and the model is saved to generate the corresponding target evaluation network.
[0087] Step 208: Obtain the mechanical state data of the transformer under test, and extract features from the mechanical state data using a target feature extractor to obtain the corresponding multimodal fault feature data.
[0088] Furthermore, the target feature extractor includes a first feature extraction branch and a second feature extraction branch, the mechanical state data includes one-dimensional state data and two-dimensional state data, and step 208 includes the following sub-steps: S21. Temporal features are extracted from the one-dimensional state data through the first feature extraction branch to obtain the corresponding temporal feature data.
[0089] The first feature extraction branch refers to the network branch in the target feature extractor that is specifically designed to process one-dimensional state data. It is constructed using a deep residual shrinking network or a one-dimensional convolutional neural network to meet the feature extraction requirements of one-dimensional time series signals.
[0090] One-dimensional state data refers to a one-dimensional time-series signal that reflects the mechanical state of a transformer, including transformer tank vibration signals, acoustic signals, and electrical signals such as load current.
[0091] Time-series feature data refers to the feature dataset obtained after extracting one-dimensional state data through the first feature extraction branch, which contains core features that can reflect the time-series changes in the mechanical state of the transformer.
[0092] In this embodiment of the invention, the first feature extraction branch extracts time-series features from the one-dimensional state data to obtain the corresponding time-series feature data. The first feature extraction branch is built using a deep residual shrinking network or a one-dimensional convolutional neural network. Differentiated network designs are made for the physical characteristics of one-dimensional time-series signals such as vibration, acoustics, and electrical signals. It can accurately capture the time-series fluctuation patterns and frequency domain distribution characteristics in the data. Through the network's multi-layer convolution, residual connections, and feature shrinking operations, redundant interference information in the data is removed, and the core time-series features that can characterize the mechanical state of the transformer are mined out. Finally, the corresponding time-series feature data is integrated.
[0093] It should be noted that the Deep Residual Shrinking Network (DRSN), a neural network combining deep residual networks and feature shrinkage mechanisms, can effectively process noisy time-series signals, achieving adaptive feature selection and removal of redundant information. One-dimensional convolutional neural networks (1D-CNNs) capture local features and sequence correlations of time-series signals through sliding calculations of one-dimensional convolutional kernels along the time dimension.
[0094] S22. Spatial features are extracted from the two-dimensional state data through the second feature extraction branch to obtain the corresponding spatial feature data.
[0095] The second feature extraction branch refers to the network branch in the target feature extractor that is specifically used to process two-dimensional state data. It is constructed using deep residual networks (ResNet) or two-dimensional convolutional neural networks (2D-CNN) to meet the feature extraction requirements of two-dimensional spatial image signals.
[0096] Two-dimensional state data refers to two-dimensional spatial image data that reflects the mechanical state of a transformer. It mainly consists of infrared thermal image video sequences of the transformer body and radiator, containing spatial state information such as temperature and texture.
[0097] Spatial feature data refers to the feature dataset obtained after extracting two-dimensional state data through the second feature extraction branch, which contains core features that can reflect the spatial distribution law of transformer mechanical state.
[0098] In this embodiment of the invention, spatial features are extracted from the two-dimensional state data through a second feature extraction branch. The second feature extraction branch is built using a deep residual network or a two-dimensional convolutional neural network. Differentiated network designs are made for the physical characteristics of two-dimensional spatial image signals such as infrared thermal imaging video sequences. It can accurately capture the spatial temperature field hotspot distribution, texture changes and regional correlation features in the two-dimensional state data. Through multi-layer two-dimensional convolution, residual connection and pooling operations of the network, background interference information in the two-dimensional state data is removed, and the core spatial features that can characterize the mechanical state of the transformer are mined out. Finally, the corresponding spatial feature data is integrated.
[0099] It should be noted that 2D convolutional neural networks (2D-CNN) are convolutional neural networks designed for 2D image data. They capture local spatial features and region correlation features of images by sliding two-dimensional convolutional kernels across the width and height dimensions of the image.
[0100] S23. Use time-series feature data and spatial feature data as multimodal fault feature data.
[0101] In this embodiment of the invention, the temporal feature data extracted by the first feature extraction branch and the spatial feature data extracted by the second feature extraction branch are subjected to dimension matching and data integration. The core information that can characterize the mechanical fault of the transformer in the two types of feature data is retained, and redundant and invalid feature dimensions are discarded to form multimodal fault feature data that integrates the temporal change law and spatial distribution characteristics of the transformer's mechanical state.
[0102] Step 209: Perform state assessment on the multimodal fault characteristic data through the target assessment network to obtain the corresponding assessment results, and conduct graded early warning based on the assessment results.
[0103] Furthermore, the target evaluation network includes a multi-head cross-attention fusion module, an interpretable diagnostic network, and a temporal prediction model. Step 209 includes the following sub-steps: S31. The multi-head cross-attention fusion module is used to perform feature fusion on the multi-modal fault feature data to obtain the corresponding attention fusion feature vector.
[0104] The multi-head cross-attention fusion module refers to a module designed based on the self-attention mechanism that can perform correlation analysis and dynamic weighted fusion of multimodal heterogeneous fault feature data, realize information complementarity of different modal features, and is a key module for multimodal deep fusion.
[0105] Attention fusion feature vector refers to a high-dimensional feature vector generated by dynamic weighted fusion through a multi-head cross-attention fusion module. It integrates the core effective information of multimodal fault features and can accurately and comprehensively characterize the mechanical fault state and attributes of transformers.
[0106] In this embodiment of the invention, a multi-head cross-attention fusion module is used to perform feature fusion on multimodal fault feature data. Based on the self-attention mechanism, the multi-head cross-attention fusion module performs feature correlation analysis on the temporal feature data corresponding to vibration, acoustics, and electrical and the spatial feature data corresponding to infrared thermography. It automatically calculates the dynamic correlation weight between different modal features, assigns high weight to key features characterizing transformer mechanical faults and low weight to redundant interference features, thereby achieving adaptive complementarity and enhancement of multimodal features. Then, through feature splicing and dimension mapping, the deep fusion of heterogeneous features is completed to obtain the corresponding attention fusion feature vector.
[0107] It should be noted that self-attention mechanism is an algorithm that assigns different weights to different parts of the input features. It can automatically identify and highlight features that are important to the task, suppress irrelevant and redundant features, and improve the effectiveness of feature fusion. Dynamic association weights refer to the weight values calculated in real time by the multi-head cross-attention fusion module based on the degree of association between multimodal features. They can adaptively adjust as the feature data changes, achieving focused attention on key fault features. Heterogeneous features refer to feature data with different sources, types, and dimensions. In this invention, they specifically refer to fault features with different attributes, such as one-dimensional temporal feature data and two-dimensional spatial feature data.
[0108] S32. Fault detection is performed on the attention fusion feature vector through an interpretable diagnostic network to obtain the corresponding fault type confidence and visual heatmap.
[0109] Furthermore, the interpretable diagnostic network includes a convolutional feature extraction layer, a fault localization branch, and a fault classification branch. S32 includes the following sub-steps: S321. The attention fusion feature vector is extracted by the convolutional feature extraction layer to obtain the corresponding attention fusion feature map and the pre-convolutional feature map.
[0110] The convolutional feature extraction layer refers to the core front-end module of the interpretable diagnostic network. It consists of multiple convolutional layers, batch normalization layers, and activation function layers stacked together. It is used to perform multi-scale, deep convolutional feature extraction on attention fusion feature vectors, providing a feature basis for fault classification and localization.
[0111] Attention fusion feature map refers to a two-dimensional feature map obtained by performing deep feature mining on the previous convolutional feature map by subsequent convolutional layers after the convolutional feature extraction layer. It contains deep correlation feature information of transformer faults.
[0112] The pre-convolutional feature map refers to the two-dimensional feature map obtained after the pre-convolutional layer of the convolutional feature extraction layer performs the first multi-scale convolution on the attention fusion feature vector, which contains the basic space and related features of transformer faults.
[0113] In this embodiment of the invention, a convolutional feature extraction layer is used to extract features from the attention fusion feature vector. The convolutional feature extraction layer consists of multiple stacked convolutional layers, batch normalization layers, and activation function layers. Differentiated designs for convolutional kernel size and stride are made to address the characteristics of high-dimensional fusion features. First, the attention fusion feature vector is subjected to multi-scale convolution and feature mapping through a pre-convolutional layer to extract a pre-convolutional feature map that can characterize the basic features of transformer faults. Then, the pre-convolutional feature map is further subjected to deep feature mining and dimensionality enhancement through subsequent convolutional layers to generate an attention fusion feature map containing richer deep fault correlation information.
[0114] It should be noted that the batch normalization layer is a network layer in the convolutional feature extraction layer used to normalize the feature data after convolution operations. This can accelerate model training convergence and improve the stability and robustness of feature extraction. The activation function layer is a network layer in the convolutional feature extraction layer used to introduce nonlinear mappings into the feature data. This can uncover nonlinear correlations between features and improve the ability of the convolutional feature extraction layer to express complex fault features.
[0115] S322. Perform fault diagnosis on the attention fusion feature map through the fault classification branch to obtain the corresponding fault category confidence. The fault classification branch includes a fully connected stacked module and a Softmax activation function connected in sequence.
[0116] Fault category confidence refers to the normalized probability value corresponding to each mechanical fault type of the transformer output by the fault classification branch. The value ranges from 0 to 1. The higher the value, the higher the confidence of the judgment result of the fault type.
[0117] A fully connected stacked module refers to a module composed of multiple layers of fully connected networks. It can perform high-dimensional feature mapping, nonlinear transformation and dimensionality compression on attention fusion feature maps, and integrate them to form a one-dimensional feature vector that can characterize the fault type.
[0118] The Softmax activation function is a function that can transform the one-dimensional feature vector output by a fully connected stacked module into probability values between 0 and 1, with the sum of all probability values being 1, thereby quantifying the confidence of each fault type.
[0119] In this embodiment of the invention, fault diagnosis is performed on the attention fusion feature map through a fault classification branch. The fault classification branch consists of a fully connected stacked module and a Softmax activation function connected in sequence. The fully connected stacked module performs depth mapping and dimensional compression of high-dimensional features on the attention fusion feature map, mines the nonlinear correlation between features and integrates them into a one-dimensional fault feature vector. Then, the Softmax activation function normalizes the feature vector and calculates the probability value corresponding to each mechanical fault type of the transformer. These probability values are used as the confidence scores of the corresponding fault categories.
[0120] S323. Fault localization is performed on the pre-convolutional feature map by fault localization branch and fault category confidence to obtain the corresponding visual heat map.
[0121] The fault localization branch refers to the core sub-branch in the interpretable diagnostic network that runs parallel to the fault classification branch. It is built based on Grad-CAM technology and is used to combine fault category confidence with pre-convolutional feature maps to achieve physical spatial localization of faults and output a visual heatmap.
[0122] A visual heatmap is a visual feature map output by the fault location branch. It uses color depth or brightness to represent the probability and location of the fault and can be overlaid on the transformer infrared image / layout map to intuitively display the suspected fault area.
[0123] In this embodiment of the invention, fault localization is performed on the pre-convolutional feature map through a fault localization branch and fault category confidence. The fault localization branch is based on gradient weighted class activation mapping technology. First, the gradient information of the target fault category corresponding to the fault category confidence is calculated and global average pooling is performed on it to obtain the feature weights corresponding to each feature layer. Then, the feature weights are weighted and summed with the pre-convolutional feature map. After superimposing the ReLU activation function to remove invalid negative feature values, an initial heat map is generated. After upsampling, the initial heat map is mapped to a size that matches the transformer infrared image or layout map to obtain the corresponding visual heat map.
[0124] It should be noted that gradient-weighted class activation mapping refers to calculating the gradient information of the target fault category on the convolutional feature map, thereby mining the regions in the feature map that contribute to fault determination and realizing a visual interpretation of fault localization. Upsampling refers to the operation of enlarging the dimensionality of the initial small-sized heatmap so that the size of the heatmap matches the transformer infrared image or layout map.
[0125] S33. Construct the corresponding multidimensional temporal feature sequence by using attention fusion feature vectors, fault type confidence and visual heatmap.
[0126] Multidimensional temporal feature sequence refers to the sequence data formed by continuously splicing together core quantitative indicators such as attention fusion feature vector, fault type confidence, and visual heatmap along the time dimension.
[0127] In this embodiment of the invention, the attention fusion feature vector, fault type confidence and visual heat map are quantified and aligned with the time dimension. The key health indicators, real-time values of fault type confidence and intensity values of key areas of visual heat map are extracted from the attention fusion feature vector. These core quantitative indicators that can characterize the mechanical state of the transformer are continuously collected and spliced according to the time series to form a multi-dimensional time series feature sequence that can comprehensively reflect the time series evolution law of the transformer state.
[0128] It should be noted that the intensity value of the key area in the visual heatmap refers to the pixel intensity quantization value of the suspected core area of the fault in the visual heatmap. The value reflects the probability and severity of the fault in that area.
[0129] S34. The state of multidimensional time-series feature sequences is predicted by a time-series prediction model to obtain the corresponding evaluation results, and graded early warning is carried out based on the evaluation results.
[0130] Temporal prediction models refer to models constructed using temporal convolutional networks (TCNs) or long short-term memory networks (LSTMs), specifically designed to uncover the temporal correlation patterns of multidimensional temporal feature sequences, thereby enabling trend prediction and health index calculation of transformer status indicators.
[0131] The evaluation result refers to the comprehensive data set output by the time-series prediction model, which includes the predicted values of various state indicators of the transformer and the comprehensive health index. It is the core basis for judging the mechanical condition of the transformer and triggering graded early warning.
[0132] In this embodiment of the invention, a multi-dimensional time-series feature sequence is input into a time-series prediction model for state prediction. The time-series prediction model is built using a time-series convolutional network or a long short-term memory network. A targeted network design is made for the time-series evolution characteristics of transformer state indicators, which can accurately capture the time correlation rules and state evolution trends of each indicator in the multi-dimensional time-series feature sequence, obtain the corresponding evaluation results, and perform graded early warning based on the evaluation results.
[0133] It is worth mentioning that the graded early warning refers to a three-level progressive intelligent early warning mechanism built based on the evaluation results output by the time series prediction model. This mechanism abandons the traditional single alarm method with a fixed threshold, and combines real-time monitoring values of multi-dimensional time series feature sequences, model prediction values, comprehensive health index and fault category confidence, visual heat map intensity changes and other multi-dimensional evaluation results for comprehensive judgment. According to the time evolution law of fault development, it divides the early warning level into three levels: attention level, abnormal level and severe level. The triggering conditions and handling suggestions of each level are progressive, realizing proactive early warning throughout the entire life cycle from early trend abnormality perception to fault confirmation. Specifically, when the deviation between the real-time value of the multi-dimensional time series feature sequence and the predicted value output by the time series prediction model continuously and significantly exceeds the preset deviation threshold. When the value is low, a warning level is triggered, indicating that the transformer has an abnormal early state trend, and equipment status monitoring and data tracking analysis need to be strengthened. When the comprehensive health index calculated based on multi-dimensional time series feature sequences falls below the preset adaptive dynamic threshold, an abnormal level warning is triggered, indicating that the transformer equipment status has undergone substantial degradation and the fault risk has significantly increased, requiring targeted equipment inspection and status retesting. When the fault category confidence exceeds the preset high-risk threshold, or the intensity value of the fault key area in the visual heat map shows a continuous and obvious deterioration trend, a severe level warning is triggered, indicating that the transformer has been confirmed to have a specific type of mechanical fault and the fault location is clear. The fault has developed to the stage requiring emergency handling, and the equipment needs to be shut down immediately and fault repair work needs to be carried out.
[0134] It should be noted that Temporal Convolutional Networks (TCNs) refer to convolutional neural networks designed for time-series data. They capture long-range dependencies in time-series sequences through causal convolution and dilated convolution, adapting to the prediction needs of transformer state time-series features. Long Short-Term Memory Networks (LSTMs) refer to an improved recurrent neural network that can effectively solve the gradient vanishing problem in long sequence training and accurately mine long-term temporal correlation features in time-series data.
[0135] In this embodiment of the invention, multiple training mechanical state data are acquired, and each training mechanical state data is preprocessed to obtain a corresponding mechanical feature set. The mechanical feature set is then used to train a preset transformer state detection model to obtain a corresponding target transformer state detection model. The target transformer state detection model includes a target feature extractor and a target evaluation network. Mechanical state data of the transformer under test is acquired, and the target feature extractor extracts features from the mechanical state data to obtain corresponding multimodal fault feature data. The target evaluation network evaluates the multimodal fault feature data to obtain corresponding evaluation results, and graded early warnings are issued based on the evaluation results. This overcomes the technical problem that existing transformer mechanical state monitoring mainly adopts a "single sensor-single parameter" monitoring mode, resulting in low confidence of diagnostic conclusions and inability to effectively locate fault positions, thus reducing the reliability of transformer mechanical fault early warning. Compared with traditional transformer mechanical condition monitoring methods, this invention uses a target evaluation network to perform state evaluation on multimodal fault characteristic data and execute hierarchical early warning. It upgrades from passive fixed threshold alarms to proactive early warning based on state time sequence evolution, realizes early fault trend perception, effectively reduces the probability of missed and false alarms, and improves the reliability of transformer mechanical fault early warning.
[0136] Please see Figure 3 , Figure 3 This is a structural block diagram of a transformer mechanical fault early warning system provided in Embodiment 3 of the present invention.
[0137] This invention provides a transformer mechanical fault early warning system, comprising: The preprocessing module 301 is used to acquire multiple training machine state data, preprocess each training machine state data, and obtain the corresponding machine feature set.
[0138] The training module 302 is used to train a preset transformer state detection model using a mechanical feature set to obtain a corresponding target transformer state detection model. The target transformer state detection model includes a target feature extractor and a target evaluation network.
[0139] The extraction module 303 is used to acquire the mechanical state data of the transformer under test, and to extract features from the mechanical state data through the target feature extractor to obtain the corresponding multimodal fault feature data.
[0140] The evaluation module 304 is used to perform state evaluation on multimodal fault characteristic data through the target evaluation network, obtain the corresponding evaluation results, and perform graded early warning based on the evaluation results.
[0141] Furthermore, training module 302 includes: The adversarial training submodule is used to perform adversarial training on the feature extractor in the preset transformer state detection model using the mechanical feature set, so as to obtain the corresponding target feature extractor.
[0142] The training submodule is used to input the mechanical feature set into the target feature extractor to obtain the corresponding training modality feature dataset.
[0143] The training modal feature dataset is input into the evaluation network of the transformer state detection model to obtain the corresponding training and evaluation data.
[0144] The loss submodule is used to calculate the training loss function value of the mechanical feature set based on the training evaluation data and a preset loss function.
[0145] The adjustment submodule is used to adjust the network parameters of the evaluation network when the training loss function value is greater than or equal to a preset loss threshold, until the training loss function value is less than the loss threshold.
[0146] When the training loss function value is less than the loss threshold, the corresponding target evaluation network is generated.
[0147] Furthermore, the adversarial training submodule includes: The enhancement unit is used to input the mechanical feature set into a preset generator to obtain the corresponding noise enhancement dataset.
[0148] The extraction unit is used to input the mechanical feature set and the noise enhancement dataset into the feature extractor in the preset transformer state detection model, respectively, to obtain the noise feature dataset and the training state feature dataset.
[0149] The adversarial training unit is used to input the noise feature dataset and the training state feature dataset into a preset domain discriminator to obtain the corresponding domain discrimination score data.
[0150] Based on the preset adversarial loss function, the corresponding adversarial loss function value is determined according to the domain discrimination score data, the noise enhancement dataset, and the mechanical feature set.
[0151] When the adversarial loss function value is greater than or equal to the preset adversarial loss threshold, the model parameters of the feature extractor and the domain discriminator are adjusted until the adversarial loss function value is less than the adversarial loss threshold.
[0152] When the adversarial loss function is less than the adversarial loss threshold, the corresponding target feature extractor is generated.
[0153] Furthermore, the target feature extractor includes a first feature extraction branch and a second feature extraction branch, the mechanical state data includes one-dimensional state data and two-dimensional state data, and the extraction module 303 includes: The first extraction submodule is used to extract time-series features from one-dimensional state data through the first feature extraction branch to obtain the corresponding time-series feature data.
[0154] The second extraction submodule is used to extract spatial features from the two-dimensional state data through the second feature extraction branch to obtain the corresponding spatial feature data.
[0155] Temporal and spatial feature data are used as multimodal fault feature data.
[0156] Furthermore, the target evaluation network includes a multi-head cross-attention fusion module, an interpretable diagnostic network, and a temporal prediction model. Evaluation module 304 includes: The feature fusion submodule is used to perform feature fusion on multimodal fault feature data through the multi-head cross-attention fusion module to obtain the corresponding attention fusion feature vector.
[0157] The fault detection submodule is used to detect faults by using an interpretable diagnostic network to analyze the attention fusion feature vectors, and obtain the corresponding fault type confidence and visual heatmap.
[0158] The prediction submodule is used to construct the corresponding multi-dimensional temporal feature sequence by using attention-fused feature vectors, fault type confidence and visual heatmap.
[0159] The state of a multidimensional time-series feature sequence is predicted using a time-series prediction model, and the corresponding evaluation results are obtained.
[0160] Furthermore, the interpretable diagnostic network includes a convolutional feature extraction layer, a fault localization branch, and a fault classification branch. The fault detection submodule includes: The pre-convolutional unit is used to extract features from the attention fusion feature vector through the convolutional feature extraction layer, so as to obtain the corresponding attention fusion feature map and the pre-convolutional feature map.
[0161] The fault diagnosis unit is used to diagnose faults in the attention fusion feature map through the fault classification branch and obtain the corresponding fault category confidence. The fault classification branch includes a fully connected stacked module and a Softmax activation function connected in sequence.
[0162] The localization unit is used to locate faults in the preceding convolutional feature map by using the fault localization branch and the fault category confidence, and obtain the corresponding visual heatmap.
[0163] Please see Figure 4 , Figure 4 This is a structural block diagram of an electronic device provided in Embodiment 4 of the present invention.
[0164] An electronic device according to an embodiment of the present invention includes: a memory 401 and a processor 402. The memory 401 stores a computer program. When the computer program is executed by the processor 402, the processor 402 executes the transformer mechanical fault early warning method as described in any of the above embodiments.
[0165] Memory 401 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Memory 401 has storage space 403 for program code 413 for performing any of the method steps described above. For example, storage space 403 for program code may include individual program codes 413 for implementing the various steps in the methods described above. This program code may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When run by a computing processing device, this code causes the computing processing device to perform the various steps in the methods described above. This program code may be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The program code may be compressed, for example, in a suitable form. When these codes are run by a computing device, they cause the computing device to perform the various steps in the transformer mechanical fault early warning method described above.
[0166] Embodiment 5 of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the transformer mechanical fault early warning method as described in any of the above embodiments.
[0167] Embodiment 6 of the present invention also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, wherein when the program instructions are executed by a computer, the computer performs the transformer mechanical fault early warning method as described in any of the above embodiments.
[0168] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0169] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0170] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0171] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0172] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0173] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for early warning of mechanical faults in transformers, characterized in that, include: Multiple training machine state data are acquired, and each training machine state data is preprocessed to obtain a corresponding machine feature set; The preset transformer state detection model is trained using the mechanical feature set to obtain the corresponding target transformer state detection model, wherein the target transformer state detection model includes a target feature extractor and a target evaluation network; The mechanical state data of the transformer under test is acquired, and the target feature extractor is used to extract features from the mechanical state data to obtain the corresponding multimodal fault feature data. The target evaluation network is used to evaluate the state of the multimodal fault feature data to obtain the corresponding evaluation results, and graded early warning is given based on the evaluation results.
2. The transformer mechanical fault early warning method according to claim 1, characterized in that, The step of training a preset transformer state detection model using the mechanical feature set to obtain a corresponding target transformer state detection model includes: The mechanical feature set is used to perform adversarial training on the feature extractor in the preset transformer state detection model to obtain the corresponding target feature extractor; The mechanical feature set is input into the target feature extractor to obtain the corresponding training modality feature dataset; The training modal feature dataset is input into the evaluation network of the transformer state detection model to obtain the corresponding training evaluation data; Based on a preset loss function, the training loss function value of the mechanical feature set is calculated according to the training evaluation data; When the training loss function value is greater than or equal to a preset loss threshold, the network parameters of the evaluation network are adjusted until the training loss function value is less than the loss threshold. When the training loss function value is less than the loss threshold, a corresponding target evaluation network is generated.
3. The transformer mechanical fault early warning method according to claim 2, characterized in that, The step of using the mechanical feature set to perform adversarial training on the feature extractor in the preset transformer state detection model to obtain the corresponding target feature extractor includes: The mechanical feature set is input into a preset generator to obtain the corresponding noise enhancement dataset; The mechanical feature set and the noise enhancement dataset are respectively input into the feature extractor in the preset transformer state detection model to obtain the noise feature dataset and the training state feature dataset; The noise feature dataset and the training state feature dataset are input into a preset domain discriminator to obtain the corresponding domain discrimination score data; Based on the preset adversarial loss function, the corresponding adversarial loss function value is determined according to the domain discrimination score data, the noise enhancement dataset, and the mechanical feature set; When the adversarial loss function value is greater than or equal to the preset adversarial loss threshold, the model parameters of the feature extractor and the domain discriminator are adjusted until the adversarial loss function value is less than the adversarial loss threshold. When the adversarial loss function is less than the adversarial loss threshold, a corresponding target feature extractor is generated.
4. The transformer mechanical fault early warning method according to claim 1, characterized in that, The target feature extractor includes a first feature extraction branch and a second feature extraction branch. The mechanical state data includes one-dimensional state data and two-dimensional state data. The step of extracting features from the mechanical state data using the target feature extractor to obtain corresponding multimodal fault feature data includes: Temporal features are extracted from the one-dimensional state data through the first feature extraction branch to obtain the corresponding temporal feature data. Spatial feature extraction is performed on the two-dimensional state data through the second feature extraction branch to obtain the corresponding spatial feature data; The temporal feature data and the spatial feature data are used as multimodal fault feature data.
5. The transformer mechanical fault early warning method according to claim 1, characterized in that, The target evaluation network includes a multi-head cross-attention fusion module, an interpretable diagnostic network, and a temporal prediction model. The step of performing state evaluation on the multimodal fault feature data through the target evaluation network to obtain the corresponding evaluation results includes: The multi-head cross-attention fusion module performs feature fusion on the multimodal fault feature data to obtain the corresponding attention fusion feature vector; The attention fusion feature vector is used to perform fault detection through the interpretable diagnostic network to obtain the corresponding fault type confidence and visual heatmap. The attention fusion feature vector, the fault type confidence, and the visual heatmap are used to construct a corresponding multidimensional temporal feature sequence; The state prediction of the multidimensional time-series feature sequence is performed using the time-series prediction model to obtain the corresponding evaluation results.
6. The transformer mechanical fault early warning method according to claim 5, characterized in that, The interpretable diagnostic network includes a convolutional feature extraction layer, a fault localization branch, and a fault classification branch. The step of performing fault detection on the attention fusion feature vector through the interpretable diagnostic network to obtain the corresponding fault type confidence and visual heatmap includes: The attention fusion feature vector is extracted through the convolutional feature extraction layer to obtain the corresponding attention fusion feature map and the pre-convolutional feature map; The attention fusion feature map is used to diagnose faults through the fault classification branch to obtain the corresponding fault category confidence. The fault classification branch includes a fully connected stacked module and a Softmax activation function connected in sequence. The fault location is performed on the pre-convolutional feature map using the fault location branch and the fault category confidence, resulting in a corresponding visual heatmap.
7. A transformer mechanical fault early warning system, characterized in that, include: The preprocessing module is used to acquire multiple training machine state data, preprocess each training machine state data, and obtain the corresponding machine feature set. The training module is used to train the preset transformer state detection model using the mechanical feature set to obtain the corresponding target transformer state detection model, wherein the target transformer state detection model includes a target feature extractor and a target evaluation network; The extraction module is used to acquire the mechanical state data of the transformer under test, and to extract features from the mechanical state data through the target feature extractor to obtain the corresponding multimodal fault feature data. The evaluation module is used to perform state evaluation on the multimodal fault feature data through the target evaluation network, obtain the corresponding evaluation results, and perform graded early warning based on the evaluation results.
8. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the transformer mechanical fault early warning method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the transformer mechanical fault early warning method as described in any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, wherein when the program instructions are executed by a computer, the computer performs the transformer mechanical fault early warning method as described in any one of claims 1-6.