A main transformer equipment anomaly detection method, system, device and storage medium
Patent Information
- Application Number
- CN202610857046.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-25
AI Technical Summary
在复杂的实际运行工况下,主变设备及其附属设施易发生两类典型缺陷:一类是表观状态异常,如阀门及密封部位的渗漏油、绝缘子(如套管绝缘子、支柱绝缘子)表面的严重污秽与破损、金属连接件(包括各类金具、销钉、螺栓等)的松脱与锈蚀等
[0016]本发明提供了一种主变设备异常检测方法、系统、设备和存储介质,通过基于跨阶段聚合的双流感知架构和基于信息回流机制的跨层级语义增强交互,能够确保主变设备微小裂纹、螺栓异常等细微几何特征在深度网络中不被丢失,提升对小尺度缺陷的定位精度;通过跨模态门控对齐融合,能够深度挖掘红外热异常与可见光表观异常之间的内在联系,降低误报率;通过跨模态一致性损失函数,约束红外热异常与可见光表观异常的物理语义关联,能够提升模型的鲁棒性,本发明在保证高精度的同时,通过优化的特征聚合结构实现了模型轻量化与检测速度的平衡,适用于变电站边缘侧巡检终端的实时监测任务。
Smart Images

Figure CN122821201A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of main transformer equipment anomaly detection technology, and in particular to a main transformer equipment anomaly detection method, system, device and storage medium. Background Technology
[0002] With the rapid development of power systems and the increasing demands for the safety of power grid equipment operation, real-time and all-weather monitoring of substation equipment has become a crucial task for ensuring the stable operation of power systems. As the core hub equipment of the power grid, main transformer equipment is exposed to an open outdoor environment for extended periods. Under complex actual operating conditions, main transformer equipment and its auxiliary facilities are prone to two typical types of defects: one is apparent abnormalities, such as oil leakage from valves and seals, severe contamination and damage to the surface of insulators (such as bushing insulators and post insulators), and loosening and corrosion of metal connectors (including various hardware, pins, bolts, etc.). The other is hidden thermal defects, such as localized overheating caused by poor contact at lead wire joints, and abnormal temperature rises caused by obstructed cooling systems or internal faults.
[0003] In current automated inspections, traditional single-vision sensors are insufficient to meet the demands for all-weather, high-reliability monitoring. While visible light (RGB) images can provide high-resolution texture details and accurately capture surface anomalies such as oil leaks, insulator damage, and minor pin and bolt abnormalities, their imaging is highly susceptible to external environmental constraints. Especially in low-light conditions at night, and in adverse weather conditions such as heavy fog or haze, visible light images often suffer from severe visibility reduction, feature loss, and background clutter interference. In contrast, infrared (IR) thermal imaging does not rely on external light sources, can penetrate darkness, and to some extent overcome the effects of haze, directly reflecting the thermal field distribution on the equipment surface, making it a powerful tool for detecting thermal defects in main transformer equipment. However, infrared images typically have lower resolution and lack geometric topology and texture details of the equipment, resulting in a lack of physical reference when locating specific heat-generating components.
[0004] Existing heterogeneous image fusion detection technologies also have significant limitations in actual substation operation and maintenance scenarios. First, due to interference from strong light reflection, alternating shadows, and high-voltage electromagnetic radiation, the acquired RGB images are prone to artifacts, while IR images are often accompanied by random noise and thermal diffusion effects. Second, RGB and IR are heterogeneous modes, and existing fusion algorithms mostly remain at the level of simple data layer stacking, lacking deep dynamic alignment of the physical semantics between the two modal features. This makes the model prone to feature conflicts when dealing with the loss of single-modal information caused by weather interference, causing minor defects such as small oil leaks or initial thermal anomalies to be swallowed by redundant background. In addition, traditional deep networks lack the maintenance of equipment geometry and topology in multi-level downsampling, making it difficult to meet the high-precision and robust monitoring requirements of industrial-grade substations. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method, system, device, and storage medium for detecting anomalies in main transformer equipment, thereby improving the accuracy and robustness of defect detection in main transformer equipment and meeting the technical requirements for industrial-grade real-time monitoring.
[0006] In a first aspect, the present invention provides a method for detecting anomalies in main transformer equipment, the method comprising: Image data of the target main transformer equipment is acquired and preprocessed. The image data includes visible light images and infrared images. The preprocessed image data is input into a preset anomaly detection model to obtain the anomaly detection result of the target main transformer equipment. The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head. The loss function of the anomaly detection model consists of a detection loss and a cross-modal consistency loss, wherein the cross-modal consistency loss is constructed based on the distance between the same defect target in the heterogeneous feature space. The semantic interaction module includes three parallel interaction sub-modules. Each interaction sub-module consists of multiple feature splicing layers and multiple feature fusion layers connected alternately. The three interaction sub-modules perform cross-scale interactive fusion of the multi-scale fusion features output by the gated fusion module through an information feedback mechanism, and output multi-scale interactive features.
[0007] Furthermore, the step of preprocessing the image data includes: The illumination component in the visible light image is removed by using the Retinex decomposition technique to obtain the preprocessed visible light image; Anisotropic diffusion filtering is used to smooth the infrared image, resulting in a preprocessed infrared image.
[0008] Furthermore, the dual-branch feature extraction module includes two parallel branch sub-modules. Each branch sub-module consists of three cascaded convolutional layers, a max pooling layer, four cross-stage aggregation modules, and an attention-based intra-scale feature interaction module. The cross-stage aggregation module is a multi-layer residual structure, consisting of a segmentation layer, a 3×3 convolutional layer, an efficient channel attention module, a feature concatenation layer, and a 1×1 convolutional layer. The first branch submodule is used to extract multi-scale features from the preprocessed visible light image to obtain the multi-scale features of the visible light image, and the second branch submodule is used to extract multi-scale features from the preprocessed infrared image to obtain the multi-scale features of the infrared image. The multi-scale features output by each branch submodule include shallow features, mid-level features, and deep features. The shallow features are the features output by the second cross-stage aggregation module, the mid-level features are the features output by the third cross-stage aggregation module, and the deep features are the features output by the adaptive feature fusion module.
[0009] Furthermore, the gated fusion module includes three parallel dual-modal unified representation modules, which are used to receive multi-scale features of visible light images and infrared images output by the dual-branch feature extraction module, and perform feature fusion using a gating mechanism to obtain multi-scale fusion features, which include shallow fusion features, medium fusion features and deep fusion features. Each dual-modal unified characterization module is used to receive the features of the visible light image and infrared image of the corresponding scale output by the dual-branch feature extraction module, and perform feature fusion to obtain the fused features of the corresponding scale.
[0010] Furthermore, each interaction submodule of the semantic interaction module is composed of two feature splicing layers and two feature fusion layers connected alternately. It is used to receive the multi-scale fusion features output by the gated fusion module, and to perform cross-scale interactive fusion of the multi-scale fusion features by adopting a dense skip connection-based information backflow mechanism to obtain multi-scale interactive features. The multi-scale interactive features include shallow interactive features, mid-level interactive features and deep interactive features.
[0011] Furthermore, the first interaction submodule of the semantic interaction module is used to receive the shallow fusion feature output by the gated fusion module and the initial mid-level fusion interaction feature output by the second interaction submodule, and perform feature splicing and feature fusion, and then perform feature splicing and feature fusion with the initial mid-level fusion interaction feature to obtain the shallow fusion interaction feature. The second interaction submodule is used to receive the mid-layer fusion feature, shallow fusion feature and initial deep fusion interaction feature output by the gated fusion module and the third interaction submodule, and perform feature splicing and feature fusion, and then perform feature splicing and feature fusion with the initial shallow fusion interaction feature and the initial deep fusion interaction feature from the first interaction submodule to obtain the mid-layer fusion interaction feature. The third interaction submodule is used to receive the deep fusion features and mid-layer fusion features output by the gated fusion module, and after performing feature splicing and feature fusion, it is then spliced and fused with the initial mid-layer fusion interaction features from the second interaction submodule to obtain the deep fusion interaction features.
[0012] Furthermore, the detection loss of the loss function includes classification loss and bounding box regression loss; The cross-modal consistency loss is constructed by maximizing the difference between the linear distance between heterogeneous feature pairs at the same location and the linear distance between heterogeneous feature pairs at different locations.
[0013] Secondly, the present invention provides a main transformer equipment anomaly detection system, the system comprising: The data processing module is used to acquire image data of the target main transformer equipment and preprocess the image data, which includes visible light images and infrared images. An anomaly detection module is used to input preprocessed image data into a preset anomaly detection model to obtain anomaly detection results for the target main transformer equipment. The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head. The loss function of the anomaly detection model consists of a detection loss and a cross-modal consistency loss, wherein the cross-modal consistency loss is constructed based on the distance between the same defect target in the heterogeneous feature space. The semantic interaction module includes three parallel interaction sub-modules. Each interaction sub-module consists of multiple feature splicing layers and multiple feature fusion layers connected alternately. The three interaction sub-modules perform cross-scale interactive fusion of the multi-scale fusion features output by the gated fusion module through an information feedback mechanism, and output multi-scale interactive features.
[0014] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0016] This invention provides a method, system, device, and storage medium for detecting anomalies in main transformer equipment. Through a dual-stream sensing architecture based on cross-stage aggregation and cross-level semantic enhancement interaction based on an information feedback mechanism, it ensures that minute geometric features such as micro-cracks and bolt anomalies in the main transformer equipment are not lost in the deep network, improving the accuracy of locating small-scale defects. Through cross-modal gating alignment fusion, it can deeply explore the intrinsic relationship between infrared thermal anomalies and visible light apparent anomalies, reducing the false alarm rate. By using a cross-modal consistency loss function to constrain the physical and semantic association between infrared thermal anomalies and visible light apparent anomalies, it can improve the robustness of the model. While ensuring high accuracy, this invention achieves a balance between model lightweighting and detection speed through an optimized feature aggregation structure, making it suitable for real-time monitoring tasks of inspection terminals on the edge side of substations. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the main transformer equipment anomaly detection method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the anomaly detection model in an embodiment of the present invention; Figure 3 This is an embodiment of the present invention. Figure 2 A schematic diagram of the CSA module; Figure 4 This is an embodiment of the present invention. Figure 2 A schematic diagram of the structure of the BUR module; Figure 5 This is an embodiment of the present invention. Figure 2 A schematic diagram of the Fusion module; Figure 6 This is a schematic diagram of the structure of the main transformer equipment anomaly detection system in an embodiment of the present invention; Figure 7 This is an internal structural diagram of the computer device in an embodiment of the present invention.
[0018] Figure label: 10. Data processing module; 20. Anomaly detection module. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Please see Figure 1 The first embodiment of the present invention proposes a method for detecting abnormalities in main transformer equipment, including steps S10 to S20: Step S10: Acquire image data of the target main transformer equipment and preprocess the image data, which includes visible light images and infrared images; Step S20: Input the preprocessed image data into the preset anomaly detection model to obtain the anomaly detection result of the target main transformer equipment.
[0021] In this embodiment, images of the target transformer are first acquired to obtain a registered visible light (RGB) image and an infrared (IR) image of the transformer. Then, the acquired image data is preprocessed. Since the transformer is often in an environment with strong light or alternating shadows, Retinex decomposition is performed on the RGB image to decompose the original RGB image into an illumination component and a reflection component. The illumination component carries information about the unevenness of ambient light, while the reflection component reflects the inherent properties of the device surface. Therefore, in this embodiment, the illumination component is removed, and only the enhanced reflection feature map is retained as the model input, thus eliminating false alarms caused by light and shadow from the data source.
[0022] For infrared images, anisotropic diffusion filtering is used for smoothing. Anisotropic diffusion filtering is an image processing technique based on partial differential equations. Its core principle is to achieve noise suppression and edge protection by dynamically adjusting the diffusion coefficient. In this embodiment, the algorithm can suppress random noise generated by high-voltage electromagnetic induction while using gradient operators to adaptively protect the thermal field edges of components such as main transformer bushings and oil conservators from being blurred.
[0023] The preprocessed image data is input into a pre-trained anomaly detection model for anomaly detection, and the model directly outputs the anomaly detection results for the target transformer. The architecture of the anomaly detection model is described in detail below; please refer to [link / reference]. Figure 2 The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head.
[0024] This embodiment addresses the characteristics of blurred surfaces and minute defects in main transformer equipment by constructing a dual-branch feature extraction module based on a dual-stream sensing architecture. This module overcomes the limitation of traditional residual block feature extraction, which cannot transmit detailed geometric features to deeper network layers. Specifically, the dual-branch feature extraction module uses a cross-stage aggregation (CSA) module as its core unit and establishes parallel first and second branch sub-modules. The two branch sub-modules have identical structures and extract features from visible light and infrared images, respectively.
[0025] like Figure 2 As shown, each branch submodule consists of three cascaded 3×3 convolutional layers (Conv), a max pooling layer (MaxPool2d), four cross-stage aggregation modules (CSA), and an attention-based intra-scale feature interaction (AIFI) module. The three 3×3 convolutional layers (Conv) are used to extract low-level features from the input image data to obtain a preliminary feature map. In the max pooling layer, the preliminary feature map is downsampled to reduce the spatial dimension of the feature map and enhance the expressive power of the features.
[0026] Please see Figure 3 The Cross-Stage Aggregation Module (CSA) is a multi-layer residual structure, consisting of a Split layer, a 3×3 convolutional layer (Conv3×3), an Efficient Channel Attention Module (ECA), a Feature Concatenation Layer (Concat), and a 1×1 convolutional layer (Conv1×1). The Efficient Channel Attention Module (ECA) comprises an Adaptive Average Pooling Layer (Adaptive AvgPool), a One-Dimensional Convolutional Layer (Adaptive 1DConv), and a Sigmoid activation function and weight application layer. This represents element-wise multiplication. The internal processing steps of the CSA module include: Feature channel splitting: The CSA module first performs channel splitting on the input feature map through the Split operation, dividing it into spatial feature maintenance branch and deep semantic extraction branch; Parallel processing of branches: The spatial feature maintenance branch features are aligned to the feature aggregation stage through residual linking operations to maintain the low-level geometric detail flow of the original image; the deep semantic extraction branch passes through a 3×3 convolutional layer and an ECA perception layer in sequence to extract discriminative defect semantic features from the noise of the complex power grid environment. Multi-dimensional aggregation: The CSA module uses the channel concatenation (Concat) operation to densely concatenate the alignment features of the spatial feature maintenance branch, the spatial convolution features of the deep semantic extraction branch, and the attention-weighted features in the channel dimension, so as to realize cross-stage collision and aggregation of features with different levels of abstraction. Global residual mapping: The aggregated features after channel concatenation are non-linearly reorganized through 1×1 convolution and connected with the original input features through global residuals to output the final evolved feature map.
[0027] In the four CSA modules composed of multi-layer residual structures, the representation of features is enhanced through residual learning, and deeper semantic information is extracted layer by layer. In this embodiment, the feature P3 output by the second layer CSA module is defined as a shallow feature, the feature P4 output by the third layer CSA module is defined as a mid-layer feature, and the feature P5 output by the fourth layer CSA module is defined as an initial deep feature.
[0028] The final layer of the branch submodule is the Attention-based Intra-scale Feature Interaction (AIFI) module. The AIFI module utilizes a self-attention mechanism to process high-level features in the image. Self-attention is a mechanism that allows the model to consider other relevant parts of the data while processing a specific part. The AIFI module focuses on internal scale interactions at the high-level feature layer. This is based on the recognition that high-level feature layers contain richer semantic concepts and can more effectively capture the relationships between conceptual entities in the image. At the same time, it avoids performing the same interactions at low-level feature layers, as low-level features lack the necessary semantic depth and may lead to duplication and confusion in data processing. The AIFI module typically consists of multi-head self-attention (MSA or deformable attention) and a feedforward network (FFN), formally equivalent to a standard Transformer encoder layer, but only operating on high-level feature layers. It should be noted that the parameters of the model architecture in this embodiment and subsequent embodiments are preferred options rather than specific limitations. The parameters of the model architecture can be flexibly adjusted based on the preferred parameters provided in this embodiment, which will not be elaborated further below.
[0029] In this embodiment, the AIFI module performs intra-scale feature interaction on the output feature P5 of the last layer CSA module. Through self-attention mechanism and feature enhancement operation, it improves the expressive power of the features and finally outputs the fused feature map F5, which is the deep feature.
[0030] In this embodiment, the shallow features P3, middle features P4, and deep features F5 of the visible light image output by the first branch submodule, and the shallow features P3, middle features P4, and deep features F5 of the infrared image output by the second branch submodule are used as the final output multi-scale features and input into the gated fusion module for dual-modal feature fusion.
[0031] In this embodiment, the visible light image output by the first branch submodule features appearance features, i.e., the surface morphology of the main transformer equipment exposed to the outside, while the infrared image output by the second branch submodule features thermal response features, i.e., the internal temperature of the main transformer equipment. These two features represent different modalities. To address the problems of large differences in modal representations, difficulty in balancing information redundancy and complementarity, and feature conflict and superposition caused by simple channel splicing during dual-modal feature fusion, and to achieve intra-scale dual-modal fusion, the gated fusion module in this embodiment adopts a parallel architecture, consisting of three parallel Bimodal Unified Representation (BUR) modules. Each BUR module is used to fuse features from the visible light image and the infrared image within the intra-scale, defining... Figure 2 The three BUR modules shown from top to bottom are the first BUR module, the second BUR module, and the third BUR module. The first BUR module is used to fuse the shallow features of the visible light image and the shallow features of the infrared image to obtain the shallow fused features. The second BUR module is used to fuse the mid-level features of the visible light image and the mid-level features of the infrared image to obtain the mid-level fused features. The third BUR module is used to fuse deep features from visible light images and infrared images to obtain deep fused features. The structure and internal processing of the BUR module will be explained in detail below.
[0032] Please see Figure 4 The BUR module consists of a feature concatenation layer (Concat), two feature transformation layers (Bblock), a gated weight generation unit, a gated fusion unit, and a third feature transformation layer (Bblock). The gated weight generation unit comprises two 3×3 convolutional layers (Conv3×3) and a sigmoid activation function. Each feature transformation layer (Bblock) consists of cascaded convolutional layers (Conv), a dynamic hyperbolic tangent function (Dynamic Tanh, DyT), and a sigmoid activation function. The gated fusion unit is based on element-wise multiplication. This can be achieved through [the following].
[0033] Inside the BUR module, the feature concatenation layer concatenates the first input feature Input1 and the second input feature Input2 along the channel dimension to obtain the concatenated feature. After the feature concatenation layer, the first feature transformation unit (Bblock) and the second feature transformation unit (Bblock) are connected in sequence. These two feature transformation units perform Bblock processing on the concatenated feature twice in sequence to extract and enhance the joint representation information of the two input features to obtain the backbone fusion feature.
[0034] In the gated weight generation unit, the first feature Input1 and the second feature Input2 are processed by 3×3 convolution, and the results of the two convolutions are added element by element. And gated weight features are generated by activating Sigmoid.
[0035] In the gated fusion unit, the backbone fusion features are multiplied element-wise with the gated weight features to perform weighted filtering of the backbone fusion features, highlighting key response regions and suppressing redundant interference information; finally, the weighted features are further processed by the third feature transformation unit to obtain the output feature, thereby forming a single fusion feature with a unified representation.
[0036] In this embodiment, the gated fusion module utilizes the BUR module to fuse bimodal features of the same scale. By calculating the mutual information correlation between the thermal response features of the infrared image and the texture features of the visible light image through gate weights, it dynamically adjusts the fusion weights of the two features, transforming the originally dispersed bimodal information into a single, discriminative fused feature representation. This eliminates redundancy in single-modal information caused by environmental interference, resulting in fused features representing complementary bimodal information and forming a unified feature pyramid. For example, when a defect manifests as significant overheating in the infrared image, while the visible light image experiences feature saturation due to reflection, the BUR module automatically increases the weight of the infrared channel and suppresses interference information in the visible light channel through network learning.
[0037] The feature pyramid generated by the gating fusion module is input into the semantic interaction module, which consists of three parallel interaction sub-modules, for cross-level semantic enhancement interaction. In this embodiment, an information feedback mechanism is constructed through dense jump connections between multiple interaction sub-modules. High-level semantic information is used as a priori guide to guide the low-level feature map for filtering.
[0038] like Figure 2 As shown, in this embodiment, the semantic interaction module (MFF) consists of three parallel interaction sub-modules. Each interaction sub-module is composed of two feature concatenation layers (Concat) and two feature fusion layers (Fusion) connected alternately. Each interaction sub-module takes the output of the corresponding BUR module as its main input and combines the fused features of adjacent interaction sub-modules to realize the reverse guidance of high-level semantic information on low-level geometric details.
[0039] like Figure 5 As shown, the feature fusion layer (Fusion) in this embodiment consists of two parallel 1×1 convolutional layers (Conv1×1), and two reparameterization modules (RepBlock) are cascaded after the second 1×1 convolutional layer. The RepBlock module is an efficient convolutional module that uses reparameterization technology to transform complex convolutional operations into simpler linear operations, thereby reducing computational complexity and improving model performance. It is usually composed of multiple RepConv modules.
[0040] After the concatenated features are input into the feature fusion layer (Fusion), the feature channels are first adjusted by point convolution through a 1×1 convolutional layer to obtain convolutional features. Then, the convolutional features are recombined through two RepBlock modules to obtain recombined features. Finally, the convolutional features and the recombined features are fused, which is done element-wise by addition. The features are fused to obtain a fused feature map. Based on the above processing steps, the Fusion operation can be represented as: In the formula, This represents a regular convolution operation; This indicates a reparameterized convolution operation.
[0041] Will as Figure 2 The MFF module shown has three interactive sub-modules defined from top to bottom as the first interactive sub-module, the second interactive sub-module, and the third interactive sub-module. The first interactive sub-module receives the shallow fusion features output by the first BUR module of the gated fusion module. And receive the initial mid-level fused interactive features output from the adjacent second interactive submodule after feature concatenation and feature fusion. Then, the initial mid-layer fusion interaction features were processed. After upsampling, the features are fused with those from the shallow layer through the first feature concatenation layer (Contcat). Feature concatenation is performed, and the concatenated features are then fused through the first feature fusion layer to obtain the initial shallow fused interactive features. Its expression is: In the formula, Indicates upsampling; Indicates feature concatenation operation; This indicates a fusion operation.
[0042] Initial shallow fusion interaction features The initial mid-layer fusion interaction features of upsampling After feature concatenation and fusion through the second feature concatenation layer and the second feature fusion layer, the final output is a shallow fused interactive feature. Its expression is: In the formula, This indicates downsampling.
[0043] The second interactive submodule receives the mid-layer fusion features output by the second BUR module in the gated fusion module. The shallow fusion features output by the first BUR module after downsampling In addition, it will receive the initial deep fusion interaction features output from the adjacent third interaction submodule after feature concatenation and feature fusion. After upsampling, the features are concatenated and fused with the other two features to obtain the initial mid-level fused interactive features. Its expression is: Then, the initial middle layer fusion interaction features are further processed through the second feature splicing layer and the second feature fusion layer. Initial shallow fusion interaction features of downsampling and the initial deep fusion interaction features of upsampling Feature splicing and fusion are performed to obtain mid-level fused interactive features. Its expression is: The third interactive submodule receives the deep fusion features output by the third BUR module of the gated fusion module. and the downsampled mid-layer fusion features output by the second BUR module The first feature concatenation layer and the first feature fusion layer concatenate and fuse these two features to obtain the initial deep fusion interactive features. Its expression is: Then, the initial deep fusion interactive features are further processed through the second feature splicing layer and the second feature fusion layer. The initial mid-layer fusion interaction features of downsampling Feature splicing and fusion are performed to obtain deep fusion interactive features. Its expression is: This embodiment constructs a cross-stage enhanced multi-scale semantic interaction module through a dense skip connection mechanism to perform cross-level feature recombination and interactive fusion of the feature pyramid, thereby integrating low-level geometric details and high-level semantic information, optimizing the feature representation of targets at different scales, and enhancing the information flow of multi-scale features through rich skip connections, thus improving the detection capability of small targets and complex backgrounds.
[0044] Finally, the multi-scale fused interactive features output by the semantic interaction module are input into the prediction head for decoding and reconstruction, outputting the detection results of the appearance and thermal anomalies of the main transformer equipment. In this embodiment, the prediction head has a fully convolutional three-scale detection structure, and its input is the three-scale fused feature map output by the feature fusion network. , and These correspond to downsampling scales of 1 / 8, 1 / 16, and 1 / 32 of the input image, respectively. Each scale has an independent detection branch, which consists of several convolutional layers and a final prediction convolutional layer. The final prediction convolutional layer outputs a 3D tensor used to encode bounding box regression parameters, target confidence, and class prediction results. For the fused feature map at the i-th scale... If the spatial dimensions are Hi×Wi, A preset bounding boxes are set at each grid position, and the number of target categories is K, then the predicted output is Oi∈R Hi×Wi×A(4+1+K)Here, 4 represents the regression parameters for the bounding box center coordinates and width and height, 1 represents the target confidence score, and K represents the category prediction result. The prediction results at the three scales are used for small-scale, medium-scale, and large-scale target detection, respectively, and are then uniformly summarized, confidence-filtered, and non-maximum suppression is performed in the post-processing stage to finally obtain the target's category, location coordinates, and detection confidence score. The final output anomaly detection results include apparent anomalies and thermal anomalies. Apparent anomalies include oil leakage from valves and seals, severe dirt and damage to insulator surfaces, and loosening and corrosion of metal connectors. Thermal anomalies include localized overheating and abnormal temperature rises.
[0045] Furthermore, to improve the robustness of the anomaly detection model, this embodiment, when training the anomaly detection model, in addition to using conventional detection loss... L det In addition to classification loss and bounding box regression loss, a cross-modal consistency loss is also introduced. Because visible light images and infrared images have physical imaging differences, simple tensor stitching can easily lead to feature conflicts. Therefore, this embodiment introduces a cross-modal consistency loss. L align By minimizing the distance between identical defective targets in the heterogeneous feature space, the model is constrained to learn the physical correlation between infrared thermal anomalies and visible light apparent anomalies. Specifically, this loss function forces a reduction in the linear distance between heterogeneous feature pairs at the same location in high-dimensional space, and makes their distance from negative samples at non-corresponding locations greater than a preset boundary. Its expression is: In the formula, The total number of target samples participating in the calculation; The preset boundary distance is used to control the separation margin between positive and negative sample pairs; This represents the specific spatial location index of the currently existing defective target, and its traversal range is from 1 to N; To perform calculations for the current position Select a non-corresponding regional spatial location index, and satisfy the following conditions: ; and They respectively indicate the same position The visible light feature vector and infrared feature vector extracted from the point are cross-modal positive sample pairs; For in position The negative sample feature vector extracted at the location; This represents the L2 norm of a vector.
[0046] Then, the detection loss was analyzed. L det and cross-modal consistency loss L alignWe perform a weighted summation to obtain the total loss function. L total Its expression is: In the formula, To balance the cross-modal constraint weights.
[0047] This embodiment uses the aforementioned loss function to train the anomaly detection model, which not only achieves deep semantic alignment and establishes a mapping logic between natural optical anomaly features and infrared image anomaly features, but also improves the model's robustness under harsh conditions (such as heavy fog and nighttime). Even with severe degradation of visible light features, the model can still accurately identify minute defects by relying on the cross-modal semantic binding established during training, significantly reducing the false negative and false positive rates caused by environmental interference.
[0048] This embodiment provides a method for detecting anomalies in main transformer equipment. Through heterogeneous feature decoupling preprocessing, it can eliminate interference from shadows, strong light, and electromagnetic induction noise, enabling the model to maintain stable detection performance under extreme weather or complex lighting conditions. By using a dual-stream sensing architecture based on cross-stage aggregation and cross-level semantic enhancement interaction based on information feedback mechanism, it ensures that subtle geometric features such as micro-cracks and bolt anomalies in the main transformer equipment are not lost in the deep network, significantly improving the accuracy of locating small-scale defects. Through cross-modal gating alignment fusion, it can deeply explore the intrinsic relationship between infrared thermal anomalies and visible light apparent anomalies, greatly reducing the false alarm rate. This embodiment achieves a balance between model lightweighting and detection speed through optimized feature aggregation structure while ensuring high accuracy, making it suitable for real-time monitoring tasks of inspection terminals on the edge side of substations.
[0049] Please see Figure 6 Based on the same inventive concept, the second embodiment of the present invention proposes a main transformer equipment anomaly detection system, comprising: The data processing module 10 is used to acquire image data of the target main transformer equipment and preprocess the image data, which includes visible light images and infrared images. Anomaly detection module 20 is used to input preprocessed image data into a preset anomaly detection model to obtain anomaly detection results of the target main transformer equipment. The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head. The loss function of the anomaly detection model consists of a detection loss and a cross-modal consistency loss, wherein the cross-modal consistency loss is constructed based on the distance between the same defect target in the heterogeneous feature space. The semantic interaction module includes three parallel interaction sub-modules. Each interaction sub-module consists of multiple feature splicing layers and multiple feature fusion layers connected alternately. The three interaction sub-modules perform cross-scale interactive fusion of the multi-scale fusion features output by the gated fusion module through an information feedback mechanism, and output multi-scale interactive features.
[0050] The technical features and effects of the main transformer equipment anomaly detection system proposed in this embodiment of the invention are the same as those of the method proposed in this embodiment of the invention, and will not be repeated here. Each module in the above-mentioned main transformer equipment anomaly detection system can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or it can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0051] Furthermore, embodiments of the present invention also propose a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0052] Please see Figure 7 The diagram illustrates the internal structure of a computer device in one embodiment. This computer device can specifically be a terminal or a server. The computer device includes a processor, memory, network interface, display, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for detecting abnormalities in the main transformer equipment. The display screen of the computer device can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0053] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computing devices may include more or fewer components than those shown in the figure, or combine certain components, or have the same component arrangement.
[0054] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0055] In summary, the present invention provides a method, system, device, and storage medium for detecting anomalies in main transformer equipment. The method acquires image data of the target main transformer equipment and preprocesses the image data, which includes visible light and infrared images. The preprocessed image data is then input into a preset anomaly detection model to obtain the anomaly detection result of the target main transformer equipment. The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head. The loss function of the anomaly detection model consists of a detection loss and a cross-modal consistency loss, which is constructed based on the distance between the same defect target and the heterogeneous feature space. The semantic interaction module includes three parallel interaction sub-modules, each composed of multiple feature splicing layers and multiple feature fusion layers alternately connected. The three interaction sub-modules perform cross-scale interactive fusion of the multi-scale fusion features output by the gated fusion module through an information feedback mechanism and output multi-scale interactive features. This invention eliminates interference from shadows, strong light, and electromagnetic noise through heterogeneous feature decoupling preprocessing, enabling the model to maintain stable detection performance under extreme weather or complex lighting conditions. Through a dual-stream sensing architecture based on cross-stage aggregation and cross-level semantic enhancement interaction based on information feedback mechanisms, it ensures that subtle geometric features such as minute cracks and bolt anomalies in the main transformer equipment are not lost in the deep network, significantly improving the accuracy of locating small-scale defects. Through cross-modal gating alignment fusion, it can deeply explore the intrinsic relationship between infrared thermal anomalies and visible light apparent anomalies, reducing the false alarm rate. While ensuring high accuracy, this invention achieves a balance between model lightweighting and detection speed through an optimized feature aggregation structure, making it suitable for real-time monitoring tasks of substation edge-side inspection terminals.
[0056] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0057] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A method for detecting anomalies in main transformer equipment, characterized in that, include: Image data of the target main transformer equipment is acquired and preprocessed. The image data includes visible light images and infrared images. The preprocessed image data is input into a preset anomaly detection model to obtain the anomaly detection result of the target main transformer equipment. The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head. The loss function of the anomaly detection model consists of a detection loss and a cross-modal consistency loss, wherein the cross-modal consistency loss is constructed based on the distance between the same defect target in the heterogeneous feature space. The semantic interaction module includes three parallel interaction sub-modules. Each interaction sub-module consists of multiple feature splicing layers and multiple feature fusion layers connected alternately. The three interaction sub-modules perform cross-scale interactive fusion of the multi-scale fusion features output by the gated fusion module through an information feedback mechanism, and output multi-scale interactive features.
2. The method for detecting abnormalities in main transformer equipment according to claim 1, characterized in that, The preprocessing steps for the image data include: The illumination component in the visible light image is removed by using the Retinex decomposition technique to obtain the preprocessed visible light image; Anisotropic diffusion filtering is used to smooth the infrared image, resulting in a preprocessed infrared image.
3. The method for detecting abnormalities in main transformer equipment according to claim 1, characterized in that, The dual-branch feature extraction module includes two parallel branch sub-modules. Each branch sub-module consists of three cascaded convolutional layers, a max pooling layer, four cross-stage aggregation modules, and an attention-based intra-scale feature interaction module. The cross-stage aggregation module is a multi-layer residual structure, consisting of a segmentation layer, a 3×3 convolutional layer, an efficient channel attention module, a feature splicing layer, and a 1×1 convolutional layer. The first branch submodule is used to extract multi-scale features from the preprocessed visible light image to obtain the multi-scale features of the visible light image, and the second branch submodule is used to extract multi-scale features from the preprocessed infrared image to obtain the multi-scale features of the infrared image. The multi-scale features output by each branch submodule include shallow features, mid-level features, and deep features. The shallow features are the features output by the second cross-stage aggregation module, the mid-level features are the features output by the third cross-stage aggregation module, and the deep features are the features output by the adaptive feature fusion module.
4. The method for detecting abnormalities in main transformer equipment according to claim 3, characterized in that, The gated fusion module includes three parallel dual-modal unified representation modules, which are used to receive multi-scale features of visible light and infrared images output by the dual-branch feature extraction module, and perform feature fusion using a gating mechanism to obtain multi-scale fusion features, which include shallow fusion features, medium fusion features and deep fusion features. Each dual-modal unified characterization module is used to receive the features of the visible light image and infrared image of the corresponding scale output by the dual-branch feature extraction module, and perform feature fusion to obtain the fused features of the corresponding scale.
5. The method for detecting abnormalities in main transformer equipment according to claim 1, characterized in that, Each interaction submodule of the semantic interaction module consists of two feature splicing layers and two feature fusion layers connected alternately. It is used to receive the multi-scale fusion features output by the gated fusion module and to perform cross-scale interactive fusion of the multi-scale fusion features by adopting a dense skip connection-based information backflow mechanism to obtain multi-scale interactive features. The multi-scale interactive features include shallow interactive features, mid-level interactive features and deep interactive features.
6. The method for detecting abnormalities in main transformer equipment according to claim 5, characterized in that, The first interaction submodule of the semantic interaction module is used to receive the shallow fusion feature output by the gated fusion module and the initial mid-level fusion interaction feature output by the second interaction submodule, and perform feature splicing and feature fusion, and then perform feature splicing and feature fusion with the initial mid-level fusion interaction feature to obtain the shallow fusion interaction feature. The second interaction submodule is used to receive the mid-layer fusion feature, shallow fusion feature and initial deep fusion interaction feature output by the gated fusion module and the third interaction submodule, and perform feature splicing and feature fusion, and then perform feature splicing and feature fusion with the initial shallow fusion interaction feature and the initial deep fusion interaction feature from the first interaction submodule to obtain the mid-layer fusion interaction feature. The third interaction submodule is used to receive the deep fusion features and mid-layer fusion features output by the gated fusion module, and after performing feature splicing and feature fusion, it is then spliced and fused with the initial mid-layer fusion interaction features from the second interaction submodule to obtain the deep fusion interaction features.
7. The method for detecting abnormalities in main transformer equipment according to claim 1, characterized in that, The detection loss of the loss function includes classification loss and bounding box regression loss; The cross-modal consistency loss is constructed by maximizing the difference between the linear distance between heterogeneous feature pairs at the same location and the linear distance between heterogeneous feature pairs at different locations.
8. A main transformer equipment anomaly detection system, characterized in that, include: The data processing module is used to acquire image data of the target main transformer equipment and preprocess the image data, which includes visible light images and infrared images. An anomaly detection module is used to input preprocessed image data into a preset anomaly detection model to obtain anomaly detection results for the target main transformer equipment. The anomaly detection model includes a cascaded dual-branch feature extraction module, a gated fusion module, a semantic interaction module, and a prediction head. The loss function of the anomaly detection model consists of a detection loss and a cross-modal consistency loss, wherein the cross-modal consistency loss is constructed based on the distance between the same defect target in the heterogeneous feature space. The semantic interaction module includes three parallel interaction sub-modules. Each interaction sub-module consists of multiple feature splicing layers and multiple feature fusion layers connected alternately. The three interaction sub-modules perform cross-scale interactive fusion of the multi-scale fusion features output by the gated fusion module through an information feedback mechanism, and output multi-scale interactive features.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.