Defect identification method, system and equipment for main transformer equipment and storage medium

By collecting and fusing multimodal image data, and using a Transformer encoder for feature transformation and temporal feature extraction, the accuracy problem of defect identification in main transformer equipment was solved, achieving a higher recognition rate and a lower false alarm rate.

CN120997746APending Publication Date: 2025-11-21STATE GRID ZHEJIANG ELECTRIC POWER CO LTD HANGZHOU POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511519155.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies cannot accurately identify defects in main transformer equipment using multimodal images, especially in complex climates and low visibility environments, resulting in low recognition rates, high false alarm rates, and weak generalization capabilities.

Method used

Real-time video data of the main transformer to be identified is collected, image data under different modalities are extracted, and feature fusion is performed using a multimodal cross-attention fusion mechanism, including infrared images, visible light images and ultraviolet images. Feature transformation and temporal feature extraction are performed through a Transformer encoder to construct a defect identification model.

Benefits of technology

It improved the accuracy of main transformer equipment defect identification, compensated for the bias of single-mode data, and enhanced the accuracy of equipment defect identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997746A_ABST
    Figure CN120997746A_ABST
Patent Text Reader

Abstract

The invention discloses a main transformer equipment defect identification method, system and equipment and a storage medium, and is applied to the technical field of main transformer equipment defect identification. Image data of to-be-identified main transformer equipment in different modes during operation are collected, and feature information corresponding to the image data in the different modes is extracted; according to the method, the main transformer equipment to be recognized is subjected to defect recognition, all feature information is converted into corresponding high-dimensional representation results, feature fusion is carried out on the basis of the obtained high-dimensional representation results, retrograde spatial features and time features, and then a defect recognition result of the main transformer equipment to be recognized is obtained. And the defect identification accuracy of the to-be-identified equipment is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid equipment defect identification, and in particular to a main transformer equipment defect identification method, system, device and storage medium. BACKGROUND

[0002] In the operation of a power system, the failure of key equipment such as transformers, circuit breakers, switch cabinets and the like often has dynamic evolution characteristics, such as hot spots, partial discharges or structural damage caused by insulation aging. If these defects cannot be identified in a timely manner, they will pose a significant threat to the safety of the power transmission and transformation system. Traditional image recognition methods are mostly based on static images or single-modal images, and are difficult to fully reflect the operating state of the equipment, especially in complex climates and low-visibility environments, with low recognition rates, high false positive rates and weak generalization capabilities. SUMMARY

[0003] To solve the above technical problems, the embodiments of the present application provide a main transformer equipment defect identification method, system, device and storage medium to solve the technical problem that the prior art cannot accurately identify defects in main transformer equipment by using multi-modal images.

[0004] The first aspect of the embodiments of the present application provides a main transformer equipment defect identification method, which comprises: Based on real-time video data of the main transformer equipment to be identified, image data under different modalities is obtained; Respectively extracting the feature information corresponding to each image data; Converting all feature information into corresponding high-dimensional representation results, wherein the high-dimensional representation results at least include time frame number information, position information and modality information; According to the common features of each high-dimensional representation result in the spatial dimension, each high-dimensional representation result is converted into a corresponding spatial feature sequence, wherein the spatial dimension at least includes the defect position of the main transformer equipment to be identified; Extracting the dynamic feature representation of the spatial feature sequence under each modality evolving over time from the real-time video data to obtain the time feature corresponding to each modality; Processing all time features based on a multi-modal cross-attention fusion mechanism, and obtaining the defect identification result of the main transformer equipment to be identified based on the processing result and the constructed defect identification model.

[0005] The method collects image data of the main transformer equipment to be identified under different modalities during operation, respectively extracts feature information corresponding to the image data under different modalities, and converts all the feature information into corresponding high-dimensional representation results. Then, spatial features and time features are extracted based on the obtained high-dimensional representation results, and then feature fusion is performed to obtain the defect recognition result of the main transformer equipment to be identified. This method makes up for the deviation of defect recognition by single modal data, and effectively improves the defect recognition accuracy of the equipment to be identified.

[0006] In a possible implementation manner of the first aspect, converting all the feature information into corresponding high-dimensional representation results comprises: performing image block division on all the feature information to obtain an image block set of each feature information; adding position embedding information and modality embedding information to each image block in each image block set to obtain a high-dimensional vector corresponding to each image block; fusing the image blocks in each image block set based on the high-dimensional vector corresponding to each image block to obtain a high-dimensional representation result corresponding to each feature information, wherein the high-dimensional representation result at least includes time frame number information, position information and modality information.

[0007] In a possible implementation manner of the first aspect, converting each high-dimensional representation result into a corresponding spatial feature sequence according to common features of the captured high-dimensional representation results in the spatial dimension comprises: performing feature conversion on the high-dimensional representation result corresponding to each feature information by using different first Transformer encoders based on common features of the high-dimensional representation results in the spatial dimension to obtain a spatial feature sequence corresponding to each high-dimensional representation result, wherein the different first Transformer encoders are a plurality of first Transformer encoders sharing parameters.

[0008] In a possible implementation manner of the first aspect, extracting a dynamic feature representation of the spatial feature sequence under each modality evolving over time from the real-time video data to obtain a time feature corresponding to each modality comprises: extracting a dynamic feature representation of each high-dimensional representation result by using different second Transformer encoders based on the real-time video data under different modalities and the spatial feature sequence corresponding to each high-dimensional representation result to obtain a time feature of each modality, wherein the different second Transformer encoders are a plurality of second Transformer encoders not sharing parameters.

[0009] In a possible implementation manner of the first aspect, all time features are processed based on a multi-modal cross-attention fusion mechanism, and a defect recognition result of the to-be-recognized main transformer equipment is obtained based on a processing result and a built defect recognition model, including: Based on the modal information corresponding to all time features, a time feature is determined as a key-value pair, and a to-be-fused time feature is obtained; The remaining time features are sequentially taken as queries, and information interaction between different modalities is performed on the to-be-fused time feature, to obtain corresponding fusion features; The fusion features are fused by using the built defect recognition model, to obtain the defect recognition result of the to-be-recognized main transformer equipment.

[0010] In a possible implementation manner of the first aspect, the image data under different modalities at least includes infrared images, visible light images, and ultraviolet images.

[0011] To solve the same technical problem, a second aspect of the embodiment of the application provides a main transformer equipment defect recognition system, including an acquisition module, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, and a fusion module, wherein, The acquisition module is configured to obtain image data under different modalities based on real-time video data of to-be-recognized main transformer equipment; The first feature extraction module is configured to extract feature information corresponding to each image data respectively; The second feature extraction module is configured to convert all feature information into corresponding high-dimensional representation results, wherein the high-dimensional representation results at least include time frame number information, position information, and modal information; The third feature extraction module is configured to convert each high-dimensional representation result into a corresponding spatial feature sequence according to common features of the captured high-dimensional representation results in a spatial dimension, wherein the spatial dimension at least includes a defect position of the to-be-recognized main transformer equipment; The fourth feature extraction module is configured to extract dynamic feature representations of spatial feature sequences under each modality evolving over time from the real-time video data, to obtain time features corresponding to each modality; The fusion module is configured to process all time features based on a multi-modal cross-attention fusion mechanism, and obtain a defect recognition result of the to-be-recognized main transformer equipment based on a processing result and a built defect recognition model.

[0012] In a possible implementation manner of the first aspect, the second feature extraction module includes a division unit, a linear mapping unit, and a high-dimensional vector fusion unit, wherein, The division unit is configured to perform image block division on all feature information, to obtain an image block set of each feature information; The linear mapping unit is configured to add position embedding information and modality embedding information to each image block in each image block set to obtain a high-dimensional vector corresponding to each image block. The high-dimensional vector fusion unit is configured to fuse the image blocks in each image block set based on the high-dimensional vectors corresponding to the image blocks to obtain a high-dimensional representation result corresponding to each feature information, wherein the high-dimensional representation result at least includes time frame number information, position information and modality information.

[0013] The third aspect of the embodiment of the present application provides a computer device, comprising: The memory is configured to store a computer program. The processor is configured to implement the steps of the main transformer equipment defect identification method according to the first aspect when executing the computer program.

[0014] The fourth aspect of the embodiment of the present application provides a storage medium, and the storage medium stores a computer program. The computer program is executed by the processor to implement the steps of the main transformer equipment defect identification method according to the first aspect.

[0015] The technical scheme of the present application has the following advantages: The present application collects image data under different modalities when the main transformer equipment to be identified is running, extracts feature information corresponding to the image data under different modalities, and converts all the feature information into corresponding high-dimensional representation results. Then, spatial features and time features are extracted based on the obtained high-dimensional representation results, and then feature fusion is performed to obtain the defect identification result of the main transformer equipment to be identified. This method makes up for the deviation of defect identification by single modality data, and effectively improves the defect identification accuracy of the equipment to be identified. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the specific embodiments of the present application or the technical scheme in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0017] Figure 1 The defect identification method flowchart of the main transformer equipment defect identification method in the embodiment of the present application; Figure 2 The multi-modal spatio-temporal structure diagram of the main transformer equipment defect identification method in the embodiment of the present application; Figure 3 The structure block diagram of the main transformer equipment defect identification system in the embodiment of the present application. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0019] The main transformer equipment defect identification method provided by the embodiments of the present application is shown in Figure 1 Figure 1 The main transformer equipment defect identification method flowchart includes steps S101 to S106, and each step is specifically as follows: S101, obtaining image data in different modalities based on real-time video data of the main transformer equipment to be identified.

[0020] In this embodiment, the specific way of obtaining image data in different modalities based on real-time video data of the main transformer equipment to be identified is not limited, for example, different types of cameras are installed on the main transformer equipment to be identified to collect image data in different modalities. Specifically, a high-definition industrial camera is installed around the main transformer equipment (such as the key parts of the transformer body, oil pillow, sleeve, etc.), an infrared thermal imager and a visible light camera are configured, so as to obtain the first type of image, the second type of image and the third type of image. The first type of image refers to an infrared image, which uses an infrared thermal imager to monitor the temperature distribution of the equipment in real time, collects thermal infrared video at a set sampling frequency, and is used to analyze local overheating, abnormal load and other thermal faults of the equipment. An infrared image is intercepted in the thermal infrared video. The second type of frame image refers to a visible light frame image. Specifically, the RGB video image of the appearance of the equipment is collected at a fixed time interval through the visible light camera, and visible defects such as oil stains, rust and deformation on the surface of the equipment are captured, and a visible light frame image is intercepted in the video. The third type of frame image refers to an ultraviolet frame image. Specifically, an ultraviolet camera is used, and a gain sensor is adjusted to collect ultraviolet light signals generated by corona discharge at a preset rate to obtain ultraviolet video, which is used to detect early faults such as insulation deterioration and sharp discharge. An ultraviolet video frame image is obtained by intercepting the ultraviolet video.

[0021] It should be noted that when collecting, a GPS time module provides a unified time reference for the three types of image collection equipment, ensuring that each frame of image is attached with a nanosecond-level time stamp. The frame images of the same time are intercepted from the thermal infrared video, the RGB video image and the ultraviolet video.

[0022] It should be further noted that the image data in different modalities at least includes infrared images, visible light images and ultraviolet images.

[0023] ​The embodiment is not limited to the specific modal type of the image data under different modalities. For example, the infrared image can capture abnormal changes in the temperature of the equipment, the ultraviolet image can detect partial discharge phenomena, and the visible light image provides surface structure information. Therefore, infrared images, visible light images, and ultraviolet images can be selected for multi-modal joint analysis, thereby providing more dimensional information support for power equipment defect detection. However, the embodiment is not limited to the above-mentioned modalities, and other modalities can also be used. S102, respectively extracting feature information corresponding to each image data.

[0024] The embodiment is not limited to the specific way of respectively extracting feature information corresponding to each image data. Specifically, different convolutional neural networks are used to extract features from image data of different modalities to obtain feature information corresponding to each image. For example, first, the size of the first type of frame image is unified, without noise reduction and brightness normalization, and then the first convolutional neural network is input into the trained first convolutional neural network for feature extraction to obtain a feature tensor, i.e., first feature information. The second type of frame image is input into the trained second convolutional neural network for feature extraction to obtain a feature map, e.g., first, the size of the second type of frame image is unified, without noise reduction and brightness normalization, and then the second convolutional neural network is input into the trained second convolutional neural network for feature extraction to obtain a feature tensor, i.e., second feature information. The third type of frame image is input into the trained third convolutional neural network for feature extraction to obtain a feature tensor, i.e., third feature information.

[0025] It should be noted that the first convolutional neural network, the second convolutional neural network, and the third convolutional neural network are all obtained by training based on the ResNet18 neural network, and the training process is the same as the existing training process of the neural network, which will not be repeated here.

[0026] S103, converting all feature information into corresponding high-dimensional representation results, wherein the high-dimensional representation results at least include time frame number information, position information, and modal information.

[0027] The embodiment is not limited to the specific way of converting all feature information into corresponding high-dimensional representation results. For example, the feature information of each modality is divided into N fixed-size image blocks, which are mapped into D-dimensional feature vectors through linear mapping to obtain high-dimensional representation results corresponding to each feature information. Specifically, after image block division and fusion of the first feature information, the second feature information, and the third feature information, the corresponding high-dimensional representation results are obtained, which provide data support for subsequent spatial feature and time feature extraction.

[0028] It should be further noted that the above-mentioned conversion of all feature information into corresponding high-dimensional representation results can include: dividing all feature information into image block sets of each feature information; adding position embedding information and modality embedding information to each image block in each image block set to obtain a high-dimensional vector corresponding to each image block; fusing image blocks in each image block set based on the high-dimensional vector corresponding to each image block to obtain a high-dimensional representation result corresponding to each feature information, wherein the high-dimensional representation result at least includes time frame number information, position information and modality information.

[0029] This embodiment does not limit the specific way of converting all feature information into corresponding high-dimensional representation results. For example, the first feature information, the second feature information and the third feature information are respectively divided into N fixed size image blocks, each image block is mapped to a D-dimensional feature vector through linear mapping, such as D=512, to obtain a high-dimensional representation result corresponding to each image block. When linear mapping is performed, a learnable position embedding (Positional Embedding) and a learnable modality embedding (Modality Embedding) are added to each high-dimensional representation result, wherein the learnable position embedding represents the spatial position information of the block; the learnable modality embedding (Modality Embedding) is used to identify whether it comes from the RGB, IR or UV modality. Then, the image blocks with spatial position information and modality information in the first feature information, the second feature information and the third feature information are respectively merged, and the high-dimensional representation results with spatial information, time frame number information and modality information are obtained after merging, that is, the image blocks with spatial position and modality embedding in the first feature information are merged to obtain a first high-dimensional representation result, the image blocks with spatial position and modality embedding in the second feature information are merged to obtain a second high-dimensional representation result, and the image blocks with spatial position and modality embedding in the third feature information are merged to obtain a third high-dimensional representation result.

[0030] S104, converting each high-dimensional representation result into a corresponding spatial feature sequence according to the common features of each high-dimensional representation result in the spatial dimension, wherein the spatial dimension at least includes the defect position of the main variable device to be identified.

[0031] The embodiment is not limited to the specific way of converting each high-dimensional representation result into a corresponding spatial feature sequence according to the common characteristics of the spatial dimensions of the captured high-dimensional representation results, for example, after obtaining the high-dimensional representation results under each modality, a spatial joint attention module (SJAM) is used to extract spatial features from each high-dimensional representation result. For example, the spatial joint attention module (SJAM) mainly includes a multi-layer Transformer encoder structure, the first layer of the Transformer encoder models the feature vector corresponding to the first high-dimensional representation result to obtain the first spatial feature sequence; the second layer of the Transformer encoder models the feature vector corresponding to the second high-dimensional representation result to obtain the second spatial feature sequence; and the third layer of the Transformer encoder models the feature vector corresponding to the third high-dimensional representation result to obtain the third spatial feature sequence. The spatial joint attention module is a multi-layer Transformer encoder structure with a parameter sharing mechanism, which is used to model the commonalities and differences of the three modalities in the spatial dimension. The spatial joint attention module includes multiple stacked Transformer encoder layers, each layer being composed of a multi-head self-attention (MHSA) sublayer and a feedforward neural network sublayer alternately; and a LayerNorm and a residual connection mechanism are used in each layer. The processing process of the Transformer encoder layer is as follows: Step 1: Calculate the query (Q), key (K), and value (V) representations for each high-dimensional representation result, and construct the attention representation by the following formula: In the formula, , , represent the weight matrices corresponding to the query, key, and value respectively, is the input vector sequence.

[0032] Step 2: Calculate the attention score using the scaled dot-product, and the expression is: In the formula, is the attention mechanism calculation method, is the dimension of the query representation and the key representation.

[0033] Step 3: The outputs of multiple attention heads are input into a linear transformation after being spliced, and a residual connection and normalization (LayerNorm) are performed with the original input: In the formula, is the output of the multi-head self-attention, is the normalization operation.

[0034] Step 4: Each token is then input into a feed-forward network (FFN): wherein, , are weight matrices, , are bias terms, is an activation function.

[0035] Residual stacking and normalization are performed again, and the final spatial feature sequence is output.

[0036] S105, extract the dynamic feature representation of the spatial feature sequence under each modality evolving over time from the real-time video data, and obtain the time feature corresponding to each modality.

[0037] The embodiment does not limit the specific way of extracting the dynamic feature representation of the spatial feature sequence under each modality evolving over time from the real-time video data, and obtaining the time feature corresponding to each modality, for example, constructing a triple-stream temporal module (TSTT), and constructing three groups of Transformer encoders that do not share weights with each other, as shown in Figure 2 The triple-stream temporal module is used to model the time dimension of the first spatial feature sequence, the second spatial feature sequence, and the third spatial feature sequence, extract the dynamic feature representation of each modality evolving over time, and obtain the time feature under each modality. For example, the dynamic feature representation of the first spatial feature sequence evolving over time is extracted to obtain the first time feature, the dynamic feature representation of the second spatial feature sequence evolving over time is extracted to obtain the second time feature, and the dynamic feature representation of the third spatial feature sequence evolving over time is extracted to obtain the third time feature.

[0038] It should be noted that the TSTT module is specially designed to process time series data and can capture the dynamic changes of each modality in the time dimension. This is crucial for identifying and understanding the development process of device defects. For example, local overheating may gradually intensify over time, or local discharge activity may occur within a certain time interval. By processing the time series of each modality separately, the TSTT can more accurately capture these dynamic changes.

[0039] S105, process all time features based on the multi-modal cross-attention fusion mechanism, and based on the processing result and the constructed defect recognition model, obtain the defect recognition result of the main transformer device to be recognized.

[0040] The embodiment is not limited to the specific manner of processing all time features based on the multi-modal cross attention fusion mechanism, and obtaining the defect recognition result of the to-be-recognized main transformer equipment based on the processing result and the constructed defect recognition model. For example, a multi-modal cross attention module (MCA) is constructed to fuse the time features of each modality to obtain corresponding fused features. Then, the multiple fused features are spliced and input into a fully connected layer to output the classification result of the equipment defects, such as main transformer local overheating, surface corrosion and damage, insulation failure, and the like.

[0041] It should be further explained that the processing of all time features based on the multi-modal cross attention fusion mechanism, and the obtaining of the defect recognition result of the to-be-recognized main transformer equipment based on the processing result and the constructed defect recognition model can include: Based on the modality information corresponding to all time features, a time feature is determined as a key-value pair to obtain a to-be-fused time feature; The remaining time features are sequentially taken as queries to interact information between different modalities with the to-be-fused time feature to obtain corresponding fused features; Each fused feature is fused by using the constructed defect recognition model to obtain the defect recognition result of the to-be-recognized main transformer equipment.

[0042] The embodiment is not limited to the specific manner of processing all time features based on the multi-modal cross attention fusion mechanism, and obtaining the defect recognition result of the to-be-recognized main transformer equipment based on the processing result and the constructed defect recognition model. For example, a multi-modal cross attention module (MCA) is constructed to fuse the time features of three types of frame images. The multi-modal cross attention module mainly realizes the deep semantic interaction between modalities through a cross-Transformer decoder structure, including the following two paths: Path one: taking the first time feature as a query and the second time feature as a key-value pair to obtain a first fused feature; Path two: taking the third time feature as a query and the second time feature as a key-value pair to obtain a second fused feature, and outputting the fused feature representation for the classification task after fusion.

[0043] The MSTT network proposed in the application is compared with existing mainstream video recognition methods on the TROPED three-modal power equipment defect identification dataset, including TSN (81.1%), I3D (83.8%), ViT (89.2%), MViTv2 (91.8%) and VideoSwin (94.6%) and other representative architectures. Under the same training conditions, the method reaches 97.3% in classification accuracy, which is 2.7% higher than that of the VideoSwin model (94.6%) with better performance in recent years, and the F1 score is increased from 0.93 to 0.96, which shows that the MSTT network proposed in the application has obvious advantages in multi-modal feature fusion and space-time modeling.

[0044] The application extracts high-dimensional features, spatial sequence features and time sequence features from image data of different modalities, mines potential space-time correlation of multi-modal data, and accurately identifies and locates key features in the defect occurrence process.

[0045] The main transformer equipment defect identification system provided by the embodiment of the application is shown in Figure 3 As shown in Figure 3 The structure block diagram of the main transformer equipment defect identification system 300 in the embodiment of the application includes an acquisition module 301, a first feature extraction module 302, a second feature extraction module 303, a third feature extraction module 304, a fourth feature extraction module 305 and a fusion module 306, wherein The acquisition module 301 is configured to obtain image data in different modalities based on real-time video data of the main transformer equipment to be identified. The first feature extraction module 302 is configured to extract feature information corresponding to each image data, respectively. The second feature extraction module 303 is configured to convert all the feature information into corresponding high-dimensional representation results, wherein the high-dimensional representation results at least include time frame number information, position information and modal information. The third feature extraction module 304 is configured to convert each high-dimensional representation result into a corresponding spatial feature sequence according to common features of the captured high-dimensional representation results in the spatial dimension, wherein the spatial dimension at least includes a defect position of the main transformer equipment to be identified. The fourth feature extraction module 305 is configured to extract dynamic feature representation of the spatial feature sequence in each modality evolving over time from the real-time video data, and obtain time features corresponding to each modality. The fusion module 306 is configured to process all the time features based on a multi-modal cross-attention fusion mechanism, and obtain a defect identification result of the main transformer equipment to be identified based on a processing result and a constructed defect identification model.

[0046] In an embodiment, the second feature extraction module comprises a division unit, a linear mapping unit and a high-dimensional vector fusion unit, wherein, The division unit is configured to divide all feature information into image block sets to obtain image blocks of each feature information. The linear mapping unit is configured to add position embedding information and modality embedding information to each image block in each image block set to obtain a high-dimensional vector corresponding to each image block. The high-dimensional vector fusion unit is configured to fuse image blocks in each image block set based on the high-dimensional vector corresponding to each image block to obtain a high-dimensional representation result corresponding to each feature information, wherein the high-dimensional representation result at least includes time frame number information, position information and modality information.

[0047] The specific implementation of the main transformer equipment defect identification system is basically the same as the specific embodiments of the main transformer equipment defect identification method described above, and will not be repeated here.

[0048] In an embodiment of the present application, a computer device is provided, which includes a memory and a processor, the memory stores a computer program, and the processor implements the above steps when executing the computer program; the computer device provided in the embodiment has similar implementation principles and technical effects to the method embodiments described above, and will not be repeated here.

[0049] In an embodiment of the present application, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the above steps; the computer readable storage medium provided in the embodiment has similar implementation principles and technical effects to the method embodiments described above, and will not be repeated here.

[0050] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0051] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only for specific embodiments of the present application and is not intended to limit the protection scope of the present application. In particular, any modification, equivalent replacement, improvement, etc. made by those skilled in the art within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for identifying defects in main transformer equipment, characterized in that, include: Based on the real-time video data of the main transformer device to be identified, image data under different modes is obtained; Extract the feature information corresponding to each of the image data; All the aforementioned feature information is converted into corresponding high-dimensional representation results, wherein the high-dimensional representation results include at least time frame number information, location information, and modality information; Based on the common features of the captured high-dimensional representation results in the spatial dimension, each high-dimensional representation result is converted into a corresponding spatial feature sequence, wherein the spatial dimension includes at least the defect location of the main transformer equipment to be identified; Extract the dynamic feature representation of the spatial feature sequence under each modality as it evolves over time from the real-time video data to obtain the temporal features corresponding to each modality; The time features are processed using a multimodal cross-attention fusion mechanism, and the defect identification results of the main transformer to be identified are obtained based on the processing results and the constructed defect identification model.

2. The method for identifying defects in main transformer equipment as described in claim 1, characterized in that, The step of converting all the feature information into corresponding high-dimensional representations includes: All the aforementioned feature information is divided into image blocks to obtain image block sets for each of the aforementioned feature information; Add positional embedding information and modality embedding information to each image block in each set of image blocks to obtain a high-dimensional vector corresponding to each image block; Based on the high-dimensional vectors corresponding to each image block, the image blocks in each set of image blocks are fused to obtain a high-dimensional representation result corresponding to each feature information, wherein the high-dimensional representation result includes at least time frame information, position information and modality information.

3. The method for identifying defects in main transformer equipment as described in claim 1, characterized in that, The step of converting each high-dimensional representation result into a corresponding spatial feature sequence based on the common features of the captured high-dimensional representation results in the spatial dimension includes: Based on the common features of the various high-dimensional representation results in the spatial dimension, different first Transformer encoders are used to perform feature transformation on the high-dimensional representation results corresponding to each feature information to obtain the spatial feature sequence corresponding to each high-dimensional representation result. The different first Transformer encoders are multiple first Transformer encoders with shared parameters.

4. The method for identifying defects in main transformer equipment as described in claim 1, characterized in that, The step of extracting the dynamic feature representation of the spatial feature sequence under each modality evolving over time from the real-time video data to obtain the temporal features corresponding to each modality includes: Based on real-time video data under different modalities and the spatial feature sequences corresponding to each of the high-dimensional representation results, dynamic feature representation extraction is performed on each of the high-dimensional representation results using different second Transformer encoders to obtain the temporal features of each modality. The different second Transformer encoders are multiple second Transformer encoders with non-shared parameters.

5. The method for identifying defects in main transformer equipment as described in claim 1, characterized in that, The process of processing all the time features based on the multimodal cross-attention fusion mechanism, and obtaining the defect identification result of the main transformer to be identified based on the processing result and the constructed defect identification model, includes: Based on the modal information corresponding to all the time features, a time feature is determined as a key-value pair to obtain the time feature to be fused. The remaining time features are used as queries in sequence, and information interaction between different modalities is performed with the time features to be fused to obtain the corresponding fused features; By using the constructed defect identification model, the various fusion features are fused to obtain the defect identification result of the main transformer equipment to be identified.

6. The method for identifying defects in main transformer equipment as described in claim 1, characterized in that, The image data in different modalities includes at least infrared images, visible light images, and ultraviolet images.

7. A defect identification system for main transformer equipment, characterized in that, It includes an acquisition module, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, and a fusion module, wherein, The acquisition module is used to obtain image data in different modes based on the real-time video data of the main transformer device to be identified; The first feature extraction module is used to extract feature information corresponding to each of the image data. The second feature extraction module is used to convert all the feature information into corresponding high-dimensional representation results, wherein the high-dimensional representation results include at least time frame information, location information and modality information; The third feature extraction module is used to convert each high-dimensional representation result into a corresponding spatial feature sequence based on the common features of each captured high-dimensional representation result in the spatial dimension, wherein the spatial dimension includes at least the defect location of the main transformer equipment to be identified. The fourth feature extraction module is used to extract the dynamic feature representation of the spatial feature sequence under each modality as it evolves over time from the real-time video data, so as to obtain the temporal features corresponding to each modality. The fusion module is used to process all the time features based on the multimodal cross-attention fusion mechanism, and obtain the defect identification result of the main transformer to be identified based on the processing result and the constructed defect identification model.

8. The main transformer equipment defect identification system as described in claim 7, characterized in that, The second feature extraction module includes a partitioning unit, a linear mapping unit, and a high-dimensional vector fusion unit, wherein, The partitioning unit is used to divide all the feature information into image blocks to obtain a set of image blocks for each feature information. The linear mapping unit is used to add position embedding information and modality embedding information to each image block in each set of image blocks to obtain a high-dimensional vector corresponding to each image block. The high-dimensional vector fusion unit is used to fuse the image blocks in the set of image blocks based on the high-dimensional vectors corresponding to each image block to obtain the high-dimensional representation result corresponding to each feature information, wherein the high-dimensional representation result includes at least time frame information, position information and modality information.

9. A computer device, characterized in that, include: Memory, used to store computer programs; A processor is configured to implement the main transformer equipment defect identification method as described in any one of claims 1 to 6 when executing the computer program.

10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the steps of the main transformer equipment defect identification method as described in any one of claims 1 to 6.