Detection method of power transformer oil leakage detection system based on multimodal prompts and multi-scale segmentation

By combining the multimodal information fusion of Mask2Former and CLIP models, high precision and robustness of power transformer oil leakage detection are achieved, solving the problem of insufficient detection accuracy in existing technologies, especially for oil leakage identification in complex environments and hidden locations.

CN119646668BActive Publication Date: 2025-10-03CHINA YANGTZE POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411791458.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-10-03
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Existing power transformer oil leakage detection technology has insufficient detection accuracy in complex environments and hidden locations. In addition, traditional methods are inefficient and costly, making it difficult to achieve efficient and accurate oil leakage identification.

Method used

A detection method based on multimodal cues and multi-scale segmentation is adopted, combined with the Mask2Former and CLIP models. By fusing image and text information, the Transformer decoder is used for feature extraction and segmentation to generate accurate segmentation and classification results of the oil spill area.

Benefits of technology

The accuracy and robustness of oil leak detection are improved, and the oil leak area can be accurately identified in complex scenarios, which reduces detection errors and improves the adaptability and versatility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119646668B_ABST
    Figure CN119646668B_ABST
Patent Text Reader

Abstract

A detection method for a power transformer oil leakage detection system based on multimodal prompts and multi-scale segmentation, through the steps of data input and preprocessing, feature extraction, multimodal feature fusion, Transformer decoder decoding, segmentation and classification, model training and fine-tuning, and output and result display, combined with Mask2Former and CLIP models, fully utilizes multimodal information to improve the accuracy and robustness of oil leakage detection, and realizes the characteristics of multimodal information fusion, high-precision target segmentation, strong robustness and strong scalability, realizes accurate identification and segmentation of oil leakage areas, and adapts to other transformer fault detection tasks by simply modifying the prompt words, with good versatility and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power equipment detection, and relates to a detection method of a power transformer oil leakage detection system based on multi-modal prompting and multi-scale segmentation. Background Art

[0002] At present, the commonly used technologies for power transformer oil leakage detection are mainly divided into manual inspection, sensor detection and automatic detection based on image processing.

[0003] 1. Manual inspection:

[0004] In the power industry, manual inspections are the traditional and most commonly used method for detecting transformer oil leaks. Inspectors regularly conduct on-site inspections of transformers, observing for oil stains on the equipment surface and oil accumulation in the transformer foundation to determine if the equipment is leaking. This method is simple, intuitive, and requires no complex equipment. However, manual inspections are inefficient, especially for large transformers, where inspection cycles are long and coverage is limited. Manual inspections are particularly challenging for hidden areas or high-altitude structures, especially in harsh working environments such as high temperature and high humidity, where inspectors' accuracy and safety are difficult to ensure. Manual inspections also carry the risk of missed or false detections, especially when initial symptoms of oil leaks are not obvious and may be overlooked, leading to further exacerbation of the problem later.

[0005] 2. Sensor-based online monitoring system:

[0006] In recent years, automated detection methods have been increasingly used to detect oil leaks in power transformers, particularly through real-time monitoring using sensors. In this approach, sensors are typically installed at key locations on the transformer, such as the oil tank interface and cooler, prone to oil leaks. These sensors can detect changes in oil level, oil pressure fluctuations, or abnormalities in the transformer's internal pressure, and transmit this data in real time to a monitoring center via a signal transmission system. When an anomaly is detected, the system triggers an alarm to alert maintenance personnel for prompt inspection. This approach offers advantages such as real-time monitoring and early warning, effectively mitigating the severity of oil leaks. However, because the sensors must be close to the transformer's surface or internal structure, installing a large number of them on large equipment increases installation complexity and maintenance costs. Furthermore, the sensors themselves are susceptible to environmental factors (such as temperature, humidity, and dust), leading to detection errors.

[0007] 3. Detection based on image processing:

[0008] With the development of computer vision and deep learning technologies, image processing techniques are increasingly being applied to oil leak detection in power transformers. This technology primarily involves installing cameras or drones to continuously capture images of transformers and combining them with image recognition algorithms to analyze the characteristics of oil stains in the images to determine if a leak has occurred. For example, algorithms such as convolutional neural networks (CNNs) can automatically extract features from images and, through multi-layer processing, gradually filter out background information to focus on possible leak locations. This approach, which does not require close proximity to the equipment, is suitable for large-scale, uninterrupted monitoring, offering significant advantages for high-altitude structures or inaccessible areas of transformers. However, image processing methods may perform poorly in complex backgrounds and are susceptible to interference from factors such as lighting and reflections from equipment surfaces. Furthermore, deep learning-based algorithms require large amounts of labeled data for training, and obtaining authentic and diverse oil leak samples in the power industry is difficult, limiting their widespread application.

[0009] Although existing oil leak detection technologies have improved the accuracy and efficiency of detection to varying degrees, they still have several defects and limitations, especially when facing complex environments, hidden locations and subtle leaks.

[0010] 1. Limitations of manual inspection:

[0011] Manual inspections are inefficient, rely heavily on personnel experience, and have a high rate of missed inspections. Transformers are complex structures, and certain high-altitude and hidden locations are difficult to inspect manually. Furthermore, the periodic nature of inspections means that oil leaks can develop significantly between inspections, leading to further damage to the equipment. Manual inspections are also limited by environmental conditions. High temperatures, humidity, and inclement weather can affect both personnel safety and inspection accuracy.

[0012] 2. Limitations of sensor technology:

[0013] While online sensor monitoring provides real-time detection, it is difficult to install and expensive to maintain, especially on large transformers, where multiple sensors must be installed. Furthermore, detection results are susceptible to environmental interference. Sensors are susceptible to damage in high-temperature and high-voltage environments and require regular calibration and maintenance. Furthermore, sensors can only monitor localized information on the transformer surface or near their mounting points, making it difficult to accurately detect oil leaks in hidden components within the transformer.

[0014] 3. Limitations of image processing technology:

[0015] Oil leak detection methods based on image processing perform poorly in complex backgrounds and with varying lighting conditions. This is particularly true when the equipment surface is reflective or the contrast between oil stains and the background color is unclear, leading to false or missed detections. Deep learning algorithms require large amounts of labeled data for training, but obtaining sufficiently diverse samples in transformer oil leak scenarios is challenging. Furthermore, existing image processing methods lack accuracy for detecting leaks in small or hidden areas. Furthermore, current technology faces challenges in efficiently processing large amounts of data and reducing computational costs. In particular, deep learning-based models require extensive labeled data for training, and obtaining labeled data for real transformer oil leaks is costly.

[0016] Therefore, how to improve the generalization ability and detection accuracy of the model when data is limited is also a major bottleneck of existing technologies. Summary of the Invention

[0017] The technical problem to be solved by the present invention is to provide a detection method for a power transformer oil leakage detection system based on multimodal prompts and multi-scale segmentation. By combining the Mask2Former and CLIP models, the accuracy and robustness of oil leakage detection are improved, and the precise identification and segmentation of oil leakage areas are achieved.

[0018] To solve the above technical problems, the technical solution adopted by the present invention is: a detection method of a power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation, which is characterized by comprising the following steps:

[0019] S1, data input and preprocessing. In the data preprocessing stage, the image is normalized and resized to meet the input size requirements of the backbone network. The input data includes images of power transformers from the site and related text prompt information. Word embedding processing is performed as the basis for subsequent multimodal feature fusion.

[0020] S2, feature extraction, backbone network Backbone feature extraction, CLIP image encoder feature extraction and CLIP text encoder feature extraction;

[0021] S3, multimodal feature fusion, the extracted image features and text features are fused through the feature fusion module;

[0022] S4, Transformer decoder decoding, the fused features are processed by multi-layer Transformer decoders. The multi-head attention mechanism of each layer refines the features layer by layer and focuses on the target area;

[0023] S5, segmentation and classification, the feature map finally output by the decoder is segmented and passed through the convolution layer to generate a binary mask, which is used to represent the segmentation result of the oil spill area; classification, the system also generates the classification result of the oil spill area and outputs the category probability through the Softmax classifier;

[0024] S6, model training and fine-tuning. In actual tasks, the system needs to fine-tune the entire network. The model's loss function is a weighted sum of segmentation loss and classification loss. This loss function is optimized through backpropagation, enabling the model to achieve higher accuracy in the specific task of power transformer oil leakage detection.

[0025] S7, output and result display, during the testing phase, the trained and optimized model accepts input images and text prompts, generates segmentation masks and classification results of the oil spill area, and outputs them to the user through the display device.

[0026] In S1, the image is collected by an on-site camera device, and the text prompt is "power transformer oil leakage detection" or "transformer oil leakage area".

[0027] In S2, the backbone network feature extraction image first performs feature extraction through a pre-trained backbone network. The backbone network selects the classic convolutional neural network ResNet or the new visual Transformer structure SwinTransformer. The network is pre-trained on the large-scale dataset ImageNet. After the image feature extraction is completed, a multi-scale feature map is output. The formula is as follows:

[0028] ;

[0029] in, is the input image, is the output multi-scale image feature, Indicates the feature layer.

[0030] In S2, the CLIP image encoder feature extraction is further processed by the image encoder in the CLIP model. The CLIP image encoder is pre-trained on the large-scale dataset OpenAI CLIP to extract image features with multimodal understanding capabilities and output multi-scale visual features:

[0031] ;

[0032] in, It is the multi-scale feature representation of the image in the CLIP model.

[0033] In S2, the prompt text in the CLIP text encoder feature extraction is processed by the text encoder in the CLIP model to obtain the embedded representation of the text; the CLIP text encoder is pre-trained on large-scale text data and is used to extract text features:

[0034] ;

[0035] in, Represents the input text, Embedding for text features.

[0036] In S3, the self-attention mechanism is used to achieve feature alignment and enhancement, which includes modeling long-range dependencies. The process is as follows:

[0037] ;

[0038] in, This is the final fused feature map, which contains the joint information of image and text, and helps to accurately locate the oil spill area.

[0039] In S4, the Transformer decoder decoding formula is expressed as follows:

[0040] ;

[0041] in, For the The decoder uses a multi-head attention mechanism with L layers stacked on top of each other to refine the features of the oil leak area layer by layer, and finally outputs a feature map with the same size as the original image for subsequent segmentation tasks.

[0042] In S5, the segmentation loss is calculated using the binary cross entropy loss function:

[0043] ;

[0044] in, is the true label, is the predicted probability.

[0045] In S5,

[0046] The classification loss is calculated using the cross entropy loss function:

[0047] ;

[0048] in, is the true category label, Predict probabilities for the categories.

[0049] In S6, the overall loss function is as follows:

[0050] ;

[0051] in, is the weight parameter of the mask loss, used to balance the mask loss the importance of is the weight parameter of the classification loss, used to balance the classification loss importance.

[0052] The main beneficial effects of the present invention are:

[0053] Multimodal information fusion: By combining the CLIP and Mask2Former models, the system can effectively fuse visual and textual information to improve the accuracy of oil leak detection, especially in complex scenarios, where textual cues can be used to enhance detection results.

[0054] High-precision target segmentation: Mask2Former effectively addresses the fine-grained segmentation of oil spill areas through multi-level feature extraction and segmentation, ensuring accurate detection of oil spill areas and overcoming the shortcomings of traditional methods in identifying complex backgrounds and small oil spill areas.

[0055] Strong Robustness: The system demonstrates excellent robustness in a variety of complex environments. Whether it is changes in illumination or angle in the image, or the diversity of oil leak locations, the system can accurately locate the oil leak area by combining multimodal features.

[0056] Strong scalability: In addition to oil leakage detection, the system can adapt to other transformer fault detection tasks by simply modifying the prompt words, and has good versatility and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The present invention will be further described below with reference to the accompanying drawings and examples.

[0058] Figure 1 It is a system diagram of the present invention. DETAILED DESCRIPTION

[0059] like Figure 1 A detection method for a power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation is provided, which is characterized by comprising the following steps:

[0060] S1, data input and preprocessing. In the data preprocessing stage, the image is normalized and resized to meet the input size requirements of the backbone network. The input data includes images of power transformers from the site and related text prompt information. Word embedding processing is performed as the basis for subsequent multimodal feature fusion.

[0061] S2, feature extraction, backbone network Backbone feature extraction, CLIP image encoder feature extraction and CLIP text encoder feature extraction;

[0062] S3, multimodal feature fusion, the extracted image features and text features are fused through the feature fusion module;

[0063] S4, step 4: Transformer decoder decoding, the fused features are processed by multi-layer Transformer decoders, and the multi-head attention mechanism of each layer refines the features layer by layer and focuses on the target area;

[0064] S5, segmentation and classification, the feature map finally output by the decoder is segmented and passed through the convolution layer to generate a binary mask, which is used to represent the segmentation result of the oil spill area; classification, the system also generates the classification result of the oil spill area and outputs the category probability through the Softmax classifier;

[0065] S6, model training and fine-tuning. In actual tasks, the system needs to fine-tune the entire network. The model's loss function is a weighted sum of segmentation loss and classification loss. This loss function is optimized through backpropagation, enabling the model to achieve higher accuracy in the specific task of power transformer oil leakage detection.

[0066] S7, output and result display, during the testing phase, the trained and optimized model accepts input images and text prompts, generates segmentation masks and classification results of the oil spill area, and outputs them to the user through the display device.

[0067] Example 1,

[0068] By combining deep learning models with multimodal features to detect oil leakage in power transformers, the following key technical solutions are proposed:

[0069] 1. Feature extraction of input image,

[0070] Using a pre-trained deep convolutional neural network (DCNN), such as ResNet, as a backbone, we perform multi-level feature extraction on the input power transformer image. The extracted image features contain multi-scale information, capturing the characteristics of oil leak areas at different scales within the image, providing high-quality feature representation for subsequent segmentation and classification.

[0071] 2. Mask2Former segmentation model,

[0072] The multi-scale features extracted by Backbone are further processed using the Mask2Former model. Mask2Former incorporates a Transformer architecture that effectively captures long-range dependencies in the image, ensuring accurate segmentation of the oil spill area. Through a multi-layer Transformer decoder, Mask2Former generates a mask image of the target area from information at different levels, classifies the target, and ultimately obtains the segmentation result of the oil spill area in the image.

[0073] 3. Multimodal matching of CLIP model,

[0074] The CLIP model is used to jointly encode input images and text information. The CLIP image encoder converts images of power transformer oil leaks into visual features, while the CLIP text encoder converts input text descriptions, such as "oil leak" or "power transformer," into text features. Through the feature alignment module, the system can match visual features with text features, helping to improve detection accuracy in different scenarios and guiding the detection process with text prompts.

[0075] 4. Multimodal feature fusion,

[0076] The multi-scale, multi-modal features from the Backbone, Mask2Former, and CLIP models are fused and integrated through the feature fusion module to achieve a deep fusion of image and text information, ensuring that the system can understand the input information from multiple perspectives. This fusion approach can better cope with complex backgrounds and blurred features in oil spill areas, improving detection accuracy and robustness.

[0077] 5. Oil leak detection and classification,

[0078] Based on the fused features, the system uses a classification module to perform final classification and labeling of the oil spill area, generating a segmentation mask with the category label and accurately marking the oil spill location in the image. The system allows users to guide the detection process through prompts, such as "oil leak phenomenon," to further enhance the localization and identification of oil spill areas.

[0079] like Figure 1 In this paper, a power transformer oil leakage detection system based on deep learning and multimodal feature fusion was developed. This system combines image and text information, and through the collaborative work of multiple modules, it achieves accurate detection and segmentation of oil leak areas. The input is an image of a power transformer. After multimodal feature extraction and fusion using the Backbone and CLIP models, the Mask2Former model is used to generate the segmentation results of the oil leak area.

[0080] Example 2,

[0081] The steps to achieve accurate identification and segmentation of oil spill areas are as follows:

[0082] Step 1: Data Input and Preprocessing. The input data includes images of power transformers from the field and associated textual information. For example, the images may be captured by on-site cameras, while the textual information may read "Power Transformer Oil Leakage Detection" or "Transformer Oil Leakage Area." During data preprocessing, the images are normalized and resized to meet the input size requirements of the backbone network. The textual information is then word-embedded to serve as the basis for subsequent multimodal feature fusion.

[0083] Step 2: Feature extraction.

[0084] Backbone feature extraction. The input image is first passed through a pre-trained backbone network for feature extraction. The backbone network can be a classic convolutional neural network, such as ResNet, or a new visual Transformer architecture, such as the Swin Transformer. This network has been pre-trained on large-scale datasets, such as ImageNet. After image feature extraction, it outputs a multi-scale feature map, as shown in the following formula:

[0085] ;

[0086] in, is the input image, is the output multi-scale image feature, Indicates the feature layer.

[0087] CLIP image encoder feature extraction. Image features are further processed by the image encoder in the CLIP model. The CLIP image encoder is pre-trained on large-scale datasets, such as OpenAI CLIP, and is used to extract image features with multimodal understanding capabilities, outputting multi-scale visual features:

[0088] ;

[0089] in, It is the multi-scale feature representation of the image in the CLIP model.

[0090] CLIP text encoder feature extraction. The prompt text is processed by the text encoder in the CLIP model to obtain an embedded representation of the text. The CLIP text encoder is also pre-trained on large-scale text data to extract text features:

[0091] ;

[0092] in, Represents the input text, Embedding for text features.

[0093] Step 3: Multimodal feature fusion. The extracted image and text features are fused through the feature fusion module. A self-attention mechanism is used to achieve feature alignment and enhancement, including modeling of long-range dependencies. The process is as follows:

[0094] ;

[0095] in, This is the final fused feature map, which contains the joint information of image and text, and helps to accurately locate the oil spill area.

[0096] Step 4: Transformer decoder decoding. The fused features are processed by a multi-layer Transformer decoder. The multi-head attention mechanism in each layer refines the features layer by layer and focuses on the target area. The formula is as follows:

[0097] ;

[0098] in, For the The decoder uses a multi-head attention mechanism with L layers stacked together to refine the features of the oil spill area layer by layer, ultimately outputting a feature map of the same size as the original image for subsequent segmentation tasks.

[0099] Step 5: Segmentation and classification.

[0100] Segmentation, the feature map output by the decoder is finally passed through the convolution layer to generate a binary mask, which is used to represent the segmentation result of the oil spill area. The segmentation loss is calculated using the binary cross entropy loss function:

[0101] ;

[0102] in, is the true label, is the predicted probability.

[0103] Classification, the system also generates classification results for the oil spill area and outputs the class probability through the Softmax classifier. The classification loss is calculated using the cross entropy loss function:

[0104] ;

[0105] in, is the true category label, Predict probabilities for the classes.

[0106] Step 6: Model training and fine-tuning. The feature extraction modules in this invention, such as Backbone and CLIP encoders, are pre-trained modules. In actual tasks, the system needs to fine-tune the entire network. The model's loss function is a weighted sum of segmentation loss and classification loss. Backpropagation optimizes this loss function, enabling the model to achieve higher accuracy in the specific task of power transformer oil leakage detection. The overall loss function is as follows:

[0107] ;

[0108] in, is the weight parameter of the mask loss, used to balance the mask loss the importance of is the weight parameter of the classification loss, used to balance the classification loss importance.

[0109] Step 7: Output and result display. During the testing phase, the trained and optimized model accepts input images and text prompts to generate segmentation masks and classification results for the oil spill area. The results can be displayed to the user on a display device.

[0110] In the above method,

[0111] Multimodal feature fusion technology effectively improves the accuracy of oil leak detection by fusing visual features with textual information. Compared with existing single-modal methods, this method can extract richer semantic features from different information sources, thereby more accurately locating the oil leak area.

[0112] The CLIP pre-trained model was used in the model construction process for feature extraction and fine-tuning. Its powerful multimodal understanding capabilities significantly improved the model's detection accuracy. Compared with traditional methods, the CLIP model can better handle complex image and text information through its vision-language alignment capabilities.

[0113] The Transformer-based multi-layer decoding architecture uses the Transformer decoder structure to decode and process the fused multimodal features, with the advantage of capturing long-range dependencies. In oil leak detection, this mechanism can more accurately focus on detailed features in the image, especially small oil stains.

[0114] The above embodiments are merely preferred technical solutions of the present invention and should not be construed as limiting the present invention. The embodiments and features in the embodiments of this application may be arbitrarily combined with each other unless they conflict. The scope of protection of the present invention shall be the technical solutions described in the claims, including equivalent alternatives to the technical features of the technical solutions described in the claims. Equivalent alternatives and improvements within this scope are also within the scope of protection of the present invention.

Claims

1. A detection method for a power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation, characterized in that: The steps include: S1, data input and preprocessing. In the data preprocessing stage, the image is normalized and resized to meet the input size requirements of the backbone network. The input data includes images of power transformers from the site and related text prompt information. Word embedding processing is performed as the basis for subsequent multimodal feature fusion. S2, feature extraction, extracts features from the backbone network (Backbone), and outputs a multi-scale feature map after image feature extraction is completed; CLIP image encoder feature extraction, extracts image features with multimodal understanding capabilities; CLIP text encoder feature extraction, extracts text features; S3, multimodal feature fusion, the extracted image features and text features are fused through the feature fusion module; S4, based on the Transformer decoder decoding in the Mask2Former model, the fused features are processed by a multi-layer Transformer decoder. The multi-head attention mechanism in each layer refines the features layer by layer and focuses on the target area; S5, segmentation and classification, the feature map finally output by the decoder is segmented and passed through the convolution layer to generate a binary mask, which is used to represent the segmentation result of the oil spill area; classification, the system also generates the classification result of the oil spill area and outputs the category probability through the Softmax classifier; S6, model training and fine-tuning. In actual tasks, the system needs to fine-tune the entire network. The model's loss function is a weighted sum of segmentation loss and classification loss. This loss function is optimized through backpropagation, enabling the model to achieve higher accuracy in the specific task of power transformer oil leakage detection. S7, output and result display, during the testing phase, the trained and optimized model accepts input images and text prompts, generates segmentation masks and classification results of the oil spill area, and outputs them to the user through the display device.

2. The detection method of the power transformer oil leakage detection system based on multi-modal prompting and multi-scale segmentation according to claim 1 is characterized in that: In S1, the image is captured by an on-site camera device, and the text prompt is "power transformer oil leakage detection" or "transformer oil leakage area".

3. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S2, the backbone network feature extraction image first uses a pre-trained backbone network to extract features. The backbone network uses the classic convolutional neural network ResNet or the new visual Transformer structure Swin Transformer. This network is pre-trained on the large-scale dataset ImageNet. After the image feature extraction is completed, a multi-scale feature map is output. The formula is as follows: ; in, is the input image, is the output multi-scale image feature, Indicates the feature layer.

4. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S2, the CLIP image encoder feature extraction is further processed by the image encoder in the CLIP model. The CLIP image encoder is pre-trained on the large-scale dataset OpenAI CLIP to extract image features with multimodal understanding capabilities and output multi-scale visual features: ; in, It is the multi-scale feature representation of the image in the CLIP model.

5. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S2, the prompt text in the CLIP text encoder feature extraction is processed by the text encoder in the CLIP model to obtain the embedded representation of the text; the CLIP text encoder is pre-trained on large-scale text data and is used to extract text features: ; in, Represents the input text, Embedding for text features.

6. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S3, the self-attention mechanism is used to achieve feature alignment and enhancement, which includes modeling long-range dependencies. The process is as follows: ; in, This is the final fused feature map, which contains the joint information of image and text, and helps to accurately locate the oil spill area.

7. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S4, the Transformer decoder decoding formula is expressed as follows: ; in, For the The decoder uses a multi-head attention mechanism with L layers stacked on top of each other to refine the features of the oil leak area layer by layer, and finally outputs a feature map with the same size as the original image for subsequent segmentation tasks.

8. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S5, the segmentation loss is calculated using the binary cross entropy loss function: ; in, is the true label, is the predicted probability.

9. The detection method of the power transformer oil leakage detection system based on multi-modal prompting and multi-scale segmentation according to claim 1 is characterized in that: In S5, The classification loss is calculated using the cross entropy loss function: ; in, is the true category label, Predict probabilities for the categories.

10. The detection method of the power transformer oil leakage detection system based on multimodal prompting and multi-scale segmentation according to claim 1 is characterized by: In S6, the overall loss function is as follows: ; in, is the weight parameter of the mask loss, used to balance the mask loss the importance of is the weight parameter of the classification loss, used to balance the classification loss importance.

Citation Information

Patent Citations

  • Zero sample image target detection method and device based on deep learning

    CN113255829A

  • CLIP guidance-based multi-scale multi-mode false information detection method and device, electronic equipment and storage medium

    CN117216709A