Infrared image power equipment fault detection method, device, equipment and medium

By introducing prompt word assisting features of the power equipment environment text information in infrared image fault detection, a convolution kernel is dynamically generated and a target positioning attention map is generated, which solves the problems of inconspicuous features and inaccurate positioning in fault detection, and achieves higher detection accuracy and positioning accuracy.

CN120408158AActive Publication Date: 2025-08-01STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510905043.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

In the existing infrared image fault detection technology, the fault thermal signal is weak, the contrast with the normal area is low, and the background noise is severe, resulting in insufficient detection accuracy and reliability. Especially in the power system, the positioning of the fault target is inaccurate, making it difficult to effectively utilize the prior knowledge and text information of the power equipment.

Method used

In the encoding stage, prompt word assist features are introduced, the text information of the power equipment environment is extracted through natural language models, and the convolution kernel is dynamically generated to extract fault features; in the decoding stage, the target positioning attention map is generated, and the prompt word assist features are used to match the low-resolution feature map to improve fault positioning capabilities.

Benefits of technology

It improves the accuracy and accuracy of infrared image fault detection, enhances the ability to extract fault characteristics, significantly improves the accuracy of fault positioning, and ensures the stable operation of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408158A_ABST
    Figure CN120408158A_ABST
Patent Text Reader

Abstract

The invention relates to an infrared image power equipment fault detection method and device, equipment and a medium, and the method comprises the following steps: obtaining environment text data allowed by power equipment as cue words, carrying out lexical element analysis on the cue words, and carrying out feature extraction through a natural language model to obtain cue word auxiliary features; an infrared image of the power equipment is obtained, the infrared image and cue word auxiliary features are input into a fault detection model based on a Unet structure together, a fault detection result is output, the fault detection model comprises a feature coding module and a feature decoding module, the feature coding module dynamically generates a convolution kernel based on the cue word auxiliary features to extract fault features, and the feature decoding module decodes the extracted fault features; in the feature decoding module, for fault features of different scales, cue word auxiliary features are matched with row features and column features of the low-resolution feature map, and a target positioning attention map is generated for auxiliary feature extraction. Compared with the prior art, the method has the advantages of improving the detection precision and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fault detection of power equipment, and particularly to an infrared image power equipment fault detection method, device, equipment and medium. Background Art

[0002] Infrared images can effectively reveal the temperature distribution on the surface and inside of an object by capturing its thermal radiation information. Since infrared imaging does not rely on visible light, it has significant advantages in many fields, such as industrial equipment monitoring, building inspection, medical imaging, military reconnaissance, etc. In the industrial field, infrared images are widely used in tasks such as equipment fault detection, quality control, equipment maintenance, and preventive diagnosis. Especially in the fault detection of electrical equipment, mechanical equipment, etc., infrared images can effectively detect temperature anomalies, thereby identifying potential problems of the equipment. With the development of deep learning technology, image segmentation and object detection have become important research directions in the field of computer vision. In many fields such as medical imaging, remote sensing imaging, and industrial inspection, image segmentation technology is widely used in automated analysis and diagnosis. However, there are still some problems with existing infrared image fault detection technologies.

[0003] In the existing infrared image fault detection technology system, the unclear infrared image features and the susceptibility of the model encoding process to background noise are the key factors limiting the detection accuracy. From a physical perspective, infrared images are formed based on the thermal radiation characteristics of objects. The thermal anomaly signals generated in the fault area are often extremely weak and show low contrast in the image. Taking the early fault detection of industrial equipment as an example, the thermal changes caused by minor electrical short circuits or mechanical wear are only manifested as slight gray-scale differences in the infrared image, making it difficult to clearly distinguish from the normal area, which poses a huge challenge to the extraction of fault features.

[0004] At the same time, the interference of background noise further exacerbates this dilemma. Background noise comes from a wide range of sources, including the dynamic change of ambient temperature, the radiation interference of other surrounding heat sources, and the electronic noise of the infrared imaging device itself. These noises are indiscriminately mixed into the input signal during the encoding stage of the model, interfering with the effective extraction of fault features by the convolutional neural network. Since the convolutional layer and pooling layer are difficult to accurately screen out the key features related to faults from the massive background information when processing these complex signals, the encoded feature representation is deviated, unable to provide a reliable basis for subsequent fault classification and judgment, thus significantly reducing the accuracy and reliability of fault detection.

[0005] For example, in the fault detection of high-voltage transmission lines, partial discharge of insulators or poor contact at joints may only appear as a very small hot spot in the infrared image. Since the characteristics of the fault target cannot be accurately extracted during the encoding stage, it is difficult for the model to accurately restore its position information from the multi-layer complex feature maps during decoding. Moreover, other heat sources in the background, such as the thermal radiation of adjacent poles and towers, and the thermal reflection of the surrounding environment, will be intertwined with the thermal signals of the fault target, interfering with the judgment of the model. The model may misjudge these background thermal signals as fault targets or have a large deviation when locating the fault position. This problem of inaccurate positioning makes it difficult for maintenance personnel to quickly find the fault point based on the detection results, which will not only delay the maintenance time but also may cause the expansion of the fault range and affect the stable operation of the power system.

[0006] Patent CN117011669A constructs an infrared small target detection method and system. By means of convolutional downsampling of single-channel images, multi-channel feature maps of different scales are obtained, details are extracted through convolutional operations, and a sparse sampling correlation attention mechanism is used to complete the conversion of the feature maps. After bilinear interpolation upsampling, multi-feature maps are fused, and the features are converted into pixel binary classification probabilities. However, the sparse sampling attention mechanism is affected by the complex background environment and noise in the infrared image, and it is difficult for the model to focus on the specified area related to the fault, which is not conducive to the detection and positioning of small targets in scenarios with insufficient information.

[0007] Patent CN119399450A constructs a processing system based on multiple network models. The low-rank background estimation network model, sparse target extraction network model, and image reconstruction network model work together, combined with the Res Blocks network model and the collaborative attention network model. Through operations such as convolutional residual blocks and collaborative attention mechanisms, the processing and extraction of the background and targets in the image are realized. However, it is difficult for collaborative attention to accurately capture the fault-related features in the complex infrared images of the power system, and it is difficult for the convolutional neural network to achieve accurate identification of faults without prior knowledge related to the power system.

[0008] Patent CN119831954A discloses a power equipment defect detection method and system based on a multi-modal large model. This method constructs a self-attention device fault detection model and sets up basic tasks for training. This method has achieved certain results in the detection of power equipment defects, but it does not fully utilize text information to assist in infrared image fault detection, resulting in limited detection performance in complex scenarios.

[0009] In summary, the main problems existing in the prior art include: First, the fault thermal signals in infrared images are often weak, with a low contrast to the normal areas, resulting in unclear fault features and affecting the accuracy of fault detection; Second, under the interference of background noise, it is difficult to accurately extract fault-related features during the model encoding stage, affecting the accuracy and reliability of fault detection; Finally, in application scenarios such as power systems, the size of fault targets in infrared images is usually small. Due to the deviation of feature extraction in the encoding stage and the interference of complex heat sources in the background, it is difficult for the model to accurately restore the position of the fault target from the feature map, and there is a problem of low accuracy in fault location during the decoding stage, which is likely to delay the maintenance time and affect the stable operation of the system. In addition, most of the infrared image fault detection methods in the prior art fail to effectively utilize the prior knowledge of power equipment and environmental information, lacking a mechanism to organically combine text information with image features, resulting in limited fault recognition and location capabilities in complex scenarios. Therefore, how to effectively use text information to guide the extraction of fault features and the precise location of targets during the feature encoding and decoding processes remains an urgent problem to be solved. Summary of the Invention

[0010] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method, device, equipment and medium for fault detection of power equipment in infrared images. During the encoding stage, prompt word-assisted features are introduced to alleviate the problems of high noise and unclear features in infrared images, and improve the ability to extract fault features; at the same time, by introducing prompt word-assisted features during the decoding stage, a target location attention map is generated to improve the decoder's ability to locate faults.

[0011] The purpose of the present invention can be achieved by the following technical solutions: According to the first aspect of the present invention, a method for fault detection of power equipment in infrared images is provided. The method includes the following steps: Obtain the allowable environment text data of the power equipment as a prompt word, perform token parsing on the prompt word, and use a natural language model for feature extraction to obtain prompt word-assisted features; Obtain an infrared image of the power equipment, and input it together with the prompt word-assisted features into a fault detection model based on the Unet structure to output a fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. Among them, in the feature encoding module, a convolutional kernel is dynamically generated based on the prompt word-assisted features to extract fault features. In the feature decoding module, for fault features of different scales, the prompt word-assisted features are respectively matched with the row features and column features of the low-resolution feature map to generate a target location attention map for assisting feature extraction.

[0012] As a preferred technical solution, the prompt words are parsed into tokens, and token embeddings are obtained using a linear mapping layer. A position embedding matrix is constructed to obtain position embeddings according to the positions of the tokens, and the token embeddings and position embeddings are added together as the input to the natural language model.

[0013] As a preferred technical solution, the natural language model uses a bidirectional encoding representation model pre-trained on a large-scale language dataset, and the vector corresponding to the CLS head of the bidirectional encoding representation model is output as the prompt word auxiliary feature.

[0014] As a preferred technical solution, the feature encoding module includes multiple encoding blocks. Each encoding block includes a first convolutional layer, a first batch normalization and activation layer, a prompt-assisted convolutional layer, a second batch normalization and activation layer, and a downsampling layer connected in sequence. The input features of each encoding block are processed by the first convolutional layer and the first batch normalization and activation layer to obtain preliminary features. The prompt-assisted convolutional layer uses a convolutional kernel dynamically generated based on the prompt word auxiliary feature to perform a convolution operation on the preliminary features to generate intermediate features, and the intermediate features are output after being processed by the downsampling layer.

[0015] As a preferred technical solution, the dynamic generation of the convolutional kernel of the prompt-assisted convolutional layer is specifically as follows: Average pooling and max pooling are respectively performed on the preliminary features to obtain average pooling features and max pooling features; The average pooling features, max pooling features, and prompt word auxiliary features are concatenated and then passed through a multi-layer perceptron to dynamically generate a convolutional kernel.

[0016] As a preferred technical solution, the feature decoding module includes multiple decoding blocks. The number of decoding blocks is the same as the number of encoding blocks. The features output from the previous decoding block or the last encoding block to the current decoding block are used as the first input features, and the intermediate features of the corresponding encoding block are used as the second input features to perform high-resolution recovery on the input features. The decoding block includes an upsampling layer, a second convolutional layer, a third batch normalization and activation layer, a target localization attention layer, and a third convolutional layer connected in sequence. After the first input features are processed by the upsampling layer, they are concatenated with the second input features, input into the second convolutional layer, and processed by the third batch normalization and activation layer to obtain a low-resolution feature map. The target localization attention layer matches the prompt word auxiliary feature with the row features and column features of the low-resolution feature map respectively for different scales of fault features to generate a target localization attention map, and the third convolutional layer performs feature extraction processing on the low-resolution feature map based on the target localization attention map and then outputs.

[0017] As a preferred technical solution, the target localization attention layer performs the following steps: Perform row slicing and column slicing on the low-resolution feature map: , , wherein, is the low-resolution feature map, W , H , C in are respectively the width, height and number of channels of the input infrared image of the power equipment, is the result of row slicing, is the result of column slicing, indicates that the shape of the feature is ; Perform dot product on the results of row slicing and column slicing respectively with the prompt-assisted feature to obtain row attention and column attention: , , wherein, is the row attention of the i -th row, is the column attention of the j -th column, is the prompt-assisted feature; Perform outer product on the row attention and column attention to generate the target location attention map: , wherein, is the target location attention map, is the row attention vector of all rows, is the column attention vector of all columns.

[0018] According to the second aspect of the present invention, there is provided a fault detection device for power equipment in infrared images, the device comprising: A data acquisition module, configured to acquire power equipment permission environment text data as a prompt word and acquire an infrared image of the power equipment; A prompt-assisted feature extraction module, configured to perform token parsing on the prompt word and perform feature extraction using a natural language model to obtain a prompt-assisted feature; A fault detection module, which is used to input the infrared image of the power equipment and the prompt word auxiliary feature into a fault detection model based on the Unet structure, and output a fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. Among them, in the feature encoding module, a convolution kernel is dynamically generated based on the prompt word auxiliary feature to extract fault features. In the feature decoding module, for fault features of different scales, the prompt word auxiliary feature is respectively matched with the row feature and column feature of the low-resolution feature map to generate a target localization attention map for assisting feature extraction.

[0019] According to the third aspect of the present invention, there is provided an electronic device, including a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the method described above is implemented.

[0020] According to the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0021] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention extracts the features related to fault detection in the prompt word through a natural language model as the prompt word auxiliary feature, introducing additional prior knowledge for the image model in an environment with high data noise and unclear features, and improving the prediction accuracy of the model.

[0022] (2) By introducing prompt word information, the problem of high noise and unclear features in the infrared image is alleviated, and the extraction ability of fault features in the encoding stage is improved. The present invention uses the prompt word auxiliary feature extracted by the natural language model to dynamically generate a convolution kernel in the feature encoding module, enhancing the extraction ability of fault features, and effectively solving the problem of unclear features in the infrared image caused by the weak fault thermal signal and low contrast with the normal area.

[0023] (3) By introducing the prompt word auxiliary feature in the decoding stage, a target localization attention map is generated, improving the localization ability of the decoder for faults. In the feature decoding module of the present invention, the prompt word auxiliary feature is respectively matched with the row feature and column feature of the low-resolution feature map to generate a target localization attention map, effectively solving the problem of low fault localization accuracy caused by the small size of the fault target in the infrared image. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of the method of the present invention; Figure 2 is a schematic diagram of the model structure of the present invention; Figure 3 is a schematic diagram of the process of obtaining the prompt word auxiliary feature of the present invention; Figure 4 Schematic diagram of the process for dynamically generating a convolution kernel according to the present invention; Figure 5 Schematic diagram of the process for generating an object localization attention map according to the present invention; Figure 6 Visual comparison of detection results of different models in an embodiment. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those of ordinary skill in the technical field to which this application belongs. The terms "a", "one", "kind", "the" and the like involved in this application do not indicate a quantity limitation and may represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0027] Embodiment 1 In this embodiment, to solve the problems of insufficient feature extraction in the encoding stage and inaccurate fault location in the decoding stage caused by high data noise and unclear features in the fault detection task of infrared images, a fault detection model integrating multi-modal text information is proposed. This model can alleviate the problems of high noise and unclear features in infrared images by introducing prompt information, improving the ability to extract fault features in the encoding stage. At the same time, by introducing prompt-assisted features in the decoding stage, a target location attention map is generated to improve the decoder's ability to locate faults.

[0028] As Figure 1 shown, this embodiment provides an infrared image power equipment fault detection method, which includes the following steps: S1, Obtain the power equipment allowable environment text data as prompts, perform token parsing on the prompts, and use a natural language model for feature extraction to obtain prompt-assisted features.

[0029] In this embodiment, first, obtain the power equipment allowable environment text data as prompts. The power equipment allowable environment text data refers to the text information describing the normal working environment conditions of power equipment, such as "The normal working temperature range of the transformer is -20°C to 40°C", "The normal working humidity of the switch cabinet does not exceed 85%", etc. These text data can be obtained from the technical manuals, specification standards, or expert knowledge bases of power equipment.

[0030] As Figure 3 shown, perform token parsing on the obtained prompts. Token parsing is the process of decomposing text data into basic language units. The prompts are parsed into tokens through token parsing, and the process is as shown in formula (1), where represents the input word, represents the parsed token. Subsequently, based on the obtained tokens T , use a linear mapping layer to obtain token embeddings. At the same time, to facilitate the model to perceive the position of the tokens in the sequence, construct a position embedding matrix , and obtain the position embedding according to the position of the tokens, where d is the embedding feature dimension. Add the token embeddings and the position embeddings as the input of the natural language model, and the process is as shown in formula (2): Specifically, the token parsing process first splits the text into words or sub-words. For example, the text "The normal operating temperature range of the transformer is from -20°C to 40°C" is split into tokens such as [transformer, normal, operating, temperature, range, from, -20°C, to, 40°C]. Then, each token is converted into a vector representation of a fixed dimension through a linear mapping layer, that is, token embedding. At the same time, a position embedding matrix is constructed to assign a position encoding to each token according to its position in the sequence to preserve the order information between tokens. Finally, the token embedding and the corresponding position embedding are added together to obtain a comprehensive representation containing semantic information and position information, which is used as the input of the natural language model.

[0031] In this embodiment, the natural language model uses a pre-trained Bidirectional Encoder Representations from Transformers (BERT) on a large-scale language dataset, and outputs the vector corresponding to the CLS head of the bidirectional encoding representation model as the prompt word auxiliary feature. The bidirectional encoding representation model is a deep learning model that can consider the context information before and after the text at the same time. Through pre-training on a large-scale language dataset, it can effectively capture the semantic features of the text. Add a special classification token [CLS] at the beginning of the input sequence of the model, and the final hidden state vector corresponding to this token can be used as the aggregated representation of the entire sequence. In this embodiment, the vector corresponding to this CLS head is output as the prompt word auxiliary feature for subsequent fault detection tasks, that is: By fusing the prompt word auxiliary feature, the model can supplement the relevant prior knowledge lacking in the image, thereby effectively assisting the subsequent encoding and decoding tasks.

[0032] S2. Obtain the infrared image of the power equipment and input it into the fault detection model based on the Unet structure together with the prompt word auxiliary feature, and output the fault detection result.

[0033] The fault detection model uses Unet as the basic architecture of the network, and obtains multi-scale information through successive downsampling and upsampling, so as to effectively capture the local features and global information in the image. Specifically, the faults in the infrared image include both large-scale overall defects and small-scale local details. In order to be able to capture this information at a single resolution, multi-scale pooling operations are used in the network so that the convolutional kernels can obtain features of different scales.

[0034] The fault detection model includes a feature encoding module and a feature decoding module. Among them, the feature encoding module includes L encoding blocks. Each encoding block extracts local features of the image through convolutional operations, reduces the resolution of the feature map through pooling operations to expand the receptive field and capture global information. At the same time, a convolutional kernel is dynamically generated based on the prompt word-assisted feature to extract fault features, improving the ability of the encoder to capture key features of infrared images. In the feature decoding module, the low-resolution feature map is upsampled through deconvolution operations to restore high-resolution detailed features. During this process, the decoding layer not only restores the image details but also retains multi-scale information by fusing with the high-level features in the feature encoding module, finally generating accurate fault detection results. At the same time, the present invention uses a natural language model to assist in fault location. For different-scale features, the prompt word-assisted features of the natural language model are respectively matched with the row features and column features of the feature map to generate a target location attention map, thus effectively improving the sensitivity of the model to the fault location.

[0035] As Figure 2 shown, each encoding block includes a first convolutional layer, a first batch normalization and activation layer, a prompt-assisted convolutional layer, a second batch normalization and activation layer, and a downsampling layer connected in sequence. The input features of each encoding block are processed by the first convolutional layer and the first batch normalization and activation layer to obtain preliminary features. The prompt-assisted convolutional layer uses the convolutional kernel dynamically generated based on the prompt word-assisted feature to perform a convolutional operation on the preliminary features to generate intermediate features, and the intermediate features are output after being processed by the downsampling layer. Encoding blocks 2 - L except encoding block 1 use the output of the previous encoding block as input features, and encoding block 1 uses the infrared image of the power equipment as input features. The overall structure of the encoding process is shown in formula (4). Among them, represents the operation of the encoding block, represents the input features of the l th encoding block, represents the l th encoding block,

[0036] In the feature encoding module, the first convolutional layer extracts conventional features through convolution, and then improves the training stability through the first batch normalization and activation layer to obtain preliminary features , as shown in formula (5). Among them, represents convolution, represents batch normalization, represents the activation function.

[0037] Subsequently, in order to accurately capture fault-related features in infrared data with complex backgrounds and unclear features, the prompt-assisted convolutional layer in this embodiment uses the prompt-assisted features of the natural language model to jointly and dynamically generate convolutional kernels with the current feature map for feature extraction. Specifically, as Figure 4 shown, the following steps are included: Perform average pooling and max pooling on the preliminary features respectively to obtain the average pooling feature and the max pooling feature; global pooling is to perform average pooling on the entire feature map to obtain the global information of each channel; max pooling extracts the most significant features of each channel. These two pooling operations respectively obtain the global statistical information and local significant information of the features, and can describe the characteristics of the feature map from different perspectives.

[0038] After concatenating the average pooling feature, the max pooling feature, and the prompt-assisted feature, dynamically generate the convolutional kernel through a Multilayer Perceptron (MLP), enabling the convolutional operation to be dynamically adjusted according to the input image and prompt information, and improving the pertinence and effectiveness of feature extraction.

[0039] The above process is shown in formula (6): where is max pooling, is average pooling, MLP is the Multilayer Perceptron, represents the convolutional operation, is the intermediate feature.

[0040] Finally, perform pooling operation on the intermediate feature for downsampling, and input the result into the next encoding block, as shown in formula (7). The feature decoding module consists of L decoding blocks. Decoding block L, corresponding to encoding block L, uses the output features of encoding block L as its first input features. Decoding blocks 1-L-1, corresponding to encoding blocks 1-L-1, use the output features of decoding blocks 2-L as their first input features. Furthermore, decoding blocks 1-L use the intermediate features of their corresponding encoding blocks 1-L as their second input features, restoring the input features at high resolution. In infrared small target detection, the decoder faces difficulties in locating fault locations. Small infrared targets are small and have less distinct features, making them difficult to discern against complex backgrounds. When processing the encoded feature information, the decoder must convert abstract features into specific target locations. However, the spatial information of small targets is easily lost during the feature mapping process, making it difficult for the decoder to accurately reconstruct their spatial location. Furthermore, infrared images contain numerous interferences, such as similar temperature regions in the background. These can easily lead the decoder to misidentify interference as the target, or cause the positioning results to deviate from the true location, making it difficult to accurately pinpoint the fault location.

[0041] By learning from massive amounts of text data, natural language models can extract accurate and rich feature information from descriptive prompts. These features cover key aspects such as target characteristics and scene associations. When integrated into the decoder for infrared small target detection, these features serve as powerful positioning guidance, helping the decoder accurately match the location information corresponding to the fault target in a complex feature space. This effectively overcomes positioning challenges caused by small targets and background interference, significantly improving fault location accuracy.

[0042] In this embodiment, Figure 2 As shown in the figure, the decoding block includes an upsampling layer, a second convolutional layer, a third batch normalization and activation layer, a target positioning attention layer and a third convolutional layer connected in sequence. After the first input feature is processed by the upsampling layer, it is spliced with the second input feature, input into the second convolutional layer, and processed by the third batch normalization and activation layer to obtain a low-resolution feature map. The target positioning attention layer matches the auxiliary features of the prompt word with the row features and column features of the low-resolution feature map for fault features of different scales to generate a target positioning attention map. The third convolutional layer extracts features from the low-resolution feature map based on the target positioning attention map and outputs it.

[0043] No. l (1≤ l < L ) decoding blocks before the output of the previous decoding block and the intermediate features of the corresponding encoding block As input, the result is passed to the next decoding block, as shown in formula (8). in, Represents the operation of decoding a block.

[0044] The decoded block is also composed of two layers of convolution. First, the features are upsampled to obtain features with a larger scale , which are concatenated with the intermediate features of the corresponding encoding block for convolution, followed by batch normalization and activation, as shown in Equation (9). To use the prompt-assisted features for auxiliary localization, this embodiment proposes a target localization attention layer based on prompt assistance, as Figure 5 shown, which performs the following steps: The low-resolution feature map is sliced by rows and columns, as shown in Equation (10): where is the low-resolution feature map, W , H , C in are the width, height, and number of channels of the input infrared image of the power equipment respectively, is the result of row slicing, is the result of column slicing, indicates that the shape of the feature is .

[0045] The results of row slicing and column slicing are respectively dot-producted with the prompt-assisted features to obtain row attention and column attention, as shown in Equation (11): ]>< / where is the row attention of the i th row, is the column attention of the j th column, is the prompt-assisted feature.

[0046] The outer product of row attention and column attention is calculated to generate a target localization attention map, as shown in Equation (12): where is the target localization attention map, is the row attention vector of all rows, , is the column attention vector of all columns, .

[0047] Row slicing and column slicing operations extract features from the low-resolution feature map row by row and column by column, obtaining feature vectors for each row and column. By calculating the similarity (dot product) between these feature vectors and the prompt word-assisted features, row attention and column attention can be obtained, indicating the degree of match between each row / column and the fault features described by the prompt words. Combining row attention and column attention through an outer product operation generates a two-dimensional target localization attention map, which highlights the areas where faults may exist and guides subsequent feature extraction to pay more attention to these areas.

[0048] Finally, the third convolutional layer outputs after performing feature extraction processing on the low-resolution feature map based on the target localization attention map, further enhancing the feature representation of the fault area, suppressing the interference of non-fault areas, and improving the accuracy of fault detection. This process is shown in Equation (13): Through prompt word feature-assisted localization, the model can accurately locate the fault area, allocate more attention to it, and thus effectively improve the accuracy of fault localization.

[0049] By introducing additional prompt word information, the network of the present invention can more accurately process fault information in infrared images under complex environments and achieve precise localization, improving the accuracy of fault detection and segmentation.

[0050] Through the above process, the fault detection model outputs fault detection results, including information such as fault type, fault location, and fault severity. These results can help power maintenance personnel promptly discover equipment faults, take corresponding maintenance measures, and ensure the safe and stable operation of the power system.

[0051] Example 2 In this example, sufficient experiments were conducted on a cable fault detection dataset. The experimental results are shown in Table 1. The Intersection over Union (IoU) of the present invention reaches 0.6042, and the Probability of Detection (PD) reaches 0.7816, significantly higher than other methods. The False Alarm (FA) rate is only 0.0029, lower than other methods, indicating that the present invention has obvious advantages in terms of fault detection performance.

[0052] Table 1 Prediction Metrics of Different Models on the Cable Infrared Fault Detection Dataset The visualization effect of the detection results of different methods is as Figure 6As shown, compared with other models, the targets detected by the present invention are more complete. The lines of the target objects in the figure basically do not show obvious breaks, and the coincidence degree with the true labels is relatively high. At the same time, there are fewer noise points and false detection areas in the detection results, and the background is basically black, indicating that it can better distinguish the target from the background, reduce the false alarm rate, and provide a reliable guarantee for the stable operation of the power system.

[0053] Embodiment 3 The above is the introduction of the method embodiment. The following further illustrates the solution of the present invention through the device embodiment.

[0054] This embodiment provides an infrared image power equipment fault detection device, which includes: A data acquisition module, configured to acquire power equipment allowable environment text data as a prompt word and acquire an infrared image of the power equipment; A prompt word assisted feature extraction module, configured to perform token parsing on the prompt word and perform feature extraction using a natural language model to obtain prompt word assisted features; A fault detection module, configured to input the infrared image of the power equipment and the prompt word assisted features into a fault detection model based on the Unet structure, and output a fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. Among them, in the feature encoding module, a convolution kernel is dynamically generated based on the prompt word assisted features to extract fault features. In the feature decoding module, for fault features of different scales, the prompt word assisted features are respectively matched with the row features and column features of the low-resolution feature map to generate a target location attention map for assisting feature extraction.

[0055] In this embodiment, the infrared image power equipment fault detection device includes three main modules: a data acquisition module, a prompt word assisted feature extraction module, and a fault detection module.

[0056] The data acquisition module is responsible for acquiring two types of data: power equipment allowable environment text data and infrared images of power equipment. The power equipment allowable environment text data is text information describing the normal working environment conditions of power equipment, such as temperature range, humidity requirements, load capacity, etc. These text data can be obtained from the technical manuals, specification standards or expert knowledge bases of power equipment and used as prompt words for subsequent processing. The infrared image of the power equipment is an image collected by an infrared thermal imaging device and can reflect the surface temperature distribution of the power equipment. The data acquisition module can connect to the infrared thermal imaging device through a wired or wireless network to obtain the infrared image of the power equipment in real time, or read historical infrared image data from an image database.

[0057] The prompt-assisted feature extraction module receives the prompts provided by the data acquisition module, performs token parsing on them, and uses a natural language model for feature extraction to obtain prompt-assisted features. Token parsing is the process of decomposing text data into basic language units, including operations such as word segmentation and tokenization. The natural language model uses a pre-trained bidirectional encoding representation model on a large-scale language dataset, such as BERT (Bidirectional Encoder Representations from Transformers). This model can understand the semantic information of the text and convert the text into a high-dimensional vector representation. The prompt-assisted feature extraction module outputs the vector corresponding to the CLS head of the bidirectional encoding representation model as the prompt-assisted feature to guide the subsequent fault detection process.

[0058] The fault detection module is the core part of the device. It inputs the infrared image of the power equipment and the prompt-assisted features into a fault detection model based on the Unet structure and outputs the fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. In the feature encoding module, convolutional kernels are dynamically generated based on the prompt-assisted features to extract fault features. In the feature decoding module, for fault features of different scales, the prompt-assisted features are respectively matched with the row features and column features of the low-resolution feature map to generate a target localization attention map for assisting feature extraction.

[0059] The feature encoding module includes multiple encoding blocks. Each encoding block includes a first convolutional layer, a first batch normalization and activation layer, a prompt-assisted convolutional layer, a second batch normalization and activation layer, and a downsampling layer connected in sequence. The prompt-assisted convolutional layer is an innovative design. It uses the prompt-assisted features to dynamically generate convolutional kernels and performs convolutional operations on the preliminary features to generate intermediate features. This way of dynamically generating convolutional kernels enables the model to adjust the feature extraction strategy according to different prompt information and improves the recognition ability for specific fault types.

[0060] The feature decoding module includes multiple decoding blocks, and the number of decoding blocks is the same as the number of encoding blocks. Each decoding block includes an upsampling layer, a second convolutional layer, a third batch normalization and activation layer, a target localization attention layer, and a third convolutional layer connected in sequence. The target localization attention layer is another innovation point. It matches the prompt-assisted features with the row features and column features of the low-resolution feature map respectively to generate a target localization attention map. This attention mechanism can guide the model to focus on the areas where faults may exist and improve the accuracy of fault detection.

[0061] The fault detection module finally outputs the fault detection results, including information such as the fault type, fault location, and fault severity. These results can be visually presented through a display device or transmitted over a network to a remote monitoring center to help power maintenance personnel promptly detect equipment faults and take corresponding repair measures.

[0062] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein.

[0063] Embodiment 4 An electronic device includes a memory and a processor. A computer program is stored on the memory. When the processor executes the program, it implements the infrared image power equipment fault detection method as described in Embodiment 1.

[0064] In this embodiment, the electronic device can be a server, a workstation, a personal computer, a laptop computer, a tablet computer, or an embedded device, etc. The electronic device includes two main hardware components, namely a memory and a processor.

[0065] The memory is used to store computer programs and data. The memory can include a non-volatile storage medium and a volatile storage medium. The non-volatile storage medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, or other types of non-volatile storage devices. The volatile storage medium can be a random access memory (RAM), which serves as an external cache memory. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM), etc.

[0066] The computer program stored in the memory includes program codes for implementing the infrared image power equipment fault detection method. These program codes include a module for obtaining the allowable environment text data of the power equipment as a prompt word, a module for performing token parsing on the prompt word, a module for feature extraction using a natural language model, a module for obtaining the infrared image of the power equipment, a fault detection model module based on the Unet structure, etc. These modules are stored in the memory in the form of program codes and are loaded and executed by the processor.

[0067] The processor can be a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor is responsible for executing the computer programs stored in the memory and implementing the various steps of the infrared image power equipment fault detection method.

[0068] When the electronic device starts and runs the infrared image power equipment fault detection program, the processor first obtains the power equipment allowable environment text data as a prompt word from the memory or an external data source, then performs token parsing on the prompt word, and uses a natural language model for feature extraction to obtain prompt word assisted features. Next, the processor obtains the infrared image of the power equipment, which may be obtained in real time from a connected infrared camera, or may read a saved infrared image from the memory or an external storage device.

[0069] The processor inputs the infrared image of the power equipment and the prompt word assisted features into a fault detection model based on the Unet structure, which is stored in the memory in the form of program code. The fault detection model includes a feature encoding module and a feature decoding module. In the feature encoding module, the processor dynamically generates convolutional kernels based on the prompt word assisted features to extract fault features. In the feature decoding module, for fault features of different scales, the processor matches the prompt word assisted features with the row features and column features of the low-resolution feature maps respectively to generate a target location attention map for assisting feature extraction.

[0070] Finally, the processor outputs the fault detection results through the fault detection model, including information such as the fault type, fault location, and fault severity. These results can be displayed on the display screen of the electronic device or transmitted to other devices or systems through a network interface for reference by power maintenance personnel.

[0071] Embodiment 5 A computer-readable storage medium stores a computer program thereon, and when the program is executed by a processor, it implements the infrared image power equipment fault detection method as described in Embodiment 1.

[0072] In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to: magnetic storage devices, such as hard disks, floppy disks, magnetic tapes; optical storage devices, such as compact discs (CDs), digital versatile discs (DVDs), Blu-ray discs; solid-state storage devices, such as solid-state drives (SSDs), flash drives, memory cards; or any combination of the above storage media.

[0073] A computer program is stored on the computer-readable storage medium, and the program includes program codes for implementing a fault detection method for power equipment in infrared images. When the program is executed by a processor, the processor performs the following operations: Obtain the power equipment allowable environment text data as a prompt word, perform token parsing on the prompt word, and use a natural language model to extract features to obtain prompt word auxiliary features. The prompt word can be text information describing the normal working environmental conditions of the power equipment. The natural language model uses a bidirectional encoding representation model pre-trained on a large-scale language dataset, and outputs the vector corresponding to the CLS head as the prompt word auxiliary feature.

[0074] Obtain the infrared image of the power equipment, and input it together with the prompt word auxiliary feature into a fault detection model based on the Unet structure to output a fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. In the feature encoding module, convolutional kernels are dynamically generated based on the prompt word auxiliary feature to extract fault features. In the feature decoding module, for fault features of different scales, the prompt word auxiliary feature is respectively matched with the row feature and column feature of the low-resolution feature map to generate a target localization attention map for assisting feature extraction.

[0075] In this way, when the program stored on the computer-readable storage medium is executed by the processor, the fault detection method for power equipment in infrared images can be implemented, and the accuracy and efficiency of power equipment fault detection can be improved.

[0076] As described above, the above is only a specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An infrared image power equipment fault detection method, characterized in that, The method includes the following steps: Obtain the allowed environmental text data of the power equipment as a prompt word, perform token parsing on the prompt word, and use a natural language model to extract features to obtain prompt word auxiliary features; Obtain the infrared image of the power equipment, and input it together with the prompt word auxiliary features into a fault detection model based on the Unet structure to output a fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. Among them, in the feature encoding module, a convolution kernel is dynamically generated based on the prompt word auxiliary features to extract fault features. In the feature decoding module, for fault features of different scales, the prompt word auxiliary features are respectively matched with the row features and column features of the low-resolution feature map to generate a target localization attention map for assisting feature extraction.

2. The infrared image power equipment fault detection method according to claim 1, characterized in that The prompt word undergoes token parsing to obtain tokens, and a token embedding is obtained using a linear mapping layer. A position embedding matrix is constructed to obtain a position embedding according to the position of the tokens, and the token embedding and the position embedding are added together as the input of the natural language model.

3. The infrared image power equipment fault detection method according to claim 1, wherein, The natural language model uses a bidirectional encoding representation model pre-trained on a large-scale language dataset, and the vector corresponding to the CLS head of the bidirectional encoding representation model is output as the prompt word auxiliary feature.

4. The infrared image power equipment fault detection method according to claim 1, characterized in that, The feature encoding module includes multiple encoding blocks. Each encoding block includes a first convolutional layer, a first batch normalization and activation layer, a prompt auxiliary convolutional layer, a second batch normalization and activation layer, and a downsampling layer connected in sequence. The input feature of each encoding block is processed by the first convolutional layer and the first batch normalization and activation layer to obtain a preliminary feature. The prompt auxiliary convolutional layer performs a convolution operation on the preliminary feature using a convolution kernel dynamically generated based on the prompt word auxiliary features to generate an intermediate feature, and the intermediate feature is output after being processed by the downsampling layer.

5. The infrared image power equipment fault detection method according to claim 4, wherein The dynamic generation of the convolution kernel of the prompt auxiliary convolutional layer is specifically as follows: Perform average pooling and max pooling on the preliminary feature respectively to obtain an average pooling feature and a max pooling feature; The average pooling feature, the max pooling feature, and the prompt word auxiliary feature are concatenated and then a convolution kernel is dynamically generated through a multi-layer perceptron.

6. A method for detecting faults in infrared image power equipment according to claim 4, characterized in that, The feature decoding module includes multiple decoding blocks. The number of decoding blocks is the same as the number of encoding blocks. The feature output from the previous decoding block or the last encoding block to the current decoding block is used as the first input feature, and the intermediate feature of the corresponding encoding block is used as the second input feature to perform high-resolution restoration on the input feature. The decoding block includes an upsampling layer, a second convolutional layer, a third batch normalization and activation layer, a target localization attention layer, and a third convolutional layer connected in sequence. After the first input feature is processed by the upsampling layer, it is concatenated with the second input feature, input into the second convolutional layer, and processed by the third batch normalization and activation layer to obtain a low-resolution feature map. The target localization attention layer matches the prompt word auxiliary features with the row features and column features of the low-resolution feature map respectively for fault features of different scales to generate a target localization attention map. The third convolutional layer performs feature extraction processing on the low-resolution feature map based on the target localization attention map and then outputs.

7. An infrared image power equipment fault detection method according to claim 6, characterized in that, The described target localization attention layer performs the following steps: Perform row slicing and column slicing on the low-resolution feature map: , , Among them, is a low-resolution feature map, W , H , C in are the width, height, and number of channels of the input infrared image of the power equipment respectively, is the result of row slicing, is the result of column slicing, indicates that the shape of the feature is ; Obtain row attention and column attention by taking the dot product of the results of row slicing and column slicing with the prompt word auxiliary feature respectively: , , Among them, is the row attention of the i th row, is the column attention of the j th column, is the prompt word auxiliary feature; Perform an outer product of the row attention and the column attention to generate a target localization attention map: , Among them, is the target location attention map, is the row attention vector for all rows, is the column attention vector for all columns.

8. An infrared image power equipment fault detection device, characterized in that, The device includes: A data acquisition module for acquiring power equipment permitted environment text data as a prompt word and acquiring an infrared image of the power equipment; A prompt word auxiliary feature extraction module for performing token parsing on the prompt word and using a natural language model for feature extraction to obtain a prompt word auxiliary feature; A fault detection module for inputting the infrared image of the power equipment and the prompt word auxiliary feature into a fault detection model based on the Unet structure to output a fault detection result. The fault detection model includes a feature encoding module and a feature decoding module. Among them, in the feature encoding module, a convolutional kernel is dynamically generated based on the prompt word auxiliary feature to extract fault features. In the feature decoding module, for fault features of different scales, the prompt word auxiliary feature is respectively matched with the row feature and the column feature of the low-resolution feature map to generate a target localization attention map for assisting feature extraction.

9. An electronic device, comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Pre-training language processing method and device

    CN114065771A

  • Auxiliary retrieval method fusing knowledge graph and large language model

    CN117633252A

  • Detection method of power transformer oil leakage detection system based on multi-mode prompt and multi-scale segmentation

    CN119646668A

  • Track prediction method and device based on pre-trained large language model

    CN119691448A

  • Power equipment defect detection method and system based on multi-modal large model

    CN119831954A