Equipment fault detection method and device
By performing multi-scale feature extraction, feature fusion and self-attention processing on infrared images, the problem of insufficient accuracy of equipment fault detection is solved, and more efficient fault recognition is achieved.
Patent Information
- Application Number
- CN202510376840.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to accurately identify equipment failures, especially in the case of complex equipment structures and diverse fault characteristics.
By performing multi-scale feature extraction and feature fusion processing on infrared images, and combining self-attention processing, the fault detection results are determined.
It improves the accuracy of equipment failure detection, can more comprehensively reflect the local details and global characteristics of the equipment, and adapts to complex structures and diverse fault characteristics.
Smart Images

Figure CN120147755A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a device fault detection method and apparatus. Background Art
[0002] Fault detection of a device based on the infrared image of the device is beneficial to timely discover possible faults or hidden dangers of the device. For example, fault detection of a power device based on the infrared image of the power device can help timely discover hidden dangers or faults existing in the power device.
[0003] However, due to the complex structure of the device and the diversity of fault feature manifestations, it is currently impossible to accurately identify possible faults of the device based on the infrared image. Summary of the Invention
[0004] On the one hand, this application provides a device fault detection method, including:
[0005] Obtain an infrared image of a device to be detected;
[0006] Perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature;
[0007] Perform self-attention processing on the infrared image to obtain a second feature;
[0008] Based on the first feature and the second feature, determine a fault detection result corresponding to the infrared image.
[0009] In a possible implementation manner, performing multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature includes:
[0010] Perform multi-scale feature extraction on the infrared image to obtain image features of multiple scales;
[0011] Perform enhancement and fusion processing on the image features of the multiple scales in the channel dimension and the spatial dimension to obtain a first feature.
[0012] In another possible implementation manner, the multiple scales include a first scale and a second scale;
[0013] The performing enhancement and fusion processing on the image features of the multiple scales in the channel dimension and the spatial dimension to obtain a first feature includes:
[0014] Perform channel enhancement processing and spatial enhancement processing on the image features corresponding to the first scale to obtain enhanced image features corresponding to the first scale;
[0015] Add the image features of the second scale to the enhanced image features corresponding to the first scale, perform channel enhancement processing and spatial enhancement processing on the added features obtained by the addition, to obtain the enhanced image features of the second scale, where the second scale is greater than the first scale;
[0016] Perform fusion processing on the enhanced image features corresponding to the first scale and the enhanced image features corresponding to the second scale to obtain a first feature.
[0017] In another possible implementation manner, the multi-scale feature extraction of the infrared image to obtain image features of multiple scales includes:
[0018] Extract the features of the infrared image by using feature extraction modules of multiple different scales to obtain initial features of multiple scales corresponding to the infrared image;
[0019] Perform normalization processing on the initial features of multiple scales in the channel dimension to obtain image features of multiple different scales.
[0020] In another possible implementation manner, the self-attention processing based on the infrared image to obtain a second feature includes:
[0021] Decompose the base layer and the detail layer of the infrared image;
[0022] Determine the base layer as the target image, and perform self-attention processing on the target image to obtain a second feature.
[0023] In another possible implementation manner, the self-attention processing based on the infrared image to obtain a second feature includes:
[0024] Perform segmentation and feature mapping processing on the infrared image to obtain multiple feature blocks corresponding to the infrared image;
[0025] Perform feature recombination on the multiple feature blocks corresponding to the infrared image to obtain a recombined feature;
[0026] Perform attention conversion on the recombined feature, and determine a second feature based on the obtained multi-head attention features after conversion.
[0027] In another possible implementation manner, the determining a second feature based on the obtained multi-head attention features after conversion includes:
[0028] Perform semantic processing on the obtained multi-head attention features to obtain semantic features of the infrared image;
[0029] Perform feature splicing processing on the multi-head attention features and the semantic features to obtain a second feature.
[0030] In yet another possible implementation, determining the fault detection result corresponding to the infrared image based on the first feature and the second feature includes:
[0031] Performing a fusion process on the first feature and the second feature to obtain a fused feature;
[0032] Determining at least one region of interest to be detected in the infrared image;
[0033] Based on the mapping relationship between the infrared image and the fused feature in the spatial dimension, determining at least one set of feature data corresponding to the at least one region of interest in the fused feature;
[0034] Based on the at least one set of feature data, determining the fault detection result of the at least one region of interest.
[0035] In yet another possible implementation, obtaining the infrared image of the device to be detected includes:
[0036] Performing denoising and upsampling on the initial infrared image of the device to be detected to obtain a high-frequency infrared image;
[0037] Performing low-pass filtering on the high-frequency infrared image to obtain a low-frequency infrared image;
[0038] Determining the difference image between the high-frequency infrared image and the low-frequency infrared image;
[0039] Performing linear amplification on the pixel values of each pixel point in the difference image to obtain an amplified high-frequency image;
[0040] Superimposing the amplified high-frequency image and the high-frequency infrared image to obtain the infrared image to be detected.
[0041] In another aspect, the present application further provides a device fault detection apparatus, including:
[0042] An image acquisition unit, configured to acquire an infrared image of a device to be detected;
[0043] A first feature determination unit, configured to perform multi-scale feature extraction and feature fusion on the infrared image to obtain a first feature;
[0044] A second feature determination unit, configured to perform self-attention processing based on the infrared image to obtain a second feature;
[0045] A fault detection unit, configured to determine the fault detection result corresponding to the infrared image based on the first feature and the second feature. Description of the Drawings
[0046] In conjunction with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.
[0047] Figure 1 It is a schematic flowchart of a device fault detection method provided by this application;
[0048] Figure 2 It is a schematic flowchart of multi-scale feature extraction and feature fusion processing for an infrared image;
[0049] Figure 3 It is a schematic diagram of the implementation principle framework for the first feature extraction branch in this application to extract the first feature of the infrared image;
[0050] Figure 4 It is a schematic implementation flowchart for obtaining the second feature by performing self-attention processing based on the infrared image in this application
[0051] Figure 5 It is another schematic flowchart of the device fault detection method provided by this application;
[0052] Figure 6 It is another schematic flowchart of the device fault detection method provided by this application;
[0053] Figure 7 It is an example diagram of the implementation principle framework of the device fault detection method of this application;
[0054] Figure 8 It is a schematic diagram of the composition structure of a device fault detection device provided by this application;
[0055] Figure 9 It is a schematic diagram of the composition structure of an electronic device provided by this application. Specific Embodiments
[0056] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments section of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0057] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of this application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0058] As Figure 1 , a schematic flowchart of a device fault detection method provided by this application is shown. The method of this embodiment can be applied to an electronic device, which can be a tablet computer, a notebook computer, a desktop computer or other intelligent terminals; it can also be a device node in a server or a cloud platform, etc., without limitation.
[0059] The method of this embodiment may include the following steps S101 to S104:
[0060] S101, obtain an infrared image of the device to be detected.
[0061] Among them, the infrared image is also called a thermal infrared image, which is obtained by an infrared sensing device such as an infrared camera or a thermal infrared scanner to collect an image of the whole or part of the device to be detected. The method of using an infrared sensing device such as an infrared camera or a thermal infrared scanner to collect an image of the device to be detected may include, but is not limited to, manually holding an infrared sensing device to take a picture, using a drone to carry an infrared sensing device to take a picture, using an inspection trolley to carry an infrared sensing device to take a picture, and using a fixed monitoring method to take a picture, etc.
[0062] Among them, the device to be detected is a device that needs to be detected for faults. There are many possibilities for the device to be detected. For example, the device to be detected can be an electrical device or a factory device, etc. According to different specific application scenarios, it can be all different. For example, the device to be detected can be an electrical device, and an infrared camera is used to collect an image of the electrical device to obtain an infrared image of the electrical device.
[0063] S102, perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature.
[0064] For example, perform multi-scale feature extraction on the infrared image and perform feature fusion processing based on the features of multiple scales extracted.
[0065] Among them, extracting multi-scale features from the infrared image is to extract hierarchical features from the overall to local details of the infrared image, so that the features of different scales can more comprehensively reflect the local detail information of the device to be detected, thus being able to adapt to the characteristics of the complex structure of the device to be detected and the diverse forms of fault feature manifestations.
[0066] S103, perform self-attention processing on the infrared image to obtain a second feature.
[0067] For example, at least the multi-head self-attention model can be used to process the infrared image to obtain a second feature.
[0068] It can be understood that performing self-attention processing on the infrared image can capture the long-range dependence relationships between different regions in the infrared image, enabling the second feature to focus on expressing the overall features of the infrared image.
[0069] S104, determine the fault detection result corresponding to the infrared image based on the first feature and the second feature.
[0070] Among them, there are multiple possibilities for the specific implementation of fault detection based on the first feature and the second feature. For example, based on the first feature and the second feature, a trained detection model can be used to determine the fault detection result. Among them, the detection model can be a deep learning model such as a neural network model, or other types of artificial intelligence models, etc., and there is no limitation on this.
[0071] Among them, the fault detection result can represent one or more of the fault type in the device to be detected, the probability of the existence of the fault type, and the fault area in the infrared image, etc. Among them, the fault type can include, but is not limited to, no fault or specific fault conditions when there is a fault, etc.
[0072] According to the different types of the device to be detected, the fault conditions will also be different. For example, taking the device to be detected as an electrical device, the fault conditions that may occur in the electrical device can include, but are not limited to, overall overheating fault, arc discharge fault, insulator damage fault, fuse melting fault, component wear fault, and insulating oil leakage fault, etc.
[0073] As can be seen from the above content, in this application, after obtaining the infrared image,
[0074] It will not only perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain the first feature, but also perform self-attention processing on the infrared image to obtain the second feature. Since the features of multiple scales extracted from the infrared image can more comprehensively reflect the local detail information of the device to be detected, the first feature obtained by fusing the features of multiple scales can comprehensively reflect the local detail information of the device to be detected. And performing self-attention processing on the infrared image can capture the long-range dependence relationship between different regions in the infrared image, so that the second feature can focus on expressing the global overall features of the infrared image and form a complement to the local image features corresponding to the first feature. Therefore, through the first feature and the second feature, the local detail features and global overall features of the infrared image can be more comprehensively reflected, so as to adapt to the characteristics of the complex structure of the device to be detected and the diverse forms of fault feature manifestations. Naturally, the fault detection result of the device to be detected can be obtained more accurately based on the first feature and the second feature.
[0075] In addition, this application only needs to focus on the feature extraction related to fault detection in the infrared image, without the need to identify the specific device type of the device to be detected, which is conducive to concentrating computing resources for feature extraction and fault detection analysis, and naturally can further improve the accuracy of device fault detection.
[0076] In this application, there can be various possibilities for the specific implementation of extracting the first feature from the infrared image.
[0077] For example, in a possible implementation manner, in order to highlight the features of the infrared image in both the channel and spatial dimensions, this application can first perform multi-scale feature extraction on the infrared image to obtain image features of multiple scales. Then, perform enhancement and fusion processing on the image features of these multiple scales in the channel dimension and the spatial dimension to obtain the first feature.
[0078] Among them, there can be various possibilities for extracting the image features of multiple scales of the infrared image. For example, multiple different-scale feature extraction modules can be directly used to perform feature extraction on the infrared image, and the features extracted by each of the multiple-scale feature extraction modules are determined as the image features of multiple scales.
[0079] In an alternative manner, in order to enable the image features at each scale to retain the features in the spatial dimension, so as to perform feature enhancement processing on the image features in both the channel dimension and the spatial dimension, in the present application, the features of the infrared image can be first extracted by using a plurality of feature extraction modules with different scales, and the initial features corresponding to the infrared image at multiple scales are obtained. On this basis, different from the conventional method of using a fully connected layer to process multi-scale features, the present application can perform normalization processing on the initial features at multiple scales in the channel dimension to obtain image features at multiple scales, so that the image features at each scale still retain the features in the spatial dimension.
[0080] For example, the features of the infrared image can be first extracted by using a plurality of different convolutional layers to obtain the initial features of the infrared image at multiple different scales. On this basis, a convolutional layer with a set receptive field size and a Rectified Linear Unit (ReLU) are successively used to perform normalization processing on the initial features at multiple scales, so as to perform normalization on the initial features at each scale in the channel dimension and obtain image features at multiple different scales.
[0081] Among them, there are also various possible implementations for performing enhancement and fusion processing on the image features at multiple scales in the channel dimension and the spatial dimension. For example, the image features at each scale are successively enhanced in the channel dimension and the spatial dimension, and the enhanced features corresponding to each scale are fused to obtain the first feature.
[0082] It can be understood that in order to further highlight the features of the image features at each scale in the channel and space, so as to better fuse the enhanced multi-scale image features, in the case where multiple scales include a first scale and a second scale, the present application can also, after performing channel enhancement processing and spatial enhancement processing on the image features of the first scale, add the obtained enhanced image features to the image features of the second scale, then perform channel enhancement processing and spatial enhancement processing on the added features, and then fuse the enhanced image features corresponding to the first scale and the second scale. Among them, the second scale is greater than the first scale.
[0083] The following is combined with Figure 2 for illustration. As Figure 2 shows a schematic flowchart of a process for multi-scale feature extraction and feature fusion processing of an infrared image in the present application, and the process may include:
[0084] S201, extracting the features of the infrared image by using a plurality of feature extraction modules with different scales to obtain the initial features corresponding to the infrared image at multiple scales.
[0085] For example, using a plurality of different convolutional layers to respectively extract the features of the infrared image to obtain the initial features at multiple scales.
[0086] In this embodiment, the multiple scales include a first scale and a second scale.
[0087] S202, perform normalization processing on the initial features of multiple scales in the channel dimension to obtain image features of multiple different scales.
[0088] For example, use a convolutional layer with a receptive field of and the ReLU activation function to perform batch normalization on the ignored features of multiple scales, obtaining image features of multiple scales with the channel dimension normalized.
[0089] It should be noted that for the convenience of understanding the implementation of extracting the first feature in this application, in Figure 2 the flowchart, an implementation method of extracting image features of multiple different scales is taken as an example for illustration, and for image features of multiple different scales extracted by other methods, it is also applicable to Figure 2 the embodiments, and will not be elaborated here specifically.
[0090] S203, perform channel enhancement processing and spatial enhancement processing on the image features corresponding to the first scale to obtain enhanced image features corresponding to the first scale.
[0091] In this application, the purpose of performing channel enhancement processing on image features is to enhance the features in the channel dimension of the image features, so as to be able to adjust the degree of attention to the features in the channel dimension during subsequent fault detection. For example, the channel attention network can be used to perform channel enhancement on the image features. For example, the Efficient Channel Attention Network (ECANet) module can be used to process the image features to perform channel enhancement on the image features.
[0092] Similarly, the purpose of performing spatial enhancement processing on image features is to highlight the features of the image features in the spatial dimension, so as to improve the degree of attention to the spatial dimension of the features during subsequent fault detection. The spatial dimension refers to the coordinates of each feature element in the height and width of the image features, that is, the position of the feature elements.
[0093] For example, the spatial attention network can be used to perform spatial enhancement on the image features that have undergone channel enhancement processing to obtain enhanced image features. For example, the Selective Kernel Networks (SKNet) attention module can be used to perform spatial enhancement on the image features that have undergone channel enhancement processing. In practical applications, the image features can be sequentially input into the cascaded channel attention network and spatial attention network to obtain the output enhanced image features.
[0094] For the sake of easy distinction, the feature obtained after the channel enhancement and spatial enhancement processing of the image feature of the first scale is called the enhanced image feature corresponding to the first scale.
[0095] S204, add the image feature of the second scale to the enhanced image feature corresponding to the first scale, and perform channel enhancement processing and spatial enhancement processing on the added feature obtained by the addition to obtain the enhanced image feature of the second scale.
[0096] Among them, the process of performing channel enhancement processing and spatial enhancement processing on the added feature after adding the enhanced image feature of the first scale and the image feature of the second scale is the same as the process of performing channel enhancement processing and spatial enhancement processing on the image feature of the first scale before, and will not be elaborated here.
[0097] In this application, multiple scales include the first scale and the second scale, and the second scale is greater than the first scale. Therefore, this application is equivalent to adding the enhanced image feature corresponding to the previous relatively smaller scale to the image feature of the next relatively larger scale, and then performing channel enhancement processing and spatial enhancement processing, so that the enhanced image feature of the later relatively larger scale can fuse the enhanced image feature of the previous relatively smaller scale.
[0098] It can be understood that if both of the two scales are included in multiple scales, then these two scales are the first scale and the second scale respectively.
[0099] If at least three scales are included in multiple scales, then in addition to including the first scale and the second scale, the at least three scales also include at least one third scale. Among them, the scale of the first scale in multiple scales is the smallest, and the scale of the second scale is the one only smaller than the first scale. On this basis, for each third scale, according to the sorting order of multiple scales from small to large, the scale that is located before the third scale and closest to the third scale is used as the reference scale corresponding to the third scale. Correspondingly, for each third scale, add the image feature of the third scale to the enhanced image feature of its corresponding reference scale, and perform channel enhancement processing and spatial enhancement processing on the added feature obtained by the addition to obtain the enhanced image feature of the third scale.
[0100] Among them, if there are multiple third scales, the image features of each third scale can be iteratively processed in the sorting order of multiple third scales from small to large until the enhanced image features corresponding to all third scales are determined.
[0101] For example, assume that multiple scales are, in ascending order, Scale 1, Scale 2, Scale 3, and Scale 4. Then Scale 1 and Scale 2 are the first scale and the second scale respectively, while Scale 3 and Scale 4 are both the third scale. Among them, the reference scale corresponding to Scale 3 is Scale 2, and the reference scale corresponding to Scale 4 is Scale 3. After determining the enhanced image features of Scale 1 and Scale 2 respectively, the image features of Scale 3 can be processed. It is necessary to add the image features of Scale 3 to the enhanced image features corresponding to Scale 2, and then perform channel enhancement processing and spatial enhancement processing on the added features obtained by the addition to obtain the enhanced image features corresponding to Scale 3. Then, add the image features of Scale 4 to the enhanced image features of Scale 3, and then perform channel enhancement processing and spatial enhancement processing on the added features obtained by the addition to obtain the enhanced image features corresponding to Scale 4.
[0102] S205, perform fusion processing on the enhanced image features corresponding to the first scale and the enhanced image features corresponding to the second scale to obtain the first feature.
[0103] Among them, if the image features of each scale are features obtained by downsampling, then the enhanced image features of the first scale and the enhanced image features of the second scale can be upsampled respectively first, and then the enhanced image features corresponding to the first scale and the enhanced image features corresponding to the second scale are subjected to fusion processing.
[0104] For ease of understanding Figure 2 The specific implementation of determining the first feature in the embodiment is described below by taking an implementation case of extracting multiple scales of an infrared image as an example. It can be understood that in the present application, an infrared image can be subjected to feature extraction through a feature extraction network. Among them, the feature extraction network includes a first feature extraction branch and a second feature extraction branch. The first feature extraction branch is used to extract a first feature from the infrared image, and the second feature extraction branch is used to extract a second feature from the infrared image. Among them, Figure 2 This is equivalent to the specific implementation of the first feature extraction branch extracting the first feature. The following takes the case where the first feature extraction branch needs to extract image features of 5 scales as an example for description.
[0105] For example Figure 3 FIG. shows an example of a schematic diagram of an implementation principle for the first feature extraction branch in the present application to extract the first feature of an infrared image.
[0106] From Figure 3It can be seen that after obtaining the infrared image, the first feature extraction branch includes 5 scale blocks for extracting features of different scales, namely the first scale block to the fifth scale block. Correspondingly, inputting the infrared image into the first scale block to the fifth scale block respectively, 5 initial features of different scales can be obtained. Among them, the sizes of the initial features extracted by the five scale blocks from the first scale block to the fifth scale block can be respectively . Among them, respectively represent a width of 256, a height of 256, and a channel number of 120, and the other initial features are similar. It can be seen that the feature extracted by the first scale block has the largest scale, while the feature extracted by the fifth scale block has the smallest scale.
[0107] Among them, each scale block can be a convolutional layer.
[0108] Then, the present application can respectively input the initial features extracted by these five scale blocks into a cascaded convolutional layer and a ReLU activation function, where the convolutional layer has a receptive field of size. Through the convolutional layer and the ReLU function, batch normalization is performed on the initial features corresponding to each scale respectively to achieve normalization of the initial features corresponding to each scale in the channel dimension, thereby obtaining image features corresponding to 5 scales. For the convenience of description, the image features corresponding to these 5 scales are respectively called the image feature 10 of the first scale block, the image feature 20 of the second scale block, the image feature 30 of the third scale block, the image feature 40 of the fourth scale block, and the image feature 50 of the fifth scale block.
[0109] In a possible implementation manner, the present application can make improvements to the VGG16 module. Specifically, the fully connected layer in the VGG16 model can be replaced with a convolutional layer with a receptive field of size ( Figure 3 the convolutional kernel in) and the ReLU activation function. Correspondingly, the improved VGG16 module can be used to extract 5 initial features of different scales corresponding to the infrared image, and batch normalization is performed using the replaced convolutional layer and ReLU activation function to obtain image features corresponding to 5 scales respectively.
[0110] It can be understood that the fully connected layer in VGG16 will reduce the features to a one-dimensional vector, resulting in the loss of the spatial dimension of the features. And the present application replaces the fully connected layer with a After the convolutional layer with large and small receptive fields and the ReLU activation function, only normalization is performed on the channel dimension, and the width and height corresponding to the image features will not be changed. Since the scale of the image features refers to the width and height, the scale sizes of the image features corresponding to each scale block will not change, so that the spatial information of each feature element in the image features can be retained, facilitating subsequent feature enhancement in the spatial dimension.
[0111] On this basis, by Figure 3 It can be seen that since the scale of the image features corresponding to the fifth scale block is the smallest, for the image features 50 corresponding to the fifth scale block, the image features 50 corresponding to the fifth scale block can be successively processed through the ECANet channel attention module and the SKNet spatial attention module to achieve channel enhancement processing and spatial enhancement processing of the image features corresponding to the fifth scale block, and obtain the enhanced image features 5 corresponding to the fifth scale block.
[0112] By Figure 3 It can be seen that after obtaining the enhanced image features 5 corresponding to the fifth scale block, the image features 40 corresponding to the fourth scale block can be added to the enhanced image features 5 corresponding to the fifth scale block, and then the added features are successively processed through the ECANet channel attention module and the SKNet spatial attention module to obtain the enhanced image features 4 corresponding to the fourth scale block.
[0113] Then, the image features 30 corresponding to the third scale block are added to the enhanced image features 4 corresponding to the fourth scale block, and the added features are successively processed through the ECANet channel attention module and the SKNet spatial attention module to obtain the enhanced image features 3 corresponding to the fourth scale block. And so on, the enhanced image features 2 corresponding to the second scale block and the enhanced image features 1 corresponding to the first scale block can be obtained respectively.
[0114] Finally, the enhanced image features 1, enhanced image features 2, enhanced image features 3, enhanced image features 4, and enhanced image features 5 corresponding to the five scale blocks are fused to obtain the first feature.
[0115] In this application, there can be multiple possible specific implementations for determining the second feature of the infrared image. Several possible situations for determining the second feature are described below.
[0116] In one possible situation, this application can perform segmentation and feature mapping processing on the infrared image to obtain multiple feature blocks corresponding to the infrared image, and then perform feature recombination on the multiple feature blocks corresponding to the infrared image to obtain the recombined features; then, perform attention conversion on the recombined features, and determine the second feature based on the obtained multi-head attention features.
[0117] Among them, segmenting the infrared image can obtain multiple image blocks. By performing feature mapping processing such as linear projection on each image block, the corresponding feature blocks of each image block (such as the embedded features of each image block) can be obtained, so as to obtain multiple feature blocks of the infrared image.
[0118] Among them, there are various specific implementations for reorganizing multiple feature blocks of the infrared image, and there is no limitation on this. For example, the shuffle window function can be used to perform a random reorganization operation on multiple feature blocks to obtain the reorganized features. In this application, there can be one or more reorganized features, and there is no specific limitation.
[0119] Among them, converting the reorganized features through attention can be to use the multi-head self-attention (MSA) model to process the reorganized features to obtain the multi-head attention features corresponding to the reorganized features. For example, based on the multi-head self-attention model and the Softmax function (normalized exponential function), the vector attention scores corresponding to the reorganized features can be calculated, and the reorganized features are weighted and summed based on the vector attention scores to obtain the multi-head attention features.
[0120] It can be understood that enhancing the diversity of features can be achieved by reorganizing the features of multiple feature blocks of the infrared image. On this basis, performing attention conversion on the reorganized features based on the multi-head self-attention mechanism can capture the long-range dependence relationships between different feature blocks in the reorganized features, so that the second feature can more comprehensively reflect the global features of the infrared image.
[0121] Among them, determining the second feature based on the multi-head attention features can be directly using the multi-head attention features as the second feature. In an alternative manner, in order to capture the high-level semantic information related to faults in the infrared image, this application can also perform semantic processing on the converted multi-head attention features to obtain the semantic features of the infrared image. Correspondingly, the converted multi-head attention features and the semantic features can be subjected to feature splicing processing to obtain the second feature.
[0122] In yet another possible case of determining the second feature, considering the characteristics of non-uniform noise and low contrast in the infrared image, in order to better adapt to the characteristics of the infrared image and effectively distinguish the noise from the real target edge in the infrared image, this application can also first decompose the infrared image into a base layer and a detail layer, and then determine the base layer as the target image and perform self-attention processing on the target image to obtain the second feature.
[0123] Among them, the detail layer mainly represents the high-frequency information in the infrared image, which includes rapidly changing information such as edges, textures, and noise in the infrared image. The base layer mainly represents the distribution of low-frequency information in the infrared image, such as the smooth area of the image and the global brightness distribution, which can represent the overall thermal radiation characteristics of the scene. Based on this, the present application performs self-attention processing on the base layer of the infrared image, which is conducive to more accurately capturing the global features of the infrared image.
[0124] Among them, the self-attention processing of the target image can be to process the target image in combination with a multi-head self-attention mechanism, and the specific processing process can be similar to the process of self-attention processing of the infrared image in the above possible situation.
[0125] It is understandable that the second feature can also be determined by combining the above two methods. Figure 4 For explanation. Figure 4 A schematic diagram of an implementation process of obtaining a second feature by performing self-attention processing based on an infrared image in the present application is shown, and the process may include:
[0126] S401, decomposing the infrared image into a base layer and a detail layer.
[0127] There is no restriction on the specific implementation of decomposing the base layer and detail layer of the infrared image.
[0128] For example, the enhanced image to be detected can be decomposed by a rolling guided filter algorithm to obtain the base layer and detail layer of the infrared image. Specifically, a window linear fitting calculation is performed on the infrared image based on a guided filter window of a preset size, and the guided filter coefficients corresponding to each guided filter window are calculated by the least square method. On this basis, the base layer of the infrared image is calculated based on the guided filter coefficients corresponding to each guided filter window and the pixel mean of the infrared image in each guided filter window, so as to decompose the base layer and detail layer of the infrared image.
[0129] S402, determining the base layer as a target image, performing segmentation and feature mapping processing on the target image, and obtaining a plurality of feature blocks corresponding to the target image.
[0130] For example, after the target image is segmented, each segmented image block is linearly projected to obtain an embedded feature block corresponding to each image block.
[0131] S403, performing feature reorganization on multiple feature blocks to obtain reorganized features.
[0132] S404, performing attention conversion on the reorganized features to obtain multi-head attention features.
[0133] The above steps S402 to S403 are similar to the process of self-attention processing based on infrared images before. For specific details, please refer to the relevant introduction above and will not be elaborated here.
[0134] S405. Perform semantic processing on the multi-head attention features to obtain the semantic features of the infrared image.
[0135] For example, the multi-head attention features can be dimensionally expanded through the linear layer and ReLU activation function in the feed-forward neural network to obtain the semantic features.
[0136] S406. Perform feature concatenation processing on the multi-head attention features and the semantic features to obtain the second feature.
[0137] In an alternative approach, the above steps S403 to S406 can be implemented through the ShuffleTransformer module. The ShuffleTransformer module is a lightweight network architecture that combines ShuffleNet and Transformer. For the specific implementation of the ShuffleTransformer module to perform feature recombination on multiple feature blocks, please refer to the relevant introduction above. The ShuffleTransformer module can include an MSA model and a Softmax function. Based on the MSA model and the Softmax function, the multi-head attention features of the recombined features can be determined. By processing the multi-head attention features through the linear layer and RelU function in the feed-forward neural network of the ShuffleTransformer module, the semantic features of the infrared image can be obtained.
[0138] It can be understood that on the premise of extracting features from the infrared image through the feature extraction network, the specific implementation of determining the second feature above is the related operation performed by the second feature extraction branch in the feature extraction network.
[0139] It can be understood that there are various specific implementations for the present application to determine the fault detection result based on the first feature and the second feature.
[0140] In a possible implementation, the present application can also first fuse the first feature and the second feature to obtain a fused feature. Then, based on the fused feature, the fault detection result is determined. The following is described in conjunction with a specific implementation.
[0141] Such as Figure 5 , which shows another schematic flowchart of the device fault detection method provided by the present application. This embodiment may include the following steps:
[0142] S501. Obtain the infrared image of the device to be detected.
[0143] S502. Perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature.
[0144] S503. Perform self-attention processing on the infrared image to obtain a second feature.
[0145] For the above steps S501 to S502, reference can be made to the relevant introductions in the previous embodiments, which will not be elaborated here.
[0146] S504. Perform fusion processing on the first feature and the second feature to obtain a fused feature.
[0147] For example, the first feature and the second feature can be fused through a feature fusion network to obtain a fused feature. Among them, there can be various possible specific implementations for the feature fusion network to fuse the first feature and the second feature.
[0148] For example, in a possible implementation manner, performing fusion processing on the first feature and the second feature based on the feature fusion network may include:
[0149] First, perform dimensional normalization on the first feature to obtain a first standard feature.
[0150] Second, perform dimensional normalization and upsampling processing on the second feature to obtain a second standard feature.
[0151] For example, both the first feature and the second feature are normalized to a preset standard scale , where the values of C, H, and W are all integers not less than 1, and the specific values can be set according to actual needs.
[0152] Performing dimensional normalization on the features can accelerate the convergence of the feature fusion network and improve the generalization ability of the model.
[0153] Finally, perform element-wise fusion processing on the first standard feature and the second standard feature to obtain a fused feature.
[0154] For example, first perform element-wise multiplication calculation on the first standard feature and the second standard feature to obtain a dual-branch fused feature, calculate the channel mean of each spatial position in the dual-branch fused feature to obtain the channel mean matrix. The dual-branch fused feature has spatial positions. Each spatial position is the position in the length and width of the dual-branch fused feature. Each spatial position includes multiple elements at the same length coordinate and width coordinate. Therefore, the channel mean of the spatial position is the average of the channel values corresponding to the feature elements at that spatial position. The global response intensity at that spatial position can be characterized by the channel mean on the spatial position.
[0155] On this basis, a cross-correlation matrix is established based on each feature element corresponding to each spatial position in the dual-branch fusion feature and the channel mean corresponding to each spatial position. The values of each element in the cross-correlation matrix are non-linearly mapped through the Sigmoid function to obtain the cross-correlation coefficient, and the fusion feature is calculated based on the cross-correlation coefficient, the first standard feature, and the second standard feature. For example, the fusion feature can be calculated as follows:
[0156] ;
[0157] where, BFM 1 represents the first standard feature, BFM 2 represents the second standard feature, represents the cross-correlation coefficient, represents element-wise multiplication calculation, represents element-wise addition calculation.
[0158] Among them, calculating the channel mean corresponding to each spatial position in the dual-branch fusion feature, establishing a cross-correlation matrix and performing non-linear mapping through the Sigmoid function to obtain the cross-correlation coefficient can adaptively adjust the fusion weight, which is beneficial to better fuse the first standard feature and the second standard feature.
[0159] S505, determine at least one region of interest to be detected in the infrared image.
[0160] Among them, the region of interest in the infrared image can be an image region that needs to be subjected to fault detection and analysis. For example, it can be an image region where the device components that need to be subjected to fault detection in the device to be detected are located in the infrared image.
[0161] There are various specific implementations for determining the region of interest from the infrared image. For example, the region of interest in the infrared image can be detected and determined by using a region of interest detection model. The region of interest detection model can be a candidate box proposal network region generation network (Region Proposal Network, RPN) model, or it can be other models, and there is no limitation on this.
[0162] S506, based on the mapping relationship between the infrared image and the fusion feature in the spatial dimension, determine at least one set of feature data corresponding to at least one region of interest in the fusion feature.
[0163] It can be understood that the fusion feature is obtained based on the infrared image. Therefore, the feature elements at each spatial position in the fusion feature have a corresponding relationship with each pixel point in the infrared image in the spatial dimension. On this basis, a set of feature data corresponding to each region of interest can be determined from the fusion feature, and each group of feature data corresponding to each region of interest can be obtained.
[0164] S507. Determine a fault detection result of at least one region of interest based on at least one set of feature data.
[0165] For example, for each region of interest, the feature data corresponding to the region of interest can be input into a fault classification model to obtain the fault detection result corresponding to the region of interest output by the fault classification model. As described above, the fault detection result at least includes the fault type corresponding to the region of interest, and may also include the probability of occurrence of this fault type, etc.
[0166] In a possible implementation, the Region of Interest Align (ROI Align) module can be used to determine the feature data corresponding to each region of interest from the fused features. On this basis, for each region of interest, feature sampling points are obtained by uniformly sampling the feature data based on the ROI Align module. Then, eigenvalue calculations are performed on the feature sampling points based on the bilinear interpolation method to obtain the sampled point feature vectors. On this basis, linear transformation and ReLU function activation processing are performed on the sampled point feature vectors through a multi-layer perceptron to obtain abstract feature vectors, and the Softmax function is used to calculate the probabilities that the region of interest belongs to each fault type. Based on the probabilities of each fault type corresponding to the region of interest, the most likely fault type of the region of interest is determined using the non-maximum suppression algorithm or other methods, and this fault type is determined as the fault detection result corresponding to the region of interest.
[0167] Among them, using ROI Align instead of the traditional region of interest pooling and calculating the eigenvalues of feature sampling points through uniform sampling and bilinear interpolation is beneficial to reducing the quantization error generated by the traditional region of interest pooling algorithm and improving the recognition accuracy.
[0168] In any of the above embodiments of the present application, considering the characteristics of uneven noise distribution, low contrast, and low resolution of the infrared image itself, in order to further improve the accuracy of fault detection of the device to be detected based on the infrared image, the infrared images processed in the above embodiments of the present application are preprocessed infrared images. For example Figure 6 , taking an implementation of obtaining an infrared image as an example, the device fault detection method of the present application will be introduced below.
[0169] For example Figure 6 , shows another flowchart of the device fault detection method provided by the present application. The method of this embodiment may include:
[0170] S601. Denoise and upsample the initial infrared image of the device to be detected to obtain a high-frequency infrared image.
[0171] For example, perform Gaussian filtering on the initial infrared image to denoise the initial infrared image. Then, perform upsampling on the initial infrared image after Gaussian filtering through bicubic interpolation to improve the resolution of the initial infrared image and obtain a high-frequency infrared image.
[0172] S602. Perform low-pass filtering on the high-frequency infrared image to obtain a low-frequency infrared image.
[0173] S603. Determine the difference image between the high-frequency infrared image and the low-frequency infrared image.
[0174] Considering the characteristic that the infrared image has a large amount of noise, the present application does not directly perform high-pass filtering on the infrared image, which can reduce the problem of noise amplification introduced by high-pass filtering. By taking the difference between the high-frequency infrared image obtained through denoising and upsampling and the low-frequency infrared image after low-pass filtering, not only can the noise be reduced, but also the edges, contours, and textures of the infrared image can be highlighted.
[0175] S604. Perform linear amplification on the pixel values of each pixel point in the difference image to obtain an amplified high-frequency image.
[0176] For example, perform linear amplification on the pixel values of each pixel point in the difference image based on a preset weight factor.
[0177] By performing linear amplification on each pixel point in the interpolated image, the detail features such as small cracks and local overheating of the device to be detected itself in the infrared image can be highlighted.
[0178] S605. Superimpose the amplified high-frequency image and the high-frequency infrared image to obtain the infrared image to be detected.
[0179] The infrared image obtained by superimposing the amplified high-frequency image and the high-frequency infrared image can have the characteristics of less noise, prominent edges, contours, and detail features, which is beneficial to more accurately detecting the device failure situation based on the infrared image in the subsequent process.
[0180] S606. Perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature.
[0181] S607. Perform self-attention processing based on the infrared image to obtain a second feature.
[0182] S608. Determine the fault detection result corresponding to the infrared image based on the first feature and the second feature.
[0183] For the above steps S606 and S608, reference can be made to the relevant introductions in the previous embodiments, and details are not described herein again.
[0184] It can be understood that, for the sake of more intuitive understanding of the specific implementation of the device fault detection method of the present application, a simple description will be given below by taking an implementation architecture of the present application as an example. As Figure 7 shows an example diagram of the implementation principle framework of the device fault detection method of the present application.
[0185] From Figure 7 it can be seen that after collecting the initial infrared image of the device to be detected by using an infrared camera or the like, a high-quality infrared image can be obtained by preprocessing the initial infrared image. Among them, the preprocessing of the initial infrared image can refer to Figure 6 the relevant introductions of steps S601 to S605 in the embodiment, which will not be elaborated here.
[0186] Then, the infrared image obtained by preprocessing is input into the feature extraction network. The feature extraction network includes a first feature extraction branch and a second feature extraction branch connected in parallel. From Figure 7 it can be seen that the first feature of the infrared image can be extracted through the first feature extraction branch, and the second feature of the infrared image can be extracted through the second feature extraction branch.
[0187] For example, the first feature extraction branch may include the aforementioned multi-scale feature extraction modules, a convolutional layer with a receptive field of size , a ReLU activation function, and an ECANet channel attention module and an SKNet spatial attention module. The specific implementation of extracting the first feature through the first feature extraction branch can refer to the previous Figure 2 and Figure 3 relevant introductions of the embodiment, which will not be elaborated here.
[0188] The second feature extraction branch may include a decomposition module for decomposing the image based on the rolling filter algorithm and a Shuffle Transformer module. The basic layer of the infrared image can be decomposed through the decomposition module, and a series of processes can be performed on the basic layer through the Shuffle Transformer module to obtain the second feature. Regarding the specific implementation of extracting the second feature by the second feature extraction branch, it can refer to the previous Figure 4 relevant introductions of the embodiment, which will not be elaborated here.
[0189] On this basis, from Figure 7 it can be seen that the first feature and the second feature are input into the feature fusion network, and the first feature and the second feature are fused through the feature fusion network to obtain the fusion feature of the infrared image. Among them, the specific implementation process of the feature fusion network for fusing the first feature and the second feature can refer to Figure 5 step S504 and other relevant introductions in the embodiment, which will not be elaborated here.
[0190] The fused features output by the feature fusion network are input into the detection and classification network. The detection and classification network first determines at least one region of interest to be detected in the infrared image, and determines at least one set of feature data corresponding to the at least one region of interest in the fused features based on the region of interest alignment. Then, based on the feature data corresponding to the region of interest, a fault classification model composed of a multi-layer perceptron and a normalization function is used to determine the fault detection result of the region of interest. For the specific implementation of the detection and classification network to determine the fault detection result, reference can be made to Figure 5 Steps S505 to S507 and other related introductions in Figure 5 , which will not be elaborated here.
[0191] It can be understood that the present application can pre-train the model parameters of the relevant models involved in the feature extraction network, feature fusion network, and detection and classification network.
[0192] For example, obtain at least one infrared image sample and the fault type label actually corresponding to each infrared image sample. Based on the fault type label corresponding to the infrared image sample, the relevant models in the above-mentioned feature extraction network, feature fusion network, and feature classification network are trained using the infrared image sample.
[0193] The training method can adopt any supervised training method, without specific limitation. For example, before training, the present application can also divide the infrared image samples into a training set, a validation set, and a test set according to a set ratio (such as 7:2:1). The error loss value between the predicted fault detection result and the fault type label of the infrared image sample is calculated through a specific loss function. Based on the Adaptive Moment Estimation (Adam) algorithm and the error loss value, the parameters of each model involved are updated by backpropagation. At the same time, the precision, recall rate, average accuracy, and mean average precision corresponding to each model are calculated through the validation set, and the model parameters are updated in combination with the calculation results. In this way, the best model parameters are obtained through continuous iteration.
[0194] Among them, the fault type label can characterize whether there is a fault in the infrared image sample and the specific fault in the case of a fault. According to the type of the device to be detected, the fault type label will also be different.
[0195] For example, taking power equipment as an example, the fault labels of infrared image samples can include no-fault label, overall overheating fault label, arc discharge fault label, insulator damage fault label, fuse melting fault label, component wear fault label, and insulating oil leakage fault label, etc.
[0196] Among them, the model parameters that need to be adjusted during the training process may include the parameters of various models mentioned above. For example, the parameters of the convolutional layer of VGG16, batch normalization parameters, the parameters of the ECANet module, the parameters of the SKNet module, the upsampling parameters in the first feature extraction branch, the parameters of the rolling guidance filtering algorithm model, the parameters of the Shuffle Transformer module, the dimension normalization parameters in the feature fusion network, the upsampling parameters of the feature fusion network, the parameters of the RPN model, the parameters of the region of interest alignment ROI Align module, and the parameters of the multi-layer perceptron.
[0197] Corresponding to a device fault detection method provided by the present application, the present application also provides a device fault detection device. As Figure 8 shown in, a schematic diagram of a composition structure of the device fault detection device provided by the present application is shown. The device in this embodiment may include:
[0198] An image acquisition unit 801, configured to acquire an infrared image of a device to be detected;
[0199] A first feature determination unit 802, configured to perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature;
[0200] A second feature determination unit 803, configured to perform self-attention processing based on the infrared image to obtain a second feature;
[0201] A fault detection unit 804, configured to determine a fault detection result corresponding to the infrared image based on the first feature and the second feature.
[0202] In a possible implementation manner, the first feature determination unit includes:
[0203] A multi-scale extraction subunit, configured to perform multi-scale feature extraction on the infrared image to obtain image features of multiple scales;
[0204] A multi-scale fusion subunit, configured to perform enhancement and fusion processing on the image features of the multiple scales in the channel dimension and the spatial dimension to obtain a first feature.
[0205] In another possible implementation manner, the multiple scales corresponding to the image features extracted by the multi-scale extraction unit include a first scale and a second scale;
[0206] The multi-scale fusion subunit includes:
[0207] A first enhancement subunit, configured to perform channel enhancement processing and spatial enhancement processing on the image features corresponding to the first scale to obtain enhanced image features corresponding to the first scale;
[0208] A second enhancer unit, configured to add the image features of the second scale to the enhanced image features corresponding to the first scale, perform channel enhancement processing and spatial enhancement processing on the added features obtained by the addition, so as to obtain the enhanced image features of the second scale, where the second scale is greater than the first scale;
[0209] A fusion processing subunit, configured to perform fusion processing on the enhanced image features corresponding to the first scale and the enhanced image features corresponding to the second scale to obtain a first feature.
[0210] In another possible implementation manner, the multi-scale extraction subunit includes:
[0211] An initial extraction subunit, configured to extract the features of the infrared image by using feature extraction modules of multiple different scales, so as to obtain initial features of multiple scales corresponding to the infrared image;
[0212] A normalization processing subunit, configured to perform normalization processing on the initial features of multiple scales in the channel dimension to obtain image features of multiple different scales.
[0213] In another possible implementation manner, the second feature determination unit includes:
[0214] An image decomposition subunit, configured to decompose the infrared image into a base layer and a detail layer;
[0215] An attention processing subunit, configured to determine the base layer as a target image, perform self-attention processing on the target image to obtain a second feature.
[0216] In another possible implementation manner, the second feature determination unit includes:
[0217] A segmentation and mapping subunit, configured to perform segmentation and feature mapping processing on the infrared image to obtain multiple feature blocks corresponding to the infrared image;
[0218] A feature recombination subunit, configured to perform feature recombination on the multiple feature blocks corresponding to the infrared image to obtain a recombined feature;
[0219] An attention conversion subunit, configured to perform attention conversion on the recombined feature, and determine a second feature based on the obtained multi-head attention feature after conversion.
[0220] In another possible implementation manner, the attention conversion subunit includes:
[0221] A semantic processing subunit, configured to perform semantic processing on the obtained multi-head attention feature after conversion to obtain the semantic feature of the infrared image;
[0222] A feature splicing subunit, configured to perform feature splicing processing on the multi-head attention feature and the semantic feature to obtain a second feature.
[0223] In another possible implementation manner, the fault detection unit includes:
[0224] A feature fusion subunit, configured to perform fusion processing on the first feature and the second feature to obtain a fusion feature;
[0225] A region determination subunit, configured to determine at least one region of interest to be detected in the infrared image;
[0226] A mapping processing subunit, configured to determine at least one set of feature data corresponding to at least one region of interest in the fusion feature based on the mapping relationship between the infrared image and the fusion feature in the spatial dimension;
[0227] A result determination subunit, configured to determine a fault detection result of the at least one region of interest based on the at least one set of feature data.
[0228] In another possible implementation manner, the image acquisition unit includes:
[0229] A denoising subunit, configured to perform denoising and upsampling processing on an initial infrared image of a device to be detected to obtain a high-frequency infrared image;
[0230] A filtering subunit, configured to perform low-pass filtering processing on the high-frequency infrared image to obtain a low-frequency infrared image;
[0231] A difference determination subunit, configured to determine a difference image between the high-frequency infrared image and the low-frequency infrared image;
[0232] A linear amplification subunit, configured to perform linear amplification processing on the pixel values of each pixel point in the difference image to obtain an amplified high-frequency image;
[0233] An overlay processing subunit, configured to perform overlay processing on the amplified high-frequency image and the high-frequency infrared image to obtain an infrared image to be detected.
[0234] An embodiment of the present application further provides an electronic device. As Figure 9 shown, it shows a schematic structural diagram of a composition of the electronic device. The electronic device includes at least a processor 901 and a memory 902;
[0235] The processor 901 is configured to execute the device fault detection method described in any one of the above embodiments;
[0236] The memory 902 is configured to store a program required for the processor to perform operations.
[0237] It can be understood that the electronic device may further include a display unit 903 and an input unit 904.
[0238] Of course, the electronic device may also have more or fewer components, and this is not limited. Figure 9 More or fewer components, and this is not restricted.
[0239] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device implements any one of the device fault detection methods provided by the embodiments of the present application.
[0240] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any one of the device fault detection methods provided by the embodiments of the present application.
[0241] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0242] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0243] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0244] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
Claims
1. A device fault detection method, comprising: Obtaining an infrared image of the device to be inspected; Performing multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature; Performing self-attention processing based on the infrared image to obtain a second feature; Based on the first feature and the second feature, a fault detection result corresponding to the infrared image is determined.
2. According to the equipment fault detection method of claim 1, performing multi-scale feature extraction and feature fusion processing on the infrared image to obtain the first feature comprises: Performing multi-scale feature extraction on the infrared image to obtain image features at multiple scales; The image features of the multiple scales are enhanced and fused in terms of channel dimension and space dimension to obtain a first feature.
3. The device fault detection method according to claim 2, wherein the plurality of scales comprises a first scale and a second scale; The step of performing channel dimension and space dimension enhancement and fusion processing on the image features of the multiple scales to obtain the first feature includes: Performing channel enhancement processing and space enhancement processing on the image features corresponding to the first scale to obtain enhanced image features corresponding to the first scale; Adding the image features of the second scale to the enhanced image features corresponding to the first scale, and performing channel enhancement processing and space enhancement processing on the added features to obtain enhanced image features of the second scale, wherein the second scale is larger than the first scale; The enhanced image features corresponding to the first scale and the enhanced image features corresponding to the second scale are fused to obtain a first feature.
4. The equipment fault detection method according to claim 2 or 3, wherein the multi-scale feature extraction is performed on the infrared image to obtain image features of multiple scales, including: Extracting features of the infrared image using feature extraction modules of multiple scales to obtain initial features of multiple scales corresponding to the infrared image; The initial features of multiple scales are normalized in the channel dimension to obtain image features of multiple different scales.
5. The device fault detection method according to claim 1, wherein the self-attention processing based on the infrared image to obtain the second feature comprises: Decomposing the infrared image into a base layer and a detail layer; The base layer is determined as a target image, and a self-attention process is performed on the target image to obtain a second feature.
6. The device fault detection method according to claim 1, wherein the self-attention processing based on the infrared image to obtain the second feature comprises: Performing segmentation and feature mapping processing on the infrared image to obtain a plurality of feature blocks corresponding to the infrared image; Performing feature reorganization on a plurality of feature blocks corresponding to the infrared image to obtain reorganized features; The recombined features are subjected to attention conversion, and a second feature is determined based on the multi-head attention features obtained by the conversion.
7. The device fault detection method according to claim 6, wherein the second feature is determined based on the multi-head attention feature obtained by conversion, comprising: Performing semantic processing on the converted multi-head attention features to obtain semantic features of the infrared image; The multi-head attention feature and the semantic feature are subjected to feature concatenation processing to obtain a second feature.
8. The device fault detection method according to claim 1, wherein determining the fault detection result corresponding to the infrared image based on the first feature and the second feature comprises: Fusing the first feature and the second feature to obtain a fused feature; Determining at least one region of interest to be detected in the infrared image; Based on the mapping relationship between the infrared image and the fused feature in the spatial dimension, determining at least one set of feature data corresponding to at least one region of interest in the fused feature; Based on the at least one set of characteristic data, a fault detection result of the at least one region of interest is determined.
9. The device fault detection method according to claim 1, wherein obtaining an infrared image of the device to be detected comprises: De-noising and up-sampling the initial infrared image of the device to be detected to obtain a high-frequency infrared image; Performing low-pass filtering on the high-frequency infrared image to obtain a low-frequency infrared image; determining a difference image between the high-frequency infrared image and the low-frequency infrared image; Performing linear amplification processing on the pixel value of each pixel point in the difference image to obtain an amplified high-frequency image; The amplified high-frequency image is superimposed on the high-frequency infrared image to obtain an infrared image to be detected.
10. An equipment fault detection device, comprising: An image acquisition unit, used for acquiring an infrared image of the device to be detected; A first feature determination unit is used to perform multi-scale feature extraction and feature fusion processing on the infrared image to obtain a first feature; A second feature determination unit, configured to perform self-attention processing based on the infrared image to obtain a second feature; The fault detection unit is used to determine a fault detection result corresponding to the infrared image based on the first feature and the second feature.
Citation Information
Cited By
Cable tunnel crack and water seepage area identification method, device, equipment, medium and product
CN120526238A