Power equipment defect detection method and system based on multi-modal image
By building an infrared image optimization module and a multimodal feature fusion device, the problem of insufficient interaction between RGB images and thermal infrared image information is solved, more accurate power equipment defect detection is achieved, false alarms and missed alarms are reduced, and the reliability and efficiency of the detection system are improved.
Patent Information
- Application Number
- CN202510171935.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing multimodal image detection methods have not fully explored the deep feature interaction between RGB images and thermal infrared images, and there is room for improvement in the accuracy of infrared images in temperature abnormality detection.
By building an infrared image optimization module, an RGB image feature extractor, a multimodal feature fusion device and a model decoder, the detection model is trained using the GPU server to optimize infrared image input and enhance information interaction.
It improves the accuracy and efficiency of power equipment defect detection, reduces false alarms and missed alarms, and improves the reliability and stability of the system.
Smart Images

Figure CN120259169A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power equipment defect detection, and in particular to a power equipment defect detection method and system based on multimodal images. Background Art
[0002] The multimodal image-based power equipment defect detection method relies on multi-source data (such as infrared images, visible light images, laser point clouds, etc.) to detect power equipment and can comprehensively analyze the equipment status. This method makes up for the limitations of single-modal information by combining the advantages of different sensors. For example, infrared images can detect temperature anomalies, visible light images can identify appearance defects, and laser point clouds help to accurately measure the equipment morphology. Based on these multimodal data, existing technologies can achieve automatic defect detection and alarm through fusion analysis using deep learning models. Compared with single-modal detection methods, multimodal methods show better performance in detection accuracy and equipment status assessment.
[0003] However, although the existing technology has achieved certain results, there are still some problems that need to be solved. First, the current multimodal methods usually simply splice or parallelize the information of RGB images (visible light images) and other modalities (such as infrared images, laser point clouds, etc.), ignoring the potential deep-level feature interactions between different modalities. This simple splicing processing method results in the failure to fully explore and utilize the synergistic relationship between the rich texture and shape information contained in the RGB image and other modal data such as infrared or laser point clouds (such as temperature and depth information), which limits the effect of multimodal data fusion.
[0004] Secondly, in terms of the technology of using RGB images to optimize infrared images, the existing power equipment defect detection methods still have room for improvement. Current methods usually directly use infrared images for abnormal temperature detection, and then combine them with RGB images for assistance. However, RGB images can provide rich information about the appearance, structure, shape and texture of the equipment. If this information can be effectively combined with infrared image data, it can further optimize the input quality of infrared images and improve the accuracy of infrared images in temperature anomaly detection. Summary of the invention
[0005] In view of the above problems existing in the prior art, the present invention is proposed.
[0006] Therefore, the technical problems to be solved by the present invention are as follows: to solve two key problems in the defect detection of power equipment. First, it focuses on how to effectively enhance the information interaction and integrated utilization between RGB images and thermal infrared images. By extracting and fusing the data of the two types of images, this patent can achieve a more accurate detection process. Second, this patent aims to make full use of the texture and edge features in RGB images to optimize the input of thermal infrared images, so that the infrared information can more accurately locate and analyze the abnormal areas of the equipment. This not only improves the accuracy and efficiency of detection, but also significantly reduces false alarms and missed detections during the detection process, and enhances the reliability and stability of the overall system.
[0007] To solve the above technical problems, the present invention provides the following technical solutions. A method for defect detection of power equipment based on multi-modal images includes: constructing an infrared image optimization module based on the spatial information of RGB images; constructing an RGB image feature extractor and an infrared image feature extractor; hierarchically constructing a multi-modal feature fusion device; constructing a model decoder based on hierarchical multi-modal features; training a detection model using a GPU server, saving model parameters and testing the detection results.
[0008] As a preferred solution of the method for defect detection of power equipment based on multi-modal images according to the present invention, wherein: the infrared image optimization module includes constructing a convolutional block connection using two consecutive convolutional layers, and optimizing the original infrared image using the convolutional block according to the multi-level features obtained from the RGB image Img R as follows: Using the convolutional block to optimize the original infrared image, expressed as:
[0009]
[0010] Img TR = CONV(Cat(F R-1 , Up2(F R-2 ), Up4(F R-3 ), Img T ))
[0011] wherein, CONV represents the convolutional block, Img R represents the RGB image, {F R-i-1 , F R-1 , F R-2 , F R-3} represents the multi-level features obtained from the RGB image, represents the average pooling layer with a kernel size of 2 i-1 ×2 i-1 , Cat represents the feature concatenation along the channel dimension, and in {Up2, Up4}, Up2 represents the bilinear interpolation upsampling with a magnification factor of 2, and Up4 represents the bilinear interpolation upsampling with a magnification factor of 4, Img TRepresents the infrared image before optimization, Img TR Represents the optimized infrared image.
[0012] As a preferred solution of the method for detecting defects in power equipment based on multimodal images described in the present invention, the RGB image feature extractor adopts ResNet50 as the feature extractor, including a 7×7 convolution kernel, an input layer f with a step size of 2 R-1 And the process is represented by the basic blocks of number 3, 4, 6, 3:
[0013]
[0014] Among them, FR i Represents the multi-level semantic information extracted from the RGB image by ResNet50, FR i-1 Represents the multi-level semantic information extracted from the RGB image by ResNet50. Represents the operations of each layer of ResNet50.
[0015] As a preferred solution of the method for detecting defects in power equipment based on multimodal images described in the present invention, the infrared image feature extractor is constructed according to the ResNet50 structure, including a 7×7 convolution kernel, an input layer f with a step size of 2 TR-1 And the number of residual convolution blocks is 3, 4, 6, 3, the process is expressed as:
[0016]
[0017] in, Represents the multi-level semantic information extracted from the infrared image, namely FRT i-1 Represents the multi-level semantic information extracted from the infrared image, Img TR represents the optimized infrared image, It represents the input layer and convolution layer group of the infrared image feature extractor, and AvgPool2 represents the average pooling layer with a kernel size of 2.
[0018] As a preferred solution of the method for defect detection of power equipment based on multimodal images described in the present invention, the multimodal feature fusion device is constructed based on the channel attention mechanism, and the global information of the input features is extracted by two operations of global average pooling AvgPool and global maximum pooling MaxPool; the global average feature and the maximum feature are sent to MLP, and the channel weight is obtained through the Sigmoid activation function.
[0019] As a preferred solution of a method for detecting defects in power equipment based on multi-modal images according to the present invention, wherein: the model decoder is constructed by arranging two residual convolutional blocks in series to construct each layer of the model decoder; the layers of the decoder are connected by bilinear interpolation upsampling operations, and the process is expressed as:
[0020]
[0021] wherein, represents each layer of the model decoder, Up2 represents bilinear interpolation upsampling with a magnification of 2, and FD i+1 represents the output features of each layer of the decoder, and Cat represents feature concatenation along the channel dimension, represents the multi-modal features after spatial weighting.
[0022] As a preferred solution of a method for detecting defects in power equipment based on multi-modal images according to the present invention, wherein: the training detection model uses random horizontal mirror flipping; wherein, the batch size in the training process is set to 8, the initial learning rate is 1×10 -3 , the number of iteration rounds is 120 rounds, and the Adam optimizer is used for parameter update; the parameters of the ResNet50 part of the encoder are initialized with the results trained on ImageNet1k; the model is trained on a GPU server based on Pytorch.
[0023] Another object of the present invention is to provide a system for detecting defects in power equipment based on multi-modal images. This system can more accurately identify and detect defects in power equipment by integrating RGB image and thermal infrared image data; through the combination of multi-modal data, it can effectively avoid the blind spots in traditional single-modal methods and improve the accuracy and efficiency of equipment detection; the system greatly improves the real-time performance, accuracy and reliability of power equipment defect detection, providing a strong guarantee for the safe operation and maintenance of power equipment.
[0024] To solve the above technical problems, the present invention provides the following technical solution: a system for detecting defects in power equipment based on multi-modal images, including: an image optimization module, a feature extraction module, a feature fusion module, a model decoding module, and a model training module;
[0025] The image optimization module constructs an infrared image optimization module based on the spatial information of the RGB image;
[0026] The feature extraction module constructs an RGB image feature extractor and an infrared image feature extractor;
[0027] The feature fusion module hierarchically constructs a multi-modal feature fusion device;
[0028] The model decoding module constructs a model decoder based on hierarchical multi-modal features;
[0029] The model training module uses a GPU server to train the detection model, saves the model parameters and tests the detection results.
[0030] A computer device includes a memory and a processor. The memory stores a computer program. The processor, when executing the computer program, implements the steps of a method for detecting defects in power equipment based on multi-modal images as described above.
[0031] A computer-readable storage medium stores a computer program. The computer program, when executed by a processor, implements the steps of a method for detecting defects in power equipment based on multi-modal images as described above.
[0032] Advantages of the present invention: First, by jointly processing RGB images and thermal infrared images, this patent solves the problem of information asymmetry in the traditional process of detecting defects in power equipment. RGB images and thermal infrared images each contain different device state information. RGB images can provide rich surface information of the device in the visible light environment, such as the device shape, color change, texture, edge features, etc., while thermal infrared images can reflect the temperature distribution of the device and capture possible abnormal hot spots inside the device. Through the image fusion algorithm of this patent, the information of these two types of images can be organically combined, not only enabling the detection system to monitor the device from multiple angles, but also avoiding the blind spots and deficiencies brought by a single image mode through complementary information fusion, thus improving the comprehensiveness and accuracy of detection.
[0033] Second, this patent particularly focuses on deeply mining and utilizing the texture and edge features of RGB images to optimize the input of thermal infrared images. In the traditional defect detection process, due to factors such as environmental interference and complex device shapes, it may be difficult to accurately locate hot spots in thermal infrared images. The texture and edge features of RGB images can provide an accurate reference for the device morphology. By combining these features with the data of thermal infrared images through the algorithm in this patent, the recognition accuracy of hot spots in infrared images can be significantly enhanced, and the area of device anomalies can be accurately located. This not only helps to reduce the detection of false hot spots (false alarms), but also effectively reduces the situation where the device has anomalies but is not recognized (missed alarms), greatly improving the robustness of detection.
[0034] In addition, the technical solution of this patent can significantly improve the efficiency of the power equipment defect detection system. Through the joint analysis of RGB and thermal infrared images, the dependence on a single image mode in traditional detection methods is reduced, the processes of repeated detection and multiple verifications are avoided, and the detection time is shortened. At the same time, the image processing algorithm of this patent has high computational efficiency and can process a large amount of image data in real time to ensure the efficient operation of the equipment detection system. This is of great significance for the online monitoring and real-time fault warning of power equipment, can detect and handle problems in the early stage of the equipment, avoid the further deterioration of equipment failures, thereby reducing maintenance costs and improving the reliability of equipment operation.
[0035] In summary, through the innovative information fusion method, this patent not only improves the accuracy and efficiency of power equipment defect detection, but also significantly reduces the incidence of false alarms and missed detections, making the detection system more reliable and stable, and having broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 It is a schematic flowchart of a power equipment defect detection method based on multi-modal images provided by an embodiment of the present invention.
[0038] Figure 2 It is a schematic diagram of an infrared image optimization module of a power equipment defect detection method based on multi-modal images provided by an embodiment of the present invention.
[0039] Figure 3 It is a schematic diagram of feature extraction of RGB images and infrared images of a power equipment defect detection method based on multi-modal images provided by an embodiment of the present invention.
[0040] Figure 4 It is a schematic diagram of multi-modal feature fusion and decoding of a power equipment defect detection method based on multi-modal images provided by an embodiment of the present invention.
[0041] Figure 5 It is a schematic diagram of comparison of detection results of a power equipment defect detection method based on multi-modal images provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0043] Example 1, referring to Figures 1 - 5 , which is an embodiment of the present invention. This embodiment provides a method for detecting power equipment defects based on multi-modal images, including:
[0044] S1: Based on the spatial information of the RGB image, construct an infrared image optimization module.
[0045] It should be noted that, as shown in S1 in Figure 1 , based on the spatial information of the RGB image, construct an infrared image optimization module for improving the imaging quality of the thermal infrared image.
[0046] Furthermore, use two consecutive convolutional layers to construct a convolutional block, with a BN layer and a ReLU activation function following each convolutional layer. The process is expressed as:
[0047] Output = Conv2(Conv1(Input))
[0048] where Input and Output respectively represent the input features and output features of the convolutional block, and Conv1 and Conv2 represent convolutional layers accompanied by a BN layer and a ReLU activation function;
[0049] As shown in Figure 2 , the convolutional blocks are connected by an average pooling layer to obtain shallow features under different receptive fields; according to the multi-level features obtained from the RGB image Img R in use the convolutional block to optimize the original infrared image to obtain the infrared image Img TR for subsequent feature extraction and interaction. The process is expressed as:
[0050]
[0051] Img TR = CONV(Cat(F R-1 , Up2(F R-2 ), Up4(F R-3 ), Img T ))
[0052] where CONV represents the convolutional block, Img RDenote the RGB image, {F R-i-1 , F R-1 , F R-2 , F R-3} denote the multi-level features obtained from the RGB image. Denote the average pooling layer with a kernel size of 2 i-1 ×2 i-1 , Cat denotes the feature concatenation along the channel dimension, {Up2, Up4} respectively denote the bilinear interpolation upsampling with magnification factors of 2 and 4, Img T , Img TR respectively denote the infrared images before and after optimization.
[0053] S2: Construct the RGB image feature extractor and the infrared image feature extractor.
[0054] It should be noted that, as shown in S2 Figure 1 , construct the RGB image feature extractor and the infrared image feature extractor respectively, which are used to extract different levels of semantic information from the RGB image and the infrared image.
[0055] Furthermore, as shown in Figure 3 , for the RGB image, ResNet50 is adopted as the feature extractor. Utilizing the residual structure, ResNet50 can converge faster during the training process and is not prone to performance degradation when deepening the network layers. Specifically, ResNet50 consists of an input layer f with a 7×7 convolutional kernel and a stride of 2 R-1 , and a convolutional layer group composed of basic blocks with quantities of 3, 4, 6, and 3 The process is expressed as:
[0056]
[0057] Among them, FR i-1 denotes the multi-level semantic information extracted by ResNet50 from the RGB image. denotes the operations of each layer of ResNet50; Given that the semantic information contained in the thermal infrared image is relatively limited, and at the same time to reduce the computational complexity of the model, therefore, based on the convolutional block in step S1, design a residual convolutional block to construct a feature extractor suitable for the thermal infrared image; Specifically, the residual convolutional block consists of two convolutional blocks connected by a residual structure, and the process is expressed as:
[0058] Output = CONV(CONV(Input) + Input)
[0059] Among them, + denotes the feature addition operation.
[0060] Furthermore, the RGB image feature extractor uses ResNet50 as the feature extractor, including a 7×7 convolution kernel and an input layer f with a stride of 2. R-1 And the number of basic blocks is 3, 4, 6, 3; the infrared image feature extractor is built according to the ResNet50 structure, including a 7×7 convolution kernel, an input layer f with a stride of 2 TR-1 And residual convolution blocks of 3, 4, 6, and 3.
[0061] According to the structure of ResNet50, a feature extractor for infrared images is constructed. The input layer contains a 7×7 convolution kernel and an input layer f with a step size of 2. TR-1 , and a convolutional layer group consisting of 3, 4, 6, and 3 residual convolutional blocks The convolutional layer group is preceded by an average pooling layer with a kernel size of 2×2 for downsampling operations. The process is expressed as:
[0062]
[0063] in, Represents the multi-level semantic information extracted from the infrared image, namely FRT i-1 represents the multi-level semantic information extracted from the infrared image, It represents the input layer and convolution layer group of the infrared image feature extractor, and AvgPool2 represents the average pooling layer with a kernel size of 2.
[0064] S3: Hierarchical construction of multimodal feature fuser.
[0065] It should be noted that if Figure 1 As shown in S3, a multimodal feature fuser is constructed hierarchically to fuse multimodal features at all levels and feed them into the model decoder.
[0066] Further, such as Figure 4 As shown in the figure, based on channel attention and spatial attention, a multimodal feature fuser is constructed. Specifically, the channel attention mechanism consists of two operations, global average pooling AvgPool and global maximum pooling MaxPool, to extract the global information of the input features; the global average feature and the maximum feature are sent to the MLP (multi-layer perceptron), and the channel weight is obtained through the Sigmoid activation function; the multi-layer perceptron consists of two 1×1 convolution kernel ReLU activation functions, and the process is expressed as:
[0067]
[0068] WC fuse-i =Sigmoid(Conv 1×1 (RELU(Conv 1×1 (Cat(F Avg-i,F Max-i )))), i = 1, 2, 3, 4, 5
[0069] Among them, AvgPool and MaxPool respectively represent average global pooling and max global pooling, respectively representing two types of global features of the fused information, Conv 1×1 represents a convolution with a convolution kernel size of 1×1, RELU represents the ReLU activation function, and Sigmoid represents the Sigmoid activation function, represents the channel weights of the fused features at each level;
[0070] Multiplying the channel weights with the concatenated multi-modal feature points results in the multi-modal feature after channel weighting The process is expressed as:
[0071]
[0072] Among them, represents the multi-modal feature after channel weighting, represents the feature multiplication operation.
[0073] Furthermore, the multi-modal feature after channel weighting is processed by the spatial attention mechanism constructed by the three convolutional blocks in step S1; by extracting the significant information at different spatial positions, the attention ability of the fused feature to the local important regions is further enhanced, and the detection accuracy of the key regions is improved. The process is expressed as:
[0074] WS fuse-i = Sigmoid(CONV(CONV(CONV(FC use-i )))), i = 1, 2, 3, 4, 5
[0075]
[0076] Among them, represents the spatial weights of the fused features at each level, represents the multi-modal feature after spatial weighting.
[0077] S4: Construct a model decoder based on the hierarchical multi-modal features.
[0078] It should be noted that, as shown in S4, two residual convolutional blocks are arranged in series to construct the model decoder at each level based on the multi-modal features at each level. Figure 1
[0079] Furthermore, based on the residual convolution block designed in step S2, a model decoder for multi-modal features at all levels is constructed. Specifically, two residual convolution blocks are arranged in series to construct the model decoder for each layer, which is used to fuse the features from the previous layer of the decoder and the fused multi-modal features of the current layer. The decoder layers are connected by bilinear interpolation upsampling operations, and the process is expressed as:
[0080]
[0081] Among them, represents the model decoder for each layer, UP2 represents bilinear interpolation upsampling with a magnification factor of 2, and FD i+1 represents the output features of each layer of the decoder.
[0082] Furthermore, based on the output features of each layer of the decoder, through convolution with a convolution kernel size of 1×1, the prediction results for each stage of the decoder are obtained The process is expressed as:
[0083] S i = Conv 1×1 (FD i )
[0084] Among them, represents the prediction results for each stage of the decoder.
[0085] S5: Use the GPU server to train the detection model, save the model parameters and test the detection results.
[0086] It should be noted that, as shown in S5 in Figure 1 , the images are uniformly scaled to 256×256, and random horizontal mirror flipping is adopted during the training process. Secondly, the batch size in the training process is set to 8, the initial learning rate is 1×10 -3 , the number of iteration rounds is 120 rounds, and the Adam optimizer is used for parameter update; the parameters of the ResNet50 part of the encoder are initialized with the results trained on ImageNet1k, and the remaining parameters of the model are randomly initialized; finally, the model is trained on the GPU server based on Pytorch, and the cross-entropy loss function is used to calculate the loss during the process.
[0087] Furthermore, as shown in Figure 5 , the first column represents the original RGB image of the transmission line, the second column represents the corresponding thermal infrared image, the third column represents the true annotation map for defect detection, and the fourth column represents the prediction result; it can be seen from Figure 5 that the prediction result can effectively fuse the information of the RGB image and the thermal infrared image, and accurately identify the defect area; compared with the ground truth map, the prediction result is highly consistent with the true annotation in the defect area.
[0088] The above is a schematic solution of a power equipment defect detection method based on multimodal images in this embodiment. It should be noted that the technical solution of the system of the power equipment defect detection method based on multimodal images belongs to the same concept as the above-mentioned technical solution of the power equipment defect detection method based on multimodal images. For the details not described in detail in the technical solution of the power equipment defect detection system based on multimodal images in this embodiment, reference can be made to the description of the technical solution of the above-mentioned power equipment defect detection method based on multimodal images.
[0089] Embodiment 2 is an embodiment of the present invention, which provides a power equipment defect detection system based on multimodal images, including: an image optimization module, a feature extraction module, a feature fusion module, a model decoding module, and a model training module;
[0090] The image optimization module constructs an infrared image optimization module based on the spatial information of the RGB image;
[0091] The feature extraction module constructs an RGB image feature extractor and an infrared image feature extractor;
[0092] The feature fusion module hierarchically constructs a multimodal feature fusion device;
[0093] The model decoding module constructs a model decoder based on the hierarchical multimodal features;
[0094] The model training module uses a GPU server to train the detection model, saves the model parameters, and tests the detection results.
[0095] This embodiment also provides a computing device applicable to the situation of a power equipment defect detection method based on multimodal images, including:
[0096] A memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement a power equipment defect detection method based on multimodal images as proposed in the above embodiment.
[0097] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by the processor, it implements a power equipment defect detection method based on multimodal images as proposed in the above embodiment.
[0098] The storage medium proposed in this embodiment and the power equipment defect detection method based on multimodal images proposed in the above embodiment belong to the same inventive concept. For the technical details not described in detail in this embodiment, reference can be made to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0099] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0100] The logics and / or steps described in other ways herein, for example, can be regarded as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0101] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field-programmable gate arrays (FPGA), etc.
[0102] It should be noted that the above embodiments are only used to illustrate the technical solution of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solution of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solution of the present invention, and all of them should be covered by the scope of the claims of the present invention.
Claims
1. A method for detecting defects of power equipment based on multimodal images, characterized in that: including: Construct an infrared image optimization module based on the spatial information of the RGB image; Construct an RGB image feature extractor and an infrared image feature extractor; Hierarchically construct a multimodal feature fusion device; Construct a model decoder based on the hierarchical multimodal features; Use a GPU server to train the detection model, save the model parameters and test the detection results.
2. The method for detecting defects of power equipment based on multimodal images according to claim 1, wherein: The infrared image optimization module includes constructing a convolutional block connection using two consecutive convolutional layers, and according to the multi-level features obtained from the RGB image Img R optimize the original infrared image using the convolutional block, which is expressed as: from the multi-level features obtained Img TR = CONV(Cat(F R-1 ,Up2(F R-2 ),Up4(F R-3 ),Img T )) Among them, CONV represents the convolutional block, Img R represents the RGB image, {F R-i-1 , F R-1 , F R-2 , F R-3} represents the multi-level features obtained from the RGB image, represents the average pooling layer with a kernel size of 2 i-1 ×2 i-1 , Cat represents the feature concatenation along the channel dimension, in {Up2, Up4}, Up2 represents the bilinear interpolation upsampling with a magnification factor of 2, Up4 represents the bilinear interpolation upsampling with a magnification factor of 4, Img T represents the infrared image before optimization, Img TR represents the infrared image after optimization.
3. The method for detecting power equipment defects based on multi-modal images according to claim 2, wherein: The RGB image feature extractor uses ResNet50 as the feature extractor, including a 7×7 convolutional kernel and an input layer f with a stride of 2 R-1 and basic blocks with quantities of 3, 4, 6, and 3. The process is expressed as: Among them, FR i represents the multi-level semantic information extracted from the RGB image by ResNet50, and FR i-1 represents the multi-level semantic information extracted from the RGB image by ResNet50, represents the operations of each layer of ResNet50.
4. The method for detecting defects of power equipment based on multimodal images according to claim 3, characterized in that: The infrared image feature extractor is constructed according to the ResNet50 structure, including a 7×7 convolutional kernel and an input layer f with a stride of 2 TR-1 and residual convolutional blocks with quantities of 3, 4, 6, and 3. The process is expressed as: Among them, represents the multi-level semantic information extracted from the infrared image, i.e., FRT i-1 represents the multi-level semantic information extracted from the infrared image, Img TR represents the optimized infrared image, represents the input layer and convolutional layer group of the infrared image feature extractor. AvgPool2 represents the average pooling layer with a kernel size of 2.
5. The method for detecting power equipment defects based on multi-modal images according to claim 4, wherein: The construction of the multimodal feature fusion device is based on the channel attention mechanism, and the global average pooling AvgPool and the global maximum pooling MaxPool are used to extract the global information of the input features; Send the global average feature and the maximum feature into the MLP, and obtain the channel weights through the Sigmoid activation function.
6. The method for detecting power equipment defects based on multimodal images according to claim 5, characterized in that: The construction of the model decoder uses two residual convolutional blocks arranged in series to construct each layer of the model decoder; the layers of the decoder are connected by the upsampling operation of bilinear interpolation, and the process is expressed as: Among them, represents the model decoders of each layer, Up2 represents bilinear interpolation upsampling with a magnification of 2, and FD i+1 represents the output features of each layer of the decoder, and Cat represents feature concatenation along the channel dimension. represents the multi-modal features after spatial weighting.
7. The method for defect detection of power equipment based on multimodal images according to claim 6, characterized in that: The training detection model uses random horizontal mirror flipping; among them, the batch size in the training process is set to 8, and the initial learning rate is 1×10 -3 , the number of iteration rounds is 120 rounds, and the Adam optimizer is used for parameter update; the parameters of the ResNet50 part of the encoder are initialized with the results trained on ImageNet1k; the model is trained on a GPU server based on Pytorch.
8. A system for detecting defects of power equipment based on multi-modal images according to any one of claims 1-7, characterized in that: including: An image optimization module, a feature extraction module, a feature fusion module, a model decoding module and a model training module; The image optimization module constructs an infrared image optimization module based on the spatial information of the RGB image; The feature extraction module constructs an RGB image feature extractor and an infrared image feature extractor; The feature fusion module hierarchically constructs a multimodal feature fusion device; The model decoding module constructs a model decoder based on the hierarchical multimodal features; The model training module uses a GPU server to train the detection model, save the model parameters and test the detection results.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of a method for detecting power equipment defects based on multimodal images according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of a method for detecting power equipment defects based on multimodal images according to any one of claims 1 to 7.