Method and system for fusing inspection infrared images and visible light images of power transmission and transformation equipment

Through the multi-level feature extraction and fusion method of the IVATFusion network, the problem that the infrared and visible light image fusion method in the prior art is difficult to dig deep information, and high-quality fusion images are generated, which improves the accuracy and detailed performance of electrical equipment failure detection.

CN120259101BActive Publication Date: 2025-08-12NANCHANG INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510740867.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-08-12
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing infrared and visible light image fusion methods are difficult to effectively dig deep information in the image, resulting in insufficient details of the fused image and unable to fully demonstrate the advantages of infrared and visible light images.

Method used

The IVATFusion network is adopted, which includes basic feature extraction module, multi-scale feature extraction module, attention-weighted fusion module, RGB information extraction module, deformable convolutional alignment module, deep feature fusion module and fine feature fusion module. The image fusion model is obtained through iterative training, global and local features are extracted, and attention-weighted and feature fusion is performed to generate high-quality fusion images.

Benefits of technology

It improves the adaptability and flexibility of image fusion, can adapt to different lighting conditions and background environments, generate high-quality fusion images, enhances the contrast and clarity of the image, and improves the accuracy and comprehensiveness of fault judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259101B_ABST
    Figure CN120259101B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for fusing inspection infrared images and visible light images of power transmission and transformation equipment. The method comprises: obtaining at least one historical target image containing the power transmission and transformation equipment; inputting the at least one historical target image into a preset IVATFusion network, iteratively training the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module; obtaining a real-time target image containing a certain power transmission and transformation equipment, inputting the real-time target image into the image fusion model, and outputting a fused image containing the certain power transmission and transformation equipment. The method can adapt to different lighting conditions and background environments, providing high-quality fused images for various electrical equipment application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to a method and system for fusing inspection infrared images and visible light images of power transmission and transformation equipment. Background Art

[0002] Power systems require equipment to be stable, reliable, and easy to operate over the long term, especially when it comes to ensuring power supply security. However, electrical equipment failures are frequent, with abnormal temperatures being a precursor to many failures. Failure to promptly identify the fault type and implement corrective measures can result in significant economic losses. Therefore, infrared and visible light image fusion technology plays a vital role in power system monitoring.

[0003] Due to the unique geographical location of high-voltage electrical equipment, power grid inspection currently relies primarily on intelligent inspection robots, drones, and manual inspections. Inspection robots and drones, incorporating multimodal imaging sensors and image processing technology, can capture infrared and visible light images generated by equipment during operation, enabling real-time monitoring and fault diagnosis.

[0004] Infrared images effectively reflect the temperature of electrical equipment, while visible light images clearly reveal the device's appearance details. However, differences in imaging mechanism, resolution, and field of view limit the effectiveness of each when used alone. Therefore, infrared and visible light image fusion technology is used to meet practical application requirements in power equipment fault detection. By fusing the features of the two images, hot spots in the equipment can be accurately identified and located, thereby improving the accuracy and comprehensiveness of fault diagnosis.

[0005] Existing image fusion methods (such as DenseFuse and FusionGAN) primarily focus on image fusion, typically requiring the images to be fused to be registered before the fusion algorithm is applied. While these methods have achieved some success in image fusion, in practice, excessively noisy input images can significantly impact the fusion effect. Furthermore, existing methods often only extract superficial image features and struggle to effectively mine deeper information. This results in insufficient detail in the fused image, failing to fully demonstrate the strengths of both infrared and visible light images. Summary of the Invention

[0006] The present invention provides a method and system for fusing inspection infrared images and visible light images of power transmission and transformation equipment, which are used to solve the technical problem that it is difficult to effectively mine deep information in images, resulting in insufficient detail expression in the fused images.

[0007] In a first aspect, the present invention provides a method for fusing inspection infrared images and visible light images of power transmission and transformation equipment, comprising:

[0008] Acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image contains a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image;

[0009] Inputting the at least one historical target image into a preset IVATFusion network, iteratively training the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module;

[0010] Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module;

[0011] The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features;

[0012] The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area;

[0013] The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature;

[0014] The feature conversion module converts the target fusion feature into a fusion image;

[0015] A real-time target image including a certain power transmission and transformation equipment is obtained, and the real-time target image is input into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation equipment.

[0016] In a second aspect, the present invention provides a system for fusing inspection infrared images and visible light images of power transmission and transformation equipment, comprising:

[0017] an acquisition module configured to acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image includes a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image;

[0018] A training module is configured to input the at least one historical target image into a preset IVATFusion network, and iteratively train the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module;

[0019] Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module;

[0020] The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features;

[0021] The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area;

[0022] The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature;

[0023] The feature conversion module converts the target fusion feature into a fusion image;

[0024] The output module is configured to obtain a real-time target image containing a certain power transmission and transformation equipment, input the real-time target image into the image fusion model, and the image fusion model outputs a fused image containing the certain power transmission and transformation equipment.

[0025] According to a third aspect, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to any embodiment of the present invention.

[0026] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the program instructions are executed by a processor, the processor executes the steps of the method for fusing inspection infrared images and visible light images of power transmission and transformation equipment of any embodiment of the present invention.

[0027] The present invention discloses a method and system for fusing inspection infrared images and visible light images of power transmission and transformation equipment. By introducing a basic feature extraction module and a multi-scale feature extraction module, the method can extract global and local features to better reflect modality-specific and modality-shared features. In addition, the use of an RGB information extraction module, an attention weighted fusion module, a deep feature fusion module, and a fine feature fusion module can effectively fuse image information of different modalities, retain more details of the source image, and enhance the contrast and clarity of the image. The method is highly adaptable and flexible, can adapt to different lighting conditions and background environments, and provide high-quality fused images for various electrical equipment application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0029] Figure 1 A flowchart of a method for fusing inspection infrared images and visible light images of power transmission and transformation equipment provided by one embodiment of the present invention;

[0030] Figure 2 A partial flow chart of a method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to a specific embodiment of the present invention is provided;

[0031] Figure 3 A structural block diagram of a basic feature extraction module of a specific embodiment is provided for one embodiment of the present invention;

[0032] Figure 4 A structural block diagram of a dynamic wavelet packet pyramid convolution block according to a specific embodiment of the present invention is provided;

[0033] Figure 5 A structural block diagram of a multi-scale feature extraction module according to a specific embodiment of the present invention is provided;

[0034] Figure 6 A structural block diagram of an attention weighted fusion module is provided for one embodiment of the present invention;

[0035] Figure 7 A structural block diagram of an RGB information extraction module according to a specific embodiment of the present invention is provided;

[0036] Figure 8A structural block diagram of a deformable convolution alignment module, a deep feature fusion module, and a fine feature fusion module according to a specific embodiment of the present invention is provided;

[0037] Figure 9 This is a structural block diagram of a system for fusing infrared images and visible light images for inspection of power transmission and transformation equipment provided by one embodiment of the present invention;

[0038] Figure 10 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0040] See also Figure 1 , which shows a flow chart of a method for fusing inspection infrared images and visible light images of power transmission and transformation equipment of the present application.

[0041] like Figure 1 As shown in FIG, the method for fusing inspection infrared images and visible light images of power transmission and transformation equipment specifically includes the following steps:

[0042] Step S101 : acquiring at least one historical target image containing power transmission and transformation equipment, wherein one historical target image includes a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image.

[0043] In this step, a historical target image includes historical inspection infrared images and historical visible light images of the same power transmission and transformation equipment.

[0044] Step S102: input the at least one historical target image into a preset IVATFusion network, iteratively train the IVATFusion network, and obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module.

[0045] In this step, see Figure 2, according to the basic feature extraction module and the multi-scale feature extraction module, the visible light image features in the visible light image and the infrared image features in the infrared image are obtained.

[0046] Specifically: multi-level features of visible light images or infrared images are extracted according to a dynamic wavelet packet transform module and a depthwise separable deformable convolution; the multi-level features are fused according to a dynamic pyramid feature fusion module to obtain a first result; the spatial size of the first result is compressed to a fixed 1x1 using a maximum pooling layer and an adaptive average pooling layer, and a feature vector containing 256 channels is output; the feature vectors are input into the first channel and the second channel respectively, the first channel outputs the second result, and the second channel outputs the third result; the second result and the third result are input into a feature fusion module, and the feature fusion module outputs visible light image features or infrared image features.

[0047] See also Figure 3 The basic feature extraction module includes: a first dynamic wavelet packet pyramid convolution block, a first 2x2 maximum pooling layer, a second dynamic wavelet packet pyramid convolution block, a second 2x2 maximum pooling layer, a third dynamic wavelet packet pyramid convolution block, and a 1x1 adaptive average pooling layer, which are connected in sequence.

[0048] See also Figure 4 The first dynamic wavelet packet pyramid convolution block, the second dynamic wavelet packet pyramid convolution block and the third dynamic wavelet packet pyramid convolution block all include: a dynamic wavelet packet transform module, a depth-separable deformable convolution and a dynamic pyramid feature fusion module connected in sequence.

[0049] Basic Feature Extraction Module: This module extracts multi-scale, high-dimensional features from the input image through a combination of depthwise separable and deformable convolution, a dynamic wavelet packet transform (DWT) module, and a dynamic pyramid feature fusion module. First, the dynamic wavelet packet transform (DWT) module and depthwise separable and deformable convolution are used to extract multi-level features from the image. The dynamic pyramid feature fusion module then integrates information at different scales. Finally, a max pooling layer and an adaptive average pooling layer are used to compress the spatial size of the feature map to a fixed 1x1 size. The module then outputs a feature vector containing 256 channels, providing the foundation for feature extraction in the subsequent network.

[0050] See also Figure 5 The multi-scale feature extraction module includes: a first channel, a second channel, and a feature fusion module connected to the first channel and the second channel respectively, wherein the first channel includes a non-local attention module and a 3x3 convolution kernel connected to the non-local attention module; the second channel includes a channel attention module and a 3x3 convolution kernel connected to the channel attention module.

[0051] It should be noted that the non-local attention module is used to capture long-range dependencies, and the CBAM channel attention is used to enhance key information. Then, ordinary 3×3 convolution and 3×3 dilated convolution are used to extract fine texture and large structure features. The two outputs are directly spliced in the channel dimension to retain scale independence. Finally, 1×1 convolution is used for adaptive compression weighting to achieve efficient and complete multi-scale feature fusion.

[0052] The non-local attention module is: Non local Block.

[0053] Multi-scale Feature Extraction Module: This module combines the non-local attention module and the channel attention module to further extract multi-scale features. By using different types of convolution kernels (standard convolution and dilated convolution), the feature map is processed to capture multi-scale information in the image. Simultaneously, the non-local and channel attention mechanisms help enhance the global and local information in the feature map, thereby improving the richness and accuracy of the feature representation.

[0054] Furthermore, the attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features.

[0055] See also Figure 6 The weighted attention fusion module includes: a first attention channel, a second attention channel, a weighted fusion module connected to the first and second attention channels, a 7x7 convolution kernel connected to the weighted fusion module, and a Sigmoid activation function connected to the 7x7 convolution kernel. Specifically, the weighted fusion module is a Convolutional Block Attention Module (CBAM).

[0056] Attention Weighted Fusion Module: This module performs a weighted fusion of features from visible light and infrared images, using an attention mechanism to enhance important features and suppress irrelevant or unimportant information. By combining two different attention mechanisms, this module selectively emphasizes the most critical parts of the image during the fusion process, improving the quality of the fused feature representation.

[0057] Furthermore, the RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high-temperature area.

[0058] RGB information extraction module: This module needs to set the normal temperature threshold while processing the original infrared image, and then perform histogram equalization on the infrared image to increase the contrast of the image and make the details clearer. It is then converted into a numpy array. Then, by retaining the areas with grayscale values greater than the normal temperature and setting the other areas to black, it is converted into an RGB image, and the areas above the normal temperature are extracted. These areas are mapped to the color of the original infrared image area, and finally a map is generated in which the other areas are set to black and the RGB information of the areas above the normal temperature is retained. Then, two specific sizes of convolution kernels, 3×3 and 5×5, are used in parallel to perform unconventional continuous stacking convolutions to achieve fine-grained multi-scale feature separation and extraction, and channel dimension splicing is used instead of element-by-element addition to retain the independence of the original scale features. Finally, the final feature fusion is completed through 1×1 convolution, which improves the representation capability while controlling the growth of computational complexity, and finally generates the RGB features of the high-temperature area.

[0059] See also Figure 7 , according to the preset temperature threshold, the historical inspection infrared image is subjected to histogram equalization processing to obtain a contrast-enhanced infrared grayscale image, and the infrared grayscale image is subjected to feature extraction by the feature extraction module, and the extracted features are respectively input into the threshold prediction branch and the region segmentation branch. The threshold prediction branch outputs a dynamic threshold in the range of 0-1 by performing pooling and full connection operations on the extracted features, and multiplies it by 255 to map it to the image grayscale threshold. The region segmentation branch outputs a segmentation probability map of the same size as the historical inspection infrared image by performing convolution and upsampling operations on the extracted features. Each pixel value in the segmentation probability map represents the probability of belonging to the high temperature area;

[0060] The threshold prediction branch includes an adaptive average pooling layer, a flattening module connected in sequence to flatten the input multi-dimensional tensor into a single dimension, a fully connected layer, a Relu activation function, a fully connected layer, and a Sigmoid activation function.

[0061] The region segmentation branch includes a 3x3 convolution kernel, a Relu activation function, an upsampling layer, a 3x3 convolution kernel, a Relu activation function, an upsampling layer, a 1x1 convolution kernel, and a Sigmoid activation function connected in sequence;

[0062] Compare each pixel value in the segmentation probability map with the image grayscale threshold;

[0063] The pixel values greater than the image grayscale threshold are set to 1, and the pixel values not greater than the image grayscale threshold are set to 0, and a binary mask matrix containing 0 and 1 is generated according to the comparison result, and the binary mask matrix is multiplied pixel by pixel by pixel by each pixel value in the historical inspection infrared image to obtain the target historical inspection infrared image, wherein the target historical inspection infrared image only contains high-temperature areas greater than the temperature threshold;

[0064] Convolving the target historical inspection infrared image according to a 3×3 convolution kernel and a 5×5 convolution kernel, respectively, to obtain a first feature map corresponding to the 3×3 convolution kernel and a second feature map corresponding to the 5×5 convolution kernel, wherein the first feature map and the second feature map have the same spatial size;

[0065] The first feature map and the second feature map are directly stacked in the channel dimension, and the stacked feature maps are convolved according to a 1×1 convolution kernel to obtain the RGB features of the high-temperature area.

[0066] Furthermore, the deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature.

[0067] Deformable convolution alignment module: Aligns the fusion features of visible light images and infrared images with the features of RGB information. By calculating the offset and using deformable convolution to adjust the spatial structure of the RGB features in the high-temperature area, it accurately aligns them with the features of the fused image, thereby achieving feature fusion of cross-modal images.

[0068] Deep Feature Fusion Module: This module performs deep processing on the fused features, further enhancing feature expression through residual blocks. It combines the RGB features of the high-temperature region with the fused features and extracts richer features through multi-layer convolution operations, thereby enhancing the detail and expressiveness of the fused image and providing more accurate feature representation for subsequent tasks.

[0069] Fine Feature Fusion Module: This module refines multi-scale features, extracting and fusing them to enhance feature expression and improve image processing or image generation tasks. By using operations such as convolution, pooling, and upsampling, it enables the model to effectively capture features at different scales, thereby improving overall image quality and detail.

[0070] See also Figure 8 ,The deformable convolution alignment module includes a 3x3 convolution kernel, a Relu activation function, and a 3x3 convolution kernel connected in sequence.

[0071] The deep feature fusion module includes a 3x3 convolution kernel, a first residual block, a second residual block and a third residual block connected in sequence, wherein the first residual block, the second residual block and the third residual block each include a 3x3 convolution kernel, a ReLU activation function and a 3x3 convolution kernel connected in sequence.

[0072] The fine feature fusion module includes a 3x3 convolution kernel, a ReLU activation function, a 3x3 convolution kernel, a 3x3 convolution kernel, an average pooling layer, and a 1x1 convolution kernel connected in sequence.

[0073] Furthermore, the feature conversion module converts the target fusion feature into a fusion image.

[0074] The feature conversion module converts the deeply processed and fused features into the final output image. Through a series of convolutional layers and upsampling operations, it decodes high-dimensional features into meaningful image outputs and ensures that the output image size and format meet the expected requirements.

[0075] It should be noted that the loss between the infrared image features in the historical inspection infrared image and the visible light image features in the historical visible light image is calculated according to the contrast loss function of the IVATFusion network, and it is determined whether the loss is less than a preset threshold. The expression for calculating the loss between the fused image and the preset true label image is:

[0076] ,

[0077] ,

[0078] ,

[0079] Where, is the loss between the infrared image feature vector and the visible light image feature vector, is a hyperparameter used to adjust the size of the loss. When >1, the loss increases, the training process becomes more constrained, and the gradient update is larger; When <1, the loss becomes smaller, the training process becomes smoother, and the update pace decreases. is the cross entropy loss, is the temperature hyperparameter, which is used to control the amplification or reduction of similarity. is the target label, a tensor containing the label corresponding to each sample itself, is the unit vector obtained after L1 normalization of the infrared image feature vector, is the infrared image feature vector, is the L1 norm of the infrared image feature vector, is the unit vector obtained after L1 normalization of the visible light image feature vector, is the visible light image feature vector, is the L1 norm of the visible light image feature vector, is transposed;

[0080] If the loss is less than the preset threshold, the training is stopped and the image fusion model is directly output; otherwise, the training is continued until the loss is less than the preset threshold.

[0081] Step S103 , obtaining a real-time target image including a certain power transmission and transformation equipment, inputting the real-time target image into the image fusion model, and the image fusion model outputting a fused image including the certain power transmission and transformation equipment.

[0082] In summary, the method of the present application, by introducing a basic feature extraction module and a multi-scale feature extraction module, can extract global and local features, better reflect modality-specific and modality-shared features. In addition, the use of an RGB information extraction module, an attention weighted fusion module, a deep feature fusion module, and a fine feature fusion module can effectively fuse image information of different modalities, retain more details of the source image, and enhance the contrast and clarity of the image. This method is highly adaptable and flexible, can adapt to different lighting conditions and background environments, and provide high-quality fused images for various electrical equipment application scenarios.

[0083] See also Figure 9 , which shows a structural block diagram of a system for fusing infrared images and visible light images for inspection of power transmission and transformation equipment of the present application.

[0084] like Figure 9 As shown, the inspection infrared image and visible light image fusion system 200 includes an acquisition module 210 , a training module 220 and an output module 230 .

[0085] The acquisition module 210 is configured to acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image includes a historical inspection infrared image and a historical visible light image corresponding to the one historical inspection infrared image;

[0086] A training module 220 is configured to input the at least one historical target image into a preset IVATFusion network, and iteratively train the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module;

[0087] Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module;

[0088] The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features;

[0089] The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area;

[0090] The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature;

[0091] The feature conversion module converts the target fusion feature into a fusion image;

[0092] The output module 230 is configured to obtain a real-time target image containing a certain power transmission and transformation equipment, input the real-time target image into the image fusion model, and the image fusion model outputs a fused image containing the certain power transmission and transformation equipment.

[0093] It should be understood that Figure 9 Modules and references documented in Figure 1 Therefore, the operations and features described above for the method and the corresponding technical effects also apply to Figure 9 The modules in it will not be described in detail here.

[0094] In other embodiments, embodiments of the present invention further provide a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor is caused to execute the method for fusing inspection infrared images and visible light images of power transmission and transformation equipment in any of the above method embodiments;

[0095] As an embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are configured as follows:

[0096] Acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image contains a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image;

[0097] Inputting the at least one historical target image into a preset IVATFusion network, iteratively training the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module;

[0098] Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module;

[0099] The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features;

[0100] The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area;

[0101] The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature;

[0102] The feature conversion module converts the target fusion feature into a fusion image;

[0103] A real-time target image including a certain power transmission and transformation equipment is obtained, and the real-time target image is input into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation equipment.

[0104] The computer-readable storage medium may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data generated based on the use of the system for fusion of infrared images and visible light images of power transmission and transformation equipment. Furthermore, the computer-readable storage medium may include high-speed random access memory and may also include storage, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include storage remote from the processor. Such remote storage may be connected to the system for fusion of infrared images and visible light images of power transmission and transformation equipment via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0105] Figure 10Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 10 As shown, the device includes: a processor 310 and a memory 320. The electronic device may also include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330 and the output device 340 may be connected via a bus or other means. Figure 10 The example of the bus connection is taken. The memory 320 is the computer-readable storage medium mentioned above. The processor 310 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 320, that is, implements the method for fusing the inspection infrared image and visible light image of the power transmission and transformation equipment in the above-mentioned method embodiment. The input device 330 can receive input digital or character information, and generate key signal input related to the user settings and function control of the inspection infrared image and visible light image fusion system of the power transmission and transformation equipment. The output device 340 may include a display device such as a display screen.

[0106] The electronic device can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided by the embodiment of the present invention.

[0107] As an embodiment, the electronic device is applied to a system for fusing infrared images and visible light images of inspections of power transmission and transformation equipment, and is used for a client, and includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:

[0108] Acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image contains a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image;

[0109] Inputting the at least one historical target image into a preset IVATFusion network, iteratively training the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module;

[0110] Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module;

[0111] The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features;

[0112] The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area;

[0113] The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature;

[0114] The feature conversion module converts the target fusion feature into a fusion image;

[0115] A real-time target image including a certain power transmission and transformation equipment is obtained, and the real-time target image is input into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation equipment.

[0116] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for fusing inspection infrared images and visible light images of power transmission and transformation equipment, characterized in that: include: Acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image contains a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image; Inputting the at least one historical target image into a preset IVATFusion network, iteratively training the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module; Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features; The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area; The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature; The feature conversion module converts the target fusion feature into a fusion image; A real-time target image including a certain power transmission and transformation equipment is obtained, and the real-time target image is input into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation equipment.

2. The method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to claim 1, characterized in that: Inputting the at least one historical target image into a preset IVATFusion network and iteratively training the IVATFusion network to obtain an image fusion model includes: The loss between the infrared image features in the historical inspection infrared image and the visible light image features in the historical visible light image is calculated according to the contrast loss function of the IVATFusion network, and it is determined whether the loss is less than a preset threshold. The expression for calculating the loss between the fused image and the preset true label image is: , , , Where, is the loss between the infrared image feature vector and the visible light image feature vector, is a hyperparameter used to adjust the size of the loss. When >1, the loss increases, the training process becomes more constrained, and the gradient update is larger; When <1, the loss becomes smaller, the training process becomes smoother, and the update pace decreases. is the cross entropy loss, is the temperature hyperparameter, which is used to control the amplification or reduction of similarity. is the target label, a tensor containing the label corresponding to each sample itself, is the unit vector obtained after L1 normalization of the infrared image feature vector, is the infrared image feature vector, is the L1 norm of the infrared image feature vector, is the unit vector obtained after L1 normalization of the visible light image feature vector, is the visible light image feature vector, is the L1 norm of the visible light image feature vector, is transposed; If the loss is less than the preset threshold, the training is stopped and the image fusion model is directly output; otherwise, the training is continued until the loss is less than the preset threshold.

3. The method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to claim 1, characterized in that: The basic feature extraction module includes: a first dynamic wavelet packet pyramid convolution block, a first 2x2 maximum pooling layer, a second dynamic wavelet packet pyramid convolution block, a second 2x2 maximum pooling layer, a third dynamic wavelet packet pyramid convolution block and a 1x1 adaptive average pooling layer connected in sequence; The multi-scale feature extraction module includes: a first channel, a second channel, and a feature fusion module connected to the first channel and the second channel respectively, wherein the first channel includes a non-local attention module and a 3x3 convolution kernel connected to the non-local attention module; the second channel includes a channel attention module and a 3x3 convolution kernel connected to the channel attention module.

4. The method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to claim 3, characterized in that: The first dynamic wavelet packet pyramid convolution block, the second dynamic wavelet packet pyramid convolution block and the third dynamic wavelet packet pyramid convolution block each include: a dynamic wavelet packet transform module, a depthwise separable deformable convolution and a dynamic pyramid feature fusion module connected in sequence; Obtaining visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module specifically includes: Extracting multi-level features of the visible light image or the infrared image according to the dynamic wavelet packet transform module and the depthwise separable deformable convolution; fusing the multi-level features according to the dynamic pyramid feature fusion module to obtain a first result; Compressing the spatial size of the first result to a fixed 1x1 using the maximum pooling layer and the adaptive average pooling layer, and outputting a feature vector containing 256 channels; Inputting the feature vector into the first channel and the second channel respectively, the first channel outputs a second result, and the second channel outputs a third result; The second result and the third result are input into the feature fusion module, and the feature fusion module outputs the visible light image feature or the infrared image feature.

5. The method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to claim 1, characterized in that: The attention weighted fusion module includes: a first attention channel, a second attention channel, a weighted fusion module connected to the first attention channel and the second attention channel, a 7x7 convolution kernel connected to the weighted fusion module, and a Sigmoid activation function connected to the 7x7 convolution kernel.

6. The method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to claim 1, characterized in that: The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high-temperature area, specifically including: A histogram equalization process is performed on the historical inspection infrared image according to a preset temperature threshold to obtain a contrast-enhanced infrared grayscale image, and feature extraction is performed on the infrared grayscale image through a feature extraction module. The extracted features are respectively input into a threshold prediction branch and a region segmentation branch. The threshold prediction branch outputs a dynamic threshold in the range of 0-1 by performing pooling and full connection operations on the extracted features, and multiplies it by 255 to map it to the image grayscale threshold. The region segmentation branch outputs a segmentation probability map of the same size as the historical inspection infrared image by performing convolution and upsampling operations on the extracted features. Each pixel value in the segmentation probability map represents the probability of belonging to a high-temperature area. Compare each pixel value in the segmentation probability map with the image grayscale threshold; The pixel values greater than the image grayscale threshold are set to 1, and the pixel values not greater than the image grayscale threshold are set to 0, and a binary mask matrix containing 0 and 1 is generated according to the comparison result, and the binary mask matrix is multiplied pixel by pixel by pixel by each pixel value in the historical inspection infrared image to obtain the target historical inspection infrared image, wherein the target historical inspection infrared image only contains high-temperature areas greater than the temperature threshold; Convolving the target historical inspection infrared image according to a 3×3 convolution kernel and a 5×5 convolution kernel, respectively, to obtain a first feature map corresponding to the 3×3 convolution kernel and a second feature map corresponding to the 5×5 convolution kernel, wherein the first feature map and the second feature map have the same spatial size; The first feature map and the second feature map are directly stacked in the channel dimension, and the stacked feature maps are convolved according to a 1×1 convolution kernel to obtain the RGB features of the high-temperature area.

7. The method for fusing inspection infrared images and visible light images of power transmission and transformation equipment according to claim 1, characterized in that: The deformable convolution alignment module includes a 3x3 convolution kernel, a Relu activation function and a 3x3 convolution kernel connected in sequence; The deep feature fusion module includes a 3x3 convolution kernel, a first residual block, a second residual block, and a third residual block connected in sequence, wherein the first residual block, the second residual block, and the third residual block each include a 3x3 convolution kernel, a ReLU activation function, and a 3x3 convolution kernel connected in sequence; The fine feature fusion module includes a 3x3 convolution kernel, a Relu activation function, a 3x3 convolution kernel, a 3x3 convolution kernel, an average pooling layer and a 1x1 convolution kernel connected in sequence.

8. A system for fusing infrared images and visible light images for inspection of power transmission and transformation equipment, characterized in that: include: an acquisition module configured to acquire at least one historical target image containing power transmission and transformation equipment, wherein one historical target image includes a historical inspection infrared image and a historical visible light image corresponding to the historical inspection infrared image; A training module is configured to input the at least one historical target image into a preset IVATFusion network, and iteratively train the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature conversion module; Acquire visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features; The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high temperature area; The deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature area RGB feature, and fuses the aligned fusion feature and the high-temperature area RGB feature through the deep feature fusion module and the fine feature fusion module to obtain the target fusion feature; The feature conversion module converts the target fusion feature into a fusion image; The output module is configured to obtain a real-time target image containing a certain power transmission and transformation equipment, input the real-time target image into the image fusion model, and the image fusion model outputs a fused image containing the certain power transmission and transformation equipment.

9. An electronic device, characterized in that: include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for identifying bimodal defects of power transformation equipment

    CN119314113A

  • Infrared and visible light image fusion method for inspection of power transmission and transformation equipment

    CN119540702A