Inspection infrared image and visible light image fusion method and system for power transmission and transformation equipment

Through the multi-level feature extraction and fusion method of the IVATFusion network, the problem that the infrared and visible light image fusion method in the prior art is difficult to dig deep information, and high-quality fusion images are generated, which improves the accuracy and detailed performance of electrical equipment failure detection.

CN120259101AActive Publication Date: 2025-07-04NANCHANG INST OF TECH

Patent Information

Application Number
CN202510740867.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The existing infrared and visible light image fusion methods are difficult to effectively dig deep information in the image, resulting in insufficient details of the fused image and unable to fully demonstrate the advantages of infrared and visible light images.

Method used

The IVATFusion network is adopted, which includes basic feature extraction module, multi-scale feature extraction module, attention-weighted fusion module, RGB information extraction module, deformable convolution alignment module, deep feature fusion module and fine feature fusion module. The image fusion model is obtained through iterative training, global and local features are extracted, and attention-weighted and feature fusion is performed to generate high-quality fusion images.

Benefits of technology

It improves the adaptability and flexibility of image fusion, can adapt to different lighting conditions and background environments, generate high-quality fusion images, enhances the contrast and clarity of the image, and improves the accuracy and comprehensiveness of fault judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259101A_ABST
    Figure CN120259101A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a system for fusing an inspection infrared image and a visible light image of power transmission and transformation equipment. The method comprises the following steps: acquiring at least one historical target image containing the power transmission and transformation equipment; and inputting the at least one historical target image into a preset IVATFusion network, and carrying out iterative training on the IVATFusion network to obtain an image fusion model, the IVATFusion network comprises a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a depth feature fusion module, a fine feature fusion module and a feature conversion module; and acquiring a certain real-time target image including a certain power transmission and transformation device, inputting the certain real-time target image into the image fusion model, and outputting the image fusion model to obtain a fusion image including the certain power transmission and transformation device. The method can adapt to different illumination conditions and background environments, and provides high-quality fusion images for various electrical equipment application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method and system for fusing infrared images and visible light images for inspection of power transmission and transformation equipment. Background Art

[0002] The power system requires that equipment has stability, reliability, and convenient operation methods during long-term operation, which is particularly crucial for ensuring power supply safety. However, electrical equipment frequently fails, and temperature anomalies are one of the precursors of many faults. If the fault type cannot be identified in time and repair measures are not taken, it may lead to serious economic losses. Therefore, infrared and visible light image fusion technology plays an important role in power system monitoring.

[0003] Due to the special geographical location of high-voltage electrical equipment, current power grid detection mainly relies on three methods: intelligent inspection robots, unmanned aerial vehicles, and manual detection. Among them, inspection robots and unmanned aerial vehicles that combine multi-modal imaging sensors and image processing technology can capture infrared and visible light images generated during equipment operation, enabling real-time monitoring and fault diagnosis of equipment.

[0004] Infrared images can effectively reflect the temperature information of electrical equipment, while visible light images can clearly present the appearance details of the equipment. However, these two types of images have differences in imaging mechanisms, resolutions, and fields of view, which limit the detection effect when used alone. Therefore, in power equipment fault detection, infrared and visible light image fusion technology is adopted to meet the actual application requirements. By fusing the features of the two images, the heat sources of the equipment can be accurately identified and located, thereby improving the accuracy and comprehensiveness of fault determination.

[0005] Existing image fusion methods (such as DenseFuse, FusionGAN, etc.) mainly focus on image fusion. Usually, the images to be fused need to be registered first, and then an algorithm is used for fusion. Although these methods have achieved certain results in image fusion, in practical applications, input images with excessive noise will significantly affect the fusion effect. In addition, existing methods often can only extract shallow features of images and are difficult to effectively mine the deep information in the images, resulting in insufficient detail performance in the fused images and unable to fully display the respective advantages of infrared images and visible light images. Summary of the Invention

[0006] The present invention provides a method and system for fusing infrared images and visible light images for inspection of power transmission and transformation equipment, which are used to solve the technical problem that it is difficult to effectively mine the deep information in the images, resulting in insufficient detail performance in the fused images.

[0007] In a first aspect, the present invention provides a method for fusing inspection infrared images and visible light images of power transmission and transformation equipment, including: Obtain at least one historical target image including power transmission and transformation equipment, where one historical target image includes one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image; Input the at least one historical target image into a preset IVATFusion network, and perform iterative training on the IVATFusion network to obtain an image fusion model, where the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features; The RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features; The deformable convolution alignment module performs deformable convolution alignment on the fused features and the high-temperature region RGB features, and performs feature fusion on the aligned fused features and high-temperature region RGB features through the deep feature fusion module and the fine feature fusion module to obtain target fused features; The feature transformation module transforms the target fused features into a fused image; Obtain a certain real-time target image including a certain power transmission and transformation equipment, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation equipment.

[0008] In a second aspect, the present invention provides a system for fusing inspection infrared images and visible light images of power transmission and transformation equipment, including: An acquisition module configured to acquire at least one historical target image including power transmission and transformation equipment, where one historical target image includes one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image; A training module, configured to input the at least one historical target image into a preset IVATFusion network, and iteratively train the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a depth feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features; The RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features; The deformable convolution alignment module performs deformable convolution alignment on the fused features and the high-temperature region RGB features, and performs feature fusion on the aligned fused features and high-temperature region RGB features through the depth feature fusion module and the fine feature fusion module to obtain target fused features; The feature transformation module transforms the target fused features into a fused image; An output module, configured to obtain a certain real-time target image including a certain power transmission and transformation device, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation device.

[0009] In a third aspect, an electronic device is provided, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method for fusing infrared images and visible light images for inspection of power transmission and transformation devices according to any embodiment of the present invention.

[0010] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program instructions are executed by a processor, the processor is enabled to execute the steps of the method for fusing infrared images and visible light images for inspection of power transmission and transformation devices according to any embodiment of the present invention.

[0011] The method and system for fusing the infrared inspection image and visible light image of the power transmission and transformation equipment of the present application can extract global and local features by introducing a basic feature extraction module and a multi-scale feature extraction module, which can better reflect the modality-specific and modality-shared features. In addition, by using an RGB information extraction module, an attention-weighted fusion module, a depth feature fusion module, and a fine feature fusion module, the image information of different modalities can be effectively fused, more details of the source images can be retained, and the contrast and clarity of the images can be enhanced. This method has high adaptability and flexibility, can adapt to different lighting conditions and background environments, and provides high-quality fused images for various electrical equipment application scenarios. Description of the Drawings

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0013] Figure 1 It is a flowchart of a method for fusing the infrared inspection image and visible light image of the power transmission and transformation equipment provided by an embodiment of the present invention; Figure 2 It is a partial flowchart block diagram of a method for fusing the infrared inspection image and visible light image of the power transmission and transformation equipment provided by a specific embodiment of the present invention; Figure 3 It is a structural block diagram of a basic feature extraction module provided by a specific embodiment of the present invention; Figure 4 It is a structural block diagram of a dynamic wavelet packet pyramid convolution block provided by a specific embodiment of the present invention; Figure 5 It is a structural block diagram of a multi-scale feature extraction module provided by a specific embodiment of the present invention; Figure 6 It is a structural block diagram of an attention-weighted fusion module provided by a specific embodiment of the present invention; Figure 7 It is a structural block diagram of an RGB information extraction module provided by a specific embodiment of the present invention; Figure 8 It is a structural block diagram of a deformable convolution alignment module, a depth feature fusion module, and a fine feature fusion module provided by a specific embodiment of the present invention; Figure 9 It is a structural block diagram of a system for fusing the infrared inspection image and visible light image of the power transmission and transformation equipment provided by an embodiment of the present invention; Figure 10It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0014] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0015] Please refer to Figure 1 , which shows a flowchart of a method for fusing inspection infrared images and visible light images of a power transmission and transformation device in this application.

[0016] As Figure 1 shown, the method for fusing inspection infrared images and visible light images of a power transmission and transformation device specifically includes the following steps: Step S101, obtaining at least one historical target image including a power transmission and transformation device, where one historical target image includes one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image.

[0017] In this step, one historical target image includes a historical inspection infrared image and a historical visible light image of the same power transmission and transformation device.

[0018] Step S102, inputting the at least one historical target image into a preset IVATFusion network, and iteratively training the IVATFusion network to obtain an image fusion model, where the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a depth feature fusion module, a fine feature fusion module, and a feature transformation module.

[0019] In this step, please refer to Figure 2 , and obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module.

[0020] Specifically: multi-level features of a visible light image or an infrared image are extracted according to a dynamic wavelet packet transform module and a depthwise separable deformable convolution; the multi-level features are fused according to a dynamic pyramid feature fusion module to obtain a first result; a maximum pooling layer and an adaptive average pooling layer are used to compress the spatial dimension of the first result to a fixed 1x1, and a feature vector containing 256 channels is output; the feature vector is respectively input into a first channel and a second channel, the first channel outputs a second result, and the second channel outputs a third result; the second result and the third result are input into a feature fusion module, and the feature fusion module outputs visible light image features or infrared image features.

[0021] Please refer to Figure 3 , the basic feature extraction module includes: a first dynamic wavelet packet pyramid convolution block, a 2x2 first maximum pooling layer, a second dynamic wavelet packet pyramid convolution block, a 2x2 second maximum pooling layer, a third dynamic wavelet packet pyramid convolution block, and a 1x1 adaptive average pooling layer connected in sequence.

[0022] Please refer to Figure 4 , the first dynamic wavelet packet pyramid convolution block, the second dynamic wavelet packet pyramid convolution block, and the third dynamic wavelet packet pyramid convolution block all include: a dynamic wavelet packet transform module, a depthwise separable deformable convolution, and a dynamic pyramid feature fusion module connected in sequence.

[0023] Basic feature extraction module: This module extracts multi-scale and high-dimensional features from the input image through the combination of a depthwise separable deformable convolution, a dynamic wavelet packet transform module, and a dynamic pyramid feature fusion module. First, multi-level features of the image are extracted through the dynamic wavelet packet transform module and the depthwise separable deformable convolution, and then information of different scales is fused through the dynamic pyramid feature fusion module. Finally, the maximum pooling layer and the adaptive average pooling layer are used to compress the spatial dimension of the feature map to a fixed 1x1, and a feature vector containing 256 channels is output, providing a basis for feature extraction of subsequent networks.

[0024] Please refer to Figure 5 , the multi-scale feature extraction module includes: a first channel, a second channel, and a feature fusion module respectively connected to the first channel and the second channel, wherein the first channel contains a non-local attention module and a 3x3 convolution kernel connected to the non-local attention module; the second channel contains a channel attention module and a 3x3 convolution kernel connected to the channel attention module.

[0025] It should be noted that the non-local attention module is used to capture long-range dependencies, and the CBAM channel attention is used to enhance key information. Subsequently, fine texture and large structure features are extracted through ordinary 3×3 convolution and 3×3 dilated convolution. The outputs of the two paths are directly concatenated in the channel dimension to preserve scale independence, and finally, 1×1 convolution is used for adaptive compression and weighting to achieve efficient and information-complete multi-scale feature fusion.

[0026] The non-local attention module is: Non local Block.

[0027] Multi-scale feature extraction module: This module combines a non-local attention module and a channel attention module to further extract multi-scale features. By using different types of convolutional kernels (standard convolution and dilated convolution), the feature map is processed to capture multi-scale information in the image. At the same time, the non-local and channel attention mechanisms help enhance the global and local information of the feature map, thereby improving the richness and accuracy of feature representation.

[0028] Furthermore, the attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fused features.

[0029] Please refer to Figure 6 , the attention weighted fusion module includes: a first attention channel, a second attention channel, a weighted fusion module connected to the first attention channel and the second attention channel, a 7x7 convolutional kernel connected to the weighted fusion module, and a Sigmoid activation function connected to the 7x7 convolutional kernel. Specifically, the weighted fusion module is a CBAM (ConvolutionalBlockAttentionModule) module.

[0030] Attention weighted fusion module: It performs weighted fusion on the features from the visible light image and the infrared image, and at the same time enhances important features through the attention mechanism and suppresses irrelevant or unimportant information. This module selectively emphasizes the most critical parts of the image during the fusion process by combining two different attention mechanisms before and after, improving the quality of the fused feature representation.

[0031] Furthermore, the RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features.

[0032] RGB Information Extraction Module: While processing the original infrared image, this module needs to set the threshold of the normal temperature. Then, perform histogram equalization on the infrared image to increase the contrast of the image, making the details clearer. Next, convert it into a numpy array. Then, by retaining the regions with gray values greater than the normal temperature and setting other regions to black, and then converting it back into an RGB image, extract the regions above the normal temperature, and map the colors of these regions to the corresponding regions in the original infrared image. Finally, generate an image with other regions set to black and the RGB information of the regions above the normal temperature retained. Then, use two specific-sized convolutional kernels of 3×3 and 5×5 in parallel for unconventional continuous stacked convolutions to achieve fine-grained multi-scale feature separation and extraction. And use channel dimension concatenation instead of element-wise addition to retain the independence of the original scale features. Finally, complete the final feature fusion through 1×1 convolution, controlling the growth of the computational amount while enhancing the representation ability, and finally generate the RGB features of the high-temperature region.

[0033] Please refer to Figure 7 , perform histogram equalization on the historical inspection infrared image according to the preset temperature threshold to obtain an infrared grayscale image with enhanced contrast. And perform feature extraction on the infrared grayscale image through the feature extraction module, and input the extracted features into the threshold prediction branch and the region segmentation branch respectively. The threshold prediction branch outputs a dynamic threshold within the range of 0-1 through pooling and fully connected operations on the extracted features, and multiplies it by 255 to map it to an image grayscale threshold. The region segmentation branch outputs a segmentation probability map with the same size as the historical inspection infrared image through convolution and upsampling operations on the extracted features. Each pixel value in the segmentation probability map represents the probability of belonging to the high-temperature region; Among them, the threshold prediction branch includes an adaptive average pooling layer, a flattening module for flattening the input multi-dimensional tensor into a single dimension, a fully connected layer, a Relu activation function, a fully connected layer, and a Sigmoid activation function connected in sequence; The region segmentation branch includes a 3x3 convolutional kernel, a Relu activation function, an upsampling layer, a 3x3 convolutional kernel, a Relu activation function, an upsampling layer, a 1x1 convolutional kernel, and a Sigmoid activation function connected in sequence; Compare each pixel value in the segmentation probability map with the image grayscale threshold; Set the pixel values greater than the image grayscale threshold to 1, and the pixel values not greater than the image grayscale threshold to 0, and generate a binary mask matrix containing 0 and 1 according to the comparison result. Then, multiply each pixel value in the binary mask matrix with each pixel value in the historical inspection infrared image to obtain the target historical inspection infrared image, where only the high-temperature regions greater than the temperature threshold are included in the target historical inspection infrared image; Convolve the target historical inspection infrared image with a 3×3 convolution kernel and a 5×5 convolution kernel respectively to obtain a first feature map corresponding to the 3×3 convolution kernel and a second feature map corresponding to the 5×5 convolution kernel, where the spatial dimensions of the first feature map and the second feature map are the same; Stack the first feature map and the second feature map directly in the channel dimension, and convolve the stacked feature maps with a 1×1 convolution kernel to obtain the high-temperature region RGB features.

[0034] Furthermore, the deformable convolution alignment module performs deformable convolution alignment on the fusion feature and the high-temperature region RGB features, and performs feature fusion on the aligned fusion feature and the high-temperature region RGB features through the depth feature fusion module and the fine feature fusion module to obtain the target fusion feature.

[0035] Deformable convolution alignment module: Align the fusion features of the visible light image and the infrared image and the features of the RGB information. By calculating the offset and using deformable convolution to adjust the spatial structure of the high-temperature region RGB features to make them accurately aligned with the features of the fusion image, so as to realize the feature fusion of cross-modal images.

[0036] Depth feature fusion module: Perform depth processing on the fused features to further enhance the feature expression ability through residual blocks. It combines the high-temperature region RGB features with the fusion features and extracts richer features through multi-layer convolution operations, so as to improve the details and expressiveness of the fusion image and provide more accurate feature representations for subsequent tasks.

[0037] Fine feature fusion module: Perform fine processing on multi-scale features. By multi-scale feature extraction and fusion, enhance the feature expression ability and improve the effect of image processing or image generation tasks. It enables the model to effectively capture features at different scales through operations such as convolution, pooling, and upsampling, thereby improving the overall quality and detail performance of the image.

[0038] Please refer to Figure 8 , the deformable convolution alignment module includes a 3x3 convolution kernel, a Relu activation function, and a 3x3 convolution kernel connected in sequence.

[0039] The depth feature fusion module includes a 3x3 convolution kernel, a first residual block, a second residual block, and a third residual block connected in sequence, where the first residual block, the second residual block, and the third residual block all include a 3x3 convolution kernel, a Relu activation function, and a 3x3 convolution kernel connected in sequence.

[0040] The fine feature fusion module includes a 3x3 convolution kernel, a Relu activation function, a 3x3 convolution kernel, a 3x3 convolution kernel, an average pooling layer, and a 1x1 convolution kernel connected in sequence.

[0041] Furthermore, the feature transformation module transforms the target fusion feature into a fused image.

[0042] Feature transformation module: Transforms the features that have undergone deep processing and fusion into the final output image. It decodes the high-dimensional features into a meaningful image output through a series of convolutional layers and upsampling operations, and ensures that the size and format of the output image meet the expectations.

[0043] It should be noted that according to the contrast loss function of the IVATFusion network, the loss between the infrared image features in the historical inspection infrared images and the visible light image features in the historical visible light images is calculated, and it is determined whether the loss is less than a preset threshold. Among them, the expression for calculating the loss between the fused image and the preset true label image is: , , , In the formula, is the loss between the infrared image feature vector and the visible light image feature vector, is a hyperparameter for adjusting the size of the loss. When > 1, the loss increases, the training process becomes more restrictive, and the gradient update is larger; while when < 1, the loss decreases, the training process is smoother, and the update step size decreases. is the cross-entropy loss, is the temperature hyperparameter used to control the amplification or reduction of similarity, is the target label, a tensor containing the corresponding label of each sample itself, is the unit vector obtained after L1 normalization of the infrared image feature vector, is the infrared image feature vector, is the L1 norm of the infrared image feature vector, is the unit vector obtained after L1 normalization of the visible light image feature vector, is the visible light image feature vector, is the L1 norm of the visible light image feature vector, is the transpose; If it is less than the preset threshold, the training is stopped and the image fusion model is directly output; otherwise, the training continues until the loss is less than the preset threshold.

[0044] Step S103: Obtain a certain real-time target image containing a certain power transmission and transformation equipment, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fusion image containing the certain power transmission and transformation equipment.

[0045] In summary, in the method of the present application, by introducing a basic feature extraction module and a multi-scale feature extraction module, this method can extract global and local features, better reflect modality-specific and modality-shared features. In addition, by using an RGB information extraction module, an attention-weighted fusion module, a depth feature fusion module, and a fine feature fusion module, it is possible to effectively fuse image information of different modalities, retain more details of the source images, enhance the contrast and clarity of the images. This method has high adaptability and flexibility, can adapt to different lighting conditions and background environments, and provides high-quality fusion images for various electrical equipment application scenarios.

[0046] Please refer to Figure 9 , which shows a structural block diagram of an inspection infrared image and visible light image fusion system for a power transmission and transformation equipment of the present application.

[0047] As Figure 9 shown, the inspection infrared image and visible light image fusion system 200 includes an acquisition module 210, a training module 220, and an output module 230.

[0048] Among them, the acquisition module 210 is configured to acquire at least one historical target image containing a power transmission and transformation equipment, where one historical target image contains one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image; The training module 220 is configured to input the at least one historical target image into a preset IVATFusion network, perform iterative training on the IVATFusion network to obtain an image fusion model, where the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention-weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a depth feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention-weighted fusion module performs attention-weighted fusion on the visible light image features and the infrared image features to obtain fusion features; The RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features; The deformable convolution alignment module performs deformable convolution alignment on the fused feature and the RGB feature of the high-temperature region, and performs feature fusion on the aligned fused feature and the RGB feature of the high-temperature region through the depth feature fusion module and the fine feature fusion module to obtain a target fused feature; The feature transformation module transforms the target fused feature into a fused image; The output module 230 is configured to obtain a certain real-time target image including a certain power transmission and transformation device, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation device.

[0049] It should be understood that Figure 9 The modules described in Figure 1 correspond to the respective steps in the method described in the reference Figure 9 Therefore, the operations, features, and corresponding technical effects described above for the method also apply to

[0050] In some other embodiments, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the program instructions are executed by a processor, the processor executes the method for fusing infrared inspection images and visible light images of power transmission and transformation equipment in any of the above method embodiments; As an implementation manner, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as follows: Obtain at least one historical target image including a power transmission and transformation device, where one historical target image includes one historical infrared inspection image and one historical visible light image corresponding to the one historical infrared inspection image; Input the at least one historical target image into a preset IVATFusion network, and perform iterative training on the IVATFusion network to obtain an image fusion model, where the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a depth feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain a fused feature; The RGB information extraction module extracts RGB information from the infrared image to obtain RGB features of the high-temperature region; The deformable convolution alignment module performs deformable convolution alignment on the fused feature and the RGB feature of the high-temperature area, and performs feature fusion on the aligned fused feature and the RGB feature of the high-temperature area through the depth feature fusion module and the fine feature fusion module to obtain a target fused feature; The feature conversion module converts the target fused feature into a fused image; Obtain a certain real-time target image containing a certain power transmission and transformation equipment, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fused image containing the certain power transmission and transformation equipment.

[0051] The computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area may store an operating system and application programs required for at least one function; the storage data area may store data created according to the use of the infrared image and visible light image fusion system for inspection of power transmission and transformation equipment. In addition, the computer-readable storage medium may include a high-speed random access memory, and may also include a memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the computer-readable storage medium may optionally include a memory remotely provided with respect to the processor, and these remote memories may be connected to the infrared image and visible light image fusion system for inspection of power transmission and transformation equipment through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0052] Figure 10 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention, as Figure 10 shown, the device includes: a processor 310 and a memory 320. The electronic device may further include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 10 taking connection through the bus as an example. The memory 320 is the above-mentioned computer-readable storage medium. The processor 310 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the infrared image and visible light image fusion method for inspection of power transmission and transformation equipment in the above method embodiment. The input device 330 may receive input digital or character information, and generate key signal inputs related to user settings and function controls of the infrared image and visible light image fusion system for inspection of power transmission and transformation equipment. The output device 340 may include a display device such as a display screen.

[0053] The above electronic device can execute the method provided by the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference may be made to the method provided by the embodiments of the present invention.

[0054] As an implementation manner, the above electronic device is applied to a fusion system of infrared inspection images and visible light images of power transmission and transformation equipment for a client, and includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Obtain at least one historical target image including power transmission and transformation equipment, wherein one historical target image includes one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image; Input the at least one historical target image into a preset IVATFusion network, and perform iterative training on the IVATFusion network to obtain an image fusion model, wherein the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a depth feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fusion features; The RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features; The deformable convolution alignment module performs deformable convolution alignment on the fusion features and the high-temperature region RGB features, and performs feature fusion on the aligned fusion features and high-temperature region RGB features through the depth feature fusion module and the fine feature fusion module to obtain target fusion features; The feature transformation module transforms the target fusion features into a fusion image; Obtain a certain real-time target image including a certain power transmission and transformation equipment, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fusion image including the certain power transmission and transformation equipment.

[0055] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A method for fusing infrared inspection images and visible light images of power transmission and transformation equipment, characterized in that, Including: Obtain at least one historical target image including power transmission and transformation equipment, where one historical target image includes one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image; Input the at least one historical target image into a preset IVATFusion network, and perform iterative training on the IVATFusion network to obtain an image fusion model, where the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fusion features; The RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features; The deformable convolution alignment module performs deformable convolution alignment on the fusion features and the high-temperature region RGB features, and performs feature fusion on the aligned fusion features and high-temperature region RGB features through the deep feature fusion module and the fine feature fusion module to obtain target fusion features; The feature transformation module transforms the target fusion features into a fusion image; Obtain a certain real-time target image including a certain power transmission and transformation equipment, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fusion image including the certain power transmission and transformation equipment.

2. The method for fusing infrared images and visible light images in the inspection of power transmission and transformation equipment according to claim 1, characterized in that The step of inputting the at least one historical target image into a preset IVATFusion network and performing iterative training on the IVATFusion network to obtain an image fusion model includes: Calculate the loss between the infrared image features in the historical inspection infrared image and the visible light image features in the historical visible light image according to the contrast loss function of the IVATFusion network, and determine whether the loss is less than a preset threshold. The expression for calculating the loss between the fusion image and a preset true label image is: , , , In the formula, is the loss between the infrared image feature vector and the visible light image feature vector, is a hyperparameter used to adjust the magnitude of the loss. When > 1, the loss increases, the training process becomes more restrictive, and the gradient update is larger; while when < 1, the loss decreases, the training process is smoother, and the update step size decreases. is the cross-entropy loss, is the temperature hyperparameter used to control the amplification or reduction of similarity, is the target label, a tensor containing the corresponding label for each sample itself, is the unit vector obtained after L1 normalization of the infrared image feature vector, is the infrared image feature vector, is the L1 norm of the infrared image feature vector, is the unit vector obtained after L1 normalization of the visible light image feature vector, is the visible light image feature vector, is the L1 norm of the visible light image feature vector, is the transpose; If it is less than the preset threshold, stop training and directly output the image fusion model; otherwise, continue training until the loss is less than the preset threshold.

3. A method for fusing infrared images and visible light images in the inspection of power transmission and transformation equipment according to claim 1, characterized in that, The basic feature extraction module includes: a first dynamic wavelet packet pyramid convolution block connected in sequence, a 2x2 first max pooling layer, a second dynamic wavelet packet pyramid convolution block, a 2x2 second max pooling layer, a third dynamic wavelet packet pyramid convolution block, and a 1x1 adaptive average pooling layer; The multi-scale feature extraction module includes: a first channel, a second channel, and a feature fusion module connected to the first channel and the second channel respectively. Among them, the first channel contains a non-local attention module and a 3x3 convolutional kernel connected to the non-local attention module; the second channel contains a channel attention module and a 3x3 convolutional kernel connected to the channel attention module.

4. The inspection infrared image and visible light image fusion method for a power transmission and transformation equipment according to claim 3, characterized in that, The first dynamic wavelet packet pyramid convolution block, the second dynamic wavelet packet pyramid convolution block, and the third dynamic wavelet packet pyramid convolution block all include: a dynamic wavelet packet transform module, a depthwise separable deformable convolution, and a dynamic pyramid feature fusion module connected in sequence; Obtain visible light image features in the visible light image and infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module, specifically including: Extract multi-level features of the visible light image or the infrared image according to the dynamic wavelet packet transform module and the depthwise separable deformable convolution; Fuse the multi-level features according to the dynamic pyramid feature fusion module to obtain a first result; Use the max pooling layer and the adaptive average pooling layer to compress the spatial size of the first result to a fixed 1x1, and output a feature vector containing 256 channels; Input the feature vector into the first channel and the second channel respectively. The first channel outputs a second result, and the second channel outputs a third result; Input the second result and the third result into the feature fusion module, and the feature fusion module outputs the visible light image features or the infrared image features.

5. A method for fusing infrared images and visible light images for inspection of power transmission and transformation equipment according to claim 1, characterized in that, The attention weighted fusion module includes: a first attention channel, a second attention channel, a weighted fusion module connected to the first attention channel and the second attention channel, a 7x7 convolutional kernel connected to the weighted fusion module, and a Sigmoid activation function connected to the 7x7 convolutional kernel.

6. A method for fusing infrared images and visible light images in the inspection of power transmission and transformation equipment according to claim 1, characterized in that, The RGB information extraction module extracts RGB information from the infrared image to obtain high-temperature region RGB features, specifically including: Perform histogram equalization processing on the historical inspection infrared image according to a preset temperature threshold to obtain a contrast-enhanced infrared grayscale image, and perform feature extraction on the infrared grayscale image through a feature extraction module, and input the extracted features into a threshold prediction branch and a region segmentation branch respectively. The threshold prediction branch outputs a dynamic threshold within the range of 0-1 through pooling and fully connected operations on the extracted features, and multiplies it by 255 to map it to an image grayscale threshold. The region segmentation branch outputs a segmentation probability map with the same size as the historical inspection infrared image through convolution and upsampling operations on the extracted features. Each pixel value in the segmentation probability map represents the probability of belonging to the high-temperature region; Compare each pixel value in the segmentation probability map with the image grayscale threshold; Set the pixel values greater than the image gray threshold to 1, and set the pixel values not greater than the image gray threshold to 0. Generate a binary mask matrix containing 0s and 1s according to the comparison results, and multiply each pixel value in the historical inspection infrared image by the binary mask matrix to obtain a target historical inspection infrared image, where only the high-temperature regions greater than the temperature threshold are included in the target historical inspection infrared image; Perform convolution on the target historical inspection infrared image according to a 3×3 convolution kernel and a 5×5 convolution kernel respectively to obtain a first feature map corresponding to the 3×3 convolution kernel and a second feature map corresponding to the 5×5 convolution kernel, where the spatial dimensions of the first feature map and the second feature map are the same; Stack the first feature map and the second feature map directly in the channel dimension, and perform convolution on the stacked feature map according to a 1×1 convolution kernel to obtain the high-temperature region RGB features.

7. A method for fusing infrared images and visible light images in the inspection of power transmission and transformation equipment according to claim 1, characterized in that, The deformable convolution alignment module includes a 3x3 convolution kernel, a Relu activation function, and a 3x3 convolution kernel connected in sequence; The deep feature fusion module includes a 3x3 convolution kernel, a first residual block, a second residual block, and a third residual block connected in sequence, where the first residual block, the second residual block, and the third residual block all include a 3x3 convolution kernel, a Relu activation function, and a 3x3 convolution kernel connected in sequence; The fine feature fusion module includes a 3x3 convolution kernel, a Relu activation function, a 3x3 convolution kernel, a 3x3 convolution kernel, an average pooling layer, and a 1x1 convolution kernel connected in sequence.

8. An inspection infrared image and visible light image fusion system for power transmission and transformation equipment, characterized in that, Include: An acquisition module configured to acquire at least one historical target image including power transmission and transformation equipment, where one historical target image includes one historical inspection infrared image and one historical visible light image corresponding to the one historical inspection infrared image; A training module configured to input the at least one historical target image into a preset IVATFusion network, and perform iterative training on the IVATFusion network to obtain an image fusion model, where the IVATFusion network includes a basic feature extraction module, a multi-scale feature extraction module, an attention weighted fusion module, an RGB information extraction module, a deformable convolution alignment module, a deep feature fusion module, a fine feature fusion module, and a feature transformation module; Obtain the visible light image features in the visible light image and the infrared image features in the infrared image according to the basic feature extraction module and the multi-scale feature extraction module; The attention weighted fusion module performs attention weighted fusion on the visible light image features and the infrared image features to obtain fusion features; The RGB information extraction module performs RGB information extraction on the infrared image to obtain high-temperature region RGB features; The deformable convolution alignment module performs deformable convolution alignment on the fused feature and the RGB feature of the high-temperature area, and performs feature fusion on the aligned fused feature and the RGB feature of the high-temperature area through the depth feature fusion module and the fine feature fusion module to obtain a target fused feature; The feature conversion module converts the target fused feature into a fused image; An output module, configured to obtain a certain real-time target image including a certain power transmission and transformation device, input the certain real-time target image into the image fusion model, and the image fusion model outputs a fused image including the certain power transmission and transformation device.

9. An electronic device, characterized in that, Comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for identifying bimodal defects of power transformation equipment

    CN119314113A

  • Infrared and visible light image fusion method for inspection of power transmission and transformation equipment

    CN119540702A

  • Image enhancement method and apparatus, device, and medium

    WO2023169582A1

Cited By

  • Target intelligent monitoring system and method based on dual-optical data

    CN120580648A

  • High-precision infrared image temperature expression method based on U-Net architecture

    CN120599432A

  • Infrared target identification method and system based on deep learning, medium and product

    CN121053371A

  • Detection method and detection device for overhead line system suspension component

    CN121147120A