A method for fusion of infrared and visible light images for power transmission and transformation equipment inspection
By using enhanced contrast preprocessing and deep learning models such as STCA and LTCS modules in infrared and visible image fusion methods, the image features of power transmission and transformation equipment are extracted and fused, and the shortcomings of existing methods in detail performance and deep information mining are solved, and high-performance image fusion and fault positioning effects are achieved.
Patent Information
- Application Number
- CN202510073715.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing infrared and visible light image fusion methods cannot effectively mine deep information in the image, resulting in the lack of detailed performance of the fused image and the inability to fully demonstrate the advantages of infrared and visible light images. At the same time, traditional methods require manual design of fusion rules, with large calculation volume and high complexity.
A method of patrol infrared and visible light image fusion of power transmission and transformation equipment is adopted, and high-frequency local feature encoder based on the LTCS module and INN module are used to achieve high-performance image fusion by enhancing infrared image contrast and removing noise preprocessing operations, combining the dual-branch shared encoder and shared decoder based on the STCA module, extracting and fusing image features, and using a low-frequency global feature encoder based on the LTCS module and the INN module, and a high-frequency local feature encoder based on the LTCS module and the INN module to achieve high-performance image fusion.
This method can well extract and fuse the features of infrared and visible light images, solve the problems of loss of details, insignificant targets and low contrast, realize the accurate positioning of thermal failures of power transmission and transformation equipment, and enhance the accuracy and comprehensiveness of fault determination.
Smart Images

Figure CN119540702B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of power transmission lines, and in particular relates to a method for fusing infrared and visible light images for power transmission and transformation equipment inspection. Background Art
[0002] The power system needs to ensure the stability and reliability of electrical equipment due to its requirements for long-term operation, power supply safety, and convenient operation. However, electrical equipment failures occur frequently, among which temperature abnormality is one of the main manifestations of failure precursors. Failure to identify the fault type in time and take maintenance measures may result in significant economic losses. Therefore, infrared and visible light image fusion technology plays a vital role in power system monitoring.
[0003] In view of the particularity of the geographical distribution of high-voltage electrical equipment, today's power grid inspection work mainly relies on three methods: intelligent inspection robots, drones and manual inspection. Among them, inspection robots and drones using multimodal imaging sensors and image processing technology can achieve real-time monitoring and fault diagnosis of equipment by capturing infrared and visible light images of power equipment in operation.
[0004] Infrared images can reflect the temperature information of electrical equipment, while visible light images can clearly present the appearance details of electrical equipment. However, there are differences in the imaging mechanism, resolution and field of view of these two images, which limits their effectiveness when used alone. In order to meet the needs of practical applications, infrared and visible light image fusion technology is introduced in the field of power equipment fault detection. By fusing the features of the two images, the identification and positioning of the hot spots of the power equipment are realized, and the accuracy and comprehensiveness of the fault judgment are enhanced. And with the development of artificial intelligence technology, deep learning-based methods have been applied to the field of image fusion. There is no need to manually design fusion rules. High-quality fused images can be obtained by training with a large number of data sets.
[0005] Although existing image fusion methods (such as DenseFuse, FusionGAN, etc.) have achieved certain results, they can only extract shallow features of images and cannot effectively mine deep information in images. This leads to a lack of detail in the fused image and the inability to fully demonstrate the respective advantages of infrared and visible light images. In addition, traditional image fusion methods require manual design of fusion rules, which is computationally intensive and complex. In addition, existing feature extraction methods (such as pyramid transform, PCA, multi-scale transform, etc.) can only focus on shallow information of images when extracting image features, but cannot deeply mine deep features of images. This is related to the complexity and diversity of image features and the differences between images, making feature extraction a difficult point in image fusion. Summary of the invention
[0006] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a method for fusing infrared and visible light images for inspection of power transmission and transformation equipment, which can well extract and fuse the features in visible light images and infrared images, thereby realizing high-performance fusion of infrared images and visible light images for inspection of power transmission and transformation equipment, and by enhancing the contrast of infrared images and removing noise preprocessing operations, as well as further enhancing the texture and contrast of high-frequency local features, thereby solving the problems of detail loss, inconspicuous targets and low contrast, thereby providing support for further realizing the precise positioning of thermal faults in power transmission and transformation equipment.
[0007] To achieve the above object, the present invention provides the following technical solution: a method for fusing infrared and visible light images for power transmission and transformation equipment inspection, comprising the following steps:
[0008] S1: Obtain infrared images and visible light images of power equipment, construct an image dataset of power equipment, and divide it into a training set and a test set;
[0009] S2: Construct the infrared and visible light image fusion network model EIVFusion, which consists of three stages;
[0010] The input of the first stage is a visible light image and an infrared image. The visible light image is processed by a shared encoder to extract visible light image features, and then the visible light image features are divided into two branches. One visible light image feature is processed by a low-frequency global feature encoder to extract low-frequency global features of the visible light image, and the other visible light image feature is processed by a high-frequency local feature encoder to extract high-frequency local features of the visible light image. The high-frequency local features of the visible light image are processed by the enhanced texture and contrast module to obtain enhanced high-frequency local features of the visible light image. The low-frequency global features of the visible light image and the enhanced high-frequency local features of the visible light image are passed through a shared decoder to obtain a reconstructed visible light image; the infrared image is processed by the infrared image preprocessing module and then the infrared image features are extracted by a shared encoder. Then the infrared image features are divided into two branches. One infrared image feature is processed by a low-frequency global feature encoder to extract low-frequency global features of the infrared image, and the other infrared image feature is processed by a high-frequency local feature encoder to extract high-frequency local features of the infrared image. The high-frequency local features of the infrared image are processed by the enhanced texture and contrast module to obtain enhanced high-frequency local features of the infrared image. The low-frequency global features of the infrared image and the enhanced high-frequency local features of the infrared image are passed through a shared decoder to obtain a reconstructed infrared image;
[0011] The input of the second stage is the reconstructed visible light image and the reconstructed infrared image. The reconstructed visible light image is passed through a shared encoder to extract the features of the reconstructed visible light image. Then the features of the reconstructed visible light image are divided into two branches. One branch is used to extract the low-frequency global features of the reconstructed visible light image through a low-frequency global feature encoder, and the other branch is used to extract the high-frequency local features of the reconstructed visible light image through a high-frequency local feature encoder. The reconstructed infrared image is passed through a shared encoder to extract the features of the reconstructed infrared image. Then the features of the reconstructed infrared image are divided into two branches. One branch is used to extract the low-frequency global features of the reconstructed infrared image through a low-frequency global feature encoder, and the other branch is used to extract the high-frequency local features of the reconstructed visible light image. The image features are extracted through a high-frequency local feature encoder to reconstruct the high-frequency local features of the infrared image; the low-frequency global features of the reconstructed visible light image and the low-frequency global features of the reconstructed infrared image are input into the low-frequency global feature fusion layer to obtain low-frequency global fusion features, the high-frequency local features of the reconstructed visible light image are processed through the texture enhancement and contrast module to obtain the enhanced high-frequency local features of the reconstructed visible light image, the high-frequency local features of the reconstructed infrared image are processed through the texture enhancement and contrast module to obtain the enhanced high-frequency local features of the reconstructed infrared image, the enhanced high-frequency local features of the reconstructed visible light image and the enhanced high-frequency local features of the reconstructed infrared image are input into the high-frequency local feature fusion layer to obtain high-frequency local fusion features;
[0012] The input of the third stage is the low-frequency global fusion feature and the high-frequency local fusion feature. The low-frequency global fusion feature and the high-frequency local fusion feature are cascaded in the channel dimension to obtain a mixed feature. The mixed feature is processed by the adaptive brightness adjustment module to obtain an enhanced mixed feature. The enhanced mixed feature is passed through the shared decoder to generate the final fused image.
[0013] S3: Using the infrared images and visible light images of the power equipment in the training set to train the infrared and visible light image fusion network model EIVFusion, a trained infrared and visible light image fusion network model EIVFusion is obtained;
[0014] S4: Input the infrared image and visible light image of the power equipment in the test set into the trained infrared and visible light image fusion network model EIVFusion to obtain a fused image.
[0015] Further preferably, the processing process of the infrared image preprocessing module is: firstly using limited contrast adaptive histogram equalization (CLAHE) to enhance the contrast of the infrared image, and then performing denoising on the infrared image.
[0016] Further preferably, the shared encoder adopts an STCA module, and the shared decoder has the opposite structure to the shared encoder. The processing process of the STCA module is: the input image first extracts local features through a convolution layer, and then performs nonlinear transformation through an activation function ReLU, uses a maximum pooling layer for spatial downsampling, and then uses a convolution kernel to further extract features, then uses a layer normalization operation to standardize the features, captures global information through a local window self-attention mechanism, and then uses global average pooling to aggregate global information of the features in the spatial dimension to generate a global description of each channel, uses a fully connected layer to process the global description, generates the importance weight of each channel through learning, uses an activation function ReLU for nonlinear transformation, outputs channel-level weight distribution, and finally uses a cross-modal attention mechanism to realize the interaction between the features of the infrared image and the features of the visible light image. By using the features of the infrared image as the query Q and the features of the visible light image as the key K and value V, the features of the infrared image are dynamically adjusted according to the feature information of the visible light image when calculating the attention, and enhanced infrared features are obtained in the output.
[0017] Further preferably, the low-frequency global feature encoder and the low-frequency global feature fusion layer both adopt the LTCS module, and the processing process of the LTCS module is: the input features are extracted with local features through the convolution layer, the extracted convolution features are flattened and input into the fully connected layer to realize the transformation of the feature dimension, and then the long-range dependency between different positions in the input features is captured through the multi-head self-attention mechanism, the features at each position are nonlinearly transformed, and then the gradient flow is improved through the Swish activation function, the feature dimension is adjusted through linear transformation, the nonlinear representation is improved through the ReLU activation function, and then the features are layer-normalized, the layer-normalized features are residually connected with the features processed by the multi-head self-attention mechanism, and finally the low-frequency global features are extracted through the joint attention mechanism of channels and spaces.
[0018] Further preferably, the high-frequency local feature encoder and the high-frequency local feature fusion layer both adopt an invertible neural network INN module.
[0019] Further preferably, the SLfusion operator is used in the texture and contrast enhancement module to enhance the strong texture features and weak texture features of the high-frequency local features, generate the high-frequency local features after texture enhancement, and then the contrast of the high-frequency local features is improved by the contrast enhancement module composed of the convolution layer. The expression of the SLfusion operator is:
[0020] ,
[0021] Where, Sobel Edge Strength represents the edge strength calculated using the Sobel operator. Represents high-frequency local features The absolute value of the convolution result with the Laplacian operator, α and β are weight parameters.
[0022] Further preferably, the processing process of the adaptive brightness adjustment module is:
[0023] The local mean of the mixed feature is calculated by sliding the window. The calculation expression is:
[0024] ,
[0025] In the formula, is a pixel value in the mixed feature, and are the horizontal and vertical axes, is the local mean of the mixed features, is the side length of the sliding window, and is the index within the window;
[0026] Calculate the global mean of the mixed features. The calculation expression is:
[0027] ,
[0028] In the formula, Represents the global mean of the mixed features, M and N are the number of rows and columns of the feature map respectively;
[0029] According to the difference between each local mean and the global mean, adaptive brightness adjustment is performed, and the adjustment expression is:
[0030] ,
[0031] In the formula, Represents the pixel value after brightness adjustment. is the adjustment factor.
[0032] Further preferably, the loss function of the first stage includes the reconstruction loss function of the infrared image and the visible light image and the feature decomposition loss function, and the loss function expression of the first stage is:
[0033] ,
[0034] In the formula, is the total loss function of the first stage, is the infrared image reconstruction loss function, is the visible light image reconstruction loss function, is the eigendecomposition loss function, ,in is the initial weight coefficient of the visible light image reconstruction loss, is a constant that controls the weight growth rate of the visible light image reconstruction loss, is the number of training rounds, is the maximum number of training rounds set, and is the tuning parameter;
[0035] The reconstruction loss expression of infrared image and visible light image is:
[0036] ,
[0037] ,
[0038] ,
[0039] ,
[0040] In the formula, is the original image, is the reconstructed image, is the weight coefficient, is the initial weight coefficient, is a constant that controls the speed at which the weight increases. is the current training round, is the first stage mean square error loss function, is a loss function based on the structural similarity index, is the structural similarity index;
[0041] The expression of the feature decomposition loss function is:
[0042] ,
[0043] In the formula, Represents high-frequency local features D The correlation measure between It is calculated by inputting the image and the target image The correlation of Represents low-frequency global features B The correlation measure between It is calculated by inputting the image and the target image The correlation of represents the correlation measurement function, is a positive value used to avoid zero denominator, and are adjustment coefficients, which control the influence of high-frequency local features and low-frequency global features respectively.
[0044] Further preferably, the loss function expression of the second stage is:
[0045] ,
[0046] ,
[0047] ,
[0048] In the formula, Represents the output image The gradient amplitude of Infrared image and visible light images The pixel-by-pixel maximum value of the gradient magnitude is is the total loss function of the second stage, is the second stage mean square error loss function, is the gradient loss function, is the height of the feature map, is the width of the feature map, represents the Sobel gradient operator, is the tuning parameter, is a hyperparameter that controls the influence of the gradient. is a constant used to avoid division by zero errors when the gradient is zero. It is the adjustment coefficient used to adjust the influence of the second-order gradient term.
[0049] Further preferably, the loss function of the third stage includes texture loss, intensity loss, color consistency loss, and feature decomposition loss. The loss function expression of the third stage is:
[0050] ,
[0051] ,
[0052] ,
[0053] ,
[0054] In the formula, is the output image At pixel position The gradient amplitude of and is the visible light image and infrared image at pixel location The gradient amplitude of is the output image At pixel position The brightness value, Indicates traversing all pixels and color channels Calculate the loss. Refers to the angular difference between color vectors, which measures the consistency of color direction. Indicates the pixel value of the visible light image in the color space and output image pixel values The color difference, is the total loss function of the third stage, is the texture loss function, is the mean square error loss function of the third stage, is the color consistency loss function, It represents the channel number of the image. It is the value of the three channels of RGB. is the tuning parameter, is a hyperparameter used to control the weight of the gradient difference, is a hyperparameter used to adjust the perception of image contrast, is the weight of the color channel, is the smoothing coefficient, is the color penalty term.
[0055] Compared with the prior art, the present invention has the following beneficial effects:
[0056] The present invention firstly improves the contrast of infrared images by preprocessing the infrared images, so as to better detect the features in the images, uses a dual-branch shared encoder and a shared decoder based on the STCA module to extract features or generate images, uses a low-frequency global feature encoder and a high-frequency local feature encoder based on the LTCS module and the INN module to extract low-frequency global features and high-frequency local features from shared features, strengthens the fine-grained representation of features and expands the perception field of the network for the high-frequency local features of visible light images and infrared images through the enhanced texture and contrast module, reduces the brightness difference between the local area and the global mean for the mixed features of visible light images and infrared images through the adaptive brightness adjustment module, and realizes the high-performance fusion of infrared images and visible light images. Through the network model, the present method can be well applied to the fusion of infrared and visible light images for inspection of power transmission and transformation equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 is a flow chart of the method of the present invention;
[0058] Figure 2 This is a schematic diagram of the first stage of the infrared and visible light image fusion network model;
[0059] Figure 3 This is a schematic diagram of the second stage of the infrared and visible light image fusion network model;
[0060] Figure 4 This is a schematic diagram of the third stage of the infrared and visible light image fusion network model;
[0061] Figure 5 It is the schematic diagram of STCA module;
[0062] Figure 6 This is a schematic diagram of the LTCS module. DETAILED DESCRIPTION
[0063] The present invention is further described below in conjunction with embodiments. It is necessary to point out that the following embodiments are only used to further illustrate the present invention and cannot be understood as limiting the scope of protection of the present invention. Some non-essential improvements and adjustments made by technical personnel in this field based on the above invention content still fall within the scope of protection of the present invention.
[0064] Reference Figure 1 , a method for fusing infrared and visible light images for power transmission and transformation equipment inspection, comprising the following steps:
[0065] S1: Obtain infrared images and visible light images of power equipment, construct an image dataset of power equipment, and divide it into a training set and a test set;
[0066] Acquisition and processing of image datasets: First, a set of images of power equipment is collected, including infrared images and visible light images of power equipment such as insulators; next, the labelimg tool is used to annotate the visible light images, focusing on the location, shape, size and other features of power equipment such as insulators. After the annotation is completed, the dataset is divided into a training set and a test set.
[0067] S2: Construct the infrared and visible light image fusion network model EIVFusion, which consists of three stages;
[0068] like Figure 2As shown, the input of the first stage is a visible light image with a resolution of 640×480 and an infrared image with a resolution of 640×480. The visible light image is subjected to a shared encoder to extract visible light image features, and then the visible light image features are divided into two branches. One visible light image feature is subjected to a low-frequency global feature encoder to extract the low-frequency global features of the visible light image, and the other visible light image feature is subjected to a high-frequency local feature encoder to extract the high-frequency local features of the visible light image. The high-frequency local features of the visible light image are processed by the enhanced texture and contrast module to obtain the enhanced high-frequency local features of the visible light image. The low-frequency global features of the visible light image and the enhanced high-frequency local features of the visible light image are passed through a shared decoder to obtain the reconstructed visible light image. The infrared image is processed by the infrared image preprocessing module and then subjected to a shared encoder to extract the infrared image features. Then the infrared image features are divided into two branches. One infrared image feature is subjected to a low-frequency global feature encoder to extract the low-frequency global features of the infrared image, and the other infrared image feature is subjected to a high-frequency local feature encoder to extract the high-frequency local features of the infrared image. The high-frequency local features of the infrared image are processed by the enhanced texture and contrast module to obtain the enhanced high-frequency local features of the infrared image. The low-frequency global features of the infrared image and the enhanced high-frequency local features of the infrared image are passed through a shared decoder to obtain the reconstructed infrared image.
[0069] like Figure 3 As shown, the input of the second stage is the reconstructed visible light image and the reconstructed infrared image. The reconstructed visible light image is subjected to a shared encoder to extract the features of the reconstructed visible light image, and then the features of the reconstructed visible light image are divided into two branches. One branch is used to extract the low-frequency global features of the reconstructed visible light image through a low-frequency global feature encoder, and the other branch is used to extract the high-frequency local features of the reconstructed visible light image through a high-frequency local feature encoder; the reconstructed infrared image is subjected to a shared encoder to extract the features of the reconstructed infrared image, and then the features of the reconstructed infrared image are divided into two branches. One branch is used to extract the low-frequency global features of the reconstructed infrared image through a low-frequency global feature encoder, and the other branch is used to extract the high-frequency local features of the reconstructed visible light image. The infrared image features are extracted through a high-frequency local feature encoder to reconstruct the high-frequency local features of the infrared image; the low-frequency global features of the reconstructed visible light image and the low-frequency global features of the reconstructed infrared image are input into the low-frequency global feature fusion layer to obtain low-frequency global fusion features, the high-frequency local features of the reconstructed visible light image are processed through the texture enhancement and contrast module to obtain the enhanced high-frequency local features of the reconstructed visible light image, the high-frequency local features of the reconstructed infrared image are processed through the texture enhancement and contrast module to obtain the enhanced high-frequency local features of the reconstructed infrared image, the enhanced high-frequency local features of the reconstructed visible light image and the enhanced high-frequency local features of the reconstructed infrared image are input into the high-frequency local feature fusion layer to obtain high-frequency local fusion features;
[0070] like Figure 4As shown in FIG. 1 , the input of the third stage is the low-frequency global fusion feature and the high-frequency local fusion feature. The low-frequency global fusion feature and the high-frequency local fusion feature are cascaded in the channel dimension to obtain a mixed feature. The mixed feature is processed by the adaptive brightness adjustment module to obtain an enhanced mixed feature. The enhanced mixed feature is processed by the shared decoder to generate the final fused image.
[0071] The processing process of the infrared image preprocessing module is as follows: first, the contrast of the infrared image is enhanced using limited contrast adaptive histogram equalization (CLAHE), and then the infrared image is denoised to further enhance the features of the infrared image and generate an infrared image with high contrast and less noise.
[0072] The present invention proposes that the STCA module is used for sharing encoders, such as Figure 5 As shown in the figure, the processing process of the STCA module is as follows: the input image first extracts local features through the convolution layer, and then performs nonlinear transformation through the activation function ReLU to enhance the nonlinear expression ability of the network. A 2×2 maximum pooling layer is used for spatial downsampling to reduce the computational complexity, and then a larger convolution kernel is used to further extract features. The extracted main features provide strong local detail support for the entire module, and then the layer normalization operation is used to standardize the features to reduce the problem of gradient disappearance or explosion. The local window self-attention mechanism is used to capture the global information of the image at a relatively small computational cost, so that the module has a good long-distance dependency modeling capability, making it particularly suitable for processing complex visual tasks, especially in complex power transmission and transformation equipment inspection environments when local and global information need to be modeled simultaneously. Then, global average pooling is used to aggregate the global information of the input features in the spatial dimension to generate a global description of each channel. The global description is processed using a fully connected layer. The importance weight of each channel is generated through learning, and the activation function ReLU is used for nonlinear transformation. The channel-level weight distribution is output, so that the module can assign a weight to each channel, thereby improving the network's attention to important features and further improving the model performance. Finally, the cross-modal attention mechanism is used to realize the interaction between the features of infrared images and visible light images. By taking the features of infrared images as queries Q and the features of visible light images as keys K and values V, the features of infrared images are dynamically adjusted according to the feature information of visible light images when calculating attention, so as to achieve the purpose of feature fusion, and then the enhanced infrared features are obtained in the output, thereby improving the performance of the overall algorithm.
[0073] In the shared decoder, the decomposed features are cascaded in the channel dimension as input, and the reconstructed original image (first stage) or fused image (third stage) is the output of the shared decoder. Since the input here involves cross-modal and multi-frequency features, the structure of the shared decoder is kept consistent with the design of the shared encoder, that is, the STCA module is used as the basic unit of the decoder. The structure of the shared decoder is opposite to that of the shared encoder. Through the shared decoder, the features fused in the first stage or the features processed by the adaptive brightness adjustment module are regenerated by the shared decoder as a reconstructed image or a fused image.
[0074] The present invention proposes that the LTCS module is used for a low-frequency global feature encoder, such as Figure 6 As shown, the processing process of the LTCS module is as follows: the input feature extracts local features through a 3×3 convolution layer, the extracted convolution features are flattened and input into the fully connected layer to achieve the transformation of feature dimensions, and then the long-range dependencies between different positions in the input features are captured through the multi-head self-attention mechanism, and the features of each position are nonlinearly transformed to increase the expression ability of the model, and then the gradient flow is improved through the Swish activation function, the feature dimension is adjusted through linear transformation, and the nonlinear representation is improved through the ReLU activation function, and then the features are layer-normalized, and the layer-normalized features are residually connected with the features processed by the multi-head self-attention mechanism to help the model retain the effective information of the initial features. Finally, through the joint attention mechanism of channels and spaces, the model is helped to focus on important feature areas, thereby improving the overall performance, so that it can efficiently extract low-frequency global features from shared features while achieving faster reasoning speed. The present invention uses a high-frequency local feature encoder based on an invertible neural network (INN module) to extract high-frequency local feature information from shared features.
[0075] Considering that the inductive deviation of the fusion of low-frequency global features and high-frequency local features is similar to the extraction of low-frequency global features and high-frequency local features in the encoder, the present invention uses the LTCS module and the INN module for the low-frequency global feature fusion layer and the high-frequency local feature fusion layer respectively.
[0076] In the texture and contrast enhancement module, the present invention proposes an SLfusion operator to enhance the strong texture features and weak texture features of high-frequency local features, thereby generating high-frequency local features after texture enhancement, and then expanding the perception field of the network through a contrast enhancement module composed of a convolutional layer to improve the contrast of high-frequency local features. The expression of the SLfusion operator is:
[0077] ,
[0078] Wherein, Sobel Edge Strength represents the edge strength calculated using the Sobel operator. The Sobel operator is a common edge detection method that can detect gradient changes in images. Represents high-frequency local features The absolute value of the convolution result with the Laplacian operator, which is a second-order derivative-based edge detection method used to detect edges and texture changes in images. α and β are weight parameters that can be optimized by the algorithm.
[0079] The processing process of the adaptive brightness adjustment module is:
[0080] The local mean of the mixed feature is calculated by sliding the window. The local mean is used to determine the brightness of a certain area. The calculation expression is:
[0081] ,
[0082] In the formula, is a pixel value in the mixed feature, and are the horizontal and vertical axes, is the local mean of the mixed features, is the side length of the sliding window, and is the index within the window;
[0083] Calculate the global mean of the mixed features. The global mean is used as the benchmark for subsequent brightness adjustment. The calculation expression is:
[0084] ,
[0085] In the formula, Represents the global mean of the mixed features, M and N are the number of rows and columns of the feature map respectively;
[0086] According to the difference between each local mean and the global mean, adaptive brightness adjustment is performed. If the brightness of the local area is lower than the global mean, the brightness is enhanced; if the brightness of the local area is too strong, the brightness is reduced. The brightness adjustment expression is:
[0087] ,
[0088] In the formula, Represents the pixel value after brightness adjustment. is an adjustment factor that controls the influence of the difference between the local mean and the global mean.
[0089] The loss function of EIVFusion (infrared and visible light image fusion network model):
[0090] The loss function of the first stage includes the reconstruction loss function of the infrared image and the visible light image and the feature decomposition loss function. The loss function expression of the first stage is:
[0091] ,
[0092] In the formula, is the total loss function of the first stage, is the infrared image reconstruction loss function, is the visible light image reconstruction loss function, is the eigendecomposition loss function, ,in is the initial weight coefficient of the visible light image reconstruction loss, is a constant that controls the weight growth rate of the visible light image reconstruction loss, is the number of training rounds, is the maximum number of training rounds set, and is the tuning parameter;
[0093] The reconstruction loss expression of infrared image and visible light image is:
[0094] ,
[0095] ,
[0096] ,
[0097] ,
[0098] In the formula, is the original image, is the reconstructed image, is the weight coefficient used to adjust the structural similarity loss As a percentage of the total loss, is the initial weight coefficient, is a constant that controls the speed at which the weight increases. is the current training round; is the first stage mean square error loss function, is a loss function based on the structural similarity index, is the structural similarity index;
[0099] The expression of the feature decomposition loss function is:
[0100] ,
[0101] In the formula, Represents high-frequency local features D The correlation measure between the input images is calculated by and the target image The correlation is obtained. Represents low-frequency global features B The correlation measure between the input images is calculated by and the target image The correlation is obtained. Represents the correlation measurement function (CorrelationCoefficient). is a small positive value used to avoid zero denominator (often used for numerical stability), and are adjustment coefficients, which control the influence of high-frequency local features and low-frequency global features respectively.
[0102] The loss function expression of the second stage is:
[0103] ,
[0104] ,
[0105] ,
[0106] In the formula, Represents the output image The gradient amplitude of Infrared image and visible light images The pixel-by-pixel maximum value of the gradient magnitude. is the total loss function of the second stage, is the second stage mean square error loss function, is the gradient loss function, is the height of the feature map, is the width of the feature map, represents the Sobel gradient operator, is the tuning parameter, It is a hyperparameter that controls the influence of the gradient and can be tuned through experiments; is a small constant used to avoid division by zero when the gradient is zero and ensure numerical stability; It is the adjustment coefficient used to adjust the influence of the second-order gradient term.
[0107] The loss function of the third stage includes texture loss, intensity loss, color consistency loss, and feature decomposition loss. The loss function expression of the third stage is:
[0108] ,
[0109] ,
[0110] ,
[0111] ,
[0112] In the formula, is the output image At pixel position The gradient amplitude of and is the visible light image and infrared image at pixel location The gradient amplitude of is the output image At pixel position The brightness value, Indicates traversing all pixels and color channels Calculate the loss. Refers to the angular difference between color vectors, which measures the consistency of color direction. Indicates the pixel value of the visible light image in the color space and output image pixel values of color differences. is the total loss function of the third stage, is the texture loss function, is the mean square error loss function of the third stage, is the color consistency loss function, It represents the channel number of the image. It is the value of the three channels of RGB. is the tuning parameter, It is a hyperparameter used to control the weight of the gradient difference. The gradient difference reflects the similarity between the restored image and the input image, which is particularly important in texture restoration. , can strengthen or weaken the contribution of gradient difference to loss; It is a hyperparameter used to adjust the perception of image contrast. Through contrast perception, the model can better restore the detailed parts of the image, such as high-contrast areas, and increase It allows high-contrast areas to have a greater impact on the loss, thereby improving the restoration quality of these areas; Cosine similarity: The difference between color channels is measured by cosine similarity, which is more stable and robust than directly calculating the angle difference; is the weight of the color channel, is the smoothing coefficient, which can be adjusted based on performance or prior knowledge during training and , enhance the training effect; is the color penalty term, introducing the color difference penalty term , which helps reduce the shift in color differences during the restoration process.
[0113] S3: Using the infrared images and visible light images of the power equipment in the training set to train the infrared and visible light image fusion network model EIVFusion, a trained infrared and visible light image fusion network model EIVFusion is obtained;
[0114] S4: Input the infrared image and visible light image of the power equipment in the test set into the trained infrared and visible light image fusion network model EIVFusion to obtain a fused image.
[0115] The present invention firstly improves the contrast of infrared images by preprocessing the infrared images, so as to better detect the features in the images, uses a dual-branch shared encoder and a shared decoder based on the STCA module to extract features or generate images, uses a low-frequency global feature encoder and a high-frequency local feature encoder based on the LTCS module and the INN module to extract low-frequency global features and high-frequency local features from shared features, strengthens the fine-grained representation of features and expands the perception field of the network for the high-frequency local features of visible light images and infrared images through the enhanced texture and contrast module, reduces the brightness difference between the local area and the global mean for the mixed features of visible light images and infrared images through the adaptive brightness adjustment module, and realizes the high-performance fusion of infrared images and visible light images. Through the network model, the present method can be well applied to the fusion of infrared and visible light images for inspection of power transmission and transformation equipment.
[0116] The present invention is a method based on deep learning, which can automatically learn and extract the feature information and fusion rules of the image, significantly improving the effect and efficiency of image fusion. At the same time, a dual-branch Transformer-CNN framework is used to extract global and local features, better reflecting the unique modality-specific and modality-shared features, thereby effectively fusing image information of different modalities, retaining more details of the source image, and enhancing the contrast and clarity of the image. It has high adaptability and flexibility, can adapt to different lighting conditions and background environments, and provides high-quality fused images for various electrical equipment application scenarios. The fusion method based on deep learning of the present invention realizes end-to-end processing from image input to fusion result output, reduces the workload of manual intervention and post-processing, and improves the automation and intelligence level of processing.
[0117] The above only expresses the preferred implementation of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosure to modify or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A method for fusing infrared and visible light images for power transmission and transformation equipment inspection, characterized in that: The following steps are involved: S1: Obtain infrared images and visible light images of power equipment, construct an image dataset of power equipment, and divide it into a training set and a test set; S2: Construct the infrared and visible light image fusion network model EIVFusion, which consists of three stages; The input of the first stage is a visible light image and an infrared image. The visible light image is processed by a shared encoder to extract visible light image features, and then the visible light image features are divided into two branches. One visible light image feature is processed by a low-frequency global feature encoder to extract low-frequency global features of the visible light image, and the other visible light image feature is processed by a high-frequency local feature encoder to extract high-frequency local features of the visible light image. The high-frequency local features of the visible light image are processed by the enhanced texture and contrast module to obtain enhanced high-frequency local features of the visible light image. The low-frequency global features of the visible light image and the enhanced high-frequency local features of the visible light image are passed through a shared decoder to obtain a reconstructed visible light image; the infrared image is processed by the infrared image preprocessing module and then the infrared image features are extracted by a shared encoder. Then the infrared image features are divided into two branches. One infrared image feature is processed by a low-frequency global feature encoder to extract low-frequency global features of the infrared image, and the other infrared image feature is processed by a high-frequency local feature encoder to extract high-frequency local features of the infrared image. The high-frequency local features of the infrared image are processed by the enhanced texture and contrast module to obtain enhanced high-frequency local features of the infrared image. The low-frequency global features of the infrared image and the enhanced high-frequency local features of the infrared image are passed through a shared decoder to obtain a reconstructed infrared image; The input of the second stage is the reconstructed visible light image and the reconstructed infrared image. The reconstructed visible light image is passed through a shared encoder to extract the features of the reconstructed visible light image. Then the features of the reconstructed visible light image are divided into two branches. One branch is used to extract the low-frequency global features of the reconstructed visible light image through a low-frequency global feature encoder, and the other branch is used to extract the high-frequency local features of the reconstructed visible light image through a high-frequency local feature encoder. The reconstructed infrared image is passed through a shared encoder to extract the features of the reconstructed infrared image. Then the features of the reconstructed infrared image are divided into two branches. One branch is used to extract the low-frequency global features of the reconstructed infrared image through a low-frequency global feature encoder, and the other branch is used to extract the high-frequency local features of the reconstructed visible light image. The image features are extracted through a high-frequency local feature encoder to reconstruct the high-frequency local features of the infrared image; the low-frequency global features of the reconstructed visible light image and the low-frequency global features of the reconstructed infrared image are input into the low-frequency global feature fusion layer to obtain low-frequency global fusion features, the high-frequency local features of the reconstructed visible light image are processed through the texture enhancement and contrast module to obtain the enhanced high-frequency local features of the reconstructed visible light image, the high-frequency local features of the reconstructed infrared image are processed through the texture enhancement and contrast module to obtain the enhanced high-frequency local features of the reconstructed infrared image, the enhanced high-frequency local features of the reconstructed visible light image and the enhanced high-frequency local features of the reconstructed infrared image are input into the high-frequency local feature fusion layer to obtain high-frequency local fusion features; The input of the third stage is the low-frequency global fusion feature and the high-frequency local fusion feature. The low-frequency global fusion feature and the high-frequency local fusion feature are cascaded in the channel dimension to obtain a mixed feature. The mixed feature is processed by the adaptive brightness adjustment module to obtain an enhanced mixed feature. The enhanced mixed feature is passed through the shared decoder to generate the final fused image. S3: Using the infrared images and visible light images of the power equipment in the training set to train the infrared and visible light image fusion network model EIVFusion, a trained infrared and visible light image fusion network model EIVFusion is obtained; S4: Input the infrared image and visible light image of the power equipment in the test set into the trained infrared and visible light image fusion network model EIVFusion to obtain a fused image; The shared encoder adopts the STCA module, and the shared decoder has the opposite structure to the shared encoder. The processing process of the STCA module is as follows: the input image first extracts local features through the convolution layer, and then performs nonlinear transformation through the activation function ReLU, uses the maximum pooling layer for spatial downsampling, and then uses the convolution kernel to further extract features, and then uses the layer normalization operation to standardize the features, and captures global information through the local window self-attention mechanism, and then uses the global average pooling to aggregate the global information of the features in the spatial dimension to generate a global description of each channel, and uses the fully connected layer to process the global description. The importance weight of each channel is generated through learning, and the activation function ReLU is used for nonlinear transformation, and the channel-level weight distribution is output. Finally, the cross-modal attention mechanism is used to realize the interaction between the features of the infrared image and the features of the visible light image. By taking the features of the infrared image as the query Q and the features of the visible light image as the key K and value V, the features of the infrared image are dynamically adjusted according to the feature information of the visible light image when calculating the attention, and the enhanced infrared features are obtained in the output; The SLfusion operator is used in the texture and contrast enhancement module to enhance the strong texture features and weak texture features of the high-frequency local features, generate the high-frequency local features after texture enhancement, and then improve the contrast of the high-frequency local features through the contrast enhancement module composed of the convolution layer. The expression of the SLfusion operator is: , Where, Sobel Edge Strength represents the edge strength calculated using the Sobel operator. Represents high-frequency local features The absolute value of the convolution result with the Laplace operator, α and β are weight parameters; The processing process of the adaptive brightness adjustment module is as follows: The local mean of the mixed feature is calculated by sliding the window. The calculation expression is: , In the formula, is a pixel value in the mixed feature, and are the horizontal and vertical axes, is the local mean of the mixed features, is the side length of the sliding window, and is the index within the window; Calculate the global mean of the mixed features. The calculation expression is: , In the formula, Represents the global mean of the mixed features, M and N are the number of rows and columns of the feature map respectively; According to the difference between each local mean and the global mean, adaptive brightness adjustment is performed, and the adjustment expression is: , In the formula, Represents the pixel value after brightness adjustment. is the adjustment factor.
2. The method for fusion of infrared and visible light images for power transmission and transformation equipment inspection according to claim 1 is characterized in that: The processing process of the infrared image preprocessing module is: firstly, the contrast of the infrared image is enhanced by using the limited contrast adaptive histogram equalization (CLAHE), and then the infrared image is denoised.
3. The method for fusion of infrared and visible light images for power transmission and transformation equipment inspection according to claim 1 is characterized in that: The low-frequency global feature encoder and the low-frequency global feature fusion layer both adopt the LTCS module. The processing process of the LTCS module is as follows: the input feature is subjected to the convolution layer to extract local features, the extracted convolution features are flattened and input into the fully connected layer to achieve the transformation of the feature dimension, and then the long-range dependency between different positions in the input feature is captured through the multi-head self-attention mechanism, the feature of each position is nonlinearly transformed, and then the gradient flow is improved through the Swish activation function, the feature dimension is adjusted through linear transformation, the nonlinear representation is improved through the ReLU activation function, and then the features are layer-normalized, the layer-normalized features are residually connected with the features processed by the multi-head self-attention mechanism, and finally the low-frequency global features are extracted through the joint attention mechanism of channels and spaces.
4. The method for fusion of infrared and visible light images for power transmission and transformation equipment inspection according to claim 1 is characterized in that: The high-frequency local feature encoder and the high-frequency local feature fusion layer both adopt an invertible neural network INN module.
5. The method for fusion of infrared and visible light images for power transmission and transformation equipment inspection according to claim 1 is characterized in that: The loss function of the first stage includes the reconstruction loss function of the infrared image and the visible light image and the feature decomposition loss function. The loss function expression of the first stage is: , In the formula, is the total loss function of the first stage, is the infrared image reconstruction loss function, is the visible light image reconstruction loss function, is the eigendecomposition loss function, ,in is the initial weight coefficient of the visible light image reconstruction loss, is a constant that controls the weight growth rate of the visible light image reconstruction loss, is the number of training rounds, is the maximum number of training rounds set, and is the tuning parameter; The reconstruction loss expression of infrared image and visible light image is: , , , , In the formula, is the original image, is the reconstructed image, is the weight coefficient, is the initial weight coefficient, is a constant that controls the speed at which the weight increases. is the current training round, is the first stage mean square error loss function, is a loss function based on the structural similarity index, is the structural similarity index; The expression of the feature decomposition loss function is: , In the formula, Represents high-frequency local features D The correlation measure between It is calculated by inputting the image and the target image The correlation of Represents low-frequency global features B The correlation measure between It is calculated by inputting the image and the target image The correlation of represents the correlation measurement function, is a positive value used to avoid zero denominator, and are adjustment coefficients, which control the influence of high-frequency local features and low-frequency global features respectively.
6. The method for fusion of infrared and visible light images for power transmission and transformation equipment inspection according to claim 5 is characterized in that: The loss function expression of the second stage is: , , , In the formula, Represents the output image The gradient amplitude of Infrared image and visible light images The pixel-by-pixel maximum value of the gradient magnitude is is the total loss function of the second stage, is the second stage mean square error loss function, is the gradient loss function, is the height of the feature map, is the width of the feature map, represents the Sobel gradient operator, is the tuning parameter, is a hyperparameter that controls the influence of the gradient. is a constant used to avoid division by zero errors when the gradient is zero. It is the adjustment coefficient used to adjust the influence of the second-order gradient term.
7. The method for fusion of infrared and visible light images for power transmission and transformation equipment inspection according to claim 6 is characterized in that: The loss function of the third stage includes texture loss, intensity loss, color consistency loss, and feature decomposition loss. The loss function expression of the third stage is: , , , , In the formula, is the output image At pixel position The gradient amplitude of and is the visible light image and infrared image at pixel location The gradient amplitude of is the output image At pixel position The brightness value, Indicates traversing all pixels and color channels Calculate the loss. Refers to the angular difference between color vectors, which measures the consistency of color direction. Indicates the pixel value of the visible light image in the color space and output image pixel values The color difference, is the total loss function of the third stage, is the texture loss function, is the mean square error loss function of the third stage, is the color consistency loss function, It represents the channel number of the image. It is the value of the three channels of RGB. is the tuning parameter, is a hyperparameter used to control the weight of the gradient difference, is a hyperparameter used to adjust the perception of image contrast, is the weight of the color channel, is the smoothing coefficient, is the color penalty term.
Citation Information
Patent Citations
Feature decomposition-based infrared image and visible light image fusion method
CN118134780A
Cited By
Visible light and infrared image difference perception dynamic cross fusion method and system
CN121213371A
Dynamic cross-fusion method and system for visible light and infrared image difference perception
CN121213371B