Power transmission line patrol infrared and visible light image fusion method based on deep learning

By combining the AE network and the generative adversarial network, introducing an image reconstruction step and a dual discriminator structure, and designing a composite loss function, the problem of poor fusion effect in infrared and visible light image fusion is solved, and high-quality image fusion effect is achieved.

CN120876253APending Publication Date: 2025-10-31JIAXING HENGCHUANG ELECTRIC EQUIP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510971815.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing infrared and visible light image fusion methods have poor fusion results when the differences are not obvious, insufficient preservation of detail and texture information under low light conditions, and poor quality of fused images.

Method used

By combining the AE network and the generative adversarial network, and through image reconstruction steps and a dual discriminator structure, a composite loss function is designed to improve the clarity and fidelity of the fused image.

Benefits of technology

It significantly improves the clarity and quality of fused images, enhances their interpretability and fidelity, and is applicable to more real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876253A_ABST
    Figure CN120876253A_ABST
Patent Text Reader

Abstract

The invention discloses a deep learning-based infrared and visible light image fusion method for power transmission line inspection, and relates to the technical field of image processing. The method comprises the following steps: S1, constructing an image fusion network: fusing an AE network structure into a generative network, and enabling a generator to realize fusion of an infrared image and a visible light image through a training model; s2, designing loss functions: respectively designing loss functions corresponding to a generator and a discriminator for the generative adversarial network; s3, constructing a data set: constructing a data set containing an infrared and visible light image pair to train and verify the proposed method, and zooming the image to ensure that the input image meets the model requirement; s4, model training and verification: training the image fusion network obtained in the step S1, and in the training process, optimizing network parameters by using the loss function designed in the step S2; after training is completed, the performance of the model is verified on a test set, and the fusion effect of the model is evaluated by comparing images before and after fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for fusing infrared and visible light images of power transmission line inspection based on deep learning. Background Technology

[0002] The primary goal of infrared and visible light image fusion is to integrate complementary information from source images and generate a high-contrast fused image that highlights prominent targets while containing rich texture details. This includes traditional image fusion frameworks, AE-based image fusion frameworks, convolutional neural network (CNN)-based frameworks, and generative adversarial network (GAN)-based frameworks. In traditional image fusion frameworks, the fusion task is typically completed in the spatial or transform domain. This method achieves image fusion by manually designing fusion rules through activity level measurements. GTF defines infrared and visible light image fusion as preserving overall intensity and texture structure in the spatial domain, generating the fused image by optimizing the objective function.

[0003] Ma et al. (Ma J, Zhou Z, Wang B, et al. Infrared and visible image fusion based on visual saliency map and weighted least square optimization[J]. Infrared Physics & Technology, 2017, 82: 8-17.) converted the source image into three independent conceptual components: signal intensity, signal structure, and average intensity, and then merged these three independent components separately to achieve multi-exposure image fusion. In the AE-based image fusion framework, this method uses an autoencoder pre-trained on a large-scale dataset as a feature extractor and image reconstructor, and then designs a special fusion strategy for deep features to achieve image fusion. Li et al. (Li H, Wu XJ. DenseFuse: A fusion approach to infrared and visible images[J]. IEEE Transactions on Image Processing, 2018, 28(5): 2614-2623.) first proposed using an encoding network to extract image features and then using a decoding network to obtain the fused image. This method can use the results of each layer in the encoding network to construct a feature map. Finally, the fused image is reconstructed using a fusion strategy and a decoding network containing four CNN layers. Jian et al. (Jian L, Yang X, Liu Z, et al. SEDRFuse: A symmetric encoder-decoder with residual block network for infrared and visible image fusion[J]. IEEE Transactions on Instrumentation and Measurement, 2020, 70: 1-15.) proposed a symmetric encoder-decoder with a residual block network and an attention map-based feature fusion method. In the CNN-based image fusion framework, one approach is to utilize a pre-trained CNN network to analyze image activity level measurements and generate partially weighted images based on handcrafted features.

[0004] Another approach is to use a CNN network image fusion framework to learn the direct mapping between the source image and the fused image (or focus image) in an end-to-end manner. Tang et al. (Tang L, Yuan J, Ma J. Image fusion in the loop of high-level vision tasks: Asemantic-aware real-time infrared and visible image fusion network[J]. Information Fusion, 2022, 82: 28-42.) introduced a semantic concatenated segmentation module and gradient residual dense blocks, which effectively improved the performance of fused images on high-level vision tasks and the ability of the fusion network to describe fine-grained spatial details. PSFusion (TANG Linfeng, ZHANG Hao, Xu Han, et al. Rethinking the necessity of image fusion in high-level vision tasks: A practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity[J]. Information Fusion, 2023, 99, 101870.) injected semantic features into the network and designed a scene fidelity path parallel to the image fusion path to constrain the fusion module to retain the complete information of the source image.

[0005] In the GAN-based image fusion framework, GAN networks can effectively model data distribution even without supervised information, a characteristic that better addresses the problem of infrared and visible light image fusion. Ma et al. (Ma J, Yu W, Liang P, et al. FusionGAN: A generative adversarial network for infrared and visible image fusion[J]. Information fusion, 2019, 48: 11-26.) proposed FusionGAN, which was the first network to apply GAN networks to the image fusion problem. In this network, the generator is responsible for processing the source image, fully integrating the latent distributions belonging to the source image into the fused image, while the discriminator is responsible for judging the input image, enhancing the distribution characteristics of the fused image generated by the generator; DDcGAN (Ma J,Xu H,Jiang J,et al.DDcGAN:Adual-discriminator conditional generative adversarial network for multi-resolutionimage fusion[J].IEEE Transactions on Image Processing,2020,29:4980-4995.) uses a densely connected convolutional network instead of U-net, strengthening the transmission of feature maps, utilizing feature maps more effectively, and improving the fusion effect.

[0006] Image fusion frameworks based on image processing (AE) typically employ an encoder to extract features, followed by simple weighted averaging or concatenation for fusion. This lacks dynamic adaptability to the complementarity and conflict of features across different modalities, leading to the loss of crucial information. Furthermore, the AE decoder may fail to effectively recover the fused feature details, especially in complex scenes where it is prone to blurring or distortion, making it difficult to retain high-frequency information.

[0007] Image fusion frameworks based on CNNs extract features through local convolutional operations, which may ignore the overall contextual information of the image, resulting in a lack of global consistency in the fusion results. Furthermore, to capture multi-scale features of different modalities, CNN network fusion frameworks require complex network structures, leading to a dramatic increase in the number of parameters and making them prone to overfitting or computational inefficiency.

[0008] Image fusion frameworks based on GANs require dynamic balancing of adversarial training between the generator and discriminator. However, the differences between visible light and infrared features can lead to gradient vanishing or mode collapse, causing the fusion result to be biased towards a single modality. Furthermore, adversarial loss may force the generator to overemphasize local regions that the discriminator is sensitive to, resulting in unnatural artifacts or noise in the fused image, especially in low signal-to-noise ratio scenarios.

[0009] To address the aforementioned problems, this invention proposes a deep learning-based method for fusing infrared and visible light images during power transmission line inspections. Unlike existing methods that directly fuse image data using only feature maps, this invention introduces an image reconstruction step. The original images are reconstructed from the infrared and visible light images separately using an AE network, and then compared with their corresponding feature maps. Figure 1 The same input generator is used to enhance the interpretability and fidelity of the fused image. Summary of the Invention

[0010] The purpose of this invention is to propose a deep learning-based method for fusing infrared and visible light images of power transmission line inspections. This addresses the problems of poor fusion performance of the ICAF network when the differences between visible and infrared images are not significant, insufficient preservation of detail and texture information when visible light intensity is weak, and poor quality of some fused images. This invention effectively improves the clarity of the fused image and enhances the quality of the generated fused image. This invention combines an AE network with a generative adversarial network, using feature maps and generation... Figure 1 Image fusion is performed to obtain an image with clearer texture information.

[0011] To achieve the above objectives, the present invention adopts the following technical solution:

[0012] A deep learning-based method for fusing infrared and visible light images of power transmission line inspections includes the following steps:

[0013] S1. Construct an image fusion network: Integrate the AE network structure into the generator network, extract image features through the AE network and reconstruct the original infrared and visible light images, and then input the extracted feature maps and the reconstructed images into the generator to generate a fused image, thereby improving the fusion effect and image clarity.

[0014] S2. Design Loss Functions: For generative adversarial networks, design loss functions for the generator and discriminator respectively; the generator's loss function is a composite loss function, including content loss, gradient loss, structural similarity loss, and a constraint on the reconstructed image's similarity to the original. Figure 1 Consistent reconstruction loss is used to ensure the quality of the fused image in terms of sharpness, structural fidelity, and detail restoration; the discriminator loss is used to identify the differences between the original image and the fused image to further improve the generator's fusion capability;

[0015] S3. Constructing the dataset: Construct a dataset containing pairs of infrared and visible light images to train and validate the proposed method. Use large power equipment data for model training and testing, and scale the images to ensure that the input images meet the model requirements.

[0016] S4. Model Training and Validation: The image fusion network obtained in S1 is trained. During the training process, the network parameters are optimized using the loss function designed in S2. After training, the model's performance is validated on the test set. The fusion effect of the model is evaluated by comparing the images before and after fusion.

[0017] Preferably, S1 specifically includes the following:

[0018] S1.1 Input the image into the AE network, extract image features, and reconstruct the original image;

[0019] S1.2. Input the extracted features and the restored image into the generator to generate a fused image. Optimize the network parameters using a dual discriminator structure to prevent the loss of texture and detail features from the visible light image and target features from the infrared image.

[0020] Preferably, the generator consists of an AE network and a fusion network, wherein the encoder of the AE network consists of four convolutional blocks, the formula of which is expressed as:

[0021] output=f(BatchNorm(conv(input)))

[0022] In the formula, input represents the input of the convolutional block, which is an image or feature information; f(·) represents the activation function; the first convolutional block of the encoder uses reflection filling to input visible light images and infrared images into the convolutional layer to extract low-level and high-level features of the image;

[0023] The decoder of the AE network consists of three convolutional blocks, which concatenate low-level features with high-level features and feed them into the convolutional layer. Each convolution input is concatenated with a low-level feature to ensure the richness and completeness of the generated image details, and outputs restored images of visible light and infrared light.

[0024] Restore the image and the features extracted by the encoder Figure 1 The image information is fused in the fusion network of the input generator, and the fused image is finally output by the generator network.

[0025] Preferably, the discriminator extracts features through four convolutional layers, integrates the feature information, and feeds it into the tanh function after passing through a linear layer to achieve the discrimination of the original visible light image, infrared image, and fused image.

[0026] Preferably, the generator loss function is expressed as follows:

[0027] L G =λ adv L adv +λ content L content +λ gradient L gradient +λ str L str

[0028] In the formula, λ adv , λ content , λ gradient , λ str L is the loss coefficient used to adjust the loss balance. adv L is the adversarial loss between the generator and the discriminator. content For content loss, L gradient For gradient loss, L str For structural loss;

[0029] The resistance loss L adv The formula used to ensure that the generator produces images that deceive the discriminator is expressed as follows:

[0030]

[0031] In the formula, D ir and D vi Indicates an infrared image and a visible light image discriminator; This represents the nth fused image;

[0032] The content loss L content The formula used to ensure that the fused image retains more content information from both visible and infrared images is expressed as follows:

[0033]

[0034] In the formula, The original infrared image; Visible light image;

[0035] The gradient loss L gradient To ensure the clarity of image edge information and prevent information loss, its formula is expressed as:

[0036]

[0037] In the formula, Let be the gradient of the nth fused image; Let be the gradient of the nth visible light image;

[0038] The structural loss L strUsed to recover the original image and perform feature extraction, ensuring the quality of the fused image, its formula is expressed as:

[0039] L str =(1-SSIM(I) f ,I vi ))+(1-SSIM(I f ,I ir ))+MSE(I rec_vi ,I vi )+MSE(I rec_ir ,I ir )

[0040] In the formula, SSIM is the structural similarity loss, and the closer SSIM is to 1, the more similar the structures are; MSE is the mean squared error; MSE(I rec_vi ,I vi ), MSE(I rec_ir ,I ir This is used to calculate the structural loss between the restored image and the original image to ensure the quality of image fusion.

[0041] Preferably, the discriminator loss function is expressed as follows:

[0042]

[0043] In the formula, L Dadv To combat losses, λ gp L is the loss coefficient. gp Gradient penalty;

[0044] The resistance loss L Dadv It consists of two parts, and its formula is expressed as:

[0045]

[0046] In the formula, D ir / vi This indicates an infrared image discriminator or a visible light image discriminator.

[0047] An infrared and visible light image fusion system based on AE networks and generative adversarial networks includes:

[0048] Image fusion network module: used to fuse the AE network structure into the generator network, and through training the model, enable the generator to fuse infrared images and visible light images;

[0049] Loss function module: For generative adversarial networks, it includes loss functions for the generator and discriminator. The generator's loss aims to retain more visible light and infrared image information to generate a clear fused image. The discriminator's loss aims to distinguish between the original image and the fused image to improve the generator's ability to fuse images.

[0050] Dataset module: Contains a dataset of infrared and visible light image pairs to train and validate the proposed method. The model is trained and tested using data from large power equipment. The images are scaled to ensure that the input images meet the model requirements.

[0051] Model training and validation module: Used to train the image fusion network. During training, the network parameters are optimized using the loss function in the loss function module. After training, the model's performance is validated on the test set. The fusion effect of the model is evaluated by comparing the images before and after fusion.

[0052] The present invention further protects a computer device, the computer device including a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set or instruction set, the instruction, program, code set or instruction set being loaded and executed by the processor to implement the above-mentioned deep learning-based method for fusing infrared and visible light images of power transmission line inspection.

[0053] The present invention further protects a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the instruction, program, code set, or instruction set is loaded and executed by a processor to implement the above-mentioned deep learning-based method for fusing infrared and visible light images of power transmission line inspection.

[0054] Compared with existing technologies, it has the following beneficial effects:

[0055] This invention proposes a deep learning-based method for fusing infrared and visible light images of power transmission line inspections, which has the following advantages:

[0056] (1) This invention introduces the generator structure of the AE network: by introducing the AE network to extract features from infrared and visible light images, the original image is restored by decoding to ensure the integrity of feature extraction; the generator generates a fused image of visible light and infrared, eliminating the need to manually design the feature fusion strategy of the network, making the method more widely applicable and able to be effectively applied in more real-world scenarios, and significantly improving the generalization ability of the model.

[0057] (2) The present invention uses a dual discriminator structure and designs the necessary loss function: Through the dual discriminator structure design, the fused image can better retain the target features of infrared and the detailed features of visible light image. By adding the necessary loss function, the quality of the fused image is significantly improved, enabling it to cope with more complex scenes and environments in practical applications.

[0058] (3) The present invention introduces an image reconstruction step before image fusion, and decodes and reconstructs infrared and visible light images respectively, which serve as auxiliary inputs to the generator to ensure that the original modal information is preserved during the fusion process and to enhance the structural integrity of the fused image.

[0059] (4) The loss function designed in this invention not only includes traditional adversarial, content, gradient and structural losses, but also incorporates structural loss evaluation of reconstructed images, further improving the realism and detail restoration capabilities of fused images. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings involved in the embodiments are now briefly described. Obviously, the drawings in the following description are merely illustrative of some embodiments of the present invention. For those skilled in the art, other forms of drawings can be constructed based on these drawings without creative effort.

[0061] Figure 1 This is a schematic diagram of the infrared and visible light image fusion network based on AE network and generative adversarial network mentioned in Embodiment 1 of the present invention;

[0062] Figure 2 This is a schematic diagram of the AE network structure mentioned in Embodiment 1 of the present invention;

[0063] Figure 3 The training and testing datasets mentioned in Embodiment 1 of the present invention are shown, wherein (a) represents a visible light image; and (b) represents an infrared image.

[0064] Figure 4 The above is a comparison of power equipment datasets mentioned in Embodiment 2 of the present invention, wherein (a) represents an infrared image, (b) represents a visible light image, (c) represents a DIDF model, (d) represents an ICAF model, and (e) represents the method of the present invention. Detailed Implementation

[0065] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0066] First, the definitions of abbreviations and key terms mentioned in this invention will be explained:

[0067] ① AE network (Auto encoder network): The AE network is an unsupervised learning neural network model, mainly used for data dimensionality reduction, feature extraction and generation tasks. Its core idea is to learn an efficient representation of data through the encoding (compression) and decoding (reconstruction) process.

[0068] ② Generative Adversarial Network (GAN): A Generative Adversarial Network (GAN) is a deep learning model that generates realistic data through adversarial training. It consists of a generator and a discriminator, which work together to optimize in a game.

[0069] ③ Infrared and visible image fusion: Infrared and visible image fusion aims to combine the complementary information of the two types of images to generate a more informative fused image. Infrared images, based on differences in thermal radiation, can clearly distinguish targets from the background in all weather conditions, low light, or complex and harsh weather conditions, effectively highlighting salient targets. However, they have lower image resolution and lack corresponding texture details. Visible images, on the other hand, rely on reflected light imaging, possessing high spatial resolution and rich texture information, and are more consistent with human visual perception. However, their images are susceptible to interference from environmental factors such as strong light and haze. Infrared and visible image fusion can significantly improve the application effects in fields such as target detection, military reconnaissance, autonomous driving, and medical diagnosis.

[0070] The following description, in conjunction with relevant accompanying drawings and specific examples, illustrates a deep learning-based method for fusing infrared and visible light images of power transmission lines.

[0071] Example 1:

[0072] This invention proposes a deep learning-based method for fusing infrared and visible light images of power transmission line inspections, comprising the following steps:

[0073] Step 1: Building an image fusion network

[0074] The network constructed in this invention integrates an infrared (AE) network structure into a generative network. Through model training, the generator achieves the fusion of infrared and visible light images. First, the image is input into the AE network to extract image features and reconstruct the original image. Then, the extracted features and the reconstructed image are input into the generator to generate a fused image. A dual discriminator structure is used to optimize network parameters, preventing the loss of visible light texture and detail features, as well as infrared target features. More specifically, it includes the following:

[0075] Please see Figure 1 The network structure proposed in this invention is as follows: Figure 1As shown, the AE network consists of an encoder and a decoder. The encoder decomposes and extracts image features, which are then input into the decoder for image restoration and reconstruction. The features are reconstructed by the decoder to obtain infrared and visible light feature images, which are then fed together with the effective features extracted by the encoder into the generative fusion network to finally generate an infrared-visible light fused image. To ensure the quality of the fused image, a dual discriminator structure is adopted. One discriminator is used to distinguish between the fused image and the visible light image to prevent the loss of more texture and detail information from the visible light image. The other discriminator is used to distinguish between the fused image and the infrared image to prevent the loss of target features in the infrared image, ultimately generating a relatively complete fused image.

[0076] (1) Generator Network Structure

[0077] The generator consists of an AE network and a fusion network, where the AE network structure is as follows: Figure 2 As shown, the encoder consists of four convolutional blocks (convolutional layer + normalization + activation function), and the convolutional block can be represented by the following formula:

[0078] output=f(BatchNorm(conv(input))) (1)

[0079] Wherein, input represents the input of the convolutional block, i.e., the image or feature information; f(·) is the activation function.

[0080] To prevent the loss of edge information and the resulting artifacts in the reconstructed image, reflection padding is used in the first convolutional block. Visible light and infrared images are input into the convolutional layer to extract low-level and high-level features. Low-level features contain more detailed information, while high-level features have rich semantic information. The decoder consists of three convolutional blocks, which concatenate low-level and high-level features and feed them into the convolutional layer. Each convolution input requires concatenation of a low-level feature to ensure the richness and completeness of the generated image's detailed information. Finally, the reconstructed images of visible light and infrared are output.

[0081] To preserve rich detail and infrared feature information, the recovered image will be compared with the features extracted by the encoder. Figure 1 The image is fed into the generator fusion network, which automatically fuses the image information, solving the problem of having to design a fusion strategy when using the AE network. The generator network consists of four convolutional blocks. The input undergoes padding, convolution, spectral normalization, batch normalization, activation functions, and other structures before outputting the fused image. The parameters of each layer are shown in Table 1, where I represents the number of input channels, O represents the number of output channels, K is the kernel size, and S is the kernel stride. The first two layers perform reflection padding to ensure the integrity of edge information.

[0082] Table 1 Fusion Network Parameters

[0083]

[0084]

[0085] (2) Discriminator network structure

[0086] The discriminator extracts features through four convolutional layers, integrates the feature information, and feeds it into the tanh function after passing through a linear layer to distinguish between the original visible light image, infrared image, and fused image. The parameters of its convolutional layers are shown in Table 2.

[0087] Table 2. Discriminator Network Parameters

[0088]

[0089] Step 2: Design the loss function

[0090] For generative adversarial networks (GANs), loss functions are designed for the generator and discriminator respectively. The generator's loss function aims to retain more information from both visible light and infrared images to generate a high-quality, clear fused image. The discriminator's loss function aims to distinguish between the original image and the fused image, ultimately improving the generator's ability to fuse images. More specifically, this includes the following:

[0091] (1) Generator loss function

[0092] The generator aims to produce a high-quality fused image that retains more infrared and visible light image features while ensuring the quality of the generated image. Its loss function is as follows:

[0093] L G =λ adv L adv +λ content L content +λ gradient L gradient +λ str L str (2)

[0094] Where, λ adv , λ content , λ gradient , λ str L is the loss coefficient used to adjust the loss balance. adv L is the adversarial loss between the generator and the discriminator. content For content loss, L gradient For gradient loss, L str This is structural loss.

[0095] The generator aims to produce an image that can fool the discriminator, and its loss is:

[0096]

[0097] Among them, D ir and D vi This represents a discriminator for infrared and visible light images. This represents the nth fused image.

[0098] To ensure that the fused image retains more content information from both visible light and infrared images, a content loss is introduced, the formula of which is:

[0099]

[0100] in, For visible-infrared fused images, The original infrared image, This is a visible light image.

[0101] To better preserve edge information, maintain edge clarity, prevent spatial gradient loss due to information loss, and avoid retaining too much infrared blur in the fused image, only the gradients of the visible light image and the fused image were calculated, and the formula is as follows:

[0102]

[0103] In the formula, Let be the gradient of the nth fused image; Let be the gradient of the nth visible light image;

[0104] To better restore the original image and achieve accurate feature extraction, while ensuring the quality of the fused image, a structural loss is introduced, the calculation formula of which is as follows:

[0105] L str =(1-SSIM(I) f ,I vi ))+(1-SSIM(I f ,I ir ))+MSE(I rec_vi ,I vi )+MSE(I rec_ir ,I ir (6)

[0106] SSIM is the structural similarity loss, and the closer SSIM is to 1, the more similar the structures are. MSE is the mean squared error. The latter two terms calculate the structural loss between the restored image and the original image to ensure the fusion quality of the images.

[0107] (2) Discriminator Loss Function

[0108] The discriminator aims to distinguish between the fused image and the original image, and its loss function is as follows:

[0109]

[0110] Among them, L Dadv Indicates the resistance loss, λ gp L is the loss coefficient. gp This is gradient penalty, used in Wasserstein GAN to stabilize training and prevent the discriminator from becoming too strong, which could lead to vanishing gradients or mode collapse.

[0111] The discriminator must accurately identify both the original real image and the fused fake image; therefore, its adversarial loss also consists of two parts, as shown in the following formula:

[0112]

[0113] In the formula, D ir / vi Indicates an infrared image discriminator or a visible light image discriminator.

[0114] Step 3: Building the dataset

[0115] To train and validate the proposed method, a dataset containing infrared and visible light image pairs needs to be constructed. A large power equipment dataset is used for model training and testing. First, the images are scaled to ensure the input images meet the model requirements. More specifically, this includes the following:

[0116] This invention is primarily used to achieve infrared-visible light image fusion for large outdoor power equipment. Both training and testing utilize a power equipment dataset, with data allocated in an 8:2 ratio. The performance of the image fusion model is tested and evaluated. Some data from the dataset are shown below. Figure 3 As shown, the first and second columns are training set images, and the third and fourth columns are test set images.

[0117] Step 4: Model Training and Validation

[0118] The constructed fusion network was trained. During training, the network parameters were optimized using the loss function designed above. After training, the model's performance was verified on the test set, and the fusion effect was evaluated by comparing the images before and after fusion. Experimental results show that the proposed method can effectively achieve the fusion of infrared and visible light images, improve image clarity, and significantly enhance image quality, specifically including the following:

[0119] The training process of the experiment designed in this invention was conducted on a large-scale power equipment dataset. During the experiment, the images were first adjusted to grayscale and then centered and cropped to 128×128 pixels. This preprocessing of the original images ensured the consistency of the input network image size, making good use of the training set. Furthermore, the Adam optimizer was used to update the model parameters, with batch size and epochs set to 24 and 80 respectively. The network underwent adaptive parameter optimization using Adam, and the generator's learning rate was set to 1×10⁻⁶. -3 The learning rate of the discriminator is set to 1×10. -3 Furthermore, the generator's learning rate decreases as the epoch increases, with the generator's learning rate decreasing by a factor of 10 every 1 / 3 of an epoch, ensuring model convergence.

[0120] Example 2:

[0121] Based on Example 1, but with a difference, this invention further illustrates the performance of the deep learning-based infrared and visible light image fusion method for power transmission line inspection proposed in this invention through characterization experiments, as detailed below.

[0122] (1) Quantitative evaluation of model performance

[0123] In infrared-visible image fusion algorithms, infrared images reflect the thermal radiation characteristics of a target, while visible images provide rich texture and contour information. The core of the algorithm lies in preserving and enhancing the complementary information of the source images, and its performance needs to be quantitatively evaluated using objective metrics. The specific evaluation metrics for infrared-visible image fusion include the following:

[0124] Information entropy (EN): Reflects the richness of image information; a higher entropy value indicates better detail preservation. Its formula is:

[0125]

[0126] Where L is the number of gray levels, p i To fuse the normalized histograms corresponding to the gray levels in the image;

[0127] Mutual Information (MI): Calculates the amount of information shared between the fused image and the source image. A higher MI value indicates a better fusion effect. The formula is:

[0128] MI = MI V,F +MI I,F (10)

[0129] MI V,F and MI I,F The information transferred from the source image to the fused image is represented by:

[0130]

[0131] In the formula P X (x) and P F (f) represents the edge histograms of the source image and the fused image, P X,F (x,f) represents the joint histogram of the source image and the fused image;

[0132] Average Gradient (AG): Reflects image sharpness and edge sharpness. A higher AG value indicates more complete detail retention. Its formula is:

[0133]

[0134] Where F(i,j) represents the pixel value of the image;

[0135] Spatial frequency (SF): Calculates the rate of change of an image in the spatial domain. The more high-frequency components, the richer the details. The formula is:

[0136]

[0137] Where RF represents row frequency and CF represents column frequency;

[0138] Standard deviation (SD): Measures the dispersion of pixel value distribution. A high SD value indicates strong contrast. The formula is:

[0139]

[0140] Where μ represents the mean of the fused images.

[0141] The model achieved a best SD of 50.817 on the test dataset, demonstrating that the generated fused images are rich in information and have significant differences in brightness.

[0142] The optimal EN value is 7.218, indicating that the generated fused image contains more detail and texture information.

[0143] The optimal SF value is 12.545, which shows that the generated fused image has clear edges and texture details, high sharpness, and obvious local changes in the image.

[0144] The optimal AG value is 9.406, indicating that the target outline in the generated fused image is clear and the details are well preserved.

[0145] (2) Qualitative results of model performance

[0146] During testing, a large-scale power equipment dataset was selected for fusion testing to demonstrate the effectiveness of the proposed model. Several representative state-of-the-art (SOTA) methods were compared with the proposed network, and the comparison results are as follows: Figure 4As shown, the images, from left to right, are IR infrared images, VIS visible light images, DIDF model fusion images, ICAF model fusion images, and the fusion image of the model proposed in this paper. It can be seen from the images that the fusion image generated by the DIDF model exhibits significant darkness, resulting in the loss of a large amount of detail and obvious errors. While the image generated by the ICAF model performs slightly better than DIDF, it still loses much detail and exhibits significant noise, resulting in a marked decrease in the quality of the fused image. In contrast, the network proposed in this paper can not only distinguish the texture details of the visible light image from the infrared target features without sacrificing image quality, but also enhances the sharpness to a certain extent, while retaining a large amount of important detail information, generating a high-quality fused image. Experimental results show that the network proposed in this paper has good generalization ability and can be well applied in practical fields, playing a greater role.

[0147] Compared to existing methods (such as DenseFuse, FusionGAN, etc.) that only extract features for fusion, this invention adds an image reconstruction process after feature extraction and guides the fusion process with the reconstructed image, further improving the realism and detail preservation of image fusion. Experimental results also show that the image fused by this method is superior to existing mainstream methods in terms of visual effect and quantitative indicators.

[0148] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for fusing infrared and visible light images of power transmission line inspection based on deep learning, characterized in that, Includes the following steps: S1. Construct an image fusion network: Integrate the AE network structure into the generator network, extract image features through the AE network and reconstruct the original infrared and visible light images, and then input the extracted feature maps and the reconstructed images into the generator to generate a fused image, thereby improving the fusion effect and image clarity. S2. Loss Function Design: For generative adversarial networks, loss functions are designed for the generator and discriminator respectively. The generator's loss function is a composite loss function, which includes content loss, gradient loss, structural similarity loss, and reconstruction loss to constrain the consistency between the reconstructed image and the original image, so as to ensure the quality of the fused image in terms of sharpness, structural fidelity, and detail restoration. The discriminator's loss is used to identify the differences between the original image and the fused image, so as to further improve the generator's fusion capability. S3. Constructing the dataset: Construct a dataset containing pairs of infrared and visible light images to train and validate the proposed method. Use large power equipment data for model training and testing, and scale the images to ensure that the input images meet the model requirements. S4. Model Training and Validation: The image fusion network obtained in S1 is trained. During the training process, the network parameters are optimized using the loss function designed in S2. After training, the model's performance is validated on the test set. The fusion effect of the model is evaluated by comparing the images before and after fusion.

2. The method for fusing infrared and visible light images of power transmission line inspection based on deep learning according to claim 1, characterized in that, S1 specifically includes the following: S1.1 Input the image into the AE network, extract image features, and reconstruct the original image; S1.

2. Input the extracted features and the restored image into the generator to generate a fused image. Optimize the network parameters using a dual discriminator structure to prevent the loss of texture and detail features from the visible light image and target features from the infrared image.

3. The method for fusing infrared and visible light images of power transmission line inspection based on deep learning according to claim 2, characterized in that, The generator consists of an AE network and a fusion network, wherein the encoder of the AE network consists of four convolutional blocks, and the formula for each convolutional block is expressed as: output=f(BatchNorm(conv(input))) In the formula, input represents the input of the convolutional block, which is an image or feature information; f(·) represents the activation function; the first convolutional block of the encoder uses reflection filling to input visible light images and infrared images into the convolutional layer to extract low-level and high-level features of the image; The decoder of the AE network consists of three convolutional blocks, which concatenate low-level features with high-level features and feed them into the convolutional layer. Each convolution input is concatenated with a low-level feature to ensure the richness and completeness of the generated image details, and outputs restored images of visible light and infrared light. The restored image and the feature map extracted by the encoder are fed into the fusion network of the generator to achieve the fusion of image information. Finally, the generator network outputs the fused image.

4. The method for fusing infrared and visible light images of power transmission line inspection based on deep learning according to claim 3, characterized in that, The discriminator extracts features through four convolutional layers, integrates the feature information, and then feeds it into the tanh function after passing through a linear layer to distinguish between the original visible light image, infrared image, and fused image.

5. The method for fusing infrared and visible light images of power transmission line inspection based on deep learning according to claim 4, characterized in that, The generator loss function is expressed as follows: L G =λ adv L adv +λ content L content +λ gradient L gradient +λ str L str In the formula, λ adv , λ content , λ gradient , λ str L is the loss coefficient used to adjust the loss balance. adv L is the adversarial loss between the generator and the discriminator. content For content loss, L gradient For gradient loss, L str For structural loss; The resistance loss L adv The formula used to ensure that the generator produces images that deceive the discriminator is expressed as follows: In the formula, D ir and D vi Indicates an infrared image and a visible light image discriminator; This represents the nth fused image; The content loss L content The formula used to ensure that the fused image retains more content information from both visible and infrared images is expressed as follows: In the formula, The original infrared image; Visible light image; The gradient loss L gradient To ensure the clarity of image edge information and prevent information loss, its formula is expressed as: In the formula, Let be the gradient of the nth fused image; Let be the gradient of the nth visible light image; The structural loss L str Used to recover the original image and perform feature extraction, ensuring the quality of the fused image, its formula is expressed as: L str =(1-SSIM(I f ,I vi ))+(1-SSIM(I f ,I ir ))+MSE(I rec_vi ,I vi )+MSE(I rec_ir ,I ir ) In the formula, SSIM is the structural similarity loss, and the closer SSIM is to 1, the more similar the structures are; MSE is the mean squared error; MSE(I rec_vi ,I vi ), MSE(I rec_ir ,I ir This is used to calculate the structural loss between the restored image and the original image to ensure the quality of image fusion.

6. The method for fusing infrared and visible light images of power transmission line inspection based on deep learning according to claim 5, characterized in that, The discriminator loss function is expressed as follows: In the formula, L Dadv To combat losses, λ gp L is the loss coefficient. gp Gradient penalty; The resistance loss L Dadv It consists of two parts, and its formula is expressed as: In the formula, D ir / vi This indicates an infrared image discriminator or a visible light image discriminator.

7. The infrared and visible light image fusion system based on AE network and generative adversarial network applied to the method described in any one of claims 1-6, characterized in that, include: Image fusion network module: used to fuse the AE network structure into the generator network, and through training the model, enable the generator to fuse infrared images and visible light images; Loss function module: For generative adversarial networks, it includes loss functions for the generator and discriminator. The generator's loss aims to retain more visible light and infrared image information to generate a clear fused image. The discriminator's loss aims to distinguish between the original image and the fused image to improve the generator's ability to fuse images. Dataset module: Contains a dataset of infrared and visible light image pairs to train and validate the proposed method. The model is trained and tested using data from large power equipment. The images are scaled to ensure that the input images meet the model requirements. Model training and validation module: Used to train the image fusion network. During training, the network parameters are optimized using the loss function in the loss function module. After training, the model's performance is validated on the test set. The fusion effect of the model is evaluated by comparing the images before and after fusion.

8. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement the deep learning-based infrared and visible light image fusion method for power transmission line inspection as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement the deep learning-based infrared and visible light image fusion method for power transmission line inspection as described in any one of claims 1-6.