A Dual-Link Neural Network Infrared Image Contrast Enhancement Method

A dual-path CNN architecture enhances infrared image contrast and detail by distinguishing target and background regions, outperforming traditional methods in PSNR and SSIM metrics.

CN114792293BActive Publication Date: 2025-07-15SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210390857.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2025-07-15
Estimated Expiration
2042-04-14

AI Technical Summary

Technical Problem

The existing infrared image enhancement algorithms have shortcomings in improving image contrast and texture details. In particular, traditional methods are prone to enhancing noise and have limited effects, making it difficult to effectively distinguish between target areas and background areas.

Method used

A dual-link convolutional neural network is built, including feature extraction module 1 and feature extraction module 2. Feature fusion is performed through feature fusion module, and the convolutional neural network learns the differences in the target and background, suppresses background noise and improves contrast and texture details.

Benefits of technology

The highlighting of the target area in the infrared image and the suppression of background noise are achieved, which significantly improves the contrast and texture details of the image, and achieves higher peak signal-to-noise ratio and structural similarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114792293B_ABST
    Figure CN114792293B_ABST
Patent Text Reader

Abstract

A method for enhancing the contrast of infrared images using a dual-link neural network. An infrared image dataset is constructed and preprocessed. Using deep learning methods, a dual-link convolutional neural network containing two feature extraction modules and one feature fusion module is designed. Feature extraction module one is used to extract the shallow feature information of infrared images. Only three layers of convolutional neural network are used to avoid the loss of detailed information due to the deepening of the network depth and the increase in the number of layers. Feature extraction module two is used to extract the high-level feature information of infrared images. Finally, the feature fusion module fuses the extracted features to enhance the features and enrich the detailed information. The final network model effectively predicts the target area and background area in the image, highlights the target, suppresses background noise, and improves the image contrast and texture details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of digital image processing, and particularly relates to a dual-link neural network infrared image contrast enhancement method. Background Art

[0002] Since an infrared imager forms an image based on temperature difference, and generally the temperature difference between the target and the environment is not large, the contrast of infrared images is low and the detail resolution ability is poor. Moreover, sensors that can obtain high-quality infrared images are often costly and expensive, which to a certain extent limits the application of infrared images in the fields of aerospace, military, medical, industrial, etc. Therefore, it is necessary to use image enhancement methods to process infrared images to improve image quality.

[0003] Relatively classic traditional algorithms in the field of image enhancement include: histogram equalization (HE) and its variants such as contrast-limited adaptive histogram equalization (CLAHE), Retinex series algorithms, etc. Histogram equalization (HE) performs well in visible light image enhancement, but for the processing of infrared images, it usually enhances noisy noise while enhancing the target, there is a problem of over-enhancement, and the detail enhancement effect is relatively limited; contrast-limited adaptive histogram equalization (CLAHE) can suppress noise to a certain extent while increasing the contrast, but it will also produce more blurred details. The Retinex theory can retain image details while reducing noise to a certain extent when processing infrared images. Due to some inherent defects of traditional infrared image enhancement algorithms, the enhancement effect on infrared images still needs to be improved. A dual-link convolutional neural network is designed for model training. The trained model effectively predicts the target area and background area in the image, highlights the target, suppresses background noise, and improves image contrast and texture details. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to overcome the deficiencies of the prior art and provide a dual-link convolutional neural network infrared image contrast enhancement method, which can effectively predict the target area and background area in the infrared image, highlight the target, suppress background noise, and improve image contrast and texture details.

[0005] To solve the problems existing in the prior art, a dual-link neural network infrared image contrast enhancement method provided by the present invention constructs a neural network composed of a feature extraction module one (1), a feature extraction module two (2), and a feature fusion module (3), where

[0006] The feature extraction module 1 (1) consists of three convolutional layers and two ReLU activation functions. The number of channels in the convolutional layers is set to 64, 32, and 32. A ReLU activation function is connected after each of the first two convolutional layers, and the output of the third convolutional layer is used for fusion with the output of the feature extraction module 2 (2).

[0007] The feature extraction module 2 (2) contains five convolutional layers, and the number of channels is 32, 64, 96, 32, and 32 in sequence. Among them, the second, third, and fourth convolutional layers form a residual structure.

[0008] The second convolutional layer of the feature extraction module 1 (1) and the third convolutional layer of the feature extraction module 2 (2) are connected using the Concat operation.

[0009] The feature fusion module (3) adds the features extracted by the feature extraction module 1 (1) and the feature extraction module 2 (2) and performs a convolutional operation.

[0010] Preferably, the feature extraction module 1 (1) does not have a pooling layer. The size of each convolutional kernel is set to 3x3, the stride is set to 1, and the padding is 1 to ensure that the size of the output image after being processed by the network is the same as the size of the input image. The feature extraction module 2 (2) does not have a pooling layer. For the convolutional layer with 64 channels, the size of its convolutional kernel is set to 5x5, the stride is 1, and the padding is 2; for the convolutional layers connected through the Concat operation in the feature extraction module 1 (1) and the feature extraction module 2 (2), the number of channels is 96, the size of its convolutional kernel is set to 7x7, the stride is 1, and the padding is 3; the size of the convolutional kernels of other convolutional layers is 3x3, the stride is set to 1, and the padding is 1 to ensure that the size of the output image after being processed by the network is the same as the size of the input image.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0012] (1) The present invention uses a convolutional neural network to learn the differences between the targets and the background in the image, enhancing the contrast, texture, and details between the detailed targets and the background in the infrared image.

[0013] (2) The present invention has a higher peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) compared with other traditional methods. Description of the Drawings

[0014] Figure 1 is a structural diagram of a dual-link neural network for enhancing the contrast of infrared images provided by the present invention.

[0015] Figure 2 is a structural diagram of the first feature extraction module for enhancing the contrast of infrared images provided by the present invention.

[0016] Figure 3 is a structural diagram of the second feature extraction module for enhancing the contrast of infrared images provided by the present invention.

[0017] Figure 4 shows the low-quality infrared image provided by the present invention.

[0018] Figure 5 shows the high-quality infrared image provided by the present invention.

[0019] Figure 6 shows the comparison chart of the enhancement effects of different methods on the infrared image provided by the present invention. Specific Embodiments

[0020] To illustrate the content and implementation steps of the present invention, the present invention will be further described in detail below with reference to the example diagrams.

[0021] The embodiments of the present invention are as follows:

[0022] Construct an infrared image dataset for training a convolutional neural network. Use the saturated and street datasets in the publicly available infrared image dataset TheLTIR dataset v1.0. The saturated dataset contains 218 infrared images, and the street dataset contains 172 infrared images. The two datasets together contain 390 infrared images.

[0023] Preprocess the data. First, uniformly scale the pictures to a size of 256X256, and then use data augmentation methods such as rotation, translation, and flipping to augment the dataset. After augmentation, the dataset contains a total of 1680 infrared images. Modify the picture contrast factor to 0.3~0.31 as low-gain pictures, and the label pictures are only scaled to a size of 256X256 without modifying the contrast factor and are used as reference pictures. The training set and the test set are divided in a ratio of 5:1, with 1400 pictures in the training set and 280 pictures in the test set.

[0024] Use the Pytorch framework to build a convolutional neural network. The network structure is as Figure 1 shown: It includes a feature extraction module one, a feature extraction module two, and a feature fusion module, a total of nine-layer convolutional neural network; use one skip link and two residual modules.

[0025] The structure of the feature extraction module one is as Figure 2As shown in the figure, it contains 3 convolutional layers, Conv1&ReLU, Conv2&ReLU, and Conv3. The number of channels is set to 64, 32, and 32. After the first two convolutional layers Conv1 and Conv2, a ReLU activation function is connected to increase non-linearity. The output of the convolutional layer Conv2 is concatenated with the convolutional layer Conv6 of Feature Extraction Module 2 to increase its number of channels and solve the problem of feature information loss in Feature Extraction Module 2 as the network depth increases. Feature Extraction Module 1 is mainly used to extract shallow feature information such as edges, geometry, and texture. The output of its convolutional layer Conv3 and the output of the convolutional layer Conv8 of Feature Extraction Module 2 are used for processing in the Feature Fusion Module. Feature Extraction Module 1 does not have a pooling layer. The size of the convolutional kernel is 3x3, the stride is set to 1, and the padding is 1 to ensure that the size of the input infrared image remains unchanged after being processed by the network.

[0026] The structure of Feature Extraction Module 2 is as Figure 3 shown. It contains 5 convolutional layers, Conv4&ReLU, Conv5&ReLU, Conv2&Conv6&ReLU, Conv7+Conv4&ReLU, and Conv8. After the first four convolutional layers, a ReLU activation function is connected. The number of channels is set to 32, 64, 96, 32, and 32 in sequence. Among them, the convolutional layers Conv5, Conv6, and Conv7 form a residual structure, which helps to solve the problems of gradient disappearance and gradient explosion, enabling the training of deeper networks while ensuring good information. Similar to Feature Extraction Module 1, Feature Extraction Module 2 does not have a pooling layer. The number of channels of the convolutional layer Conv5 is 64, the size of the convolutional kernel is 5x5, the stride is 1, and the padding is 2. The number of channels of the convolutional layer Conv2&Conv6 is 96, the size of the convolutional kernel is 7x7, the stride is 1, and the padding is 3. The size of the convolutional kernel of other convolutional layers is 3x3, the stride is set to 1, and the padding is 1 to ensure that the size of the input infrared image remains unchanged after being processed by the network.

[0027] The Feature Fusion Module adds and fuses the features extracted by Feature Extraction Module 1 and Feature Extraction Module 2, and outputs after convolution operation through the convolutional layer Conv9 to ensure that the features of the original image are fully learned. The detailed parameter settings of the convolutional layers of the entire network are shown in Table 1:

[0028]

[0029] Set the loss function. The greater the difference between the true value and the predicted value, the greater the Loss. The goal of optimization is to reduce Loss. The SmoothL1Loss function can be used as the loss function, and its formula is as follows:

[0030]

[0031] Among them, x represents the enhanced infrared image, and y represents the high-quality reference image. The low-quality infrared image provided by the present invention is shown in Figure 4, and the high-quality infrared image is as Figure 5 shown.

[0032] Train the convolutional neural network for infrared image contrast enhancement. Instantiate the built neural network, and then instantiate the data loader. Load the constructed infrared image training set into the network, and construct an Adam optimizer object Optimizer to save the current state and be able to update the network parameters, learning rate, and weight decay according to the calculated gradients. The initial learning rate is 0.0001, the weight decay is set to 0.0001, the number of training epochs is 400, the Batch Size is set to 4, the number of iteration epochs is 350, and the optimal model is saved after training.

[0033] Test the trained model on the test set and save the result pictures generated by the test. In addition to using the dual-link neural network infrared image contrast enhancement method described in the present invention, other methods such as histogram equalization (HE), dual-platform histogram equalization (DPHE), single-scale Retinex (SSR), etc. are also used for comparison of enhancement effects. The comparison diagrams of enhancement effects of different methods are as Figure 6 shown. It can be seen from the effect diagrams that histogram equalization (HE) and histogram equalization algorithm with limited contrast have the effect of improving image clarity, but it is not obvious. Dual-platform histogram equalization (DPHE) improves the brightness of the target, but the background is over-suppressed. The multi-scale Retinex (MSR) algorithm has obvious enhancement effect and is already very close to the high-quality image. Compared with the above algorithms, the method described in the present invention suppresses background noise and improves image contrast closer to the high-quality image.

[0034] Calculate the evaluation indicators using the generated result pictures and reference pictures, that is, calculate the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). The calculation formulas are as follows:

[0035]

[0036] The comparison results of the evaluation indicators of the enhancement effects of different methods are shown in Table 2:

[0037]

[0038] Result analysis: Among the above several traditional algorithms, the peak signal-to-noise ratio (PSNR) of the multi-scale Retinex (MSR) algorithm is the highest, which is 22.91. The single-scale Retinex (SSR) algorithm and the multi-scale Retinex (MSR) algorithm have the same structural similarity (SSIM), which is 0.94. Among all the above methods, the peak signal-to-noise ratio (PSNR) of the enhancement method proposed in this patent is the highest, which is 25.56, and the structural similarity (SSIM) is 0.95.

[0039] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for enhancing the contrast of infrared images in a dual-link neural network, characterized in that It includes the following steps: S1: Construct an infrared image dataset; S2: Preprocess the infrared image dataset; S3: Build a dual-link neural network. The constructed dual-link neural network includes a feature extraction module one (1), a feature extraction module two (2), and a feature fusion module (3); where The feature extraction module one (1) extracts shallow feature information including edges, geometry, and texture, and is composed of three convolutional layers and two ReLU activation functions. The number of channels in the convolutional layers is set to 64, 32, 32; ReLU activation functions are connected after the first two convolutional layers, and the output of the third convolutional layer is used for fusion with the output of the feature extraction module two (2); The feature extraction module two (2) extracts high-level feature information of the infrared image, including 5 convolutional layers, and the number of channels is 32, 64, 96, 32, 32 in sequence, where the second, third, and fourth convolutional layers form a residual structure; The second convolutional layer of the feature extraction module one (1) and the third convolutional layer of the feature extraction module two (2) are connected using the Concat operation and input to the fourth convolutional layer of the feature extraction module two (2) for further processing; The feature fusion module (3) adds the features extracted by the feature extraction module one (1) and the feature extraction module two (2) and performs a convolutional operation to ensure full learning of the features of the original image; S4: Train the constructed dual-link neural network through the infrared image training set and save the optimal model; S5: Perform testing through the optimal model to obtain the infrared image after contrast enhancement.

2. The dual-link neural network infrared image contrast enhancement method according to claim 1, wherein The feature extraction module one (1) does not set a pooling layer. The size of each convolutional kernel is set to 3x3, the stride is set to 1, and the padding is 1 to ensure that the size of the output image after network processing is the same as the size of the input image; the feature extraction module two (2) does not set a pooling layer. For the convolutional layer with 64 channels, the size of its convolutional kernel is set to 5x5, the stride is 1, and the padding is 2; for the convolutional layers connected through the Concat operation in the feature extraction module one (1) and the feature extraction module two (2), the number of channels is 96, the size of its convolutional kernel is set to 7x7, the stride is 1, and the padding is 3; the size of the convolutional kernels of other convolutional layers is 3x3, the stride is set to 1, and the padding is 1 to ensure that the size of the output image after network processing is the same as the size of the input image.