Infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth
By using semantic segmentation and artifact detection technology to remove infrared pedestrian artifacts when fusing infrared and visible light images, the problem of artifact interference in the prior art is solved, and the resolution and quality of the fusion image are improved.
Patent Information
- Application Number
- CN202510160329.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, when infrared and visible light images are fused, it is difficult to effectively remove infrared pedestrian artifacts, resulting in a degradation in the quality of the final fusion result.
The method of removing the convolution depth of the feature of infrared pedestrian artifacts is adopted to obtain the joint area of the character and artifact through semantic segmentation and artifact detection, and a fused image is generated by combining feature maps and masks to reduce the adverse impact of artifacts on the final result.
The resolution and quality of infrared and visible light images are improved, the interference of artifacts on the fusion results is reduced, and the interpretability and accuracy of the image are enhanced.
Smart Images

Figure CN120013777A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth. Background Art
[0002] Image fusion refers to the technology of merging images from different sensors or different modalities into a single image. Through image fusion, more comprehensive and rich information can be obtained, so that the synthetic image has better visualization and more details. At the same time, image fusion can also help reduce data redundancy, improve the accuracy of image analysis and recognition, and enhance the quality and interpretability of images.
[0003] Visible light images have rich color information and can present rich color information, including the combination of the three primary colors of red, green, and blue, which can provide the natural color of objects in the real world; they have good visual perception, and visible light is easy for the human eye to perceive and understand; they have high resolution, clear detail texture information, and contain rich detail information. However, its imaging effect is affected by external factors such as light, rain, fog, haze, and smoke, and it is highly dependent on the external environment.
[0004] Infrared images are formed by the thermal radiation of objects. They are not strongly dependent on the environment and have strong penetrating power. Even in low-light environments or foggy weather, they still have good imaging effects. However, infrared images can only provide the general location of the target and lack detailed texture information. In addition, infrared thermal radiation often produces serious reflection phenomena, resulting in the appearance of thermal radiation artifacts. Artifacts may contain false information and carry a large amount of redundant information. Fusion of this information into other images may lead to results that are inconsistent with the actual situation, and may very likely lead to a decrease in the quality of the fused image, because artifacts will interfere with the accuracy of the fusion algorithm, resulting in defects or distortion in the final fusion result.
[0005] Based on the retrieval of the above materials, an infrared and visible light fusion method with deep convolution feature to remove infrared pedestrian artifacts is proposed. The resolution of the fused image is improved by strengthening the extraction of image edge features. When fusing infrared and visible light images, the artifacts in the infrared image are processed to reduce the adverse effects on the final fusion result. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention provides an infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth, which solves the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for fusion of infrared and visible light to remove infrared pedestrian artifact feature convolution depth, specifically comprising the following steps:
[0008] Step 1: Input the visible light source image and the infrared source image into the image enhancement module to obtain an enhanced visible light image and an enhanced infrared image;
[0009] Step 2: Input the visible light image, enhanced visible light image, infrared source image and enhanced infrared image into a feature extraction module to obtain a feature map;
[0010] Step 3: perform semantic segmentation and artifact detection on the enhanced infrared image, obtain the joint area of the person and the artifact, and generate a person mask mask1 and an artifact mask mask2;
[0011] Step 4: Combine the feature map, the person mask mask1 and the artifact mask mask2 to obtain a fused image.
[0012] The present invention is further configured as follows: the method for acquiring the enhanced visible light image in step 1 includes:
[0013] The method of reflection-guided histogram equalization plus comparison parameter approximation is used to process the visible light source image to obtain the enhanced visible light image.
[0014] The present invention is further configured as follows: the method for acquiring the enhanced infrared image in step 1 includes:
[0015] The infrared source image is processed using the infrared method of slope histogram to obtain an enhanced infrared image.
[0016] The present invention is further configured as follows: the feature map in step 2 includes a visible light detail feature map, a visible light background feature map, an infrared light detail feature map and an infrared light background feature map.
[0017] The present invention is further configured as follows: the feature extraction module in step 2 includes five direction-oriented depth expansion residual modules and a convolution layer 1;
[0018] The number of input channels of the five directional depth expansion residual modules is 1, and the number of output channels is 1;
[0019] The convolution kernel format of the convolution layer 1 is 3×3, and the step size, input channel number and output channel number are all 1.
[0020] The present invention is further configured as follows: the five direction-oriented depth expansion residual modules each include four direction-oriented convolution modules and one convolution layer two;
[0021] The activation functions of the four directional convolution modules are all ReLU, and every two are divided into a group, and the number of input and output channels of the two directional convolution modules in each group are oppositely set, specifically [1, 64][64, 1] and [2, 64][64, 2];
[0022] The convolution kernel format of the convolution layer 2 is 3×3, the number of input channels is 4, and the step size and the number of output channels are both 1.
[0023] The present invention is further configured as follows: the method of obtaining the fused image in step 4 includes:
[0024] Decompose complementary features in an image using a difference algorithm:
[0025]
[0026] In the formula, f i is the infrared image feature, f v is the visible light image feature, f i and f v The format is (H×W×4);
[0027] The complementary features are compressed into a vector with the help of global average pooling GAP, normalized to [0, 1] using the Sigmoid function, and the channel weights are generated. The complementary features are multiplied by the channel weights and added to the original complementary features as modal supplementary information to generate features. and The calculation formula includes:
[0028]
[0029] In the formula, and The format is (H×W×4), GAP is global average pooling, Conv is a 3×3 convolutional layer, BN is a regularization layer, and Sig is a Sigmoid loss function;
[0030] Input all features into the fusion network to obtain the fused image:
[0031] I f =Sig(BN(Conv(f s +f c )))
[0032] In the formula, f s Yes c The features obtained by three layers of convolution after semantic screening, f s and f c The format is (H×W), I f is the output fused image in the format of (H×W).
[0033] The present invention provides an infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth. It has the following beneficial effects:
[0034] The invention processes the artifacts in the infrared image when fusing the infrared and visible light images to reduce the adverse effects on the final fusion result, and adopts an iterative edge expansion feature extraction network to strengthen the extraction of image edge features to improve the resolution of the fused image. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0036] Figure 2 A schematic diagram of feature extraction in an embodiment of the present invention;
[0037] Figure 3 Schematic diagram of feature fusion in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0039] See also Figure 1-3 The embodiment of the present invention provides the following technical solution: a method for fusion of infrared and visible light to remove infrared pedestrian artifact feature convolution depth, specifically comprising the following steps:
[0040] Step 1: Input the visible light source image and the infrared source image into the image enhancement module, use the reflection guided histogram equalization plus comparison parameter approximation method to process the visible light source image, improve the amount of image detail information, and obtain an enhanced visible light image. Use the slope histogram infrared method to process the infrared source image, improve the image quality, and obtain an enhanced infrared image.
[0041] Step 2: Input the visible light image, enhanced visible light image, infrared source image and enhanced infrared image into a feature extraction module to obtain a feature map, wherein the feature map includes a visible light detail feature map, a visible light background feature map, an infrared detail feature map and an infrared background feature map, wherein the feature extraction module includes five directional depth expansion residual modules and a convolution layer one, the five directional depth expansion residual modules have an input channel number of 1 and an output channel number of 1, the convolution kernel format of the convolution layer one is 3×3, and the step size, input and output channel numbers are all 1, the five directional depth expansion residual modules each include four directional-oriented convolution modules and a convolution layer two, the activation functions of the four directional-oriented convolution modules are all ReLU, and every two are divided into a group, and the input and output channel numbers of the two directional-oriented convolution modules in each group are set oppositely, specifically [1, 64][64, 1] and [2, 64][64, 2], the convolution kernel format of the convolution layer two is 3×3, the input channel number is 4, and the step size and output channel number are both 1.
[0042] Further explanation: for the feature extraction module, let its input be f in , its output expression is as follows:
[0043] f out =PReLU(F c +F es +F sob +F lap )
[0044] F es =conv 3×3 (conv 1×1 (f in ))
[0045]
[0046] g(i,j)=4f(i,j)-f(i+1,j)-f(i-1,j)-f(i,j+1)-f(i,j-1)
[0047]
[0048] In the formula, f out is the output of the direction-oriented convolutional module, F c Yes in The intermediate result after 3×3 convolution; F es Yes in The result after first performing 1×1 channel processing and then 3×3 convolution, F sob Yes in The results obtained after the horizontal and vertical Sobel operators are shown above. lap Yes in First, the channel is combed and then the Laplacian filter is processed. Finally, all the above output results are input into the PreLU through the adder to get the output. All the above intermediate variables such as F c 、F es Equivalent to the final output f out The format is (H×W×3);
[0049] As attached Figure 2 As shown in FIG. 1 , the iterative deep unfolding feature extraction network consists of two residual structures. Each residual block is mainly composed of two direction-oriented convolution modules, where the input / output channels of component 1 and component 2 are set to [1, 64] and [64, 1] respectively, while the input / output channels of component 3 and component 4 are set to [2, 64] and [64, 2]. The calculation expression of the direction-oriented deep unfolding residual module is as follows:
[0050]
[0051] In the formula, x k is the output of the k-th oriented depth-expanded residual module, which has the format of (H×W), and y1 is x k-1 and x k The combined two-dimensional features, θ1, θ2, η1 and η2 are learnable vector parameters, and η1 and η2 obey the Gaussian distribution of N (0.1, 0.032). The iterative edge expansion feature extraction network is composed of 5 iterative deep expansion feature extraction networks. The core process of this module is to send the input image to the submodules for processing in turn, and splice the intermediate results including the input through jump connections. Finally, the 6-dimensional data is convolved 3×3 (kernl_size=3,srtide=1,in_channels=1,out_channels=1) to get the output. Let the original input image be I, and the output of the direction-oriented deep expansion residual module be x k ,[k=1,2,…,5],EURM k is the kth module, and the result of splicing is [I, x 1 、x 2 、x 3 、x 4 、x 5 ], where I and x k , k=1,…,5 are all three-dimensional feature maps, and the formula is described as follows:
[0052] f=Conv(I,x 1 ,x 2 ,…,x 5 )
[0053] x k =EURM k (x k-1 ),k=1,2,…,5
[0054] Step 3: Perform semantic segmentation and artifact detection on the enhanced infrared image. Use semantic segmentation technology to obtain the joint area of the person and the artifact. Use artifact detection to generate person mask mask1 and artifact mask mask2, where the format of person mask mask1 and artifact mask mask2 are both (H×W). The mask is a grayscale image with the same size as the corresponding image. 1 is used to represent the identified image, and the remaining part is filled with 0. Infrared feature restoration technology is used to reduce information loss caused by weather factors or occlusion. The characteristics of this technology are: introducing a learnable auxiliary context reconstruction branch / loss to encourage the generator network to use appropriate known areas as references to fill in missing areas; using an attention-free restoration generator, which is jointly trained with traditional restoration loss and auxiliary context reconstruction loss.
[0055] It is further explained that the convolution-based semantic segmentation deep learning module is used to extract the artifact area of the infrared image and remove it before feature extraction, which effectively improves the overall quality of the fused image. Specifically, it includes: manually extracting the portrait artifact area with the help of annotation tools to generate a data set, using the PaddlePaddle AI training platform to build a deep learning model, and training to obtain the semantic segmentation module.
[0056] Artifact detection is processed in two stages. In the first stage, a convolutional neural network is used to remove the background of the infrared image and extract the joint area of the portrait-artifact. In the second stage, the artifact area of the infrared image is detected and removed from the results extracted in the first stage.
[0057] Step 4: Combine the feature map, the person mask mask1 and the artifact mask mask2 to obtain a fused image, including:
[0058] Decompose complementary features in an image using a difference algorithm:
[0059]
[0060] In the formula, f i is the infrared image feature, f v is the visible light image feature, f i and f v The format is (H×W×4);
[0061] The complementary features are compressed into a vector with the help of global average pooling GAP, normalized to [0, 1] using the Sigmoid function, and the channel weights are generated. The complementary features are multiplied by the channel weights and added to the original complementary features as modal supplementary information to generate features. and The calculation formula includes:
[0062]
[0063] In the formula, and The format is (H×W×4), GAP is global average pooling, Conv is a 3×3 convolutional layer, BN is a regularization layer, and Sig is a Sigmoid loss function;
[0064] Input all features into the fusion network to obtain the fused image:
[0065] I f =Sig(BN(Conv(f s +f c )))
[0066] In the formula, f s Yesc The features obtained by three layers of convolution after semantic screening, f s and f c The format is (H×W), I f is the output fused image in the format of (H×W).
[0067] During model training, 200 pairs of infrared and visible light images from the Roadscene dataset were randomly selected as training sets, and the remaining 21 pairs of data were used as test sets. 20 pairs of data were randomly selected from the MSRS dataset for indicator detection. 3 20 pairs of images were randomly selected from the FD dataset for generalization testing. Since there is no dataset specifically for infrared artifacts, a self-made infrared artifact dataset with infrared artifact information was created. The dataset has 200 pairs of training images and 35 pairs of test images. The format of these images is 640×640×3.
[0068] During training, the parameters of the directional deep residual expansion module are set to different initial values according to the function and are independent of each other. The iterative directional expansion feature extraction network for background feature extraction randomly initializes the parameters [η1, η2] to the normal distribution N(0.1, 0.032), and [θ1, θ2] to [10 -3 ,10 -3 ]; The iterative directional expansion feature extraction network used to extract detail features randomly initializes the parameters [η1,η2] to the normal distribution N(0.1,0.032), and [θ1,θ2] to [1,1].
[0069] The network was trained for 100 epochs with a batch size of 32. The learning rate for the first 50 epochs was 10 -2 , the learning rate for the remaining cycles is reduced to 10 -3 , the training samples are randomly cropped to 128×128, and finally the trained network is obtained.
[0070] When in use, the visible light source image and the infrared source image are input into the network to obtain a fused image.
[0071] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for fusion of infrared and visible light by removing infrared pedestrian artifact feature convolution depth, characterized by: The specific steps include: Step 1: Input the visible light source image and the infrared source image into the image enhancement module to obtain an enhanced visible light image and an enhanced infrared image; Step 2: Input the visible light image, enhanced visible light image, infrared source image and enhanced infrared image into a feature extraction module to obtain a feature map; Step 3: perform semantic segmentation and artifact detection on the enhanced infrared image, obtain the joint area of the person and the artifact, and generate a person mask and an artifact mask; Step 4: Combine the feature map, person mask and artifact mask to obtain a fused image.
2. The infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth according to claim 1 is characterized in that: The method for acquiring the enhanced visible light image in step 1 includes: The method of reflection-guided histogram equalization plus comparison parameter approximation is used to process the visible light source image to obtain the enhanced visible light image.
3. The infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth according to claim 1 is characterized in that: The method for acquiring the enhanced infrared image in step 1 includes: The infrared source image is processed using the infrared method of slope histogram to obtain an enhanced infrared image.
4. The infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth according to claim 1 is characterized in that: The feature maps in step 2 include a visible light detail feature map, a visible light background feature map, an infrared light detail feature map, and an infrared light background feature map.
5. The infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth according to claim 1 is characterized in that: The feature extraction module in step 2 includes five direction-oriented depth expansion residual modules and a convolution layer 1; The number of input channels of the five directional depth expansion residual modules is 1, and the number of output channels is 1; The convolution kernel format of the convolution layer 1 is 3×3, and the step size, input channel number and output channel number are all 1.
6. The infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth according to claim 5 is characterized in that: The five direction-oriented depth unfolded residual modules each include four direction-oriented convolution modules and one convolution layer two; The activation functions of the four direction-guided convolution modules are all ReLU, and every two are divided into a group, and the number of input and output channels of the two direction-guided convolution modules in each group are set oppositely; The convolution kernel format of the convolution layer 2 is 3×3, the number of input channels is 4, and the step size and the number of output channels are both 1.
7. The infrared and visible light fusion method for removing infrared pedestrian artifact feature convolution depth according to claim 1 is characterized by: The method of obtaining the fused image in step 4 includes: Decompose complementary features in an image using a difference algorithm: In the formula, f i is the infrared image feature, f v is the visible light image feature, f i and f v The format is (H×W×4); The complementary features are compressed into a vector with the help of global average pooling GAP, normalized to [0, 1] using the Sigmoid function, and the channel weights are generated. The complementary features are multiplied by the channel weights and added to the original complementary features as modal supplementary information to generate features. and The calculation formula includes: In the formula, and The format is (H×W×4), GAP is global average pooling, Conv is a 3×3 convolutional layer, BN is a regularization layer, and Sig is a Sigmoid loss function; Input all features into the fusion network to obtain the fused image: I f =Sig(BN(Conv(f s +f c ))) In the formula, f s Yes c The features obtained by three layers of convolution after semantic screening, f s and f c The format is (H×W), I f is the output fused image in the format of (H×W).
Citation Information
Cited By
Electric power machine patrol image enhancement method and device based on semantic consistency and reinforcement learning
CN122089600A