An adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment

Through the adaptive low-light visible light and infrared image fusion method, using the fusion module of visible light encoder, infrared encoder and cross-attention mechanism, the problem of poor image fusion effect of intelligent reconnaissance equipment under low light conditions is solved, and clear target information retention and efficient analysis are achieved in low-light environments.

CN119850440BActive Publication Date: 2025-10-03GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510024412.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-10-03
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

Existing intelligent reconnaissance equipment has poor infrared-visible light image fusion effect under low-light conditions, which limits target detection and tracking performance. In addition, existing image enhancement and fusion methods are prone to cause color distortion, affecting the fusion effect.

Method used

An adaptive low-illumination visible light and infrared image fusion method is adopted. Deep features are extracted through the visible light encoder and infrared encoder, and a fusion module and decoder with a cross-attention mechanism are combined to generate a fused image with visible light color information.

Benefits of technology

It effectively retains clear and accurate target information in low-light environments, improving the accuracy and speed of analysis and decision-making of intelligent reconnaissance equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850440B_ABST
    Figure CN119850440B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology and proposes an adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment, comprising the following steps: collecting visible light images and infrared images through the intelligent reconnaissance equipment; inputting the visible light image into a trained visible light encoder to obtain deep visible light features; inputting the infrared image into a trained infrared encoder to obtain deep infrared features; inputting the deep visible light features and deep infrared features into a fusion module based on a cross-attention mechanism to obtain fused features; inputting the fused features into a decoder for decoding to obtain grayscale features; splicing the grayscale features with the visible light color channel, and then performing channel conversion to obtain a fused image with visible light color information. The present invention achieves visible light and infrared image fusion in low-light environments, effectively retaining clear and accurate target information, and helps intelligent reconnaissance equipment improve the accuracy and speed of analysis and decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to an adaptive low-illumination visible light and infrared image fusion method for intelligent reconnaissance equipment. Background Art

[0002] Intelligent reconnaissance equipment is a detection device that integrates high-performance optoelectronic devices such as visible light cameras, infrared thermal imagers, and laser rangefinders. During use, it fuses captured visible light images with infrared images, providing high-quality fused image data in various lighting conditions. However, the current fusion effect of infrared and visible light images captured by intelligent reconnaissance equipment in low-light conditions is poor, significantly limiting the equipment's target detection and tracking performance in low-light scenarios.

[0003] Currently, infrared-visible light image fusion mainly relies on infrared image information to compensate for the loss of scene information in visible light images caused by the decrease in illumination. This results in the rich scene information in nighttime visible light images being unable to be expressed in the fused image, deviating from the original intention of the image fusion task. Another intuitive solution is to use advanced low-light enhancement algorithms to pre-enhance visible light images, and then merge the source images through fusion methods. However, treating image enhancement and image fusion as separate tasks often leads to incompatibility issues. Specifically, due to the weak light of night scenes, nighttime visible light images will have slight color distortion. Applying low-light enhancement algorithms to them will change the color distribution of the light source, further amplifying the color distortion of the entire image to a certain extent, and ultimately resulting in poor image fusion effects. Summary of the Invention

[0004] In order to overcome the defect of poor image fusion effect of the above-mentioned infrared-visible light image fusion method under low-light conditions, the present invention provides an adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment.

[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:

[0006] An adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment includes the following steps:

[0007] Collect visible light and infrared images through intelligent reconnaissance equipment;

[0008] Inputting the visible light image into a trained visible light encoder to obtain deep visible light features;

[0009] Inputting the infrared image into a trained infrared encoder to obtain deep infrared features;

[0010] Inputting the deep visible light feature and the deep infrared feature into a fusion module based on a cross attention mechanism to obtain a fused feature;

[0011] Input the fused features into a decoder for decoding to obtain grayscale features;

[0012] After the grayscale features are spliced ​​with the visible light color channel, a fused image with visible light color information is obtained through channel conversion.

[0013] Furthermore, the present invention also proposes a device comprising a memory and a processor, wherein the memory stores computer-readable instructions, wherein when the computer-readable instructions are executed by the processor, the processor performs the steps of the anti-interference vehicle detection method proposed in the present invention.

[0014] Furthermore, the present invention also proposes a storage medium having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by a processor, implement the steps of the anti-interference vehicle detection method proposed in the present invention.

[0015] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0016] The present invention combines a visible light encoder, an infrared encoder, a fusion module based on a cross-attention mechanism, and a decoder to form a visible light-infrared image fusion module with a "dual-channel encoding-fusion-decoding" architecture, ensuring that the final fused image has higher contrast and richer texture details;

[0017] The present invention splices the grayscale features fused and decoded by the cross-attention mechanism with the visible light color channel to obtain a fused image with visible light color information, realizing the fusion of visible light and infrared images in low-light environments, effectively retaining clear and accurate target information, and helping intelligent reconnaissance equipment to improve the accuracy and speed of analysis and decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The figure is a flow chart of a method for adaptively fusing low-illumination visible light and infrared images according to one embodiment of the present invention.

[0019] Figure 2 FIG. 4 is an architecture diagram of a visible light-infrared image fusion module according to an embodiment of the present invention.

[0020] Figure 3 FIG. 1 is a structural diagram of a brightness adjustment unit according to an embodiment of the present invention.

[0021] Figure 4 FIG. 1 is an architecture diagram of a converged network according to an embodiment of the present invention.

[0022] Figure 5 FIG. 4 is a schematic diagram showing image fusion results of an MSRS dataset according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0024] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0026] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] Example 1

[0028] This embodiment proposes an adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment, such as Figure 1 FIG. 1 is a flow chart of the adaptive low-illumination visible light and infrared image fusion method of this embodiment.

[0029] The adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment proposed in this embodiment includes the following steps:

[0030] S1. Collect visible light and infrared images through intelligent reconnaissance equipment;

[0031] S2. Inputting the visible light image into a trained visible light encoder to obtain deep visible light features;

[0032] Inputting the infrared image into a trained infrared encoder to obtain deep infrared features;

[0033] S3. Inputting the deep visible light feature and the deep infrared feature into a fusion module based on a cross attention mechanism to obtain a fusion feature;

[0034] S4. Input the fusion feature into a decoder for decoding to obtain a grayscale feature;

[0035] S5. After splicing the grayscale features with the visible light color channel, a fused image with visible light color information is obtained through channel conversion.

[0036] This embodiment combines a visible light encoder, an infrared encoder, a fusion module based on a cross-attention mechanism, and a decoder to form a visible light-infrared image fusion module with a "dual-channel encoding-fusion-decoding" architecture. Figure 2 Figure 2 shows the architecture of the visible-infrared image fusion module of this embodiment. The visible light encoder is used to extract deep visible light features, which contain rich texture details in the visible light image, while the infrared encoder is used to extract contrast information between the target and background in the infrared image, thereby ensuring that the final fused image has higher contrast and richer texture details.

[0037] By extracting deep visible light features and deep infrared features from the image, decoding them after fusion based on the cross-attention mechanism, and splicing them with the visible light color channel, a fused image with visible light color information is obtained. This realizes the fusion of visible light and infrared images in low-light environments, effectively retaining clear and accurate target information, and helping intelligent reconnaissance equipment improve the accuracy and speed of analysis and decision-making.

[0038] In an optional embodiment, the visible light encoder includes several encoding layers consisting of a brightness adjustment unit (BAU) and a LeakyReLU activation function. The brightness adjustment unit includes a 3×3 convolutional layer and a 1×1 convolutional layer. After the input visible light features are subjected to 3×3 convolution operations and 1×1 convolution operations respectively, the extracted features are element-wise multiplied to obtain the visible light features.

[0039] For example, Figure 3 The figure shows the architecture of the brightness adjustment unit of this embodiment. In the brightness adjustment unit, the input features are subjected to 1×1 and 3×3 convolution operations with a step size of 1. When the convolution kernel is 1×1, no padding is performed, while when the convolution kernel is 3×3, the mapping padding is 1. The features obtained from the two branches are then element-wise multiplied.

[0040] In this example, by setting the brightness adjustment unit BAU, it is possible to extract deep features containing rich visible light image texture details from low-illuminance visible light images collected under adverse weather or nighttime conditions.

[0041] Exemplarily, the visible light encoder includes four groups of encoding layers consisting of brightness adjustment units and LeakyReLU activation functions.

[0042] Furthermore, in step S2, the visible light image is input into a trained visible light encoder, comprising the following steps:

[0043] Perform a balanced histogram operation on the collected visible light image, enhance the overall pixels of the original low-light visible light image, and obtain the visible light color channel CrCb and the original grayscale features. Enhanced feature information;

[0044] The original grayscale features The input is the visible light encoder, and after being processed by four sets of coding layers consisting of brightness adjustment units and LeakyReLU activation functions, the deep features F containing rich visible light image texture details are extracted. vi .

[0045] In an optional embodiment, the infrared encoder includes several groups of encoding layers consisting of 3×3 convolutional layers and LeakyReLU activation functions.

[0046] Exemplarily, the infrared encoder includes four encoding layers consisting of 3×3 convolutional layers and LeakyReLU activation functions, wherein the stride of the convolutional layers is 1 and the mapping padding is 1.

[0047] In this embodiment, the collected infrared image is preprocessed to obtain infrared image information I ir And as the input of the infrared encoder, after processing by the encoder, the deep infrared feature F rich in contrast information can be extracted ir .

[0048] In an optional embodiment, the fusion network includes a feature concatenation layer, a pooling layer and a cross-attention fusion layer.

[0049] The feature concatenation layer is used to concatenate the input deep visible light features F vi and deep infrared features F ir Perform splicing;

[0050] The pooling layer is used to perform global average pooling (GAP) on the splicing features output by the feature splicing layer to obtain a one-dimensional vector representation of the splicing features;

[0051] The cross attention fusion layer includes a first branch and a second branch; wherein the first branch performs a convolution operation and a sigmoid function activation on the input splicing feature to obtain a first activation vector v1, and combines it with the deep visible light feature F vi Multiply to obtain the first screening feature;

[0052] The second branch performs convolution operation and sigmoid function activation on the input splicing features to obtain a second activation vector v2, and combines it with the deep infrared feature F ir Multiply to obtain the second screening feature;

[0053] The first screening feature and the second screening feature are spliced ​​together to obtain a fusion feature

[0054] For example, Figure 4 FIG. 1 is a schematic diagram of the structure of the fusion network of this embodiment.

[0055] The convolution operation is used to extract mixed features and reduce channel dimensions, while the sigmoid function is used to activate and filter the corresponding key feature vectors. The cross-attention fusion operation utilizes channel attention to enhance the deep features of visible light and infrared by multiplying the corresponding elements of the two channels. The combined fusion is then performed, which helps to enhance the representation power and texture details of the fused features.

[0056] In an optional embodiment, the decoder includes several groups of decoding layers consisting of 3×3 convolutional layers and LeakyReLU activation functions.

[0057] Exemplarily, the decoder includes four decoding layers consisting of 3×3 convolutional layers and LeakyReLU activation functions, wherein the convolutional layer stride is 1 and the mapping padding is 1.

[0058] In this embodiment, the infrared encoder increases the number of image information channels from 1 to 128, while the decoder reduces the number of image information channels from 128 to 1. Therefore, the decoder can not only extract deep features but also reduce the channel dimension, restoring the image information to the dimension of a normal image.

[0059] Furthermore, in an optional embodiment, the method further comprises the following steps:

[0060] Obtain a visible light image dataset and an infrared image dataset and preprocess them to obtain a training dataset;

[0061] Initializing all weights and bias items in the visible light infrared image fusion module composed of the visible light encoder and the infrared encoder;

[0062] Construct a loss function that combines scene loss, perceptual loss, and gradient loss;

[0063] The visible light infrared image fusion module is optimized and trained based on the AdamW optimizer and the loss function, and the parameters of the visible light infrared image fusion module are saved.

[0064] For example, in this embodiment, the MSRS dataset is selected as the training dataset, wherein the images cover various scenes such as nature, city, and indoor.

[0065] Exemplarily, in this embodiment, the Kaiming initialization method is used to initialize all weights and bias items in the network.

[0066] For example, in this embodiment, the AdamW optimizer is selected for model optimization. Among them, the hyperparameters in the AdamW optimizer are set according to experience, such as momentum factors β1 and β2, initial learning rate ∈, weight decay coefficient λ and decay rate γ. The Bayesian hyperparameter optimization method is adopted, and the weighted combination of average gradient (AG), mutual information (MI) and entropy (EN) used to evaluate the fusion effect is used as the evaluation index. By repeatedly training the model, the model hyperparameters are continuously optimized according to the obtained evaluation index results, and finally a set of relatively optimal hyperparameters is obtained. The optimized hyperparameters are used to guide the training process of the model to ensure the optimization of model performance.

[0067] Furthermore, in an optional embodiment, preprocessing the visible light image dataset and the infrared image dataset includes the following steps:

[0068] Crop all images to a uniform size;

[0069] Divide the image pixel value by 255 so that the pixel value range is normalized to [0, 1];

[0070] Perform histogram equalization operation on visible light images;

[0071] All images are rotated, flipped, and randomly cropped.

[0072] For example, all images are cropped to a uniform size of 256×256.

[0073] Furthermore, in an optional embodiment, in the loss function of constructing the joint scene loss, perceptual loss and gradient loss, the scene loss includes the scene loss in the encoding stage. and scene loss in the feature fusion stage Its expression is:

[0074]

[0075] Among them, α1 and α2 are hyperparameters used to adjust the order of magnitude, L1(·) is the L1 norm loss function, that is, the absolute value deviation of the corresponding position element in the vector; F vi Represents the visible light characteristics, F ir Indicates infrared characteristics; Represents grayscale features; Indicates fusion features; I ir Indicates the infrared image information extracted from the input infrared image; max(·,·) is used to select the maximum value from the pixel points;

[0076] The perceptual loss l perce Calculated by the trained VGG19 network; its expression is:

[0077]

[0078] Among them, L vgg (·) represents the VGG19 network, which is used to calculate the loss value between the generated features and the target features;

[0079] The gradient loss l grad Including the gradient loss in the encoding stage and the gradient loss in the feature fusion stage Its expression is:

[0080]

[0081] Among them, β1 and β2 are hyperparameters; is the Sobel operator;

[0082] Then, the loss function l join The expression is:

[0083] l join =λ1·l scene +λ2·l perce +λ3·l grad

[0084] Among them, λ1, λ2, and λ3 are hyperparameters used to adjust the scene loss, perceptual loss, and gradient loss to the same order of magnitude.

[0085] For example, the training of the visible light infrared image fusion module is divided into two stages, the first stage is the encoding stage, and the second stage is the fusion network and decoder stage. The two training stages are based on the scene loss l scene and perceptual loss l perce The features generated by the encoder are adaptively adjusted to normal brightness, and the gradient loss l grad The edge information is extracted, forcing the edge texture to approach the original image, so that the generated image has more texture information.

[0086] In this embodiment, the loss function is constructed by combining scene loss, perceptual loss and gradient loss, which can guide the model to obtain reasonable lighting components, rich scene content and texture details from the fused features.

[0087] As an example, in order to visually observe the fusion performance of different algorithms on the MSRS dataset, this embodiment selects a pair of infrared and visible light images, and the image fusion results are as follows: Figure 5 shown.

[0088] Among them, due to the illumination degradation problem of night images, an excellent night fusion algorithm should gather meaningful information from the source images and provide a bright scene with high contrast. Figure 5 It can be seen that in the image fusion effects based on the GANMcC and TarDAL algorithms, the entire scene is submerged in darkness; the fusion results based on algorithms such as Gan-FM, SwinFusion, UMF-CMGR and DIVFusion are relatively blurred and a large amount of edge detail information is lost. Although SeAFusion, STDFusion and SuperFuse have more texture information, the brightness of the entire scene is low. Obviously, the existing image fusion methods cannot provide visual experience that meets the user's needs. The image fusion results of this embodiment can obtain bright scenes and rich textures even in the dark. It is worth emphasizing that none of the nine SOTA methods can display rich details in dark areas, while this embodiment can easily capture texture information without considering dark areas.

[0089] Furthermore, in terms of quantitative evaluation, this embodiment selects six indicators, including average gradient (AG), mutual information (MI), entropy (EN), spatial frequency (SF), standard deviation (SD), and visual information fidelity (VIF), to objectively evaluate the fusion effect. Among them, AG reflects the richness of image texture information, MI reflects the amount of mutual information, EN evaluates the amount of information contained in the fused image from the perspective of information theory, SF reflects the spatial frequency information contained in the fused image, SD statistically represents the distribution and contrast of the fused image, and VIF measures the fidelity of information from the perspective of human visual perception. The larger the value of the above six indicators, the better the fusion effect. This embodiment selects 30 pairs of images from the MSRS dataset to objectively evaluate the SD, MI, VIF, AG, EN and SF indicators. The fusion result indicator values ​​are shown in Table 1 below.

[0090] Table 1 Six index values ​​of the fusion results of 10 methods on the MSRS dataset

[0091]

[0092] As shown in the table above, the method proposed in this embodiment ranks in the top two across six metrics compared to other existing algorithms. The best results for MI, VIF, and SF indicate that the method's results have richer texture details, more scene information, and better visual perception, respectively. Furthermore, this embodiment also performs well in terms of SD, AG, and EN metrics, demonstrating that, with the help of the loss function and brightness adjustment module, the fusion results of this embodiment are more consistent with the source image and appear more natural.

[0093] This embodiment significantly improves the fusion of visible and infrared images at night or in low-light environments, promoting the development of advanced visual tasks such as target detection and tracking. This means that intelligent reconnaissance equipment can still provide clear and accurate target information at night or in low-light conditions, improving the efficiency of all-weather monitoring.

[0094] Example 2

[0095] This embodiment proposes a device including a memory and a processor, wherein the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes all or part of the steps of the adaptive low-illumination visible light and infrared image fusion method proposed in Example 1.

[0096] Example 3

[0097] This embodiment proposes a storage medium having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by a processor, implement all or part of the steps of the adaptive low-illumination visible light and infrared image fusion method proposed in Example 1.

[0098] Exemplarily, the storage medium includes but is not limited to a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0099] Exemplarily, the instructions, programs, code sets or instruction sets may be implemented using conventional programming languages.

[0100] Exemplarily, the processor includes but is not limited to a smart phone, a personal computer, a server, a network device, etc., and is used to execute all or part of the steps of the adaptive low-illumination visible light and infrared image fusion method described in Example 1.

[0101] The terms in the drawings are for illustrative purposes only and are not to be construed as limiting the present invention;

[0102] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. An adaptive low-light visible light and infrared image fusion method for intelligent reconnaissance equipment, characterized in that: The following steps are involved: Collect visible light and infrared images through intelligent reconnaissance equipment; Inputting the visible light image into a trained visible light encoder to obtain deep visible light features; Inputting the infrared image into a trained infrared encoder to obtain deep infrared features; Inputting the deep visible light feature and the deep infrared feature into a fusion module based on a cross attention mechanism to obtain a fused feature; Input the fused features into a decoder for decoding to obtain grayscale features; After splicing the grayscale features with the visible light color channel, a fused image with visible light color information is obtained through channel conversion; The fusion module includes a feature splicing layer, a pooling layer and a cross-attention fusion layer; wherein: The feature stitching layer is used to stitch the input deep visible light features and deep infrared features; The pooling layer is used to perform maximum average pooling on the splicing features output by the feature splicing layer to obtain a one-dimensional vector representation of the splicing features; The cross-attention fusion layer includes a first branch and a second branch; wherein the first branch performs a convolution operation and a sigmoid function activation on the input splicing feature to obtain a first activation vector, and multiplies the first activation vector with the deep visible light feature to obtain a first screening feature; The second branch performs a convolution operation and a sigmoid function activation on the input splicing feature to obtain a second activation vector, and multiplies the second activation vector with the deep infrared feature to obtain a second screening feature; The first screening feature and the second screening feature are spliced ​​together to obtain a fusion feature.

2. The adaptive low-illumination visible light and infrared image fusion method according to claim 1, characterized in that: The visible light encoder includes several groups of encoding layers consisting of brightness adjustment units and LeakyReLU activation functions; The brightness adjustment unit includes a 3×3 convolution layer and a 1×1 convolution layer. After the input visible light features undergo 3×3 convolution operations and 1×1 convolution operations respectively, the extracted features are element-wise multiplied to obtain visible light features.

3. The adaptive low-illumination visible light and infrared image fusion method according to claim 1, characterized in that: The infrared encoder includes several groups of encoding layers consisting of 3×3 convolutional layers and LeakyReLU activation functions.

4. The adaptive low-illumination visible light and infrared image fusion method according to claim 1, characterized in that: The decoder includes several sets of decoding layers consisting of 3×3 convolutional layers and LeakyReLU activation functions.

5. The adaptive low-illumination visible light and infrared image fusion method according to any one of claims 1 to 4, characterized in that: The method further comprises the following steps: Obtain a visible light image dataset and an infrared image dataset and preprocess them to obtain a training dataset; Initializing all weights and bias items in the visible light infrared image fusion module composed of the visible light encoder and the infrared encoder; Construct a loss function that combines scene loss, perceptual loss, and gradient loss; The visible light infrared image fusion module is optimized and trained based on the AdamW optimizer and the loss function, and the parameters of the visible light infrared image fusion module are saved.

6. The adaptive low-illumination visible light and infrared image fusion method according to claim 5, characterized in that: Preprocessing the visible light image dataset and the infrared image dataset includes the following steps: Crop all images to a uniform size; Divide the image pixel value by 255 so that the pixel value range is normalized to [0, 1]; Perform histogram equalization operation on visible light images; All images are rotated, flipped, and randomly cropped.

7. The adaptive low-illumination visible light and infrared image fusion method according to claim 5, characterized in that: In the loss function of constructing the joint scene loss, perceptual loss and gradient loss, the scene loss includes the scene loss in the encoding stage and scene loss in the feature fusion stage , whose expression is: in, and is a hyperparameter used to adjust the order of magnitude, is the L1 norm loss function, that is, the absolute value deviation of the corresponding position element in the vector; Represents the visible light characteristics, Indicates infrared characteristics; Represents grayscale features; represents fusion features; Indicates the infrared image information extracted from the input infrared image; Used to select the maximum value from the pixels; The perceptual loss Calculated by the trained VGG19 network; its expression is: in, Represents the VGG19 network, which is used to calculate the loss value between the generated features and the target features; The gradient loss Including the gradient loss in the encoding stage and the gradient loss in the feature fusion stage , whose expression is: in, and is a hyperparameter; is the Sobel operator; Then, the loss function The expression is: in, 、 and is a hyperparameter used to adjust the scene loss, perceptual loss, and gradient loss to the same order of magnitude.

8. A device comprising a memory and a processor, wherein the memory stores computer-readable instructions, wherein: When the computer-readable instructions are executed by the processor, the processor performs all or part of the steps of the adaptive low-illumination visible light and infrared image fusion method according to any one of claims 1 to 7.

9. A storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by a processor, all or part of the steps of the adaptive low-illumination visible light and infrared image fusion method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Infrared visible light fusion recognition system and method based on modal difference feature guidance

    CN114898189A

  • Infrared-visible light fusion method and system based on feature enhancement and readable storage medium

    CN116258934A