A method for fusion of RGB and near-infrared images under weak light conditions for intelligent reconnaissance equipment
By combining deep structure extraction and multi-scale fusion modules, the problem of fusion of RGB and near-infrared images in intelligent reconnaissance equipment under low-light conditions is solved, higher quality and more stable image fusion effects are achieved, and the imaging and recognition performance of intelligent reconnaissance equipment is improved.
Patent Information
- Application Number
- CN202510145162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing RGB and near-infrared image fusion algorithms in intelligent reconnaissance equipment have problems such as image quality degradation, difficulty in matching and fusion, and poor robustness, especially in low-light conditions.
The deep structure extraction module is used to extract the deep structural features of RGB and near-infrared images. The inconsistency of feature information is calculated by combining the depth inconsistency prior algorithm. The multi-scale deep fusion module is used to perform image fusion and output a structurally consistent fusion result.
More robust target RGB and near-infrared image fusion is achieved under low-light conditions, which improves image quality and stability and enhances the imaging and recognition capabilities of intelligent reconnaissance equipment.
Smart Images

Figure CN120070204B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and more specifically, to a method for fusing RGB and near-infrared images under weak light conditions of intelligent reconnaissance equipment. Background Art
[0002] With the widespread use of drones and other types of aircraft, intelligent reconnaissance equipment has become a crucial payload component. Intelligent reconnaissance equipment integrates multiple sensors, such as visible light cameras and infrared thermal imagers, to perform aerial target search and identification. However, single-band signal imaging often suffers from incomplete target information and limited recognition accuracy. Visible light imaging can capture comprehensive information about a target, including color and detail. However, this performance decreases when ambient lighting conditions deteriorate. Infrared thermal imagers can sense a target's thermal radiation and capture images at night or in low-light conditions, but this can lose information such as color and texture. Direct information-level fusion struggles to address the mismatch in signal distribution.
[0003] When intelligent reconnaissance equipment performs missions, especially in poor lighting conditions at night, the quality of RGB and near-infrared images often degrades due to noise and interference. This complicates the matching and fusion of the two, negatively impacting the effectiveness of multimodal imaging. Therefore, it is necessary to develop a robust matching and fusion method to achieve high-quality multimodal imaging and provide effective support for subsequent intelligent analysis and decision-making.
[0004] At present, RGB and infrared low-light image fusion algorithms directly targeting intelligent reconnaissance equipment are still relatively rare. The main existing technical routes include: (1) Simple pixel-level weighted superposition: This method has the problem of structural mismatch between different modalities and is prone to image distortion. (2) Fusion based on traditional image processing algorithms: For example, image enhancement is performed first and then superposition; this type of method has a strong dependence on source data and poor robustness. (3) Fusion based on shallow machine learning models: This type of method has difficulty in modeling complex nonlinear mapping relationships and has the problem of insufficient generalization ability. (4) Deep learning fusion model trained with a general data set: Since it cannot reflect the unique perception characteristics of the electric pod, the effect is poor. (5) Two-stage method, first restore the image of each modality separately and then fuse it; but this method also has difficulties in the single image restoration process. Summary of the Invention
[0005] In order to overcome the defects of the prior art described above, the present invention provides a method for fusing RGB and near-infrared images under weak light conditions of intelligent reconnaissance equipment.
[0006] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0007] A method for fusing RGB and near-infrared images under low-light conditions for intelligent reconnaissance equipment includes:
[0008] Use intelligent reconnaissance equipment to collect RGB images and near-infrared images under low-light conditions and perform preprocessing;
[0009] The preprocessed RGB image and near-infrared image are input into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image respectively;
[0010] Inputting the structural features of the RGB image and the near-infrared image into a deep structure extraction module to obtain deep structure feature maps of the RGB image and the near-infrared image;
[0011] Performing a deep structure inconsistency prior calculation on the deep structure feature maps of the RGB image and the near-infrared image to obtain a deep structure inconsistency prior feature map, and then performing a product operation on the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure;
[0012] The near-infrared feature map and RGB image with consistent structure are input into a multi-scale deep fusion module to obtain a fused image.
[0013] Furthermore, the present invention also proposes a system for fusion of RGB and near-infrared images under low-light conditions for intelligent reconnaissance equipment, which applies the method for fusion of RGB and near-infrared images under low-light conditions proposed by the present invention. The system includes:
[0014] Image acquisition module, including intelligent reconnaissance equipment, used to collect RGB images and near-infrared images under low-light conditions and perform pre-processing;
[0015] The restoration module is equipped with an RGB restoration network and a near-infrared restoration network, which are used to input the preprocessed RGB image and the near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image respectively;
[0016] The deep structure extraction module is used to extract the deep structure of the structural features of the input RGB image and near-infrared image to obtain the deep structure feature maps of the RGB image and near-infrared image;
[0017] a structural consistency generation module for performing a deep structural inconsistency prior calculation on the deep structural feature maps of the RGB image and the near-infrared image to obtain a deep structural inconsistency prior feature map; and then performing a product operation on the deep structural inconsistency prior feature map and the deep structural feature map of the near-infrared image to output a structurally consistent near-infrared feature map;
[0018] The multi-scale deep fusion module is used to deeply fuse the input near-infrared feature map and RGB image with consistent structure to obtain a fused image.
[0019] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0020] The present invention extracts deep structure feature maps from RGB and near-infrared images through a deep structure extraction module, introduces prior knowledge at the same time, uses a deep inconsistency prior algorithm to calculate the inconsistency of deep structure feature information in RGB and near-infrared images, multiplies the deep structure inconsistency prior feature map and the deep structure feature map of the near-infrared image, suppresses the part inconsistent with the RGB deep structure, and outputs a near-infrared feature map with consistent structure; the near-infrared feature map with consistent structure and the RGB feature map are input into a multi-scale deep fusion module, and the final fusion result map is output, so that intelligent reconnaissance equipment can achieve more robust target RGB and near-infrared image fusion in complex environments such as low-light conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 The figure is a flow chart of a method for fusing RGB and near-infrared images according to one embodiment of the present invention.
[0022] Figure 2 The figure is a schematic diagram of the framework of a method for fusing RGB and near-infrared images according to one embodiment of the present invention.
[0023] Figure 3 FIG. 4 is an architecture diagram of a deep structure extraction module according to an embodiment of the present invention.
[0024] Figure 4 FIG. 1 is a diagram illustrating an architecture of an autoencoder according to an embodiment of the present invention.
[0025] Figure 5 FIG. 4 is an architecture diagram of a multi-scale deep fusion module according to an embodiment of the present invention.
[0026] Figure 6 FIG. 4 is an architecture diagram of an RGB and near-infrared image fusion system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0027] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.
[0028] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."
[0030] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] Example 1
[0032] This embodiment proposes a method for fusing RGB and near-infrared images under weak light conditions for intelligent reconnaissance equipment. Figure 1 、 2 , which is a flow chart and framework diagram of the RGB and near-infrared image fusion method of this embodiment.
[0033] The method for fusing RGB and near-infrared images under low-light conditions for intelligent reconnaissance equipment proposed in this embodiment includes the following steps:
[0034] S1. Use intelligent reconnaissance equipment to collect RGB images and near-infrared images under low-light conditions and perform preprocessing;
[0035] S2. Input the preprocessed RGB image and near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image respectively;
[0036] S3, inputting the structural features of the RGB image and the near-infrared image into a deep structure extraction module to obtain deep structure feature maps of the RGB image and the near-infrared image;
[0037] S4, performing a deep structure inconsistency prior calculation on the deep structure feature maps of the RGB image and the near-infrared image to obtain a deep structure inconsistency prior feature map, and then performing a product operation on the deep structure feature map with the near-infrared image to output a near-infrared feature map with consistent structure;
[0038] S5. Inputting the near-infrared feature map and the RGB image with consistent structure into a multi-scale deep fusion module to obtain a fused image.
[0039] In this embodiment, the photoelectric detection equipment in the intelligent reconnaissance equipment is used to collect RGB and near-infrared images of the target, and the deep structure feature maps in the RGB and near-infrared images are extracted by the deep structure extraction module. At the same time, prior knowledge is introduced, and the inconsistency of the deep structure feature information in the RGB and near-infrared images is calculated using the deep inconsistency prior algorithm, and a deep structure inconsistency prior feature map is output; the deep structure inconsistency prior feature map and the deep structure feature map of the near-infrared image are multiplied to suppress the part inconsistent with the RGB deep structure, and a near-infrared feature map with consistent structure is output; the near-infrared feature map with consistent structure and the RGB feature map are input into the multi-scale deep fusion module, and the final fusion result map is output, so that the intelligent reconnaissance equipment can achieve more robust target RGB and near-infrared image fusion in complex environments such as low-light conditions.
[0040] In an optional embodiment, the method further comprises the following steps:
[0041] Intelligent reconnaissance equipment is used to collect reference image pairs in normal light and low light environments, and they are preprocessed to obtain training data for pre-training the RGB and near-infrared image fusion model composed of an RGB restoration network, a near-infrared restoration network, a deep structure extraction module, and a multi-scale deep fusion module.
[0042] For example, in a specific implementation process, in order to make the fused image closer to the real image, synthetic noise is added to the reference image pair for training and testing, and when pre-training the image fusion model, 5k pairs of reference images (256*256) are used as the training set, and another 1k reference image pair (256*256) and 10 additional real noise image pairs (1920*1080) are used for testing.
[0043] It should be noted that the image pair mentioned in this embodiment includes an RGB image and a near-infrared image captured at the same time and angle.
[0044] Furthermore, in an optional embodiment, the step of preprocessing the acquired image includes:
[0045] (1) Perform noise removal on the collected images;
[0046] (2) Align the RGB and NIR image pairs using a matching algorithm with manual double checking;
[0047] (3) The image is collected into image blocks of a preset size.
[0048] Optionally, in order to obtain a high-quality reference, for image pairs collected from normal light environments, multiple static captures are averaged to remove noise, and a matching algorithm with manual double checking is used to ensure the alignment of the image pairs. The collected images are cropped into 256*256 image blocks.
[0049] For a real noisy image pair collected from a low-light environment, the preprocessing steps are the same as the steps for collecting the reference image pair, and also include not performing multi-frame averaging on the noisy RGB image.
[0050] In an optional embodiment, in step S3, the deep structure extraction module includes several convolutional layers, and the output of each convolutional layer is summed and activated by an activation function to output a multi-scale deep structure feature map; wherein, the convolutional layer for outputting the structural features of one scale is connected to a residual channel attention block for performing summation operations with the structural features of other scales.
[0051] For example, Figure 3 The figure shows the architecture of the deep structure extraction module of this embodiment. In this embodiment, the structural features of one scale are passed through a residual channel attention block, and the output is sequentially summed with the structural features of other scales to obtain a multi-scale deep structure feature map.
[0052] The deep structure extraction module in this embodiment can enable intelligent reconnaissance equipment to better analyze and represent the intrinsic structure of the image, and can extract effective image structure as the basis for subsequent processing even in extremely dim lighting environments.
[0053] Furthermore, in an optional embodiment, the deep structure extraction module includes the following steps during pre-training:
[0054] After the training data is input into the RGB restoration network and the near-infrared restoration network, the output structural features are input into the deep structure extraction module for training; wherein the loss function of the deep structure extraction module is:
[0055]
[0056]
[0057] Among them, N i Indicates the number of channels of the deep structure of the i-th scale; str i,c Represents the predicted structure graph of the cth channel of the deep structure of the i-th scale; The deep structure representing the true value of the cth channel of the deep structure of the i-th scale; y d and z dis the pixel value of the dth pixel in the deep structure of the predicted structure graph and the true value, and N is the number of image pixels.
[0058] It should be noted that it is very difficult for the deep structure extraction module to generate a noise-free structure graph alone, so this embodiment uses a noise-free true value as the supervisory signal. Optionally, the supervisory signal is generated by an encoder.
[0059] Further, in an optional embodiment, the real structure feature map is generated by an autoencoder; the encoder includes a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block, and a sixth convolution block connected in sequence, wherein the outputs of the first convolution block, the fourth convolution block, and the fifth convolution block are respectively connected to a residual channel attention block; the input ends of the second convolution block and the third convolution block are respectively connected to a residual channel attention block; at least two residual channel attention blocks are connected between the third convolution block and the fourth convolution block;
[0060] The second convolution block and the third convolution block include a downsampling layer and a convolution layer; the fourth convolution block and the fifth convolution block include an upsampling layer and a convolution layer;
[0061] The input feature map of the fourth convolution block, the output feature map of the fourth convolution block, and the output feature map of the fifth convolution block are taken as the supervision signal of the deep structure extraction module to generate a structure map of the real value.
[0062] For example, Figure 4 , which is a diagram showing the architecture of the autoencoder of this embodiment.
[0063] This embodiment extracts multi-scale decoder deep features from the autoencoder i,c , and calculate the supervisory signal:
[0064]
[0065] in, express The nth pixel of is the sobel operator, express The nth pixel in ; express After subtracting the two, we use 0 as the threshold and set the corresponding pixel value in the structure map to 0 for pixels less than or equal to 0, and 1 for pixels otherwise, to obtain the deep structure of the final true value.
[0066] In this embodiment, the supervisory signal is generated by the autoencoder, which makes the structural feature extraction more accurate, provides a guarantee for stable imaging under complex lighting changes, and further improves the robustness of imaging and recognition of intelligent reconnaissance equipment.
[0067] In an optional embodiment, performing a deep structure inconsistency prior calculation on the deep structure feature maps of the RGB image and the near-infrared image includes the following steps:
[0068] According to the deep structure feature maps of RGB images and near-infrared images, a binary edge map is extracted from each feature channel and its inconsistency is calculated to obtain a deep structure inconsistency prior feature map; its expression is:
[0069]
[0070] in, and represents the cth channel of the deep structural features of the i-th scale in the noisy RGB image and the near-infrared image respectively; β∈(0,1) is a hyperparameter indicating that there is no significant inconsistency; U i,c Represents the eigenvalue of the cth channel of the deep structure feature of the i-th scale in the deep structure inconsistent prior feature map;
[0071] The deep structure is inconsistent with the prior feature map U i,c Perform a product operation with the deep structure feature map of the near-infrared image, discard inconsistent structures in the image, update the corresponding deep feature values, and obtain a near-infrared feature map with consistent structure; its expression is:
[0072]
[0073] in, Represents the feature value of the cth channel of the deep structure feature of the i-th scale in the deep structure feature map of the updated near-infrared image.
[0074] In this embodiment, we introduce the prior knowledge of deep structure inconsistency and design the deep structure inconsistency feature U by combining the multi-scale deep structure feature map. i,c , and further multiply the result with the original structure to discard the deep structure feature map of the near-infrared image The structure in the image is inconsistent with that in the RGB image.
[0075] Exemplarily, the hyperparameter β in this embodiment is set to 0.5.
[0076] In this embodiment, by using a depth inconsistency prior algorithm, intelligent reconnaissance equipment can detect structural and information differences between RGB and near-infrared images, further guiding subsequent fusion and optimization. This enables a more coordinated and unified effect of multi-source heterogeneous images even in low-light environments.
[0077] In an optional embodiment, the structurally consistent near-infrared feature map and RGB image are input into a multi-scale deep fusion module to obtain a fused image, wherein the multi-scale deep fusion module includes a convolutional layer based on residual connections, an upsampling layer, and a downsampling layer.
[0078] The RGB image input into the multi-scale deep fusion module is processed by convolution and residual channel attention to obtain a coarse RGB structural feature map after noise reduction;
[0079] The near-infrared feature map with consistent structure input to the multi-scale deep fusion module is subjected to convolution processing, and then subjected to a residual connection-based addition operation with the coarse RGB structural feature map after noise reduction. After downsampling and upsampling processing, a fused image is generated through a convolution layer and an activation function.
[0080] For example, Figure 5 , which is an architecture diagram of the multi-scale deep fusion module of this embodiment.
[0081] Furthermore, in an optional embodiment, during the pre-training process, the multi-scale deep fusion module uses the training data as model input and calculates the total loss function of the model based on the fused image output by the model; the total loss function is expressed as:
[0082]
[0083] in, and is the loss function for RGB and near-infrared deep structure prediction, where α and γ are preset coefficients; and are the reconstruction losses of the fused image, coarse RGB, and near-infrared image, respectively, and are Charbonnier losses, which are in the form of:
[0084]
[0085] Among them, X and X z denote the network output and the corresponding true value supervision respectively, and θ is a hyperparameter.
[0086] Exemplarily, θ is set to 1 / 1000.
[0087] The multi-scale deep fusion module in this embodiment effectively integrates multi-level information such as image structure and text, constructing a richer and more complete image representation under low-light conditions, which can significantly improve the performance of subsequent intelligent analysis such as target recognition and classification.
[0088] The method proposed in this embodiment was repeatedly applied to image acquisition and processing experiments under mobile conditions of intelligent reconnaissance equipment. The results showed that the output image quality, stability and other indicators of the present invention were superior to the simple fusion method of direct superposition, and the robustness of the method of this embodiment was verified. It was also verified that intelligent reconnaissance equipment can achieve more robust target RGB and near-infrared image fusion in complex environments.
[0089] Example 2
[0090] This embodiment proposes a RGB and near-infrared image fusion system under weak light conditions for intelligent reconnaissance equipment, which applies the RGB and near-infrared image fusion method under weak light conditions proposed in Example 1. Figure 6 FIG. 1 is an architecture diagram of the RGB and near-infrared image fusion system under weak light conditions of this embodiment.
[0091] The RGB and near-infrared image fusion system for intelligent reconnaissance equipment under low-light conditions proposed in this embodiment includes:
[0092] Image acquisition module, including intelligent reconnaissance equipment, used to collect RGB images and near-infrared images under low-light conditions and perform pre-processing;
[0093] The restoration module is equipped with an RGB restoration network and a near-infrared restoration network, which are used to input the preprocessed RGB image and the near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image respectively;
[0094] The deep structure extraction module is used to extract the deep structure of the structural features of the input RGB image and near-infrared image to obtain the deep structure feature maps of the RGB image and near-infrared image;
[0095] a structural consistency generation module for performing a deep structural inconsistency prior calculation on the deep structural feature maps of the RGB image and the near-infrared image to obtain a deep structural inconsistency prior feature map; and then performing a product operation on the deep structural inconsistency prior feature map and the deep structural feature map of the near-infrared image to output a structurally consistent near-infrared feature map;
[0096] The multi-scale deep fusion module is used to deeply fuse the input near-infrared feature map and RGB image with consistent structure to obtain a fused image.
[0097] It can be understood that the system of this embodiment corresponds to the method of the above-mentioned embodiment 1, and the options in the above-mentioned embodiment 1 are also applicable to this embodiment, so they will not be described again here.
[0098] The terms in the drawings are for illustrative purposes only and are not to be construed as limiting the present invention;
[0099] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A method for fusing RGB and near-infrared images under low-light conditions for intelligent reconnaissance equipment, characterized in that: The following steps are involved: Use intelligent reconnaissance equipment to collect RGB images and near-infrared images under low-light conditions and perform preprocessing; The preprocessed RGB image and near-infrared image are input into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image respectively; Inputting the structural features of the RGB image and the near-infrared image into a deep structure extraction module to obtain deep structure feature maps of the RGB image and the near-infrared image; Performing a deep structure inconsistency prior calculation on the deep structure feature maps of the RGB image and the near-infrared image to obtain a deep structure inconsistency prior feature map, and then performing a product operation on the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure; The near-infrared feature map and RGB image with consistent structure are input into a multi-scale deep fusion module to obtain a fused image.
2. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 1, characterized in that: The deep structure extraction module includes several convolutional layers. The output of each convolutional layer is summed and then activated by an activation function to output a multi-scale deep structure feature map. Among them, the convolutional layer used to output the structural features of one scale is connected to a residual channel attention block for summing the structural features of other scales.
3. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 1, characterized in that: The deep structure inconsistency prior calculation of the deep structure feature maps of the RGB image and the near-infrared image comprises the following steps: According to the deep structure feature maps of RGB images and near-infrared images, a binary edge map is extracted from each feature channel and its inconsistency is calculated to obtain a deep structure inconsistency prior feature map; its expression is: in, and represents the cth channel of the deep structure feature of the i-th scale in the noisy RGB image and the near-infrared image respectively; β∈(0,1) is a hyperparameter; U i,c Represents the eigenvalue of the cth channel of the deep structure feature of the i-th scale in the deep structure inconsistent prior feature map; The deep structure is inconsistent with the prior feature map U i,c Perform a product operation with the deep structure feature map of the near-infrared image, discard inconsistent structures in the image, update the corresponding deep feature values, and obtain a near-infrared feature map with consistent structure; its expression is: in, Represents the feature value of the cth channel of the deep structure feature of the i-th scale in the deep structure feature map of the updated near-infrared image.
4. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 1, characterized in that: The multi-scale deep fusion module includes a convolutional layer based on residual connection, an upsampling layer and a downsampling layer; The RGB image input into the multi-scale deep fusion module is processed by convolution and residual channel attention to obtain a coarse RGB structural feature map after noise reduction; The near-infrared feature map with consistent structure input to the multi-scale deep fusion module is subjected to convolution processing, and then subjected to a residual connection-based addition operation with the coarse RGB structural feature map after noise reduction. After downsampling and upsampling processing, a fused image is generated through a convolution layer and an activation function.
5. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 1, characterized in that: The steps for preprocessing the acquired images include: Perform noise removal on the collected images; The RGB and NIR image pairs were aligned using a matching algorithm with manual double checking; Capture images into blocks of a preset size.
6. The method for fusion of RGB and near-infrared images under low-light conditions according to any one of claims 1 to 5, characterized in that: The method further comprises the following steps: Intelligent reconnaissance equipment is used to collect RGB images and near-infrared images in normal light and low light environments and preprocess them to obtain training data for pre-training the RGB and near-infrared image fusion model composed of an RGB restoration network, a near-infrared restoration network, a deep structure extraction module, and a multi-scale deep fusion module.
7. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 6, characterized in that: Pre-training the RGB and near-infrared image fusion model includes the following steps: The training data is used as the model input, and the total loss function of the model is calculated based on the fused image output by the model; its expression is: in, and is the loss function for RGB and near-infrared deep structure prediction, where α and γ are preset coefficients; and are the reconstruction losses of the fused image, coarse RGB, and near-infrared image, respectively, and are Charbonnier losses.
8. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 7, characterized in that: The method further comprises the following steps: After the training data is input into the RGB restoration network and the near-infrared restoration network, the output structural features are input into the deep structure extraction module for training; wherein the loss function of the deep structure extraction module is: Among them, N i Indicates the number of channels of the deep structure of the i-th scale; str i,c Represents the predicted structure graph of the cth channel of the deep structure of the i-th scale; The deep structure representing the true value of the cth channel of the deep structure of the i-th scale; y d and z d is the pixel value of the dth pixel in the deep structure of the predicted structure graph and the true value, and N is the number of image pixels.
9. The method for fusion of RGB and near-infrared images under low-light conditions according to claim 8, characterized in that: The real value structure feature map is generated by an autoencoder; The encoder includes a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block, and a sixth convolution block connected in sequence, wherein the outputs of the first convolution block, the fourth convolution block, and the fifth convolution block are respectively connected to a residual channel attention block; the input ends of the second convolution block and the third convolution block are respectively connected to a residual channel attention block; at least two residual channel attention blocks are connected between the third convolution block and the fourth convolution block; The second convolution block and the third convolution block include a downsampling layer and a convolution layer; the fourth convolution block and the fifth convolution block include an upsampling layer and a convolution layer; The input feature map of the fourth convolution block, the output feature map of the fourth convolution block, and the output feature map of the fifth convolution block are taken as the supervision signal of the deep structure extraction module to generate a structure map of the real value.
10. A system for fusion of RGB and near-infrared images under low-light conditions for intelligent reconnaissance equipment, applying the method for fusion of RGB and near-infrared images under low-light conditions according to any one of claims 1 to 9, characterized in that: include: Image acquisition module, including intelligent reconnaissance equipment, used to collect RGB images and near-infrared images under low-light conditions and perform pre-processing; The restoration module is equipped with an RGB restoration network and a near-infrared restoration network, which are used to input the preprocessed RGB image and the near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image respectively; The deep structure extraction module is used to extract the deep structure of the structural features of the input RGB image and near-infrared image to obtain the deep structure feature maps of the RGB image and near-infrared image; The structure consistency generation module is used to perform deep structure inconsistency prior calculation on the deep structure feature maps of RGB images and near-infrared images to obtain deep structure inconsistency prior feature maps; Then, a product operation is performed on the deep structure inconsistent prior feature map and the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure; The multi-scale deep fusion module is used to deeply fuse the input near-infrared feature map and RGB image with consistent structure to obtain a fused image.
Citation Information
Patent Citations
Infrared and visible light fusion imaging method based on deep learning
CN113487530A
Infrared and visible light image fusion method with regional attention
CN114782298A