RGB and near-infrared image fusion method of intelligent reconnaissance equipment under weak light condition
The deep structure extraction module and depth inconsistent prior algorithm process RGB and near-infrared images of intelligent reconnaissance equipment under low light conditions are solved, and the problem of image quality reduction and fusion difficulty is achieved, achieving a more robust image fusion effect.
Patent Information
- Application Number
- CN202510145162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In low-light conditions, the quality of RGB images and near-infrared images decreases, resulting in increased matching and fusion difficulty, affecting the multimodal imaging effect.
The deep structure extraction module is used to extract the deep structure feature maps of RGB and near-infrared images, and the inconsistency of feature information is calculated through the depth inconsistency prior algorithm, and the near-infrared feature map with a consistent structure is output, and then it is fused with the RGB image input multi-scale deep fusion module.
It realizes more robust RGB and near-infrared image fusion under low-light conditions, improving the imaging quality and recognition accuracy of intelligent reconnaissance equipment.
Smart Images

Figure CN120070204A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, to a method for fusing RGB and near-infrared images under low-light conditions for an intelligent reconnaissance equipment. Background Art
[0002] With the wide use of unmanned aerial vehicles and various aircraft, intelligent reconnaissance equipment has become one of the important load components. The intelligent reconnaissance equipment integrates various sensors such as visible light cameras and infrared thermal imagers to achieve functions such as aerial target search and identification. However, the signal mapping of a single band often has problems such as incomplete target information and limited recognition accuracy. Visible light cameras can obtain all-round information such as the color and details of the target. However, when the ambient light conditions deteriorate, the visible light imaging effect decreases. Infrared thermal imagers can sense the thermal radiation of the target and obtain images under night or low-illumination conditions, but will lose information such as color and texture. Direct information-level fusion is difficult to solve the problem of mismatched signal distributions.
[0003] When the intelligent reconnaissance equipment is performing tasks, especially under poor night lighting conditions, due to the influence of noise and interference, the quality of RGB images and near-infrared images often deteriorates. This increases the difficulty of matching and fusing between the two, and has an adverse impact on the effect of multi-modal imaging. Therefore, it is necessary to study a robust matching and fusion method to achieve high-quality multi-modal imaging and provide effective support for subsequent intelligent analysis and decision-making.
[0004] Currently, RGB and infrared low-light image fusion algorithms directly for intelligent reconnaissance equipment are relatively rare. The existing main technical routes include: (1) Simple pixel-level weighted superposition: This method has the problem of mismatched structures between different modalities and is prone to image distortion. (2) Fusion based on traditional image processing algorithms: Such as first performing image enhancement and then superposition; these methods are highly dependent on the source data and have poor robustness. (3) Fusion based on shallow machine learning models: These methods are difficult to model complex non-linear mapping relationships and have problems with insufficient generalization ability. (4) Deep learning fusion models trained using general datasets: Since they cannot reflect the unique sensing characteristics of the electro-optical pod, the effects are not good. (5) Two-stage methods, first separately restoring each modal image and then performing fusion; but this method also has difficulties in the process of separately restoring images. Summary of the Invention
[0005] The present invention aims to overcome the above-mentioned defects of the prior art and provides a method for fusing RGB and near-infrared images under low-light conditions for an intelligent reconnaissance equipment.
[0006] To solve the above technical problems, the technical solution of the present invention is as follows:
[0007] A method for fusing RGB and near-infrared images under low-light conditions of an intelligent reconnaissance equipment, comprising:
[0008] Using the intelligent reconnaissance equipment to collect RGB images and near-infrared images under low-light conditions and perform preprocessing;
[0009] Inputting the preprocessed RGB image and near-infrared image into an RGB restoration network and a near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image;
[0010] Inputting the structural features of the RGB image and the near-infrared image into a deep structure extraction module to obtain deep structure feature maps of the RGB image and the near-infrared image;
[0011] Performing a prior calculation of deep structure inconsistency on the deep structure feature maps of the RGB image and the near-infrared image to obtain a prior feature map of deep structure inconsistency, and then performing a multiplication operation on the prior feature map of deep structure inconsistency and the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure;
[0012] Inputting the near-infrared feature map with consistent structure and the RGB image into a multi-scale deep fusion module to obtain a fused image.
[0013] Furthermore, the present invention also proposes a system for fusing RGB and near-infrared images under low-light conditions of an intelligent reconnaissance equipment, applying the method for fusing RGB and near-infrared images under low-light conditions proposed by the present invention. Among them, the system includes:
[0014] An image acquisition module, including an intelligent reconnaissance equipment, for collecting RGB images and near-infrared images under low-light conditions and performing preprocessing;
[0015] A restoration module, on which an RGB restoration network and a near-infrared restoration network are mounted, for inputting the preprocessed RGB image and near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image;
[0016] A deep structure extraction module, for performing deep structure extraction on the structural features of the input RGB image and near-infrared image to obtain deep structure feature maps of the RGB image and the near-infrared image;
[0017] A consistent structure generation module, for performing a prior calculation of deep structure inconsistency on the deep structure feature maps of the RGB image and the near-infrared image to obtain a prior feature map of deep structure inconsistency; and then performing a multiplication operation on the prior feature map of deep structure inconsistency and the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure;
[0018] A multi-scale deep fusion module is used to deeply fuse the input near-infrared feature map with consistent structure and the RGB image to obtain a fused image.
[0019] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0020] The present invention extracts deep structure feature maps from RGB and near-infrared images through a deep structure extraction module. At the same time, prior knowledge is introduced, and the depth inconsistency prior algorithm is used to calculate the inconsistency of the deep structure feature information in the RGB and near-infrared images. Multiply the deep structure inconsistency prior feature map and the deep structure feature map of the near-infrared image to suppress the part inconsistent with the RGB deep structure, and output a near-infrared feature map with consistent structure; input the near-infrared feature map with consistent structure and the RGB feature map into the multi-scale deep fusion module to output the final fused result map, so that the intelligent reconnaissance equipment can achieve more robust fusion of target RGB and near-infrared images in complex environments such as low-light conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a flowchart of the RGB and near-infrared image fusion method shown according to an embodiment of the present invention.
[0022] Figure 2 It is a schematic framework diagram of the RGB and near-infrared image fusion method shown according to an embodiment of the present invention.
[0023] Figure 3 It is an architecture diagram of the deep structure extraction module shown according to an embodiment of the present invention.
[0024] Figure 4 It is an architecture diagram of the autoencoder shown according to an embodiment of the present invention.
[0025] Figure 5 It is an architecture diagram of the multi-scale deep fusion module shown according to an embodiment of the present invention.
[0026] Figure 6 It is an architecture diagram of the RGB and near-infrared image fusion system shown according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0028] The terms used in the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "said", and "the" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0030] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0031] Embodiment 1
[0032] This embodiment proposes a method for fusing RGB and near-infrared images under low-light conditions for an intelligent reconnaissance device, as Figure 1 、 2 shown, which is the flowchart and framework diagram of the RGB and near-infrared image fusion method of this embodiment.
[0033] The method for fusing RGB and near-infrared images under low-light conditions for the intelligent reconnaissance device proposed in this embodiment includes the following steps:
[0034] S1. Use the intelligent reconnaissance device to collect RGB images and near-infrared images under low-light conditions and perform preprocessing;
[0035] S2. Input the preprocessed RGB image and near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image;
[0036] S3. Input the structural features of the RGB image and the near-infrared image into the deep structure extraction module to obtain the deep structure feature maps of the RGB image and the near-infrared image;
[0037] S4. Perform a prior calculation of deep structure inconsistency on the deep structure feature maps of the RGB image and the near-infrared image to obtain a prior feature map of deep structure inconsistency, and then perform a multiplication operation on it and the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure;
[0038] S5. Input the near-infrared feature maps and RGB images with consistent structures into the multi-scale deep fusion module to obtain a fused image.
[0039] In this embodiment, an optoelectronic detection device in an intelligent reconnaissance equipment is used to collect the RGB and near-infrared images of a target. Deep structure feature maps in the RGB and near-infrared images are extracted through a deep structure extraction module. Meanwhile, prior knowledge is introduced, and a depth inconsistency prior algorithm is used to calculate the inconsistency of the deep structure feature information in the RGB and near-infrared images, and a deep structure inconsistency prior feature map is output. Multiply the deep structure inconsistency prior feature map and the deep structure feature map of the near-infrared image to suppress the part inconsistent with the RGB deep structure, and output a near-infrared feature map with consistent structures. Input the near-infrared feature map with consistent structures and the RGB feature map into the multi-scale deep fusion module, and output a final fused result map, so that the intelligent reconnaissance equipment can achieve more robust fusion of target RGB and near-infrared images in complex environments such as low-light conditions.
[0040] In an alternative embodiment, the method further includes the following steps:
[0041] Use the intelligent reconnaissance equipment to collect reference image pairs in normal light environments and low-light environments, and preprocess them to obtain training data for pre-training an RGB and near-infrared image fusion model composed of an RGB restoration network, a near-infrared restoration network, a deep structure extraction module, and a multi-scale deep fusion module.
[0042] Exemplarily, in a specific implementation process, in order to make the fused image closer to the real image, synthetic noise is added to the reference image pairs for training and testing, and when pre-training the image fusion model, 5k pairs of reference images (256*256) are used as the training set, and another 1k pairs of reference images (256*256) and 10 additional pairs of real noise images (1920*1080) are used for testing.
[0043] It should be noted that the image pairs mentioned in this embodiment include the RGB image and the near-infrared image collected at the same time and the same angle.
[0044] Furthermore, in an alternative embodiment, the steps for preprocessing the collected images include:
[0045] (1) Perform noise removal processing on the collected images;
[0046] (2) Align the RGB and near-infrared image pairs using a manual double-check matching algorithm;
[0047] (3) Collect the images into image patches of a preset size.
[0048] Optionally, to obtain high-quality references, for image pairs collected in a normal light environment, multiple static captures are averaged to remove noise, and a manually double-checked matching algorithm is used to ensure the alignment of the image pairs. The collected images are cropped into 256*256 image patches.
[0049] For real noisy image pairs collected in a low-light environment, the preprocessing steps are the same as those for collecting the reference image pairs, and further include not performing multi-frame averaging on the noisy RGB images.
[0050] In an optional embodiment, in step S3, the deep structure extraction module includes several convolutional layers. After the output of each convolutional layer is subjected to an addition operation, it is activated by an activation function to output a multi-scale deep structure feature map. Among them, a convolutional layer for outputting the structural features of one scale is connected to a residual channel attention block for performing an addition operation with the structural features of other scales.
[0051] Exemplarily, as Figure 3 shown, it is the architecture diagram of the deep structure extraction module of this embodiment. In this embodiment, the structural features of one scale need to pass through a residual channel attention block, and the output is sequentially added to the structural features of other sizes to obtain a multi-scale deep structure feature map.
[0052] The deep structure extraction module in this embodiment can enable intelligent reconnaissance equipment to better analyze and represent the internal structure of images. Even in an extremely dim light environment, effective image structures can be extracted as the basis for subsequent processing.
[0053] Furthermore, in an optional embodiment, during the pre-training process of the deep structure extraction module, it includes the following steps:
[0054] After inputting the training data into the RGB restoration network and the near-infrared restoration network, the structural features at their output ends are input into the deep structure extraction module for training. Among them, the loss function of the deep structure extraction module is:
[0055]
[0056]
[0057] Among them, N i represents the number of channels of the deep structure of the i-th scale; str i,c represents the predicted structure diagram of the c-th channel of the deep structure of the i-th scale; represents the deep structure of the true value of the c-th channel of the deep structure of the i-th scale; y d and z dis the pixel value of the d-th pixel on the deep structure of the predicted structure diagram and the true value, and N is the number of image pixels.
[0058] It should be noted that it is very difficult for the deep structure extraction module to generate a noise-free structure diagram alone. Therefore, in this embodiment, the noise-free true value is selected as the supervision signal. Optionally, the supervision signal is generated by the encoder.
[0059] Furthermore, in an optional embodiment, the true structure feature map is generated by an autoencoder; the encoder includes a first convolutional block, a second convolutional block, a third convolutional block, a fourth convolutional block, a fifth convolutional block, and a sixth convolutional block connected in sequence. Among them, a residual channel attention block is respectively connected to the outputs of the first convolutional block, the fourth convolutional block, and the fifth convolutional block; a residual channel attention block is respectively connected to the inputs of the second convolutional block and the third convolutional block; at least 2 residual channel attention blocks are connected between the third convolutional block and the fourth convolutional block;
[0060] The second convolutional block and the third convolutional block include a downsampling layer and a convolutional layer; the fourth convolutional block and the fifth convolutional block include an upsampling layer and a convolutional layer;
[0061] Take the input feature map of the fourth convolutional block, the output feature map of the fourth convolutional block, and the output feature map of the fifth convolutional block as the supervision signal of the deep structure extraction module for generating the structure diagram of the true value.
[0062] Exemplarily, as Figure 4 shown, is the architecture diagram of the autoencoder of this embodiment.
[0063] This embodiment extracts the deep features de of the multi-scale decoder from the autoencoder i,c , and calculates the supervision signal:
[0064]
[0065] Among them, represents the n-th pixel of is the sobel operator, represents the n-th pixel in represents the global average pooling of . After subtracting the two, with 0 as the threshold, for pixels less than or equal to 0, set the corresponding pixel value in the structure diagram to 0, otherwise set it to 1, to obtain the final deep structure of the true value.
[0066] In this embodiment, a supervision signal is generated by an autoencoder, making the structural feature extraction more accurate, providing guarantee for stable imaging under complex illumination changes, and further improving the robustness of imaging and recognition of intelligent reconnaissance equipment.
[0067] In an alternative embodiment, a prior calculation of deep structure inconsistency for the deep structure feature maps of RGB images and near-infrared images includes the following steps:
[0068] According to the deep structure feature maps of RGB images and near-infrared images, binary edge maps are extracted from each feature channel, and their inconsistency is calculated to obtain a prior feature map of deep structure inconsistency; its expression is:
[0069]
[0070] Where and respectively represent the c-th channel of the deep structure features of the i-th scale in the noisy RGB image and near-infrared image; β ∈ (0, 1) is a hyperparameter, indicating no significant inconsistency; U i,c represents the feature value of the c-th channel of the deep structure features of the i-th scale in the prior feature map of deep structure inconsistency;
[0071] The prior feature map of deep structure inconsistency U i,c is multiplied by the deep structure feature map of the near-infrared image, the inconsistent structures in the image are discarded, and the corresponding deep feature values are updated to obtain a structurally consistent near-infrared feature map; its expression is:
[0072]
[0073] Where represents the feature value of the c-th channel of the deep structure features of the i-th scale in the updated deep structure feature map of the near-infrared image.
[0074] In this embodiment, prior knowledge of deep structure inconsistency is introduced, combined with multi-scale deep structure feature maps, to design a deep structure inconsistent feature U i,c , and further multiply the result by the original structure to discard the structures in the deep structure feature map of the near-infrared image that are inconsistent with the RGB image.
[0075] Exemplarily, the hyperparameter β in this embodiment is set to 0.5.
[0076] In this embodiment, by adopting the depth inconsistency prior algorithm, the intelligent reconnaissance equipment can detect the differences in structure and information between the RGB image and the near-infrared image, which is further used to guide subsequent fusion and optimization. This enables multi-source heterogeneous images in low-light environments to also achieve a more coordinated and unified effect.
[0077] In an alternative embodiment, the structurally consistent near-infrared feature map and the RGB image are input into a multi-scale deep fusion module to obtain a fused image. Among them, the multi-scale deep fusion module includes a convolutional layer based on residual connection, as well as an upsampling layer and a downsampling layer.
[0078] Among them, the RGB image input into the multi-scale deep fusion module undergoes convolution and residual channel attention processing to obtain a denoised coarse RGB structural feature map;
[0079] The structurally consistent near-infrared feature map input into the multi-scale deep fusion module undergoes convolution processing, and then a summation operation based on residual connection is performed with the denoised coarse RGB structural feature map. After downsampling and upsampling processing, a fused image is generated through a convolutional layer and an activation function.
[0080] Exemplarily, as Figure 5 shown, it is the architecture diagram of the multi-scale deep fusion module of this embodiment.
[0081] Furthermore, in an alternative embodiment, during the pre-training process of the multi-scale deep fusion module, the training data is used as the model input, and based on the fused image output by the model, the total loss function of the model is calculated; its expression is:
[0082]
[0083] Among them, and are the loss functions for RGB and near-infrared deep structure prediction, and α and γ are preset coefficients; and are the reconstruction losses of the fused image, the coarse RGB, and the near-infrared image respectively, and are Charbonnier losses, and their form is:
[0084]
[0085] Among them, X and X z represent the network output and the corresponding true value supervision respectively, and θ is a hyperparameter.
[0086] Exemplarily, θ is set to 1 / 1000.
[0087] In this embodiment, the multi-scale deep fusion module effectively unifies and fuses multi-level information such as the structure and text of the image, constructs a richer and more complete image representation under low-light conditions, and can significantly improve the performance of subsequent intelligent analysis such as target recognition and classification.
[0088] Under the moving conditions of intelligent reconnaissance equipment, the method proposed in this embodiment is repeatedly applied for image acquisition and processing experiments. The results show that the output image quality, stability and other indicators of the present invention are superior to the simple fusion method of direct superposition, and the robustness of the method in this embodiment is verified, and it is verified that the intelligent reconnaissance equipment can achieve more robust fusion of target RGB and near-infrared images in the face of complex environments.
[0089] Embodiment 2
[0090] This embodiment proposes an RGB and near-infrared image fusion system for intelligent reconnaissance equipment under low-light conditions, which applies the RGB and near-infrared image fusion method proposed in Embodiment 1. As Figure 6 shown, it is the architecture diagram of the RGB and near-infrared image fusion system under low-light conditions in this embodiment.
[0091] In the RGB and near-infrared image fusion system for intelligent reconnaissance equipment proposed in this embodiment, it includes:
[0092] An image acquisition module, including intelligent reconnaissance equipment, is used to collect RGB images and near-infrared images under low-light conditions and perform preprocessing;
[0093] A restoration module, on which an RGB restoration network and a near-infrared restoration network are mounted, is used to input the preprocessed RGB image and near-infrared image into the RGB restoration network and the near-infrared restoration network respectively to obtain the structural features of the RGB image and the near-infrared image;
[0094] A deep structure extraction module is used to perform deep structure extraction on the structural features of the input RGB image and near-infrared image to obtain deep structure feature maps of the RGB image and the near-infrared image;
[0095] A structure-consistent generation module is used to perform a priori calculation of deep structure inconsistency on the deep structure feature maps of the RGB image and the near-infrared image to obtain a deep structure inconsistency prior feature map; then perform a multiplication operation on the deep structure inconsistency prior feature map and the deep structure feature map of the near-infrared image, and output a structure-consistent near-infrared feature map;
[0096] A multi-scale deep fusion module is used to perform deep fusion on the input structure-consistent near-infrared feature map and the RGB image to obtain a fused image.
[0097] It can be understood that the system of this embodiment corresponds to the method of Embodiment 1 above. The optional items in Embodiment 1 above are equally applicable to this embodiment, so they will not be described repeatedly here.
[0098] The terms in the accompanying drawings are only for illustrative purposes and should not be construed as limiting the present invention.
[0099] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for fusing RGB and near-infrared images under weak light conditions of intelligent reconnaissance equipment, characterized in that: The following steps are involved: Use intelligent reconnaissance equipment to collect RGB images and near-infrared images under low-light conditions and perform preprocessing; The preprocessed RGB image and near-infrared image are input into the RGB restoration network and the near-infrared restoration network respectively, and the structural features of the RGB image and the near-infrared image are obtained respectively; Inputting the structural features of the RGB image and the near-infrared image into a deep structure extraction module to obtain deep structure feature maps of the RGB image and the near-infrared image; Performing a deep structure inconsistency prior calculation on the deep structure feature maps of the RGB image and the near-infrared image to obtain a deep structure inconsistency prior feature map, and then performing a product operation on the deep structure feature map of the near-infrared image to output a near-infrared feature map with consistent structure; The near-infrared feature map and the RGB image with consistent structure are input into a multi-scale deep fusion module to obtain a fused image.
2. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 1, characterized in that: The deep structure extraction module includes several convolutional layers. The output of each convolutional layer is summed and then activated by an activation function to output a multi-scale deep structure feature map. Among them, the convolutional layer used to output the structural features of one scale is connected to a residual channel attention block for summing the structural features with other scales.
3. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 1, characterized in that: The deep structure inconsistency prior calculation of the deep structure feature map of the RGB image and the near infrared image comprises the following steps: According to the deep structure feature maps of RGB images and near-infrared images, the binary edge map is extracted from each feature channel, and its inconsistency is calculated to obtain the deep structure inconsistency prior feature map; its expression is: in, and represents the cth channel of the deep structure feature of the i-th scale in the noisy RGB image and the near-infrared image respectively; β∈(0,1) is a hyperparameter; U i,c Represents the feature value of the cth channel of the deep structure feature of the i-th scale in the deep structure inconsistent prior feature map; The deep structure is inconsistent with the prior feature map U i,c Perform a product operation with the deep structure feature map of the near-infrared image, discard inconsistent structures in the image, update the corresponding deep feature values, and obtain a near-infrared feature map with consistent structure; its expression is: in, Represents the feature value of the cth channel of the deep structure feature of the i-th scale in the deep structure feature map of the updated near-infrared image.
4. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 1, characterized in that: The multi-scale deep fusion module includes a convolutional layer based on residual connection, an upsampling layer and a downsampling layer; The RGB image input into the multi-scale deep fusion module is processed by convolution and residual channel attention to obtain a denoised coarse RGB structural feature map; The near infrared feature map with consistent structure input into the multi-scale deep fusion module is subjected to convolution processing, and then subjected to residual connection-based addition operation with the coarse RGB structural feature map after noise reduction, and after downsampling and upsampling processing, a fused image is generated through a convolution layer and an activation function.
5. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 1, characterized in that: The steps of preprocessing the acquired images include: Perform noise removal on the collected images; The RGB and NIR image pairs were aligned using a matching algorithm with manual double checking; Capture images into blocks of preset size.
6. The method for fusion of RGB and near-infrared images under weak light conditions according to any one of claims 1 to 5, characterized in that: The method further comprises the following steps: Intelligent reconnaissance equipment is used to collect RGB images and near-infrared images in normal light and low-light environments and preprocess them to obtain training data for pre-training the RGB and near-infrared image fusion model composed of an RGB restoration network, a near-infrared restoration network, a deep structure extraction module and a multi-scale deep fusion module.
7. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 6, characterized in that: Pre-training the RGB and near-infrared image fusion model includes the following steps: The training data is used as the model input, and the total loss function of the model is calculated based on the fused image output by the model; its expression is: in, and It is the loss function for RGB and near-infrared deep structure prediction, α and γ are preset coefficients; and are the reconstruction losses of the fused image, coarse RGB, and near-infrared image, respectively, and are Charbonnier losses.
8. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 7, characterized in that: The method further comprises the following steps: After the training data is input into the RGB restoration network and the near infrared restoration network, the output structural features are input into the deep structure extraction module for training; wherein the loss function of the deep structure extraction module is: Among them, N i Indicates the number of channels of the deep structure of the i-th scale; str i,c Represents the predicted structure graph of the cth channel of the deep structure at the i-th scale; The deep structure representing the true value of the cth channel of the deep structure at the i-th scale; y d and z d is the pixel value of the dth pixel in the deep structure of the predicted structure graph and the true value, and N is the number of image pixels.
9. The method for fusion of RGB and near-infrared images under weak light conditions according to claim 8, characterized in that: The real structure feature map is generated by an autoencoder; The encoder includes a first convolution block, a second convolution block, a third convolution block, a fourth convolution block, a fifth convolution block and a sixth convolution block connected in sequence, wherein the outputs of the first convolution block, the fourth convolution block and the fifth convolution block are respectively connected to a residual channel attention block; the input ends of the second convolution block and the third convolution block are respectively connected to a residual channel attention block; at least two residual channel attention blocks are connected between the third convolution block and the fourth convolution block; The second convolution block and the third convolution block include a downsampling layer and a convolution layer; the fourth convolution block and the fifth convolution block include an upsampling layer and a convolution layer; The input feature map of the fourth convolution block, the output feature map of the fourth convolution block, and the output feature map of the fifth convolution block are taken as the supervision signal of the deep structure extraction module to generate a structure map of the true value.
10. A system for fusion of RGB and near-infrared images under weak light conditions of intelligent reconnaissance equipment, using the method for fusion of RGB and near-infrared images under weak light conditions as claimed in any one of claims 1 to 9, characterized in that: include: Image acquisition module, including intelligent reconnaissance equipment, for collecting RGB images and near-infrared images under low-light conditions and preprocessing them; A restoration module, which is equipped with an RGB restoration network and a near-infrared restoration network, is used to input the preprocessed RGB image and the near-infrared image into the RGB restoration network and the near-infrared restoration network, respectively, to obtain the structural features of the RGB image and the near-infrared image; A deep structure extraction module is used to perform deep structure extraction on the structural features of the input RGB image and near-infrared image to obtain deep structure feature maps of the RGB image and near-infrared image; The structure consistency generation module is used to perform deep structure inconsistency prior calculation on the deep structure feature maps of RGB images and near-infrared images to obtain deep structure inconsistency prior feature maps; Then, the deep structure inconsistent prior feature map and the deep structure feature map of the near infrared image are multiplied to output a near infrared feature map with consistent structure; The multi-scale deep fusion module is used to deeply fuse the input near-infrared feature map and RGB image with consistent structure to obtain a fused image.
Citation Information
Patent Citations
Infrared and visible light fusion imaging method based on deep learning
CN113487530A
Infrared and visible light image fusion method with regional attention
CN114782298A
Infrared and visible light image fusion method based on multi-scale features
CN116258936A
Fusion method and system of infrared image and visible light image
CN117474782A
Image processing method, device and equipment and computer readable storage medium
CN118055302A