Brightness perception multi-mode intelligent reconnaissance equipment image fusion system

By designing a multimodal intelligent reconnaissance equipment image fusion system with brightness perception, using lighting perception network and feature extraction fusion technology, the problems of inconsistency in image fusion of different sensors are solved, and high-quality image fusion is achieved in complex environments.

CN120070203APending Publication Date: 2025-05-30GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510144729.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing images from different sensors, the prior art faces problems such as inconsistent resolution, poor information complementarity, and poor adaptability of fusion algorithms. Especially in dynamically changing reconnaissance environments, it is difficult to quickly and accurately fusion multi-source images.

Method used

A multimodal intelligent reconnaissance equipment image fusion system is designed. Through the image acquisition module, lighting evaluation module, feature extraction module, feature fusion module and image fusion module, the illumination perception network is used to evaluate the image brightness conditions, calculate the fusion weight, and adaptively fuse the image information of different sensors through feature extraction and feature fusion.

Benefits of technology

It realizes the generation of clear and rich synthetic images under different lighting conditions, improves the target reconnaissance capabilities of intelligent reconnaissance equipment in complex environments, ensures the quality of image fusion, and is suitable for various complex environments, including extreme weather and different geographical environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070203A_ABST
    Figure CN120070203A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides a brightness-sensing multi-mode intelligent reconnaissance equipment image fusion system, which comprises an image acquisition module used for acquiring an infrared image and a visible light image in the same scene and at the same time point; the illumination evaluation module is used for evaluating brightness conditions of the infrared image and the visible light image by using an illumination sensing network to obtain an illumination probability, and calculating a fusion weight of the infrared image and the visible light image based on the illumination probability; the feature extraction module is used for extracting feature maps of the infrared image and the visible light image; the feature fusion module is used for performing complementary information interaction fusion on the feature map according to the fusion weight, generating fusion features and inputting the fusion features into the next-layer feature extraction module; the fusion module is used for fusing the feature maps output by the last layer in the feature extraction module to generate fusion features; and the image fusion module is used for carrying out feature reconstruction according to the fusion features and converting a reconstructed feature map into an image domain to generate a fusion image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more particularly, to a brightness-aware multimodal intelligent reconnaissance equipment image fusion system. Background Art

[0002] Intelligent reconnaissance equipment is an advanced airborne reconnaissance device, widely used on fixed-wing aircraft and rotary-wing aircraft in both military and civilian fields. They are mainly used for long-range reconnaissance and surveillance of ground or sea targets. These pods integrate a variety of optoelectronic sensors, including but not limited to infrared thermal imagers, laser rangefinders, television cameras, and low-light night vision devices, capable of providing day and night reconnaissance capabilities. These sensors each have unique advantages. For example, an infrared thermal imager can detect targets under low illumination conditions, while a low-light night vision device can provide a higher spatial resolution. However, due to the imaging principles and performance limitations of the sensors, a single sensor is difficult to provide comprehensive target information. For example, although infrared images can prominently display heat source targets, they often lack sufficient texture details; while visible light images have rich texture details but are difficult to capture targets at night or under poor lighting conditions.

[0003] Existing image processing technologies often face problems such as inconsistent resolutions, weak information complementarity, and poor adaptability of fusion algorithms when processing images from different sensors. Especially in a dynamically changing reconnaissance environment, how to quickly and accurately fuse multi-source images to generate high-quality images for analysis is the main challenge faced by existing technologies. At the same time, lighting conditions have a significant impact on the quality and information content of images. In the application scenarios of intelligent reconnaissance equipment, the lighting conditions of targets may change rapidly, which requires the image fusion algorithm to be able to adapt to different lighting environments. Summary of the Invention

[0004] The present invention provides a brightness-aware multimodal intelligent reconnaissance equipment image fusion system to overcome the defect that the image fusion algorithm in the above-mentioned prior art is difficult to adapt to different lighting environments.

[0005] To solve the above technical problems, the technical solution of the present invention is as follows:

[0006] A brightness-aware multimodal intelligent reconnaissance equipment image fusion system, comprising:

[0007] An image acquisition module for acquiring infrared images and visible light images of the same scene at the same time point;

[0008] An illumination evaluation module for evaluating the brightness conditions of the infrared images and visible light images using an illumination perception network to obtain illumination probabilities, and calculating the fusion weights of the infrared images and visible light images based on the illumination probabilities;

[0009] A feature extraction module, which includes several convolutional layers thereon, for extracting feature maps of the infrared image and the visible light image;

[0010] A feature fusion module, for performing complementary information interaction and fusion on the feature maps output by each layer in the feature extraction module according to the fusion weight, generating a fusion feature and inputting it into the next layer of the feature extraction module; and, for fusing the feature maps output by the last layer in the feature extraction module to generate a fusion feature;

[0011] An image fusion module, for performing feature reconstruction according to the fusion feature and converting the reconstructed feature map into the image domain to generate a fused image.

[0012] Furthermore, the present invention also proposes an image fusion method for a brightness-aware multi-modal intelligent reconnaissance equipment, which is applied to the brightness-aware multi-modal intelligent reconnaissance equipment image fusion system proposed by the present invention. Among them, the method includes the following steps:

[0013] Collect an infrared image and a visible light image of the same scene at the same time point;

[0014] Use an illumination perception network to evaluate the brightness conditions of the infrared image and the visible light image, obtain an illumination probability, and calculate the fusion weights of the infrared image and the visible light image based on the illumination probability;

[0015] Use convolutional layers to extract features from the infrared image and the visible light image. At the same time, perform complementary information interaction and fusion on the feature maps output by each convolutional layer according to the fusion weight, generate a fusion feature and input it into the next convolutional layer;

[0016] Fuse the feature maps output by the last layer in the feature extraction module to generate a fusion feature;

[0017] Perform feature reconstruction according to the fusion feature and convert the reconstructed feature map into the image domain to generate a fused image.

[0018] Furthermore, the present invention also proposes a storage medium, on which computer-readable instructions are stored. Among them, when the computer-readable instructions are executed by a processor, all or part of the steps of the multi-modal intelligent reconnaissance equipment image fusion method proposed by the present invention are implemented.

[0019] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0020] The present invention utilizes an illumination perception network to evaluate the brightness conditions of the infrared image and the visible light image, and adjusts the image fusion strategy according to the brightness conditions; moreover, through feature extraction and feature fusion, it adaptively fuses the image information from different sensors to generate a clear and content-rich synthetic image, effectively improving the target reconnaissance ability of intelligent reconnaissance equipment in complex environments;

[0021] The present invention can ensure the quality of image fusion regardless of strong light, weak light or mixed lighting conditions, enabling the system to work stably in various complex environments, including extreme weather conditions and different geographical environments, providing strong technical support for field operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is an architecture diagram of an image fusion system for a multi-modal intelligent reconnaissance equipment shown according to an embodiment of the present invention.

[0023] Figure 2 It is an architecture diagram of a feature extraction module and a feature fusion module shown according to an embodiment of the present invention.

[0024] Figure 3 It is an architecture diagram of an image fusion module shown according to an embodiment of the present invention.

[0025] Figure 4 It is a flowchart of a multi-modal intelligent reconnaissance equipment image fusion method shown according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0027] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0029] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] Embodiment 1

[0031] This embodiment proposes an image fusion system for a brightness-aware multi-modal intelligent reconnaissance equipment, as Figure 1 shown, which is the architecture diagram of the multi-modal intelligent reconnaissance equipment image fusion system of this embodiment.

[0032] In the brightness-aware multi-modal intelligent reconnaissance equipment image fusion system proposed in this embodiment, it includes:

[0033] An image acquisition module, configured to acquire infrared images and visible light images of the same scene at the same time point;

[0034] An illumination evaluation module, configured to evaluate the brightness conditions of the infrared images and visible light images by using an illumination perception network to obtain illumination probabilities, and calculate the fusion weights of the infrared images and visible light images based on the illumination probabilities;

[0035] A feature extraction module, which includes a number of convolutional layers, configured to extract feature maps of the infrared images and visible light images;

[0036] A feature fusion module, configured to perform complementary information interaction and fusion on the feature maps output by each layer in the feature extraction module according to the fusion weights to generate fusion features and input them into the next layer of the feature extraction module; and, configured to fuse the feature maps output by the last layer in the feature extraction module to generate fusion features;

[0037] An image fusion module, configured to perform feature reconstruction according to the fusion features and convert the reconstructed feature maps into the image domain to generate a fused image.

[0038] In this embodiment, the illumination perception network is used to evaluate the brightness conditions of the infrared image and the visible light image, and the image fusion strategy is adjusted according to the brightness conditions; moreover, through feature extraction and feature fusion, the image information from different sensors is adaptively fused to generate a clear and content-rich synthetic image, effectively improving the target reconnaissance ability of the intelligent reconnaissance equipment in complex environments. This embodiment ensures the quality of image fusion regardless of strong light, weak light or mixed lighting conditions, enabling the system to work stably in various complex environments, including extreme weather conditions and different geographical environments, providing strong technical support for field operations.

[0039] In an alternative embodiment, the image acquisition module includes:

[0040] An acquisition unit for acquiring infrared images and visible light images;

[0041] An image pairing unit for registering the acquired infrared image and visible light image in terms of size and coordinates to ensure pixel-level correspondence.

[0042] Wherein, when performing the registration operation, the image pairing unit executes the following steps:

[0043] Use a feature extraction module composed of a convolutional neural network and an attention mechanism to perform image feature extraction on the input image, which can pay more attention to the key regions in the image, thereby improving the expression ability of features; then use the sobel operator combined with the maximum response detection to capture the key points K I and K V ;

[0044] For the key points K I and K V Use the fast approximate nearest neighbor algorithm to obtain the feature descriptors D I and D V , and then match the extracted infrared image feature key points with the visible light image feature key points; this process can be expressed as:

[0045] {D I ,D V}={FLANN(K I ),FLANN(K V )};

[0046] Wherein, FLANN(·) represents the fast approximate nearest neighbor algorithm;

[0047] Obtain the estimated transformation matrix through direct linear transformation (DLT), and use the estimated transformation matrix to transform the infrared image into the same coordinate system as the visible light image to obtain the registered infrared image Ip and the visible light image V p 。

[0048] Further optionally, the acquisition unit is a multi-modal sensor array mounted on intelligent reconnaissance equipment.

[0049] Further optionally, the multi-modal sensor array includes an infrared thermal imager, a visible light camera, a near-infrared camera, etc., for capturing images of the same scene under different spectra.

[0050] It should be noted that in the image acquisition module, it is necessary to synchronously trigger each sensor to ensure that the captured images correspond to the same time point for subsequent fusion processing.

[0051] Further, in an optional embodiment, the image acquisition module further includes a preprocessing unit for preprocessing the acquired infrared image and visible light image, including denoising and enhancement operations.

[0052] In an optional embodiment, the lighting evaluation module includes:

[0053] A brightness probability calculation unit equipped with a lighting perception network for calculating the lighting probabilities P d and P n ; where P d represents the probability that the image belongs to daytime, and P n represents the probability that the image belongs to night;

[0054] A lighting condition evaluation unit for calculating the fusion weights W d and W n of the infrared image and the visible light image; i and W v 。

[0055] Among them, when the brightness probability calculation unit calculates the lighting probabilities of daytime and night, the following steps are performed:

[0056] The lighting perception network extracts features from the input visible light image through a series of convolutional layers, adds non-linearity through the ReLU activation function after the convolutional layers, and then compresses the spatial information through the global average pooling layer to obtain lighting features;

[0057] The lighting features are used to calculate the probabilities of the current image belonging to daytime and night through 2 fully connected layers, and the output of the fully connected layer is converted into probability values P d and P n 。

[0058] Among them, the probability values satisfy Pd +P n = 1, and P d ≥ 0, P n ≥ 0.

[0059] When calculating the fusion weight, the lighting condition evaluation unit performs the following steps:

[0060] Based on the lighting probability P d and P n , calculate the lighting perception weight through the comprehensive feature function; its expression is:

[0061] f IR = [F texture (I p ), F contrast (I p ), F edge (I p )]

[0062] f VI = [F texture (V p ), F contrast (V p ), F edge (V p )]

[0063]

[0064] Among them, I p and V p represent the registered infrared image and visible light image; F texture (I), F contrast (I), F edge (I) respectively represent the functions for calculating the texture feature, contrast feature, and edge information of the image; f IR and f VI respectively represent the texture, contrast, and edge information feature vectors of the infrared image and visible light image; W(f IR , f VI ) is the weight adjustment function, which is used to adjust the weight of the image according to the feature vector.

[0065] Further optionally, the lighting perception network uses the cross-entropy loss function to train the lighting perception network to accurately classify whether the image belongs to day or night. Its expression is:

[0066]

[0067] Among them, y i is the true label, y i = 1 indicates that the image is during the day, y i= 0 indicates that the image is in darkness; z i is the prediction output of the lighting perception network for each category; σ(·) is the softmax function.

[0068] Furthermore, according to P d and P n values, the lighting condition of the current image is evaluated. When P d > P n , it is considered that the image is in a daytime scene; otherwise, it is a nighttime scene.

[0069] The lighting condition evaluation unit of this embodiment uses the lighting probability output by the lighting perception network to calculate the lighting perception weight through the comprehensive feature function, so as to represent the contribution degree of the infrared image and the visible light image in the fusion process. Among them, the comprehensive feature function integrates other image features such as texture, contrast, and edge information into the weight calculation of the infrared image and the visible light image to more comprehensively evaluate the lighting condition of the image, and finally obtains the weight W i of the infrared image and the weight W v of the visible light image.

[0070] In an alternative embodiment, the feature extraction module is composed of multiple convolutional layers and a Mamba architecture layer; among them, the convolutional layer is used to extract features at different levels of the image; the Mamba architecture layer is used to establish connections between different channels of the feature map.

[0071] Among them, the infrared image and the visible light image collected by the image acquisition module are respectively input into the feature extraction module, and the infrared image feature map and the visible light image feature map are respectively extracted. This process is expressed as:

[0072] F conv = ReLU(W i * Y i-1 + b)

[0073] F manba = MambaLayer(Y i-1 , θ)

[0074] Among them, F conv represents the convolution operation process of the convolutional layer; F manba represents the Mamba architecture layer; W i represents the convolution kernel of the i-th convolutional layer, * represents the convolution operation, Y i-1 represents the feature map output by the (i - 1)-th convolutional layer, b is the bias term; θ is the learnable parameter.

[0075] Furthermore, the feature maps of each layer output by the feature extraction module are input into the feature fusion module for interactive fusion to compensate for the difference information between different modalities.

[0076] Further optionally, when the feature fusion module performs complementary information interaction fusion on the extracted feature maps according to the fusion weights of the infrared image and the visible light image, the following steps are executed:

[0077] Calculate the common feature weight according to the infrared image feature and the visible light image feature output by the feature extraction module and the complementary feature weight The expression is:

[0078]

[0079] where α, β, γ, θ are dynamic weights; represents the infrared image feature output by the i-th convolutional layer, represents the visible light image feature output by the i-th convolutional layer;

[0080] Calculate the importance weight of each feature channel; the expression is:

[0081]

[0082] where, represents the common feature importance weight of the i-th convolutional layer, represents the complementary feature importance weight of the i-th convolutional layer;

[0083] Combine the importance weights to perform interactive fusion on the infrared image feature and the visible light image feature output by the feature extraction module to obtain the infrared image fusion feature and the visible light image fusion feature and transmit them to the next layer of the feature extraction module; the expression is:

[0084]

[0085] where, represents the element addition operation, ⊙ represents the channel multiplication operation; δ(δ) represents the sigmoid function, and GAP(·) represents the global average pooling operation.

[0086] The feature fusion module in this embodiment is used to introduce dynamic weights for the common part and the complementary part of different modality features, and these weights can be adaptively adjusted according to the importance of the features.

[0087] Further optionally, when the feature fusion module performs fusion on the feature maps output by the last layer in the feature extraction module, the following steps are included:

[0088] The infrared image feature extracted by the last layer of the feature extraction module and the visible light image feature In the input attention aggregation layer, the global semantic features of the spatial and channel dimensions of the infrared image and the visible light image are further extracted and fused. Finally, the infrared and visible light features are concatenated at the channel level to obtain the fused feature map F f ; Its expression is:

[0089]

[0090] where AG(·) represents the attention aggregation layer, and Concat(·) represents the channel-level concatenation operation.

[0091] Exemplarily, as Figure 2 shown, it is the architecture diagram of the feature extraction and fusion module of this embodiment. Among them, CIFIU represents the feature fusion module.

[0092] In this embodiment, the feature extraction module and the feature fusion module form a cascaded structure to realize the fusion and interaction of complementary information, make full use of the complementary information of different sensors, and generate a synthetic image with richer information. The feature fusion module introduces dynamic weights for features of different modalities to achieve adaptive adjustment according to the feature importance, so as to compensate for the differential information between different modalities. Finally, the features extracted from the last layer of the feature extraction module are input into the attention aggregation layer, and the global semantic features of spatial attention and channel attention are fused. Finally, the infrared and visible light features are concatenated at the channel level to integrate common and complementary information.

[0093] Furthermore, in an optional embodiment, when the image fusion module performs feature reconstruction according to the fused features, the following steps are executed:

[0094] The fused features are passed through at least 4 convolutional units composed of 3×3 convolution and activation functions for feature reconstruction to obtain the reconstructed enhanced feature I f , and the enhanced feature I f is converted back to the image domain to generate the final enhanced image output.

[0095] Exemplarily, as Figure 3 shown, it is the architecture diagram of the image fusion module of this embodiment. In this embodiment, the image fusion module further processes the extracted image features to enhance the useful information in the image, such as edges, textures, and contrasts, thereby improving the readability and interpretability of the image, and making the synthetic image clearer and more accurate visually.

[0096] Exemplarily, compared with common image fusion methods, the AG index of the image fusion system of this embodiment is improved by 21.8% on 50 pairs of images in the MSRS dataset.

[0097] Embodiment 2

[0098] This embodiment proposes an image fusion method for a brightness-aware multi-modal intelligent reconnaissance equipment, which applies the multi-modal intelligent reconnaissance equipment image fusion system proposed in Embodiment 1. As Figure 4 shown, it is a flowchart of the image fusion method for the brightness-aware multi-modal intelligent reconnaissance equipment in this embodiment.

[0099] The image fusion method for the brightness-aware multi-modal intelligent reconnaissance equipment proposed in this embodiment includes the following steps:

[0100] Collect infrared images and visible light images of the same scene at the same time point;

[0101] Use the illumination perception network to evaluate the brightness conditions of the infrared images and visible light images, obtain the illumination probability, and calculate the fusion weights of the infrared images and visible light images based on the illumination probability;

[0102] Use the convolutional layer to extract features from the infrared images and visible light images. At the same time, according to the fusion weights, perform complementary information interaction fusion on the feature maps output by each convolutional layer, generate fusion features and input them into the next convolutional layer;

[0103] Fuse the feature maps output by the last layer in the feature extraction module to generate fusion features;

[0104] Perform feature reconstruction according to the fusion features, and convert the reconstructed feature maps into the image domain to generate a fused image.

[0105] Embodiment 3

[0106] This embodiment proposes a device, including a memory and a processor. When the computer-readable instructions stored in the memory are executed by the processor, the processor executes all or part of the steps of the multi-modal intelligent reconnaissance equipment image fusion method proposed in Embodiment 2.

[0107] Embodiment 4

[0108] This embodiment proposes a storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, all or part of the steps of the multi-modal intelligent reconnaissance equipment image fusion method proposed in Embodiment 2 are implemented.

[0109] Exemplarily, the storage medium includes but is not limited to various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0110] Exemplarily, the instructions, programs, code sets or instruction sets can be implemented in a conventional programming language.

[0111] Exemplarily, the processor includes but is not limited to smartphones, personal computers, servers, network devices, etc., and is used to execute all or part of the steps of the multi-modal intelligent reconnaissance equipment image fusion method described in Embodiment 2.

[0112] The terms in the accompanying drawings are only for illustrative purposes and should not be construed as limiting the present invention.

[0113] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A brightness-aware multi-modal intelligent reconnaissance equipment image fusion system, characterized in that: include: An image acquisition module is used to acquire infrared images and visible light images of the same scene and at the same time point; an illumination assessment module, configured to assess the brightness conditions of the infrared image and the visible light image using an illumination perception network, obtain an illumination probability, and calculate a fusion weight of the infrared image and the visible light image based on the illumination probability; A feature extraction module, comprising a plurality of convolutional layers, for extracting feature maps of the infrared image and the visible light image; A feature fusion module, used for interactively fusing complementary information of the feature maps output by each layer in the feature extraction module according to the fusion weight, generating fused features and inputting them into the feature extraction module of the next layer; and for fusing the feature maps output by the last layer in the feature extraction module to generate fused features; The image fusion module is used to reconstruct features according to the fusion features, and convert the reconstructed feature map into an image domain to generate a fused image.

2. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to claim 1 is characterized in that: The image acquisition module comprises: A collection unit, used for collecting infrared images and visible light images; The image pairing unit is used to align the size and coordinates of the acquired infrared image and visible light image; wherein: The feature extraction module composed of convolutional neural network and attention mechanism is used to extract image features from the input image, and then the key points K of infrared image and visible light image are captured by using Sobel operator combined with maximum response detection. I and K V ; For the key point K I and K V Using the fast approximate nearest neighbor algorithm, we can get the feature descriptors D of infrared and visible images. I and D V ; The estimated transformation matrix is ​​obtained by direct linear transformation, and the infrared image is transformed into the same coordinate system as the visible light image using the estimated transformation matrix to obtain the completed registered infrared image I p and the visible light image V p .

3. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to claim 2 is characterized in that: The image acquisition module also includes a preprocessing unit for preprocessing the acquired infrared image and visible light image, including denoising and enhancement operations.

4. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to claim 1, characterized in that: The lighting assessment module comprises: A brightness probability calculation unit, on which a lighting perception network is mounted, is used to calculate the lighting probability of daytime and nighttime according to the input visible light image; wherein: The lighting perception network extracts features from the input visible light image through a series of convolutional layers, adds nonlinearity through the ReLU activation function after the convolutional layer, and then obtains the lighting features through a global average pooling layer; The lighting features are passed through two fully connected layers to calculate the probability of the current image input being day or night, and the output of the fully connected layer is converted into a probability value P through a softmax function. d and P n ; A lighting condition evaluation unit is configured to evaluate the lighting condition according to the lighting probability P d and P n Calculate the fusion weight W of infrared image and visible light image i and W v ; Wherein, according to the lighting probability P d and P n , the lighting perception weight is calculated by comprehensive feature function; its expression is: f IR =[F texture (I p ),F contrast (I p ),F edge (I p )] f VI =[F texture (V p ),F contrast (V p ),F edge (V p )] Among them, I p and V p represents the registered infrared image and visible light image; F texture (I), F contrast (I), F edge (I) represent the functions used to calculate the texture features, contrast features and edge information of the image respectively; f IR and f VI Represent the texture, contrast, and edge information feature vectors of infrared images and visible light images respectively; W(f IR ,f VI ) is a weight adjustment function, which is used to adjust the weight of the image according to the feature vector.

5. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to claim 1, characterized in that: The feature extraction module is composed of multiple convolutional layers and Mamba architecture layers; wherein the convolutional layers are used to extract features at different levels of the image; and the Mamba architecture layer is used to establish connections between different channels of the feature map; its expression is: F conv =ReLU(W i *Y i-1 +b) F manba =MambaLayer(Y i-1 ,θ) Among them, F conv Represents the convolution operation process of the convolution layer; F manba Represents the Mamba architecture layer; W i represents the convolution kernel of the i-th convolutional layer, * represents the convolution operation, and Y i-1 Represents the feature map output by the i-1th convolutional layer, b is the bias term, and θ is a learnable parameter.

6. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to claim 5, characterized in that: When the feature fusion module performs complementary information interactive fusion on the extracted feature maps according to the fusion weights of the infrared image and the visible light image, the feature fusion module performs the following steps: Calculate the common feature weights based on the infrared image features and visible light image features output by the feature extraction module and complementary feature weights Its expression is: Among them, α, β, γ, and θ are dynamic weights; represents the infrared image features output by the i-th convolutional layer, Represents the visible light image features output by the i-th convolutional layer; Calculate the importance weight of each feature channel; its expression is: in, represents the common feature importance weight of the i-th convolutional layer, represents the importance weight of the complementary features of the i-th convolutional layer; Combined with the importance weight, the infrared image features and the visible light image features output by the feature extraction module are interactively fused to obtain the infrared image fusion feature. Fusion features with visible light images And transmitted to the next layer of feature extraction module; its expression is: in, represents the element addition operation, ⊙ represents the channel multiplication operation; δ(·) represents the sigmoid function, and GAP(·) represents the global average pooling operation.

7. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to claim 6, characterized in that: When the feature fusion module performs fusion of the feature graph outputted by the last layer in the feature extraction module, the following steps are included: The infrared image features extracted by the last layer of the feature extraction module are and visible light image features The input is sent to the attention aggregation layer to further extract the global semantic features of the space and channels of the infrared image and the visible light image and fuse them. Finally, the infrared and visible light features are connected at the channel level to obtain the fused feature map F f ; Its expression is: Among them, AG(·) represents the attention aggregation layer, and Concat(·) represents the channel-level connection operation.

8. The brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to any one of claims 1 to 7, characterized in that: When the image fusion module performs feature reconstruction according to the fusion feature, the image fusion module performs the following steps: The fusion feature is reconstructed through at least 4 convolution units consisting of 3×3 convolution and activation function to obtain the reconstructed enhanced feature I f , and the enhanced feature I f Convert back to the image domain to generate the final enhanced image output.

9. A brightness-aware multi-modal intelligent reconnaissance equipment image fusion method, applied to the brightness-aware multi-modal intelligent reconnaissance equipment image fusion system according to any one of claims 1 to 8, characterized in that: The following steps are involved: Collect infrared images and visible light images of the same scene and at the same time point; Using an illumination perception network to evaluate the brightness conditions of the infrared image and the visible light image, obtain an illumination probability, and calculate a fusion weight of the infrared image and the visible light image based on the illumination probability; The infrared image and the visible light image are subjected to feature extraction by using a convolution layer, and at the same time, complementary information is interactively fused on the feature maps output by each convolution layer according to the fusion weight, so as to generate fusion features and input them into the next convolution layer; Fusing the feature map output by the last layer in the feature extraction module to generate fused features; Feature reconstruction is performed according to the fusion features, and the reconstructed feature map is converted into an image domain to generate a fused image.

10. A storage medium having computer-readable instructions stored thereon, characterized in that: When the computer-readable instructions are executed by the processor, all or part of the steps of the multi-modal intelligent reconnaissance equipment image fusion method according to claim 9 are implemented.