Low-light aerial photography imaging method for monochromatic-infrared image fusion
By constructing an ALMIN model through a low-light aerial imaging method based on monochrome-infrared image fusion, and utilizing an illumination prior extractor and a dual-modal fusion restorer, the problems of underexposure and noise enhancement in low-light aerial images are solved, generating high-quality aerial images and improving image quality and detail richness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-27
AI Technical Summary
In low-light environments at night, low-light aerial images are prone to underexposure, increased noise, and severe loss of target information. Traditional color infrared image fusion methods are difficult to cope with the characteristics of low-light aerial images, thus affecting image quality.
A low-light aerial imaging method based on monochrome-infrared image fusion is proposed. By constructing the ALMIN method based on monochrome-infrared cameras, the method integrates monochrome and infrared features using an illumination prior extractor and a dual-modal fusion restorer to reconstruct illumination and reflectivity components, thereby generating aerial images with good exposure and rich texture details.
High-quality aerial images were generated, improving image brightness, suppressing noise, and enhancing aerial imaging quality, providing reliable visual data support for automated detection and analysis.
Smart Images

Figure CN121746199A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-light aerial imaging technology, and in particular to a low-light aerial imaging method for monochrome-infrared image fusion. Background Technology
[0002] Aerial imaging technology plays a crucial role in acquiring clear target information on the ground. Compared to ground images, aerial images can capture a wider range of spatial and denser texture information, and are widely used in various fields such as urban planning, drone target tracking, agricultural monitoring, and small target detection. However, in low-light environments at night, low-light aerial images are prone to underexposure, increased noise, and severe loss of target information under extreme lighting conditions, seriously affecting image quality.
[0003] Due to the presence of numerous small, inconspicuous targets in low-light aerial images, significant differences in illumination between regions, and motion blur caused by the movement of drones, traditional color infrared image fusion methods struggle to address these characteristics. Therefore, developing a specific multimodal fusion method tailored to low-light aerial images has become a key research direction for improving the quality of nighttime aerial imaging. Effective low-light aerial imaging methods can improve the overall brightness of images, suppress noise, and thus enhance the practicality of aerial imagery, providing more reliable visual data support for automated detection, monitoring, and analysis. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a low-light aerial imaging method for monochrome-infrared image fusion, which can generate aerial images with good exposure and rich texture details.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a low-light aerial imaging method for monochrome-infrared image fusion, comprising the following steps:
[0006] Step S1: Use the monochrome and infrared cameras mounted on the aerial photography equipment to acquire the required paired unregistered low-light monochrome infrared image dataset at low altitude at night.
[0007] Step S2: Construct the overall structure of ALMIN, an aerial low-light imaging method based on a monochrome-infrared camera;
[0008] Step S3: Construct an illumination prior extractor in ALMIN. Use the pre-trained model to obtain the local illumination prior for the low-light monochrome image obtained in step S1, and use the local illumination prior to initially enhance the texture details of the low-light monochrome image.
[0009] Step S4: Construct a dual-modal fusion restorer in ALMIN, use the local illumination prior obtained in step S3 for restoration guidance, integrate monochromatic features and infrared features to reconstruct illumination and reflectivity components, and generate aerial images.
[0010] In a preferred embodiment, step S1 specifically involves:
[0011] Step S11: In a low-altitude nighttime environment, use the monochrome and infrared cameras mounted on the aerial photography equipment to collect unregistered monochrome and infrared images.
[0012] Step S12: Use a pre-trained network model to detect matching feature points in unregistered monochrome and infrared image pairs; then, use the obtained matching points to estimate the homography matrix using the RANSAC algorithm. This matrix is used to perform perspective transformation on the monochrome image to obtain registered monochrome infrared image pairs.
[0013] Step S13: Divide the obtained n pairs of registered monochrome and infrared images into training set and test set.
[0014] In a preferred embodiment, in step S3, given a low-light monochrome image In this study, a feature decomposition network based on Retienx theory is used to analyze low-light monochromatic images. Perform illumination component Estimate the estimated illumination components This is used to enhance the details of low-light monochrome images, resulting in a preliminarily enhanced low-light monochrome image. The entire process is described as follows:
[0015] (1)
[0016] in This represents the noise introduced into the image during atmospheric propagation. This represents the illumination disturbance factor. This represents the perturbation factor of the reflection component. Representing richly detailed reflection components and This indicates the light component that provides even exposure.
[0017] In a preferred embodiment, in step S4, the low-light monochromatic image obtained in step S3 is preliminarily enhanced. and infrared images Simultaneously perform noise reduction processing. The process is represented as follows:
[0018] (2)
[0019] (3)
[0020] in, This represents the denoised low-light monochrome image. This represents the denoised infrared image.
[0021] In a preferred embodiment, in step S4, the dual-modal fusion repairer is used to perform... and The operation yields the illumination component of the low-light monochrome image. and reflection component and global features of infrared images With local features The process is represented as follows:
[0022] (4)
[0023] (5)
[0024] (6)
[0025] (7).
[0026] In a preferred embodiment, in step S3, the monochromatic features and infrared features obtained in step S4 are used to estimate the illumination component. Conducting guided integration This allows for the acquisition of well-exposed light components. and rich in texture details The process is represented as follows:
[0027] (8)
[0028] (9)
[0029] (10).
[0030] in This indicates an enhanced low-light aerial image.
[0031] In a preferred embodiment, this is achieved by designing a frequency domain denoising module. Functionality: First, the input image is projected into Q, K, and V using standard convolutional layers; second, Fourier transforms are performed on Q, K, and V to map them from the spatial domain to the frequency domain; third, in the frequency domain, dynamic filters are designed, and pointwise convolution and channel aggregation are performed; finally, an inverse Fourier transform is used to map the frequency domain representation back to the spatial domain. , and This is used to calculate attention scores; the specific formula is as follows:
[0032] (11)
[0033] (12)
[0034] (13)
[0035] in Indicates a dynamic filter. This indicates multi-head attention calculation. Indicates Fourier transform. Indicates the inverse Fourier transform. This represents the enhanced monochrome image and the initial infrared image acquired by the device. This represents three standard convolutions.
[0036] In a preferred embodiment, this is achieved through a global illumination sensing module. Functionality: First, a bidirectional sequential scan is performed using a multi-head SSD, employing parallel recursive control to extract the illumination component of low-light monochromatic images. Global features of infrared images The relative position mask matrix Effectively utilize the spatial locality of this feature information to extract global features; as shown below:
[0037] (14)
[0038] in The definition is as follows:
[0039]
[0040] in This represents a sequence reversal operation. These are learnable parameters;
[0041] Next, depthwise separable convolution is applied to the processed 2D feature map to model the local neighborhood relationships between feature channels, while adaptively adjusting local features through a gating mechanism.
[0042] In a preferred embodiment, this is achieved through a fine-grained reflectivity sensing module. Function: First, perform wavelet transform on the image. This process yields three high-frequency components and one low-frequency component. Subsequently, the three high-frequency components are stacked and summed along the channel dimension to obtain a unified feature representation. This is then achieved through global average pooling. Extract global context information and utilize convolution and non-linear activation. Channel-by-channel dimensionality reduction is performed; then, a standard convolutional layer is used to generate an attention vector for each high-frequency component, which is multiplied by the original high-frequency component, thereby simultaneously suppressing noise and enhancing details; at the same time, a reversible neural network module is used for the low-frequency components. The process involves extracting the lighting-invariant detail components from the low-frequency components.
[0043] (15)
[0044] (16)
[0045] (17)
[0046] (18).
[0047] In a preferred embodiment, this is achieved through a light-guided fusion module. Function; First, standard convolutional layers are applied to estimate the illumination information. Convolutional operations are performed to align the feature channels of the two modalities; then a lightweight multi-level perceptron is introduced to capture multi-scale contextual information and adaptively model the spatial distribution of brightness and contrast variations; finally, the generated fusion weights are normalized to the range [0,1] using a Sigmoid activation function.
[0048] (19)
[0049] (20)
[0050] (twenty one)
[0051] (twenty two)
[0052] in This represents the set of fusion weights.
[0053] Compared with the prior art, the present invention has the following advantages: the present invention can accurately generate high-quality aerial images with good exposure and rich texture details. Attached Figure Description
[0054] Figure 1 This is an overall workflow diagram of a preferred embodiment of the present invention;
[0055] Figure 2 This is a diagram of a low-light aerial imaging model based on a monochrome infrared camera, representing a preferred embodiment of the present invention. Detailed Implementation
[0056] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0057] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0058] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0059] Please refer to Figure 1-2 This invention provides a monochrome-infrared dual-camera drone aerial imaging method for low-light conditions, comprising the following steps:
[0060] Step S1: Use the monochrome and infrared cameras mounted on the aerial photography equipment to acquire the required paired unregistered low-light monochrome infrared image dataset at low altitude at night.
[0061] Step S2: Construct the overall structure of ALMIN, an aerial low-light imaging method based on a monochrome-infrared camera;
[0062] Step S3: Construct an illumination prior extractor in ALMIN. Use the pre-trained model to obtain a local illumination prior for the low-light monochrome image obtained in step S1. Use this local illumination prior to initially enhance the texture details of the low-light monochrome image, effectively preventing the loss of details in the subsequent fusion process.
[0063] Step S4: Construct a dual-modal fusion restorer in ALMIN, using the local illumination prior obtained in step S3 for restoration guidance, effectively integrating monochromatic features and infrared features to reconstruct illumination and reflectivity components, generating high-quality aerial images with good exposure and rich texture details.
[0064] In this embodiment, step S1 specifically includes:
[0065] Step S11: In a low-altitude nighttime environment, use the monochrome and infrared cameras mounted on the aerial photography equipment to collect unregistered monochrome and infrared images.
[0066] Step S12: Detect matching feature points in the unregistered monochrome and infrared image pairs using a pre-trained network. Then, the obtained matching points are used to estimate the homography matrix using the RANSAC algorithm. This matrix is used to perform perspective transformation on the monochrome image, thereby obtaining the registered monochrome infrared image pair.
[0067] Step S13: Divide the obtained n pairs of registered monochrome and infrared images into training set and test set.
[0068] In this embodiment, reference Figure 2 The Aerial Low-Light Imaging Model Based on Monochrome-Infrared Camera (ALMIN) employs a strategy guided by prior illumination information to fuse monochrome and infrared images, mitigating low-light degradation in aerial photography. Specifically, the ALMIN model, composed of a prior illumination information extractor and a dual-modal fusion inoculator, generates high-quality aerial images.
[0069] Given a low-light monochrome image In this study, a feature decomposition network based on Retienx theory is used to analyze low-light monochromatic images. Perform illumination component Estimate the estimated illumination components This is used to enhance the details of low-light monochrome images, resulting in a preliminarily enhanced low-light monochrome image. The entire process is described as follows:
[0070] (1)
[0071] in This represents the noise introduced into the image during atmospheric propagation. This represents the illumination disturbance factor. This represents the perturbation factor of the reflection component. Representing richly detailed reflection components and This indicates the light component that provides even exposure.
[0072] Subsequently, a preliminary enhanced low-light monochromatic image will be obtained. and infrared images Simultaneously perform noise reduction processing. The process is represented as follows:
[0073] (2)
[0074] (3)
[0075] Next, the dual-modal fusion repair device was used to perform separate procedures. and The operation yields the illumination component of the low-light monochrome image. and reflection component and global features of infrared images With local features The process is represented as follows:
[0076] (4)
[0077] (5)
[0078] (6)
[0079] (7)
[0080] Finally, the obtained monochromatic and infrared features are used to estimate the illumination component. Conducting guided integration This allows for the acquisition of well-exposed light components. and rich in texture details The process is represented as follows:
[0081] (8)
[0082] (9)
[0083] (10)
[0084] in This indicates an enhanced low-light aerial image.
[0085] In this embodiment, reference Figure 2 In step S4, a frequency domain denoising module is designed to achieve this. Functionality. First, we project the input image into Q, K, and V using standard convolutional layers. Second, we perform Fourier transforms on Q, K, and V, mapping them from the spatial domain to the frequency domain. Third, in the frequency domain, we design dynamic filters, performing pointwise convolution and channel aggregation to effectively suppress noise in the amplitude and phase components. Finally, we use an inverse Fourier transform to map the frequency domain representation back to the spatial domain. , and This module is used to calculate attention scores. The specific formula for this module is as follows:
[0086] (11)
[0087] (12)
[0088] (13)
[0089] in Indicates a dynamic filter. This indicates multi-head attention calculation. Indicates Fourier transform. Indicates the inverse Fourier transform. This represents the enhanced monochrome image and the initial infrared image acquired by the device. This represents three standard convolutions.
[0090] In this embodiment, reference Figure 2 In step S4, the global illumination sensing module is used to achieve this. Functionality. First, a bidirectional sequential scan is performed using a multi-head SSD, employing parallel recursive control to extract the illumination component of the low-light monochromatic image. Global features of infrared images The relative position mask matrix This effectively leverages the spatial locality of this feature information to extract global features. This process can be represented as follows:
[0091] (14)
[0092] in The definition is as follows:
[0093]
[0094] in This represents a sequence reversal operation. These are learnable parameters.
[0095] Next, we apply depthwise separable convolutions to the processed 2D feature maps to explicitly model the local neighborhood relationships between feature channels, while adaptively adjusting these local features through a gating mechanism. This design provides additional spatial context by supplementing global modeling, thereby ensuring the consistency of the extracted feature space.
[0096] In this embodiment, reference Figure 2 In step S4, the fine-grained reflectivity sensing module is used to achieve this. Function. First, perform wavelet transform on the image. This yields three high-frequency components and one low-frequency component. Subsequently, the three high-frequency components are stacked and summed along the channel dimension to obtain a unified feature representation. Global average pooling is then applied. Extract global context information and utilize convolution and non-linear activation. Channel-by-channel dimensionality reduction is performed. Next, a standard convolutional layer is used to generate an attention vector for each high-frequency component, which is then multiplied by the original high-frequency component, thus simultaneously suppressing noise and enhancing details. Meanwhile, a reversible neural network module is used for the low-frequency components. The process involves extracting the lighting-invariant detail components from the low-frequency components.
[0097] (15)
[0098] (16)
[0099] (17)
[0100] (18)
[0101] In this embodiment, reference Figure 2 In step S3, the illumination-guided fusion module is used to achieve this. Functionality. First, we apply standard convolutional layers to the estimated illumination information. A convolutional operation is performed to align the convolutional values with the feature channels of both modalities. A lightweight multi-level perceptron block is then introduced to capture multi-scale contextual information, adaptively modeling the spatial distribution of brightness and contrast variations. Finally, the generated fusion weights are normalized to the range [0,1] using a sigmoid activation function.
[0102] (19)
[0103] (20)
[0104] (twenty one)
[0105] (twenty two)
[0106] in The set representing the fusion weights
[0107] The above are preferred embodiments of the present invention. Any changes made to the technical solution of the present invention that do not exceed the scope of the technical solution of the present invention shall fall within the protection scope of the present invention.
Claims
1. A low-light aerial imaging method for monochrome-infrared image fusion, characterized in that, Includes the following steps: Step S1: Use the monochrome and infrared cameras mounted on the aerial photography equipment to acquire the required paired unregistered low-light monochrome infrared image dataset at low altitude at night. Step S2: Construct the overall structure of ALMIN, an aerial low-light imaging method based on a monochrome-infrared camera; Step S3: Construct an illumination prior extractor in ALMIN. Use the pre-trained model to obtain the local illumination prior for the low-light monochrome image obtained in step S1, and use the local illumination prior to initially enhance the texture details of the low-light monochrome image. Step S4: Construct a dual-modal fusion restorer in ALMIN, use the local illumination prior obtained in step S3 for restoration guidance, integrate monochromatic features and infrared features to reconstruct illumination and reflectivity components, and generate aerial images.
2. The low-light aerial imaging method for monochrome-infrared image fusion according to claim 1, characterized in that, Step S1 is as follows: Step S11: In a low-altitude nighttime environment, use the monochrome and infrared cameras mounted on the aerial photography equipment to collect unregistered monochrome and infrared images. Step S12: Use a pre-trained network model to detect matching feature points in unregistered monochrome and infrared image pairs; then, use the obtained matching points to estimate the homography matrix using the RANSAC algorithm. This matrix is used to perform perspective transformation on the monochrome image to obtain registered monochrome infrared image pairs. Step S13: Divide the obtained n pairs of registered monochrome and infrared images into training set and test set.
3. The low-light aerial imaging method for monochrome-infrared image fusion according to claim 1, characterized in that, In step S3, given a low-light monochrome image In this study, a feature decomposition network based on Retienx theory is used to analyze low-light monochromatic images. Perform illumination component Estimate the estimated illumination components This is used to enhance the details of low-light monochrome images, resulting in a preliminarily enhanced low-light monochrome image. The entire process is described as follows: (1) in This represents the noise introduced into the image during atmospheric propagation. This represents the illumination disturbance factor. This represents the perturbation factor of the reflection component. Representing richly detailed reflection components and This indicates the light component that provides even exposure.
4. The low-light aerial imaging method for monochrome-infrared image fusion according to claim 2, characterized in that, In step S4, the low-light monochromatic image obtained in step S3 is preliminarily enhanced. and infrared images Simultaneously perform noise reduction processing. The process is represented as follows: (2) (3) in, This represents the denoised low-light monochrome image. This represents the denoised infrared image.
5. A low-light aerial imaging method for monochrome-infrared image fusion according to claim 3, characterized in that, In step S4, the dual-modal fusion repair device is used to perform the following steps respectively: and The operation yields the illumination component of the low-light monochromatic image. and reflection component and global features of infrared images With local features The process is represented as follows: (4) (5) (6) (7)。 6. A low-light aerial imaging method for monochrome-infrared image fusion according to claim 4, characterized in that, In step S3, the monochromatic features and infrared features obtained in step S4 are used to estimate the illumination component. Conducting guided integration This allows for the acquisition of well-exposed light components. and rich in texture details The process is represented as follows: (8) (9) (10)。 in This indicates an enhanced low-light aerial image.
7. A low-light aerial imaging method for monochrome-infrared image fusion according to claim 5, characterized in that, This is achieved by designing a frequency domain noise reduction module. Functionality: First, the input image is projected into Q, K, and V using standard convolutional layers; second, Fourier transforms are performed on Q, K, and V to map them from the spatial domain to the frequency domain; third, in the frequency domain, dynamic filters are designed, and pointwise convolution and channel aggregation are performed; finally, an inverse Fourier transform is used to map the frequency domain representation back to the spatial domain. , and This is used to calculate attention scores; the specific formula is as follows: (11) (12) (13) in Indicates a dynamic filter. This indicates multi-head attention calculation. Indicates Fourier transform. Indicates the inverse Fourier transform. This represents the enhanced monochrome image and the initial infrared image acquired by the device. This represents three standard convolutions.
8. A low-light aerial imaging method for monochrome-infrared image fusion according to claim 6, characterized in that, Implemented through a global illumination sensing module. Functionality: First, a bidirectional sequential scan is performed using a multi-head SSD, employing parallel recursive control to extract the illumination component of low-light monochromatic images. Global features of infrared images Perform global feature extraction; as shown below: (14) in The definition is as follows: in This represents a sequence reversal operation. These are learnable parameters; Next, depthwise separable convolution is applied to the processed 2D feature map to model the local neighborhood relationships between feature channels, while adaptively adjusting local features through a gating mechanism.
9. A low-light aerial imaging method for monochrome-infrared image fusion according to claim 7, characterized in that, Implemented through a fine-grained reflectivity sensing module. Function: First, perform wavelet transform on the image. This yields three high-frequency components and one low-frequency component. Subsequently, the three high-frequency components are stacked and summed along the channel dimension to obtain a unified feature representation; Global average pooling Extract global context information and utilize convolution and non-linear activation. Channel-by-channel dimensionality reduction is performed; then, a standard convolutional layer is used to generate an attention vector for each high-frequency component, which is multiplied by the original high-frequency component, thereby simultaneously suppressing noise and enhancing details; at the same time, a reversible neural network module is used for the low-frequency components. The process involves extracting the lighting-invariant detail components from the low-frequency components. (15) (16) (17) (18)。 10. A low-light aerial imaging method for monochrome-infrared image fusion according to claim 8, characterized in that, Achieved through a light-guided fusion module Function; First, standard convolutional layers are applied to estimate the illumination information. Perform a convolution operation to align it with the feature channels of the two modalities; Then, a lightweight multi-level perceptron is introduced to capture multi-scale contextual information and adaptively model the spatial distribution of brightness and contrast variations; finally, the generated fusion weights are normalized to the range [0,1] using a Sigmoid activation function. (19) (20) (21) (22) in This represents the set of fusion weights.