Multi-spectral fusion night low-illumination image enhancement and occlusion compensation method

Through multispectral fusion technology, visible light, infrared, thermal imaging and millimeter wave radar are used to collaboratively collect data to solve the problems of image noise and occlusion in low light at night, generate high-definition and complete target images, and restore the details and contours of dynamic scenes.

CN120689257APending Publication Date: 2025-09-23MINAMI ACOUSTICS LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510635572.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In low-light environments at night, image acquisition equipment is prone to high noise, loss of details, and perception blind spots caused by dynamic occlusion. Existing technologies are difficult to simultaneously solve the problems of low-light noise, loss of details, and information recovery in occluded areas.

Method used

Visible light cameras, infrared sensors, thermal imaging sensors and millimeter-wave radars are used to collaboratively collect multimodal data. Image fusion and occlusion compensation are performed through a multi-scale transformation algorithm and u-Het deep learning network. Combined with temperature distribution information and motion information, the texture details of the occluded area are predicted and filled to generate a complete target image.

Benefits of technology

It significantly improves the clarity and integrity of nighttime low-light images, restores the outline and texture of occluded objects, solves the loss of details caused by occlusion in traditional methods, and achieves high signal-to-noise ratio and seamless reconstruction of occluded areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689257A_ABST
    Figure CN120689257A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses a multispectral fusion night low-illumination image enhancement and shielding compensation method, which comprises the following steps of: cooperatively acquiring multi-modal data through a visible light camera, an infrared sensor, a thermal imaging sensor and a millimeter wave radar; comprising a low-illumination basic image, dark light texture details, target temperature distribution and contour and motion information of an object behind the shelter; fusing visible light and infrared textures by adopting a multi-scale transformation algorithm to generate a transition fusion image, and performing illumination compensation based on temperature distribution; predicting texture details of the occlusion area through a u-Het deep learning network in combination with detection data of the millimeter wave radar; and filling the missing texture according to the shielding contour, and generating a final target image through edge optimization and illumination smoothing technologies. According to the invention, the definition of a night low-illumination image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular to a multi-spectral fusion method for nighttime low-light image enhancement and occlusion compensation. Background Art

[0002] Currently, image acquisition and processing in low-light environments at night face two major technical challenges: first, visible light cameras are prone to high noise, loss of details and insufficient contrast under dark conditions, resulting in a significant decrease in image quality; second, dynamic occlusions (such as trees, buildings or moving objects) will block the target area, and traditional visual sensors (such as cameras or infrared sensors) cannot penetrate the occlusions to obtain the contour and texture information of the occluded object, resulting in perception blind spots.

[0003] In the existing technology, although some studies have enhanced imaging capabilities in low-light environments through infrared sensors or thermal imaging technology, it is difficult to simultaneously solve the problems of low-light noise, missing details and information recovery in occluded areas using a single sensor.

[0004] From the above, we can see that how to improve the clarity of low-light images at night still needs to be solved. Summary of the Invention

[0005] In order to improve the clarity of nighttime low-light images, the present application provides a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method.

[0006] In the first aspect, the present application provides a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method, which adopts the following technical solutions:

[0007] A multispectral fusion method for nighttime low-light image enhancement and occlusion compensation, comprising:

[0008] The visible light camera, infrared sensor, thermal imaging sensor, and millimeter-wave radar collaboratively collect corresponding multimodal data. The multimodal data includes: basic images acquired by the visible light camera in low-light environments, dark light texture details enhanced by the infrared sensor, temperature distribution information of the corresponding target object collected by the thermal imaging sensor, and detection results data from the millimeter-wave radar, which includes the outline and motion information of objects behind the obstruction;

[0009] Based on the base image and the dark light texture details, a multi-scale transformation algorithm is used to fuse them to generate a corresponding transition fusion image, and the transition fusion image is subjected to illumination compensation in combination with the temperature distribution information of the target object; the detection result data and the transition fusion image after illumination compensation are used to predict the predicted texture detail data of the occluded area through the u-Het deep learning network;

[0010] Based on the predicted texture detail data and the outline of the object behind the obstruction in the detection result data, the corresponding missing information is determined, and the missing information is filled into the transition fusion image after illumination compensation. The filled transition fusion image is used to generate the final target image through edge fusion and illumination smoothing technology, wherein the final target image contains the complete outline, texture details, visible light color, infrared enhanced texture and temperature characteristics of the object behind the obstruction.

[0011] Optionally, in the multimodal data acquisition step, the visible light camera, infrared sensor, thermal imaging sensor, and millimeter wave radar are jointly calibrated to achieve spatial and temporal synchronization, and the method further includes:

[0012] The visible light camera and the millimeter wave radar are spatially aligned using Zhang's calibration method, and the conversion error between the pixel coordinate system of the visible light camera and the polar coordinate system of the millimeter wave radar is controlled within 0.5 pixels;

[0013] The temperature distribution information of the thermal imaging sensor is mapped to the grayscale value of the visible light image through a thermal radiation model, and the mapping relationship is iteratively optimized through a Bayesian optimization algorithm;

[0014] The multimodal data is time synchronized via GPS timestamps or external trigger signals.

[0015] Optionally, the multi-scale transformation method further includes:

[0016] When fusing the base image and the dark light texture details, the weight coefficient is dynamically adjusted according to the illumination intensity of the base image and the texture clarity of the dark light texture details:

[0017] When the illumination intensity of the visible light image is lower than the preset threshold, the weight of the dark light texture details in the high-frequency detail band is increased; the weight coefficient is further adjusted according to the gradient intensity of the base image.

[0018] Optionally, the u-Het deep learning network includes:

[0019] Combining the motion information and temperature distribution information of the millimeter-wave radar, the dynamic features of the occluded area are extracted through spatiotemporal convolution; the RGB channels of the transition fusion image are fused with the temperature distribution heat map at the channel level, and the cross-modal feature association of the occluded area is strengthened through the self-attention mechanism; based on the outline of the object behind the occluder and the direction of the temperature gradient, the gradient transfer algorithm is used to optimize the lighting consistency between the predicted texture detail data and the environment.

[0020] Optionally, the edge fusion and illumination smoothing technology method further includes:

[0021] A physical lighting model of the occluded area is constructed through differentiable rendering technology. The physical lighting model is based on the temperature distribution of thermal imaging and the outline of the object behind the occluder, and the lighting consistency of the predicted texture detail data is optimized through backpropagation. The style of the final target image is transferred through a generative adversarial network. The transition fusion image and the predicted texture detail data are jointly optimized, and Gaussian bilateral filtering is used to eliminate lighting discontinuities. The U-Net edge detection model is used to perform sub-pixel optimization on the boundaries of the occluded area.

[0022] Optionally, when processing the motion information detected by the millimeter-wave radar, the method further includes:

[0023] Perform 3D target detection on the millimeter-wave radar's reflected signal and segment the Doppler spectrum of objects behind the obstruction using a point cloud clustering algorithm.

[0024] The radial velocity of the object behind the obstruction is calculated by combining the Doppler shift and the radar carrier frequency. The future trajectory of the object is predicted through joint optimization of Kalman filtering and profile data.

[0025] The future motion trajectory is used as a dynamic input to update the predicted texture detail data of the u-Het network in real time.

[0026] Optionally, during the process of performing illumination compensation on the temperature distribution information, the method further includes:

[0027] The temperature distribution information and the brightness channel of the transition fusion image are constructed into a joint feature map, and the contrast of the high temperature area is amplified and the noise of the low temperature area is suppressed by the exponential function.

[0028] The thermal radiation gradient direction of the thermal imaging data is matched with the gradient direction of the visible light image, and the edge sharpness of the transition fusion image is optimized through the gradient transfer algorithm;

[0029] Dynamically adjust the lighting compensation intensity based on the scene's average temperature and temperature standard deviation.

[0030] In a second aspect, the present application provides a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, which adopts the following technical solutions:

[0031] A multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, comprising:

[0032] A multimodal data acquisition module, which uses a visible light camera, infrared sensor, thermal imaging sensor, and millimeter-wave radar to collaboratively collect corresponding multimodal data. The multimodal data includes: basic images acquired by the visible light camera in low-light environments, dark light texture details enhanced by the infrared sensor, temperature distribution information of the corresponding target object collected by the thermal imaging sensor, and detection results data from the millimeter-wave radar, which includes the outline and motion information of objects behind the obstruction;

[0033] A transition fusion image generation module is configured to fuse the base image and the dark light texture details using a multi-scale transformation algorithm to generate a corresponding transition fusion image, and to perform illumination compensation on the transition fusion image in combination with the temperature distribution information of the target object; and to use the u-Het deep learning network to predict texture detail data of the occluded area using the detection result data and the illumination-compensated transition fusion image.

[0034] The final target image generation module determines the corresponding missing information based on the predicted texture detail data and the outline of the object behind the occluder in the detection result data, fills the missing information into the transition fusion image after illumination compensation, and uses the filled transition fusion image to generate the final target image through edge fusion and illumination smoothing technology, wherein the final target image contains the complete outline, texture details, visible light color, infrared enhanced texture and temperature characteristics of the object behind the occluder.

[0035] In a third aspect, the present application provides a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, which adopts the following technical solutions:

[0036] A multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system includes a processor running a program of any one of the multi-spectral fusion nighttime low-light image enhancement and occlusion compensation methods described above.

[0037] In a fourth aspect, the present application provides a storage medium, which adopts the following technical solution:

[0038] A storage medium stores a program for the multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method described in any one of the above.

[0039] In summary, this application includes at least one of the following beneficial technical effects:

[0040] First, the base image from the visible light camera and the enhanced low-light texture details from the infrared sensor are fused through an improved multi-scale transformation algorithm, dynamically adjusting the weighting coefficients of the visible light and infrared images. In low-light areas, the high-frequency details of the infrared image are weighted more highly, enhancing texture clarity in low-light scenes. Simultaneously, the temperature distribution information from the thermal image is combined for illumination compensation. By optimizing the correlation between temperature gradients and image brightness, the contrast in high-temperature areas is amplified and low-temperature noise is suppressed, further reducing image blur and noise in low-light environments. This process not only preserves the color information of visible light but also, through the complementary properties of infrared and thermal imaging, significantly improves the overall image detail visibility and signal-to-noise ratio, resolving the problem of distortion in traditional single-sensor imaging in low light.

[0041] To address the problem of dynamic occlusion, the contours and motion information of objects behind the occluders detected by millimeter-wave radar are combined with the predictive capabilities of the u-Het deep learning network to achieve accurate reconstruction of the texture details of the occluded area through multimodal feature fusion and physical constraint optimization. For example, the u-Het network extracts the motion characteristics of dynamic occluders through spatiotemporal convolution, and uses the temperature gradient of thermal imaging to match the direction of the visible light gradient to ensure the consistency of the predicted texture with the lighting of the surrounding environment. In addition, edge fusion and illumination smoothing techniques (such as Gaussian bilateral filtering and U-Net boundary optimization) further eliminate the boundary differences between the reconstructed area and the original image, so that the details of the occluded area and the non-occluded area of ​​the final image are seamlessly connected. This process not only restores the contours and textures of the occluded objects, but also avoids the loss of details caused by occlusion in traditional methods through the synergistic effect of multi-sensor data, thereby comprehensively improving the clarity and integrity of night images. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 The present invention is a flowchart showing a method for nighttime low-light image enhancement and occlusion compensation using multi-spectral fusion according to an exemplary embodiment.

[0043] Figure 2 The figure is a structural block diagram of a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system according to an exemplary embodiment. DETAILED DESCRIPTION

[0044] Embodiments of the present application are described in detail below, examples of which are illustrated in the accompanying drawings.

[0045] Throughout this specification, reference to the terms "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0046] The present application embodiment discloses a multi-spectral fusion night low-light image enhancement and occlusion compensation method, referring to Figure 1 ,include:

[0047] S100 uses a visible light camera, infrared sensor, thermal imaging sensor, and millimeter-wave radar to collaboratively collect corresponding multimodal data. The multimodal data includes: basic images acquired by the visible light camera in low-light environments, dark light texture details enhanced by the infrared sensor, temperature distribution information of the corresponding target object collected by the thermal imaging sensor, and detection result data from the millimeter-wave radar. The detection result data includes the outline and motion information of objects behind the obstruction.

[0048] The visible light camera uses a low-light-optimized sensor (such as a high-sensitivity CMOS) and supports manual or automatic adjustment of exposure time, gain, and white balance. In low-light environments (such as at night, in tunnels, and in shadowed areas), the camera acquires the corresponding base image by extending the exposure time (e.g., from 1 / 10 second to 1 second) and increasing the gain. Dynamic adjustments are made based on ambient light intensity to ensure that dark areas are not over- or under-exposed, and color temperature is adjusted based on ambient light sources (such as streetlights or moonlight) to avoid color deviation. It should be noted that the base image contains visible light color information, but may contain noise and loss of detail due to low light conditions.

[0049] The infrared sensor operates in the near-infrared band (850-940nm) and is equipped with a filter that allows only infrared light to pass, preventing interference from visible light. Through active infrared illumination (such as invisible infrared fill light) or passive reception of ambient infrared radiation, texture details in dimly lit areas are enhanced. The sensor outputs a grayscale image, with highly reflective areas (such as metal and light-colored objects) appearing bright and low-reflective areas (such as dark objects and shadows) appearing dark. Automatically adjusting to ambient lighting, it avoids overexposure and blurred object outlines, enhancing dark texture detail and compensating for the low contrast of visible light images.

[0050] Thermal imaging sensors operate in the long-wave infrared band (8-14μm) and generate temperature distribution images by detecting thermal radiation from an object. The sensor receives thermal radiation from the target object's surface and converts it into a matrix of temperature values ​​(T). With a temperature resolution of up to 0.1°C, it can distinguish subtle temperature differences on the surface of an object (such as the temperature difference between the human body, engine, and the environment). The temperature distribution information is presented as a heat map to assist with illumination compensation and texture prediction.

[0051] Millimeter-wave radar operates in the 24-77GHz frequency band, has Doppler effect analysis capabilities, and can penetrate obstructions (such as trees and rocks). The radar transmits electromagnetic waves and receives reflected signals, which are processed through the following steps: calculating the distance to the object through the pulse echo time difference and the radial velocity through the Doppler frequency shift; performing point cloud clustering on the reflected signal to segment the outline of the object behind the obstruction; and combining continuous frame data with Kalman filtering to predict the future trajectory of the object. The detection result data includes the outline, motion trajectory and distance of the object behind the obstruction.

[0052] It should be noted here that in the multimodal data acquisition step, the visible light camera, infrared sensor, thermal imaging sensor, and millimeter wave radar are jointly calibrated to achieve spatial and temporal synchronization. The specific method also includes:

[0053] At step S110, the visible light camera and the millimeter-wave radar are spatially aligned using the Zhang calibration method, and the conversion error between the visible light camera's pixel coordinate system and the millimeter-wave radar's polar coordinate system is controlled within 0.5 pixels.

[0054] The specific implementation process includes: using a standard checkerboard calibration plate (black and white, size such as 20×15cm), whose geometric structure is known (such as the side length of each grid is 1cm); placing the calibration plate in the camera's field of view, taking multiple visible light images at different angles and positions, and recording the pixel coordinates of the calibration plate's corner points; placing the calibration plate within the radar detection range, recording the polar coordinate data (distance and angle) of the radar reflection signal, and calculating the polar coordinates of the calibration plate's center point.

[0055] The camera's focal length, principal point coordinates, and distortion parameters (such as barrel or pincushion distortion caused by the lens) are calculated using Zhang's calibration method. The calibration plate's 3D coordinates in the world coordinate system are mapped to the pixel coordinates of visible light, and the camera's extrinsic parameters (rotation angle and translation distance) are calculated using perspective transformation (PnP algorithm). The radar's polar coordinates are converted to 3D coordinates in the world coordinate system, and the relationship between the radar and camera's extrinsic parameters (rotation angle and translation distance) is fitted using the least squares method.

[0056] Then, a conversion rule is established between the camera pixel coordinate system and the radar polar coordinate system. For example, the distance and angle of objects detected by the radar are converted to pixel positions in the visible light image. The conversion error is verified using calibration plate data. If the error exceeds 0.5 pixels, the extrinsic parameters (such as rotation angle and translation) are iteratively adjusted until the error is ≤ 0.5 pixels. This ensures that the visible light image and the outline of the object behind the occluder detected by the radar are precisely aligned in space, avoiding fusion misalignment caused by coordinate system deviation. In addition, the outline of the object detected by the millimeter-wave radar can be directly mapped to the pixel position of the visible light image, improving the accuracy of subsequent texture prediction of the occluded area.

[0057] S120 , establishing a mapping relationship between the temperature distribution information of the thermal imaging sensor and the grayscale value of the visible light image through a thermal radiation model, and iteratively optimizing the mapping relationship through a Bayesian optimization algorithm.

[0058] The specific implementation process includes: Based on thermal radiation theory, a relationship is established between an object's surface temperature and radiated energy: hotter objects radiate more infrared light, while cooler objects radiate less. It is also assumed that there is a nonlinear relationship between the grayscale values ​​of the visible light image and the temperature distribution of the thermal image (for example, high-temperature areas may appear brighter in visible light).

[0059] In the same scene, visible light images and thermal imaging temperature distributions are collected simultaneously to form a training dataset. The parameter space is initialized, for example, the mapping relationship between temperature and grayscale values ​​may include coefficients, exponents, and other parameters. By continuously trying different parameter combinations, the parameters that closest the predicted grayscale value to the actual grayscale value are found. A Gaussian process is used to predict the error of the current parameter combination, and the parameter that minimizes the error is selected as the sampling point for the next iteration. The iteration is repeated until the error drops below a threshold. After obtaining the optimal parameter combination, a mapping rule between thermal radiation and visible light grayscale is formed, for example, "the higher the temperature, the brighter the corresponding visible light grayscale value."

[0060] A quantitative relationship is established between the temperature distribution of thermal imaging and the grayscale value of visible light. For example, high-temperature areas (such as heated objects) may appear as high grayscale values ​​(bright areas) in visible light, while low-temperature areas (shadows) have lower grayscale values. This can also provide a physical correlation between temperature and image brightness for subsequent steps, such as adjusting the contrast of visible light images through temperature gradients to enhance details in dark areas.

[0061] S130 , the multimodal data is time synchronized through a GPS timestamp or an external trigger signal.

[0062] The specific implementation process of the hardware synchronization solution includes:

[0063] GPS timestamp: equip each sensor with a GPS module to record the UTC time of each frame of data (accuracy ≤ 1ms); external trigger signal: synchronize the collection time of all sensors through the hardware trigger line (such as GPIO signal), for example, the main control chip sends a pulse signal to trigger the start of collection at the same time.

[0064] Timestamp matching: For GPS solutions, align the timestamps of each sensor with the master time source to ensure that the time difference is ≤5ms. For trigger signal solutions, ensure that the acquisition start time of all sensors is consistent by recording the arrival time of the trigger signal.

[0065] Data registration: Pair multimodal data in chronological order according to timestamps (such as visible light images, infrared images, thermal imaging heat maps, and radar point clouds at the same moment); for frames with time differences, compensate through interpolation or nearest neighbor method.

[0066] Verify synchronization through scenario testing: For example, if the visible light image and the occlusion outline detected by the radar are not synchronized in time, it may lead to incorrect positioning of the occlusion area. Ensure that the time difference is ≤5ms by adjusting the trigger signal delay or GPS clock offset.

[0067] Spatial alignment ensures precise physical matching of visible light and radar data, providing a reliable coordinate basis for contour positioning and texture prediction of occluded areas; thermal-optical correlation establishes a physical mapping between temperature distribution and visible light grayscale values ​​through thermal radiation model and Bayesian optimization, enhancing detail visibility in dark areas and optimizing illumination compensation; time synchronization ensures strict consistency of multimodal data in the time dimension, avoiding information dislocation in dynamic scenes, and ultimately providing a high-precision, real-time multimodal data basis for image enhancement and occlusion compensation in low-light environments at night, significantly improving target recognition and imaging quality in complex scenes.

[0068] S200, based on the basic image and dark light texture details through a multi-scale transformation algorithm to fuse, generate the corresponding transition fusion image, combined with the temperature distribution information of the target object to perform illumination compensation on the transition fusion image; the detection result data and the transition fusion image after illumination compensation are used to predict the predicted texture detail data of the occluded area through the u-Het deep learning network.

[0069] The fusion process of the multi-scale transformation algorithm includes:

[0070] 1. Multi-scale decomposition: The visible light base image and dark light texture details are decomposed at multiple scales (e.g., using wavelet transform or high-pass / low-pass filtering) to extract high-frequency details (e.g., edges and textures) and low-frequency background (overall brightness and structure). The high-frequency details preserve the image's texture and edge information (e.g., details in dark areas of infrared images); the low-frequency background preserves the image's overall brightness and color information (e.g., the color of visible light).

[0071] 2. Dynamic Weight Coefficient Adjustment: The average brightness value of the base image corresponding to visible light (e.g., the mean of the RGB channels) is calculated. If it is below a preset threshold (e.g., 50 / 255), it is determined to be a "low-light area." In low-light areas, the high-frequency details of the dark texture details corresponding to the infrared image are weighted more (e.g., a weight coefficient ≥ 0.7) to enhance texture clarity in the dark areas.

[0072] The gradient intensity (edge ​​sharpness) of the visible light image is calculated using an edge detection algorithm (such as the Sobel operator). Regions with high gradient intensity (such as object outlines) retain more visible light color information, while regions with low gradient intensity (such as dark, smooth areas) receive a higher weight for infrared texture. The weight coefficient follows the logical formula: high-frequency weight = infrared weight × gradient intensity compensation factor (the compensation factor increases when the gradient intensity is low).

[0073] 3. A weighted fusion strategy is used to combine the low-frequency base of visible light with the high-frequency details of infrared light. In the low-frequency portion, visible light brightness is prioritized, preserving color information. In the high-frequency portion, dynamic weighting coefficients are used to balance infrared texture with visible light details. The resulting transition fusion image is then output to generate a high-definition, low-noise intermediate image that preserves both visible light color and infrared texture details.

[0074] The process of illumination compensation (combined with temperature distribution information) includes:

[0075] 1. Temperature and illumination association: The temperature distribution information of the thermal imaging is combined with the brightness channel of the transition fusion image to form a joint feature map.

[0076] 2. Temperature gradient analysis: High-temperature areas (such as heating objects) generally require contrast enhancement (e.g., increasing brightness); low-temperature areas (such as shadows) require noise suppression (e.g., reducing brightness fluctuations).

[0077] 3. Compensation Strategy: In high-temperature areas, histogram equalization or contrast stretching is used to enhance detail visibility; in low-temperature areas, Gaussian filtering or median filtering is used to reduce noise. The compensation strength is dynamically adjusted based on the scene's average temperature and temperature standard deviation (e.g., higher compensation in high-temperature areas).

[0078] 4. Output the illumination-compensated image: preserve the texture and color of the transition-fused image while optimizing the contrast and signal-to-noise ratio in the dark areas.

[0079] Among them, the texture prediction corresponding process of the u-Het deep learning network includes:

[0080] 1. Multimodal input:

[0081] Transitional fused image after illumination compensation: contains visible light color information and infrared enhanced texture details, and is a high-quality intermediate image optimized through illumination compensation.

[0082] Millimeter-wave radar detection data: occlusion outline, the boundary position of objects behind the occlusion detected by the millimeter-wave radar (such as the outline of vehicles and pedestrians); motion information (M), the radial velocity and motion trajectory of the object (such as the future position predicted by the Kalman filter).

[0083] The occlusion contours and motion trajectories are converted into binary masks or heat maps that match the resolution of the transition fusion image; the temperature distribution information (T) of the thermal imaging is input in the form of a heat map and aligned with the RGB channels of the transition fusion image.

[0084] 2. Network structure and process:

[0085] First, feature extraction and multi-scale fusion are performed. Convolution kernels of different sizes are used to extract multi-scale texture features from the transition-fused image: small convolution kernels (3×3) capture local details (such as edges and textures), while large convolution kernels (5×5 and 7×7) extract global structure (such as object shape and regional features). Multi-scale feature maps are concatenated into composite features to enhance the network's ability to perceive textures in occluded areas.

[0086] Next, dynamic occlusion feature enhancement is performed: The millimeter-wave radar motion information and occlusion outline are input, and the dynamic features of the occluded area are extracted through spatiotemporal convolution (3D convolution). In the spatial dimension, the pixel position of the occluded area is located based on the occlusion outline (C). In the temporal dimension, the motion trajectory (M) of consecutive frames is combined to predict the future position and dynamic changes of objects behind the occluder. Attention mechanisms (such as channel attention or spatial attention) are used to enhance the feature weights of the occluded area, ensuring that the network focuses on dynamically changing occluded areas.

[0087] Furthermore, cross-modal feature fusion is performed: channel-level fusion of RGB and temperature distribution, where the RGB channels of the transition-fused image are concatenated with the temperature distribution heatmap (T) of the thermal imaging. Through a self-attention mechanism (such as the Transformer block), correlations between different channels are calculated, such as the correlation between the texture features of high-temperature areas (such as the hood) and the metal material in visible light. This can strengthen cross-modal feature correlations in occluded areas (such as matching temperature and material).

[0088] Furthermore, semantic consistency optimization is performed: temperature gradient and lighting consistency. According to the temperature gradient direction of thermal imaging (such as the steepness of temperature change at the edge of the high-temperature area), the edge sharpness of the predicted texture is adjusted through the gradient transfer algorithm. This ensures that the lighting of the predicted texture is consistent with the surrounding environment (for example, the brightness of the high-temperature area is higher than the low-temperature background).

[0089] Finally, the predicted texture detail data is output: a texture detail map is generated, and the network outputs a high-resolution texture detail map (P) that matches the occluded area, including information such as the surface material of the object (such as metal, fabric), edge contours, etc.; the predicted texture is aligned with the original image resolution through an upsampling layer (such as transposed convolution).

[0090] Through multi-scale transformation algorithms and dynamic weight adjustment, the texture clarity of dark areas is significantly improved in low-light environments while retaining the color information of visible light. The temperature distribution of thermal imaging is combined for illumination compensation, dynamically optimizing the contrast of high-temperature areas and noise suppression in low-temperature areas to ensure global visual consistency. Through the u-Het deep learning network, the occlusion contour and motion information of the millimeter-wave radar, as well as temperature-driven semantic associations, are used to accurately predict and reconstruct the texture details of the occluded area, ultimately generating high-quality images with high definition, natural colors, and complete occlusion information.

[0091] S300, based on the predicted texture detail data and the outline of the object behind the occluder in the detection result data, the corresponding missing information is determined, the missing information is filled into the transition fusion image after illumination compensation, and the filled transition fusion image is used to generate the final target image through edge fusion and illumination smoothing technology.

[0092] The final target image contains the complete outline of the object behind the occluder, texture details, visible light color, infrared enhanced texture and temperature characteristics of thermal imaging.

[0093] First, determine and fill in missing information, including:

[0094] 1. Occlusion area positioning: Based on the detection result data of the millimeter-wave radar (occlusion outline and motion information), the pixel position of the occlusion area is marked in the transition fusion image after illumination compensation.

[0095] 2. Definition of missing information: Texture details and color information in the occluded area are missing and need to be filled by texture details predicted by the u-Het network.

[0096] 3. Filling strategy: Texture matching, matching the predicted texture details with the surrounding environment (such as color and lighting) of the occluded area to ensure that the filled texture is consistent with the surrounding style; spatial alignment, according to the boundary of the occlusion outline, the predicted texture is accurately filled into the occluded area to avoid exceeding or missing.

[0097] 4. Output filled image: Generates an intermediate image that contains texture details of the occluded area, but may have edge discontinuities or sudden changes in lighting.

[0098] Then, for edge fusion and lighting smoothing technology.

[0099] 1. Edge optimization:

[0100] The first is sub-pixel edge detection (U-Net model). The specific execution process includes: inputting the padded transition fusion image, in which the occluded areas have been filled with the texture (P) predicted by the u-Het network; extracting multi-scale edge features through the encoder-decoder structure of the U-Net edge detection model, and outputting a high-resolution edge map (such as a 1-pixel wide boundary line);

[0101] Upsampling layers (such as transposed convolution) are used to increase the edge map resolution to the same level as the original image, accurately distinguishing the boundaries between occluded and unoccluded areas. Pixels along the occluded edges need to be sharpened (such as with high-pass filtering) to improve edge clarity, and blur or burrs need to be removed (such as through non-local mean denoising). This ensures that the boundary between the occluded area and the surrounding environment is clear and natural, avoiding "jaggies" or "blurring" effects.

[0102] Next comes gradient matching (thermal imaging and visible light). The specific execution process includes: thermal imaging temperature gradient, calculating the gradient direction of the thermal imaging temperature distribution (T) (such as the steepness of temperature change at the edge of the high-temperature area); visible light gradient direction, extracting the gradient direction of the visible light image through the Sobel operator or Canny edge detection; angle matching: aligning the temperature gradient direction of the thermal imaging with the gradient direction of the visible light to ensure that the texture gradient at the occluded edge is consistent with the surrounding environment. For example, the temperature gradient direction at the edge of the high-temperature area should match the brightness change direction at the visible light edge. This can eliminate texture mutations at the edge of the occluded area and make the transition between the predicted texture and the texture of the surrounding environment natural.

[0103] 2. Lighting consistency optimization:

[0104] The first is differentiable rendering and physical lighting model. The specific execution process includes: lighting model construction, based on the temperature distribution of thermal imaging and the occlusion contour of millimeter-wave radar, building a physical lighting model of the occluded area. The surface reflectivity and light intensity of high-temperature areas (such as heating objects) should be higher than the low-temperature background; the shadow relationship between the occluder and the object behind it is defined according to the occlusion contour.

[0105] Backpropagation optimization involves inputting the predicted texture into the illumination model and calculating the difference in illumination (such as brightness and contrast) between the predicted texture and the surrounding environment. Backpropagation is then used to adjust parameters such as the reflectivity and shadow distribution of the predicted texture to align with global illumination conditions. This ensures that the brightness and contrast of the occluded area match the surrounding environment, avoiding artifacts such as "too bright" or "too dark."

[0106] Next, Gaussian bilateral filtering is performed. The specific implementation process involves applying Gaussian bilateral filtering at the junction of occluded and unoccluded areas. In the spatial domain, this filter preserves edge structure (such as the boundaries detected by U-Net) and in the intensity domain, it smooths out lighting differences (such as the brightness gradient between the occluded area and the background). By adjusting filter parameters (such as spatial radius and intensity threshold) to balance lighting continuity and edge sharpness, "seam" artifacts can be eliminated, resulting in a natural transition between the occluded area and the surrounding lighting.

[0107] 3. Style transfer and detail enhancement

[0108] The first is generative adversarial network style transfer. The specific execution process includes: the input is the padded image and the intermediate result after lighting optimization. It should be pointed out here that the role of the generative adversarial network is: it can preserve style, through adversarial training, to ensure that the output image retains the original color style of visible light (such as white balance and color temperature); it can eliminate artifacts, with the generator learning to predict high-frequency details of the texture, eliminating reconstruction artifacts such as block noise and texture distortion; it can also provide feedback to the discriminator to distinguish between real images and generated images, and optimize the fidelity of the generator output. This can improve the visual realism of the predicted texture and make it consistent with the style of the original image.

[0109] Next comes joint optimization and detail enhancement. The specific implementation process includes: simultaneously optimizing the transition-fused image and predicted texture to ensure global consistency in color, texture, and lighting; balancing edge sharpness, lighting matching, and detail realism through loss functions (such as perceptual loss and gradient loss); and using a super-resolution network to increase the resolution of textures in occluded areas so that their detail levels match those of unoccluded areas. This further eliminates residual artifacts and improves the detail clarity and overall quality of the final image.

[0110] Through the coordinated optimization of edge fusion and illumination smoothing techniques, U-Net edge detection and gradient matching ensure that the boundaries of occluded areas are clear and naturally transition with the surrounding texture, eliminating edge blur and burrs. The differentiable rendering model combines thermal imaging temperature distribution with Gaussian bilateral filtering to optimize illumination consistency in occluded areas, avoiding sudden brightness changes and "seam" artifacts. The generative adversarial network preserves the original color style of visible light while eliminating reconstruction noise in the predicted texture, enhancing the realism of details. The resulting image not only contains the complete outline of the occluded area, material details, and temperature characteristics of the thermal imaging, but also seamlessly integrates multimodal information, providing high-precision and highly realistic visual output for applications such as target recognition and obstacle detection in complex environments.

[0111] When processing the motion information detected by the millimeter-wave radar in the embodiment of the present application, the method further includes:

[0112] Firstly, three-dimensional target detection is performed on the reflected signal of the millimeter-wave radar, and the Doppler spectrum of the object behind the occluder is segmented using the point cloud clustering algorithm.

[0113] Millimeter-wave radar transmits electromagnetic waves and receives reflected signals, generating three-dimensional point cloud data containing the object's position and reflection intensity. An algorithm analyzes the changes in reflection frequency (Doppler effect) of different objects in the point cloud, distinguishing between obstructions and moving objects behind them and segmenting the area behind them into a separate area. This allows the position and speed of objects behind obstructions to be accurately located, avoiding missed detections due to obstructions.

[0114] Then, the radial velocity of the object behind the obstruction is calculated by combining the Doppler shift and the radar carrier frequency, and the future motion trajectory of the object is predicted through joint optimization of Kalman filtering and contour data.

[0115] The method calculates the object's velocity based on the frequency variation of electromagnetic wave reflections (the Doppler effect), and combines this with a Kalman filter algorithm to predict the object's future position and path. Furthermore, it incorporates contour data from visible light or thermal imaging to optimize predictions and reduce errors. This improves the accuracy of object motion prediction and provides a real-time motion reference for texture prediction.

[0116] Furthermore, the future motion trajectory is used as dynamic input to update the predicted texture detail data of the u-Het network in real time.

[0117] The predicted object motion trajectory (e.g., direction and speed) is fed into a deep learning network (u-Het network) in real time, dynamically adjusting the rules for generating texture details in occluded areas. For example, if an object accelerates, the network updates the texture of the occluded area to match the dynamic change. This ensures that the texture prediction is consistent with the object's actual motion, avoiding prediction errors caused by object movement.

[0118] Through millimeter-wave radar's three-dimensional target detection and point cloud clustering and segmentation, the system can accurately identify and locate the position and velocity of objects behind obstructions; combined with the Doppler effect and Kalman filter algorithm, it can predict the future motion trajectory of the object in real time, improving the accuracy of motion prediction in dynamic environments; at the same time, the predicted trajectory information is dynamically input into the deep learning network (such as the u-Het network), and the texture detail generation rules of the obstructed area are adjusted in real time to ensure that the texture prediction is synchronized with the actual movement of the object.

[0119] In the embodiment of the present application, during the process of performing illumination compensation based on temperature distribution information, the method further includes:

[0120] Firstly, the temperature distribution information and the brightness channel of the transition fusion image are constructed as a joint feature map, and the contrast of the high temperature area is amplified and the noise of the low temperature area is suppressed by an exponential function.

[0121] The system combines the temperature distribution map of the thermal image (high-temperature areas are bright, low-temperature areas are dim) with the brightness channel (grayscale image) of the transition-fused image into a joint feature map. High-temperature area processing uses exponential amplification to enhance contrast in hot areas (such as heated objects), making details (such as edges and textures) clearer. Low-temperature area processing uses exponential decay to reduce brightness fluctuations in low-temperature areas (such as shadows and cold backgrounds), suppressing noise in low light conditions. This results in brighter, more detailed high-temperature areas and cleaner, less noisy low-temperature areas.

[0122] Then, the thermal radiation gradient direction of the thermal imaging data is matched with the gradient direction of the visible light image, and the edge sharpness of the transition fusion image is optimized through the gradient transfer algorithm;

[0123] The "thermal radiation gradient direction" (the direction of temperature change at the edge of the high-temperature region) is calculated from the temperature changes in the thermal image, and the "brightness gradient direction" (the direction of the visible light edge) is calculated from the visible light image. The thermal radiation gradient direction and the visible light gradient direction are then aligned to ensure that the two edge directions are consistent. Based on the matched gradient directions, an algorithm is used to enhance the edge sharpness of the transition-fused image (e.g., sharper edges in high-temperature regions and smoother edges in low-temperature regions). This eliminates blur or artifacts caused by the inconsistency between the temperature and visible light edge directions, resulting in a natural edge transition.

[0124] Furthermore, the illumination compensation intensity is dynamically adjusted according to the scene average temperature and temperature standard deviation.

[0125] This is achieved by calculating the average temperature (overall ambient temperature) and temperature fluctuation (temperature differences) of the current scene. It's important to note that in high-temperature, high-fluctuation scenes (such as near a vehicle engine), contrast compensation is enhanced to make the high-temperature areas stand out more clearly. In low-temperature, low-fluctuation scenes (such as on a cold night), contrast enhancement is reduced to avoid overexposure and enhance noise suppression. Light compensation adapts to different ambient temperatures, avoiding a "one-size-fits-all" approach that results in overbrightness or darkness.

[0126] By using illumination compensation technology based on temperature distribution information, the system enhances contrast in high-temperature areas (for example, highlighting details of heat-generating objects like vehicle engines and human bodies), suppresses noise in low-temperature areas (for example, reducing granular noise in shadows or low-light environments), and optimizes image edge sharpness by matching thermal radiation with the direction of visible light edges. The system also dynamically adjusts the compensation intensity based on the scene temperature to ensure consistent illumination across different environments. The resulting image is crisp in high-temperature areas and clear in low-temperature areas, providing high-quality, highly stable visual input for applications such as object recognition and occlusion reconstruction in complex scenes.

[0127] The present application embodiment discloses a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, referring to Figure 2 ,include:

[0128] Multimodal data acquisition module 001, which uses a visible light camera, infrared sensor, thermal imaging sensor, and millimeter wave radar to collaboratively collect corresponding multimodal data. The multimodal data includes: basic images acquired by the visible light camera in low-light environments, dark light texture details enhanced by the infrared sensor, temperature distribution information of the corresponding target object collected by the thermal imaging sensor, and detection result data from the millimeter wave radar, which includes the outline and motion information of objects behind the obstruction;

[0129] The transition fusion image generation module 002 is based on the fusion of the basic image and the dark light texture details through a multi-scale transformation algorithm to generate a corresponding transition fusion image, and the transition fusion image is light-compensated in combination with the temperature distribution information of the target object; the detection result data and the transition fusion image after light compensation are used to predict the predicted texture detail data of the occluded area through the u-Het deep learning network;

[0130] The final target image generation module 003 determines the corresponding missing information based on the predicted texture detail data and the outline of the object behind the occluder in the detection result data, fills the missing information into the transition fusion image after illumination compensation, and uses the filled transition fusion image to generate the final target image through edge fusion and illumination smoothing technology, wherein the final target image contains the complete outline of the object behind the occluder, texture details, visible light color, infrared enhanced texture and temperature characteristics of thermal imaging.

[0131] An embodiment of the present application also discloses a multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, comprising a processor in which a program of any one of the multi-spectral fusion nighttime low-light image enhancement and occlusion compensation methods described above is run.

[0132] An embodiment of the present application also discloses a storage medium storing a program of any one of the multi-spectral fusion nighttime low-light image enhancement and occlusion compensation methods described above.

[0133] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A multi-spectral fusion method for nighttime low-light image enhancement and occlusion compensation, characterized in that: include: The visible light camera, infrared sensor, thermal imaging sensor, and millimeter-wave radar collaboratively collect corresponding multimodal data. The multimodal data includes: basic images acquired by the visible light camera in low-light environments, dark light texture details enhanced by the infrared sensor, temperature distribution information of the corresponding target object collected by the thermal imaging sensor, and detection results data from the millimeter-wave radar, which includes the outline and motion information of objects behind the obstruction; Based on the base image and the dark light texture details, a multi-scale transformation algorithm is used to fuse them to generate a corresponding transition fusion image, and the transition fusion image is subjected to illumination compensation in combination with the temperature distribution information of the target object; the detection result data and the transition fusion image after illumination compensation are used to predict the predicted texture detail data of the occluded area through the u-Het deep learning network; Based on the predicted texture detail data and the outline of the object behind the obstruction in the detection result data, the corresponding missing information is determined, and the missing information is filled into the transition fusion image after illumination compensation. The filled transition fusion image is used to generate the final target image through edge fusion and illumination smoothing technology, wherein the final target image contains the complete outline, texture details, visible light color, infrared enhanced texture and temperature characteristics of the object behind the obstruction.

2. The multispectral fusion nighttime low-light image enhancement and occlusion compensation method according to claim 1, characterized in that: In the multimodal data acquisition step, the visible light camera, infrared sensor, thermal imaging sensor, and millimeter wave radar are jointly calibrated to achieve spatial and temporal synchronization. The method further includes: The visible light camera and the millimeter wave radar are spatially aligned using Zhang's calibration method, and the conversion error between the pixel coordinate system of the visible light camera and the polar coordinate system of the millimeter wave radar is controlled within 0.5 pixels; The temperature distribution information of the thermal imaging sensor is mapped to the grayscale value of the visible light image through a thermal radiation model, and the mapping relationship is iteratively optimized through a Bayesian optimization algorithm; The multimodal data is time synchronized via GPS timestamps or external trigger signals.

3. The multispectral fusion nighttime low-light image enhancement and occlusion compensation method according to claim 1, characterized in that: The multi-scale transformation method further includes: When fusing the base image and the dark light texture details, the weight coefficient is dynamically adjusted according to the illumination intensity of the base image and the texture clarity of the dark light texture details: When the illumination intensity of the visible light image is lower than the preset threshold, the weight of the dark light texture details in the high-frequency detail band is increased; the weight coefficient is further adjusted according to the gradient intensity of the base image.

4. The multispectral fusion nighttime low-light image enhancement and occlusion compensation method according to claim 1, characterized in that: The u-Het deep learning network includes: Combining the motion information and temperature distribution information of the millimeter-wave radar, the dynamic features of the occluded area are extracted through spatiotemporal convolution; the RGB channels of the transition fusion image are fused with the temperature distribution heat map at the channel level, and the cross-modal feature association of the occluded area is strengthened through the self-attention mechanism; based on the outline of the object behind the occluder and the direction of the temperature gradient, the gradient transfer algorithm is used to optimize the lighting consistency between the predicted texture detail data and the environment.

5. The multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method according to claim 4, characterized in that: The edge fusion and illumination smoothing technology and method further include: A physical lighting model of the occluded area is constructed through differentiable rendering technology. The physical lighting model is based on the temperature distribution of thermal imaging and the outline of the object behind the occluder, and the lighting consistency of the predicted texture detail data is optimized through backpropagation. The style of the final target image is transferred through a generative adversarial network. The transition fusion image and the predicted texture detail data are jointly optimized, and Gaussian bilateral filtering is used to eliminate lighting discontinuities. The U-Net edge detection model is used to perform sub-pixel optimization on the boundaries of the occluded area.

6. The multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method according to claim 5, characterized in that: When processing the motion information detected by the millimeter wave radar, the method further includes: Perform 3D target detection on the millimeter-wave radar's reflected signal and segment the Doppler spectrum of objects behind the obstruction using a point cloud clustering algorithm. The radial velocity of the object behind the obstruction is calculated by combining the Doppler shift and the radar carrier frequency. The future trajectory of the object is predicted through joint optimization of Kalman filtering and profile data. The future motion trajectory is used as dynamic input to update the predicted texture detail data of the u-Het network in real time.

7. The multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method according to claim 6, characterized in that: During the process of performing illumination compensation using the temperature distribution information, the method further includes: The temperature distribution information and the brightness channel of the transition fusion image are constructed into a joint feature map, and the contrast of the high temperature area is amplified and the noise of the low temperature area is suppressed by the exponential function. The thermal radiation gradient direction of the thermal imaging data is matched with the gradient direction of the visible light image, and the edge sharpness of the transition fusion image is optimized through the gradient transfer algorithm; Dynamically adjust the lighting compensation intensity based on the scene's average temperature and temperature standard deviation.

8. A multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, characterized in that: include: A multimodal data acquisition module, which uses a visible light camera, infrared sensor, thermal imaging sensor, and millimeter-wave radar to collaboratively collect corresponding multimodal data. The multimodal data includes: basic images acquired by the visible light camera in low-light environments, dark light texture details enhanced by the infrared sensor, temperature distribution information of the corresponding target object collected by the thermal imaging sensor, and detection results data from the millimeter-wave radar, which includes the outline and motion information of objects behind the obstruction; A transition fusion image generation module is configured to fuse the base image and the dark light texture details using a multi-scale transformation algorithm to generate a corresponding transition fusion image, and to perform illumination compensation on the transition fusion image in combination with the temperature distribution information of the target object; and to use the u-Het deep learning network to predict texture detail data of the occluded area using the detection result data and the illumination-compensated transition fusion image. The final target image generation module determines the corresponding missing information based on the predicted texture detail data and the outline of the object behind the occluder in the detection result data, fills the missing information into the transition fusion image after illumination compensation, and uses the filled transition fusion image to generate the final target image through edge fusion and illumination smoothing technology, wherein the final target image contains the complete outline, texture details, visible light color, infrared enhanced texture and temperature characteristics of the object behind the occluder.

9. A multi-spectral fusion nighttime low-light image enhancement and occlusion compensation system, characterized in that: The method comprises a processor running a program of the multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: A program is stored for the multi-spectral fusion nighttime low-light image enhancement and occlusion compensation method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Object recognition method and related device

    CN121074530A

  • Marine ship intelligent detection and three-dimensional positioning method based on multi-modal data cooperation

    CN121259288A

  • Visual feature extraction method and system based on multi-modal image enhancement

    CN121437908A

  • Visual feature extraction method and system based on multi-modal image enhancement

    CN121437908B

  • Panoramic depth and high dynamic range imaging system and method based on compound eye structure

    CN121563855A