High-precision image processing method and system based on illumination adaptive compensation
By combining multi-scale residual networks and generative adversarial networks with ambient light sensors and scene semantic segmentation models, the dynamic range and detail preservation problems of image processing under complex lighting conditions are solved, achieving high-precision adaptive lighting compensation that can adapt to real-time compensation and stability in different lighting scenarios.
Patent Information
- Application Number
- CN202511185094.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies struggle to handle complex lighting variations, resulting in insufficient dynamic range and extreme lighting handling in image processing, weak algorithm generalization ability, and impact on image detail preservation and segmentation accuracy.
A high-precision image processing method based on adaptive illumination compensation is adopted. The image is decomposed by multi-scale residual network, combined with ambient light sensor and scene semantic segmentation model to generate dynamic compensation parameters. The dynamic range expansion of illumination layer and noise suppression are performed by physical illumination model and generative adversarial network. The tone mapping is performed by combining cross-modal fusion and human visual characteristics.
It achieves high dynamic range and detail preservation of images under complex lighting conditions, improves the accuracy and stability of image processing, adapts to real-time compensation in different lighting scenarios, and conforms to human visual perception habits.
Smart Images

Figure CN121032846A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and image processing, in particular to a high-precision image processing method and system based on adaptive compensation of illumination. BACKGROUND
[0002] In the image acquisition process, the uncertainty of the illumination condition is one of the core factors affecting the image quality, which is specifically manifested as: uneven illumination, such as bright and dark patches caused by multiple indoor light sources, outdoor shadow shielding; insufficient dynamic range, loss of details in strong or weak light scenes; color temperature deviation, color distortion caused by different light sources; real-time requirement, in the scenarios of automatic driving, industrial pipeline detection, etc., real-time images need to be processed quickly.
[0003] In the scenarios of medical image analysis, unmanned aerial vehicle inspection, intelligent security, etc., the accuracy of image details directly affects the reliability of decision-making. Illumination interference may cause target edge blur, affecting segmentation accuracy; noise amplification, reducing feature extraction reliability; color distortion, leading to classification model misjudgment.
[0004] Traditional methods are difficult to cope with complex illumination changes, while deep learning technology performs well in image enhancement, but there are problems such as insufficient dynamic range and extreme illumination processing, existing methods are difficult to retain high light details and dark area information at the same time, leading to image hierarchy loss in extreme illumination scenes; weak algorithm generalization ability, fixed parameter model unstable in multi-light source and complex reflection scenes, frequent manual intervention is required, etc. Therefore, researching adaptive illumination compensation algorithm has become a key direction to improve image processing accuracy. SUMMARY
[0005] To solve the above technical problems, a high-precision image processing method and system based on adaptive compensation of illumination are provided, which solve the problems of insufficient dynamic range and extreme illumination processing, and weak algorithm generalization ability.
[0006] To achieve the above purposes, the technical scheme adopted by the present application is: A high-precision image processing method based on adaptive compensation of illumination, comprising: S1: inputting an original image, decomposing the image into a high-frequency edge layer, a medium-frequency texture layer and a low-frequency illumination layer through a multi-scale residual network; S2: using an ambient light sensor and a scene semantic segmentation model to obtain illumination intensity, color temperature and scene category in real time, and generating dynamic compensation parameters; S3: performing dynamic range expansion on the low-frequency illumination layer based on a physical illumination model, and adjusting the weight of high light suppression and dark area enhancement through an adaptive S-shaped exposure curve; S4: A dual-branch generative adversarial network is used to perform noise suppression and super-resolution reconstruction on the high-frequency layer, and texture detail enhancement on the mid-frequency layer. S5: By using the cross-modal fusion module, the data from the depth camera and infrared sensor are aligned with the visible light image to constrain the physical rationality of illumination compensation; S6: Based on the characteristics of human visual perception, the fused image is tone-mapped to output an enhanced image with high dynamic range and detail preservation.
[0007] Preferably, S1 specifically includes: The original image is normalized to eliminate color differences between devices, and an estimate of the noise level is calculated based on local variance analysis. Three convolutional kernels of different scales are used in parallel to extract features at different scales: small kernel to capture high-frequency edges; medium kernel to extract mid-frequency textures; and large kernel to model low-frequency illumination. The output consists of three branches. After unifying the number of channels of the outputs of the three branches, element-wise addition and fusion are performed to preserve multi-scale details. An improved SENet module is introduced to dynamically allocate the weights of features at each scale. This includes global average pooling of the fused features to generate channel description vectors; learning the dependencies between channels through a two-layer fully connected network to output the normalized weights of each branch; and weighted summing of the original branch features according to the weights to obtain the final fused features.
[0008] Preferably, S1 specifically includes: High-frequency edge layer extraction is performed, and the fused features are enhanced by Laplacian operator convolution to improve edge response; adaptive hyperbolic tangent activation dynamically adjusts the slope of the activation function according to the noise level to suppress noise interference; Mid-frequency texture layer separation, design of multi-directional Gabor filter bank to constrain the main directionality and periodicity of texture: adversarial training orthogonalizes mid-frequency features and low-frequency illumination features; Low-frequency illumination layer reconstruction, anisotropic diffusion loss, constraining the low-frequency layer to be smooth and retaining soft shadow transitions during the training phase, fusing color temperature and brightness data from the ambient light sensor, and correcting color deviations in the illumination layer through convolution.
[0009] Preferably, S3 specifically includes: Light propagation equation modeling simplifies scene illumination distribution based on radiative transfer equation; Monte Carlo ray tracing simulates complex reflections and multiple scatterings to generate low-noise illumination estimates. Sensor data fusion and correction, ambient light parameter injection, mapping the color temperature and light intensity obtained in step S2 to RGB three-channel gain coefficients; light source direction is used to estimate the physical rationality boundary of the shadow area; dynamic parameter initialization, loading predefined model parameter templates according to scene semantic tags; fine-tuning BRDF parameters through online learning module to adapt to unknown material reflection characteristics; The dynamic range of the low-frequency illumination layer is expanded, and an adaptive logarithmic mapping is used to normalize the pixel values of the low-frequency layer; highlight details are preserved, and excessive compression of the highlight area is limited by gradient clipping. Construct a piecewise S-curve, define the curve function, generate dynamic parameters, perform differentiable optimization, and build a loss function; smooth the shadow-highlight transition, apply anisotropic diffusion filtering to the lighting layer after adjusting the S-curve, and maintain soft shadow gradation; implement cross-modal depth constraints and segment the foreground / background regions using the depth map from step S5. Reflection consistency verification, bidirectional reflection distribution function verification, real-time calculation of material reflection properties, if an anomaly is detected, triggering parameter rollback mechanism; anomaly handling, switching to backup model.
[0010] Preferably, S4 specifically includes: A dual-branch GAN architecture is used, with a high-frequency branch generator (including a U-Net variant), an encoder-decoder structure, and embedded residual dense blocks. A noise perception module takes the fused high-frequency layer and noise estimation map as input and dynamically suppresses the feature responses of noisy regions through an adaptive gating mechanism. A discriminator-based multi-scale spectral discrimination module takes the high-frequency layer and the generated result as input and extracts features through downsampling. Spectral normalization is applied to stabilize adversarial training. The output is a realism probability map at each scale, guiding the generator to retain edge sharpness. The generator for the mid-frequency branch includes a direction-sensitive convolutional group and employs an 8-direction learnable Gabor filter bank; it embeds a self-attention mechanism to model long-range texture correlation; it evaluates texture complexity based on a discriminator and calculates the similarity between local binary pattern features and generated textures; and it introduces gradient orientation histogram loss to constrain texture naturalness. The high-frequency branch loss function includes noise suppression loss and super-resolution reconstruction loss. The super-resolution reconstruction loss includes multi-scale structural similarity loss, which preserves edge structures. The perceptual loss is based on VGG-19 feature map alignment. The mid-frequency branch loss function includes texture adversarial loss and orientation consistency loss. Cross-branch collaborative training: the underlying convolutional kernels of the high-frequency and mid-frequency generators are shared to extract basic features; dynamic weight allocation: the gradient backpropagation ratio between branches is adjusted according to the scene classification label; joint adversarial training: the outputs of the two discriminators are fused into a global realism loss. The model is lightweight by using dynamic channel pruning. During the training phase, channel importance scoring is introduced, and redundant channels are pruned during inference. The pruning intensity of high-frequency branches is higher than that of mid-frequency branches. Quantization-aware training is used, with generator weights quantized to 8-bit specific points and the discriminator retaining FP16 accuracy. A quantization error compensation layer is introduced. Heterogeneous computing acceleration, dedicated NPU core, high-frequency branch generator mapped to the matrix acceleration unit of NPU, mid-frequency branch direction convolution using programmable DSP; seamless data transmission between layers is achieved through a double buffering mechanism.
[0011] Preferably, S5 specifically includes: Spatiotemporal synchronization calibration corrects the non-uniformity of infrared images by aligning the exposure times of the depth camera, infrared sensor, and visible light camera using trigger signals. The multi-sensor extrinsic parameter matrix is calculated based on a checkerboard calibration board, and the depth / infrared data is mapped to the visible light image coordinate system; bilinear interpolation and edge-guided upsampling are used. Physical feature extraction, calculation of scene surface normal vectors, identification of occlusion boundaries and shadow casting areas; construction of a 3D spatial illumination attenuation model; Infrared data analysis is used to segment areas with significant temperature, and the radiance is corrected by combining the material emissivity table. The effect of ambient light on the surface temperature of the object is inferred by using the heat conduction equation. The multimodal feature fusion network, a graph neural network, includes node definition, with visible light image patches, depth regions, and infrared regions as heterogeneous nodes; edge weights, which dynamically calculate the association strength based on physical relationships; and inter-layer propagation, which updates node features through message passing to generate a fused feature map. Adversarial training optimization involves the discriminator judging whether the input fused features conform to physical laws; physical rationality constraints include prohibiting illumination compensation in areas where the depth map shows an object's distance is beyond the effective range of the light source; adjusting local exposure gain based on surface normal vectors; inferring material roughness from infrared data to limit the compensation range of highly reflective surfaces; and detecting specular reflection areas by fusing polarized light data to avoid over-enhancement leading to flare artifacts.
[0012] Preferably, S2 specifically includes: An RGB ambient light sensor measures ambient light intensity, color temperature, and spectral distribution; a ToF sensor helps determine the distance and direction of the light source; multiple cameras are synchronized, and sensor data and image frame timestamps are aligned via hardware trigger signals. The relationship between sensor output and actual illumination is fitted using a multinomial regression model to complete nonlinear correction; the dynamic range is expanded by applying logarithmic compression to bright areas and exponential stretching to dark areas. Temporal illumination analysis: statistically analyzes the variance of illumination intensity fluctuations within a sliding window to detect sudden light sources; analyzes the periodicity of illumination changes based on FFT; physical parameter mapping: converts color temperature into CIE1931 color space coordinates to quantify the impact of ambient light on white balance; and estimates a scene illumination attenuation model by combining the direction of the light source. The lightweight DeepLabv3+ was used, and the backbone network was replaced with MobileNetv4. An attention mechanism was introduced to enhance the segmentation accuracy of lighting-related regions. Lighting-related scene labels include indoor, outdoor, and mixed light sources. Object semantic labels include reflective surfaces and objects with light-absorbing materials that affect the lighting compensation strategy. The dataset was synthesized using Unreal Engine 5 to simulate extreme lighting scenes; the reflection characteristics of complex materials were simulated through physical rendering; real data was labeled using a semi-automatic labeling tool to extend the lighting labels of the Cityscapes and ADE20K datasets; and a polarized light camera was introduced to capture specular reflection areas. Multimodal data feature-level fusion encodes sensor data into 32-dimensional vectors; scene segmentation results are mapped to 64-dimensional vectors through an embedding layer; cross-modal feature interaction is achieved using a Transformer encoder; output parameters include: exposure compensation weights, color temperature correction matrix, and local enhancement mask; The adaptive learning strategy employs meta-learning initialization, where the model learns to quickly adapt to new lighting scenarios from a small number of samples during the pre-training phase. It combines loss functions, including lighting parameter error and semantic segmentation cross-entropy loss. Online reinforcement learning optimization defines a reward function and provides real-time feedback based on image quality assessment metrics. Finally, it updates the generated policy network parameters online using the PPO algorithm.
[0013] Preferably, S6 specifically includes: Brightness adaptation modeling, local brightness adaptation curves, based on Stevens' power law, design regional brightness mapping functions, and divide the foreground / background into near and far scenes according to the depth map in step S5; Color perception optimization, based on the CIE170-1 color matching function, converts the image from RGB to CIEXYZ color space to simulate the response of human eye cone cells; asymmetric enhancement is applied to the red and green channels. By constraining the contrast sensitivity function, a frequency domain filter bank is constructed to suppress high-frequency noise and low-frequency color difference that are insensitive to the human eye; the detail enhancement amplitude is controlled by the JND threshold. Dynamic tone mapping curve, global S-curve, local detail preservation, using bilateral filters to separate the primary color layer and detail layer, compressing the dynamic range of the primary color layer, and superimposing and enhancing the detail layer; adaptive color gamut compression, color gamut boundary mapping, detecting colors outside the display device's color gamut, and compressing them along the CIELAB uniform chromaticity axis; prioritizing the preservation of memory colors, sacrificing less important tones; Local contrast enhancement is achieved by applying Gaussian filtering to the luminance channel based on the multi-scale Retinex algorithm to extract multi-scale reflection components; dynamic weighted fusion enhances small-scale details in the foreground and improves large-scale brightness balance in the background. Color adaptation and white balance correction: CAT02 color adaptation model, based on color temperature data from step S2, converts the light source color to D65 standard white point; applies color shift compensation to highlight areas; eliminates halo effect; detects brightness abrupt change boundaries; applies anisotropic diffusion filtering to transition areas within 5 pixels; and constrains gradient change rate.
[0014] A high-precision image processing system based on adaptive illumination compensation is provided to implement the high-precision image processing method based on adaptive illumination compensation as described above, comprising: Multi-scale residual network module: used to decompose the input image into a high-frequency edge layer, a mid-frequency texture layer and a low-frequency illumination layer, which includes an asymmetric multi-scale convolutional kernel group and a noise estimation subnetwork; Ambient light sensor module: integrates RGB sensor and ToF sensor to acquire real-time data on light intensity, color temperature and light source direction; Scene semantic segmentation module: Based on the lightweight DeepLabv3+ model, it outputs scene category labels and material reflection properties; Dynamic compensation parameter generation module: By fusing sensor data and semantic segmentation results through the Transformer encoder, exposure compensation weights, color temperature correction matrix and local enhancement mask are generated; The dual-branch generative adversarial network module includes a high-frequency noise suppression and super-resolution branch, a mid-frequency texture enhancement branch, and shares the underlying convolutional parameters. Cross-modal fusion module: Aligns depth camera or infrared sensor data with visible light images and achieves feature fusion under physical constraints through graph neural networks; Tone mapping module: Based on the CIE170-1 color matching function and contrast sensitivity function, dynamically generate a perceptually optimized S-shaped mapping curve.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention proposes a cross-modal graph neural network with physical constraints to deeply fuse multi-source data such as visible light, depth, and infrared on the basis of spatiotemporal alignment. Combined with scene semantic segmentation and dynamic parameter generation, it achieves physical rationality and environmental adaptability of illumination compensation. In complex backlighting scenes, the system can jointly suppress overexposed areas with depth information and infer the influence of thermal radiation through infrared data, avoiding the distortion caused by traditional algorithms ignoring physical laws.
[0016] This system combines physical optics models with human visual characteristics. On one hand, it achieves precise compensation based on material reflection properties and light source attenuation models; on the other hand, it uses strategies such as CIE170-1 color adaptation and memory color priority preservation to ensure that the processing results conform to human visual perception habits. Through online learning and semantically driven parameter adjustment, the system can adapt in real time to complex lighting changes from extreme low light to direct strong light. The lightweight DeepLabv3+ segmentation network quickly identifies scene types and dynamically switches compensation strategies; the dual-branch GAN network addresses the differences between high-frequency noise and mid-frequency texture, resolving the contradiction between dynamic range compression and detail preservation in traditional methods. Attached Figure Description
[0017] Figure 1 A flowchart of a high-precision image processing method based on adaptive illumination compensation; Figure 2 This is an internal framework diagram of a high-precision image processing system based on adaptive illumination compensation. Detailed Implementation
[0018] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0019] Reference Figure 1 As shown, a high-precision image processing method based on adaptive illumination compensation includes: S1: Input the original image, and decompose the image into a high-frequency edge layer, a mid-frequency texture layer, and a low-frequency illumination layer through a multi-scale residual network; S2: Utilize ambient light sensors and scene semantic segmentation models to acquire real-time light intensity, color temperature, and scene category, and generate dynamic compensation parameters; S3: Based on the physical illumination model, the low-frequency illumination layer is dynamically expanded, and the weights of highlight suppression and shadow enhancement are adjusted by an adaptive S-shaped exposure curve. S4: A dual-branch generative adversarial network is used to perform noise suppression and super-resolution reconstruction on the high-frequency layer, and texture detail enhancement on the mid-frequency layer. S5: By using the cross-modal fusion module, the data from the depth camera and infrared sensor are aligned with the visible light image to constrain the physical rationality of illumination compensation; S6: Based on the characteristics of human visual perception, the fused image is tone-mapped to output an enhanced image with high dynamic range and detail preservation.
[0020] It should be noted that in the joint optimization of the frequency domain and the physical domain, the low-frequency illumination layer relies on the physical illumination model (radiative transfer equation) to constrain the global dynamic range expansion and dynamically calibrate parameters based on sensor data; the high-frequency / mid-frequency layer achieves noise-detail decoupling through GAN adversarial training and forms a closed-loop feedback with the physical verification of cross-modal data.
[0021] To ensure temporal consistency, for video stream processing, optical flow estimation is used to align the decomposition layers of adjacent frames to avoid flicker effects; the dynamic compensation parameters adopt a sliding window mechanism (window size of 5 frames) to balance real-time performance and stability.
[0022] With dual adaptation to human vision and machine perception, and display optimization, the S6's tone mapping module is adapted to the human eye color matching function (CIE170-1) to ensure natural colors; it is also compatible with machine vision, and the output image retains high-frequency details (such as edges and text), supporting subsequent OCR, object detection and other tasks.
[0023] The dynamic range extension employs a physical-data dual-drive mechanism, with physical model constraints and Monte Carlo ray tracing simulation of complex reflections in S3, addressing the failure of traditional Retinex theory for specular highlights; data-driven correction, with an online learning module optimizing BRDF parameters in real time to adapt to the lighting response of unknown materials.
[0024] A noise-texture decoupled generative adversarial architecture with a high-frequency branch and a noise mask-guided local suppression strategy to avoid edge blurring caused by global filtering; and a mid-frequency branch with a direction-sensitive Gabor filter bank to enhance texture orientation consistency, surpassing the limitations of traditional LBP features.
[0025] Cross-modal depth reliability enhancement, physical rule embedding, and surface normal vector constraint of shadow compensation range in S5 to suppress long-range enhancement artifacts; heterogeneous data alignment, dynamic weighted fusion of ToF depth map and infrared radiation data in GNN to improve the signal-to-noise ratio in low-light scenes.
[0026] S1 specifically includes: The original image is normalized to eliminate color differences between devices, and an estimate of the noise level is calculated based on local variance analysis. Three convolutional kernels of different scales are used in parallel to extract features at different scales: small kernel to capture high-frequency edges; medium kernel to extract mid-frequency textures; and large kernel to model low-frequency illumination. The output consists of three branches. After unifying the number of channels of the outputs of the three branches, element-wise addition and fusion are performed to preserve multi-scale details. An improved SENet module is introduced to dynamically allocate the weights of features at each scale. This includes global average pooling of the fused features to generate channel description vectors; learning the dependencies between channels through a two-layer fully connected network to output the normalized weights of each branch; and weighted summing of the original branch features according to the weights to obtain the final fused features. High-frequency edge layer extraction is performed, and the fused features are enhanced by Laplacian operator convolution to improve edge response; adaptive hyperbolic tangent activation dynamically adjusts the slope of the activation function according to the noise level to suppress noise interference; Mid-frequency texture layer separation, design of multi-directional Gabor filter bank to constrain the main directionality and periodicity of texture: adversarial training orthogonalizes mid-frequency features and low-frequency illumination features; Low-frequency illumination layer reconstruction, anisotropic diffusion loss, constraining the low-frequency layer to be smooth and retaining soft shadow transitions during the training phase, fusing color temperature and brightness data from the ambient light sensor, and correcting color deviations in the illumination layer through convolution.
[0027] It should be noted that normalization and noise estimation are involved: Device-independent color calibration, cross-device mapping, converting input images to the standard color gamut based on ICC Profile, and eliminating sensor differences through color adaptation transformation (CAT16); dynamic white balance, for low color temperature scenes (<3000K), adopts color temperature-hue curve fitting to avoid the problem of cool tones turning blue. Noise level estimation, local variance analysis, calculation of the variance of image blocks (16×16), and separation of signal-related noise by combining Poisson-Gaussian mixture model; Deep learning assists in the pre-training of a noise estimation network (MobileNetv4 lightweight architecture) to output a noise mask, which guides subsequent adaptive activation function adjustments.
[0028] Multi-scale feature extraction and fusion: The asymmetric convolution kernel group design includes a small kernel (3×3) with embedded edge-guided convolution, which constrains the high-frequency edge localization accuracy through Sobel gradient; a medium kernel (5×5) with split convolution (1×5+5×1) to reduce computation and introduce dynamic sparse connections (80% channel activation); and a large kernel (7×7) with dilated convolution (dilation=2) to expand the receptive field and cover a larger area of illumination gradient.
[0029] The feature fusion strategy involves cross-scale interaction, achieving complementary enhancement of small kernel edge features and large kernel illumination features through a cross-attention mechanism; residual compensation involves residual connections between the original input and the fused features, preserving details that were not captured.
[0030] Improved SENet module optimizations: Dynamic channel weight allocation, global context modeling, and the use of StripPooling instead of traditional average pooling to extract long-range spatial dependencies in images; nonlinear enhancement, with the fully connected network using the Swish activation function to improve the expressive power of nonlinear relationships between channels; multi-task constraints, jointly optimizing channel weights and frequency domain decomposition loss during the training phase to avoid weight allocation bias.
[0031] Physical-data dual-drive mechanism for frequency domain separation: High-frequency edge layer extraction, Laplacian operator enhancement, and convolution kernel weight matrix are as follows: [[0,1,0],[1,-4,1],[0,1,0]], the response threshold is dynamically adjusted by the noise level (σ=0.1~0.3). Adaptive hyperbolic tangent activation, where the slope parameter k is controlled by the noise estimate, as shown in the formula:
[0032] In the formula, α=0.05 is an empirical coefficient, and σ is the noise standard deviation.
[0033] Mid-frequency texture layer separation, multi-directional Gabor filter bank: preset 8 directions (0°, 45°, 90°, ..., 315°) and 3 scales (λ=4, 8, 16 pixels), dynamically select the dominant direction; Adversarial orthogonalization training is used to design a discriminator network that distinguishes between mid-frequency and low-frequency features. The loss function is:
[0034] In the formula, D is the discriminator network used to distinguish between mid-frequency and low-frequency features; G is the generator network used to generate mid-frequency features; mid is the mid-frequency feature input into the generator G, which processes or transforms it; low is the low-frequency feature, which is directly input into the discriminator D for discrimination.
[0035] Low-frequency illumination layer reconstruction, anisotropic diffusion loss, gradient smoothing constrained by the Perona-Malik equation, formula:
[0036] In the formula, ∇ is the gradient operator, representing the gradient of the low-frequency illumination layer. Gradient calculation is performed to obtain the rate of change and direction of light intensity in space; L_low is the low-frequency illumination layer, representing low-frequency illumination information in the image; exp: exponential function, used to calculate the weighting factor and control the intensity of diffusion; k is a parameter that controls the diffusion intensity and determines the degree of influence of the gradient magnitude on diffusion. Sensor data fusion maps the ambient light color temperature into a 3×3 color correction matrix, and adjusts the color of the illumination layer through 1×1 convolution.
[0037] Computation graph optimization and operator fusion combine convolution, normalization, and activation into a single kernel function, reducing memory read / write operations; Winograd acceleration applies the F(4×4,3×3) algorithm to 3×3 convolutions, reducing computational load.
[0038] The memory management strategy involves dynamic block caching, dividing the image into 512×512 blocks based on the GPU memory capacity, and pipelined processing; mixed precision training, where feature maps are retained using FP16 and weight updates use FP32, balancing accuracy and speed.
[0039] The hardware adaptation solution involves exporting the model to ONNX format, which can then be compiled and optimized using Apache TVM to generate TVM runtime modules (.so / .tar), thereby reducing computational load and memory consumption and enabling real-time inference.
[0040] S2 specifically includes: An RGB ambient light sensor measures ambient light intensity, color temperature, and spectral distribution; a ToF sensor helps determine the distance and direction of the light source; multiple cameras are synchronized, and sensor data and image frame timestamps are aligned via hardware trigger signals. The relationship between sensor output and actual illumination is fitted using a multinomial regression model to complete nonlinear correction; the dynamic range is expanded by applying logarithmic compression to bright areas and exponential stretching to dark areas. Temporal illumination analysis: statistically analyzes the variance of illumination intensity fluctuations within a sliding window to detect sudden light sources; analyzes the periodicity of illumination changes based on FFT; physical parameter mapping: converts color temperature into CIE1931 color space coordinates to quantify the impact of ambient light on white balance; and estimates a scene illumination attenuation model by combining the direction of the light source. The lightweight DeepLabv3+ was used, and the backbone network was replaced with MobileNetv4. An attention mechanism was introduced to enhance the segmentation accuracy of lighting-related regions. Lighting-related scene labels include indoor, outdoor, and mixed light sources. Object semantic labels include reflective surfaces and objects with light-absorbing materials that affect the lighting compensation strategy. The dataset was synthesized using Unreal Engine 5 to simulate extreme lighting scenes; the reflection characteristics of complex materials were simulated through physical rendering; real data was labeled using a semi-automatic labeling tool to extend the lighting labels of the Cityscapes and ADE20K datasets; and a polarized light camera was introduced to capture specular reflection areas. Multimodal data feature-level fusion encodes sensor data into 64-dimensional vectors; scene segmentation results are mapped to 64-dimensional vectors through an embedding layer; cross-modal feature interaction is achieved using a Transformer encoder; output parameters include: exposure compensation weights, color temperature correction matrix, and local enhancement mask; The adaptive learning strategy employs meta-learning initialization, where the model learns to quickly adapt to new lighting scenarios from a small number of samples during the pre-training phase. It combines loss functions, including lighting parameter error and semantic segmentation cross-entropy loss. Online reinforcement learning optimization defines a reward function and provides real-time feedback based on image quality assessment metrics. Finally, it updates the generated policy network parameters online using the PPO algorithm.
[0041] It should be noted that multi-sensor collaborative calibration: The spatiotemporal synchronization mechanism uses hardware-triggered alignment to synchronize the data streams of the RGB sensor, ToF, and polarized light camera by generating a global timestamp, with a maximum time deviation of ≤0.5ms. Spatial calibration compensation calculates the extrinsic parameter matrix of multiple sensors based on a checkerboard calibration board and uses B-spline interpolation to compensate for minor offsets caused by mechanical vibration.
[0042] The dynamic range expansion strategy employs a segmented mapping function, using logarithmic compression (L_out=log10(L_in / L_max)) for bright areas (>10^4 lux) and exponential stretching (L_out=(L_in)^γ, γ=1.8~2.2) for dark areas (<1 lux). An adaptive segmentation point is used, dynamically adjusting the compression / stretching interval through histogram peak-valley detection to avoid artifacts in the transition zone caused by manually set thresholds.
[0043] Modeling of physical parameters of illumination: Light source direction estimation: ToF depth data combined with point cloud normal vector analysis to construct the probability distribution of light source direction (Gaussian mixture model); polarized light data is used to analyze the mirror reflection direction and derive the azimuth angle of the main light source in reverse. Color temperature-color space mapping maps the color temperature to the CIE1931xy coordinate system and constrains the white balance adjustment range through the MacAdam ellipse; a 3×3 color temperature correction matrix is dynamically generated, with the constraint condition being to minimize the ΔE2000 color difference of the memory colors (skin and green leaves).
[0044] Lightweight DeepLabv3+ optimized design: The backbone network adopts MobileNetV4, which integrates grouped convolution and channel shuffling (ShuffleNetV2 strategy) to reduce computation; a spatial-channel dual attention CBAM module is introduced to dynamically adjust weights through an adaptive learning mechanism; Multi-task tag fusion, lighting scene tags, such as "direct sunlight - noon" and "mixed light source - dusk", guides the initialization of dynamic compensation parameters; material attribute tags, annotating reflection characteristics (specular metal, matte plastic, etc.), affect the generation of local enhancement masks.
[0045] Synthetic-real data joint training, Unreal Engine 5 simulation, to construct extreme lighting scenes, such as strong reflections in snow and sudden changes in light at tunnel entrances and exits, and generate physically realistic specular highlights through ray tracing; the material library contains 500+ BRDF parameters and supports anisotropic reflection modeling. The semi-automatic annotation tool uses SAM (SegmentAnythingModel) pre-annotated boundaries, requiring manual correction of only 10% of key areas; polarized light data provides reflection intensity labels to assist in annotating specular areas.
[0046] Multimodal feature interaction mechanism: Transformer encoder design, input layer: sensor data (32-dimensional) + segmentation embedding (64-dimensional) + temporal features (16-dimensional); multi-head attention (8 heads) to calculate cross-modal association weights, focusing on regions with abrupt changes in illumination; output parameters: exposure weight map (1 / 8 resolution original image), 3×3 color temperature matrix, local enhancement mask (binarization threshold 0.7).
[0047] Adaptive learning strategy: Meta-learning initialization uses the MAML algorithm during the pre-training phase to enable the model to adapt to new scenes within 5 iterations; online reinforcement learning uses the reward function R=0.6SSIM+0.3NIQE+0.1*ΔE, and the PPO algorithm updates the policy network every 10 frames.
[0048] Model lightweighting and deployment: The model undergoes knowledge distillation to reduce the number of parameters; hierarchical dynamic quantization is used, with backbone network weights and activation values dynamically quantized to INT8 (inverse quantization calculation at runtime); position encoding employs INT8 lookup table method + dynamic inverse quantization (INT8 compression during storage, temporary conversion to FP16 during computation) to reduce memory usage; and the FP16 accuracy of the output layer is preserved to maintain prediction stability. This approach meets the limitations of embedded devices while maintaining good performance.
[0049] S3 specifically includes: Light propagation equation modeling simplifies scene illumination distribution based on radiative transfer equation; Monte Carlo ray tracing simulates complex reflections and multiple scatterings to generate low-noise illumination estimates. Sensor data fusion and correction, ambient light parameter injection, mapping the color temperature and light intensity obtained in step S2 to RGB three-channel gain coefficients; light source direction is used to estimate the physical rationality boundary of the shadow area; dynamic parameter initialization, loading predefined model parameter templates according to scene semantic tags; fine-tuning BRDF parameters through online learning module to adapt to unknown material reflection characteristics; The dynamic range of the low-frequency illumination layer is expanded, and an adaptive logarithmic mapping is used to normalize the pixel values of the low-frequency layer; highlight details are preserved, and excessive compression of the highlight area is limited by gradient constraints. Construct a piecewise S-curve, define the curve function, generate dynamic parameters, perform differentiable optimization, and construct a loss function; smooth the shadow-highlight transition by anisotropically diffusing the gradient field of the illumination layer before adjusting the S-curve to constrain the slope change of the S-curve and maintain soft shadow gradation; and implement cross-modal depth constraints by segmenting the foreground / background regions using the depth map from step S5. Reflection consistency verification, bidirectional reflection distribution function verification, real-time calculation of material reflection properties, if an anomaly is detected, triggering parameter rollback mechanism; anomaly handling, switching to backup model.
[0050] It should be noted that the modeling and real-time simulation of light propagation are as follows: Scene lighting is separated into direct lighting (analytical calculation) and indirect lighting (pre-integrated environment mapping) to adapt to real-time rendering on mobile devices; Monte Carlo importance sampling increases sampling density for specular reflection areas and reduces noise iterations; noise reduction acceleration integrates OptiXDenoiser and combines motion vectors and normal buffers to achieve single-frame noise reduction.
[0051] Sensor-semantic joint calibration, RGB gain mapping, based on the CIE1931 color matching function, converts color temperature (CCT) into RGB gain coefficients; shadow physical verification verifies shadow boundaries and depth through light projection. Figure 1 For consistency, a dynamic threshold is used (tolerance pixel count = object distance × 5). If the region meets the predetermined threshold, it is marked as an abnormal area. BRDF online learning, material reflection modeling, adopts Cook-Torrance micro-surface model, dynamically updates roughness and Fresnel terms; meta-learning strategy, pre-trained network initialized on 5 basic materials (metal, plastic, fabric, etc.), and new material adaptation is achieved through online gradient descent.
[0052] Adaptive logarithmic mapping: Nonlinear compression, defining a piecewise function:
[0053] In the formula, L_out is the output brightness value, which is the result after nonlinear compression; L_in is the input brightness value, which is the brightness of each pixel in the original image; ε is a very small constant used to avoid taking log(0) in the logarithmic function and prevent numerical calculation errors; α is a scaling factor, which scales the result after taking the logarithm when L_in>T to adjust the output brightness range; β is a linear scaling factor, which linearly scales the input brightness value when L_in≤T to control the output brightness of low brightness areas; T is a threshold used to distinguish different brightness areas in the image, which is dynamically set according to the scene semantic label, such as T=0.7 for indoor scenes and T=0.35 for outdoor scenes, so as to adapt to the characteristics of brightness distribution in different scenes and make the brightness mapping more in line with the needs of the actual scene.
[0054] Gradient constraint: Apply gradient clipping (|∇L|≤0.1) to the highlight region (L_in>0.9) to prevent texture loss.
[0055] The S-curve is constructed using differentiable parameterization, and the curve function is designed as follows:
[0056] In the formula, S(x) represents the function value of the S-shaped curve, with the output value in the interval (0,1). It is used to perform nonlinear mapping on the input variables, so that the data has an S-shaped distribution characteristic. x is the input variable, which can be the gray value, brightness value, or other feature quantity of the image that needs to be nonlinearly transformed. k is a parameter that controls the steepness of the curve, which is determined by the dynamic range ratio (DR_ratio=L_max / L_min). The larger the value of k, the steeper the curve changes near the inflection point; the smaller the value of k, the smoother the curve.
[0057] The loss function is designed to jointly optimize the contrast-sensitive loss (CSF-weighted MSE) and the color difference loss (ΔE2000).
[0058] S4 specifically includes: A dual-branch GAN architecture is used, with a high-frequency branch generator (including a U-Net variant), an encoder-decoder structure, and embedded residual dense blocks. A noise perception module takes the fused high-frequency layer and noise estimation map as input and dynamically suppresses the feature responses of noisy regions through an adaptive gating mechanism. A discriminator-based multi-scale spectral discrimination module takes the high-frequency layer and the generated result as input and extracts features through downsampling. Spectral normalization is applied to stabilize adversarial training. The output is a realism probability map at each scale, guiding the generator to retain edge sharpness. The generator for the mid-frequency branch employs an 8-directional learnable Gabor filter bank; it embeds a self-attention mechanism to model long-range texture correlation; it uses a discriminator-based texture complexity evaluation to calculate the similarity between local binary pattern features and generated textures; and it introduces gradient orientation histogram loss to constrain texture naturalness. The high-frequency branch loss function includes noise suppression loss and super-resolution reconstruction loss. The super-resolution reconstruction loss includes multi-scale structural similarity loss, which preserves edge structures. The perceptual loss is based on VGG-19 feature map alignment. The mid-frequency branch loss function includes texture adversarial loss and orientation consistency loss. Cross-branch collaborative training: the underlying convolutional kernels of the high-frequency and mid-frequency generators are shared to extract basic features; dynamic weight allocation: the gradient backpropagation ratio between branches is adjusted according to the scene classification label; joint adversarial training: the outputs of the two discriminators are fused into a global realism loss. The model is lightweight by using dynamic channel pruning. During the training phase, channel importance scoring is introduced, and redundant channels are pruned during inference. The pruning intensity of high-frequency branches is higher than that of mid-frequency branches. Quantization-aware training is used, with generator weights quantized to 8-bit specific points and the discriminator retaining FP16 accuracy. A quantization error compensation layer is introduced. Heterogeneous computing acceleration, dedicated NPU core, high-frequency branch generator mapped to the matrix acceleration unit of NPU, mid-frequency branch direction convolution using programmable DSP; seamless data transmission between layers is achieved through a double buffering mechanism.
[0059] It should be noted that high-frequency branching, noise suppression, and super-resolution reconstruction are involved. The generator architecture is optimized with an improved U-Net design. The residual dense blocks contain 6 convolutional layers and employ cross-layer dense connections to improve feature reuse efficiency. Noise-aware gating is used to dynamically mask the feature responses of noisy regions by inputting high-frequency layers and noise estimation maps through Sigmoid gating weights (0~1).
[0060] A multi-scale spectral discriminator with pyramid downsampling and 4 scale levels (original image → 1 / 2 → 1 / 4 → 1 / 8), each level containing 3×3 convolution + spectral normalization (SN-GANs); adversarial loss fusion, weighted summation of realism probability maps at each scale (weights 0.4 / 0.3 / 0.2 / 0.1), guides the generator to balance edge preservation and noise suppression.
[0061] High-frequency loss function design, noise suppression loss: smoothL1 loss based on noise mask, which only calculates the penalty for high-confidence noise regions (confidence > 0.7); multi-scale SSIM loss, which calculates structural similarity at 1×, 2×, and 4× downsampling scales, preserving edge topology; perceptual alignment loss, which extracts VGG-19 features and constrains the high-dimensional feature distance between the generated image and the ground truth.
[0062] Mid-frequency branch, texture enhancement and orientation consistency: Orientation-sensitive texture modeling can learn Gabor filter banks, initialized as complex numerical filters with 8 directions (0°~157.5° interval of 22.5°) and 3 frequencies (λ=4, 8, 16 pixels); during the training phase, the orientation angle θ and frequency λ are jointly optimized to adapt to different texture periodic characteristics.
[0063] The self-attention mechanism is enhanced by calculating the query-key relevance matrix and focusing on long-range periodic textures, such as fabric folds and brick walls; positional encoding is introduced to learn positional information and improve performance.
[0064] Low-frequency loss function design, texture adversarial loss: the discriminator extracts VGG-19 depth features and measures the difference between the generated and real texture distribution by the slice Wasserstein distance (SWD); gradient direction loss: the cosine similarity of the generated image and the gradient direction histogram (HOG) of the ground plane is calculated at 1×, 2×, and 4× downsampling scales to constrain the naturalness of edge orientation.
[0065] Joint training mechanism: The underlying parameters are shared, and the weights of the first three convolutional kernels are shared to extract general edge / texture primitive features; dynamic frequency modulation is used to generate adaptive branch weights based on illumination intensity and scene through a gated attention network; global realism constraints are applied, and the loss of high-frequency and mid-frequency discriminators is dynamically fused through illumination perception coefficients.
[0066] Dynamic channel pruning calculates channel importance during the training phase using activation sensitivity and adaptively determines the pruning ratio based on scenario analysis (10-40% for high-frequency branches and 15-30% for mid-frequency branches); this significantly reduces the number of model parameters and improves inference speed.
[0067] Hybrid precision quantization is used, with generator weights employing 8-bit specific points (INT8) and activation values retaining FP16 precision, resulting in a quantization error of <1.5%; the discriminator uses FP16 to maintain adversarial training stability.
[0068] S5 specifically includes: Spatiotemporal synchronization calibration corrects the non-uniformity of infrared images by aligning the exposure times of the depth camera, infrared sensor, and visible light camera using trigger signals. The multi-sensor extrinsic parameter matrix is calculated based on a checkerboard calibration board, and the depth / infrared data is mapped to the visible light image coordinate system; bilinear interpolation and edge-guided upsampling are used. Physical feature extraction is performed to calculate scene surface normal vectors and identify occlusion boundaries and shadow projection areas. A physical illumination attenuation model is constructed, integrating light source radiance (Io), material reflectivity (p), medium scattering coefficient (β), and distance attenuation (1 / d). 2This enables radiometric-level brightness prediction.
[0069] Infrared data analysis is used to segment areas with significant temperature, and the radiance is corrected by combining the material emissivity table. The effect of ambient light on the surface temperature of the object is inferred by using the heat conduction equation. The multimodal feature fusion network, a graph neural network, includes node definition, with visible light image patches, depth regions, and infrared regions as heterogeneous nodes; edge weights, which dynamically calculate the association strength based on physical relationships; and inter-layer propagation, which updates node features through message passing to generate a fused feature map. Adversarial training optimization involves the discriminator judging whether the input fused features conform to physical laws; physical rationality constraints include prohibiting illumination compensation in areas where the depth map shows an object's distance is beyond the effective range of the light source; adjusting local exposure gain based on surface normal vectors; inferring material roughness from infrared data to limit the compensation range of highly reflective surfaces; and detecting specular reflection areas by fusing polarized light data to avoid over-enhancement leading to flare artifacts.
[0070] It should be noted that the hardware-level synchronization mechanism triggers and controls the generation of a global exposure pulse signal to synchronize the depth camera, infrared and visible light sensors, with a maximum timing deviation of ≤0.2ms; it dynamically compensates for mechanical delay, predicts motion blur based on gyroscope data, and adjusts the sensor exposure start time. Non-uniformity correction: pixel-by-pixel gain / offset calibration of the infrared sensor, generating a correction matrix through a blackbody radiation source; dynamic temperature drift compensation: non-uniformity parameters are updated every 10 frames.
[0071] Multimodal data alignment: The extrinsic parameter calibration was optimized by using the ArUco enhanced checkerboard grid (12×9 corner points) to jointly solve the depth-visible-infrared extrinsic parameter matrix, with a reprojection error ≤1.5 pixels; edge-guided upsampling was performed by applying guided filtering to interpolate low-resolution infrared data (320×240) to 1080P, preserving the sharpness of thermal boundaries.
[0072] 3D physics reconstruction: Surface normal vector calculation: Construct a Poisson surface based on the depth map and calculate vertex normal vectors; Occlusion boundary detection: Combine visible light edges and depth discontinuities to construct a binary mask; Light attenuation modeling, physical attenuation formula:
[0073] In the formula, θ is the angle between the normal vector and the light source, d is the distance, and ε = 0.1 (divide by zero). Dynamic light source estimation: The position of the main light source is inferred from the infrared high-temperature region, with an accuracy of ±0.5m@10m.
[0074] Infrared physical analysis: Significant temperature region segmentation: Regions with temperature differences >2°C are segmented based on the UNet lightweight network (2.1M parameters); Emissivity correction: The true emissivity of the material is obtained by looking up a table (common material database), and the temperature is corrected by the law of radiation.
[0075] Heat conduction inversion was performed, and the one-dimensional thermal equation was solved using the finite difference method to inversely deduce the influence coefficient of ambient light on surface temperature (α = 0.05~0.8). Heterogeneous node definition and interaction: Node feature encoding: visible light nodes: 512-dimensional ResNet-34 features; depth nodes: 64-dimensional PointNet features; infrared nodes: 32-dimensional temporal features. Dynamic edge weight calculation, physical relationship weights, and modal design: Depth - Visible Edge:
[0076] Infrared-Visible Edge:
[0077] In the formula, σ is the standard deviation; α and β are weighting parameters used to adjust the influence of IoU (Intersection over Union) and OCC (Occlusion) on the depth-visible edge weight; γ is the infrared-visible edge weighting parameter; ϵ is the emissivity, used to describe the ability of an object's surface to emit infrared radiation; ∆T is the temperature difference. Spatial s near weight, adaptive high break kernel:
[0078] In the formula, An adaptive Gaussian kernel is used for calculating spatial proximity weights; 动态 Dynamically adjusted based on scene depth; Adversarial training and physical rule embedding: Physical discriminator design, input: fused feature map + surface normal vector + illumination attenuation coefficient; Rule base validation: If the depth value d > 3 * effective distance of the light source (calculated by S2), then the compensation weight of that area will be forced to zero.
[0079] Specular reflection suppression and polarized light data analysis: Calculate the polarization degree threshold > 0.6 and mark it as a specular area; Dynamic gain limitation: Set the upper limit of exposure gain for specular areas to 1.2 times to avoid flare overexposure.
[0080] S6 specifically includes: Brightness adaptation modeling is performed by designing local brightness mapping curves based on Stevens' law and combining them with a retinal model to simulate spatial brightness adaptation. Color perception optimization: Based on the CE170-1 color matching function, the image is converted to the CEXVZ color space to maintain color balance, and the color components of the insufficient saturation area are uniformly enhanced in the CIELAB space. By constraining the contrast sensitivity function, a frequency domain filter bank is constructed to suppress high-frequency noise and low-frequency color difference that are insensitive to the human eye; the detail enhancement amplitude is controlled by the JND threshold. Dynamic tone mapping curve, global S-curve, local detail preservation, using bilateral filters to separate the primary color layer and detail layer, compressing the dynamic range of the primary color layer, and superimposing and enhancing the detail layer; adaptive color gamut compression, color gamut boundary mapping, detecting colors outside the display device's color gamut, and compressing them to the target color gamut along the visual uniformity axis in the IPT color space; establishing memory color priority (skin tone ΔE<3, sky blue ΔE<4), and allowing low-priority hues to have ΔE<8. Local contrast enhancement is achieved by applying Gaussian filtering to the luminance channel based on the multi-scale Retinex algorithm to extract multi-scale reflection components; dynamic weighted fusion enhances small-scale details in the foreground and improves large-scale brightness balance in the background. Color adaptation and white balance correction: The CAM16-UCS color adaptation model is used to convert the light source color to the D65 standard white point; color shift in the highlight area is compensated by separating the specular reflection component through polarization; halo effect elimination: the boundary of brightness change is detected, anisotropic diffusion filtering is applied to the transition area within 5 pixels, and the gradient change rate is constrained.
[0081] It should be noted that brightness adaptation modeling and region mapping are involved: Brightness adaptation modeling: global mapping design based on Stevens' power law; Regional optimization: Spatial frequency weighting enhances details in the mid-band (4-6 cpd) (gain × 1.3); Depth continuous modulation, using the Sgmoid function to smooth the depth of field effect of the liquid ( Fogging suppression is achieved by applying adaptive filtering (intensity α depth) to low-frequency components (<0.5cpd). Color perception optimization and physiological constraints: Color space conversion and channel enhancement, CIE170-1 color matching, RGB→XYZ conversion matrix:
[0082] Simulate the L / M / S cone response and generate the perceived brightness Y by weighted fusion.
[0083] Asymmetric red-green enhancement: Red channel gain coefficient R_gain=1+0.3*Sigmoid(Y) enhances warm color memory; Green channel dynamic compression G_out=G_in^(1 / 1.2) alleviates vegetation oversaturation.
[0084] With contrast sensitivity constraints, a frequency domain filter bank was designed, and six Butterworth filters were constructed to cover the spatial frequency range of 0.5 to 30 cpd (cycles / degree). 90% of the energy was retained in the human eye sensitive frequency band (4 to 8 cpd), while the energy in the insensitive frequency band (>16 cpd) was attenuated to 30%.
[0085] JND threshold control, dynamically adjusting the enhancement magnitude based on the JND model: ΔL_max=JND_base*(1+0.5*log(L_avg)) When the local contrast ΔL > ΔL_max, soft shearing is enabled to prevent over-enhancement.
[0086] Global and local co-mapping: The S-curve is parameterized, with the center point x0 determined by the average image brightness: x0 = 0.3 * μ_Y + 0.7 * Y_median; the slope k is related to the dynamic range DR: k = 1.5 / (1 + e^{-0.1 * (DR-10)}).
[0087] Bilateral filtering decomposition, spatial domain σ_space=3 pixels, luminance domain σ_range=0.1*Y_max, separating the primary color layer (low frequency) and detail layer (high frequency); the primary color layer is compressed to the range of 0.3~0.7, and the detail layer gain is increased by 1.8 times before being superimposed.
[0088] Color gamut adaptive compression: Color gamut boundary detection: calculate the color gamut boundary in CIELAB space based on the display device's ICC Profile; compress colors that exceed the boundary along a constant hue line, and preferentially retain memory colors (ΔE<3).
[0089] Prioritization strategy: the compression weight of skin color (LAB: L=50~75, a=5~20, b=15~30) and vegetation green (a=-30~-10) is reduced by 50%; low saturation areas (C<15) are allowed a greater compression.
[0090] Multi-scale Retinex optimization: Gaussian kernel parameters, scale set σ=[15,80,250] pixels, corresponding to detail, mid-frequency, and global luminance components; dynamic weighting, near-field weight [0.6,0.3,0.1], far-field weight [0.2,0.3,0.5].
[0091] Reflection component fusion is performed, and edge optimization is achieved through guided filtering to avoid halo artifacts (filter radius r = 5 pixels, λ = 0.01).
[0092] Halo effect elimination: Abrupt boundary detection employs the Sobel operator to calculate gradients, followed by nonmaximum suppression to refine edges. Dual thresholding (high threshold = 85th percentile of gradient histogram, low threshold = high threshold × 0.4) preserves strong edges and weak connected edges; Transition region generation: Generate a Gaussian decay mask (σ=1.5 pixels) centered on the edge to replace the dilation operation.
[0093] Anisotropic diffusion: Diffusion coefficient:
[0094] Adaptive iteration (1-5 times), termination condition:
[0095] Reference Figure 2 As shown, a high-precision image processing system based on adaptive illumination compensation includes: Multi-scale residual network module: used to decompose the input image into a high-frequency edge layer, a mid-frequency texture layer and a low-frequency illumination layer, which includes an asymmetric multi-scale convolutional kernel group and a noise estimation subnetwork; Ambient light sensor module: integrates RGB sensor and ToF sensor to acquire real-time data on light intensity, color temperature and light source direction; Scene semantic segmentation module: Based on the lightweight DeepLabv3+ model, it outputs scene category labels and material reflection properties; Dynamic compensation parameter generation module: By fusing sensor data and semantic segmentation results through the Transformer encoder, exposure compensation weights, color temperature correction matrix and local enhancement mask are generated; The dual-branch generative adversarial network module includes a high-frequency noise suppression and super-resolution branch, a mid-frequency texture enhancement branch, and shares the underlying convolutional parameters. Cross-modal fusion module: Aligns depth camera or infrared sensor data with visible light images and achieves feature fusion under physical constraints through graph neural networks; Tone mapping module: Based on the CIE170-1 color matching function and contrast sensitivity function, dynamically generate a perceptually optimized S-shaped mapping curve.
[0096] It should be noted that the multi-scale residual network module includes: The hierarchical decomposition strategy is as follows: High-frequency layer: 1×1, 3×3, and 5×5 convolutional kernel groups are used to extract edge details (maximizing the difference in receptive field), and dilated convolution (dilation=2) is fused to capture long-range edge correlations; Mid-frequency layer: Direction-sensitive convolution (initialized with an 8-axis Gabor filter) is used to enhance periodic textures, and deformable convolution is combined to adapt to irregular surfaces; Low-frequency layer: 7×7 large kernel convolution + global average pooling is used to extract the illumination distribution basis function.
[0097] The noise estimation subnetwork is a lightweight U-Net structure (4 layers of encoding and decoding). It takes high-frequency layer features as input and outputs a noise confidence heatmap (0~1). Synthetic noise (Gaussian-Poisson mixed noise) is introduced during the training phase to simulate the background noise characteristics of the sensor.
[0098] Ambient light sensor module: Spatiotemporal alignment optimization and hardware-triggered synchronization: Generate a global timestamp, align RGB and ToF data streams, with timing deviation ≤0.3ms; ToF data correction: Depth calibration based on temperature compensation to eliminate multipath interference; Light source direction modeling: combining ToF point cloud normal vectors and RGB image shadow analysis to construct the light source direction probability distribution (Gaussian mixture model); polarization sensor assistance: analyzing the polarization angle of mirror reflection to improve the accuracy of direction estimation.
[0099] Scene semantic segmentation module: The lightweight DeepLabv3+ is improved with backbone network optimization and MobileNetv4 architecture: dynamic sparse convolution is introduced to reduce computational cost; dual attention is enhanced, spatial attention (CBAM) focuses on highlight / shadow regions, and channel attention (SE-Net) enhances material features.
[0100] Multi-label joint learning, scene category labels: lighting scene, such as "direct sunlight - noon" and "mixed light source - dusk", guides the initialization of dynamic parameters; material property prediction, outputs reflectivity, such as mirror metal and matte plastic, and generates BRDF parameters in conjunction with the physics engine.
[0101] Dynamic compensation parameter generation module: Feature encoding and fusion: sensor encoding includes mapping RGB / ToF data to a 64-dimensional vector via MLP and adding position encoding; semantic encoding includes mapping the segmentation results to a 64-dimensional vector via an embedding layer; cross-modal attention includes an 8-head attention mechanism to calculate sensor-semantic association weights, with a focus on regions with abrupt changes in illumination. Parameter generation and optimization, the output layer structure includes an exposure weight map (original image with 1 / 8 resolution), a 3×3 color temperature correction matrix, and a local enhancement mask (binarization threshold 0.7). The joint loss function consists of L1 loss (parameter accuracy), perceptual loss (VGG-19 feature alignment), and physical constraint loss (depth-light source distance verification).
[0102] Dual-branch generative adversarial network module: High-frequency and mid-frequency synergistic enhancement, high-frequency branch (noise suppression + super-resolution), generator, improved U-Net (residual dense block + noise gating) to suppress noise while preserving real details; discriminator, multi-scale spectrum analysis (4-level downsampling) to constrain the frequency domain distribution to conform to the characteristics of natural images.
[0103] Mid-frequency branch (texture enhancement): Generator, direction-sensitive convolution + self-attention, enhances complex textures such as fabric wrinkles and wood grain; Discriminator, LBP feature comparison, ensures that the generated texture conforms to local statistical regularities.
[0104] Dynamic parameter sharing, with a 70% sharing rate of the underlying convolutional kernels and independent optimization of the higher-level networks; adaptive weight allocation, dynamically adjusting the branch loss weights based on scene labels.
[0105] Cross-modal fusion module: Physically constrained graph neural network, node and edge definition: visible light nodes, 512-dimensional ResNet features; depth nodes, 64-dimensional PointNet features; infrared nodes: 64-dimensional temporal features; edge weight calculation: physical relationship weight (depth-light attenuation model) + spatial proximity weight (Gaussian kernel function).
[0106] Anomaly handling mechanism: reflectivity verification, real-time calculation of material BRDF parameters, if the deviation from the preset range is >20%, parameter rollback is triggered; thermodynamic constraints, infrared data is used to infer the surface temperature change rate, limiting unreasonable light compensation intensity.
[0107] Tone mapping module; Color gamut compression and memory color protection: CIELAB space compression projects the color gamut along a constant hue line, prioritizing the preservation of skin tones (ΔE<2) and vegetation green (ΔE<3); JND constraint enhancement: the contrast variation is limited to within the perceptible difference threshold to avoid over-processing.
[0108] In summary, the advantages of this invention are: Integrating data from multiple sensors including visible light, depth, and infrared, the system achieves cross-modal fusion under physical constraints through spatiotemporal alignment and graph neural networks. By combining physical parameters such as heat conduction equations and material reflectivity, the system dynamically suppresses and compensates for distortions, ensuring that the processed results conform to optical laws while also adapting to complex environmental changes.
[0109] Based on the CIE170-1 color matching function and contrast sensitivity characteristics, a brightness-color dynamic mapping model is constructed. By constraining the detail enhancement amplitude with JND threshold and prioritizing the preservation of memory colors through LAB color gamut compression, the output image is made closer to human visual preferences on a scientifically accurate basis, avoiding the distortion and incongruity caused by over-processing in traditional algorithms.
[0110] A lightweight semantic segmentation network rapidly identifies scene types, driving a dual-branch GAN network to differentiate high-frequency noise and mid-frequency texture. Combined with a Transformer to dynamically generate compensation parameters, it achieves global adaptive adjustment from extremely dark to bright light, resolving the inherent contradiction between dynamic range compression and detail preservation in traditional methods.
[0111] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A high-precision image processing method based on adaptive illumination compensation, characterized in that, include: S1: Input the original image, and decompose the image into a high-frequency edge layer, a mid-frequency texture layer, and a low-frequency illumination layer through a multi-scale residual network; S2: Utilize ambient light sensors and scene semantic segmentation models to acquire real-time light intensity, color temperature, and scene category, and generate dynamic compensation parameters; S3: Based on the physical illumination model, the low-frequency illumination layer is dynamically expanded, and the weights of highlight suppression and shadow enhancement are adjusted by an adaptive S-shaped exposure curve. S4: A dual-branch generative adversarial network is used to perform noise suppression and super-resolution reconstruction on the high-frequency layer, and texture detail enhancement on the mid-frequency layer. S5: By using the cross-modal fusion module, the data from the depth camera and infrared sensor are aligned with the visible light image to constrain the physical rationality of illumination compensation; S6: Based on the characteristics of human visual perception, the fused image is tone-mapped to output an enhanced image with high dynamic range and detail preservation.
2. The high-precision image processing method based on adaptive illumination compensation according to claim 1, characterized in that, S1 specifically includes: The original image is normalized, and an estimate of the noise level is calculated based on local variance analysis. Three convolutional kernels of different scales are used in parallel to extract features at different scales: small kernel to capture high-frequency edges; medium kernel to extract mid-frequency textures; and large kernel to model low-frequency illumination. The output consists of three branches. After unifying the number of channels of the outputs of the three branches, element-wise addition and fusion are performed to preserve multi-scale details. An improved SENet module is introduced to dynamically allocate the weights of features at each scale. This includes global average pooling of the fused features to generate channel description vectors; learning the dependencies between channels through a two-layer fully connected network to output the normalized weights of each branch; and weighted summing of the original branch features according to the weights to obtain the final fused features.
3. The high-precision image processing method based on adaptive illumination compensation according to claim 2, characterized in that, S1 specifically includes: High-frequency edge layer extraction is performed, and the fused features are enhanced by Laplacian operator convolution to improve edge response; adaptive hyperbolic tangent activation dynamically adjusts the slope of the activation function according to the noise level to suppress noise interference; Mid-frequency texture layer separation is achieved by designing a multi-directional Gabor filter bank to constrain the main directionality and periodicity of the texture, and adversarial training is used to orthogonalize the mid-frequency features and low-frequency illumination features. Low-frequency illumination layer reconstruction, anisotropic diffusion loss, constraining the low-frequency layer to be smooth and retaining soft shadow transitions during the training phase, fusing color temperature and brightness data from the ambient light sensor, and correcting color deviations in the illumination layer through convolution.
4. The high-precision image processing method based on adaptive illumination compensation according to claim 3, characterized in that, S3 specifically includes: Light propagation equation modeling simplifies scene illumination distribution based on radiative transfer equation; Monte Carlo ray tracing simulates complex reflections and multiple scatterings to generate low-noise illumination estimates. Sensor data fusion and correction, ambient light parameter injection, mapping the color temperature and light intensity obtained in step S2 to RGB three-channel gain coefficients; light source direction is used to estimate the physical rationality boundary of the shadow area; dynamic parameter initialization, loading predefined model parameter templates according to scene semantic tags; fine-tuning BRDF parameters through online learning module to adapt to unknown material reflection characteristics; The dynamic range of the low-frequency illumination layer is expanded, and an adaptive logarithmic mapping is used to normalize the pixel values of the low-frequency layer; highlight details are preserved, and excessive compression of the highlight area is limited by gradient clipping. Construct a piecewise S-curve, define the curve function, generate dynamic parameters, perform differentiable optimization, and build a loss function; smooth the shadow-highlight transition, apply anisotropic diffusion filtering to the lighting layer after adjusting the S-curve, and maintain soft shadow gradation; implement cross-modal depth constraints and segment the foreground / background regions using the depth map from step S5. Reflection consistency verification, bidirectional reflection distribution function verification, real-time calculation of material reflection properties, if an anomaly is detected, triggering parameter rollback mechanism; anomaly handling, switching to backup model.
5. The high-precision image processing method based on adaptive illumination compensation according to claim 4, characterized in that, S4 specifically includes: A dual-branch GAN architecture is used, with a high-frequency branch generator (including a U-Net variant), an encoder-decoder structure, and embedded residual dense blocks. A noise perception module takes the fused high-frequency layer and noise estimation map as input and dynamically suppresses the feature responses of noisy regions through an adaptive gating mechanism. A discriminator-based multi-scale spectral discrimination module takes the high-frequency layer and the generated result as input and extracts features through downsampling. Spectral normalization is applied to stabilize adversarial training. The output is a realism probability map at each scale, guiding the generator to retain edge sharpness. The generator for the mid-frequency branch includes a direction-sensitive convolutional group and employs an 8-direction learnable Gabor filter bank; it embeds a self-attention mechanism to model long-range texture correlation; it evaluates texture complexity based on a discriminator and calculates the similarity between local binary pattern features and generated textures; and it introduces gradient orientation histogram loss to constrain texture naturalness. The high-frequency branch loss function includes noise suppression loss and super-resolution reconstruction loss. The super-resolution reconstruction loss includes multi-scale structural similarity loss, which preserves edge structures. The perceptual loss is based on VGG-19 feature map alignment. The mid-frequency branch loss function includes texture adversarial loss and orientation consistency loss. Cross-branch collaborative training: the underlying convolutional kernels of the high-frequency and mid-frequency generators are shared to extract basic features; dynamic weight allocation: the gradient backpropagation ratio between branches is adjusted according to the scene classification label; joint adversarial training: the outputs of the two discriminators are fused into a global realism loss. The model is lightweight by using dynamic channel pruning. During the training phase, channel importance scoring is introduced, and redundant channels are pruned during inference. The pruning intensity of high-frequency branches is higher than that of mid-frequency branches. Quantization-aware training is used, with generator weights quantized to 8-bit specific points and the discriminator retaining FP16 accuracy. A quantization error compensation layer is introduced. Heterogeneous computing acceleration, dedicated NPU core, high-frequency branch generator mapped to the matrix acceleration unit of NPU, mid-frequency branch direction convolution using programmable DSP; seamless data transmission between layers is achieved through a double buffering mechanism.
6. A high-precision image processing method based on adaptive illumination compensation according to claim 5, characterized in that, S5 specifically includes: Spatiotemporal synchronization calibration corrects the non-uniformity of infrared images by aligning the exposure times of the depth camera, infrared sensor, and visible light camera using trigger signals. The multi-sensor extrinsic parameter matrix is calculated based on a checkerboard calibration board, and the depth / infrared data is mapped to the visible light image coordinate system; bilinear interpolation and edge-guided upsampling are used. Physical feature extraction, calculation of scene surface normal vectors, identification of occlusion boundaries and shadow casting areas; construction of a 3D spatial illumination attenuation model; Infrared data analysis is used to segment areas with significant temperature, and the radiance is corrected by combining the material emissivity table. The effect of ambient light on the surface temperature of the object is inferred by using the heat conduction equation. The multimodal feature fusion network, a graph neural network, includes node definition, with visible light image patches, depth regions, and infrared regions as heterogeneous nodes; edge weights, which dynamically calculate the association strength based on physical relationships; and inter-layer propagation, which updates node features through message passing to generate a fused feature map. Adversarial training optimization involves the discriminator judging whether the input fused features conform to physical laws; physical rationality constraints include prohibiting illumination compensation in areas where the depth map shows an object's distance is beyond the effective range of the light source; adjusting local exposure gain based on surface normal vectors; inferring material roughness from infrared data to limit the compensation range of highly reflective surfaces; and detecting specular reflection areas by fusing polarized light data to avoid over-enhancement leading to flare artifacts.
7. A high-precision image processing method based on adaptive illumination compensation according to claim 6, characterized in that, S2 specifically includes: An RGB ambient light sensor measures ambient light intensity, color temperature, and spectral distribution; a ToF sensor helps determine the distance and direction of the light source; multiple cameras are synchronized, and sensor data and image frame timestamps are aligned via hardware trigger signals. The relationship between sensor output and actual illumination is fitted using a multinomial regression model to complete nonlinear correction; the dynamic range is expanded by applying logarithmic compression to bright areas and exponential stretching to dark areas. Temporal illumination analysis: statistically analyzes the variance of illumination intensity fluctuations within a sliding window to detect sudden light sources; analyzes the periodicity of illumination changes based on FFT; physical parameter mapping: converts color temperature into CIE1931 color space coordinates to quantify the impact of ambient light on white balance; and estimates a scene illumination attenuation model by combining the direction of the light source. The lightweight DeepLabv3+ was used, and the backbone network was replaced with MobileNetv4. An attention mechanism was introduced to enhance the segmentation accuracy of lighting-related regions. Lighting-related scene labels include indoor, outdoor, and mixed light sources. Object semantic labels include reflective surfaces and objects with light-absorbing materials that affect the lighting compensation strategy. The dataset was synthesized using Unreal Engine 5 to simulate extreme lighting scenes; the reflection characteristics of complex materials were simulated through physical rendering; real data was labeled using a semi-automatic labeling tool to extend the lighting labels of the Cityscapes and ADE20K datasets; and a polarized light camera was introduced to capture specular reflection areas. Multimodal data feature-level fusion encodes sensor data into 32-dimensional vectors; scene segmentation results are mapped to 64-dimensional vectors through an embedding layer; cross-modal feature interaction is achieved using a Transformer encoder; output parameters include: exposure compensation weights, color temperature correction matrix, and local enhancement mask; The adaptive learning strategy employs meta-learning initialization, where the model learns to quickly adapt to new lighting scenarios from a small number of samples during the pre-training phase. It combines loss functions, including lighting parameter error and semantic segmentation cross-entropy loss. Online reinforcement learning optimization defines a reward function and provides real-time feedback based on image quality assessment metrics. Finally, it updates the generated policy network parameters online using the PPO algorithm.
8. A high-precision image processing method based on adaptive illumination compensation according to claim 7, characterized in that, S6 specifically includes: Brightness adaptation modeling, local brightness adaptation curves, based on Stevens' power law, design regional brightness mapping functions, and divide the foreground / background into near and far scenes according to the depth map in step S5; Color perception optimization, based on the CIE170-1 color matching function, converts the image from RGB to CIEXYZ color space to simulate the response of human eye cone cells; asymmetric enhancement is applied to the red and green channels. By constraining the contrast sensitivity function, a frequency domain filter bank is constructed to suppress high-frequency noise and low-frequency color difference that are insensitive to the human eye; the detail enhancement amplitude is controlled by the JND threshold. Dynamic tone mapping curve, global S-shaped curve, local detail preservation, using bilateral filters to separate the primary color layer and detail layer, compressing the dynamic range of the primary color layer, and superimposing and enhancing the detail layer; adaptive color gamut compression, color gamut boundary mapping, detecting colors outside the display device's color gamut, and compressing them along the CIELAB uniform chromaticity axis; prioritizing the preservation of memory colors, sacrificing less important tones.
9. A high-precision image processing method based on adaptive illumination compensation according to claim 8, characterized in that, S6 specifically includes: Local contrast enhancement is achieved by applying Gaussian filtering to the luminance channel based on the multi-scale Retinex algorithm to extract multi-scale reflection components; dynamic weighted fusion enhances small-scale details in the foreground and improves large-scale brightness balance in the background. Color adaptation and white balance correction: CAT02 color adaptation model, based on color temperature data from step S2, converts the light source color to D65 standard white point; applies color shift compensation to highlight areas; eliminates halo effect; detects brightness abrupt change boundaries; applies anisotropic diffusion filtering to transition areas within 5 pixels; and constrains gradient change rate.
10. A high-precision image processing system based on adaptive illumination compensation, comprising the high-precision image processing method based on adaptive illumination compensation according to claims 1-8, characterized in that, include: Multi-scale residual network module: used to decompose the input image into a high-frequency edge layer, a mid-frequency texture layer and a low-frequency illumination layer, which includes an asymmetric multi-scale convolutional kernel group and a noise estimation subnetwork; Ambient light sensor module: integrates RGB sensor and ToF sensor to acquire real-time data on light intensity, color temperature and light source direction; Scene semantic segmentation module: Based on the lightweight DeepLabv3+ model, it outputs scene category labels and material reflection properties; Dynamic compensation parameter generation module: By fusing sensor data and semantic segmentation results through the Transformer encoder, exposure compensation weights, color temperature correction matrix and local enhancement mask are generated; The dual-branch generative adversarial network module includes a high-frequency noise suppression and super-resolution branch, a mid-frequency texture enhancement branch, and shares the underlying convolutional parameters. Cross-modal fusion module: Aligns depth camera or infrared sensor data with visible light images and achieves feature fusion under physical constraints through graph neural networks; Tone mapping module: Based on the CIE170-1 color matching function and contrast sensitivity function, dynamically generate a perceptually optimized S-shaped mapping curve.
Citation Information
Cited By
Multi-domain voxel reconstruction three-dimensional display method and device
CN121235932A
A multi-domain voxel reconstruction three-dimensional display method and device
CN121235932B
Visual inspection method and system for grade sorting of molybdenum oxide products
CN121258987A
A visual inspection method and system for grade sorting of molybdenum oxide products.
CN121258987B
Sensing-driven enhanced network method and system for unmarked microscopic cell image
CN121280268A