Low-light image quality enhancement system and method based on enhanced night scene modeling
By combining generative adversarial networks and dynamic weighting strategies, the problem of insufficient sample coverage in nighttime scenes was solved, improving the model accuracy and real-time performance of low-light image enhancement, achieving high-quality image enhancement results, and meeting the real-time and safety requirements of autonomous driving systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INNER MONGOLIA UNIVERSITY
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-03
AI Technical Summary
Existing low-light image enhancement techniques are prone to pattern collapse and gradient vanishing when the model learns the noise distribution in dark areas due to insufficient sample coverage in nighttime scenes. Furthermore, there is a contradiction between computational complexity and real-time performance, making it difficult to meet the real-time requirements of scenarios such as autonomous driving.
By employing a hybrid generative adversarial network and a dynamic weighting strategy, enhanced low-light scene images are generated through a dual-channel hybrid generative network architecture. An optimized training sample library is constructed by combining an attention module and a probabilistic graphical model. The model is trained using a region-adaptive weighted loss function and combined with dynamic adaptive closed-loop adjustment to achieve real-time estimation of illumination and noise distribution and image enhancement.
It significantly improved the training data coverage for nighttime scenes from 30% to 68.7%, reduced the noise prediction error rate by 72.7%, improved PSNR by 2.7dB under extreme conditions, significantly improved image enhancement quality and detail restoration capabilities, and achieved inference latency of less than 18.3ms, meeting the real-time requirements of in-vehicle systems and possessing high reliability and functional safety.
Smart Images

Figure CN122335591A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for enhancing the quality of low-light images, specifically to a system and method for enhancing the quality of low-light images based on enhanced nighttime scene modeling. Background Technology
[0002] In the field of computer vision and image processing, image quality enhancement technology under low-light conditions is a fundamental link supporting key applications such as autonomous driving, security monitoring, and medical imaging. With the rapid development of deep learning technology, neural network-based image enhancement methods have gradually replaced traditional algorithms based on histogram equalization or Retinex theory, becoming a research hotspot for solving low-light imaging problems. Early works such as LIME (Low-light Image Enhancement) achieved basic enhancement effects by establishing a decomposition model of illumination map and reflectance map, combined with Retinex theory. However, its shortcomings in noise suppression led to a loss rate of up to 35% of details in dark areas (X. Guo, Y. Li and H. Ling, "LIME: Low-Light ImageEnhancement via Illumination Map Estimation," in IEEE Transactions on Image Processing, vol. 26, no. 2, pp. 982-993, Feb. 2017). Recent generative adversarial network (GAN) based methods (such as EnlightenGAN) have improved the peak signal-to-noise ratio (PSNR) to 28.6 dB by introducing physically guided adversarial training, but this improvement is limited in extremely low light conditions. Even in scenes with obvious color cast and artifacts (Y. Jiang et al., "EnlightenGAN: Deep Light Enhancement Without Paired Supervision," in IEEE Transactions on Image Processing, vol. 30, pp. 2340-2349, 2021).
[0003] In current technical solutions, the coupling between noise modeling and illumination estimation is the core bottleneck restricting performance improvement. End-to-end methods, such as SID (See in Dark), achieve noise suppression on the Sony IMX sensor dataset by jointly optimizing noise suppression and illumination compensation tasks (e.g., the method reported in X. Xu et al., "SNR-Aware Low-Light Image Enhancement," CVPR 2022, pp. 17714-17724, achieving a noise suppression ratio of approximately 17.3 dB in a specific frequency band). However, according to the latest benchmark tests (such as the reproduction test based on Y. Wang et al., "Low-Light Image Enhancement via Structure Modeling and Guidance," CVPR 2023), the prediction error rate for noise in dark areas remains above 15%, especially in complex scenarios such as nighttime fog, rain, and snow, where the error variance is as high as [missing information]. Far exceeding the ideal threshold Further experiments and analyses conducted by the applicant indicate that the modeling bias of existing illumination distribution estimation algorithms for non-uniform noise has become a major bottleneck. The root cause lies in the insufficient coverage of nighttime scene samples (currently, nighttime samples account for <30% of publicly available datasets). This data bias leads to mode collapse in the model's learning of noise distribution in dark areas. Experimental data show that this problem manifests as gradient vanishing (the illumination signal-to-noise ratio can decrease by approximately 42% in extremely low illumination) and a significant artifact multiplication effect (PSNR loss can reach 4.7 dB), severely impacting nighttime imaging quality.
[0004] Further technical analysis reveals that existing methods suffer from irreversible information loss during dynamic range compression. Traditional illumination mapping functions... While maintaining overall brightness balance, the detail recovery rate in dark areas is low. The fixed selection results in a loss of local contrast of up to (Measured by CIEDE2000 color difference). On the other hand, although physical modeling-based joint optimization methods (such as KinD++) improve the color shift problem through separation decomposition and enhancement steps, the computational redundancy caused by multi-stage processing increases its inference latency to 48ms / frame, which is difficult to meet the real-time requirements of vehicle systems (latency threshold ≤33ms@4K) (Y. Zhang et al., "Beyond Brightening Low-lightImages," International Journal of Computer Vision (IJCV), vol. 129, pp. 1013–1037, 2021).
[0005] The practical impact of these technical shortcomings is particularly significant in the field of autonomous driving. When onboard cameras operate in scenarios with sudden changes in lighting, such as tunnel exits or dawn / dusk, the dynamic adjustment lag of existing algorithms (response time > 80ms) can cause the target detection box coordinate offset error to exceed [a certain threshold]. Far exceeding the safety threshold of ADAS systems (Refer to the definition of perception limitations in ISO 21448:2022 "Road vehicles — Safety of the intended functionality"). Furthermore, the cumulative effect of illumination estimation errors can also cause white balance imbalances. According to the applicant's statistical analysis based on NHTSA (National Highway Traffic Safety Administration) nighttime traffic accident data, such environmental perception biases may increase the probability of misjudging traffic light colors. .
[0006] In summary, existing low-light image enhancement technologies face the following technical bottlenecks: (1) The core bottleneck of existing technologies lies in the difficulty of solving the deep coupling problem between noise modeling and illumination estimation. Deep learning-based methods heavily rely on training data; however, the coverage of nighttime scene samples in current public datasets is severely insufficient (less than 30%), which makes the model prone to pattern collapse and gradient vanishing problems when learning the noise distribution in dark areas, especially in extremely low illumination (below 30%). Under these conditions, the noise prediction error rate remains high (above 15%), far exceeding the ideal threshold.
[0007] (2) Existing methods have inherent limitations in dynamic range processing. Traditional illumination mapping functions or dynamic range compression algorithms usually have globally fixed or locally adaptive parameters, which means that while improving the overall brightness, the local contrast of dark area details is often sacrificed, resulting in irreversible information loss, and is prone to color shift and artifacts in scenes with sudden changes in illumination;
[0008] (3) There is a contradiction between the computational complexity and real-time performance of existing enhancement algorithms. Although some multi-stage processing methods that separate decomposition and enhancement steps have improved some enhancement effects, their inference latency is generally long (e.g., more than 48ms / frame), which is difficult to meet the stringent requirements for processing speed in scenarios such as autonomous driving and real-time monitoring (the latency threshold usually needs to be less than 33ms).
[0009] These technical deficiencies directly lead to safety hazards in critical application areas. For example, in autonomous driving systems, the lag in dynamic adjustment of the algorithm and errors in illumination estimation can cause the coordinates of the target detection box to deviate beyond the safety threshold, and even cause misjudgment of the color of critical targets such as traffic lights, posing a serious threat to driving safety.
[0010] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0011] The purpose of this invention is to provide a low-light image quality enhancement system and method based on enhanced night scene modeling. This invention solves the problem that insufficient coverage of night scene samples in the prior art leads to pattern collapse and gradient vanishing when the model learns the noise distribution in dark areas. This invention significantly improves the coverage of effective training data for night scenes by using a hybrid generative adversarial network and a dynamic weighting strategy.
[0012] To achieve the above objectives, this invention provides a low-light image quality enhancement system based on enhanced nighttime scene modeling. This system includes: a sample data construction module, which generates enhanced low-light scene images based on an initial image dataset using a dual-channel hybrid generative network architecture, and dynamically weights the enhanced low-light scene images according to content analysis results to construct an optimized training sample library; and a hybrid enhancement model, constructed based on an attention module and a probabilistic graphical model module. During the training phase, the hybrid enhancement model uses the optimized training sample library as input and employs a region-adaptive weighted loss function for parameter optimization. During the enhancement processing phase, the optimized model receives the low-light image to be processed, estimates the illumination and noise distribution of the low-light image, and generates the enhanced image. The dual-channel hybrid generative network architecture includes: a first generative adversarial network unit for scene style transfer and a second generative adversarial network unit for multi-condition generation; the attention module is a Transformer encoder, which captures long-range dependencies between image patches through a multi-head attention mechanism; the probabilistic graphical model module is a Bayesian probabilistic graphical model, which is used to establish a probability transfer function between illumination intensity and noise variance, and to quantify the uncertainty of the model's prediction results.
[0013] Preferably, the first generative adversarial network unit adopts CycleGAN generative adversarial network, and the second generative adversarial network unit adopts ACGAN generative adversarial network.
[0014] Preferably, the system further includes at least one of the following modules: a dark area confidence estimation module, which evaluates the confidence of dark areas in the input low-light image, providing a basis for adjusting the weights of the loss function used by the model training module; a hardware interface module, which connects to the image sensor and receives initial image data; and a distribution consistency monitoring module, which monitors the Bayesian probability branch in the hybrid enhancement model in real time and calculates the relationship between the current predicted distribution and the prior distribution in real time. KL Divergence serves as an early warning indicator of system uncertainty to prevent severe distortion of the model in extreme unknown scenarios; the output optimization module is used to perform dynamic range mapping operation on the enhanced image of the low-light image to be processed.
[0015] More preferably, the dark area confidence estimation module uses ResNeXt-101 as its basic architecture and introduces a multi-scale feature pyramid pooling module in the fourth residual stage.
[0016] More preferably, the dynamic range mapping operation adopts the Sigmoid-Curve dimming curve equation, as shown in formula (14): ; In equation (14), To output pixel values; The brightness component of the input image is used for enhancement. The maximum brightness threshold supported by the display device; The slope parameter of the dimming curve controls the compressive diffraction rate of the contrast ratio. This is the brightness center offset, used to define the equilibrium point of the curve; Brightness compensation reference value.
[0017] A second objective of this invention is to provide a low-light image quality enhancement method based on enhanced nighttime scene modeling, the method comprising the following steps: (S100) Sample data construction: Based on an initial image dataset, enhanced low-light scene images are generated through a dual-channel hybrid generative network architecture. Based on the content analysis results of the enhanced low-light scene images, the images are dynamically weighted to construct an optimized training sample library. (S200) Model training: Using the optimized training sample library as input, the hybrid augmentation model is trained using a region-adaptive weighted loss function. The optimal weight parameters of the model are determined through offline iteration to obtain the optimized hybrid augmentation model. (S300) Image enhancement: The low-light image to be processed is input into the optimized hybrid enhancement model, which estimates the light and noise distribution based on the learned distribution rules and generates the final enhanced image; (S400) Dynamic adaptive closed-loop adjustment: Real-time monitoring of the quality index of the enhanced image and dynamic adjustment of the system's operating parameter configuration (such as sensor exposure parameters or dynamic range mapping parameters) through a feedback mechanism.
[0018] Preferably, in step (S100), the dual-channel hybrid generative network architecture includes: a first generative adversarial network unit for scene style transfer and a second generative adversarial network unit for multi-condition generation.
[0019] The first generative adversarial network (GAN) unit is configured with several layers of ResNet generators to achieve scene style transfer. This ResNet generator includes a front-end encoding module, an intermediate transformation module, and a back-end decoding module. The entire network uses the LeakyReLU activation function. The intermediate transformation module consists of several residual block groups, each containing a convolutional layer and a normalization layer. Skip connections are introduced within each residual block to preserve low-frequency features. Its cycle consistency loss function is defined as: (1) In equation (1), For Cycle Consistency Loss; For mathematical expectation; A generator used to map a source domain image to a target domain; A generator used to map a target domain image back to the source domain; The input source domain image sample; These are the weighting coefficients for the identity mapping loss. This is used to maintain the identity mapping property between inputs and outputs; This is the Identity Mapping Loss term. The second generative adversarial network unit is configured with several layers of DCGAN discriminators to generate multiple conditions, and a scene classification head is attached, using several scene labels with one-hot encoding as condition inputs.
[0020] Preferably, in step (S100), the dynamic weight allocation strategy achieves optimal sample allocation through the dynamic weight formula (2); (2) In equation (2), For the first i Percentage of pixels in the dark area of each sample; Temperature coefficient; For the first j The percentage of dark area pixels in each sample; For the first i Attention weights for each sample; N sample Indicates the total number of samples; Preferably, in step (S100), sample validity control is achieved through three-level quality monitoring, sample distribution offset detection adopts K-means clustering algorithm, and when sample distribution offset triggers the set conditions, a stereo depth camera is used to collect boundary samples.
[0021] Preferably, in step (S200), the hybrid enhancement model adopts a Bayesian-Transformer hybrid architecture, which couples a hierarchical Bayesian probabilistic graphical model through a Transformer encoder with an attention mechanism. The spatial feature extraction process of the attention mechanism and the hierarchical probabilistic graphical model formed by the Gamma prior distribution are mathematically expressed as follows: (3) in, In equation (3), the hybrid enhancement model takes the deterministic feature representation henc extracted by the Transformer encoder as input, and maps it to the posterior distribution parameters of the latent variable z through a multilayer perceptron (MLP), i.e. For and Let Gamma be the distribution of the parameter, where the shape parameter is... With rate parameter All are predicted by nonlinear transformation of the deterministic feature representation henc.
[0022] Preferably, in step (S200), the hybrid enhancement model is constructed with a light intensity mapping function, as shown in formula (4): (4) In equation (4), This represents the total number of samples. For the first The sampling weights of each sample follow a normal distribution. ; Input pixel values; For bias terms; These are measured parameters; The hybrid enhancement model also constructs a probability transfer function between illumination intensity and noise variance, as shown in formula (5): (5) In equation (5), The standard deviation of noise; Light intensity; These are the sensor response parameters; This is a dynamic adjustment coefficient; The dynamic adjustment coefficient Real-time optimization is achieved through adaptive update rules, which are shown in formula (6): (6) In equation (6), For dynamic adjustment coefficients, subscript For training steps; The learning rate; This is the total loss function; The standard deviation of noise; To prevent the stability constant from having a denominator of zero.
[0023] Preferably, in step (S200), the Bayesian probability branch in the hybrid enhancement model is monitored in real time by the distribution consistency monitoring module to prevent severe distortion of the model under extreme unknown scenarios. The distribution consistency monitoring module calculates the relationship between the current predicted distribution and the prior distribution in real time. KL divergence As an early warning indicator of system uncertainty, it is shown in formula (7): (7) In equation (7), This represents the total number of steps. The prior target distribution (i.e., the baseline distribution) learned by the model during the training phase. This represents the predicted distribution of the hybrid augmentation model output at the current time.
[0024] More preferably, according to the KL divergence Determine whether the noise pattern of the current input image deviates significantly from the knowledge range learned by the model. If it does, determine that the current enhancement result is unreliable and trigger the model self-correction process, including: (1) Parameter freezing and activation: Temporarily freeze the backbone weights of the Transformer encoder and activate only the dynamic adjustment coefficients in the hybrid enhancement model. As trainable parameters; (2) Gradient backpropagation: based on the current KL The divergence bias is calculated using the adaptive update rule defined in formula (6) to calculate the gradient. ; (3) Online parameter update: dynamically adjust the coefficients along the gradient direction Perform single-step or multi-step fine-tuning to adjust the predicted distribution. Rapidly approximate the prior target distribution Thus KL The divergence has been pulled back to a safe range.
[0025] Preferably, in step (S200), the optimization training phase designs a region-weighted loss function, as shown in formula (8): (8) In equation (8), These are the pixel domain loss weight coefficients; For structural loss weights; The pixel values of the target image (true value); The pixel values of the enhanced image (predicted value); Structural feature representation of the target image; To enhance the structural feature representation of the image; Preferably, in step (S200), in the implementation of the region-weighted loss function, the confidence level of the dark area is generated, and the dynamic weight is calculated based on the dynamic weight calculation function, the mathematical expression of which is: (9) In equation (9), For the pixel domain loss, there are dynamic weighting coefficients. Confidence level for the dark area.
[0026] Preferably, in step (S200), during the training phase, the parameters are updated along the gradient direction by the optimizer to improve the performance of the hybrid augmentation model; the training phase adopts a three-stage learning rate scheduling, with the learning rate initially decaying linearly and then converted to cosine annealing; gradient management implements a layered pruning strategy, constraining the gradient norm of the layers. ,in For gradient, The L2 norm of the gradient. This is the weight matrix. The Frobenius norm is the weight.
[0027] Preferably, in step (S400), during the closed-loop regulation implementation, the quality index is calculated. The formula is: (10) In equation (10), The total number of samples; As weight; For indicator functions; For the sample ; This is a nighttime dataset; For example The noise variance; This is the noise sensitivity coefficient.
[0028] Preferably, in step (S400), in the closed-loop regulation implementation, based on the noise level... With system uncertainty The correlation equations are used to establish the following dynamic system model: First, define the state variables. , T Its evolution process is described by discrete-time state equations: (11) In equation (11), This is the system matrix. Each row in the matrix represents the evolution logic of the system state variables at the next time step, and each column represents the linear contribution weight of the current state variable to the system evolution. The off-diagonal elements of the system matrix reflect... Causal chain; input matrix Corresponding ISO - Exposure Control ; for The system state vector at time t; For system x, process noise; The calculation uses hyperbolic tangent normalization: (12) In equation (12), k is The slope adjustment coefficient of the normalization function; Uncertainty Then through three-dimensional vectors The comprehensive characterization incorporates three indicators: spatial gradient energy, temporal KL divergence, and intensity relative error.
[0029] Preferably, in step (S400), in the closed-loop adjustment implementation, the transfer function block diagram of the ISO-exposure joint control mechanism includes a dual path of feedforward and feedback: the feedforward path adjusts according to the real-time light intensity. The baseline ISO is determined by looking up a table; the feedback path adjusts the exposure time using a PID controller. : (13) In equation (13), For error signals, ; This is the proportional gain coefficient; This is the integral gain coefficient; This is the differential gain coefficient.
[0030] The low-light image quality enhancement system and method based on enhanced night scene modeling of the present invention solves the problem of insufficient coverage of night scene samples in the prior art, which easily leads to mode collapse and gradient vanishing when the model learns the noise distribution in dark areas. It has the following advantages: (1) This invention solves the data bottleneck and improves the modeling accuracy: by combining generative adversarial networks and dynamic weighting strategies, the effective training data coverage for nighttime scenes is significantly increased from the baseline of 30% to 68.7%, fundamentally alleviating the data bias problem. Based on this, the Bayesian-Transformer hybrid architecture achieves extremely high modeling accuracy, reducing the noise prediction error rate to 4.1% on authoritative datasets such as NuScenes-Night, a reduction of 72.7% compared to existing technologies;
[0031] (2) This invention significantly improves image enhancement quality and detail restoration capability: the region-weighted loss function guides the model to effectively focus on dark area optimization, making it more effective in low-light conditions. Under extreme conditions, the peak signal-to-noise ratio (PSNR) in the dark area was improved by 2.7 dB. The overall solution achieved a noise suppression rate of up to 28 dB, while effectively covering... With its wide dynamic range, the output image performs excellently in terms of brightness, contrast, color fidelity, and detail clarity.
[0032] (3) The present invention achieves the unity of high performance and low latency: Through algorithm optimization and hardware co-design, the present invention achieves excellent performance of end-to-end inference latency of less than 18.3ms when processing 4K resolution images on edge computing platforms such as NVIDIA Jetson AGX Orin, which is far lower than the real-time requirement of 33ms, while the power consumption is controlled within 8W, meeting the stringent constraints of embedded scenarios such as automotive. (4) The present invention has high reliability and functional safety: the dynamic adaptive closed-loop adjustment module ensures the robustness of the solution under complex scenarios such as sudden changes in lighting. Through more than 2,000 kilometers of real vehicle road testing, the false alarm rate of the system is less than 0.8% / 1,000 kilometers. The overall solution has reached the functional safety standard of ASIL-B, providing a reliable technical solution for night vision applications with extremely high safety requirements such as autonomous driving and intelligent security. Attached Figure Description
[0033] Figure 1 This is a flowchart illustrating the overall process of the low-light image enhancement system and method based on enhanced nighttime scene modeling, as described in this invention.
[0034] Figure 2 This is a schematic diagram of the Bayesian-Transformer hybrid noise modeling framework of the present invention.
[0035] Figure 3 This is a flowchart illustrating the optimization of the region-weighted loss function and the three-stage training strategy of this invention.
[0036] Figure 4 This is a schematic diagram of the dark area confidence estimation module of the present invention. Detailed Implementation
[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] It should be noted that: Unless otherwise specified in the examples, conditions should be followed according to standard conditions or the manufacturer's recommendations. Instruments whose manufacturers are not specified are all commercially available products. Raw materials and reagents whose manufacturers are not specified are all commercially available goods or can be prepared using known methods.
[0039] In this invention, all features defined in the form of numerical ranges or percentage ranges, such as numerical values, quantities, contents, and concentrations, are used only for simplicity and convenience. Accordingly, the description of numerical ranges or percentage ranges should be considered as covering and specifically disclosing all possible sub-ranges and individual numerical values (including integers and fractions) within those ranges.
[0040] The features mentioned in this invention can be combined arbitrarily, and all possible combinations should be considered within the scope of this specification, provided that there is no contradiction in the combination of these features. Each feature disclosed in the specification can be replaced by any alternative feature that provides the same, equivalent, or similar purpose. Therefore, unless otherwise specified, the disclosed features are merely general examples of equivalent or similar features.
[0041] In the description of this invention, it should be noted that the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0042] Example 1 A low-light image quality enhancement method based on enhanced nighttime scene modeling achieves end-to-end optimization from data generation to system verification by constructing a four-level technical closed-loop system. This method addresses core issues in existing technologies, such as high noise prediction error rates (15% vs. target ≤5%) due to insufficient nighttime scene sample coverage and gradient vanishing (signal-to-noise ratio decrease of 42%) caused by dynamic range compression, by employing a multi-level collaborative optimization strategy. Figure 1 As shown, the specific implementation includes the following steps:
[0043] (S100) Sample data construction: Based on an initial image dataset (such as a public dataset or a self-collected dataset containing clear daytime images), enhanced low-light scene images are generated through a dual-channel hybrid generative network architecture. Based on the content analysis results of the enhanced low-light scene images, the images are dynamically weighted to construct an optimized training sample library. (S200) Offline Model Training Phase: Only the optimized training sample library constructed in step (S100) is used as the input data for the hybrid augmentation model. The hybrid augmentation model is constructed based on an attention module for extracting global spatial dependencies and a probabilistic graphical model module for quantifying uncertainty. Iterative training is performed using a region-adaptive weighted loss function to determine the optimal weight parameters and obtain the optimized hybrid augmentation model. Specifically, the region-adaptive weighted loss function includes a pixel domain loss term and a structural similarity loss term. The pixel domain loss term is constructed using the Charbonnier loss function to ensure differentiability when the error is close to zero and to enhance robustness to outliers; the structural similarity loss term is calculated in the YUV color space, and the decoupling of brightness and chroma makes the augmentation effect more consistent with human visual characteristics.
[0044] (S300) Image Enhancement: This step is the online application stage; the low-light image to be processed is input into the optimized hybrid enhancement model, which estimates the illumination and noise distribution of the low-light image and generates the enhanced image accordingly. Simultaneously, during processing, the system calculates the predicted distribution of the current frame in real time through the distribution consistency monitoring module. Compared with prior target distribution KL divergence between As an early warning indicator of system uncertainty, it prevents the model from generating nonlinear distortion in extreme unknown scenarios; (S400) Dynamic adaptive closed-loop adjustment: This is achieved by constructing a dynamic adaptive closed-loop adjustment module, which includes a quality index constructed based on the Boltzmann distribution principle. An observer is used to quantify the degree of order in the enhanced image. The system establishes state variables. The dynamic evolution equation, where the uncertainty is... It is characterized by a three-dimensional vector consisting of spatial gradient energy, temporal KL divergence, and intensity relative error.
[0045] In step (S100), at the data augmentation level, an augmented dataset (i.e., a multi-level nighttime scene sample library) is constructed by using a dual-channel hybrid generative adversarial network architecture and a dynamic weight allocation strategy to cover foggy days, rainy nights, and extremely low illumination (illuminance <0.1 lux) scenarios. This strategy increases the nighttime scene coverage from the baseline of 30% to 68.7%, and the PSNR (peak signal-to-noise ratio) of the generated samples reaches 29.4 dB.
[0046] The dual-channel architecture includes a style transfer channel centered on CycleGAN and a category-aware generation channel centered on ACGAN, forming a complementary dual-path architecture. (1) Style transfer channel: used to realize style transfer across domain scenes (e.g., mapping daytime clear scene images to nighttime low light style, or mapping sunny scene to rainy and foggy weather style), thereby providing scene samples under different environmental conditions for the sample library and solving the problem of single style level; CycleGAN is the core component of this channel to realize the generation of diverse style dimensions. (2) Category-aware generation channel: This channel is responsible for generating diverse category dimensions for nighttime scenes, making up for the shortcomings of CycleGAN, which only focuses on style and lacks sufficient control over semantic categories. This channel provides nighttime scene samples of different categories and targets to the sample library, solving the problem of single scene content levels. For example, it generates multi-level samples containing pedestrians in different poses, various types of vehicles (such as cars and trucks), and various traffic signs.
[0047] Specifically, the CycleGAN branch uses a ResNet-based generator to achieve scene style transfer. This generator employs a three-stage topology of "encoding-transformation-decoding," which includes:
[0048] (1) Front-end encoding module: contains 3 convolutional layers, with the number of channels increasing layer by layer in the order of [64, 128, 256], used to extract shallow texture features of the image; (2) Intermediate transformation module: It consists of 9 (or 6) cascaded residual blocks. Each residual block contains two 3×3 convolutional layers and an instance normalization layer, and skip connections are introduced within the block to preserve high-frequency detail information. (3) Back-end decoding module: Contains 2 deconvolutional layers (or upsampling layers) and 1 output convolutional layer, with the number of channels decreasing layer by layer in the order [128, 64, 3], reconstructing the feature map into the target style image. The entire network uses LeakyReLU ( The activation function, except for the output layer, is configured with either batch normalization or instance normalization layers. Specifically, the generator introduces a skip connection structure to preserve low-frequency features, and its cycle consistency loss function is defined as:
[0049] (1) In equation (1), This is the Cycle Consistency Loss function; For mathematical expectation; A generator used to map a source domain image to a target domain; A generator used to map a target domain image back to the source domain; The input source domain image sample; These are the weighting coefficients for the identity mapping loss. This is used to maintain the identity mapping property between inputs and outputs; This is the Identity Mapping Loss term.
[0050] Specifically, in terms of hardware acceleration, the CycleGAN branch reduces the time taken for a single iteration from 78ms to 42ms by using 8-bit quantization training technology from the NVIDIA A100 Tensor Core GPU.
[0051] Specifically, the ACGAN branch works in parallel with the CycleGAN branch. The ACGAN branch is configured with a 9-layer fully convolutional discriminator architecture for multi-condition generation. Regarding the specific structure of the discriminator, this architecture aims to... The input image is progressively downsampled to feature vectors, without using pooling layers. Instead, downsampling is achieved through convolutional layers with a stride of 2. The detailed hierarchical connections are as follows:
[0052] (1) Input layer: receiving Image data; (2) Feature extraction layer (layers 1-8): 8 convolutional blocks are stacked consecutively, each block containing one feature extraction layer. The system uses convolutional kernels (stride = 2, padding = 2), spectral normalization layers, and LeakyReLU activation function (slope set to 0.2). As the number of layers increases, the feature map size is halved from 256 to 1 layer, and the number of channels is doubled from 64 to 512 layer by layer.
[0053] (3) Output layer (9th layer): The feature map of the last layer is flattened and divided into two branches.
[0054] The scene classification head attached to the ACGAN branch is essentially an auxiliary classifier embedded at the end of the discriminator. It shares the convolutional feature extraction weights of the first 8 layers with the real / fake discrimination module, and is separated only in the 9th layer. This classification head uses 12 one-hot encoded nighttime scene labels (such as foggy night, rainy night, snowy night, etc.) as conditional input, forcing the generator to not only meet the realism constraint when generating images, but also to possess the semantic features of specific scenes.
[0055] Furthermore, to prevent mode collapse when generating high-frequency nighttime noise (such as rain streaks or halos), this embodiment uses Wasserstein GAN-GP as the training strategy. Wasserstein GAN-GP is related to the aforementioned architecture in that it serves as an optimization strategy for the loss function and is applied during the training of the 9-layer discriminator. Here, the Lipschitz constraint refers to the requirement that the discriminator function... satisfy This embodiment introduces a gradient penalty term and sets a coefficient. =10, forcing the gradient norm of the discriminator to be close to 1 (i.e., satisfying the 1-Lipschitz constraint). This mechanism effectively limits drastic changes in the discriminator and solves the gradient vanishing problem in traditional GANs during deep network training.
[0056] To verify the effectiveness of the improved architecture, the inventors conducted ablation experiments. The baseline model used a standard 5-layer DCGAN without gradient penalty, while the proposed model employed the aforementioned 9-layer architecture combined with the WGAN-GP strategy. The experimental results are shown in Table 1.
[0057] Table 1. Comparison of ablation experiments with improved discriminator architecture Experiments show that this hybrid architecture generates When dealing with high-resolution images, the combination of deep structure and Lipschitz constraints effectively improves the diversity and realism of the samples, and the FID score is finally optimized to 21.3.
[0058] In step (S100), the dynamic weight allocation strategy achieves sample optimization allocation through the dynamic weight formula (2).
[0059] (2) In equation (2), For the first i The percentage of dark area pixels in each sample can be calculated using OTSU threshold segmentation. For the temperature coefficient, it is preferably set to [value] in this example. The settings have been rigorously verified: when T < 0.5, it will cause the weight distribution to become overly sharp (i.e., the model will over-focus on a very small number of extreme samples), while T > 1.0 will cause the attention to be scattered (tending to a uniform distribution). For the first j The percentage of dark area pixels in each sample is used for normalized summation calculation of the denominator; For the first i Attention weights for each sample; N sample This indicates the total number of samples in the current batch or dataset.
[0060] By employing the dynamic weight allocation strategy of this invention, the utilization rate of dark area samples in 2,975 nighttime driving videos in the Cityscapes dataset was effectively increased from 38% to 72%.
[0061] Furthermore, sample validity control is achieved through a hierarchical "three-level quality control" mechanism, specifically including: (1) First-level signal integrity monitoring: detect the signal-to-noise ratio and Laplacian sharpness of the samples, and remove artifact images that fail to generate; (2) Second-level semantic consistency monitoring: Use a pre-trained semantic segmentation network to verify the target structure constraints (such as the closure of vehicle contours) in the generated image. (3) Third-level distribution diversity monitoring: K-means clustering algorithm is used to detect sample distribution shift.
[0062] The three are arranged in a sequential cascade relationship to ensure that the samples entering the database are both high-quality and highly diverse. In the third level of monitoring, when K-means clustering detects the distance between the scene centroids... (in When the value is less than the standard deviation of the distribution, it indicates that there are data blind spots in the current scene, and the system will trigger the "active learning module". This module communicates with the hardware interface to control the Intel RealSense D455 stereo depth camera to perform targeted supplementary sampling of boundary samples and feed the supplementary data back to the sample construction module.
[0063] At the data management level, the storage architecture adopts a tiered storage strategy with hot and cold data layering, specifically configured with a combination of 2TB NVMe SSD solid-state drives and 20TB HDD hard disk drives to accommodate a streaming data ingestion speed of 2000 samples per second. The NVMe SSD serves as the hot data caching layer, leveraging its high IOPS characteristics to handle the high-speed writing of real-time generated samples; the HDD serves as the cold data archiving layer, used for long-term storage of historical sample libraries after three levels of screening.
[0064] Specifically, the sample distribution shift detection uses an improved K-means clustering algorithm (k=15). The so-called "improvement" means that the K-means++ strategy is used in the initialization phase to optimize the selection of the initial centroids and avoid getting trapped in local optima; and in terms of distance metric, the following statistical threshold determination mechanism is combined, which transforms the algorithm from a simple clustering tool into a distribution monitoring tool with anomaly detection capabilities.
[0065] The trigger condition is set to the distance from the center of mass. The mathematical principle is as follows: It is assumed that the feature vectors of various nighttime scenes follow a multivariate Gaussian distribution in the latent space. According to statistics... According to the Three-Sigma rule and its confidence interval derivation, under the assumption of a normal distribution, approximately 97.7% of normal sample points should fall within a distance from the center of the distribution. ( Within the range of the standard deviation of the distribution (one-sided upper confidence limit). That is:
[0066] Therefore, this embodiment selects This serves as a decision boundary. When the centroid offset of a newly acquired sample exceeds this threshold, it indicates that its distribution characteristics have significantly deviated from the core coverage area of the training set's statistical properties (falling within the remaining...). The long-tailed distribution range is identified, thus determining that a "sample distribution shift" has occurred, which triggers the active learning module to perform supplementary sampling.
[0067] To verify the above mechanism's ability to capture rare scenes and its distribution optimization effect, the inventors conducted a sample distribution consistency verification experiment. The experiment selected a test set containing rare long-tail scenes such as "aurotic nighttime" (accounting for <3.8% of the total dataset), and compared the "random sampling strategy" with the present invention's "active learning-based boundary supplementation sampling strategy." The experimental results are shown in Table 2:
[0068] Table 2. Experimental comparison of the active learning module for optimizing sample distribution. Real-world testing data shows that this standard can accurately detect rare scenes (such as aurora nights) with a coverage rate of <3.8% in the sample library. After triggering the active learning module, boundary samples acquired by the Intel RealSense D455 depth camera successfully reduced the KL divergence between the generated training set and the real-world distribution from 2.38 bits to 1.40 bits (a reduction of 41.2%), significantly improving the model's generalization ability in long-tailed scenes.
[0069] At the data throughput level, this system has constructed a dynamic update mechanism for the sample library. This mechanism adopts the Apache Arrow columnar storage format and supports a streaming data ingestion rate of 2000 samples / second with zero-copy technology. Combined with a hybrid architecture of 2TB NVMe SSD primary storage (hot data layer) and 20TB HDD cold backup (cold data layer), a storage access hit rate of 98.7% was achieved in a stress test involving 100,000 random sample reads.
[0070] Furthermore, in step (S300), before formally inputting the acquired image to be processed into the hybrid enhancement model, the original sensor data (such as RAW format data) is first parsed through a dedicated ISP (Image Signal Processor) preprocessing pipeline. This parsing process includes hardware-level demosaic processing and bit-width recalibration. To eliminate hardware read / write latency, the system employs a double-buffered asynchronous transmission mechanism to send the parsed image data stream into the computing unit memory, ensuring that the front-end data throughput rate and the back-end model inference rate are matched in real time.
[0071] Furthermore, in step (S200), to address interference from non-uniform illumination or localized strong light (such as industrial arc light or strong reflection), the hybrid enhancement model supports dual-modal feature fusion at the input end. In addition to the visible light sensor, the system optionally integrates auxiliary sensors such as an infrared thermal imager. It extracts the correlation information between visible light features and thermal imaging features through a cross-modal attention mechanism (as shown in Equation 16), and dynamically adjusts the fusion weights according to the ambient light and temperature distribution. This is to compensate for areas of overexposure or loss of detail in visible light.
[0072] Meanwhile, the hybrid enhancement model integrates a frequency domain denoising submodule in the noise suppression stage. This submodule first uses wavelet packet transform (such as the Daubechies-9 wavelet) to decompose the image into multiple high-frequency and low-frequency subbands, and then applies a Bayesian threshold-based denoising algorithm (as shown in Equation 15) to specific subbands. This algorithm uses the noise distribution information estimated by the aforementioned probabilistic graphical model to calculate the Bayesian threshold in real time. This enables the synergistic suppression of high-frequency pulse noise and low-frequency texture noise.
[0073] Subsequently, the low-light image to be processed (i.e., real-time image acquired in a real-world application scenario) is input into a pre-trained hybrid enhancement model. This model is built upon an attention module for extracting global spatial dependencies and a probabilistic graphical model module for quantifying uncertainty. Specifically, the probabilistic graphical model module is built on a variational autoencoder (VAE) framework. To achieve real-time inference in an in-vehicle environment, the hybrid enhancement model undergoes deep optimization at the computational level: First, layer fusion technology is used to merge matrix multiplication and Softmax operations in the attention module into a single computational operator, increasing computational density by 1.8 times; second, layered hybrid precision quantization is implemented, maintaining FP16 high precision for the probabilistic graphical model module to ensure the accuracy of uncertainty quantization, while INT8 quantization is used for the feature extraction path, leveraging heterogeneous computing resources to improve processing efficiency. This hybrid enhancement model is ultimately used to estimate the illumination and noise distribution of the low-light image and generate the enhanced image accordingly.
[0074] Furthermore, to enhance the system's robustness in complex, non-uniform lighting environments such as industrial settings, the hybrid enhancement model supports dual-modal data fusion of visible light and infrared thermal imaging at the input. The system utilizes a cross-modal attention mechanism (as shown in Equation 16) to perform feature compensation for locally bright or extremely dark areas using thermal imaging information. During processing, the model integrates a frequency domain denoising submodule. This module maps the image to the frequency domain using wavelet packet transform (such as Daubechies wavelet decomposition) and, combined with the aforementioned noise distribution estimation results, applies a Bayesian threshold denoising algorithm (as shown in Equation 15) to accurately filter out noise in specific frequency bands.
[0075] In step (S200), the aforementioned hybrid enhancement model employs a Bayesian-Transformer hybrid architecture. The core of this architecture lies in constructing a two-stream coupling mechanism between "deterministic feature encoding" and "probabilistic uncertainty inference." For example... Figure 2 The diagram shown illustrates the Bayesian-Transformer hybrid noise modeling framework of this invention. It details the technical aspects of extracting globally deterministic features using a Transformer encoder and coupling it with a hierarchical Bayesian probabilistic graphical model for estimating illumination and noise distribution. Specifically, the architecture comprises the following two cascaded sub-modules:
[0076] (1) Transformer encoder (deterministic path) as backbone network: This module first segments the input low-light image into fixed sizes (e.g., Image patches are linearly projected and positionally encoded before being fed into a deep network containing a 12-head self-attention mechanism. Specifically, to ensure the training stability of the deep network, this module integrates a spectral normalization layer into all linear projection layers. This layer dynamically adjusts the singular values of the weight matrix to constrain the Lipschitz constant of each layer. Within a certain range, this effectively suppresses gradient explosion. This module utilizes the global receptive field of the attention mechanism to capture long-range dependencies between different image subspaces in parallel, outputting a high-dimensional deterministic feature representation rich in global contextual information. .
[0077] (2) Hierarchical Bayesian Probabilistic Graphical Model (Probabilistic Path): This module is "coupled" to the output of the Transformer encoder, and its construction is based on a generative framework of Variational Autoencoder (VAE). Specifically, it treats the features output by the Transformer as the encoding result, and represents the above features through multiple multilayer perceptron (MLP) heads. Mapping to latent variables The posterior distribution parameters (such as the shape parameters of the Gamma distribution) Sum parameter In the hierarchical probabilistic graph structure established here, the top-level latent variables represent the uncertainty of the global illumination distribution, while the bottom-level latent variables represent the local noise variance constrained by the top-level variables. This hierarchical design enables the model to dynamically output enhancement results with confidence intervals based on the degree of blurring of the image content, thereby achieving accurate modeling and suppression of complex nighttime noise patterns.
[0078] Specifically, the spatial feature extraction process of the 12-head attention mechanism is related to Gamma ( The hierarchical Bayesian probabilistic graphical model formed by the prior distribution is mathematically expressed as follows: (3) In equation (3), Indicates that in a given input image x Under the condition of latent variables z The posterior distribution; Transformer ( x ) indicates that the Transformer encoder processes the input image x The feature representation obtained after feature extraction is transformed into statistical parameters of the distribution through a mapping layer; Indicated byα and β The Gamma distribution with parameters is used to quantify the uncertainty of illumination and noise distribution, where the shape parameter... α with rate parameter β All of these are predicted by the feature representation Transformer(x) through nonlinear transformation.
[0079] Specifically, the query-key matching strategy in the multi-head attention mechanism is used to capture long-range dependencies between image patches, while Bayesian layers are used to quantify the uncertainty of feature representations based on the principle of variational inference. These processes together construct a hierarchical Bayesian probabilistic graphical model, enabling it to more accurately describe complex noise patterns in low-light nighttime images, providing a more reliable basis for subsequent image enhancement.
[0080] To verify the accuracy advantage of the proposed Bayesian-Transformer hybrid architecture in noise distribution modeling, the inventors conducted a distribution fit comparison experiment. The authoritative NuScenes-Night dataset (containing approximately 35,000 nighttime street view images) was used to compare the proposed model with a baseline traditional method (a ResNet-based deterministic noise estimation network). The evaluation metric used was KL divergence (KL divergence), which measures the noise distribution predicted by the model. With respect to the actual noise distribution The lower the value, the more accurate the modeling.
[0081] The experimental results are shown in Table 3: Table 3 Comparison of Noise Distribution Modeling Accuracy Experimental Results Experimental results show that traditional methods often produce large distribution biases (KL divergence as high as 2.3 bits) when dealing with non-uniform noise at night due to their lack of ability to model uncertainty. However, the hybrid architecture of this invention, by combining global feature extraction of Transformer with probabilistic inference of Bayesian network, successfully reduces the KL divergence of noise distribution modeling to 0.8 bits. This means that the model can more realistically reproduce the physical characteristics of noise in complex nighttime environments.
[0082] The above hybrid enhancement model has a light intensity mapping function, as shown in formula (4): (4) In equation (4), Total number of samples for Monte Carlo sampling; For the first The random sampling weights of each sample follow a normal distribution with a mean of 0.5 and a standard deviation of 0.1. ; The input is the normalized pixel value; This is the sensor's black level offset, used to correct dark current noise. Let be the sensor response slope parameter, where The measured data of the photoelectric conversion characteristic curve from the Sony IMX490 sensor are set as follows: This value was obtained through laboratory calibration: under illuminance 300 sets of raw RAW data were collected within the specified range, and the response slope was obtained by least squares fitting. (95% confidence interval) This parameter reflects the non-linear mapping relationship between pixel values and actual illuminance.
[0083] The above hybrid enhancement model also constructs a probability transfer function between illumination intensity and noise variance, as shown in formula (5): ; In equation (5), The predicted noise standard deviation; Light intensity; This is the sensor noise response coefficient. This parameter, also based on darkroom calibration data of the Sony IMX490 sensor, reflects the exponential growth characteristic of the sensor's thermal noise with changes in light intensity. It has been measured... ; This is a dynamic adjustment coefficient. .
[0084] The above dynamic adjustment coefficient Real-time optimization is achieved through adaptive update rules, which are shown in formula (6): ; In equation (6), For dynamic adjustment coefficients, subscript For training steps; For learning rate, ; This is the total loss function; The standard deviation of noise; To prevent the stability constant from having a denominator of zero, .
[0085] To verify the accuracy of the Bayesian-Transformer hybrid architecture in modeling noise distribution under extremely low illumination conditions, the inventors designed a specific comparative experiment based on the NuScenes-Night dataset.
[0086] 1. Experimental setup Dataset: 5,000 extreme nighttime scene images with illumination below 1 lux were selected from the NuScenes-Night dataset as the test set.
[0087] Truth acquisition: The noise variance of the static region is extracted using the multi-frame temporal averaging method as the "Ground Truth".
[0088] Baseline: The current mainstream CNN-based deterministic noise estimation network (ResNet-Est) was selected as the control group. 2. Evaluation Indicators Error Rate (MAPE): Measures the mean absolute percentage error (MAPE) between the predicted variance and the true variance. KL divergence: measures the predicted noise probability distribution. With the true distribution The information difference between them is expressed in bits.
[0089] 3. Experimental Results The experimental results are shown in Table 4: Table 4 Comparison of Noise Distribution Modeling Performance 4. Results Analysis Real-world data shows that the control group, due to its deterministic point estimation, fails to effectively fit the long-tailed characteristics of the distribution when facing complex non-uniform noise at night (such as halo noise under streetlights), resulting in a KL divergence as high as 2.3 bits. In contrast, the hybrid architecture of this invention utilizes a Transformer to capture global illumination dependence and combines it with Bayesian inference of output distribution parameters, successfully reducing the noise prediction error rate from 15.3% to 4.1% and optimizing the KL divergence to 0.8 bits. This demonstrates that this model can not only accurately predict noise intensity but also precisely quantify the uncertainty (confidence) of the prediction results, thus effectively solving the "mode collapse" problem that traditional methods easily encounter in dark areas.
[0090] Furthermore, in step (S200), the system integrates a distribution consistency monitoring module. This module performs real-time monitoring of the Bayesian probability branch in the hybrid augmentation model, aiming to prevent the model from producing "illusions" or severe distortions in extremely unknown scenarios. Its core mechanism is to calculate the Kullback-Leibler divergence between the current predicted distribution and the prior distribution in real time. As shown in formula (7):
[0091] (7) In equation (7), This represents the number of sampling steps within the time window. The target distribution (i.e., the statistical prior distribution of nighttime scenes in the training set); This represents the predicted distribution of the hybrid enhancement model's output for the current frame image.
[0092] The technical role of real-time calculation of KL divergence is to serve as an early warning indicator of system uncertainty. When... If the noise pattern of the current input image deviates significantly from the range of knowledge learned by the model (i.e., a distribution shift has occurred), the system will determine that the current enhancement result is unreliable and trigger the ModelSelf-Calibration Process. This process specifically includes the following online adaptation steps:
[0093] 1. Parameter Freeze and Activation: Temporarily freeze the backbone weights of the Transformer encoder and activate only the dynamic adjustment coefficients in the hybrid augmentation model. As trainable parameters; 2. Gradient backpropagation: Based on the current KL divergence bias, the gradient is calculated using the adaptive update rule defined in the aforementioned formula (6). ; 3. Online parameter update: updating parameters along the gradient direction. Perform single-step or multi-step fine-tuning to adjust the predicted noise distribution. Rapidly approaching the target distribution This brings the KL divergence back to a safe range of 0.8 bits. During this process, the system simultaneously applies hierarchical regularization techniques (such as spectral normalization constraints) to prevent excessive parameter updates from causing the model to diverge.
[0094] Furthermore, to prevent model overfitting and gradient explosion, this embodiment employs hierarchical regularization. This technique applies constraints to different components of the hybrid enhancement model from the following three key levels:
[0095] (1) In the spatial dimension (for the Transformer encoder): By applying spectral normalization constraints to the multi-head projection matrix in the 12-head attention mechanism, the maximum singular value of the weights in each layer is limited, thereby ensuring the Lipschitz constant of each layer. This ensures the stability of the gradient in deep networks; (2) In the probability dimension (for Bayesian probabilistic graphical models): Based on the theoretical framework of variational autoencoders (VAEs), the KL divergence term calculated by the aforementioned formula (7) is... It is directly added as a regularization loss term to the total loss function, where the prior distribution The distribution is set to be isotropic Gamma distribution, thereby constraining the posterior distribution to not deviate from the preset physical laws; (3) In the feature dimension (for multi-head attention mechanism): design consistency loss for cross-head attention. Its mathematical expression is (where F is the Frobenius norm), by penalizing the similarity between different attention heads, each attention head is forced to focus on different image subspace features, thus preventing the learning of redundant features.
[0096] The combined effect of these three regularization methods allows the model to maintain strong representational capabilities while also... By keeping the bit depth below 0.8 bits, the distribution matching accuracy is improved by 65.2% compared to the baseline model.
[0097] The combined effect of these three regularization methods allows the model to maintain strong representational capabilities while also... KL The divergence was controlled below 0.8 bits, resulting in a 65.2% improvement in distribution matching accuracy compared to the baseline model.
[0098] In step (S300), the region-weighted loss function is designed for the optimization training phase, as shown in formula (8): (8) In equation (8), the first term is the pixel domain loss term. In this embodiment, the Charbonnier loss function (a robust variant of L1 loss) is used for construction, i.e. Compared to ordinary L2 loss, it can handle outliers better; the second term is the structural similarity loss term. In this embodiment, the SSIM index is calculated in the YUV color space. That is, the RGB image is first converted to YUV format, and the structural differences are calculated only for the Y channel (luminance) and UV channel (chrominance) to conform to the visual characteristics of the human eye being more sensitive to luminance.
[0099] In step (S300), in the implementation of the region-weighted loss function, the dark area confidence estimation module adopts ResNeXt-101 (32×8d) as its basic architecture, which innovatively introduces a multi-scale feature pyramid in the fourth residual stage. For example... Figure 4The diagram shows the structural principle of the dark area confidence estimation module of this invention. It focuses on the topology of introducing a pyramid pooling module (PPM) between the third and fourth residual stages of the ResNeXt-101 architecture, as well as the fusion method of the identity mapping path and multi-scale parallel branches. Specifically, the input image first passes through a 7×7 convolutional layer (stride 2) and a downsampling layer, then enters four residual stages, each containing 3, 4, 23, and 3 bottleneck blocks, respectively. In particular, a pyramid pooling module is added after the third stage, capturing multi-scale contextual information through four parallel average pooling branches: 1×1, 2×2, 3×3, and 6×6. The outputs of each branch are upsampled to a uniform size after being compressed through a 1×1 convolutional channel and then stitched together. This structural design enables the network to simultaneously perceive local details and global illumination distribution. Experimental results show that the localization accuracy of dark area boundaries is improved by 12.7% compared to single-scale methods.
[0100] The key to generating confidence in the dark area lies in the design of the dynamic weight calculation function, which is mathematically expressed as: (9) In equation (9), For the pixel domain loss, there are dynamic weighting coefficients. Confidence level for the dark area.
[0101] The construction principle of the above dynamic weight calculation function is based on three considerations: (1) A base value of 0.3 ensures that even in non-dark areas ( Moderate oversight will still be maintained. (2) The steepness coefficient of the Sigmoid function is set to 5 and the center point to be 0.7. This has been verified experimentally. The time weight quickly approaches 1.2, which aligns with the physical intuition that dark pixels require stronger constraints; (3) An amplitude coefficient of 0.9 will precisely control the output range within the interval [0.3, 1.2].
[0102] Gradient analysis shows that this dynamic weight calculation function... It has the maximum gradient at point This effectively promotes the network's focus on learning in uncertain areas.
[0103] In step (S300), during the training phase, the optimizer updates the parameters along the gradient direction to improve the performance of the hybrid augmentation model. The learning rate is a core hyperparameter controlling the step size of parameter updates during model training, and the gradient indicates the direction and magnitude of the parameter updates. Figure 3 The diagram shown is an optimization flowchart of the region-weighted loss function and the three-stage training strategy of this invention, illustrating how to optimize based on the confidence level of dark areas. Dynamically adjust pixel domain loss weights The training process employs a hierarchical adaptive gradient clipping approach. The training phase utilizes a three-stage learning rate scheduling, starting with the initial stage... linear decay to Then converted to cosine annealing to Simultaneously, gradient management implements a layer-wise adaptive clipping strategy, individually constraining the gradient norm of each layer in the neural network, with the constraints shown in the inequality: ,in This is the gradient vector of this layer. The L2 norm of the gradient. This is the weight matrix. Let Frobenius norm be the weight. This strategy will... A dynamic clipping threshold is set, directly coupling the clipping intensity to the magnitude (Frobenius norm) of the layer's weight matrix, thereby achieving adaptive gradient stabilization control between layers. Combined with three-stage learning rate scheduling and hierarchical gradient clipping, this achieves stable control under varying illumination conditions. Under these conditions, the PSNR in the dark area is increased by 2.7 dB.
[0104] Compared with traditional L1 or L2 loss, the combined loss function based on Charbonnier and YUV-SSIM designed in this invention has advantages in the following three aspects: (1) At the gradient propagation level, the improved Charbonnier loss (i.e. It maintains differentiability when the error is close to zero, effectively avoiding training oscillations caused by the non-differentiability of ordinary L1 loss at zero, so that the model can converge to a finer minimum value during the fine-tuning stage. (2) Regarding color fidelity, the SSIM loss calculation in the YUV space In the formula, The pixel mean of the target image (true value); The pixel mean of the enhanced image (predicted value); The variance of the target image; To enhance the variance of the image; The covariance between the target image and the enhanced image; , To prevent the stability constant from having a denominator of zero, a certain value is usually set. , Compared to direct calculations in the RGB space, this calculation method based on the YUV space decouples luminance and chrominance, which is more in line with the characteristics of the human visual system (HVS) being highly sensitive to changes in luminance;
[0105] (3) Regarding training stability, the dynamic weight mechanism (i.e., in Formula 9) This method increases the gradient contribution in dark areas from 18% to 43% compared to traditional methods, accelerating model convergence by 1.8 times. Experimental data show that this loss combination improves PSNR by 2.7 dB under extremely dark conditions (0.1 lux) while controlling the gradient variance within a certain range. The stable range.
[0106] To comprehensively evaluate the robustness and practicality of the system, this embodiment constructs a multi-level verification process (i.e., the "three-dimensional system") that includes theoretical stability analysis, quantitative performance evaluation, and real-vehicle scenario verification. The specific construction and implementation of this process are as follows:
[0107] 1. Theoretical stability analysis (theoretical dimension): Based on the principle of spectral normalization, the Lipschitz constant of the feature space is mathematically derived and verified. The gradient stability of the model during deep feature extraction was mathematically proven, ensuring that the algorithm does not diverge under extreme input conditions. 2. Quantitative Performance Evaluation (Experimental Dimension): Black-box testing based on a standard test set verifies that the system's mean absolute error (MAE) for dark area pixel recovery is stably controlled within a certain range. This demonstrates the model's accuracy in restoring details in low-light conditions; 3. Real-world vehicle scenario verification (application dimension): Deploy the system on an onboard computing platform and conduct 2000 km of real-world road testing to verify its false alarm rate under real-world lighting change scenarios. Thousands of kilometers.
[0108] Through comprehensive verification across the three dimensions mentioned above, it was finally confirmed that the system achieved a noise suppression rate of 28 dB and a dynamic range coverage of 120 lux, with end-to-end processing latency controlled at 18.3 ms @ 4K resolution, fully meeting the stringent requirements of automotive applications for real-time performance and reliability.
[0109] In step (S400), the system dynamically adjusts the sensor's exposure time and sensitivity based on uncertainty error feedback using a PID controller (as shown in Equation 13). The enhanced image is further fed into the output optimization module, where a Sigmoid-Curve dimming curve (Equation 14) is applied for dynamic range mapping. This enhances details in dark areas while utilizing the progressive saturation characteristics of the high-brightness region to prevent highlight clipping. Finally, the system ensures that its operation meets ASIL-B functional safety standards through a multi-level monitoring system consisting of a hardware limiting circuit, an algorithm backup model, and vehicle bus linkage. Specifically, the multi-level monitoring system calculates the uncertainty propagation model in real time using Monte Carlo sampling (as shown in Equation 17). When uncertainty or gradient abrupt changes (as shown in Equation 18) exceed a threshold, an early warning is triggered, and an emergency processing unit is activated. During the system's recovery to the main enhancement model, a progressive weighted mixing strategy (as shown in Equation 19) is employed to ensure the continuity of visual feedback.
[0110] In step (S400), in the specific implementation of the dynamic adaptive closed-loop adjustment module, the module constructs a feedback control loop that connects the image enhancement result with the sensor hardware parameters.
[0111] In this loop, the quality index Designed as the core observation state variable of the system, it is used to quantify the "order" of the current image quality. Its construction is based on the Boltzmann distribution principle in statistical mechanics, treating noise as "thermal motion" interference of the system. The core formula is as follows:
[0112] (10) In equation (10), The total number of samples; As weight; For indicator functions; For the sample ; This is a nighttime dataset; For example The noise variance; This is the noise sensitivity coefficient.
[0113] The above quality index The core formula can be viewed as the expression of the generalized partition function in a discrete sample space. Among them, the exponential term... Analogous to energy probability distribution, temperature parameter The control operator controls the system's sensitivity to noise variance. Validated through Monte Carlo sampling, this exponential form enables... right The change exhibits an S-shaped response curve, with its inflection point located at... At this location, it can match the typical noise variance characteristics of the Sony IMX490 sensor at a high ISO setting of 12800. This matching mechanism ensures that the dynamic adaptive closed-loop adjustment module can generate sensitive status feedback in a timely manner when the sensor enters the high-noise operating region.
[0114] The detailed structure and function of this dynamic adaptive closed-loop control module are described below: System Definition and Function: This system is a feedback control system that includes a state observer and a PID controller. Its function is to monitor the quality index of the enhanced image in real time. With uncertainty Based on this, the physical parameters of the front-end image sensor (ISO sensitivity and exposure time) are dynamically adjusted. This is to maintain stable image quality in scenarios with sudden changes in lighting, such as tunnel exits.
[0115] Quality Index Uses: As a feedback input variable of the control system, it is used to characterize the probability that the current scene sample belongs to the "effective night scene". A higher value indicates lower image noise and more effective information; conversely, a lower value triggers subsequent exposure parameter compensation.
[0116] 3. The specific control structure of the system: State Evolution Layer: System establishment based on The state-space equations (as described in Equation 11 below) utilize Derive the current input noise level ; Dual-path control layer: includes a feedforward path and a feedback path. The feedforward path adjusts according to ambient light intensity. The table is used to set the baseline ISO; the feedback path then calculates the uncertainty error. The exposure time is finely adjusted using a PID controller (as described in Formula 13 below). This forms a complete closed-loop control.
[0117] In step (S400), in the implementation of the dynamic adaptive closed-loop adjustment module, based on the noise level... With system uncertainty The correlation equations are used to establish the following dynamic system model. This model, expressed in state space, feeds back the output of the enhancement algorithm to the hardware control layer to achieve predictive control of the sensor parameters.
[0118] First, define the state variables. , T Its evolution process is described by discrete-time state equations: (11) In equation (11), This is the system matrix. Each row in the matrix represents the evolution logic of the system state variables at the next time step, and each column represents the linear contribution weight of the current state variable to the system evolution. In this example, the system matrix... Its off-diagonal elements reflect The causal chain shows that changes in image quality index can cause fluctuations in noise level, thus affecting the uncertainty of the system. Input matrix. Corresponding ISO - Exposure Control .in, The baseline control gain is obtained by querying a preset sensor response lookup table based on real-time illuminance. for The system state vector at time t; This represents system process noise, used to compensate for modeling errors.
[0119] The calculation uses hyperbolic tangent normalization: (12) slope parameter This value was obtained by logistic regression fitting of contrast test curves at different noise levels, ensuring that... Within the range It exhibits an approximately linear response, and this range covers 90% of the actual operating points.
[0120] Uncertainty Then through three-dimensional vectors The comprehensive characterization, specifically calculated as follows: Spatial gradient energy The degree of loss of spatial details is characterized by calculating the variance and gradient distribution entropy of local regions in the enhanced image. Time KL divergence The temporal stability of the video stream is characterized by calculating the KL divergence of the predicted distribution between adjacent frames (as shown in Equation 7). relative strength error : Calculate the deviation between the enhanced pixel mean and the physical light intensity mapping value (Formula 4) to characterize the accuracy of brightness restoration.
[0121] In step (S400), the transfer function block diagram of the ISO-exposure control mechanism includes both feedforward and feedback paths. The feedback path adjusts the exposure time via a PID controller. :
[0122] (13) In equation (13), For error signals, ; The proportional gain coefficient is set to 0.8. The integral gain coefficient is set to 0.15. The differential gain coefficient is set to 0.05; the control parameters are determined by the Ziegler-Nichols tuning method to balance response speed and overshoot risk.
[0123] The dynamic adaptive closed-loop control module exhibited the following characteristics in the step response test: rise time Frame, overshoot steady-state error This meets the requirements for real-time processing. In terms of hardware implementation, parallel computing is achieved through the PL (Programmable Logic) side of the Xilinx Zynq UltraScale+, ensuring that the control loop latency is <0.8ms.
[0124] Example 2 This invention presents a specific implementation example of the low-light image quality enhancement system based on enhanced night scene modeling in a vehicle environment. This embodiment uses the Sony IMX490 image sensor as the hardware foundation and elaborates on the system deployment scheme and performance optimization strategies as follows: In terms of storage architecture, a 2TB NVMe SSD (Samsung 980 Pro) is used as the main storage device, and a 20TB HDD (Seagate Exos X20) is configured to form a RAID5 array for cold backup storage. The stripe size is set to 256KB, and the parity algorithm uses Reed-Solomon encoding. The actual test shows that this configuration can achieve a continuous write speed of 2000 samples / second under the Apache Arrow columnar storage format, while ensuring that the data reconstruction time does not exceed 4.5 hours in the event of a single disk failure.
[0125] The processing platform chosen is the NVIDIA Jetson AGX Orin, whose core acceleration strategy includes three levels of optimization: (1) First, the matrix multiplication and Softmax operation in the 12-head attention mechanism of the Transformer encoder are merged into a single operator by using the layer fusion technology of TensorRT 8.6, which improves the computational density by 1.8 times; (2) Secondly, hybrid precision quantization is implemented to maintain FP16 precision for the hierarchical Bayesian probabilistic graphical model, while INT8 quantization is used for the spatial attention path, and heterogeneous computing is achieved by utilizing the 2048 CUDA Cores and 64 Tensor Cores of the Orin chip. (3) Finally, dynamic shape optimization was used to process the changing image resolution in the vehicle scene, achieving a real-time processing capability of 42 FPS at 1080p input. Quantitative tests showed that the conversion from FP32 to INT8 reduced the model size from 3.2GB to 1.8GB, while the processing latency decreased from 48.7ms to 31.2ms.
[0126] The sensor adaptation phase focuses on resolving the RAW12 data stream parsing issue of the IMX490, and a dedicated ISP preprocessing pipeline is designed, as follows: First, analog-to-digital conversion is performed using a 14-bit ADC (quantization step size). Then, the ON Semiconductor AR0234 chip was used to implement hardware-level de-mosaic, and finally, a double buffering mechanism was used to send the processed image into DDR4 memory.
[0127] To match the in-vehicle environment, the dynamic range control implements an ISO-exposure interlock strategy: when the ambient illuminance is below 2 lux, it automatically switches to ISO 12800 mode and extends the exposure to 50ms. The Sigmoid-Curve dimming curve utilizes its progressive saturation characteristics in the high-brightness region to prevent highlight blowout. Real-world testing data shows that in abrupt changes such as entering or exiting tunnels, this solution reduces the system response time to only 2.3 frames, a 67% improvement over traditional methods.
[0128] The Sigmoid-Curve dimming curve equation serves as the dynamic range mapping function, as shown in formula (14): ; In equation (14), To output pixel values; The brightness component of the input image is used for enhancement. The maximum brightness threshold supported by the display device (e.g., 255); In this embodiment, it is preferably set to 0.25, which is used to control the companding rate of the contrast. This is the brightness center offset, used to define the equilibrium point of the curve; Brightness compensation reference value.
[0129] To ensure reliability, a multi-level monitoring system is constructed, as follows: Specifically, the output optimization module is executed after the hybrid enhancement model has completed image generation and before the image data is output to the display terminal or downstream algorithm.
[0130] The signal transmission and correlation between this module and other modules are as follows: The input end of the output optimization module is connected to the output end of the hybrid enhancement model to receive enhanced high dynamic range (HDR) or floating-point image data; at the same time, the control end of this module is electrically connected to the dynamic adaptive closed-loop adjustment module (step S400) to receive dynamic parameters (such as the slope kslope and center offset I0 in Formula 14) calculated by the closed-loop system based on the ambient light intensity (Iv) and uncertainty (Uout).
[0131] After performing dynamic range mapping, the main changes to the image are as follows: 1. Numerical mapping: Non-linearly maps high bit-width pixel values (such as >255 or floating-point numbers) that may overflow after enhancement to the standard display range (such as 0-255) to ensure data format compatibility; 2. Visual optimization: The highlight areas are "soft clipped" using the Sigmoid-Curve dimming curve (Formula 14) to prevent overexposure (highlight clipping) in strong light areas such as tunnel exits, while maintaining a high contrast slope in the dark areas. This solves the common problems of "washed-out" or "artifacts" in night-enhanced images while preserving enhanced details in the dark areas.
[0132] (1) At the hardware level, an analog limiting circuit is constructed using the TI LM3886 chip, and the uncertainty threshold U=0.45 is set to 3.3V; (2) At the algorithm layer, the three-dimensional monitoring vector is calculated in real time. When 3 consecutive frames Activate the MobileNetV3 lightweight backup model at the time; (3) At the system level, it is linked with the vehicle ECU via CAN bus to trigger 850nm infrared supplementary light in an emergency.
[0133] 2000 km road tests show that the proposed multi-level monitoring system deployment scheme has low total system power consumption. Under the constraints, a noise suppression rate of 28 dB was achieved, keeping the PSNR in the dark area stable above 29 dB. Simultaneously, through coordinated monitoring at the hardware, algorithm, and system levels, the probability of system failure was effectively reduced, enabling it to meet the ASIL-B level automotive functional safety standard.
[0134] Example 2 verified the robustness of the system through multiple extreme scenarios. In extremely dark environments with illumination below 0.1 lux, a three-dimensional noise distribution analysis of the image (i.e., statistical analysis with pixel coordinates as the base and noise variance as the height axis) revealed the following:
[0135] 1. The noise variance surface of the traditional histogram equalization method exhibits a large number of severe pulse-like spikes, and the mean variance rises rapidly to above 0.25, resulting in severe visual noise spots. 2. In rainy, high-humidity environments, the RetinexNet method exhibits significant color cast due to inaccurate estimation of light components, resulting in a noticeable color difference value. The mean is greater than 12.5; 3. This solution leverages the Bayesian-Transformer hybrid modeling framework to accurately capture noise uncertainty, maintaining a smooth 3D noise surface with a stable variance within the range of 0.08 to 0.12 and a low color bias. It dropped below 3.2.
[0136] Quantitative test data show that this solution achieves excellent performance in foggy weather with a PSNR of 29.4 dB and an SSIM of 0.893, representing an improvement of 42.6% (PSNR) and 23.8% (SSIM) compared to the comparison method. Furthermore, the cross-scenario volatility is less than 5%, demonstrating outstanding stability.
[0137] In particular, during 2000km of real-world road testing, this system demonstrated rapid adaptive capabilities in the typical extreme scenario of tunnel entry. The time-series analysis shows that when a vehicle enters a tunnel at 80km / h (illuminance from...),... sudden drop The system completes parameter adjustments within 2.3 frames (approximately 76ms), as detailed below:
[0138] First, the ISO-exposure interlocking mechanism switches to ISO12800 mode in the first frame and extends the exposure to 50 ms. Then, in the second frame, the output optimization module is activated, and the Sigmoid-Curve dimming curve (with a dimming slope of 0.25) is applied to perform dynamic range mapping on the enhanced image to prevent highlight clipping. Finally, in the third frame, an enhanced image with noise suppression of 28dB is output.
[0139] Spatial uncertainty in the entire above process The power consumption remained consistently below 0.35, verifying the system's robustness. Hardware monitoring data showed that the Orin chip's peak power consumption in this scenario was only 7.2W, fully meeting the energy efficiency requirements of automotive equipment.
[0140] Example 3 This invention presents a specific implementation example of the low-light image quality enhancement system based on enhanced nighttime scene modeling in an industrial security scenario. This embodiment uses a factory with non-uniform lighting as a typical application scenario, focusing on analyzing the enhancement process of the system under complex industrial lighting conditions, as detailed below: In terms of hardware configuration, the ON Semiconductor AR0820AT image sensor is used as the front-end acquisition device, which features an ultra-high dynamic range of 83dB and 4K resolution. After conversion by a 12-bit ADC, it is connected to the Xilinx Versal ACAP adaptive computing acceleration platform. In particular, to cope with the local strong light interference commonly found in factory environments (such as welding arc light), the system integrates a FLIR A315 infrared thermal imager as an auxiliary sensor to construct a visible light-thermal imaging dual-modal data fusion channel.
[0141] Regarding frequency domain noise suppression, this embodiment achieves precise removal of dark area noise through a combination of wavelet packet transform and an attention mechanism. Specifically, after performing a 4-layer Daubechies-9 wavelet decomposition on the input image, a Bayesian thresholding denoising algorithm is applied to the third high-frequency sub-band. The threshold calculation formula is as follows:
[0142] (15) In equation (15), Real-time estimation using a Bayesian-Transformer hybrid noise modeling framework; The signal length; This is the Bayesian threshold.
[0143] Frequency domain analysis shows that this solution achieves a noise suppression ratio of 28.3dB in the 0.5~10kHz frequency band, which is 9.7dB higher than that of traditional ISP processing pipelines.
[0144] The multimodal feature fusion module is implemented through a cross-modal attention mechanism, and its core operation is as follows: (16) In equation (16), , These are the query vector and key vector for visible light features, respectively. The feature dimension of the key vector (used to scale the dot product result to prevent gradient vanishing); , These are the query vector and key vector for thermal imaging features, respectively; Its feature dimension; The fusion weights (or attention weight matrices) for visible light features and thermal imaging features are dynamically adjusted through Softmax normalization or gating mechanisms. This represents the matrix transpose operation.
[0145] Experimental data show that the design of the cross-modal attention mechanism enables high-temperature regions ( The detail retention rate is improved by 42%, while avoiding the "heat blurring" phenomenon commonly found in traditional methods.
[0146] The system was deployed in an automotive welding production line for field verification, a scenario that presents the following typical challenges: localized welding arc light ( ) and equipment shadow area ( The solution addresses the extreme dynamic range of welding defects, high-frequency noise from sparks, and reflection interference from high-temperature metal surfaces. Performance tests show that, while maintaining real-time processing at 30fps, the signal-to-noise ratio in dark areas increased from 14.2dB to 28.5dB, and the false positive rate in high-temperature areas decreased to 3.2%. Thermal imaging data confirms that the attention mechanism assigns a weight range of 0.35 to 0.45 to high-temperature areas, effectively suppressing overexposure while preserving key details, resulting in a 19.8 percentage point improvement in accuracy compared to traditional methods for welding defect detection.
[0147] Further detailed verification of the system's power consumption characteristics was conducted, and an energy consumption optimization model based on a 4nm process NPU was constructed. At the hardware implementation level, the high-performance computing unit integrated into the NVIDIA Jetson AGX Orin platform was adopted, and precise power consumption control was achieved through DVFS (Dynamic Voltage and Frequency Scaling) power management technology. Real-time power consumption curves were collected from 720p to 4K resolutions during testing. Data shows that when running the dynamic adaptive anchor frame generation algorithm at 1920×1080 resolution, the system's average power consumption was 6.8W ± 0.3W, with peak power consumption not exceeding 8W. This excellent performance is mainly attributed to three key design features:
[0148] (1) First, the 4nm NPU dedicated instruction set has been hardware-level optimized for matrix multiplication and attention calculation, enabling the Transformer encoder to achieve a computational density of 312 GOPS / W; (2) Secondly, the dynamic anchor box generation algorithm optimizes the feature aggregation process through a spatial pyramid pooling strategy, reducing the computational complexity from the traditional O( ) decreased to O( This effectively curbs the non-linear growth of computation time when processing high-resolution images, ensuring the system's real-time response capability under large input sizes.
[0149] (3) Finally, the intelligent task scheduler dynamically adjusts the activation module of the processing pipeline according to the inter-frame difference, which can save 23.7% of computing power during the stable period of the scene.
[0150] The power consumption verification experiment set up three typical operating modes for comparative testing, as follows: (1) In performance mode, the NPU runs at a main frequency of 1.8GHz, and the power consumption is 9.2W when processing 4K video streams; (2) In balanced mode, by dynamically downclocking to 1.2GHz and combining INT8 quantization, the power consumption is reduced to 6.5W while maintaining a processing speed of 32FPS; (3) In energy efficiency mode, the inter-frame differential detection algorithm is activated, and full-resolution processing is performed only on the motion area, so that the average power consumption is further reduced to 5.3W.
[0151] Specifically, the hardware accelerator for the dynamic anchor frame generation algorithm adopts a dataflow architecture design, reducing the energy consumption per unit anchor frame generation from 2.1 mJ to 0.7 mJ by processing 32 anchor frame candidate regions in parallel. Experimental data shows that when the system continuously processes 1080p video at 30 FPS, the NPU core temperature remains stable at [temperature value missing]. This verified the ability to operate sustainably under an 8W power consumption constraint.
[0152] This paper further elaborates on the failure mode analysis mechanism of this system, focusing on the mathematical modeling process of the three-level early warning system and its performance in actual deployment. The core of the system uncertainty propagation model lies in constructing dynamic monitoring equations:
[0153] (17) In equation (17), This represents the uncertainty estimate (variance) of the model's prediction results. Number of Monte Carlo samples; Indicates the first k Predicted output from sub-Monte Carlo sampling; This is for predicting the mean.
[0154] The physical significance of this dynamic monitoring equation lies in quantifying the differences in the output of multiple variants of the model for the same input. Its calculation process is implemented through the parallel computing unit of the Jetson AGX Orin platform, and a single evaluation takes only 1.2ms.
[0155] By establishing a joint probability distribution model of uncertainty propagation and gradient mutation, the system determines when... A level-two warning is triggered when the model predicts an anomaly probability exceeding 92.3%.
[0156] Gradient mutation As shown in formula (18): (18) In equation (18), for Gradient at time; for Gradient at time.
[0157] The Monte Carlo simulation verification step involved performing 10,000 random samples on the NuScenes-Night test set. The results showed that... The threshold setting enables the system to issue an early warning an average of 2.7 frames (90ms@30fps) before a real anomaly occurs, keeping the false alarm rate below 1.2%.
[0158] The system behavior after the emergency response unit is activated follows a strict time sequence, as follows: The MobileNetV3 lightweight model is loaded within the first frame after triggering (takes 8.3ms), the sensor parameters are reset in the second frame (including ISO being reduced to 6400 and exposure time being adjusted to 30ms), and the output of stable enhanced results is completed from the third frame onwards.
[0159] Actual test data shows that the average time from warning triggering to complete system recovery is 4.2 frames (140ms), with 95% of cases recovering within 5 frames. Table 1 below shows the recovery time distribution under different failure scenarios.
[0160] Table 1. Distribution of recovery time under different failure scenarios The system recovery process employs a progressive weighted hybrid strategy, mathematically expressed as follows: (19) In equation (19), The mixing coefficient, The linear increment from 0 to 1 takes 5 frames to ensure visual continuity; The final fused image output by the system; Output images for emergency backup models; Output image for the restored master augmentation model.
[0161] Hardware-level protection mechanisms include: upon detection of... The system immediately cuts off the overload voltage supply to the NPU, reducing the core voltage from 1.2V to 0.9V using a TI BQ25703 power management chip. Simultaneously, it activates the infrared illumination module (850nm LED, 15mW / sr) to improve input signal quality. Road test data shows that this emergency mechanism successfully intercepted 97.3% of potential failure events during 2000km of cumulative operation, maximizing system reliability.
[0162] Although the present invention has been described in detail through the preferred embodiments above, it should be understood that the above description should not be considered as a limitation of the present invention. Various modifications and substitutions to the present invention will be apparent to those skilled in the art after reading the above description. Therefore, the scope of protection of the present invention should be defined by the appended claims.
Claims
1. A low-light image quality enhancement system based on enhanced night scene modeling, characterized in that, The system includes: The sample data construction module generates enhanced low-light scene images based on the initial image dataset through a dual-channel hybrid generative network architecture, and dynamically weights the enhanced low-light scene images according to the content analysis results to construct an optimized training sample library. as well as A hybrid enhancement model is constructed based on an attention module and a probabilistic graphical model module. During the training phase, the hybrid enhancement model takes the optimized training sample library as input and uses a region-adaptive weighted loss function for parameter optimization. During the enhancement processing phase, the optimized model receives the low-light image to be processed, estimates the illumination and noise distribution of the low-light image, and generates the enhanced image. The dual-channel hybrid generative network architecture includes: a first generative adversarial network unit for scene style transfer and a second generative adversarial network unit for multi-condition generation; The attention module is a Transformer encoder, which captures long-range dependencies between image patches through a multi-head attention mechanism. The probabilistic graphical model module is a Bayesian probabilistic graphical model, which is used to establish the probability transfer function between light intensity and noise variance, and to quantify the uncertainty of the model's prediction results.
2. The low-light image quality enhancement system based on enhanced night scene modeling of claim 1, wherein, The first generative adversarial network unit uses CycleGAN generative adversarial network, and the second generative adversarial network unit uses ACGAN generative adversarial network.
3. The low-light image quality enhancement system based on enhanced night scene modeling according to claim 1 or 2, characterized in that, The system also includes at least one of the following modules: The dark area confidence estimation module evaluates the confidence of dark areas in the input low-light image, providing a basis for adjusting the weights of the loss function used by the model training module. A hardware interface module, which is used to connect to the image sensor and receive initial image data; a distribution consistency monitoring module for monitoring the Bayesian probability branch in the hybrid enhanced model in real time, calculating the divergence between the current prediction distribution and the prior distribution in real time KL divergence as an uncertainty warning indicator of the system to prevent the model from being severely distorted in an extreme unknown scenario; The output optimization module is used to perform dynamic range mapping operation on the enhanced image of the low-light image to be processed.
4. The low-light image quality enhancement system based on enhanced nighttime scene modeling according to claim 3, characterized in that, The dark area confidence estimation module uses ResNeXt-101 as its basic architecture and introduces a multi-scale feature pyramid pooling module in the fourth residual stage. Or / and, the dynamic range mapping operation adopts the Sigmoid-Curve dimming curve equation, as shown in formula (14): ; In equation (14), To output pixel values; The brightness component of the input image is used for enhancement. The maximum brightness threshold supported by the display device; The slope parameter of the dimming curve; This represents the brightness center offset. Brightness compensation reference value.
5. A low-light image quality enhancement method based on enhanced nighttime scene modeling, characterized in that, The method includes the following steps: (S100) Sample data construction: Based on an initial image dataset, enhanced low-light scene images are generated through a dual-channel hybrid generative network architecture. Based on the content analysis results of the enhanced low-light scene images, the images are dynamically weighted to construct an optimized training sample library. (S200) Model training: Using the optimized training sample library as input, the hybrid augmentation model is trained using a region-adaptive weighted loss function. The optimal weight parameters of the model are determined through offline iteration to obtain the optimized hybrid augmentation model. (S300) Image enhancement: The low-light image to be processed is input into the optimized hybrid enhancement model, which estimates the light and noise distribution based on the learned distribution rules and generates the final enhanced image; (S400) Dynamic adaptive closed-loop adjustment: Real-time monitoring of the quality index of the enhanced image and dynamic adjustment of the system's operating parameters through a feedback mechanism.
6. The low-light image quality enhancement method based on enhanced nighttime scene modeling according to claim 5, characterized in that, In step (S100), the dual-channel hybrid generative network architecture includes: a first generative adversarial network unit for scene style transfer and a second generative adversarial network unit for multi-condition generation; The first generative adversarial network (GAN) unit is configured with several layers of ResNet generators to achieve scene style transfer. This ResNet generator includes a front-end encoding module, an intermediate transformation module, and a back-end decoding module. The entire network uses the LeakyReLU activation function. The intermediate transformation module consists of several residual block groups, each containing a convolutional layer and a normalization layer. Skip connections are introduced within each residual block to preserve low-frequency features. Its cycle consistency loss function is defined as: (1) In equation (1), The cycle consistency loss function; For mathematical expectation; A generator used to map a source domain image to a target domain; A generator used to map a target domain image back to the source domain; The input source domain image sample; These are the weighting coefficients for the identity mapping loss; This is the loss term for the identity mapping; The second generative adversarial network unit is configured with several layers of DCGAN discriminators to generate multiple conditions, and a scene classification head is attached, using several one-hot encoded scene labels as conditional inputs; Or / and, in step (S100), the dynamic weight allocation strategy achieves sample optimization allocation through the dynamic weight formula (2); (2) In equation (2), For the first i Percentage of pixels in the dark area of each sample; Temperature coefficient; For the first j The percentage of dark area pixels in each sample; For the first i Attention weights for each sample; N sample Indicates the total number of samples; Or / and, in step (S100), sample validity control is achieved through three-level quality monitoring, sample distribution offset detection adopts K-means clustering algorithm, and when sample distribution offset triggers the set conditions, a stereo depth camera is used to re-collect boundary samples.
7. The low-light image quality enhancement method based on enhanced nighttime scene modeling according to claim 5, characterized in that, In step (S200), the hybrid enhancement model adopts a Bayesian-Transformer hybrid architecture, coupling a hierarchical Bayesian probabilistic graphical model through an attention-based Transformer encoder. The spatial feature extraction process of the attention mechanism and the hierarchical probabilistic graphical model formed by the Gamma prior distribution are mathematically expressed as follows: (3) in, In equation (3), the hybrid enhancement model takes the deterministic feature representation henc extracted by the Transformer encoder as input, and maps it to the posterior distribution parameters of the latent variable z through a multilayer perceptron (MLP), i.e. For and Let Gamma be the distribution of the parameter, where the shape parameter is... With rate parameter All are predicted by nonlinear transformation of the deterministic feature representation henc; Or / and, in step (S200), the hybrid enhancement model is constructed with a light intensity mapping function, as shown in formula (4): (4) In equation (4), This represents the total number of samples. For the first The sampling weights of each sample follow a normal distribution. ; Input pixel values; For bias terms; These are measured parameters; The hybrid enhancement model also constructs a probability transfer function between illumination intensity and noise variance, as shown in formula (5): (5) In equation (5), The standard deviation of noise; Light intensity; These are the sensor response parameters; This is a dynamic adjustment coefficient; The dynamic adjustment coefficient Real-time optimization is achieved through adaptive update rules, which are shown in formula (6): (6) In equation (6), For dynamic adjustment coefficients, subscript For training steps; The learning rate; This is the total loss function; The standard deviation of noise; To prevent the stability constant from having a denominator of zero; Or / and, in step (S300), the Bayesian probability branch in the hybrid enhancement model is monitored in real time by the distribution consistency monitoring module, and the KL divergence between the predicted distribution and the prior distribution of the current frame image is calculated in real time. As an early warning indicator of system uncertainty, to prevent the model from being severely distorted in extreme unknown scenarios, as shown in formula (7): (7) In equation (7), This represents the total number of steps. The prior target distribution learned by the model during the training phase; This represents the predicted distribution of the hybrid augmentation model output at the current time.
8. The low-light image quality enhancement method based on enhanced nighttime scene modeling according to claim 6, characterized in that, According to the above KL divergence Determine whether the noise pattern of the current input image deviates significantly from the knowledge range learned by the model. If it does, determine that the current enhancement result is unreliable and trigger the model self-correction process, including: (1) Parameter freezing and activation: Temporarily freeze the backbone weights of the Transformer encoder and activate only the dynamic adjustment coefficients in the hybrid enhancement model. As trainable parameters; (2) Gradient backpropagation: based on the current KL The divergence bias is calculated using the adaptive update rule defined in formula (6) to calculate the gradient. ; (3) Online parameter update: dynamically adjust the coefficients along the gradient direction Perform single-step or multi-step fine-tuning to adjust the predicted distribution. Rapidly approximate the prior target distribution Thus KL The divergence has been pulled back to a safe range.
9. The low-light image quality enhancement method based on enhanced nighttime scene modeling according to claim 5, characterized in that, In step (S200), the region-weighted loss function is designed during the optimization training phase, as shown in formula (8): (8) In equation (8), These are the pixel domain loss weight coefficients; For structural loss weights; The pixel values of the target image; To enhance the pixel values of the image; Structural feature representation of the target image; To enhance the structural feature representation of the image; Or / and, in step (S200), in the implementation of the region-weighted loss function, the confidence level of the dark area is generated, and the dynamic weight is calculated based on the dynamic weight calculation function, the mathematical expression of which is: (9) In equation (9), For the pixel domain loss, there are dynamic weighting coefficients. Confidence level for the dark area; Or / and, in step (S200), during the training phase, the parameters are updated along the gradient direction by the optimizer to improve the performance of the hybrid augmentation model; The training phase employs a three-stage learning rate scheduling, with the initial learning rate undergoing linear decay followed by cosine annealing; gradient management utilizes a layered pruning strategy, constraining the gradient norm of each layer. ,in For gradient, The L2 norm of the gradient. This is the weight matrix. The Frobenius norm is the weight.
10. The low-light image quality enhancement method based on enhanced nighttime scene modeling according to claim 5, characterized in that, In step (S400), during closed-loop regulation, the quality index is calculated. The formula is: (10) In equation (10), The total number of samples; As weight; For indicator functions; For the sample ; This is a nighttime dataset; For example The noise variance; This is the noise sensitivity coefficient; Or / and, in step (S400), in the closed-loop regulation implementation, based on the noise level With system uncertainty The correlation equations are used to establish the following dynamic system model: First, define the state variables. , T Its evolution process is described by discrete-time state equations: (11) In equation (11), This is the system matrix. Each row in the matrix represents the evolution logic of the system state variables at the next time step, and each column represents the linear contribution weight of the current state variable to the system evolution. The off-diagonal elements of the system matrix reflect... Causal chain; input matrix Corresponding ISO - Exposure Control ; for The system state vector at time t; For system x, process noise; The calculation uses hyperbolic tangent normalization: (12) In equation (12), k This is the slope adjustment coefficient of the normalization function; Uncertainty Then through three-dimensional vectors The comprehensive characterization incorporates three indicators: spatial gradient energy, temporal KL divergence, and relative intensity error. Or / and, in step (S400), in the closed-loop adjustment implementation, the transfer function block diagram of the ISO-exposure control mechanism includes a dual-pathway system of feedforward and feedback: the feedforward path is based on the real-time illumination intensity. The benchmark ISO was determined by looking up a table. The feedback path adjusts the exposure time via a PID controller. : (13) In equation (13), For error signals, ; This is the proportional gain coefficient; This is the integral gain coefficient; This is the differential gain coefficient.