Low-bit Depth Image Enhancement Method Based on Noise Shaping and Latent Space Modeling
Through the Sigma-Delta modulator and hidden space modeling method, the problem of low-bit image reconstruction performance degradation is solved, high-precision image recovery is achieved, and the image quality and robustness are improved.
Patent Information
- Application Number
- CN202510623478.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The prior art uses a dramatic degradation of reconstruction performance when processing extremely low bit quantized images, dynamic range loss and context information are insufficient, and robustness is limited by fixed offset mode and local context constraints.
The Sigma-Delta modulator is used for noise shaping quantization, and the layered feature discovery module generates low-bit depth-low resolution image pairs, and model the relationship between low-resolution images and high-bit depth images in hidden space. The adaptive weight fusion module is used to fuse multiple low-bit depth-low resolution image pairs to achieve high-precision recovery of the image.
The reconstruction quality of low-bit depth images is significantly improved. Through the synergy between noise shaping and sparse driving decoding, the redundancy and reversibility of the image are optimized. The adaptive weight fusion module dynamically adjusts the weight to restore details and global consistency, and improves the spatial and bit dimension feature interaction effect of the image.
Smart Images

Figure CN120125466B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a low-bit-depth image enhancement method based on noise shaping and latent space modeling. Background Art
[0002] With the rise of deep learning, convolutional neural network-based bit-depth enhancement methods have significantly improved reconstruction quality through data-driven strategies. The BE-CALF (Bit-depth enhancement by concatenating all level features) network uses residual learning to directly predict missing bit planes. The IRFRN (Iterative residual feature refinement network for bit-depth enhancement) incorporates information from different frequencies into residual features, but their limited bit input design limits model flexibility. Multi-scale methods such as RMFNet (Residual-guided multiscale fusion network for bit-depth enhancement) fuse multi-resolution features through channel shuffling, but the downsampling operation leads to irreversible loss of high-frequency information.
[0003] Attempts to model the problem as frequency-domain phase estimation using implicit neural representations have proven difficult to effectively separate frequency components for extremely low-bit inputs. While these methods perform well in conventional bit conversion scenarios (e.g., 8-bit to 16-bit), when processing sensor-side low-bit inputs, such as 1-bit quantized data, dynamic range loss and insufficient context lead to a sharp degradation in reconstruction performance.
[0004] The PSD (planned sensor distortion) algorithm, based on the strong correlation of neighborhood pixel values and the theory of multi-observation signal reconstruction, has opened up a new path through joint reconstruction and acquisition optimization. The classic PSD-DR (deterministic reconstruction) algorithm injects a preset offset into the ADC front-end and leverages the Markov characteristics of adjacent pixels to construct multiple descriptors. However, its reliance on a manually designed 3×3 template limits its contextual awareness. The PSD-block algorithm proposes a block-based adaptive offset strategy, while the auto-reset PSD algorithm combines the characteristics of analog-to-digital sensors to restore high dynamic range. However, both require complex hardware support. PSD-CNN, for the first time, introduces a residual dense network into the reconstruction stage. It implicitly integrates multi-scale molecular dynamics information through dynamic range-preserving mid-tread quantization, achieving significantly better results than traditional PSD. However, the robustness of existing methods to extreme quantization noise is still limited by the fixed offset pattern and local context constraints. Summary of the Invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a low-bit-depth image enhancement method based on noise shaping and latent space modeling, thereby improving the quality of image restoration.
[0006] The present invention adopts the following technical solutions to achieve the above-mentioned objectives. The present invention provides a low-bit-depth image enhancement method based on noise shaping and latent space modeling, comprising:
[0007] S1, noise shaping quantization is performed through Sigma-Delta modulator to construct low bit depth image;
[0008] S2, generating low bit depth-low resolution image pairs through a hierarchical feature discovery module;
[0009] S3, modeling the relationship between low-resolution images and high-bit-depth images in the latent space;
[0010] S4. Fusion of multiple low bit depth-low resolution image pairs via an adaptive weight fusion module.
[0011] Furthermore, step S1 specifically includes:
[0012] The Sigma-Delta modulator combines information embedding with signal processing through the synergistic mechanism of noise shaping technology and high-frequency pulse sequences, constructs low-bit-depth images with redundancy and reversibility that meet the set conditions under low-resolution and low-bit-depth conditions, encodes the original information in the latent space, and provides a potential conditional posterior distribution for the super-resolution module.
[0013] Furthermore, for low bit depth images, the quantization results and state variable matrix are updated according to the following rules:
[0014] ;
[0015] ;
[0016] Where, Indicates the output of the quantized result, represents the quantization function, represents the cumulative state variable, i represents the row index, j represents the column index, is the upper adjacent pixel, representing the state variable of row i−1 and column j, is the left adjacent pixel, representing the state variable of row i and column j−1, is the diagonally adjacent pixel, representing the state variable of the i-1th row and j-1th column, Represents the input value of the current pixel;
[0017] The quantization error satisfies , D is the first-order forward difference matrix, through high-order quantization , concentrating the noise in the high-frequency area, represents the quantization order, represents the transpose of the matrix, represents the original input image, Represents the quantitative results, represents the state variable matrix.
[0018] Furthermore, step S1 specifically includes: using TV regularization to exploit gradient sparsity in the decoding stage. For a sparse image in the column direction, the reconstructed image is obtained by solving the following optimization problem:
[0019] ;
[0020] ;
[0021] Where, represents the reconstructed image, represents the L1 norm, used to promote sparsity, represents the order of gradient sparsity, D is the first-order forward difference matrix, express Order difference matrix, used to calculate the image Step gradient, represents the quantization step size, represents the inverse matrix transformation, Z represents the image matrix to be reconstructed, Indicates the quantitative results;
[0022] For two-dimensional sparse gradients, the optimization objective is expanded to:
[0023] ;
[0024] ;
[0025] For edges that meet the minimum separation condition, additional boundary constraints are introduced to improve accuracy.
[0026] Furthermore, the Frobenius norm error of the column-wise sparse image satisfies:
[0027] ;
[0028] Where, represents the reconstructed image, represents the original input image matrix, F represents the Frobenius norm, C represents the constant factor, s represents the sparsity, and N represents the size of the image. Indicates the quantization step size;
[0029] The two-dimensional sparse gradient error bound is optimized as:
[0030] ;
[0031] For the edge that meets the minimum separation condition, the separation condition is achieved:
[0032] ;
[0033] Where M represents the minimum interval parameter, N represents the size of the image, represents the quantization order, Indicates the order of gradient sparsity.
[0034] Furthermore, step S2 specifically includes:
[0035] The hierarchical feature discovery module consists of a sampling module and a feature encoding module. The sampling module is defined as follows:
[0036] , ;
[0037] , ;
[0038] Where, represents the downsampling operation, Indicates that low bit depth images are processed i The result obtained after sampling is represents the upsampling operation, Indicates upsampling of low bit depth images j The result obtained after multiplication;
[0039] The same dimension and Connect to obtain a low bit depth-low resolution image pair, freeze the parameters of the super resolution module through the encoding module, and copy the super resolution module as a A training copy of , which takes a low bit depth-low resolution image pair as input.
[0040] Furthermore, step S3 specifically includes:
[0041] Introducing conditional posterior distribution , modeling the relationship between low-resolution and high-bit-depth images in the latent space, assuming that there are different and , in the natural state, high bit depth images are obtained by the real information of the image through the prior distribution Get the parameters that will be obtained from the training copy Adding latent space and using generative models Simulating conditional posterior distribution ; Prior distribution Refers to obtaining a low-resolution image by reducing the resolution and bit depth of a high-bit-depth image. represents a high bit depth image, represents a low-resolution image, represents the nth low-resolution image, Represents the nth high bit depth image.
[0042] Furthermore, step S4 specifically includes:
[0043] Multiple low-bit-depth-low-resolution image pairs are fused through an adaptive weight fusion module. Low-bit-depth-low-resolution image pairs contain different feature information. The adaptive weight fusion module dynamically fuses according to the amount of feature information. The weight is dynamically adjusted according to the specific features of the input image pairs. For image pairs with information loss or compression artifacts greater than the set value, the weight of the low-bit-depth-low-resolution image pair branch is increased. For image pairs with complete information or compression artifacts less than the set value, the weight of the low-resolution image branch is increased.
[0044] The input of the adaptive weight fusion module is the low-bit-depth-low-resolution image pair obtained from the latent space and the low-resolution image processed by the super-resolution module. The input is passed to the attention branch, the non-attention branch, and the fusion branch respectively. The attention branch and the non-attention branch are responsible for extracting complementary features from the input image. The attention branch focuses on details, and the non-attention branch captures global image information. The fusion branch is responsible for adaptively weighting the contributions of the attention branch and the non-attention branch and applying Softmax normalization weights.
[0045] The weighted outputs of the attention branch, the non-attention branch, and the fusion branch are combined, and the contribution of each branch is adjusted according to the calculated weights. The processed weighted output is fused with the original input image to restore the features of the high bit-depth image, and the final output image is generated by combining the weighted representations of all image pairs.
[0046] The beneficial effects of the present invention are:
[0047] The present invention first introduces a Sigma-Delta modulator to implement noise-shaped quantization. Through a first-order error feedback mechanism, the quantization noise is pushed to the high-frequency region. Low-pass filtering is then combined to retain low-frequency information, thereby generating a low-redundancy, highly reversible low-bit-depth image input. A high-bit-depth image reconstruction method based on super-resolution and bit-depth enhancement is then constructed. To better achieve cross-dimensional feature interaction between space and bits, the present invention designs a hierarchical feature discovery module that generates multi-scale low-bit-depth-low-resolution image pairs through a cascaded downsampling-upsampling structure. Starting from the initial low-bit-depth image, multi-scale features are obtained through three levels of controllable downsampling. The lowest-resolution features are then upsampled to generate a low-resolution image sequence, thereby modeling the statistical dependence of low-resolution and high-bit-depth images in the latent space. Inspired by the visual attention mechanism, the present invention proposes an adaptive weight fusion module that dynamically allocates fusion weights based on the information entropy of the feature map. For high-entropy regions with severe artifacts, the low-bit-depth image branch is enhanced to restore details; for smooth low-entropy regions, the super-resolution branch is relied upon to maintain global consistency. This mechanism is implemented through gated convolution and channel attention, and its weight distribution can be formalized as a sigmoid-activated gating function to ensure the adaptive interaction of spatial and bit-dimensional features. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a flow chart of a low-bit-depth image enhancement method based on noise shaping and latent space modeling provided by the present invention.
[0049] Figure 2 This is a flowchart of processing low bit depth images provided by the present invention. DETAILED DESCRIPTION
[0050] To make the objectives, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0051] The present invention provides a low bit depth image enhancement method based on noise shaping and latent space modeling, referring to Figure 1 and Figure 2 , the method comprising:
[0052] S1, noise shaping quantization is performed through Sigma-Delta modulator to construct low bit depth image;
[0053] The Sigma-Delta modulator combines information embedding with signal processing through a synergistic mechanism involving noise shaping techniques and high-frequency pulse sequences. This allows for the construction of low-redundancy, reversible low-bit-depth images at low resolution and low bit depth, while preserving the original details to the greatest extent possible. This information can be encoded in a latent space, providing a potential conditional posterior distribution for the super-resolution module, enabling effective restoration of the target image.
[0054] For the original input image , N represents the image size, and the quantization results and state variable matrix are updated according to the following rules:
[0055] ;
[0056] ;
[0057] Where, Indicates the output of the quantized result, represents the quantization function, represents the cumulative state variable, i represents the row index, j represents the column index, is the upper adjacent pixel, representing the state variable of row i−1 and column j, is the left adjacent pixel, representing the state variable of row i and column j−1, is the diagonally adjacent pixel, representing the state variable of the i-1th row and j-1th column, Represents the input value of the current pixel;
[0058] The quantization error satisfies , D is the first-order forward difference matrix, through high-order quantization , concentrating the noise in the high-frequency area, represents the quantization order, represents the transpose of the matrix, Represents the quantitative results, represents the state variable matrix.
[0059] In the decoding stage, TV regularization is used to exploit gradient sparsity. For a sparse image in the column direction, the reconstructed image is obtained by solving the following optimization problem:
[0060] ;
[0061] ;
[0062] Where, represents the reconstructed image, represents the L1 norm, used to promote sparsity, represents the order of gradient sparsity, D is the first-order forward difference matrix, express Order difference matrix, used to calculate the image Step gradient, represents the quantization step size, represents the inverse matrix transformation, and Z represents the image matrix to be reconstructed;
[0063] For two-dimensional sparse gradients, the optimization objective is expanded to:
[0064] ;
[0065] ;
[0066] For edges that meet the minimum separation condition, additional boundary constraints are introduced to improve accuracy.
[0067] The Frobenius norm error for column-wise sparse images satisfies:
[0068] ;
[0069] Where, Represents the reconstructed image, F represents the Frobenius norm, C represents the constant factor, and s represents the sparsity;
[0070] The two-dimensional sparse gradient error bound is optimized as:
[0071] ;
[0072] For the edge that meets the minimum separation condition, the separation condition is achieved:
[0073] ;
[0074] Where M represents the minimum interval parameter, which means the minimum interval requirement of the image edge or jump discontinuity on the discretized annulus, N represents the size of the image, represents the quantization order, Indicates the order of gradient sparsity.
[0075] The above method significantly outperforms other quantization methods through the synergy of noise shaping and sparse-driven decoding.
[0076] S2, generating low bit depth-low resolution image pairs through a hierarchical feature discovery module;
[0077] The hierarchical feature discovery module consists of a sampling module and a feature encoding module.
[0078] The sampling module first uses a step-by-step downsampling method, starting with a low-bit-depth image and downsampling it to half the original resolution three times. The downsampling is repeated three times, resulting in low-bit images with downsampling levels of 2x, 4x, and 8x, respectively. The final 8x downsampled image is used as the starting point for upsampling, and then step-by-step upsampling is continued, each time upsampling to twice the original resolution, ultimately resulting in low-resolution images with upsampling levels of 2x, 4x, and 8x.
[0079] The sampling module is defined as follows:
[0080] , ;
[0081] , ;
[0082] Where, represents the downsampling operation, Indicates that low bit depth images are processed i The result obtained after sampling is represents the upsampling operation, Indicates upsampling of low bit depth images j The result obtained after multiplication;
[0083] The same dimension and To add the low bit depth-low resolution image pair to the feature encoding module, the parameters of the super resolution module are frozen by the feature encoding module, and the super resolution module is copied as a A training copy of the pre-trained model takes a low-bit-depth, low-resolution image pair as input. When this structure is applied to hidden information extraction, the locked super-resolution module retains good super-resolution processing capabilities for general images, while the trainable copy reuses the pre-trained model to build a model that can handle more specialized downstream tasks with additional requirements.
[0084] In the latent space, the present invention utilizes Designing a neural network to train a generative model , multiple low bit depth-low resolution image pairs with different information content are output and handed over to the next step for processing.
[0085] S3, modeling the relationship between low-resolution images and high-bit-depth images in the latent space;
[0086] This paper introduces the conditional posterior distribution for the first time , modeling the relationship between low-resolution and high-bit-depth images in the latent space, assuming that there are different and , in the natural state, high bit depth images are obtained by the real information of the image through the prior distribution Get the parameters that will be obtained from the training copy Adding latent space and using generative models Simulating conditional posterior distribution ; Prior distribution Refers to obtaining a low-resolution image by reducing the resolution and bit depth of a high-bit-depth image. represents a high bit depth image, represents a low-resolution image, represents the nth low-resolution image, Represents the nth high bit depth image.
[0087] Here’s how it works:
[0088] Assume that HBD is a continuous random variable, there exists For some distribution, the parameter Under the condition of Represented as a deterministic variable , where t is independently distributed in LR and has independent function value distribution Auxiliary variables. At this time, it can be expressed as Regarding expectations for generating HBD:
[0089] ;
[0090] in, To generate the quality evaluation function of HBD, , , then the expectation can be expressed as:
[0091] ;
[0092] Discretize the integral, and we have:
[0093] ,Right now: ;
[0094] in The geometric meaning obtained by describing it in words is the generative model The expectation of HBD generation effect can be achieved by some method Approximately the true expression of t for HBD.
[0095] S4. Fusion of multiple low bit depth-low resolution image pairs via an adaptive weight fusion module.
[0096] The adaptive weighted fusion module is responsible for dynamically fusing multiple low-bit-depth and low-resolution image pairs. These pairs have been processed using a latent space variable algorithm and contain features with varying amounts of information. The adaptive weighted fusion module dynamically fuses features based on their information content, achieving efficient interaction and adaptive fusion of spatial and bit information. This mechanism aims to adjust the contribution of each image pair based on the information content in the latent space representation. To achieve this, the network learns a set of weights that are dynamically adjusted based on the specific characteristics of the input image pair, such as quality and level of detail. For image pairs with significant information loss or compression artifacts, the weight of the low-bit-depth and low-resolution image branch is increased, leveraging the information from the attention branch to help recover lost details. Conversely, for high-quality, artifact-free image pairs, the weight of the low-resolution image branch is increased, resulting in a result that is more closely aligned with the output of the non-attention weighted branch, preserving the advantages of the super-resolution algorithm as much as possible.
[0097] The input to the adaptive weight fusion module is a low-bit-depth-low-resolution image pair obtained from the latent space and a low-resolution image processed by a super-resolution algorithm. The input is passed to the attention branch, the non-attention branch, and the fusion branch, respectively. The attention and non-attention branches extract complementary features from the input image: the attention branch focuses on fine details, while the non-attention branch captures global image information. The fusion branch adaptively weights the contributions of the attention and non-attention branches and applies softmax normalization to the weights. The weighted outputs of the attention, non-attention, and fusion branches are combined, and the contribution of each branch is adjusted based on the calculated weights. Finally, the final image is refined and fused with the original input image to recover the features of the high-bit-depth image. The final output is generated by combining the weighted representations of all image pairs. This image fully utilizes the low-frequency information of the low-resolution input and the high-frequency details of the low-bit-depth input. The output processed by the adaptive weight fusion module effectively integrates the strengths of each image pair to maximize image restoration quality.
[0098] This paper rigorously validates the feasibility of Sigma-Delta modulation for high-precision quantization in LBD image processing, conducting comparative experiments against traditional memoryless scalar quantization and PSD. All three methods perform 1-bit quantization and 16-bit reconstruction, and are quantitatively evaluated using PSNR (dB) and SSIM metrics. The Sigma-Delta quantization method achieves significant performance improvements on test images, reaching 31.04 dB PSNR and 0.9013 SSIM, representing absolute improvements of 15.07 dB and 0.2363, respectively, over the MSQ method, and 8.47 dB and 0.1061, respectively, over the PSD method. This advantage stems from the inherent noise shaping mechanism of Sigma-Delta modulation, which redistributes quantization error toward high-frequency spectral components, where perceptual sensitivity is lower. Subjective image comparisons show that MSQ quantization produces images with grayscale distortion, while PSD quantization produces images with poor detail recovery and exhibits significant pseudo-contouring in smooth gradients. This method effectively suppresses these artifacts through spectral error diffusion. This visual advantage is consistent with the quantitative metrics, demonstrating the perceptual effectiveness of the noise shaping strategy.
[0099] To explore the necessity of replicating a trainable super-resolution module with parameters θ, we conducted an ablation experiment. In this experiment, the replicated SR (Super-Resolution) module was set to non-trainable mode. In two comparative experiments, the original LBD image and a low-bit-depth-low-resolution image pair were fed into the replicated SR module. This setting excluded the replicated module from training, allowing the contribution of the trainable module to be evaluated by comparing the performance of the module with and without trained parameters. In the ablation experiment, the replicated SR module received the same input as the original model, but its parameters were fixed and not learned. This modification demonstrates the importance of replicating the additional trainable SR module, especially when leveraging the SR module for downstream tasks. In the bit-depth enhancement task, without trainable parameters, the replicated SR module was unable to adapt to the characteristics of the input data, resulting in a significant performance degradation. Performance also significantly degraded when the replicated SR module was non-trainable. This demonstrates that replicating an SR module with trainable parameters is crucial for effectively recovering high-resolution features, especially when processing low-bit-depth and low-resolution images.
[0100] To demonstrate the core advantage of utilizing latent space models—the abstract transformation of the relationship between LR, LBD, and HBD images in the latent space—this paper aims to identify the conditional posterior distribution relationship between super-resolution LR images and target HBD images. It is hypothesized that in the latent space, the mapping between LR images processed by the SR module and the target HBD images follows a certain conditional posterior distribution, which can be effectively learned, thereby improving super-resolution performance. In ablation experiments, the latent space module architecture was modified to remove the entanglement between different variables. The experiments included three different setups: LR, LBD, and LR-LBD image pairs were fed into the super-resolution algorithm separately, while keeping all other model configurations unchanged. The performance of each setup was evaluated using loss curves and the subjective and objective quality of the generated HBD images. Visual performance and objective metrics were significantly lower when using LR, LBD, or LBD-LR image pairs alone than when using the latent space module. This is because the latent space abstracts and extracts intrinsic image features, which are crucial for enhancing high-frequency details and low-frequency structure. The latent space essentially simplifies the complex problem of high-resolution reconstruction. By learning a more compact representation of the underlying image structure, it can more accurately map the relationship between the LR image and the HBD image. In contrast, without the latent space module, the SR module has difficulty adapting to the complex relationships in the data, resulting in failure to recover important image details, especially in complex areas of the image.
[0101] Experimental results demonstrate the importance of the latent space module for the downstream bit-depth enhancement task of the SR module. The adaptive nature of the latent space module not only improves the quality of the SR output image but also accelerates model training and convergence when processing low-bit-depth and low-resolution inputs. These experimental results demonstrate the key role of latent space abstraction in improving performance in bit-depth enhancement tasks.
[0102] This invention adaptively assigns weights based on the amount of feature information in low-bit-depth / low-resolution image pairs, thereby improving the quality of the final super-resolution output. Specifically, the impact of using an adaptive weight fusion module was tested by comparing it to a baseline experiment, in which all branches were weighted equally. In this experiment, the present invention removed the dynamic weight allocation mechanism in the adaptive weight fusion module and set the weights of all branches (attention branch, non-attention branch, and fusion branch) to be equal. This configuration ensures that each branch contributes equally to the final output, thereby objectively demonstrating the effectiveness of the adaptive weight allocation mechanism. The model with equal weights for all branches performed significantly worse than the version using adaptive weights. The model with equal weights had lower PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure), indicating poorer recovery of high-resolution features. In contrast, the model using adaptive weights demonstrated faster convergence and lower final loss, indicating that dynamically adjusting weights contributes to a more accurate super-resolution process. Experimental results show that the adaptive weight allocation mechanism of the adaptive weight fusion module significantly improves the maximum image quality of SR and improves the overall restoration performance, thereby improving the BDE (Bit-Depth Enhancement) performance.
[0103] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A low-bit-depth image enhancement method based on noise shaping and latent space modeling, characterized in that: include: S1, noise shaping quantization is performed through Sigma-Delta modulator to construct low bit depth image; S2, generating low bit depth-low resolution image pairs through a hierarchical feature discovery module; The hierarchical feature discovery module consists of a sampling module and a feature encoding module. The sampling module is defined as follows: , ; , ; Where, represents the downsampling operation, Indicates the low bit depth image t i The result obtained after sampling is represents the upsampling operation, Indicates upsampling of the low bit depth image t j The result obtained after multiplication; The same dimension and Connect to obtain a low bit depth-low resolution image pair, freeze the parameters of the super resolution module through the feature encoding module, and copy the super resolution module as a A training copy of that takes a low bit depth-low resolution image pair as input; S3, modeling the relationship between low-resolution images and high-bit-depth images in the latent space; Introducing conditional posterior distribution , modeling the relationship between low-resolution and high-bit-depth images in the latent space, assuming that there are different and , in the natural state, high bit depth images are obtained by the real information of the image through the prior distribution Get the parameters that will be obtained from the training copy Adding latent space and using generative models Simulating conditional posterior distribution ; Prior distribution Refers to obtaining a low-resolution image by reducing the resolution and bit depth of a high-bit-depth image. represents a high bit depth image, represents a low-resolution image, represents the nth low-resolution image, represents the nth high bit depth image; S4, fusing multiple low bit depth-low resolution image pairs through an adaptive weight fusion module; Multiple low-bit-depth-low-resolution image pairs are fused through an adaptive weight fusion module. Low-bit-depth-low-resolution image pairs contain different feature information. The adaptive weight fusion module dynamically fuses according to the amount of feature information. The weight is dynamically adjusted according to the specific features of the input image pairs. For image pairs with information loss or compression artifacts greater than the set value, the weight of the low-bit-depth-low-resolution image pair branch is increased. For image pairs with complete information or compression artifacts less than the set value, the weight of the low-resolution image branch is increased. The input of the adaptive weight fusion module is the low-bit-depth-low-resolution image pair obtained from the latent space and the low-resolution image processed by the super-resolution module. The input is passed to the attention branch, the non-attention branch, and the fusion branch respectively. The attention branch and the non-attention branch are responsible for extracting complementary features from the input image. The attention branch focuses on details, and the non-attention branch captures global image information. The fusion branch is responsible for adaptively weighting the contributions of the attention branch and the non-attention branch and applying Softmax normalization weights. The weighted outputs of the attention branch, the non-attention branch, and the fusion branch are combined, and the contribution of each branch is adjusted according to the calculated weights. The processed weighted output is fused with the original input image to restore the features of the high bit-depth image, and the final output image is generated by combining the weighted representations of all image pairs.
2. The low-bit-depth image enhancement method based on noise shaping and latent space modeling according to claim 1, characterized in that: Step S1 specifically includes: The Sigma-Delta modulator combines information embedding with signal processing through the synergistic mechanism of noise shaping technology and high-frequency pulse sequences, constructs low-bit-depth images with redundancy and reversibility that meet the set conditions under low-resolution and low-bit-depth conditions, encodes the original information in the latent space, and provides a potential conditional posterior distribution for the super-resolution module.
3. The low-bit-depth image enhancement method based on noise shaping and latent space modeling according to claim 1, characterized in that: Step S1 specifically also includes: In the decoding stage, TV regularization is used to exploit gradient sparsity. For a sparse image in the column direction, the reconstructed image is obtained by solving the following optimization problem: ; ; Where, represents the reconstructed image, represents the L1 norm, used to promote sparsity, represents the order of gradient sparsity, D is the first-order forward difference matrix, express Order difference matrix, used to calculate the image Step gradient, represents the quantization step size, represents the inverse matrix transformation, Z represents the image matrix to be reconstructed, Represents the quantitative results, Represents the transpose of a matrix; For two-dimensional sparse gradients, the optimization objective is expanded to: ; ; For edges that meet the minimum separation condition, additional boundary constraints are introduced to improve accuracy.
4. The low-bit-depth image enhancement method based on noise shaping and latent space modeling according to claim 3, characterized in that: The Frobenius norm error for column-wise sparse images satisfies: ; Where, represents the reconstructed image, represents the original input image, F represents the Frobenius norm, C represents the constant factor, s represents the sparsity, and N represents the size of the image. Indicates the quantization step size; The two-dimensional sparse gradient error bound is optimized as: ; For the edge that meets the minimum separation condition, the separation condition is achieved: ; Where M represents the minimum interval parameter, N represents the size of the image, represents the quantization order, Indicates the order of gradient sparsity.
5. The low-bit-depth image enhancement method based on noise shaping and latent space modeling according to claim 1, characterized in that: For low bit depth images, the quantization results and state variable matrix are updated according to the following rules: ; ; Where, Indicates the output of the quantized result, represents the quantization function, represents the cumulative state variable, i represents the row index, j represents the column index, is the upper adjacent pixel, representing the state variable of row i−1 and column j, is the left adjacent pixel, representing the state variable of row i and column j−1, is the diagonally adjacent pixel, representing the state variable of the i-1th row and j-1th column, Represents the input value of the current pixel; The quantization error satisfies , D is the first-order forward difference matrix, through high-order quantization , concentrating the noise in the high-frequency area, represents the quantization order, represents the transpose of the matrix, represents the original input image, Represents the quantitative results, represents the state variable matrix.
Citation Information
Patent Citations
Pig image BDE reconstruction system and method based on differential image rate filtering
CN119991857A