Intelligent image compression encoder based on conditional reversible neural network
Through an intelligent image compression encoder based on conditional reversible neural network, the multi-expanded channel refiner, an expanded residual attention module and a reversible multi-frequency fusion network are used to solve the performance bottleneck of the image compression algorithm in the prior art in complex texture and low-code rate scenarios, and efficient image compression and high-fidelity reconstruction are achieved.
Patent Information
- Application Number
- CN202510444011.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-08
AI Technical Summary
Existing image compression algorithms have block effects, high-frequency details loss, and prominent contradictions between complexity and compression performance when dealing with complex textures or low-code rate scenarios. It is difficult for existing end-to-end models to effectively capture the correlation between local details and global structures of images across scales, resulting in poor texture continuity and edge jagging.
An intelligent image compression encoder based on conditional reversible neural network is adopted, including an image enhancement module, a reversible multi-frequency fusion network and an entropy model. Through the multi-expanded channel refinement module, an expanded residual attention module, a reversible multi-frequency fusion network and entropy coding technology, multi-level nonlinear mapping and multi-scale feature extraction are realized, and combined with wavelet downsampling and entropy coding, feature representation and reconstruction quality are optimized.
Maintaining high-fidelity reconstruction at low bit rate breaks through the performance bottlenecks of traditional solutions, solving the bottlenecks in bit rate-distortion balance, high-frequency detail retention and computing efficiency, and achieving more efficient image compression and reconstruction quality.
Smart Images

Figure CN120281925A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to end-to-end intelligent image compression, belonging to the field of image / video compression, and particularly relates to an intelligent image compression method combining image enhancement, wavelet transform, and conditional reversible neural network. Background Art
[0002] With the continuous growth of the amount of network image and video data, the limitations of traditional compression algorithms are gradually emerging. Technologies represented by JPEG, JPEG2000, HEVC, and VVC are based on artificially designed linear transforms (such as DCT, wavelet transform). Although the compression efficiency has been continuously improved, problems such as blocking artifacts and loss of high-frequency details still exist when dealing with complex textures or low-bitrate scenarios, and the contradiction between complexity and compression performance is prominent. For example, the intra-frame coding efficiency of the new generation standard VVC (H.266) is about 50% higher than that of its predecessors, but its complexity is high and it relies on fixed transformation rules. Traditional image compression methods may use fixed receptive fields or processing at the same scale, resulting in the inability to effectively capture details under complex textures or low bitrates; existing reversible image compression models rely on single-scale feature transforms and cannot effectively capture the cross-scale local detail and global structure correlation of images, resulting in poor texture continuity (such as broken repeated patterns) and edge jagging. There is a lack of conditional coupling mechanism between the receptive fields of each scale in existing image compression models, resulting in insufficient collaborative optimization of local details and global structures. For example, in high compression ratio scenarios, low-frequency components (such as smooth regions) require large receptive fields to maintain structural consistency, while high-frequency components (such as edges) require small receptive fields for accurate positioning.
[0003] In recent years, end-to-end compression methods based on deep learning have developed rapidly. Such technologies use neural networks to automatically learn the non-linear transformation of images and combine adaptive entropy coding to compress images more efficiently. The compression quality (such as PSNR, MS-SSIM metrics) of some end-to-end models has exceeded traditional algorithms, and at the same time, the decoding speed has been significantly improved.
[0004] The end-to-end methods based on deep learning optimize non-linear transformations through neural networks. Representative technologies include the variational autoencoder (VAE) framework and the CNN / Transformer hybrid framework. Among them, the VAE model estimates the distribution of latent variables through a hyperprior network and jointly optimizes the rate-distortion trade-off. However, there is irreversible loss of high-frequency information during the encoding and decoding processes, and the edges of the decoded images are blurred. The CNN / Transformer hybrid framework (such as SwinT-Compress) combines local feature extraction and non-local attention mechanisms. Although it improves the compression efficiency of complex textures, its window partitioning strategy is difficult to capture cross-scale long-range dependencies, and the global attention calculation complexity is high, resulting in the failure of attention weights and the destruction of texture continuity at low bitrates. Summary of the Invention
[0005] To solve the above technical problems existing in the prior art, the present invention provides an intelligent image compression encoder based on a conditional reversible neural network, including an image enhancement module, a reversible multi-frequency fusion network RMFFN, and an entropy model, characterized in that: the image enhancement module includes a multi-dilated channel refinement module MDCR and a dilated residual attention module RDAM; the reversible multi-frequency fusion network RMFFN includes wavelet downsampling, 1x1 reversible convolution, and 4 hierarchical reversible blocks HIB; the image enhancement module optimizes the input features and constructs a conditional reversible neural network using a flow model, that is, realizes multi-level non-linear mapping by stacking 4 reversible multi-frequency fusion networks; the hyperprior codec of the entropy model extracts latent variable features by stacking multi-scale residual attention blocks and introduces a discrete Gaussian mixture likelihood model in the entropy coding stage.
[0006] Furthermore, the image enhancement module first maps the original RGB image to a high-dimensional feature space through a dense residual block, extracts local texture features with multiple receptive fields using a densely connected multi-layer convolutional network, and forms an initial feature basis; subsequently, the RDAM module takes the hybrid dilated convolution as the core, combines residual skip connections and channel attention mechanisms, and enhances the global context modeling ability while maintaining the spatial resolution; next, the MDCR module performs parallel multi-scale processing on the feature channels through grouped dilated convolution and reorganizes the channel information across branches to fuse semantic features of different scales in a lightweight manner; after further deep optimization by the dense residual module, the final receptive field dense block reconstructs the high-dimensional features into a 3-channel RGB image.
[0007] Furthermore, the RDAM module first performs multi-level dilated convolution feature extraction. The input feature map X passes through convolutional layers with different dilation rates to expand the receptive field while maintaining the resolution, capture local details and global context, and then cascades the inverse dilation rate convolution to refine the features; then, fuses the features at all levels, applies channel attention calibration to dynamically adjust the channel weights; finally, fuses the initial input x with the attention calibration result.
[0008] Furthermore, the MDCR realizes image enhancement through cascaded multi-scale feature operations: after the feature map output by the RDAM module is input, it is first evenly divided into 4 sub-tensors along the channel dimension and respectively input into grouped convolution branches with dilation rates of 1, 6, 12, and 18 for parallel processing; subsequently, a cross-branch channel shuffling strategy is adopted to dynamically align the features of adjacent branches, and the multi-scale features are adaptively fused by combining learnable gating weights; the recombined features are compressed for redundancy through pointwise convolution and then concatenated along the channel dimension, and finally the cross-scale information is integrated through 1×1 convolution and the channel dimension is restored.
[0009] Further, the input feature map is normalized and then input into the reversible multi-frequency fusion network, followed by a downsampling operation based on the discrete wavelet transform (DWT); DWT decomposes the input feature into four orthogonal sub-band components, which can be expressed by the following formula:
[0010]
[0011] where LL (k) is the low-frequency sub-band (approximate component) at the k-th level, LH (k) , HL (k) , HH (k) are the vertical, horizontal, and diagonal detail components of the high-frequency sub-bands;
[0012] At the decoder end, the sub-bands are fused level by level through IDWT to restore the original resolution, which can be expressed as:
[0013]
[0014] Then, frequency approximation guidance is performed through a reversible 1×1 convolution. The weight matrix W is guaranteed to be reversible through PLU decomposition, which is expressed as:
[0015] Y = W·X, X = W -1 ·Y#(7)
[0016] W = P·L·U#(8)
[0017] where P is a permutation matrix, and L and U are lower / upper triangular matrices;
[0018] Finally, the image after wavelet downsampling is input into the HIB. After decomposing the four components of different frequencies, an affine transformation is performed and the frequency features are fused. Each layer in the decoder end network can be restored through the same inverse transformation of the HIB.
[0019] Further, the HIB evenly divides the features after wavelet downsampling and 1x1 reversible convolution according to the number of channels, obtaining four groups of frequency approximation components: low-frequency approximation, horizontal approximate high-frequency, vertical approximate high-frequency, and diagonal approximate high-frequency. The mathematical expression is:
[0020]
[0021] Among them, the low-frequency dominant component is mainly the LL sub-band, mixed with a small amount of high-frequency information; the high-frequency dominant component: is mainly the LH / HL / HH sub-bands, mixed with information from other frequency bands, and the frequency band energy distribution is still mainly in the original decomposition direction;
[0022] The high-frequency sub-bands are combined into The output is generated through an affine transformation, which is expressed by the mathematical formula:
[0023]
[0024]
[0025] where σ(·) is the activation function, and G i , H i are the displacement factor and the scaling factor generation module, respectively.
[0026] Furthermore, in the hyperprior encoding and decoding process of the entropy model, the input image is gradually downsampled through residual blocks and stride convolutions to reduce the input resolution to 1 / 16, generating a compact latent representation; the quantized latent representation is gradually upsampled through residual blocks and subpixel convolutions to restore the original resolution and reconstruct the image; in the entropy encoding stage, the discrete Gaussian mixture likelihood is used to model the probability of the latent variable as a mixture of K Gaussian distributions, which can be expressed as:
[0027]
[0028] where π i,k is the mixing weight, dynamically generated by the hyperprior network, and the number of Gaussian mixtures K is set to 3. The joint encoding is performed by combining the context model to measure the current pixel distribution.
[0029] The MDCR module of the present invention enhances the feature representation through channel splitting and recombination; the RDAM module transmits low-frequency information based on residual jumps and dynamically weights high-frequency details through channel attention; the image enhancement module reduces the redundancy of the input features, making it more focused on key information and providing a more compact input for the subsequent reversible network. The enhanced image features enter the reversible multi-frequency fusion network, and the image resolution is reduced through wavelet downsampling. The HIB decomposes the image into multi-level subbands through channel splitting, and the high-frequency approximation and low-frequency approximation subbands are coupled hierarchically for compression. Subsequently, in the decoder stage, the compressed low-frequency features are fused with the high-frequency subbands, and the resolution is restored through wavelet upsampling. Through modular design and reversible-multi-frequency collaborative optimization, a breakthrough balance is achieved among compression efficiency, reconstruction quality, and computational resource consumption.
[0030] The present invention still maintains high-fidelity reconstruction at low bitrates, breaking through the performance bottleneck of traditional schemes; it solves the bottlenecks of traditional methods in bitrate-distortion balance, high-frequency detail retention, and computational efficiency, providing theoretical and technical support for the next-generation intelligent image compression standard. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a schematic diagram of the intelligent image compression encoder framework based on the conditional reversible neural network of the present invention;
[0032] Figure 2 is a block diagram of the image compression encoding and decoding system;
[0033] Figure 3 is the working principle diagram of the image enhancement module
[0034] Figure 4 is the working principle diagram of the MDCR module;
[0035] Figure 5 is the working principle diagram of the RDAM module;
[0036] Figure 6 is the working principle diagram of the HIB conditional reversible block;
[0037] Figure 7 is the working principle diagram of the scaling / displacement factor generator;
[0038] Figure 8 is the block diagram of the mixture of Gaussian hyperprior network system. Specific implementation manners
[0039] The present invention will be further described below with reference to the accompanying drawings.
[0040] End-to-end image compression jointly optimizes the encoding, quantization, entropy encoding, and decoding modules through a deep neural network, breaking through the limitations of traditional step-by-step optimization. Among them, in the encoding stage, the input image is converted into a compact latent representation to remove spatial redundancy. Usually, a convolutional neural network (CNN) or an autoencoder structure is used to extract features through multiple layers of convolution, downsampling, and non-linear activation functions. In the quantization stage, the continuous feature y is discretized into integers To avoid the problem of non-differentiable gradients, during training, additive uniform noise is used to simulate the quantization effect, which is expressed as:
[0041]
[0042] Hyperprior variable Its statistical characteristics are modeled through a hyperencoder, which is expressed as:
[0043]
[0044] The hyperdecoder generates conditional parameters to enhance the flexibility of probability modeling. Usually, it is assumed that the latent variable obeys a Gaussian distribution, and its probability density function is:
[0045]
[0046] where μ i and is a conditional parameter. At the same time, a parallel context model is used to predict the current pixel distribution using the decoded adjacent pixels. Although the surrounding elements have been used as the input of the context model, the parametric distribution cannot make good use of the context information and additional bits, which may be limited by the fixed shape of a single Gaussian model. Therefore, a more flexible Gaussian mixture model is considered later.
[0047] The decoding module includes an entropy decoding step that analyzes the bitstream according to the probability model at the encoding end to recover the discrete latent variables The inverse quantization stage converts the discrete values into continuous values y ′ , and the non-linear synthesis transformation uses deconvolution, upsampling and inverse activation functions to reconstruct the image
[0048] As Figure 1 and Figure 2 shown, the intelligent image compression encoder based on the conditional reversible neural network of the present invention includes an image enhancement module, a reversible multi-frequency fusion network (RMFFN) and an entropy coding model. The image enhancement module includes a multi-dilated channel refinement module (MDCR) and a dilated residual attention module (RDAM). The reversible multi-frequency fusion network (RMFFN) includes wavelet downsampling, 1x1 reversible convolution and 4 hierarchical reversible blocks (HIB). The main encoder and decoder combine the image enhancement module to optimize the input features, and use the flow model to construct a conditional reversible neural network, that is, a multi-level non-linear mapping is realized by stacking 4 reversible multi-frequency fusion networks to enhance the analysis and synthesis transformation capabilities; the hyperprior encoder-decoder extracts the latent variable features by stacking multi-scale residual attention blocks, and introduces a discrete Gaussian mixture likelihood model in the entropy coding stage, and combines context autoregressive modeling to improve the probability estimation accuracy, with both high fidelity and computational efficiency advantages.
[0049] Image enhancement module
[0050] As Figure 3As shown, as the core unit of end-to-end image preprocessing, through cascaded multi-scale feature operations and residual learning mechanisms, efficient enhancement and detail restoration of the input image are achieved. This module first maps the original RGB image (3 channels) to a high-dimensional feature space through dense residual blocks, and uses a densely connected multi-layer convolutional network to extract local texture features with multiple receptive fields to form an initial feature base; subsequently, the RDAM module takes hybrid dilated convolution as the core, combines residual skip connections and channel attention mechanism (CAB), enhances the global context modeling ability while maintaining the spatial resolution, and effectively distinguishes noise from effective signals; the subsequent MDCR module performs parallel multi-scale processing on the feature channels through grouped dilated convolution, and reorganizes the channel information across branches to fuse semantic features of different scales in a lightweight manner; after being deeply optimized by the dense residual module again, the final receptive field dense block reconstructs the high-dimensional features into a 3-channel RGB image, suppresses the amplification of high-frequency noise through residual connections, and retains the low-frequency authenticity of the input image.
[0051] Residual Dilated Attention Module (RDAM)
[0052] As Figure 5 shown, first, multi-level dilated convolution feature extraction is performed. The input feature map X passes through convolutional layers with different dilation rates (1, 2, 3, 4), expands the receptive field while maintaining the resolution, captures local details and global context, and continues to cascade convolutional layers with reverse dilation rates (3, 2, 1) to refine the features. Dilated convolution can expand the receptive field by increasing the gap between convolutional kernels. Let the input feature map be X, the convolutional kernel weight be ω, and the dilation rate be r. Then, each position (i, j) of the output feature map Y is calculated as follows:
[0053]
[0054] where denotes rounding down. When r = 1, the dilated convolution degenerates into a standard convolution. The size of the equivalent convolutional kernel after dilation is k d = k + (k - 1)(r - 1), and the receptive field area is F = k d × k d .
[0055] Then, the features at each level are fused, and then channel attention calibration is applied to dynamically adjust the channel weights. Finally, the initial input x is fused with the attention calibration result. The feature advantage of this module is that it covers multi-granularity information from local details to global semantics through the cascading of alternately increasing-decreasing dilation rate convolutional layers, optimizing the multi-scale feature representation. Multi-path residual connections are introduced to alleviate gradient disappearance and retain high-frequency details. The channel attention module (CAB) generates channel weights through global average pooling and fully connected layers to suppress noise channels.
[0056] Multi-Diffusion Channel Refiner Module (MDCR)
[0057] like Figure 4 As shown in the figure, the multi-dilated channel refiner module (MDCR) achieves image enhancement through cascaded multi-scale feature operations: after the feature map output by the dilated residual attention module (RDAM) is input, it is first divided into 4 sub-tensors along the channel dimension, and input into the grouped convolution branches with dilation rates of 1, 6, 12, and 18 for parallel processing. Among them, the 3×3 convolution with dilation rate 1 focuses on local texture details, the branches with dilation rates 6 and 12 capture medium and long-distance context information, and the branch with dilation rate 18 models global semantics through a 37×37 equivalent receptive field. Subsequently, a cross-branch channel shuffling strategy is used to dynamically align adjacent branch features, and multi-scale features are adaptively fused in combination with learnable gated weights to enhance the complementarity of local and global information. The reorganized features are concatenated along the channel dimension after point-by-point convolution to compress redundancy, and finally 1×1 convolution is used to integrate cross-scale information and restore the channel dimension.
[0058] Reversible Multi-Frequency Fusion Network (RMFFN)
[0059] The input feature map is normalized and then input into the reversible multi-frequency fusion network (RMFFN), followed by a downsampling operation based on discrete wavelet transform (DWT). According to the Nyquist sampling theorem, standard downsampling requires anti-aliasing low-pass filtering before downsampling. DWT decomposes the input features into four orthogonal sub-band components, which can be expressed as follows:
[0060]
[0061] Among them LL (k) is the k-th low-frequency subband (approximate component), LH (k) ,HL (k) ,HH (k) is the high frequency subband (vertical, horizontal, and diagonal detail components). After each decomposition, the resolution is reduced to The number of channels is expanded to 4C. At the decoder end, the sub-bands are fused step by step through IDWT to restore the original resolution, which can be expressed as:
[0062]
[0063] After each reconstruction, the resolution is increased to H×W×4, and the number of channels is reduced to
[0064] Next, the frequency approximation is guided by a reversible 1×1 convolution, and the weight matrix W is decomposed by PLU to ensure reversibility, which is expressed as:
[0065] Y=W·X,X=W -1 ·Y#(7)
[0066] W = P·L·U#(8)
[0067] where P is a permutation matrix, and L and U are lower / upper triangular matrices.
[0068] Finally, the image after wavelet downsampling is input into the HIB. After decomposing the four components of different frequencies, an affine transformation is performed and the frequency features are fused. Each layer in the decoder network can be restored through the inverse transformation of the same HIB. Each stage of the reversible multi-frequency fusion network contains a wavelet downsampling module, 1x1 reversible convolution, and 4 HIB modules, and a total of four identical operations are included, which together with the image preprocessing module constitute the encoder and decoder. Among them, DWT decomposes the image into different frequency bands. The low frequency retains the main structure, and the high frequency captures details to achieve targeted compression. The multi-level wavelet approximation decomposition allows optimizing the compression ratio and quality balance at multiple resolution levels. For example, higher compression is performed on the deep low-frequency subbands. The reversible convolution and coupling layer support symmetric operations of forward compression and reverse reconstruction, avoiding the problem of gradient disappearance. RMFFN realizes the efficient separation and lossless reconstruction of multi-scale features in image compression through multi-level decomposition in the wavelet domain, the strict information retention mechanism of the reversible network, and the joint optimization of high-frequency and low-frequency subbands.
[0069] Hierarchical Invertible Block (HIB)
[0070] As Figure 6 shown, the features after wavelet downsampling and 1x1 reversible convolution are evenly divided by the number of channels to obtain four groups of frequency approximation components: low-frequency approximation (LL, x1), horizontal approximation high-frequency (LH, x2), vertical approximation high-frequency (HL, x3), and diagonal approximation high-frequency (HH, x4). The mathematical expression is:
[0071]
[0072] Among them, the low-frequency dominant component is mainly in the LL subband, mixed with a small amount of high-frequency information. The high-frequency dominant components: are mainly in the LH / HL / HH subbands, mixed with information from other frequency bands, and the frequency band energy distribution is still mainly in the original decomposition direction. The high-frequency subbands are combined into The output is generated through an affine transformation, which can be expressed by the mathematical formula:
[0073]
[0074] where σ(·) is the activation function, G i , H i are the displacement factor and the scaling factor generation module respectively.
[0075] The 1x1 invertible convolution combined with band-guided channel splitting enables the low-frequency and high-frequency feature weights to be adaptive according to the input content, preserves the smooth features of the LL sub-band, directly transmits them to the deep network, focuses on the details of LH / HL / HH, and compresses redundancy through affine transformation. At the same time, the dynamic mixing of cross-band features is achieved through linear transformation. The adjusted high-frequency sub-bands (X2 - X4) are concatenated to form the joint feature X5, and the input network generates affine parameters. This design realizes cross-direction interaction and the joint modeling of horizontal, vertical, and diagonal high frequencies. The activation function is used to adaptively adjust the contribution degree of each band and suppress redundant high frequencies. Finally, a new type of affine transformation chain structure (y1→y2→y3→y4) is designed to allow the high-frequency noise to decay step by step and the low-frequency features to be enhanced gradually, and the nonlinearity is enhanced through cascaded affine transformations.
[0076] The adaptive modeling of the scaling factor and displacement factor is achieved through the multi-scale feature fusion and dynamic parameter generation mechanism. The multi-scale feature extraction uses standard convolution, dilated convolution (dilation = 2), and large kernel convolution (5×5) in parallel to capture local and global features under different receptive fields. Among them, the dilated convolution enhances the long-distance dependence modeling, and the large kernel convolution strengthens the low-frequency information extraction, providing multi-scale statistical characteristics for the dynamic factor generation. After concatenating the original features and the multi-branch fusion features, the channels are compressed through 1×1 convolution to generate the fusion features containing multi-scale information. The introduction of dilated convolution and large kernel convolution enables the network to model local details (small receptive fields) and global structures (large receptive fields) simultaneously within a single layer, and the statistical characteristics of the generated factors are more in line with the image content.
[0077] Entropy coding model
[0078] As Figure 8 shown, in the hyperprior encoding and decoding process, the input image is gradually downsampled through residual blocks and stride convolutions. This multi-level residual transformation reduces the input resolution to 1 / 16, generating a compact latent representation; the quantized latent representation is gradually upsampled through residual blocks and sub-pixel convolutions to restore the original resolution and reconstruct the image. In the entropy coding stage, the probability of the latent variable is modeled as a mixture of K Gaussian distributions, which can be expressed as:
[0079]
[0080] where π i,k is the mixing weight, dynamically generated by the hyperprior network, and the number of Gaussian mixtures K is set to 3. The joint encoding is performed by combining the context model to measure the current pixel distribution.
[0081] Through the Gaussian mixture model guided by the hyperprior network and multi-stage context modeling, the dynamic probability distribution estimation and the fusion of global-local statistical information are achieved, significantly improving the compression efficiency and reconstruction quality.
[0082] The intelligent image compression encoder based on conditional reversible neural network of the present invention, wherein MDCR optimizes feature representation through multi-scale dynamic convolution, RDAM combines dilated residual and attention mechanism to enhance high-frequency details, and RMFFN realizes multi-frequency fusion by using wavelet downsampling and hierarchical reversible blocks. In the entropy coding stage, a compact bitstream is generated through quantization and context model, and the computational efficiency and reconstruction quality are significantly improved, which is applicable to high-fidelity scenarios.
Claims
1. An intelligent image compression encoder based on a conditional reversible neural network, comprising an image enhancement module, a reversible multi-frequency fusion network RMFFN, and an entropy model, characterized in that: The image enhancement module includes a multi-dilated channel refiner module MDCR and a dilated residual attention module RDAM. The reversible multi-frequency fusion network RMFFN includes wavelet downsampling, 1x1 reversible convolution, and 4 hierarchical invertible blocks HIB. The image enhancement module optimizes the input features and constructs a conditional reversible neural network using a flow model, that is, realizes multi-level non-linear mapping by stacking 4 reversible multi-frequency fusion networks. The hyperprior encoder-decoder of the entropy model extracts latent variable features by stacking multi-scale residual attention blocks and introduces a discrete Gaussian mixture likelihood model in the entropy coding stage.
2. The intelligent image compression encoder based on the conditional reversible neural network according to claim 1, wherein: The image enhancement module first maps the original RGB image to a high-dimensional feature space through a dense residual block, and uses a densely connected multi-layer convolutional network to extract local texture features with multiple receptive fields to form an initial feature basis. Subsequently, the RDAM module takes the hybrid dilated convolution as the core, combines the residual skip connection and the channel attention mechanism to enhance the global context modeling ability while maintaining the spatial resolution. Next, the MDCR module performs parallel multi-scale processing on the feature channels through grouped dilated convolution, and reorganizes the channel information across branches to fuse semantic features of different scales in a lightweight manner. After being deeply optimized by the dense residual module again, the final receptive field dense block reconstructs the high-dimensional features into a 3-channel RGB image.
3. The intelligent image compression encoder based on a conditional reversible neural network according to claim 2, wherein: The RDAM module first performs multi-level dilated convolution feature extraction. The input feature map X passes through convolutional layers with different dilation rates to expand the receptive field while maintaining the resolution, capture local details and global context, and then cascades the inverse dilation rate convolution to refine the features. Then, the features at all levels are fused, and channel attention calibration is applied to dynamically adjust the channel weights. Finally, the initial input x is fused with the attention calibration result.
4. The intelligent image compression encoder based on the conditional reversible neural network according to claim 2, wherein: The MDCR realizes image enhancement through cascaded multi-scale feature operations: after the feature map output by the RDAM module is input, it is first evenly divided into 4 sub-tensors along the channel dimension and respectively input into the grouped convolution branches with dilation rates of 1, 6, 12, and 18 for parallel processing. Subsequently, a cross-branch channel shuffling strategy is adopted to dynamically align the features of adjacent branches, and multi-scale features are adaptively fused in combination with learnable gating weights. The recombined features are compressed for redundancy through pointwise convolution and then concatenated along the channel dimension. Finally, 1×1 convolution is used to integrate cross-scale information and restore the channel dimension.
5. The intelligent image compression encoder based on the conditional reversible neural network according to claim 1, wherein: The input feature map is input into the reversible multi-frequency fusion network after being normalized, and then a downsampling operation based on the discrete wavelet transform (DWT) is performed. DWT decomposes the input features into four orthogonal sub-band components, which can be expressed by the following formula: where LL (k) is the k-th low-frequency subband (approximate component), LH (k) , HL (k) , HH (k) are the vertical, horizontal, and diagonal detail components of the high-frequency subbands; At the decoder end, the sub-bands are fused step by step through IDWT to restore the original resolution, which can be expressed as: Next, frequency approximation guidance is performed through reversible 1×1 convolution, and the weight matrix W is guaranteed to be reversible through PLU decomposition, which is expressed as: Y = W·X, X = W -1 ·Y#(7) W = P·L·U#(8) where P is a permutation matrix, and L and U are lower / upper triangular matrices. Finally, the image after wavelet downsampling is input into the HIB. After decomposing the four components of different frequencies, an affine transformation is performed and the frequency features are fused. Each layer in the decoder-side network can be restored through the inverse transformation of the same HIB.
6. The intelligent image compression encoder based on the conditional reversible neural network according to claim 1, wherein: The HIB evenly divides the features after wavelet downsampling and 1x1 invertible convolution according to the number of channels, obtaining four groups of frequency approximate components: low-frequency approximation, horizontal approximate high-frequency, vertical approximate high-frequency, and diagonal approximate high-frequency. The mathematical expression is: Among them, the low-frequency dominant component is mainly the LL sub-band, mixing a small amount of high-frequency information; the high-frequency dominant component: mainly the LH / HL / HH sub-bands, mixing information of other frequency bands, and the frequency band energy distribution is still mainly in the original decomposition direction; High-frequency subbands are combined into An output is generated through an affine transformation, which is expressed by the mathematical formula as follows: where σ(·) is the activation function, G i , H i are the displacement factor and the scaling factor generation module respectively.
7. The intelligent image compression encoder based on conditional reversible neural network according to claim 1, characterized in that: During the hyperprior encoding and decoding process, the entropy model gradually downsamples the input image through residual blocks and stride convolutions, reducing the input resolution to 1 / 16 to generate a compact latent representation; and gradually upsamples the quantized latent representation through residual blocks and sub-pixel convolutions to restore the original resolution and reconstruct the image; in the entropy coding stage, the discrete Gaussian mixture likelihood is used to model the probability of the latent variable as a mixture of K Gaussian distributions, which can be expressed as: where π i,k is the mixing weight, dynamically generated by the hyperprior network, and the number K of Gaussian mixtures is set to 3. The current pixel distribution is jointly encoded in combination with the context model.
Citation Information
Cited By
Facial expression recognition method, system and equipment based on space channel convolution and enhanced compression incentive attention
CN120452047A
Super capacitor-lithium battery hybrid energy storage life optimization method based on variational Bayesian inference
CN121279125A
Video time-varying sampling method and device based on reversible network, and medium
CN121284242A
Multi-scale semantic and edge prior fused extensible image coding system and method
CN121924261A