End-to-End Image Compression Method and Device Based on Vision-Mamba CNN
Through the end-to-end image compression method based on Vision-Mamba CNN, the block effect, high computational complexity, and insufficient balance between bit rate and mass in the prior art are solved, and efficient image compression and reconstruction on the mobile terminal are realized.
Patent Information
- Application Number
- CN202510562636.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing image compression technology has problems such as block effect and detail loss, high computational complexity, difficulty in deploying on the mobile side, and insufficient optimization of compression code rate and reconstruction quality balance.
Using the end-to-end image compression method based on Vision-Mamba CNN, a dynamic parameterized Gaussian distribution modeling framework is constructed by introducing the fusion structure of the Vision-Mamba module and the CNN module, combining the wavelet convolution layer and channel context model, to achieve multi-scale feature information fusion and refined balance of compression code rate and reconstruction quality.
It significantly improves the long-distance dependency capture capability, reduces computing complexity, reduces high-frequency details loss and structural distortion, meets the mobile deployment needs, and achieves a refined balance between compression code rate and reconstruction quality.
Smart Images

Figure CN120088349B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to an end-to-end image compression method and device based on Vision-Mamba CNN. Background Art
[0002] With the rapid development of multimedia technology, the storage and transmission requirements of image data have increased exponentially. Traditional image compression technologies (such as JPEG, HEVC, etc.) are mainly based on discrete cosine transform (DCT) and entropy coding, and obvious block effects and detail losses will occur at high compression rates.
[0003] In recent years, deep learning-based image compression methods (such as VAE, GAN, Transformer, etc.) have significantly improved compression performance through end-to-end learning, but still face the following technical challenges:
[0004] (1) Traditional convolutional neural networks (CNNs) are limited by local receptive fields and are difficult to effectively model global dependencies in images, resulting in the loss of high-frequency details and structural distortion;
[0005] (2) Transformer-based models capture long-range dependencies through self-attention mechanisms, but have a large number of parameters and high computational complexity (O(N²)), making it difficult to be deployed on mobile devices; existing methods usually only utilize single-scale features and lack the fusion of multi-scale context information, affecting compression efficiency and reconstruction quality;
[0006] (3) Traditional hyperprior networks (such as Variational-Autoencoder) do not fully consider the dynamic changes of feature distributions when generating side information, resulting in insufficient balance optimization between compression bitrate and reconstruction quality. Summary of the Invention
[0007] The purpose of this application is to propose an end-to-end image compression method and device based on Vision-Mamba CNN for the above-mentioned technical problems.
[0008] In the first aspect, the present invention provides an end-to-end image compression method based on Vision-Mamba CNN, including the following steps:
[0009] Construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module;
[0010] Obtain the image to be compressed and input it into the trained image compression model. The image first undergoes a non-linear transformation through a non-linear transformation network to obtain the corresponding latent representation; the latent representation is input into the first quantizer to obtain the quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain the latent representation of the hyperprior transformation; the latent representation of the hyperprior transformation is input into the second quantizer to obtain the quantized latent representation of the hyperprior transformation; the quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain the corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain the second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution, and the quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain the corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image.
[0011] Preferably, the non-linear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent downsampling modules; the non-linear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent upsampling modules; the hyperprior transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent downsampling modules; the hyperprior inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent upsampling modules.
[0012] Preferably, the downsampling module includes a first convolutional layer, a first division normalization layer, a second convolutional layer, a first activation function layer, a third convolutional layer, and a second activation function layer connected in sequence; wherein, the first convolutional layer, the second convolutional layer, and the third convolutional layer are all convolutional structures with a convolutional kernel size of and an output channel number of The first division normalization layer adopts a division normalization operation, and the first activation function layer and the second activation function layer both adopt activation function; the input feature of the downsampling module is input into the downsampling module and sequentially passes through the first convolutional layer, the first division normalization layer, the second convolutional layer, the first activation function layer, the third convolutional layer, and the second activation function layer. The output feature of the first division normalization layer and the second The output features of the activation function layer are added together to obtain the output features of the downsampling module, as shown in the following formula:
[0013] ;
[0014] Among them, represents the input features of the downsampling module, represents the function corresponding to the downsampling module, represents the output features of the downsampling module, represents a convolution structure with a convolution kernel size of , and an output channel number of corresponding function, represents the division normalization operation; represents activation function;
[0015] The upsampling module includes a transposed convolutional layer, a second division normalization layer, a fourth convolutional layer, a third activation function layer, a fifth convolutional layer, and a fourth activation function layer connected in sequence; among them, the transposed convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are all convolution structures with a convolution kernel size of , and an output channel number of ; the second division normalization layer uses the division normalization operation, and the third activation function layer and the fourth activation function layer both use activation function; the input features of the upsampling module are input into the upsampling module and sequentially pass through the transposed convolutional layer, the second division normalization layer, the fourth convolutional layer, the third activation function layer, the fifth convolutional layer, and the fourth activation function layer. The output features of the second division normalization layer are added to the output features of the fourth activation function layer to obtain the output features of the upsampling module, as shown in the following formula:
[0016] ;
[0017] Among them, represents the input features of the upsampling module, represents the function corresponding to the upsampling module, represents the output features of the upsampling module, represents a convolution structure with a convolution kernel size of , and an output channel number of corresponding function of the transposed convolutional layer.
[0018] Preferably, the Vision-Mamba CNN module includes a sixth convolutional layer, a segmentation layer, a Vision-Mamba module, a CNN module, a first concatenation layer, and a seventh convolutional layer; the input features of the Vision-Mamba CNN module are input into the Vision-Mamba CNN module, and sequentially pass through the sixth convolutional layer and the segmentation layer to obtain a first intermediate feature and a second intermediate feature. The first intermediate feature and the second intermediate feature are respectively input into the Vision-Mamba module and the CNN module to obtain global features and local features. The global features and the local features are input into the first concatenation layer for concatenation to obtain a first concatenated feature. The first concatenated feature is input into the seventh convolutional layer to obtain the output features of the seventh convolutional layer. The output features of the seventh convolutional layer are added to the input features of the Vision-Mamba CNN module to obtain the output features of the Vision-Mamba CNN module; the CNN module includes an eighth convolutional layer, a fifth activation function layer, a ninth convolutional layer, and a sixth activation function layer. The input features of the CNN module are input into the CNN module and sequentially pass through the eighth convolutional layer, the fifth activation function layer, the ninth convolutional layer, and the sixth activation function layer. The output features of the sixth activation function layer are added to the input features of the CNN module to obtain local features; the convolutional kernel sizes of the sixth convolutional layer and the seventh convolutional layer are both ; the convolutional kernel sizes of the eighth convolutional layer and the ninth convolutional layer are both .
[0019] Preferably, the context module includes a second concatenation layer, a wavelet convolutional layer, a channel context model, and an entropy parameter layer connected in sequence. The quantized latent representation and the second latent representation features of the hyperprior transformation are input into the second concatenation layer to obtain a second concatenated feature. The second concatenated feature sequentially passes through the wavelet convolutional layer and the channel context model to obtain a third intermediate feature. The third intermediate feature is input into the entropy parameter layer to obtain a Gaussian distribution; the channel context model includes a tenth convolutional layer, a first activation function layer, an eleventh convolutional layer, a second activation function layer, and a twelfth convolutional layer; data block operations are adopted in the entropy parameter layer.
[0020] Preferably, the first quantizer and the second quantizer adopt adding a uniform distribution on the input features of the first quantizer and the input features of the second quantizer during the training stage of the image compression model During the inference phase of the image compression model, perform a rounding operation on the input features of the first quantizer and the input features of the second quantizer; the first encoder and the first decoder use arithmetic encoders, and the first decoder and the second decoder use arithmetic decoders.
[0021] In a second aspect, the present invention provides an end-to-end image compression device based on Vision-Mamba CNN, including:
[0022] A model construction module configured to construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module;
[0023] A compression module configured to obtain an image to be compressed and input it into the trained image compression model. The image first undergoes non-linear transformation through the non-linear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain a latent representation of the hyperprior transformation; the latent representation of the hyperprior transformation is input into the second quantizer to obtain a quantized latent representation of the hyperprior transformation; the quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain a second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain a corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image.
[0024] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect.
[0025] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.
[0026] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The end-to-end image compression method based on Vision-Mamba CNN proposed by the present invention introduces a fusion structure of the Vision-Mamba module and the CNN module, which significantly improves the long-distance dependence capture ability while maintaining low computational complexity, breaks through the local receptive field limitation of traditional CNNs, effectively captures the global dependence relationship of images, reduces the loss of high-frequency details and structural distortion; by replacing the traditional Transformer with Vision-Mamba CNN, the complexity of long-distance dependence modeling is reduced from O(N²) to linear complexity, significantly reducing the computational cost to meet the requirements of mobile deployment.
[0029] (2) The end-to-end image compression method based on Vision-Mamba CNN proposed by the present invention designs a context module that fuses the wavelet convolutional layer and the channel context model to achieve deep fusion of multi-scale feature information to enhance the feature representation ability; and constructs a dynamically parameterized Gaussian distribution modeling framework, combines the side information generated by the hyperprior network to adjust the potential feature distribution parameters in real time, and realizes a refined balance between the compression bitrate and the reconstruction quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0031] Figure 1 It is a schematic flow chart of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;
[0032] Figure 2 It is a schematic diagram of the image compression model of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;
[0033] Figure 3 It is a schematic diagram of the downsampling module of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;
[0034] Figure 4 It is a schematic diagram of the upsampling module of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;
[0035] Figure 5Schematic diagram of the Vision-Mamba module of the end-to-end image compression method based on Vision-Mamba CNN according to an embodiment of the present application;
[0036] Figure 6 Schematic diagram of the non-linear transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to an embodiment of the present application;
[0037] Figure 7 Schematic diagram of the hyperprior transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to an embodiment of the present application;
[0038] Figure 8 Schematic diagram of the hyperprior inverse transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to an embodiment of the present application;
[0039] Figure 9 Schematic diagram of the context module of the end-to-end image compression method based on Vision-Mamba CNN according to an embodiment of the present application;
[0040] Figure 10 Schematic diagram of the non-linear inverse transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to an embodiment of the present application;
[0041] Figure 11 Schematic diagram of the end-to-end image compression device based on Vision-Mamba CNN according to an embodiment of the present application;
[0042] Figure 12 Schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Figure 1 An end-to-end image compression method based on Vision-Mamba CNN provided by an embodiment of the present application is shown, including the following steps:
[0045] S1. Construct an image compression model and train it to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module.
[0046] In a specific embodiment, the non-linear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent downsampling modules; the non-linear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent upsampling modules; the hyperprior transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent downsampling modules; the hyperprior inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent upsampling modules.
[0047] Specifically, referring to Figure 2 , the image compression model in the embodiment of the present application is composed of a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module. In the non-linear transformation network, the combination of the downsampling module and the Vision-Mamba CNN module converts the image from the pixel space to a feature space that is more conducive to compression, removing the spatial correlation between image pixels and making the data distribution more convenient for subsequent processing. In the hyperprior transformation network, the combination of the downsampling module and the Vision-Mamba CNN module extracts the latent representation of the hyperprior transformation, which contains the statistical features and prior information of the image. These information can help perform entropy coding more effectively subsequently and improve the compression performance. In the hyperprior inverse transformation network, the combination of the upsampling module and the Vision-Mamba CNN module performs a hyperprior inverse transformation on the obtained second latent representation feature, restoring part of the information related to the statistical characteristics of the original image, and its structure is a symmetric structure with the hyperprior transformation network. In the non-linear inverse transformation network, the combination of the upsampling module and the Vision-Mamba CNN module decompresses the first latent representation feature and restores the compressed image, and its structure is a symmetric structure with the non-linear transformation network.
[0048] S2. Obtain the image to be compressed and input it into the trained image compression model. The image first undergoes a non-linear transformation through a non-linear transformation network to obtain the corresponding latent representation. The latent representation is input into the first quantizer to obtain the quantized latent representation. The latent representation is input into the hyperprior transformation network to obtain the latent representation of the hyperprior transformation. The latent representation of the hyperprior transformation is input into the second quantizer to obtain the quantized latent representation of the hyperprior transformation. The quantized latent representation of the hyperprior transformation successively passes through the second encoder and the second decoder to obtain the corresponding second latent representation feature. The second latent representation feature is input into the hyperprior inverse transformation network to obtain the second latent representation feature of the hyperprior transformation. The quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation successively passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain the corresponding first latent representation feature. The first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image.
[0049] Specifically, deploy the trained image compression model. During the inference process, input the image to be compressed into the trained image compression model. After processing, obtain the compressed image. The specific details will be described one by one later.
[0050] In a specific embodiment, the downsampling module includes a first convolutional layer, a first division normalization layer, a second convolutional layer, a first activation function layer, a third convolutional layer, and a second activation function layer connected in sequence; where the first convolutional layer, the second convolutional layer, and the third convolutional layer are all convolutional structures with a convolutional kernel size of and an output channel number of . The first division normalization layer uses a division normalization operation. The first activation function layer and the second activation function layer both use activation functions. The input feature of the downsampling module is input into the downsampling module and successively passes through the first convolutional layer, the first division normalization layer, the second convolutional layer, the first activation function layer, the third convolutional layer, and the second activation function layer. The output feature of the first division normalization layer is added to the output feature of the second activation function layer to obtain the output feature of the downsampling module, as shown in the following formula:
[0051] ;
[0052] where represents the input feature of the downsampling module, represents the corresponding function of the downsampling module, Represents the output feature of the downsampling module, Represents that the convolution kernel size is , and the number of output channels is The function corresponding to the convolution structure, Represents the division normalization operation; Represents Activation function;
[0053] The upsampling module includes a transposed convolutional layer, a second division normalization layer, a fourth convolutional layer, a third Activation function layer, a fifth convolutional layer, and a fourth Activation function layer; among them, the transposed convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are all convolutional structures with a convolution kernel size of , and the number of output channels is The second division normalization layer adopts the division normalization operation, and the third Activation function layer and the fourth Activation function layer both adopt Activation function; the input feature of the upsampling module is input into the upsampling module and passes through the transposed convolutional layer, the second division normalization layer, the fourth convolutional layer, the third Activation function layer, the fifth convolutional layer, and the fourth Activation function layer in sequence. The output feature of the second division normalization layer is added to the output feature of the fourth Activation function layer to obtain the output feature of the upsampling module, as shown in the following formula:
[0054] ;
[0055] Among them, Represents the input feature of the upsampling module, Represents the function corresponding to the upsampling module, Represents the output feature of the upsampling module, Represents that the convolution kernel size is , and the number of output channels is The function corresponding to the transposed convolutional layer.
[0056] Specifically, the downsampling module is denoted as , referring to Figure 3 , the downsampling module consists of a first convolutional layer, a first division normalization layer, a second convolutional layer, a first Activation function layer, a third convolutional layer, and a second Activation function layer, and the output feature of the first division normalization layer and the output feature of the second Activation function layer form a residual connection. Among them, the first convolutional layer adopts a convolution kernel size of , a convolutional structure with a stride of 2 and a padding of 1. Both the second convolutional layer and the third convolutional layer use a convolutional kernel size of , a convolutional structure with a stride of 1 and a padding of 1; the module above is denoted as , referring to Figure 4 , the upsampling module consists of a transposed convolutional layer, a second division normalization layer, a fourth convolutional layer, a third activation function layer, a fifth convolutional layer, and a fourth activation function layer, and the output features of the second division normalization layer and the fourth activation function layer form a residual connection. Among them, the convolutional kernel size of the transposed convolutional layer is , the stride is 2, the padding is 1, and the output padding is 1; both the fourth convolutional layer and the fifth convolutional layer use a convolutional kernel size of , a convolutional structure with a stride of 1 and a padding of 1.
[0057] In a specific embodiment, the Vision-Mamba CNN module includes a sixth convolutional layer, a segmentation layer, a Vision-Mamba module, a CNN module, a first splicing layer, and a seventh convolutional layer; the input features of the Vision-Mamba CNN module are input into the Vision-Mamba CNN module, and successively pass through the sixth convolutional layer and the segmentation layer to obtain a first intermediate feature and a second intermediate feature. The first intermediate feature and the second intermediate feature are respectively input into the Vision-Mamba module and the CNN module to obtain global features and local features. The global features and local features are input into the first splicing layer for splicing to obtain a first spliced feature. The first spliced feature is input into the seventh convolutional layer to obtain the output features of the seventh convolutional layer. The output features of the seventh convolutional layer are added to the input features of the Vision-Mamba CNN module to obtain the output features of the Vision-Mamba CNN module; the CNN module includes an eighth convolutional layer, a fifth activation function layer, a ninth convolutional layer, and a sixth activation function layer. The input features of the CNN module are input into the CNN module and successively pass through the eighth convolutional layer, the fifth activation function layer, the ninth convolutional layer, and the sixth activation function layer. The output features of the sixth activation function layer are added to the input features of the CNN module to obtain local features; the convolutional kernel sizes of the sixth convolutional layer and the seventh convolutional layer are both ; the convolutional kernel sizes of the eighth convolutional layer and the ninth convolutional layer are both .
[0058] Referring to Figure 5 , the Vision-Mamba CNN module is briefly denoted as , The specific process is as follows: The input features of the Vision-Mamba CNN module first pass through a sixth convolutional layer with a convolutional kernel size of , a stride of 1, and a padding of 1. After that, they are split into two parts by a split layer and , enter the Vision-Mamba module for processing, and the output is ; then successively pass through two convolutional layers, the eighth and ninth convolutional layers, each with a convolutional kernel size of , a stride of 1, and a padding of 1. After the eighth and ninth convolutional layers, there is an activation function layer, namely the fifth activation function layer and the sixth activation function layer. Then, the results of the two convolutions are added to to obtain . Then, and are concatenated at the first concatenation layer. After concatenation, they pass through a seventh convolutional layer with a convolutional kernel size of , a stride of 1, and a padding of 1. The output features obtained are added to the input features of the original Vision-Mamba CNN module to obtain . The role of using the Vision-Mamba module is that it can perform multi-scale and multi-level feature extraction on images, accurately capture the global structural details of images, and thus more precisely retain important information during the compression process. The reason for using the CNN module is that it has good local feature extraction ability.
[0059] Furthermore, referring to Figure 6 , the image to be compressed is input into the non-linear transformation network to obtain the latent representation , as shown in the following formula:
[0060] ;
[0061] where represents the input image to be compressed, represents a tensor with 3 channels and a height × width of 256×256, represents the output of the non-linear transformation network, represents a tensor with 128 channels and a height × width of 16×16.
[0062] Referring to Figure 7 , the latent representation is input into the hyperprior transformation network to calculate the latent representation of the hyperprior transformation , as shown in the following formula:
[0063] ;
[0064] Among them, is expressed as the output of the hyperprior transformation network, and is a tensor with 128 channels and a height×width of 4×4.
[0065] In a specific embodiment, the first quantizer and the second quantizer add noise uniformly distributed in during the training phase of the image compression model to the input features of the first quantizer and the input features of the second quantizer, and perform a rounding operation on the input features of the first quantizer and the input features of the second quantizer during the inference phase of the image compression model; the first encoder and the first decoder use arithmetic encoders, and the first decoder and the second decoder use arithmetic decoders.
[0066] Specifically, the latent representation is input into the first quantizer to obtain the quantized latent representation, as shown in the following formula:
[0067] ;
[0068] The latent representation of the hyperprior transformation is input into the second quantizer to obtain the quantized latent representation of the hyperprior transformation, as shown in the following formula:
[0069] ;
[0070] Among them, represents the quantized latent representation of the hyperprior transformation, is a tensor with 128 channels and a height×width of 4×4, represents the quantized latent representation, and is a tensor with 128 channels and a height×width of 16×16 which are the functions corresponding to the first quantizer and the second quantizer respectively. During the training phase, noise uniformly distributed in [-0.5, 0.5] is added to the input features, and during the inference phase, the input features are rounded.
[0071] Furthermore, the second encoder performs arithmetic coding on the quantized latent representation of the hyperprior transformation to obtain the second bitstream. The second decoder performs arithmetic decoding on the second bitstream to obtain the corresponding second latent representation feature ; referring to Figure 8 the second latent representation feature is input into the hyperprior inverse transformation network to obtain the second latent representation feature of the hyperprior transformation, as shown in the following formula:
[0072] ;
[0073] Among them, represents the second latent representation feature of the hyperprior inverse transform, which is represented as a tensor with 128 channels and a height × width of 16 × 16.
[0074] In a specific embodiment, the context module includes a second concatenation layer, a wavelet convolution layer, a channel context model, and an entropy parameter layer connected in sequence. The quantized latent representation and the second latent representation feature of the hyperprior transform are input into the second concatenation layer to obtain a second concatenation feature. The second concatenation feature passes through the wavelet convolution layer and the channel context model in sequence to obtain a third intermediate feature. The third intermediate feature is input into the entropy parameter layer to obtain a Gaussian distribution. The channel context model includes a tenth convolution layer, a first activation function layer, an eleventh convolution layer, a second activation function layer, and a twelfth convolution layer.
[0075] Specifically, referring to Figure 9 , the calculation process of the context module is shown in the following formula:
[0076] ;
[0077] Among them, represents the mean, represents the variance, C represents the second concatenation layer, and the second concatenation layer performs concatenation according to the channel dimension. WC represents the function corresponding to the wavelet convolution (Wavelet Convolutions) layer, CCM represents the function corresponding to the channel context model (Channel Context Model), and EP represents the entropy parameter (Entropy Parameters) layer. Among them, the wavelet convolution layer combines the multi-scale analysis ability of wavelet transform and the local feature extraction ability of convolution operation to discard redundant information. The wavelet convolution layer belongs to an existing structure, and its specific structure will not be elaborated here. The role of the channel context model is to adjust the number of channels, and the number of channels changes from 256 to 512 and then to 256. The specific operation is as follows:
[0078] ;
[0079] Among them, represents the output feature of the wavelet convolution layer, represents the output feature of the channel context model, 3 represents the convolution kernel size of , 256 or 512 represents the number of channels, represents The activation function, and the channel context model does not change the data size; the entropy parameter layer calculates the Gaussian distribution and uses the Gaussian distribution to adjust the probability model of the arithmetic codec, making the compression ratio closer to the entropy of the data and improving the compression efficiency. A dynamic probability model is established through the Gaussian distribution to adaptively adjust the symbol probability to meet the compression requirements of different data. The entropy parameter layer uses data chunking operation to divide the data into 。
[0080] The first encoder performs arithmetic coding on the quantized latent representation During the coding process, the Gaussian distribution calculated by the entropy parameter layer is used to dynamically adjust the probability model of the arithmetic encoder, and then the first bitstream is obtained;
[0081] The first decoder decodes the first bitstream. During the decoding process, the Gaussian distribution calculated by the entropy parameter layer is used to dynamically adjust the probability model of the arithmetic decoder to obtain the first latent representation feature corresponding to the first bitstream 。
[0082] Refer to Figure 10 The first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image, as shown in the following formula:
[0083] ;
[0084] Among them, represents the output of the non-linear inverse transformation network, that is, the compressed image, represents a tensor with 3 channels and a height×width of 256×256.
[0085] Specifically, the end-to-end image compression method based on Vision-Mamba CNN proposed in the embodiments of the present application can significantly improve the coding efficiency through its unique network architecture design. This method not only fully utilizes the spatial information capture ability of the Vision-Mamba module in the feature extraction stage. In this embodiment, to ensure the efficient operation and stable training of the algorithm, the software environment of this method is carefully configured to use the Python3.10 programming language, combined with the CUDA11.8 acceleration library, the pytorch2.1.1 deep learning framework, and the torchvision image processing library. This combination not only provides powerful computing capabilities but also ensures the efficient execution and maintainability of the code.
[0086] The image compression training in the embodiments of the present invention is carried out on input images with a size of 256×256 pixels. To make full use of computing resources and accelerate the training process, the batch size is set to 4, that is, 4 images are processed simultaneously in each iteration. The total number of iterations is set to 1000 times to ensure that the model can fully learn the internal laws of the image data. The initial learning rate is set to 0.005, which has been verified through multiple experiments to achieve a relatively fast convergence rate while ensuring the stability of training. During the training process, the Adam optimizer is adopted, which can adaptively adjust the learning rate to further optimize the training effect. In terms of the dataset, a total of 6000 images are collected, and these images are carefully divided into a training set, a test set, and a validation set according to a ratio of 8:1:1. To increase the diversity of the data and improve the generalization ability of the model, data augmentation operations of horizontal and vertical flipping are performed on the training set data, which not only increases the number of training samples but also enables the model to learn more diverse image features. In the training stage, the loss function is as follows:
[0087] ;
[0088] The loss function adopted in this embodiment comprehensively considers the balance between the reconstruction error and the bit rate, and includes a weight coefficient used to adjust the relative importance between the two. Specifically, the loss function consists of two parts: D is the mean square error (MSE), which is used to measure the difference between the compressed image and the original image; R is the bit rate, which is calculated by the number of bits of the second bitstream output by the second encoder and the number of bits of the first bitstream output by the first encoder, and reflects the number of bits occupied during the compression coding process. By continuously optimizing this loss function, the model can achieve a lower reconstruction error while ensuring a higher compression ratio, thereby improving the overall compression performance.
[0089] For further reference Figure 11 , as an implementation of the methods shown in the above figures, an embodiment of an end-to-end image compression device based on Vision-Mamba CNN is provided in the present application. This device embodiment corresponds to the Figure 1 method embodiment shown, and this device can be specifically applied to various electronic devices.
[0090] An embodiment of an end-to-end image compression device based on Vision-Mamba CNN is provided in the embodiments of the present application, including:
[0091] The model construction module 1 is configured to construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module;
[0092] The compression module 2 is configured to obtain an image to be compressed and input it into the trained image compression model. The image first undergoes non-linear transformation through the non-linear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain a latent representation of the hyperprior transformation; the latent representation of the hyperprior transformation is input into the second quantizer to obtain a quantized latent representation of the hyperprior transformation; the quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain a second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain a corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image.
[0093] Figure 12 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. As Figure 12 shown, the electronic device of this embodiment includes: a processor 1201 and a memory 1202; wherein the memory 1202 is used to store computer execution instructions; the processor 1201 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0094] Optionally, the memory 1202 can be either independent or integrated with the processor 1201.
[0095] When the memory 1202 is independently set, the electronic device further includes a bus 1203 for connecting the memory 1202 and the processor 1201.
[0096] An embodiment of the present invention also provides a computer storage medium. Computer execution instructions are stored in the computer storage medium. When the processor 1201 executes the computer execution instructions, the above method is implemented.
[0097] An embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 1201, the above method is implemented.
[0098] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.
[0099] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0100] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.
[0101] The above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 1201 to execute some steps of the methods in various embodiments of the present application.
[0102] It should be understood that the above processor 1201 can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 1201 can also be any conventional processor 1201, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the hardware processor 1201, or can be implemented by a combination of hardware and software modules in the processor 1201.
[0103] The memory 1202 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.
[0104] The bus 1203 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 1203 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 1203 in the drawings of this application is not limited to only one bus 1203 or one type of bus 1203.
[0105] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0106] An exemplary storage medium is coupled to the processor 1201, so that the processor 1201 can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor 1201. The processor 1201 and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor 1201 and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0107] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An end-to-end image compression method based on Vision-Mamba CNN, characterized in that, The method includes the following steps: Construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module. The non-linear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent downsampling modules; the non-linear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent upsampling modules; the hyperprior transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent downsampling modules; the hyperprior inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent upsampling modules; Obtain an image to be compressed and input it into the trained image compression model. The image first undergoes non-linear transformation through the non-linear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain a hyperprior-transformed latent representation; the hyperprior-transformed latent representation is input into the second quantizer to obtain a quantized hyperprior-transformed latent representation; the quantized hyperprior-transformed latent representation sequentially passes through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain a second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The context module includes a second concatenation layer, a wavelet convolutional layer, a channel context model, and an entropy parameter layer connected in sequence. The quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the second concatenation layer to obtain a second concatenated feature. The second concatenated feature sequentially passes through the wavelet convolutional layer and the channel context model to obtain a third intermediate feature. The third intermediate feature is input into the entropy parameter layer to obtain the Gaussian distribution; the channel context model includes a tenth convolutional layer, a first GeLU activation function layer, an eleventh convolutional layer, a second GeLU activation function layer, and a twelfth convolutional layer connected in sequence; data chunking operation is adopted in the entropy parameter layer; the quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain a corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain a compressed image.
2. The end-to-end image compression method based on Vision-Mamba CNN according to claim 1, characterized in that, The downsampling module includes a first convolutional layer, a first division normalization layer, a second convolutional layer, a first Leaky ReLU activation function layer, a third convolutional layer, and a second Leaky ReLU activation function layer connected in sequence; wherein, the first convolutional layer, the second convolutional layer, and the third convolutional layer are all convolutional structures with a convolutional kernel size of 3×3 and an output channel number of c, the first division normalization layer uses division normalization operation, and the first Leaky ReLU activation function layer and the second Leaky ReLU activation function layer both use the Leaky ReLU activation function; the input feature of the downsampling module is input into the downsampling module and sequentially passes through the first convolutional layer, the first division normalization layer, the second convolutional layer, the first Leaky ReLU activation function layer, the third convolutional layer, and the second Leaky ReLU activation function layer, and the output feature of the first division normalization layer is added to the output feature of the second Leaky ReLU activation function layer to obtain the output feature of the downsampling module, as shown in the following formula: RES(x) = GDN(Conv (3,c) (x)) + LR(Conv (3,c) (LR(Conv (3,c) (GDN(Conv (3,c) (x)))))); Among them, x represents the input feature of the downsampling module, RES(·) represents the function corresponding to the downsampling module, RES(x) represents the output feature of the downsampling module, and Conv (3,c) (·) represents the function corresponding to the convolutional structure with a convolutional kernel size of 3×3 and an output channel number of c, GDN(·) represents the division normalization operation; LR(·) represents the LeakyReLU activation function; The upsampling module includes a transposed convolutional layer, a second division normalization layer, a fourth convolutional layer, a third Leaky ReLU activation function layer, a fifth convolutional layer, and a fourth Leaky ReLU activation function layer connected in sequence; wherein, the transposed convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are all convolutional structures with a convolutional kernel size of 3×3 and an output channel number of c, the second division normalization layer uses division normalization operation, and the third Leaky ReLU activation function layer and the fourth Leaky ReLU activation function layer both use the Leaky ReLU activation function; the input feature of the upsampling module is input into the upsampling module and sequentially passes through the transposed convolutional layer, the second division normalization layer, the fourth convolutional layer, the third Leaky ReLU activation function layer, the fifth convolutional layer, and the fourth Leaky ReLU activation function layer, and the output feature of the second division normalization layer is added to the output feature of the fourth Leaky ReLU activation function layer to obtain the output feature of the upsampling module, as shown in the following formula: RET(x′) = GDN(TConv (3,c) (x')) + LR(Conv (3,c) (LR(Conv (3,c) (GDN(TConv (3,c) ((x')))))); Among them, x′ represents the input feature of the upsampling module, RET(·) represents the function corresponding to the upsampling module, RET(x′) represents the output feature of the upsampling module, and TConv (3,c) (·) represents the function corresponding to the transposed convolutional layer with a convolutional kernel size of 3×3 and an output channel number of c.
3. The end-to-end image compression method based on Vision-Mamba CNN according to claim 1, characterized in that The Vision-Mamba CNN module includes a sixth convolutional layer, a segmentation layer, a Vision-Mamba module, a CNN module, a first concatenation layer, and a seventh convolutional layer; the input features of the Vision-Mamba CNN module are input into the Vision-Mamba CNN module, and sequentially pass through the sixth convolutional layer and the segmentation layer to obtain a first intermediate feature and a second intermediate feature. The first intermediate feature and the second intermediate feature are respectively input into the Vision-Mamba module and the CNN module to obtain a global feature and a local feature. The global feature and the local feature are input into the first concatenation layer for concatenation to obtain a first concatenated feature. The first concatenated feature is input into the seventh convolutional layer to obtain the output features of the seventh convolutional layer. The output features of the seventh convolutional layer are added to the input features of the Vision-Mamba CNN module to obtain the output features of the Vision-Mamba CNN module; the CNN module includes an eighth convolutional layer, a fifth LeakyReLU activation function layer, a ninth convolutional layer, and a sixth LeakyReLU activation function layer connected in sequence. The input features of the CNN module are input into the CNN module and sequentially pass through the eighth convolutional layer, the fifth LeakyReLU activation function layer, the ninth convolutional layer, and the sixth LeakyReLU activation function layer. The output features of the sixth LeakyReLU activation function layer are added to the input features of the CNN module to obtain the local feature; the convolutional kernels of the sixth convolutional layer and the seventh convolutional layer are both 1×1; the convolutional kernels of the eighth convolutional layer and the ninth convolutional layer are both 3×3.
4. The end-to-end image compression method based on Vision-Mamba CNN according to claim 1, wherein, During the training stage of the image compression model, the first quantizer and the second quantizer add noise uniformly distributed in [-0.5, 0.5] to the input features of the first quantizer and the second quantizer. During the inference stage of the image compression model, rounding operations are performed on the input features of the first quantizer and the second quantizer; the first encoder and the first decoder use arithmetic encoders, and the first decoder and the second decoder use arithmetic decoders.
5. An end-to-end image compression device based on Vision-Mamba CNN, characterized in that, It includes: A model construction module, configured to construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module. The non-linear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent downsampling modules; the non-linear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent upsampling modules; the hyperprior transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent downsampling modules; the hyperprior inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent upsampling modules; A compression module, configured to obtain an image to be compressed and input it into the trained image compression model. The image first undergoes non-linear transformation through the non-linear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain a latent representation of hyperprior transformation; the latent representation of hyperprior transformation is input into the second quantizer to obtain a quantized latent representation of hyperprior transformation; the quantized latent representation of hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain a second latent representation feature of hyperprior transformation; the quantized latent representation and the second latent representation feature of hyperprior transformation are input into the context module to obtain a Gaussian distribution. The context module includes a second concatenation layer, a wavelet convolutional layer, a channel context model, and an entropy parameter layer connected in sequence. The quantized latent representation and the second latent representation feature of hyperprior transformation are input into the second concatenation layer to obtain a second concatenated feature. The second concatenated feature sequentially passes through the wavelet convolutional layer and the channel context model to obtain a third intermediate feature. The third intermediate feature is input into the entropy parameter layer to obtain the Gaussian distribution; the channel context model includes a tenth convolutional layer, a first GeLU activation function layer, an eleventh convolutional layer, a second GeLU activation function layer, and a twelfth convolutional layer connected in sequence; data chunking operation is adopted in the entropy parameter layer; the quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain a corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain a compressed image.
6. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-4 is implemented.
8. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Visual representation method and device based on bidirectional state space model
CN117876845A
Multi-scale Mama transformation unmanned aerial vehicle visual gesture flight control method
CN119206795A