Vision-Mamba CNN-based end-to-end image compression method and device

By adopting an end-to-end method based on Vision-Mamba CNN in image compression, combining the fusion structure of the Vision-Mamba module and the CNN module and the multi-scale context module, the problem of global dependency modeling and computing complexity in image compression is solved, and efficient image compression and refined balance reconstruction quality is achieved.

CN120088349AActive Publication Date: 2025-06-03HUAQIAO UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510562636.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively model the global dependence of images in image compression, resulting in high-frequency details loss and structural distortion; the model parameters based on Transformer are huge and the computational complexity is high, making it difficult to deploy on mobile devices; the traditional hyper-priori network does not fully consider the dynamic changes in feature distribution when generating edge information, resulting in insufficient balance optimization of compression code rate and reconstruction quality.

Method used

The end-to-end image compression method based on Vision-Mamba CNN is adopted, and the fusion structure between the Vision-Mamba module and the CNN module is introduced to improve the long-distance dependency capture capability and reduce the computational complexity; a context module that fuses wavelet convolution layer and channel context model is designed to realize the deep fusion of multi-scale feature information and dynamic parameterized Gaussian distribution modeling, and balance the bit rate and reconstruction quality are refined.

Benefits of technology

It significantly improves image compression performance, reduces high-frequency detail loss and structural distortion, reduces computing costs to meet mobile deployment needs, and achieves a refined balance between compression code rate and reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088349A_ABST
    Figure CN120088349A_ABST
Patent Text Reader

Abstract

The invention discloses an end-to-end image compression method and device based on a Vision-Mamba CNN, and relates to the field of image processing, and the method comprises the steps: obtaining a to-be-compressed image, inputting the to-be-compressed image into a trained image compression model, obtaining potential representations through a nonlinear transformation network, inputting the potential representations into a first quantizer and a super-prior transformation network, and obtaining a second quantizer; obtaining a quantized potential representation and a potential representation of hyper-priori transformation; quantizing the potential representation of the hyper-priori transformation, and sequentially passing through a second encoder, a second decoder and a hyper-priori inverse transformation network to obtain a second potential representation feature of the hyper-priori transformation; the quantized potential representation and the second potential representation feature of the hyper-prior transformation are input into a context module to obtain Gaussian distribution, the quantized potential representation sequentially passes through a first encoder and a first decoder and is combined with the Gaussian distribution to obtain a first potential representation feature, and the first potential representation feature is input into a nonlinear inverse transformation network to obtain a compressed image; the problems of low compression efficiency and low reconstruction quality are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to an end-to-end image compression method and device based on Vision-Mamba CNN. Background Art

[0002] With the rapid development of multimedia technology, the storage and transmission requirements of image data have increased exponentially. Traditional image compression technologies (such as JPEG, HEVC, etc.) are mainly based on discrete cosine transform (DCT) and entropy coding, and obvious block effects and detail losses will occur at high compression rates.

[0003] In recent years, deep learning-based image compression methods (such as VAE, GAN, Transformer, etc.) have significantly improved compression performance through end-to-end learning, but still face the following technical challenges:

[0004] (1) Traditional convolutional neural networks (CNNs) are limited by local receptive fields and are difficult to effectively model global dependencies in images, resulting in the loss of high-frequency details and structural distortion;

[0005] (2) Transformer-based models capture long-range dependencies through self-attention mechanisms, but have a large number of parameters and high computational complexity (O(N²)), making it difficult to be deployed on mobile devices; existing methods usually only utilize single-scale features and lack the fusion of multi-scale context information, affecting compression efficiency and reconstruction quality;

[0006] (3) Traditional hyperprior networks (such as Variational-Autoencoder) do not fully consider the dynamic changes of feature distributions when generating side information, resulting in insufficient balance optimization between compression bitrate and reconstruction quality. Summary of the Invention

[0007] The purpose of this application is to propose an end-to-end image compression method and device based on Vision-Mamba CNN for the above-mentioned technical problems.

[0008] In a first aspect, the present invention provides an end-to-end image compression method based on Vision-Mamba CNN, including the following steps:

[0009] Construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module;

[0010] Obtain the image to be compressed and input it into the trained image compression model. The image first undergoes a non-linear transformation through a non-linear transformation network to obtain the corresponding latent representation; the latent representation is input into the first quantizer to obtain the quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain the latent representation of the hyperprior transformation; the latent representation of the hyperprior transformation is input into the second quantizer to obtain the quantized latent representation of the hyperprior transformation; the quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain the corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain the second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain the corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image.

[0011] Preferably, the non-linear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent downsampling modules; the non-linear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent upsampling modules; the hyperprior transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent downsampling modules; the hyperprior inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent upsampling modules.

[0012] Preferably, the downsampling module includes a first convolutional layer, a first division normalization layer, a second convolutional layer, a first activation function layer, a third convolutional layer, and a second activation function layer connected in sequence; wherein, the first convolutional layer, the second convolutional layer, and the third convolutional layer are all convolutional structures with a convolutional kernel size of and an output channel number of . The first division normalization layer adopts a division normalization operation. The first activation function layer and the second activation function layer both adopt activation function; the input feature of the downsampling module is input into the downsampling module and sequentially passes through the first convolutional layer, the first division normalization layer, the second convolutional layer, the first activation function layer, the third convolutional layer, and the second activation function layer. The output feature of the first division normalization layer and the second The output features of the activation function layer are added together to obtain the output features of the downsampling module, as shown in the following formula:

[0013] ;

[0014] Among them, represents the input features of the downsampling module, represents the function corresponding to the downsampling module, represents the output features of the downsampling module, represents that the convolution kernel size is , and the number of output channels is the function corresponding to the convolution structure of represents the division normalization operation; represents the activation function;

[0015] The upsampling module includes a transposed convolutional layer, a second division normalization layer, a fourth convolutional layer, a third activation function layer, a fifth convolutional layer, and a fourth activation function layer connected in sequence; among them, the transposed convolutional layer, the fourth convolutional layer, and the fifth convolutional layer are all convolution structures with a convolution kernel size of , and the number of output channels is ; the second division normalization layer uses the division normalization operation, and the third activation function layer and the fourth activation function layer both use the activation function; the input features of the upsampling module are input into the upsampling module and sequentially pass through the transposed convolutional layer, the second division normalization layer, the fourth convolutional layer, the third activation function layer, the fifth convolutional layer, and the fourth activation function layer. The output features of the second division normalization layer are added to the output features of the fourth activation function layer to obtain the output features of the upsampling module, as shown in the following formula:

[0016] ;

[0017] Among them, represents the input features of the upsampling module, represents the function corresponding to the upsampling module, represents the output features of the upsampling module, represents that the convolution kernel size is , and the number of output channels is the function corresponding to the transposed convolutional layer of

[0018] Preferably, the Vision-Mamba CNN module includes a sixth convolutional layer, a segmentation layer, a Vision-Mamba module, a CNN module, a first concatenation layer, and a seventh convolutional layer; the input features of the Vision-Mamba CNN module are input into the Vision-Mamba CNN module, and successively pass through the sixth convolutional layer and the segmentation layer to obtain a first intermediate feature and a second intermediate feature. The first intermediate feature and the second intermediate feature are respectively input into the Vision-Mamba module and the CNN module to obtain a global feature and a local feature. The global feature and the local feature are input into the first concatenation layer for concatenation to obtain a first concatenated feature. The first concatenated feature is input into the seventh convolutional layer to obtain the output features of the seventh convolutional layer. The output features of the seventh convolutional layer are added to the input features of the Vision-Mamba CNN module to obtain the output features of the Vision-Mamba CNN module; the CNN module includes an eighth convolutional layer, a fifth activation function layer, a ninth convolutional layer, and a sixth activation function layer. The input features of the CNN module are input into the CNN module and successively pass through the eighth convolutional layer, the fifth activation function layer, the ninth convolutional layer, and the sixth activation function layer. The output features of the sixth activation function layer are added to the input features of the CNN module to obtain the local features; the convolutional kernels of the sixth convolutional layer and the seventh convolutional layer both have a size of ; the convolutional kernels of the eighth convolutional layer and the ninth convolutional layer both have a size of .

[0019] Preferably, the context module includes a second concatenation layer, a wavelet convolutional layer, a channel context model, and an entropy parameter layer connected in sequence. The quantized latent representation and the second latent representation features of the hyperprior transformation are input into the second concatenation layer to obtain a second concatenated feature. The second concatenated feature successively passes through the wavelet convolutional layer and the channel context model to obtain a third intermediate feature. The third intermediate feature is input into the entropy parameter layer to obtain a Gaussian distribution; the channel context model includes a tenth convolutional layer, a first activation function layer, an eleventh convolutional layer, a second activation function layer, and a twelfth convolutional layer; data chunking operations are used in the entropy parameter layer.

[0020] Preferably, the first quantizer and the second quantizer adopt adding a uniform distribution on the input features of the first quantizer and the input features of the second quantizer during the training stage of the image compression model During the inference stage of the image compression model, perform a rounding operation on the input features of the first quantizer and the input features of the second quantizer for the noise; the first encoder and the first decoder use arithmetic encoders, and the first decoder and the second decoder use arithmetic decoders.

[0021] In a second aspect, the present invention provides an end-to-end image compression device based on Vision-Mamba CNN, including:

[0022] A model construction module, configured to construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module;

[0023] A compression module, configured to obtain an image to be compressed and input it into the trained image compression model. The image first undergoes non-linear transformation through the non-linear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain a latent representation of the hyperprior transformation; the latent representation of the hyperprior transformation is input into the second quantizer to obtain a quantized latent representation of the hyperprior transformation; the quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain a second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain a corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain a compressed image.

[0024] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0025] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0026] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] (1) The end-to-end image compression method based on Vision-Mamba CNN proposed by the present invention introduces a fusion structure of the Vision-Mamba module and the CNN module, which significantly improves the long-distance dependence capture ability while maintaining low computational complexity, breaks through the local receptive field limitation of traditional CNNs, effectively captures the global dependence relationship of images, reduces the loss of high-frequency details and structural distortion; by replacing the traditional Transformer with Vision-Mamba CNN, the complexity of long-distance dependence modeling is reduced from O(N²) to linear complexity, significantly reducing the computational cost to meet the requirements of mobile deployment.

[0029] (2) The end-to-end image compression method based on Vision-Mamba CNN proposed by the present invention designs a context module that fuses the wavelet convolutional layer and the channel context model to achieve deep fusion of multi-scale feature information to enhance the feature representation ability; and constructs a dynamically parameterized Gaussian distribution modeling framework, which combines the side information generated by the hyperprior network to adjust the potential feature distribution parameters in real time, realizing a refined balance between the compression bitrate and the reconstruction quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0031] Figure 1 It is a schematic flowchart of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0032] Figure 2 It is a schematic diagram of the image compression model of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0033] Figure 3 It is a schematic diagram of the downsampling module of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0034] Figure 4 It is a schematic diagram of the upsampling module of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0035] Figure 5Schematic diagram of the Vision-Mamba module of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0036] Figure 6 Schematic diagram of the non-linear transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0037] Figure 7 Schematic diagram of the hyperprior transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0038] Figure 8 Schematic diagram of the hyperprior inverse transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0039] Figure 9 Schematic diagram of the context module of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0040] Figure 10 Schematic diagram of the non-linear inverse transformation network of the end-to-end image compression method based on Vision-Mamba CNN according to the embodiment of the present application;

[0041] Figure 11 Schematic diagram of the end-to-end image compression device based on Vision-Mamba CNN according to the embodiment of the present application;

[0042] Figure 12 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Figure 1 An end-to-end image compression method based on Vision-Mamba CNN provided by the embodiment of the present application is shown, including the following steps:

[0045] S1. Construct an image compression model and train it to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module.

[0046] In a specific embodiment, the non-linear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules disposed between two adjacent downsampling modules; the non-linear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules disposed between two adjacent upsampling modules; the hyperprior transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module disposed between two adjacent downsampling modules; the hyperprior inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module disposed between two adjacent upsampling modules.

[0047] Specifically, referring to Figure 2 , the image compression model in the embodiment of the present application is composed of a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module. In the non-linear transformation network, the combination of the downsampling module and the Vision-Mamba CNN module converts the image from the pixel space to a feature space more conducive to compression, removing the spatial correlation between image pixels and making the data distribution more convenient for subsequent processing. In the hyperprior transformation network, the combination of the downsampling module and the Vision-Mamba CNN module extracts the latent representation of the hyperprior transformation. The latent representation of the hyperprior transformation contains the statistical features and prior information of the image, and this information can help perform entropy coding more effectively subsequently and improve the compression performance. In the hyperprior inverse transformation network, the combination of the upsampling module and the Vision-Mamba CNN module performs a hyperprior inverse transformation on the obtained second latent representation feature, restoring part of the information related to the statistical characteristics of the original image, and its structure is a symmetric structure with the hyperprior transformation network. In the non-linear inverse transformation network, the combination of the upsampling module and the Vision-Mamba CNN module decompresses the first latent representation feature and restores the compressed image, and its structure is a symmetric structure with the non-linear transformation network.

[0048] S2. Obtain the image to be compressed and input it into the trained image compression model. The image first undergoes a non-linear transformation through a non-linear transformation network to obtain the corresponding latent representation. The latent representation is input into the first quantizer to obtain the quantized latent representation. The latent representation is input into the hyperprior transformation network to obtain the latent representation of the hyperprior transformation. The latent representation of the hyperprior transformation is input into the second quantizer to obtain the quantized latent representation of the hyperprior transformation. The quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain the corresponding second latent representation feature. The second latent representation feature is input into the hyperprior inverse transformation network to obtain the second latent representation feature of the hyperprior transformation. The quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain the corresponding first latent representation feature. The first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image.

[0049] Specifically, deploy the trained image compression model. During the inference process, input the image to be compressed into the trained image compression model. After processing, the compressed image is obtained. The specific details will be described one by one later.

[0050] In a specific embodiment, the downsampling module includes a first convolutional layer, a first division normalization layer, a second convolutional layer, a first activation function layer, a third convolutional layer, and a second activation function layer connected in sequence; where the first convolutional layer, the second convolutional layer, and the third convolutional layer are all convolutional structures with a convolutional kernel size of and an output channel number of . The first division normalization layer adopts a division normalization operation. The first activation function layer and the second activation function layer both adopt activation function; the input feature of the downsampling module is input into the downsampling module and sequentially passes through the first convolutional layer, the first division normalization layer, the second convolutional layer, the first activation function layer, the third convolutional layer, and the second activation function layer. The output feature of the first division normalization layer is added to the output feature of the second activation function layer to obtain the output feature of the downsampling module, as shown in the following formula:

[0051] ;

[0052] where, represents the input feature of the downsampling module, represents the corresponding function of the downsampling module, Represents the output feature of the downsampling module, Represents that the convolution kernel size is , and the number of output channels is The function corresponding to the convolution structure, Represents the division normalization operation; Represents Activation function;

[0053] The upsampling module includes a transposed convolution layer, a second division normalization layer, a fourth convolution layer, a third Activation function layer, a fifth convolution layer, and a fourth Activation function layer; among them, the transposed convolution layer, the fourth convolution layer, and the fifth convolution layer all have a convolution kernel size of , and the number of output channels is The convolution structure of, the second division normalization layer uses the division normalization operation, and the third Activation function layer and the fourth Activation function layer both use Activation function; the input feature of the upsampling module is input into the upsampling module, and successively passes through the transposed convolution layer, the second division normalization layer, the fourth convolution layer, the third Activation function layer, the fifth convolution layer, and the fourth Activation function layer, and the output feature of the second division normalization layer is added to the output feature of the fourth Activation function layer to obtain the output feature of the upsampling module, as shown in the following formula:

[0054] ;

[0055] Among them, Represents the input feature of the upsampling module, Represents the function corresponding to the upsampling module, Represents the output feature of the upsampling module, Represents that the convolution kernel size is , and the number of output channels is The function corresponding to the transposed convolution layer of.

[0056] Specifically, the downsampling module is denoted as , referring to Figure 3 , the downsampling module consists of a first convolution layer, a first division normalization layer, a second convolution layer, a first Activation function layer, a third convolution layer, and a second Activation function layer, and the output feature of the first division normalization layer and the second Activation function layer's output feature form a residual connection. Among them, the first convolution layer uses a convolution kernel size of , a convolutional structure with a stride of 2 and a padding of 1. Both the second convolutional layer and the third convolutional layer adopt a convolutional structure with a kernel size of , a stride of 1, and a padding of 1; the module above is denoted as , referring to Figure 4 , the upsampling module consists of a transposed convolutional layer, a second division normalization layer, a fourth convolutional layer, and a third activation function layer, a fifth convolutional layer, and a fourth activation function layer, and the output features of the second division normalization layer and the fourth activation function layer form a residual connection. Among them, the convolutional kernel size of the transposed convolutional layer is , a stride of 2, a padding of 1, and an output padding of 1; both the fourth convolutional layer and the fifth convolutional layer adopt a convolutional structure with a kernel size of , a stride of 1, and a padding of 1.

[0057] In a specific embodiment, the Vision-Mamba CNN module includes a sixth convolutional layer, a segmentation layer, a Vision-Mamba module, a CNN module, a first splicing layer, and a seventh convolutional layer; the input features of the Vision-Mamba CNN module are input into the Vision-Mamba CNN module and sequentially pass through the sixth convolutional layer and the segmentation layer to obtain a first intermediate feature and a second intermediate feature. The first intermediate feature and the second intermediate feature are respectively input into the Vision-Mamba module and the CNN module to obtain global features and local features. The global features and local features are input into the first splicing layer for splicing to obtain a first spliced feature. The first spliced feature is input into the seventh convolutional layer to obtain the output features of the seventh convolutional layer. The output features of the seventh convolutional layer are added to the input features of the Vision-Mamba CNN module to obtain the output features of the Vision-Mamba CNN module; the CNN module includes an eighth convolutional layer, a fifth activation function layer, a ninth convolutional layer, and a sixth activation function layer connected in sequence. The input features of the CNN module are input into the CNN module and sequentially pass through the eighth convolutional layer, the fifth activation function layer, the ninth convolutional layer, and the sixth activation function layer. The output features of the sixth activation function layer are added to the input features of the CNN module to obtain local features; the convolutional kernel sizes of the sixth convolutional layer and the seventh convolutional layer are both ; the convolutional kernel sizes of the eighth convolutional layer and the ninth convolutional layer are both .

[0058] Referring to Figure 5 , the Vision-Mamba CNN module is briefly denoted as , the specific process is as follows: The input features of the Vision-MambaCNN module first pass through a sixth convolutional layer with a convolutional kernel size of , a stride of 1, and a padding of 1. After that, they are split into two parts by a split layer and , enter the Vision-Mamba module for processing, and the output is ; then successively pass through two convolutional layers, the eighth and ninth convolutional layers, each with a convolutional kernel size of , a stride of 1, and a padding of 1. After the eighth and ninth convolutional layers, there is an activation function layer, namely the fifth activation function layer and the sixth activation function layer. Then, the results of the two convolutions are added to to obtain . Then, and are concatenated in the first concatenation layer. After concatenation, they pass through a seventh convolutional layer with a convolutional kernel size of , a stride of 1, and a padding of 1. The output features obtained are added to the input features of the original Vision-MambaCNN module to obtain . The role of using the Vision-Mamba module is that it can perform multi-scale and multi-level feature extraction on images, accurately capture the global structural details of images, and thus more precisely retain important information during the compression process. The role of using the CNN module is that it has good local feature extraction ability.

[0059] Furthermore, referring to Figure 6 , the image to be compressed is input into the non-linear transformation network to obtain the latent representation , as shown in the following formula:

[0060] ;

[0061] Among them, represents the input image to be compressed, represents a tensor with 3 channels and a height × width of 256×256, represents the output of the non-linear transformation network, represents a tensor with 128 channels and a height × width of 16×16.

[0062] Referring to Figure 7 , the latent representation is input into the hyperprior transformation network to calculate the latent representation of the hyperprior transformation, as shown in the following formula:

[0063] ;

[0064] Among them, is expressed as the output of the hyperprior transformation network, and is a tensor with 128 channels and a height×width of 4×4.

[0065] In a specific embodiment, the first quantizer and the second quantizer add noise uniformly distributed in during the training phase of the image compression model to the input features of the first quantizer and the input features of the second quantizer, and perform rounding operations on the input features of the first quantizer and the input features of the second quantizer during the inference phase of the image compression model; the first encoder and the first decoder use arithmetic encoders, and the first decoder and the second decoder use arithmetic decoders.

[0066] Specifically, the latent representation is input into the first quantizer to obtain the quantized latent representation, as shown in the following formula:

[0067] ;

[0068] The latent representation of the hyperprior transformation is input into the second quantizer to obtain the quantized latent representation of the hyperprior transformation, as shown in the following formula:

[0069] ;

[0070] Among them, represents the quantized latent representation of the hyperprior transformation, is a tensor with 128 channels and a height×width of 4×4, represents the quantized latent representation, and is a tensor with 128 channels and a height×width of 16×16 are the functions corresponding to the first quantizer and the second quantizer respectively. During the training phase, noise uniformly distributed in [-0.5, 0.5] is added to the input features, and during the inference phase, rounding operations are performed on the input features.

[0071] Furthermore, the second encoder performs arithmetic coding on the quantized latent representation of the hyperprior transformation to obtain a second bitstream. The second decoder performs arithmetic decoding on the second bitstream to obtain the corresponding second latent representation feature ; Referring to Figure 8 , the second latent representation feature is input into the hyperprior inverse transformation network to obtain the second latent representation feature of the hyperprior transformation, as shown in the following formula:

[0072] ;

[0073] Among them, represents the second latent representation feature of the hyperprior inverse transform, which is represented as a tensor with 128 channels and a height×width of 16×16.

[0074] In a specific embodiment, the context module includes a second concatenation layer, a wavelet convolution layer, a channel context model, and an entropy parameter layer connected in sequence. The quantized latent representation and the second latent representation feature of the hyperprior transform are input into the second concatenation layer to obtain a second concatenated feature. The second concatenated feature passes through the wavelet convolution layer and the channel context model in sequence to obtain a third intermediate feature. The third intermediate feature is input into the entropy parameter layer to obtain a Gaussian distribution. The channel context model includes a tenth convolution layer, a first activation function layer, an eleventh convolution layer, a second activation function layer, and a twelfth convolution layer.

[0075] Specifically, referring to Figure 9 , the calculation process of the context module is shown in the following formula:

[0076] ;

[0077] Among them, represents the mean, represents the variance, C represents the second concatenation layer, and the second concatenation layer concatenates along the channel dimension. WC represents the function corresponding to the wavelet convolution (Wavelet Convolutions) layer, CCM represents the function corresponding to the channel context model (Channel Context Model), and EP represents the entropy parameter (Entropy Parameters) layer. Among them, the wavelet convolution layer combines the multi-scale analysis ability of wavelet transform and the local feature extraction ability of convolution operation to discard redundant information. The wavelet convolution layer belongs to an existing structure, and its specific structure will not be elaborated here. The role of the channel context model is to adjust the number of channels, and the number of channels changes from 256 to 512 and then to 256. The specific operation is as follows:

[0078] ;

[0079] Among them, represents the output feature of the wavelet convolution layer, represents the output feature of the channel context model, 3 represents the convolution kernel size of , 256 or 512 represents the number of channels, represents The activation function, and the channel context model does not change the data size; the entropy parameter layer calculates the Gaussian distribution, uses the Gaussian distribution to adjust the probability model of the arithmetic codec, makes the compression ratio closer to the entropy of the data, improves the compression efficiency, establishes a dynamic probability model through the Gaussian distribution, and adaptively adjusts the symbol probability to meet the compression requirements of different data. The entropy parameter layer uses data chunking operations to divide the data into pieces.

[0080] The first encoder is used to perform arithmetic coding on the quantized latent representation During the coding process, the Gaussian distribution calculated by the entropy parameter layer is used to dynamically adjust the probability model of the arithmetic encoder, and then the first bitstream is obtained;

[0081] The first decoder is used to decode the first bitstream. During the decoding process, the Gaussian distribution calculated by the entropy parameter layer is used to dynamically adjust the probability model of the arithmetic decoder to obtain the first latent representation feature corresponding to the first bitstream pieces.

[0082] Referring to Figure 10 the first latent representation feature is input into the non-linear inverse transformation network to obtain the compressed image, as shown in the following formula:

[0083] ;

[0084] Wherein, represents the output of the non-linear inverse transformation network, which is the compressed image, represents a tensor with 3 channels and a height×width of 256×256.

[0085] Specifically, the end-to-end image compression method based on Vision-Mamba CNN proposed in the embodiments of the present application can significantly improve the coding efficiency through its unique network architecture design. This method not only makes full use of the spatial information capture ability of the Vision-Mamba module in the feature extraction stage. In this embodiment, to ensure the efficient operation and stable training of the algorithm, the software environment of this method is carefully configured to use the Python3.10 programming language, combined with the CUDA11.8 acceleration library, the pytorch2.1.1 deep learning framework, and the torchvision image processing library. This combination not only provides powerful computing capabilities but also ensures the efficient execution and maintainability of the code.

[0086] The image compression training in the embodiments of the present invention is performed on input images with a size of 256×256 pixels. To make full use of computing resources and accelerate the training process, the batch size is set to 4, that is, 4 images are processed simultaneously in each iteration. The total number of iterations is set to 1000 times to ensure that the model can fully learn the internal laws of the image data. The initial learning rate is set to 0.005. This value has been verified through multiple experiments and can achieve a relatively fast convergence rate while ensuring the stability of training. During the training process, the Adam optimizer is adopted, which can adaptively adjust the learning rate and further optimize the training effect. In terms of the dataset, a total of 6000 images are collected, and these images are carefully divided into a training set, a test set, and a validation set according to a ratio of 8:1:1. To increase the diversity of the data and improve the generalization ability of the model, data augmentation operations of horizontal and vertical flipping are performed on the training set data. This not only increases the number of training samples but also enables the model to learn more diverse image features. In the training stage, the loss function is as follows:

[0087] ;

[0088] The loss function adopted in this embodiment comprehensively considers the balance between the reconstruction error and the bit rate, and includes a weight coefficient used to adjust the relative importance between the two. Specifically, the loss function consists of two parts: D is the mean square error (MSE), which is used to measure the difference between the compressed image and the original image; R is the bit rate, which is calculated by the number of bits of the second bitstream output by the second encoder and the number of bits of the first bitstream output by the first encoder, and reflects the number of bits occupied during the compression coding process. By continuously optimizing this loss function, the model can achieve a lower reconstruction error while ensuring a high compression ratio, thereby improving the overall compression performance.

[0089] Further referring to Figure 11 , as an implementation of the methods shown in the above figures, an embodiment of an end-to-end image compression device based on Vision-Mamba CNN is provided in the present application. This device embodiment corresponds to the Figure 1 shown method embodiment, and this device can be specifically applied to various electronic devices.

[0090] An embodiment of an end-to-end image compression device based on Vision-Mamba CNN is provided in the embodiments of the present application, including:

[0091] The model construction module 1 is configured to construct and train an image compression model to obtain a trained image compression model. The image compression model includes a non-linear transformation network, a hyperprior transformation network, a non-linear inverse transformation network, a hyperprior inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module;

[0092] The compression module 2 is configured to obtain an image to be compressed and input it into the trained image compression model. The image first undergoes non-linear transformation through the non-linear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the hyperprior transformation network to obtain a latent representation of the hyperprior transformation; the latent representation of the hyperprior transformation is input into the second quantizer to obtain a quantized latent representation of the hyperprior transformation; the quantized latent representation of the hyperprior transformation sequentially passes through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the hyperprior inverse transformation network to obtain a second latent representation feature of the hyperprior transformation; the quantized latent representation and the second latent representation feature of the hyperprior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation sequentially passes through the first encoder and the first decoder and combines with the Gaussian distribution to obtain a corresponding first latent representation feature; the first latent representation feature is input into the non-linear inverse transformation network to obtain a compressed image.

[0093] Figure 12 It is a schematic hardware structure diagram of the electronic device provided by the embodiment of the present invention. As Figure 12 shown, the electronic device of this embodiment includes: a processor 1201 and a memory 1202; wherein the memory 1202 is used to store computer execution instructions; the processor 1201 is used to execute the computer execution instructions stored in the memory to implement the various steps executed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0094] Optionally, the memory 1202 can be either independent or integrated with the processor 1201.

[0095] When the memory 1202 is independently set, the electronic device further includes a bus 1203 for connecting the memory 1202 and the processor 1201.

[0096] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 1201 executes the computer execution instructions, the above method is implemented.

[0097] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 1201, the above method is implemented.

[0098] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0099] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0100] In addition, each functional module in various embodiments of the present invention can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in a unit. The unit formed by the above modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.

[0101] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 1201 to execute some steps of the methods in various embodiments of the present application.

[0102] It should be understood that the above processor 1201 can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor 1201 can also be any conventional processor 1201, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by the hardware processor 1201, or executed by a combination of hardware and software modules in the processor 1201.

[0103] The memory 1202 may include high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk, or an optical disc, etc.

[0104] The bus 1203 can be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 1203 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the bus 1203 in the drawings of this application does not limit that there is only one bus 1203 or one type of bus 1203.

[0105] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0106] An exemplary storage medium is coupled to the processor 1201, so that the processor 1201 can read information from the storage medium and can write information to the storage medium. Of course, the storage medium can also be a component of the processor 1201. The processor 1201 and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor 1201 and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0107] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: ROM, RAM, magnetic disk, or optical disc and other media that can store program codes.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An end-to-end image compression method based on Vision-Mamba CNN, characterized in that: The following steps are involved: Constructing and training an image compression model to obtain a trained image compression model, wherein the image compression model includes a nonlinear transformation network, a super a priori transformation network, a nonlinear inverse transformation network, a super a priori inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module; The image to be compressed is obtained and input into the trained image compression model, the image is first subjected to nonlinear transformation by the nonlinear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the super-a priori transformation network to obtain a latent representation of the super-a priori transformation; the latent representation of the super-a priori transformation is input into the second quantizer to obtain a quantized latent representation of the super-a priori transformation; the latent representation of the quantized super-a priori transformation is sequentially passed through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the super-a priori inverse transformation network to obtain a second latent representation feature of the super-a priori transformation; The quantized latent representation and the second latent representation features of the super-prior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation passes through the first encoder and the first decoder in sequence and is combined with the Gaussian distribution to obtain a corresponding first latent representation feature. The first latent representation feature is input into the nonlinear inverse transformation network to obtain a compressed image.

2. The end-to-end image compression method based on Vision-Mamba CNN according to claim 1, characterized in that: The nonlinear transformation network includes 4 downsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent downsampling modules; the nonlinear inverse transformation network includes 4 upsampling modules and 3 Vision-Mamba CNN modules arranged between two adjacent upsampling modules; the super priori transformation network includes 2 downsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent downsampling modules; the super priori inverse transformation network includes 2 upsampling modules and 1 Vision-Mamba CNN module arranged between two adjacent upsampling modules.

3. The end-to-end image compression method based on Vision-Mamba CNN according to claim 2, characterized in that: The downsampling module includes a first convolution layer, a first division normalization layer, a second convolution layer, a first Activation function layer, the third convolutional layer and the second Activation function layer; wherein the first convolution layer, the second convolution layer and the third convolution layer all have a convolution kernel size of , the number of output channels is The convolution structure of the first division normalization layer adopts the division normalization operation, and the first The activation function layer and the second The activation function layer uses activation function; the input features of the downsampling module are input into the downsampling module, and sequentially pass through the first convolution layer, the first division normalization layer, the second convolution layer, the first Activation function layer, the third convolutional layer and the second activation function layer, the output features of the first division normalization layer and the second The output features of the activation function layer are added to obtain the output features of the downsampling module, as shown in the following formula: ; in, represents the input features of the downsampling module, Represents the function corresponding to the downsampling module, represents the output features of the downsampling module, Indicates that the convolution kernel size is , the number of output channels is The function corresponding to the convolution structure is: Represents the division normalization operation; express Activation function; The upsampling module includes a transposed convolution layer, a second division normalization layer, a fourth convolution layer, a third Activation function layer, fifth convolutional layer and fourth Activation function layer; wherein the transposed convolution layer, the fourth convolution layer and the fifth convolution layer all have a convolution kernel size of , the number of output channels is The convolution structure of the second division normalization layer adopts the division normalization operation, and the third The activation function layer and the fourth The activation function layer uses activation function; the input features of the upsampling module are input into the upsampling module, and sequentially pass through the transposed convolution layer, the second division normalization layer, the fourth convolution layer, the third Activation function layer, fifth convolutional layer and fourth activation function layer, the output features of the second division normalization layer and the fourth The output features of the activation function layer are added to obtain the output features of the upsampling module, as shown in the following formula: ; in, represents the input features of the upsampling module, Represents the function corresponding to the upsampling module, represents the output features of the upsampling module, Indicates that the convolution kernel size is , the number of output channels is The function corresponding to the transposed convolutional layer.

4. The end-to-end image compression method based on Vision-Mamba CNN according to claim 2, characterized in that: The Vision-Mamba CNN module includes a sixth convolutional layer, a segmentation layer, a Vision-Mamba module, a CNN module, a first splicing layer and a seventh convolutional layer; the input features of the Vision-Mamba CNN module are input into the Vision-Mamba CNN module, and sequentially pass through the sixth convolutional layer and the segmentation layer to obtain a first intermediate feature and a second intermediate feature, the first intermediate feature and the second intermediate feature are respectively input into the Vision-Mamba module and the CNN module to obtain a global feature and a local feature, the global feature and the local feature are input into the first splicing layer for splicing to obtain a first splicing feature, the first splicing feature is input into the seventh convolutional layer to obtain an output feature of the seventh convolutional layer, the output feature of the seventh convolutional layer is added to the input feature of the Vision-Mamba CNN module to obtain an output feature of the Vision-Mamba CNN module; the CNN module includes an eighth convolutional layer, a fifth convolutional layer and a CNN module connected in sequence; Activation function layer, ninth convolution layer and sixth The input features of the CNN module are input into the CNN module and sequentially pass through the eighth convolutional layer, the fifth Activation function layer, ninth convolution layer and sixth Activation function layer, the sixth The output features of the activation function layer are added to the input features of the CNN module to obtain the local features; the convolution kernel sizes of the sixth and seventh convolution layers are both ; The convolution kernel sizes of the eighth and ninth convolution layers are both .

5. The end-to-end image compression method based on Vision-Mamba CNN according to claim 1, characterized in that: The context module includes a second splicing layer, a wavelet convolution layer, a channel context model and an entropy parameter layer connected in sequence. The quantized potential representation and the second potential representation feature of the super prior transformation are input to the second splicing layer to obtain a second splicing feature. The second splicing feature passes through the wavelet convolution layer and the channel context model in sequence to obtain a third intermediate feature. The third intermediate feature is input to the entropy parameter layer to obtain the Gaussian distribution. The channel context model includes a tenth convolution layer, a first Activation function layer, eleventh convolution layer, second activation function layer and the twelfth convolutional layer; the entropy parameter layer adopts data blocking operation.

6. The end-to-end image compression method based on Vision-Mamba CNN according to claim 1, characterized in that: The first quantizer and the second quantizer use a method of adding uniformly distributed features to the input features of the first quantizer and the second quantizer during the training phase of the image compression model. noise, and rounding the input features of the first quantizer and the input features of the second quantizer in the inference stage of the image compression model; the first encoder and the first decoder adopt arithmetic encoders, and the first decoder and the second decoder adopt arithmetic decoders.

7. An end-to-end image compression device based on Vision-Mamba CNN, characterized in that: include: A model building module is configured to build and train an image compression model to obtain a trained image compression model, wherein the image compression model includes a nonlinear transformation network, a super a priori transformation network, a nonlinear inverse transformation network, a super a priori inverse transformation network, a first quantizer, a second quantizer, a first encoder, a first decoder, a second encoder, a second decoder, and a context module; A compression module is configured to obtain an image to be compressed and input it into the trained image compression model, wherein the image is first subjected to a nonlinear transformation through the nonlinear transformation network to obtain a corresponding latent representation; the latent representation is input into the first quantizer to obtain a quantized latent representation; the latent representation is input into the super-a priori transformation network to obtain a latent representation of the super-a priori transformation; the latent representation of the super-a priori transformation is input into the second quantizer to obtain a quantized latent representation of the super-a priori transformation; the latent representation of the quantized super-a priori transformation is sequentially passed through the second encoder and the second decoder to obtain a corresponding second latent representation feature; the second latent representation feature is input into the super-a priori inverse transformation network to obtain a second latent representation feature of the super-a priori transformation; The quantized latent representation and the second latent representation features of the super-prior transformation are input into the context module to obtain a Gaussian distribution. The quantized latent representation passes through the first encoder and the first decoder in sequence and is combined with the Gaussian distribution to obtain a corresponding first latent representation feature. The first latent representation feature is input into the nonlinear inverse transformation network to obtain a compressed image.

8. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • End-to-end image compression method based on context clustering transformation

    CN117456017A

  • Visual representation method and device based on bidirectional state space model

    CN117876845A

  • Hyperspectral and multispectral remote sensing image fusion method based on multi-level collaborative mapping

    CN118898545A

  • Multi-scale Mama transformation unmanned aerial vehicle visual gesture flight control method

    CN119206795A

  • High-resolution remote sensing image change detection method based on multi-scale convolution decoding, electronic equipment and storage medium

    CN119418216A