End-to-end image compression method based on mixed attention and SwinV2 entropy model
By introducing hybrid attention and SwinV2 entropy model into the end-to-end image compression method, the problem of unutilized channel importance and high computational complexity is solved, and more efficient image compression and reconstruction effects are achieved.
Patent Information
- Application Number
- CN202510102238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
When existing end-to-end image compression methods deal with multi-scale structures and complex image data, there are problems such as channel importance differences and high computational complexity.
Using an image compression method based on hybrid attention and SwinV2 entropy model, the channel importance is adjusted adaptively through the hybrid attention module, and the attention module is moved to the channel autoregressive entropy model, and the channel-level attention mechanism is used for channel-by-channel dependency modeling.
The compression performance and spatial feature perception capabilities of the model are improved, the complexity of the model is significantly reduced, and more efficient image compression and reconstruction are achieved.
Smart Images

Figure CN119946271A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of Internet big data and new generation information technology, and in particular to an end-to-end image compression method based on hybrid attention and SwinV2 entropy model. Background Art
[0002] With the rapid development of information technology and big data, multimedia data has shown a sharp growth trend in many fields. Image compression is an important research direction in image processing. With the rapid growth of image data, lossy image compression technology plays a vital role in efficient storage and transmission. Although traditional image compression methods, such as JPEG, JPEG2000, HEVC and VVC standards, have achieved remarkable results in rate-distortion performance, they still have certain limitations when facing increasingly complex image data.
[0003] With the development of deep learning technology, end-to-end learned image compression (LIC) technology can more effectively balance bit rate and distortion in the compression and reconstruction process through a global optimization framework, and has gradually become a research hotspot in the current field. Recently, some LIC works have surpassed the current most advanced classical coding standard VVC in terms of peak signal-to-noise ratio and multi-scale structural similarity, demonstrating its application potential in the next generation of image compression.
[0004] The LIC framework mainly consists of four parts: transformation network, entropy model, quantization and super prior network. At present, researchers have mainly proposed transformation networks based on CNN and Transformer architectures. The input image is usually converted into a large number of channel features, and each channel is processed with equal importance. However, each channel is not equally important for image tasks, so the image compression performance of existing methods has certain limitations. In addition, in end-to-end image compression, it is crucial to build a superior entropy model to accurately estimate the probability distribution of the potential representation. A large number of studies have shown that the attention module can help the model learn complex areas in the image. Therefore, when designing the entropy model of the LIC framework, the introduction of the attention mechanism can more accurately estimate the probability distribution of the potential representation. However, if the attention module is directly applied to the main branch and the super prior path, the computational overhead will be significantly increased, especially the large input size of the main branch will lead to the complexity of the model.
[0005] In summary, how to design an end-to-end image compression method that can improve the compression performance of the model and significantly reduce the complexity of the model is a technical problem that needs to be solved urgently. Summary of the invention
[0006] In view of the deficiencies of the above-mentioned prior art, the technical problem to be solved by the present invention is: how to provide an end-to-end image compression method based on hybrid attention and SwinV2 entropy model, which can adaptively adjust the importance of channels and focus on the extraction of important features, improve the modeling capabilities of global and local features and the spatial feature perception capabilities, and at the same time move the attention module to the channel autoregressive entropy model, and use the channel-level attention mechanism to model channel-by-channel dependencies, thereby improving the compression performance of the model and significantly reducing the complexity of the model.
[0007] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0008] An end-to-end image compression method based on hybrid attention and SwinV2 entropy model, comprising:
[0009] S1: Get the image to be compressed;
[0010] S2: Input the image to be compressed into the trained image compression model and output the corresponding reconstructed image;
[0011] The training steps of the image compression model include:
[0012] S201: Acquire an original image used as a training sample as an input of an image compression model;
[0013] S202: extracting several channel features generated by the convolution transformation of the original image x through the main encoder, and adaptively focusing on the channel features important to the compression task to generate a potential representation y;
[0014] S203: Capture redundant information between potential representations y through a hyper-prior network to generate a potential representation z; calculate a Gaussian distribution (μ, σ) based on the potential representation z;
[0015] S204: Model the potential representation y using the Gaussian probability model combined with the Gaussian distribution (μ, σ) through the S2CAEM entropy model to obtain a probability estimate of the potential representation; generate the potential representation based on the probability estimate
[0016] S205: The main decoder decodes the latent representation Perform decoding and reconstruction to generate a reconstructed image;
[0017] S206: Calculate the loss function based on the difference between the reconstructed image and the original image and the number of code stream bits generated by compressing the original image, and reversely optimize the parameters of the image compression model
[0018] S207: Repeat S201 to S206 to iteratively train the image compression model until the model converges or reaches a maximum number of iterations;
[0019] S3: Using the reconstructed image as an end-to-end image compression result of the image to be compressed.
[0020] Preferably, in steps S202 and S205, the main encoder includes a first convolution layer, a first GDN module, a second convolution layer, a second GDN module, a first MAB module, a third convolution layer, a third GDN module, a fourth convolution layer and a second MAB module connected end to end in sequence;
[0021] The main decoder includes a third MAB module, a fifth convolutional layer, a first IGDN module, a sixth convolutional layer, a fourth MAB module, a second IGDN module, a seventh convolutional layer, a third IGDN module, and an eighth convolutional layer connected end to end in sequence;
[0022] The MAB module is used to make the model focus on channel features that are more important for compression tasks.
[0023] Preferably, the MAB module includes a plurality of RHAB modules connected end to end in sequence and a convolutional layer; the input features of the MAB module are residually connected with the output of the convolutional layer to obtain the output features;
[0024] The RHAB module includes several MWAB modules, CWCB modules and convolutional layers connected end to end in sequence; the input features of the RHAB module are residually connected with the output of the convolutional layer to obtain the output features;
[0025] The MWAB module is used to calculate the channel attention weights through global information.
[0026] Preferably, the MWAB module comprises a first LN layer connected end to end, a first OCSA module and a W-MSA module connected in parallel, a second LN layer, a first multilayer perceptron, a third LN layer, a second OCSA module and a SW-MSA module connected in parallel, a fourth LN layer, and a second multilayer perceptron;
[0027] The workflow of the MWAB module: the input feature X passes through the first LN layer to obtain the intermediate feature X N ;X N After passing through the first OCSA module and the W-MSA module respectively, the output features of the first OCSA module and the W-MSA module are added together, and then added to αX to obtain the intermediate feature X M ;X M After the second LN layer and the first multi-layer perceptron, αX M Add together to get the intermediate feature X′; X′ passes through the third LN layer to get the intermediate feature X N ′;X NAfter ′ passes through the second OCSA module and the SW-MSA module respectively, the output features of the second OCSA module and the SW-MSA module are added together, and then added to αX′ to obtain the intermediate feature X M ′;X M ′ passes through the fourth LN layer and the second multi-layer perceptron and is combined with αX M ′Add together to get the output feature Y;
[0028] The formula is:
[0029] X N = LN(X);
[0030] X M =W-MSA(X N )+βOCSA(X N )+αX;
[0031] X′=MLP(LN(X M ))+αX M ;
[0032] X N ′ = LN(X′);
[0033] X M ′=SW-MSA(X N ′)+βOCSA(X N ′)+αX′
[0034] Y = MLP(LN(X M ′))+αX M ′;
[0035] Where: α represents the learnable parameter; β represents a constant.
[0036] Preferably, the calculation formula of the OCSA module is expressed as:
[0037] Weight=Conv(DW-D-Conv(DW-Conv(X)));
[0038]
[0039] Where: X′ represents the output feature; X represents the input feature; Weight represents the weight generated by the LKSA module in the OCSA module; DW-Conv represents a 5×5 deep convolutional layer; DW-D-Conv represents a 7×7 deep dilated convolutional layer; Conv represents a 1×1 convolution;
[0040] The LKSA module includes a deep convolutional layer, a deep dilated convolutional layer and a 1×1 convolutional layer which are connected end to end in sequence.
[0041] Preferably, the CWCB module comprises a first LN layer, a CWB module, a second LN layer and a multilayer perceptron connected end to end in sequence;
[0042] The workflow of the CWCB module: After the input features pass through the first LN layer and the CWB module, the obtained features are added to the input features to obtain the intermediate features; after the intermediate features pass through the second LN layer and the multi-layer perceptron, the obtained features are added to the intermediate features to obtain the final output features;
[0043] The CWB module calculates self-attention by dividing the standard window and the overlapping window, and the sizes of the standard window and the overlapping window are different.
[0044] Preferably, in step S203, the processing steps of the super prior network include:
[0045] S2031: Through the super prior encoder h a Encode the latent representation y to obtain the latent representation z of the super prior network;
[0046] The formula is:
[0047] z=h a (y; φ h );
[0048] Where: φ h represents learnable parameters;
[0049] S2032: quantize the potential representation z into a potential representation through a quantizer Q
[0050] The formula is:
[0051]
[0052] S2033: Potential Representation via Fully Decomposed Entropy After the probability estimation, the super prior decoder h s Potential Decode and obtain Gaussian distribution (μ, σ);
[0053] The formula is:
[0054]
[0055] Where: θ h represents learnable parameters;
[0056] The super prior encoder h a and the super prior decoder h sThey all include five convolutional layers connected end to end in sequence, and GELU activation functions are connected between adjacent convolutional layers.
[0057] Preferably, in step S204, the processing steps of the S2CAEM entropy model include:
[0058] S2041: Potential representation Divided into s slices
[0059] S2042: Calculate each slice using the S2CAEM entropy model combined with a Gaussian distribution (μ, σ) The residual r i , that is, to obtain a probability estimate of the potential representation;
[0060] The formula is:
[0061]
[0062] Where: Φ i represents the probability distribution parameters; Decoded slices;
[0063] S2043: Through the residual r i Calculate each slice Get slices with smaller errors
[0064] S2044: All slices Combination to form a potential representation
[0065] Preferably, the S2CAEM entropy model includes inputs that are both Gaussian distributions (μ, σ) and slices The three branches of
[0066] The first branch and the second branch both include a connection layer, a SwinT V2 module, and a parameters layer connected end to end; the SwinT V2 module is used to calculate the parameters according to the Gaussian distribution (μ, σ) and the slice Calculate probability estimates; the parameters layer includes three convolutional layers connected end to end, and a GELU activation function is set between adjacent convolutional layers;
[0067] The third branch includes a connection layer and an LRP layer which are connected end to end in sequence; the LRP layer includes three convolutional layers which are connected end to end in sequence, and a GELU activation function is set between adjacent convolutional layers.
[0068] Preferably, in step S206, the loss function is calculated as follows:
[0069]
[0070] Where: L represents the loss function of the model; R represents the potential representation of the main encoder and the latent representation of the hyper-prior network The sum of the resulting bit rates, potentially representing and is the quantized value of the potential representation z and y; D represents the original image x and the reconstructed image Distortion; Express expectation; p x represents the probability distribution of the original image x; d represents the original image x and the reconstructed image The difference between λ and λ represents the trade-off factor between bit rate and distortion. and Respectively represent the potential representation based on the full decomposition entropy and the S2CAEM entropy model and probability estimate.
[0071] Compared with the prior art, the end-to-end image compression method based on hybrid attention and SwinV2 entropy model in the present invention has the following beneficial effects:
[0072] In the current end-to-end image compression transformation network, a large number of channel features are treated with equal importance, which shows limitations in CNN and Transformer structures. The present invention designs a hybrid attention (MAB) module in the transformation network part of the image compression model. The module integrates the orthogonal channel-space attention (OCSA) and the window self-attention mechanism based on Transformer, which can extract features from multiple channel features of the image and adaptively adjust the channel features that are most important for the compression task, thereby achieving complementary advantages and spatial feature extraction in global feature and local detail modeling, and obtaining a more compact and rich feature expression, thereby improving the compression performance of the model. At the same time, the orthogonal channel-space attention (OCSA) used by the hybrid attention module realizes the adaptive allocation of channel weights by initializing an orthogonal filter. Due to the orthogonality of the filter, the filter can extract information from the orthogonal subspace of the feature space and focus on unique features, thereby further ensuring the compression performance of the model. And the present invention introduces large-core spatial attention in the orthogonal channel-space attention, which can further enhance the spatial perception ability of the model.
[0073] In view of the problem that applying the attention module directly to the main branch and the hyper-prior path in the end-to-end image compression transformation network will significantly increase the computational overhead. The present invention designs a channel autoregressive entropy model based on the SwinTransformer V2 architecture, namely the S2CAEM entropy model, in the entropy model part of the image compression model, and moves the attention module to the S2CAEM entropy model, so that the input size is only 1 / 16 of the main branch, which significantly reduces the complexity of the model. At the same time, the entropy model improves the probability estimation accuracy of the potential representation by using the channel-level attention mechanism to model the channel-by-channel dependencies, thereby enhancing the compression performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:
[0075] Figure 1 This is the network structure diagram of the image compression model.
[0076] Figure 2 Figure 2 is the network structure diagram of the hybrid attention (MAB) module.
[0077] Figure 3 Figure 2. Network structure diagram of the mixed window attention (MWAB) module.
[0078] Figure 4 The network structure diagram of the Orthogonal Channel-Spatial Attention (OCSA) module.
[0079] Figure 5 This is the network structure diagram of the LKSA module.
[0080] Figure 6 This is the network structure diagram of the cross-window connection (CWCB) module.
[0081] Figure 7 This is the network structure diagram of the S2CAEM entropy model.
[0082] Figure 8 This is the rate-distortion performance comparison of this experiment on three datasets. DETAILED DESCRIPTION
[0083] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.
[0084] The following is a further detailed description through specific implementation methods:
[0085] Example:
[0086] This embodiment discloses an end-to-end image compression method based on hybrid attention and SwinV2 entropy model.
[0087] An end-to-end image compression method based on hybrid attention and SwinV2 entropy model, comprising:
[0088] S1: Get the image to be compressed;
[0089] S2: Input the image to be compressed into the trained image compression model and output the corresponding reconstructed image (i.e. the compressed image);
[0090] Combination Figure 1 As shown, the training steps of the image compression model include:
[0091] S201: Acquire an original image used as a training sample as an input of an image compression model;
[0092] In this embodiment, 30,000 images are randomly selected from the OpenImages dataset as training samples, and in the dataset preprocessing stage, the images are cropped into samples with a size of 256 × 256. The batch size of each training iteration is set to 16.
[0093] S202: extracting several channel features generated by the convolution transformation of the original image x through the main encoder, and adaptively focusing on the channel features important to the compression task to generate a potential representation y;
[0094] S203: Capture redundant information between potential representations y through a hyper-prior network to generate a potential representation z; calculate a Gaussian distribution (μ, σ) based on the potential representation z;
[0095] S204: Use the S2CAEM (Swin Transformer V2-based Channel-wise Autoregressive Entropy Model) entropy model to model the potential representation y using a Gaussian probability model combined with a Gaussian distribution (μ, σ) to obtain a probability estimate of the potential representation; generate a potential representation based on the probability estimate
[0096] S205: The main decoder decodes the latent representation Perform decoding and reconstruction to generate a reconstructed image;
[0097] S206: Calculate a loss function based on the difference between the reconstructed image and the original image and the number of code stream bits generated by compressing the original image, and reversely optimize the parameters of the image compression model;
[0098] S207: Repeat S201 to S206 to iteratively train the image compression model until the model converges or reaches a maximum number of iterations;
[0099] S3: Using the reconstructed image as an end-to-end image compression result of the image to be compressed.
[0100] In view of the fact that a large number of channel features are treated with equal importance in the current end-to-end image compression transformation network, the method shows limitations in the CNN and Transformer structures. The present invention designs a mixed attention block (MAB) module in the transformation network part of the image compression model. The module integrates the orthogonal channel-spatial attention (OCSA) and the window self-attention mechanism based on the Transformer, which can extract features from multiple channel features of the image and adaptively adjust the channel features that are most important for the compression task, thereby achieving complementary advantages and spatial feature extraction in global feature and local detail modeling, and obtaining a more compact and rich feature expression, thereby improving the compression performance of the model. At the same time, the orthogonal channel-spatial attention used by the mixed attention module realizes the adaptive allocation of channel weights by initializing an orthogonal filter. Due to the orthogonality of the filter, the filter can extract information from the orthogonal subspace of the feature space, focus on unique features, and further ensure the compression performance of the model. In addition, the present invention introduces large-core spatial attention in the orthogonal channel-spatial attention, which can further enhance the spatial perception ability of the model.
[0101] In view of the problem that applying the attention module directly to the main branch and the hyper-prior path in the end-to-end image compression transformation network will significantly increase the computational overhead. The present invention designs a channel autoregressive entropy model based on the SwinTransformer V2 architecture, namely the S2CAEM entropy model, in the entropy model part of the image compression model, and moves the attention module to the S2CAEM entropy model, so that the input size is only 1 / 16 of the main branch, which significantly reduces the complexity of the model. At the same time, the entropy model improves the probability estimation accuracy of the potential representation by using the channel-level attention mechanism to model the channel-by-channel dependencies, thereby enhancing the compression performance of the model.
[0102] In summary, the present invention significantly improves the quality of the reconstructed image while maintaining a high compression rate, and improves the limitation of the existing method that does not fully consider the difference in the importance of channel features.
[0103] To better introduce the technical solution of the present invention, this embodiment is described through the following parts.
[0104] 1. Main encoder and main decoder
[0105] The relationship between the main encoder, quantizer and main decoder is expressed as:
[0106]
[0107] Where: x represents the input image, g a represents the main encoder, y represents the potential representation, and y is quantized by Q to form a discrete potential representation g s represents the main decoder, Represents the compressed image, that is, the reconstructed image.
[0108] In this embodiment, the hybrid attention (MAB) module is used to adjust the channel weights through learning to highlight the important channels. At the same time, the large-core spatial attention is used to enhance the spatial perception ability of the model and obtain a more compact image feature expression. Therefore, a transformation network based on hybrid attention is designed. a ,g s ) module to address the limitation of treating each channel with equal importance in the end-to-end image compression transformation network.
[0109] Combination Figure 1 As shown, the main encoder g a It includes the first convolution layer, the first GDN (generalized divisor normalization) module, the second convolution layer, the second GDN module, the first MAB module, the third convolution layer, the third GDN module, the fourth convolution layer and the second MAB module connected end to end in sequence;
[0110] Main decoder g s It includes a third MAB module, a fifth convolutional layer, a first IGDN module, a sixth convolutional layer, a fourth MAB module, a second IGDN module, a seventh convolutional layer, a third IGDN module, and an eighth convolutional layer connected end to end in sequence;
[0111] The GDN module is a normalization layer suitable for generation tasks such as image reconstruction. Unlike the traditional batch normalization (BN) layer, the GDN module does not introduce noise and is more suitable for generation tasks that require clear images. The IGDN module is the inverse process of the GDN module and is used in the image synthesis module to restore the original image from the representation coefficients and reconstruct the correlation of the data.
[0112] 1. Mixture Attention Block (MAB) module
[0113] In this embodiment, the MAB module is used to make the model focus on channel features that are more important for the compression task.
[0114] Combination Figure 2 As shown in FIG, the MAB module includes a series of Residual Hybrid Attention Blocks (RHABs) connected end to end and a convolutional layer; the input features of the MAB module are residually connected with the output of the convolutional layer to obtain the output features;
[0115] The RHAB module consists of several Mixture Window Attention Block (MWAB) modules, Cross-window Connections Block (CWCB) modules and convolutional layers connected end to end in sequence; the input features of the RHAB module are residually connected with the output of the convolutional layer to obtain the output features.
[0116] In the present invention, the design structure of the main decoder not only strengthens the attention mechanism in the channel dimension, but also enhances the feature fusion in the spatial dimension, thereby optimizing the feature representation in the image compression task.
[0117] 2. MWAB module
[0118] In this embodiment, the MWAB module is used to calculate the channel attention weights through global information. Specifically, the MWAB module combines the orthogonal channel-spatial attention (OCSA) into the standard Swin Transformer block to enhance the representation ability of the network.
[0119] Combination Figure 3As shown in FIG. 1 , the MWAB module includes the first LayerNorm (LN) layer connected end to end, the first OCSA module and the window-based multi-head self-attention (W-MSA) module in parallel, the second LN layer, the first multi-layer perceptron (MLP), the third LN layer, the second OCSA module and the shifted window-based multi-head self-attention (SW-MSA) module in parallel, the fourth LN layer and the second multi-layer perceptron;
[0120] The workflow of the MWAB module: the input feature X passes through the first LN layer to obtain the intermediate feature X N ;X N After passing through the first OCSA module and the W-MSA module respectively, the output features of the first OCSA module and the W-MSA module are added together, and then added to αX to obtain the intermediate feature X M ;X M After the second LN layer and the first multi-layer perceptron, αX M Add together to get the intermediate feature X′; X′ passes through the third LN layer to get the intermediate feature X N ′;X N After ′ passes through the second OCSA module and the SW-MSA module respectively, the output features of the second OCSA module and the SW-MSA module are added together, and then added to αX′ to obtain the intermediate feature X M ′;X M ′ passes through the fourth LN layer and the second multi-layer perceptron and is combined with αX M ′Add together to get the output feature Y;
[0121] The formula is:
[0122] X N = LN(X);
[0123] X M =W-MSA(X N )+βOCSA(X N )+αX;
[0124] X′=MLP(LN(X M ))+αX M ;
[0125] X N ′ = LN(X′);
[0126] X M ′=SW-MSA(X N ′)+βOCSA(XN ′)+αX′
[0127] Y = MLP(LN(X M ′))+αX M ′;
[0128] Where: α represents the learnable parameter; β represents a constant.
[0129] Transformer-based structures usually require a large number of channels for tag embedding, so enabling the model to adaptively allocate channel weights is crucial to improving performance. The OCSA used by the MAB model of the present invention achieves adaptive allocation of channel weights by initializing an orthogonal filter. Due to the orthogonality of the filter, the filter can extract information from the orthogonal subspace of the feature space, thereby focusing on unique features. Compared with traditional channel attention, orthogonal channel attention avoids the disadvantage of global average pooling that it is easy to lose low-frequency information, and can extract a richer representation of each feature map.
[0130] 3. Orthogonal-Channel Spatial Attention (OCSA) module
[0131] Combination Figure 4 As shown in the figure, spatial attention based on large kernel convolution is used in the OCSA module. The large kernel spatial attention consists of three parts: spatial local convolution with a convolution kernel size of 5×5, spatial global convolution with a convolution kernel size of 7×7, and channel convolution with a convolution kernel size of 1×1.
[0132] The calculation formula of the OCSA module is expressed as:
[0133] Weight=Conv(DW-D-Conv(DW-Conv(X)));
[0134]
[0135] Where: X′ represents the output feature; X represents the input feature; Weight represents the weight generated by the LKSA module in the OCSA module; DW-Conv represents a 5×5 deep convolutional layer; DW-D-Conv represents a 7×7 deep dilated convolutional layer; Conv represents a 1×1 convolution;
[0136] Combination Figure 5As shown in Figure 1, the LKSA module includes a deep convolutional layer, a deep dilated convolutional layer, and a 1×1 convolutional layer connected end to end in sequence. Unlike common attention methods, the LKSA module we use does not require additional normalization functions such as sigmoid and softmax, but directly multiplies the learned attention weight value Weight with the input X to obtain the final output feature map X′.
[0137] 5. Window self-attention calculation
[0138] For the calculation of self-attention, the input feature of size H×W×C is first divided into M×M The self-attention is then calculated in each local window. For the connection between non-overlapping windows, a shift window partitioning method is also used, and the shift size is set to half the size of the window.
[0139] In the W-MSA module and the SW-MSA module, the calculation formula of the window self-attention is expressed as:
[0140]
[0141] Where Q, K, V represent the query, key, and value of the attention head respectively; d represents the dimension of the query key; B represents the relative position encoding.
[0142] 6. Cross-window Connections Block (CWCB) module
[0143] In order to realize feature interaction between adjacent windows and enhance the representation ability of window self-attention, the present invention uses the CWCB module to realize cross-window connection.
[0144] Combination Figure 6 As shown in the figure, the CWCB module includes the first LN layer, the CWB module, the second LN layer and the multi-layer perceptron which are connected end to end in sequence; the working process of the CWCB module is as follows: after the input features pass through the first LN layer and the CWB module, the obtained features are added to the input features to obtain the intermediate features; after the intermediate features pass through the second LN layer and the multi-layer perceptron, the obtained features are added to the intermediate features to obtain the final output features.
[0145] The CWB module calculates self-attention by dividing the standard window and the overlapping window. The sizes of the standard window and the overlapping window are different. For the input feature X, X Q For size M×M Standard windows, X K and X V For size M O ×M O of Overlapping windows. O =(1+γ)×M, γ is a constant related to the overlap size, and this module sets γ=0.5. Different from the standard window attention, CWCB extracts keys and values from a larger range to obtain more useful information in each window.
[0146] 2. Super Prior Network
[0147] In the end-to-end image compression model, the super prior network is used to capture the redundancy between potential representations and improve the prediction probability of entropy coding for the potential representation. The processing steps of the super prior network include:
[0148] S2031: Through the super prior encoder h a Encode the latent representation y to obtain the latent representation z of the super prior network;
[0149] The formula is:
[0150] z=h a (y; φ h );
[0151] Where: φ h represents learnable parameters;
[0152] S2032: quantize the potential representation z into a potential representation through a quantizer Q
[0153] The formula is:
[0154]
[0155] S2033: Potential Representation via Fully Decomposed Entropy After the probability estimation, the super prior decoder h s Potential Decoding is performed to obtain a Gaussian distribution (μ, σ); "AE" and "AD" in the figure represent an arithmetic encoder and an arithmetic decoder, respectively, which are used to encode the potential representation into a code stream and decode the code stream into the potential representation.
[0156] The formula is:
[0157]
[0158] Where: θ h represents learnable parameters;
[0159] Combination Figure 1 As shown, the super prior encoder h a and the super prior decoder h sEach of them includes five convolutional layers connected end to end in sequence, and the adjacent convolutional layers are connected with GELU activation function. The first, second and fourth are convolutional layers with a convolution kernel size of 3×3 and a stride of 1, and the third and fifth are convolutional layers with a convolution kernel size of 3×3 and a stride of 2; the decoder part corresponds to the encoder part.
[0160] 3. S2CAEM (Swin Transformer V2-based Channel-wise Autoregressive Entropy Model) entropy model
[0161] In the end-to-end image compression model, entropy coding uses a Gaussian probability model to model each potential representation for rate estimation and entropy coding. In order to better balance the network rate-distortion performance and speed, the present invention uses Swin Transformer V2-based Channel-wise Autoregressive Entropy Model (S2CAEM) entropy coding in the entropy coding part.
[0162] S2CAEM performs channel-level slicing on the input potential representation, and then inputs the Gaussian distribution (σ, μ) obtained by the hyper-prior network and the input slices into the SwinT V2 module to obtain accurate probability estimates. S2CAEM is a channel autoregressive model. When encoding subsequent slices, the previously decoded channel slices can be used as context information to obtain better probability estimates and improve the compression efficiency of the model. Divided into s slices The encoding process of each slice uses the encoded slice as context information to improve the encoding of subsequent slices. After being processed sequentially by the entropy model, it is decoded into In this process, the decoded slices and the current slice Input into the entropy model network together to better obtain the estimated probability distribution parameter Φ i , which helps subsequent slices to be better encoded into bitstreams. The entropy model models each slice of the received potential representation as having a mean μ i and variance σ i The Gaussian distribution of is, the mean and variance are the distribution characteristics of a potential representation point. The mean determines the center position of the distribution, and the variance represents the degree of dispersion of the center position. The formula is:
[0163]
[0164] Since quantization will inevitably produce quantization error during the encoding process, this quantization error will cause the decoded image to be distorted, so the present invention uses local residual prediction to predict this quantization error. As the number of slices increases, the estimation of the entropy model parameters becomes more accurate. In addition, the input size of the attention mechanism applied to the entropy model in the present invention is only 1 / 16 of that of the main encoder, which greatly reduces the computational complexity.
[0165] Specifically, the processing steps of the S2CAEM entropy model include:
[0166] S2041: Potential representation Divided into s slices
[0167] S2042: Calculate each slice using the S2CAEM entropy model combined with a Gaussian distribution (μ, σ) The residual r i , that is, to obtain a probability estimate of the potential representation;
[0168] The formula is:
[0169]
[0170] Where: Φ i represents the probability distribution parameters; Decoded slices;
[0171] where Φ i Used for AE and AD to generate slices Previous slice Inputting the S2CAEM entropy model allows the Φ of subsequent slices to be i More accurate.
[0172] S2043: Through the residual r i Calculate each slice Get slices with smaller errors
[0173] S2044: All slices Combination to form a potential representation
[0174] Combination Figure 7 As shown, the S2CAEM entropy model includes inputs of Gaussian distribution (μ, σ) and slices The three branches of
[0175] The first branch and the second branch both include a connection layer, a SwinT V2 module, and a parameters layer connected end to end; the SwinT V2 module is used to calculate the parameters according to the Gaussian distribution (μ, σ) and the slice Calculate probability estimates; the parameters layer includes three convolutional layers connected end to end, and a GELU activation function is set between adjacent convolutional layers;
[0176] The third branch includes a connection layer and an LRP layer which are connected end to end in sequence; the LRP layer includes three convolutional layers which are connected end to end in sequence, and a GELU activation function is set between adjacent convolutional layers.
[0177] 4. Loss Function
[0178] In this embodiment, the loss function is calculated as follows:
[0179]
[0180] Where: L represents the loss function of the model; R represents the potential representation of the main encoder and the latent representation of the hyper-prior network The sum of the resulting bit rates, potentially representing and is the quantized value of the potential representation z and y; D represents the original image x and the reconstructed image Distortion; Express expectation; p x represents the probability distribution of the original image x; d represents the original image x and the reconstructed image The difference between λ and λ represents the trade-off factor between bit rate and distortion. and Respectively represent the potential representation based on the full decomposition entropy (from the hyper-prior network) and the S2CAEM entropy model and probability estimate.
[0181] In the model verification phase, the model with the best performance evaluation is selected for testing the performance of the method proposed in the present invention. The Peak Signal-to-Noise Ratio (PSNR) indicator is used to evaluate the quality of image reconstruction, and bpp measures the number of bits after compression. BD-rate evaluates the bitrate saving rate of each method compared with the benchmark algorithm at the same quality.
[0182] 5. Experimental Description
[0183] 30,000 images were randomly selected from the OpenImages dataset as the training set and 1,000 images as the validation set. In the dataset preprocessing stage, the images were cropped into samples of size 256×256. The batch size of each training iteration was set to 8, and the Adam optimizer was used for model training. The model was trained based on a single NVIDIA GeForce RTX4090 GPU. First, 400 rounds of training were performed at a learning rate of 1×10-4, and then 100 rounds of training were performed at a learning rate of 1×10-5. The reason for selecting a smaller learning rate to train the model later is to make the model update the network parameters with a smaller amplitude when approaching the relative optimal solution, so as to avoid skipping the relative optimal solution in the parameter space.
[0184] This experiment uses Mean Square Error (MSE) and Multi-scale Structural Similarity (MS-SSIM) as image quality evaluation parameters for the optimization model. The loss function of this experiment is L = R + λ·D. When the model is optimized by MSE, λ is set to {0.0018, 0.0035, 0.0067, 0.0130, 0.025, 0.0483}, and when the model is optimized by MS-SSIM, λ is set to {2.4, 4.58, 8.73, 16.64, 31.73, 60.50}.
[0185] 1. Evaluation
[0186] This experiment tests the proposed method on three datasets, namely the Kodak image set with an image size of 768×512, the Tecnick test set with an image size of 1200×1200, and the CLIC professional validation dataset with a resolution of 2K. This experiment uses PSNR and MS-SSIM to measure the distortion, and uses Bpp to evaluate the bitrate.
[0187] 2. Rate-distortion performance
[0188] This experiment compares the proposed method with some LIC methods that have achieved SOTA performance and the classic image compression codec VVC. Figure 8 The rate-distortion performance of each method on the Kodak dataset is shown. This experiment tested PSNR and MS-SSIM on Kodak, and the results verified the stability of the algorithm proposed in the present invention. In order to facilitate a clearer comparison, this experiment converted MS-SSIM to -10log10(1-MS-SSIM). As shown in the results, at the same bit rate, compared with some existing SOTA methods, the model proposed in the present invention shows superior performance in both PSNR and MS-SSIM.
[0189] In order to obtain quantitative results, this experiment calculated BD-rate as a quantitative indicator and used the rate-distortion performance of VVC (VTM 14.0) as a benchmark. The BD-rate of the model of the present invention on the Kodak, Tecnick and CLIC Pro datasets increased by 16.67%, 15.15% and 14.96% respectively. Table 1 shows the results on the Kodak dataset.
[0190] Table 1 Comparison of BD-Rate of various methods on three datasets based on VVC
[0191]
[0192] 3. Comparison of encoding and decoding time
[0193] Table 2 shows the encoding and decoding time comparison between the proposed method and other state-of-the-art methods on the Kodak dataset.
[0194] The method of the present invention adopts a channel-wise autoregressive strategy to divide the potential representation into 10 slices. This experiment uses average time for comparison. The experimental results show that VTM14.0 requires the longest encoding time. Among these compared models, Cheng et al. use a context entropy model, which results in a longer encoding and decoding time. In addition, although Fu et al. achieved relatively superior performance, it was at the expense of the complexity of the model, and the encoding and decoding time was significantly slower than other methods. The proposed method is superior to the current SOTA method in encoding and decoding time.
[0195] Table 2 Comparison of encoding and decoding time
[0196]
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.
Claims
1. An end-to-end image compression method based on hybrid attention and SwinV2 entropy model, characterized in that: include: S1: Get the image to be compressed; S2: Input the image to be compressed into the trained image compression model and output the corresponding reconstructed image; The training steps of the image compression model include: S201: Acquire an original image used as a training sample as an input of an image compression model; S202: extracting several channel features generated by the convolution transformation of the original image x through the main encoder, and adaptively focusing on the channel features important to the compression task to generate a potential representation y; S203: Capture redundant information between potential representations y through a hyper-prior network to generate a potential representation z; calculate a Gaussian distribution (μ, σ) based on the potential representation z; S204: Model the potential representation y using the Gaussian probability model combined with the Gaussian distribution (μ, σ) through the S2CAEM entropy model to obtain a probability estimate of the potential representation; generate the potential representation based on the probability estimate S205: The main decoder decodes the latent representation Perform decoding and reconstruction to generate a reconstructed image; S206: Calculate the loss function based on the difference between the reconstructed image and the original image and the number of code stream bits generated by compressing the original image, and reversely optimize the parameters of the image compression model S207: Repeat S201 to S206 to iteratively train the image compression model until the model converges or reaches a maximum number of iterations; S3: Using the reconstructed image as an end-to-end image compression result of the image to be compressed.
2. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 1, characterized in that: In steps S202 and S205, the main encoder includes a first convolution layer, a first GDN module, a second convolution layer, a second GDN module, a first MAB module, a third convolution layer, a third GDN module, a fourth convolution layer and a second MAB module connected end to end in sequence; The main decoder includes a third MAB module, a fifth convolutional layer, a first IGDN module, a sixth convolutional layer, a fourth MAB module, a second IGDN module, a seventh convolutional layer, a third IGDN module, and an eighth convolutional layer connected end to end in sequence; The MAB module is used to make the model focus on channel features that are more important for compression tasks.
3. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 2, characterized in that: The MAB module includes several RHAB modules connected end to end and a convolutional layer. The input features of the MAB module are residually connected with the output of the convolutional layer to obtain the output features. The RHAB module includes several MWAB modules, CWCB modules and convolutional layers connected end to end in sequence; the input features of the RHAB module are residually connected with the output of the convolutional layer to obtain the output features; The MWAB module is used to calculate the channel attention weights through global information.
4. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 3, characterized in that: The MWAB module includes the first LN layer connected end to end, the first OCSA module and W-MSA module in parallel, the second LN layer, the first multi-layer perceptron, the third LN layer, the second OCSA module and SW-MSA module in parallel, the fourth LN layer, and the second multi-layer perceptron; The workflow of the MWAB module: the input feature X passes through the first LN layer to obtain the intermediate feature X N ;X N After passing through the first OCSA module and the W-MSA module respectively, the output features of the first OCSA module and the W-MSA module are added together, and then added to αX to obtain the intermediate feature X M ;X M After the second LN layer and the first multi-layer perceptron, αX M Add together to get the intermediate feature X′; X′ passes through the third LN layer to obtain the intermediate feature X N ′;X N After ′ passes through the second OCSA module and the SW-MSA module respectively, the output features of the second OCSA module and the SW-MSA module are added together, and then added to αX′ to obtain the intermediate feature X M ′;X M ′ passes through the fourth LN layer and the second multi-layer perceptron and is combined with αX M ′Add together to get the output feature Y; The formula is: X N =LN(X); X M =W-MSA(X N )+βOCSA(X N )+αX; X′=MLP(LN(X M ))+αX M ; X N ′=LN(X′); X M ′=SW-MSA(X N ′)+βOCSA(X N ′)+αX′ Y=MLP(LN(X M ′))+αX M ′; Where: α represents the learnable parameter; β represents a constant.
5. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 4, characterized in that: The calculation formula of the OCSA module is expressed as: Weight=Conv(DW-D-Conv(DW-Conv(X))); Where: X′ represents the output feature; X represents the input feature; Weight represents the weight generated by the LKSA module in the OCSA module; DW-Conv represents a 5×5 deep convolutional layer; DW-D-Conv represents a 7×7 deep dilated convolutional layer; Conv represents a 1×1 convolution; The LKSA module includes a deep convolutional layer, a deep dilated convolutional layer and a 1×1 convolutional layer which are connected end to end in sequence.
6. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 3, characterized in that: The CWCB module includes the first LN layer, the CWB module, the second LN layer, and the multilayer perceptron connected end to end in sequence; Workflow of the CWCB module: After the input features pass through the first LN layer and the CWB module, the obtained features are added to the input features to obtain the intermediate features; After the intermediate features pass through the second LN layer and the multi-layer perceptron, the obtained features are added to the intermediate features to obtain the final output features; The CWB module calculates self-attention by dividing the standard window and the overlapping window, and the sizes of the standard window and the overlapping window are different.
7. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 1, characterized in that: In step S203, the processing steps of the super prior network include: S2031: Through the super prior encoder h a Encode the latent representation y to obtain the latent representation z of the super prior network; The formula is: z=h a (y;φ h ); Where: φ h represents learnable parameters; S2032: quantize the potential representation z into a potential representation through a quantizer Q The formula is: S2033: Potential Representation via Fully Decomposed Entropy After the probability estimation, the super prior decoder h s Potential Decode and obtain Gaussian distribution (μ, σ); The formula is: Where: θ h represents learnable parameters; The super prior encoder h a and the super prior decoder h s They all include five convolutional layers connected end to end in sequence, and GELU activation functions are connected between adjacent convolutional layers.
8. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 1, characterized in that: In step S204, the processing steps of the S2CAEM entropy model include: S2041: Potential representation Divided into s slices S2042: Calculate each slice using the S2CAEM entropy model combined with a Gaussian distribution (μ, σ) The residual r i , that is, to obtain a probability estimate of the potential representation; The formula is: Where: Φ i represents the probability distribution parameters; Decoded slices; S2043: Through the residual r i Calculate each slice Get slices with smaller errors S2044: All slices Combination to form a potential representation 9. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 1, characterized in that: The S2CAEM entropy model includes inputs of Gaussian distribution (μ, σ) and slices The three branches of The first branch and the second branch both include a connection layer, a SwinT V2 module, and a parameters layer connected end to end; the SwinT V2 module is used to calculate the parameters according to the Gaussian distribution (μ, σ) and the slice Calculate probability estimates; the parameters layer includes three convolutional layers connected end to end, and a GELU activation function is set between adjacent convolutional layers; The third branch includes a connection layer and an LRP layer which are connected end to end in sequence; the LRP layer includes three convolutional layers which are connected end to end in sequence, and a GELU activation function is set between adjacent convolutional layers.
10. The end-to-end image compression method based on hybrid attention and SwinV2 entropy model as claimed in claim 1, characterized in that: In step S206, the loss function is calculated as follows: Where: L represents the loss function of the model; R represents the potential representation of the main encoder and the latent representation of the hyper-prior network The sum of the resulting bit rates, potentially representing and is the quantized value of the potential representation z and y; D represents the original image x and the reconstructed image Distortion; Express expectation; p x represents the probability distribution of the original image x; d represents the original image x and the reconstructed image The difference between λ and λ represents the trade-off factor between bit rate and distortion. and Respectively represent the potential representation based on the full decomposition entropy and the S2CAEM entropy model and probability estimate.
Citation Information
Patent Citations
Image compression model based on graph attention and asymmetric convolutional network
CN115512199A
Land utilization classification method based on multiple attention semantic segmentation
CN115908946A
Transform-based electron microscope pollen image target detection method
CN117197632A
End-to-end image compression method based on context clustering transformation
CN117456017A
Image super-resolution reconstruction method based on lightweight hybrid attention network
CN117745541A