Image compression method and device based on deep learning

By embedding the fusion attention module, fusion channel and window attention mechanism in the image encoding and decoding model, the problem of failing to fully utilize image feature information in the prior art is solved, and more efficient image compression effect and higher compression rate are achieved.

CN115484459BActive Publication Date: 2025-05-16TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211080617.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2025-05-16
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

When image compression is carried out by deep learning in the prior art, the channel correlation and global local correlation of image compression characteristics cannot be fully considered, resulting in unsatisfactory compression effect and poor compression rate.

Method used

By embedding the fusion attention module, the fusion channel attention mechanism and the window attention mechanism in the image encoding and decoding model, the feature information of the image data is better used for compression processing.

Benefits of technology

The rate distortion performance of deep learning image compression is improved, the image compression effect is significantly improved, and the compression rate is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115484459B_ABST
    Figure CN115484459B_ABST
Patent Text Reader

Abstract

The present invention provides an image compression method and device based on deep learning. The method comprises: determining image data to be processed; inputting the image data into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing, and obtaining a compressed target image output by the image encoding and decoding model; the fusion attention module is a processing module that fuses a channel attention mechanism and a window attention mechanism; the image encoding and decoding model is a deep learning model trained based on a sample image and an image processing result corresponding to the sample image. The image compression method based on deep learning provided by the present invention can better utilize channel correlation information and information of windows and shifted windows for image compression, improve the rate-distortion performance of deep learning image compression, effectively improve the image compression effect, and have a higher compression rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and more particularly to a method and device for image compression based on deep learning. In addition, the present invention also relates to an electronic device and a processor-readable storage medium. Background Art

[0002] In recent years, with the rapid development of computer technology, deep learning neural network technology has been widely used, and image encoding and decoding technology (i.e., image compression technology) based on deep learning has also emerged. However, the existing technology of using deep learning to compress images has high limitations, and does not take into account the channel correlation and global-local correlation of image compression features, so the compression effect is not ideal, and the compression rate and compression effect are poor. Therefore, how to provide a more effective image compression scheme based on deep learning to improve the compression rate and its compression effect has become a problem that needs to be solved urgently. Summary of the invention

[0003] To this end, the present invention provides an image compression method and device based on deep learning to solve the defect that the deep learning image compression scheme in the prior art has high limitations, resulting in poor compression rate and image compression effect.

[0004] In a first aspect, the present invention provides an image compression method based on deep learning, comprising:

[0005] determining image data to be processed;

[0006] The image data is input into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing to obtain a compressed target image output by the image encoding and decoding model; wherein the fusion attention module is a processing module that fuses the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on a sample image and the image processing results corresponding to the sample image.

[0007] Furthermore, before the image to be processed is input into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, it also includes:

[0008] Determine an initial image coding and decoding model; the initial image coding and decoding model includes an image coding network and an image decoding network; wherein the image coding network is used to encode the input image data to obtain a potential representation of the image data; the image decoding network is used to process the acquired potential representation after arithmetic decoding to obtain image reconstruction features;

[0009] The fusion attention module is respectively embedded into the corresponding target positions in the image encoding network and the image decoding network to obtain a first image encoding and decoding model to be trained; wherein the fusion attention module includes a channel attention module for obtaining a window attention feature map and a window attention module for obtaining a channel attention weight;

[0010] Obtain sample images; the sample images include sample images in a training set and sample images in a test set; train the first image encoding and decoding model based on the sample images in the training set, and verify the effect of the trained first image encoding and decoding model based on the sample images in the test set, to obtain a corresponding image encoding and decoding model pre-embedded with a fusion attention module.

[0011] Furthermore, the initial image coding and decoding model further includes a super-a priori coding network and a super-a priori decoding network; the super-a priori coding network is used to perform super-a priori coding processing on the input potential representation to obtain a corresponding super-a priori potential representation; the super-a priori decoding network is used to process the acquired arithmetic decoded super-a priori potential representation to obtain a super-a priori reconstruction feature;

[0012] After embedding the fused attention module into corresponding target positions in the image encoding network and the image decoding network respectively, the method further includes: embedding the fused attention module into corresponding target positions in the super a priori encoding network and the super a priori decoding network respectively to obtain a second image encoding and decoding model to be trained; training the second image encoding and decoding model based on sample images in the training set, and verifying the effect of the trained second image encoding and decoding model based on sample images in the test set to obtain a corresponding image encoding and decoding model pre-embedded with the fused attention module.

[0013] Furthermore, the initial image encoding and decoding model further includes a context module; the context module is used to extract corresponding context features based on the quantized potential representation;

[0014] After obtaining the context feature and the super a priori reconstruction feature, the method also includes: fusing the context feature and the super a priori reconstruction feature, and processing the fusion result using a linear transformation module to obtain Gaussian distribution parameter information corresponding to the latent representation; and arithmetically encoding and arithmetically decoding the quantized latent representation based on the Gaussian distribution parameter information.

[0015] Further, the image data is input into an image encoding and decoding model pre-embedded in a fusion attention module for image encoding and decoding processing to obtain a compressed target image output by the image encoding and decoding model, specifically including:

[0016] Input the image data into the image coding network pre-embedded with a fusion attention module for coding processing to obtain a potential representation of the image data; input the potential representation into the fusion attention module to obtain a window attention feature map and the channel attention weight; obtain a new potential representation based on the potential representation, the window attention feature map and the channel attention weight;

[0017] Quantizing the new latent representation based on a quantization module to obtain a quantized latent representation;

[0018] Based on the arithmetic coding module and the algorithmic decoding module, respectively, the quantized latent representation is subjected to arithmetic coding processing and algorithmic decoding processing to obtain an arithmetic-decoded latent representation;

[0019] Based on the image decoding network pre-embedded with a fusion attention module, corresponding decoding processing is performed on the potential representation after the arithmetic decoding to obtain image reconstruction features;

[0020] A compressed target image is obtained based on the image reconstruction feature.

[0021] Furthermore, the latent representation is input into the fusion attention module to obtain the window attention feature map and the channel attention weight, specifically including: inputting the latent representation into the channel attention module in the fusion attention module, and obtaining the channel attention feature map based on the convolutional layer and Gelu activation function in the channel attention module; obtaining the channel attention weight based on the sigmoid function in the channel attention module and the channel attention feature map; inputting the latent representation into the window attention module, and obtaining the window attention feature map based on the window attention mechanism submodule, Gelu activation function and shifted window attention mechanism submodule in the window attention module.

[0022] Furthermore, the new latent representation is obtained based on the latent representation, the window attention feature map and the channel attention weight, specifically including: multiplying the window attention feature map with the channel attention weight to obtain the residual of the latent representation; adding the residual to the latent representation to obtain a new latent representation.

[0023] In a second aspect, the present invention further provides an image compression device based on deep learning, comprising:

[0024] An image data determination unit, used to determine image data to be processed;

[0025] An image compression processing unit is used to input the image data into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing, and obtain a compressed target image output by the image encoding and decoding model; wherein the fusion attention module is a processing module that fuses a channel attention mechanism and a window attention mechanism; and the image encoding and decoding model is a deep learning model trained based on a sample image and an image processing result corresponding to the sample image.

[0026] Furthermore, before the image to be processed is input into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, it also includes: a first model construction and training module; the first model construction and training module is specifically used to:

[0027] Determine an initial image coding and decoding model; the initial image coding and decoding model includes an image coding network and an image decoding network; wherein the image coding network is used to encode the input image data to obtain a potential representation of the image data; the image decoding network is used to process the acquired potential representation after arithmetic decoding to obtain image reconstruction features;

[0028] The fusion attention module is respectively embedded into the corresponding target positions in the image encoding network and the image decoding network to obtain a first image encoding and decoding model to be trained; wherein the fusion attention module includes a channel attention module for obtaining a window attention feature map and a window attention module for obtaining a channel attention weight;

[0029] Obtain sample images; the sample images include sample images in a training set and sample images in a test set; train the first image encoding and decoding model based on the sample images in the training set, and verify the effect of the trained first image encoding and decoding model based on the sample images in the test set, to obtain a corresponding image encoding and decoding model pre-embedded with a fusion attention module.

[0030] Furthermore, the initial image coding and decoding model further includes a super-a priori coding network and a super-a priori decoding network; the super-a priori coding network is used to perform super-a priori coding processing on the input potential representation to obtain a corresponding super-a priori potential representation; the super-a priori decoding network is used to process the acquired arithmetic decoded super-a priori potential representation to obtain a super-a priori reconstruction feature;

[0031] After embedding the fused attention module into corresponding target positions in the image encoding network and the image decoding network respectively, the device also includes: a second model construction and training module; the second model construction and training module is specifically used to: embed the fused attention module into corresponding target positions in the super a priori encoding network and the super a priori decoding network respectively to obtain a second image encoding and decoding model to be trained; train the second image encoding and decoding model based on sample images in the training set, and verify the effect of the trained second image encoding and decoding model based on sample images in the test set to obtain a corresponding image encoding and decoding model pre-embedded with the fused attention module.

[0032] Furthermore, the initial image encoding and decoding model further includes a context module; the context module is used to extract corresponding context features based on the quantized potential representation;

[0033] After obtaining the context feature and the super a priori reconstruction feature, the device also includes: a feature fusion module, used to fuse the context feature and the super a priori reconstruction feature, and use a linear transformation module to process the fusion result to obtain Gaussian distribution parameter information corresponding to the potential representation; an arithmetic encoding and decoding module, used to perform arithmetical encoding and arithmetical decoding on the quantized potential representation based on the Gaussian distribution parameter information.

[0034] Furthermore, the image compression processing unit is specifically used for:

[0035] Input the image data into the image coding network pre-embedded with a fusion attention module for coding processing to obtain a potential representation of the image data; input the potential representation into the fusion attention module to obtain a window attention feature map and the channel attention weight; obtain a new potential representation based on the potential representation, the window attention feature map and the channel attention weight;

[0036] Quantizing the new latent representation based on a quantization module to obtain a quantized latent representation;

[0037] Based on the arithmetic coding module and the algorithmic decoding module, respectively, the quantized latent representation is subjected to arithmetic coding processing and algorithmic decoding processing to obtain an arithmetic-decoded latent representation;

[0038] Based on the image decoding network pre-embedded with a fusion attention module, corresponding decoding processing is performed on the potential representation after the arithmetic decoding to obtain image reconstruction features;

[0039] A compressed target image is obtained based on the image reconstruction feature.

[0040] Furthermore, the latent representation is input into the fusion attention module to obtain the window attention feature map and the channel attention weight, specifically including: inputting the latent representation into the channel attention module in the fusion attention module, and obtaining the channel attention feature map based on the convolutional layer and Gelu activation function in the channel attention module; obtaining the channel attention weight based on the sigmoid function in the channel attention module and the channel attention feature map; inputting the latent representation into the window attention module, and obtaining the window attention feature map based on the window attention mechanism submodule, Gelu activation function and shifted window attention mechanism submodule in the window attention module.

[0041] Furthermore, the new latent representation is obtained based on the latent representation, the window attention feature map and the channel attention weight, specifically including: multiplying the window attention feature map with the channel attention weight to obtain the residual of the latent representation; adding the residual to the latent representation to obtain a new latent representation.

[0042] In a third aspect, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the deep learning-based image compression method as described in any one of the above are implemented.

[0043] In a fourth aspect, the present invention further provides a processor-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the deep learning-based image compression method as described in any one of the above are implemented.

[0044] The image compression method based on deep learning provided by the present invention determines the image data to be processed; inputs the image data into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, and obtains the compressed target image output by the image encoding and decoding model; the fusion attention module is a processing module that integrates the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on the sample image and the image processing result corresponding to the sample image. It can better utilize the channel correlation information and the information of the window and the shift window for image compression, improve the rate-distortion performance of deep learning image compression, effectively improve the image compression effect, and have a higher compression rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0046] Figure 1 is a flowchart of an image compression method based on deep learning provided by an embodiment of the present invention;

[0047] Figure 2 It is a schematic diagram of an image encoding process in the training and reasoning test process of an image encoding and decoding model pre-embedded with a fusion attention module provided by an embodiment of the present invention;

[0048] Figure 3 It is a schematic diagram of an image decoding process in the inference test process of an image encoding and decoding model pre-embedded with a fusion attention module provided by an embodiment of the present invention;

[0049] Figure 4 It is a schematic diagram of an image encoding and decoding structure based on deep learning provided by an embodiment of the present invention;

[0050] Figure 5 is a schematic diagram of the structure of a fusion attention module provided by an embodiment of the present invention;

[0051] Figure 6 It is a complete flowchart of the deep learning-based image compression method provided by an embodiment of the present invention;

[0052] Figure 7 is a schematic diagram of the structure of an image compression device based on deep learning provided by an embodiment of the present invention;

[0053] Figure 8 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0055] The following describes in detail the embodiment of the image compression method based on deep learning according to the present invention. Figure 1As shown, it is a flow chart of the image compression method based on deep learning provided by an embodiment of the present invention, and the specific process includes the following steps:

[0056] Step 101: Determine image data to be processed.

[0057] Specifically, the image data is pre-prepared image data to be compressed.

[0058] Step 102: Input the image data into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing to obtain a compressed target image output by the image encoding and decoding model; the fusion attention module is a processing module that fuses the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on sample images and image processing results corresponding to the sample images.

[0059] Specifically, the image data is first input into the image coding network pre-embedded with a fusion attention module for coding processing to obtain a potential representation of the image data; the potential representation is then input into the fusion attention module to obtain a window attention feature map and the channel attention weight; then a new potential representation is obtained based on the potential representation, the window attention feature map and the channel attention weight; further, the new potential representation is quantized based on a quantization module to obtain a quantized potential representation; then, the quantized potential representation is arithmetically encoded and algorithmically decoded based on an arithmetic coding module and an algorithmic decoding module respectively to obtain an arithmetically decoded potential representation; and the arithmetically decoded potential representation is correspondingly decoded based on the image decoding network pre-embedded with a fusion attention module to obtain image reconstruction features; finally, a compressed target image is obtained based on the image reconstruction features.

[0060] like Figure 6 As shown in the figure, for the input image data x, during the encoding process, it is input into the image encoding network to obtain the potential representation y of x, and after quantization and entropy coding (i.e. arithmetic coding) a binary code stream is obtained. The binary code stream is entropy decoded (i.e. arithmetic decoding), dequantized and input into the image decoding network to obtain the decoded image That is the compressed target image. Figure 6The position pointed by the black arrow is the target position. In this application, the fusion attention module can be placed in the middle position (i.e., the target position) between the first GDN (Generalized Divisive Normalization) and the second convNx4x4 / 2↓ from the left in the image coding network 601. At this time, the input feature X input by the fusion attention module is the output of the first GDN from the left; it can also be placed in the image coding network 601 after the last conv Nx4x4 / 2↓ from the left (i.e., the target position) to achieve embedding of the image coding network 601. At this time, the input feature X input by the fusion attention module is the output of the last conv Nx4x4 / 2↓. For example, the input feature X can be the potential representation y of the image data output by the last conv Nx4x4 / 2↓. That is, after the entire image coding network 601, or after the activation function or normalization in the middle of the image coding network 601, the output of the previous submodule inserted at the target position is the input feature X. Correspondingly, the fusion attention module can be placed in the middle position (i.e., the target position) between the first IGDN (Inverse GeneralizedDivisive Normalization) and the second conv Nx4x4 / 2↓ from the left in the image decoding network 602. At this time, the input feature X input to the fusion attention module is the output of the second conv Nx4x4 / 2↓ from the left; it can also be placed in the image decoding network 602 behind the last conv Nx4x4 / 2↓ from the left (i.e., the target position) to achieve embedding in the image decoding network 602. At this time, the input feature X input to the fusion attention module is the output of AD (arithmetic coding module). For example, the input feature X can be the potential representation of the arithmetic coding output of AD. In addition, Q represents the quantization module, AE represents the arithmetic decoding module, Bits represents the bitstream module, and 5x5 mask 2N represents the context module. It should be noted that when the fusion attention module is placed, it needs to be placed in the corresponding positions of the image coding network 601 and the image decoding network 602.

[0061] Among them, the potential representation is input into the fusion attention module to obtain the window attention feature map and the channel attention weight. The corresponding implementation process includes: inputting the potential representation into the channel attention module in the fusion attention module, and obtaining the channel attention feature map based on the convolutional layer and Gelu activation function in the channel attention module; obtaining the channel attention weight based on the sigmoid function in the channel attention module and the channel attention feature map; inputting the potential representation into the window attention module, and obtaining the window attention feature map based on the window attention mechanism submodule, Gelu activation function and shifted window attention mechanism submodule in the window attention module.

[0062] like Figure 5 As shown, the channel attention module 501 (i.e., channel attention mechanism) sequentially includes the first 1×1 convolution layer (i.e., Conv: convolutional neural network), the first Gelu activation function, the 3×3 convolution layer, the second Gelu activation function, the second 1×1 convolution layer, and the sigmoid function (which implements the function of the sigmoid function and can be defined as 1 / (1+e -x )). The window attention module 502 (i.e., the window attention mechanism) includes a window attention mechanism submodule (WindowAttention, WA), a Gelu activation function, a 1×1 convolution layer, a shift window attention submodule (Shift WindowAttention, SWA) and a 3×3 convolution layer. The Gelu activation function is an activation function in a neural network and can be defined as

[0063] In addition, based on the potential representation, the window attention feature map and the channel attention weight, a new potential representation is obtained, and the corresponding implementation process includes: multiplying the window attention feature map with the channel attention weight to obtain the residual of the potential representation; adding the residual to the potential representation to obtain a new potential representation. Specifically, for the original input feature X of the fusion attention module, its dimension representation is (N, C, H, W), where N is the number, C is the channel, H is the height, and W is the width. The original input feature X (at this time, the potential representation y) is retained. Input feature X is input to the channel attention module, after 1×1 convolution (output dimension (N, C / / 2, H, W)), Gelu activation function, 3×3 convolution (input and output dimensions remain unchanged, (N, C / / 2, H, W)), Gelu activation function, 1×1 convolution (input channel number C / / 2, output channel number C), channel attention feature map is obtained, and then after sigmoid, channel attention weight X1 (output dimension (N, C / / 2, H, W)) is obtained. In the entire channel attention module, the input and output dimensions are consistent. Only the size of C changes during the process, and the sizes of H and W remain unchanged. Input feature X is input to the window attention module, after WA window attention mechanism submodule, Gelu activation function, SWA shift window attention mechanism submodule, Gelu activation function, 3×3 convolution to obtain the window attention feature map. The input and output dimensions are consistent throughout the process. In WA, we divide the features into multiple subwindows and calculate the self-attention feature map within the window. In SWA, we shift each window by multiple units and perform self-attention feature maps on the shifted windows. The window size and the shift unit size of the shifted window are pre-set parameters (usually the window size is set to 8 and the shift size is set to 4 or the window size is set to 4 and the shift size is set to 2). Multiply the window attention feature map X2 with the channel attention weight X1 to get the residual X1·X2 of the input feature X. Add the residual X1·X2 to X to get the output of the fusion attention module. The sizes of X1 and X2 are both (N, C, H, W).

[0064] It should be noted that, in the specific implementation process of the present invention, before the image to be processed is input into the image coding and decoding model pre-embedded with the fusion attention module for image coding and decoding processing, it is also necessary to predetermine the initial image coding and decoding model and perform model training. The initial image coding and decoding model includes an image coding network and an image decoding network. Among them, the image coding network is used to encode the input image data to obtain the potential representation of the image data; the image decoding network is used to process the acquired latent representation after arithmetic decoding to obtain image reconstruction features. By embedding the fusion attention module into the corresponding target positions in the image coding network and the image decoding network respectively, the first image coding and decoding model to be trained can be obtained. Among them, the fusion attention module includes a channel attention module for obtaining a window attention feature map and a window attention module for obtaining a channel attention weight. Furthermore, in the model training process, sample images are first obtained; the sample images include sample images in the training set and sample images in the test set; the first image encoding and decoding model is trained based on the sample images in the training set, and the effect of the trained first image encoding and decoding model is verified based on the sample images in the test set to obtain the corresponding image encoding and decoding model pre-embedded in the fusion attention module, that is, the model parameters with the best verification effect are saved to obtain the image encoding and decoding model pre-embedded in the fusion attention module, which can be used to encode and decode the image data to be processed. Specifically, entering the model inference application process, in the image coding network (i.e., the coding network), the image encoding and decoding model pre-embedded in the fusion attention module determined by the saved model parameters with the best verification effect is used for encoding the test set, and the binary code stream file of the encoded output image data is obtained. Figure 3 As shown, in the image decoding network (i.e., decoding network), the input encoded binary code stream file is obtained, and the saved image encoding and decoding model pre-embedded with the fusion attention module is used for decoding, and the decoded image data, i.e., the compressed target image, is output.

[0065] That is, the image encoding and decoding model in the present application solution may include at least two parts: image encoding process and image decoding process. Among them, the image encoding process can be divided into two parts: model training and model reasoning. Figure 2As shown: During the model training process: collect sample images, and divide them into training set, validation set and test set. The training set and validation set can use open source data, such as open-images. The test set is the pictures that need to be tested, and Kodak24 can also be used as a test. During the image encoding process of model inference: load the saved model parameters, that is, use the image encoding and decoding model to encode the image data to be processed on the test set, and output the encoded binary code stream. During the image decoding process of model inference: obtain the input encoded binary code stream, that is, use the trained image encoding and decoding model to decode, and output the decoded target image until the decoding is completed.

[0066] like Figure 4 As shown, in the actual implementation of the present application, in order to further improve the image compression effect based on deep learning, the initial image encoding and decoding model may also include a super-prior encoding network and a super-prior decoding network. The super-prior encoding network is used to perform super-prior encoding processing on the input potential representation to obtain a corresponding super-prior potential representation; the super-prior decoding network is used to process the acquired super-prior potential representation after arithmetic decoding to obtain a super-prior reconstruction feature. After the fusion attention module is respectively embedded in the corresponding target positions in the image encoding network and the image decoding network, the method also includes: embedding the fusion attention module in the corresponding target positions in the super-prior encoding network and the super-prior decoding network, so as to obtain a second image encoding and decoding model to be trained; at this time, the corresponding positions of the image encoding network, the image decoding network, the super-prior encoding network and the super-prior decoding network are all embedded with the fusion attention module. The second image encoding and decoding model is trained based on the sample images in the training set, and the effect of the trained second image encoding and decoding model is verified based on the sample images in the test set to obtain the corresponding image encoding and decoding model pre-embedded with the fusion attention module, and the image encoding and decoding model can also be used to perform subsequent encoding and decoding processing on the image data to be processed. In addition, in order to obtain accurate representation modeling, a context module and a super priori module (i.e., including a super priori encoding network and a super priori decoding network, etc.) are introduced to learn a Gaussian distribution to more accurately simulate the distribution of the potential representation. That is, the initial image encoding and decoding model may also include a context module. The context module is used to extract corresponding context features based on the quantized potential representation. After obtaining the context features and the super priori reconstruction features, the method also includes: fusing the context features and the super priori reconstruction features, and processing the fusion results using a linear transformation module to obtain Gaussian distribution parameter information corresponding to the potential representation, i.e., N(μ, δ 2); through the arithmetic coding module and the arithmetic decoding module, the quantized potential representation is arithmetically encoded and arithmetically decoded respectively based on the Gaussian distribution parameter information. Furthermore, it should be noted that in the specific implementation process, the fusion attention module can also be simply embedded in the corresponding position of the super prior encoding network and the super prior decoding network to obtain the corresponding image encoding and decoding model pre-embedded with the fusion attention module, and the model is trained, and then the optimal image encoding and decoding model pre-embedded with the fusion attention module obtained by training is used to compress the image data. The specific process will not be described in detail here, and the above specific description can be referred to.

[0067] Figure 6 , the position pointed by the black arrow is the target position. The fused attention module can be placed in the middle position between LeakyRelu and the second conv Nx4x4 / 2↓ from the left in the super-prior encoding network 603 (i.e., the target position). At this time, the input feature X input by the fused attention module is the output of LeakyRelu, so as to embed the super-prior encoding network 603. Correspondingly, the fused attention module can be placed in the middle position between LeakyRelu and the first conv Nx4x4 / 2↓ from the right in the super-prior decoding network 604 (i.e., the target position). At this time, the input feature X input by the fused attention module is the output of the first conv Nx4x4 / 2↓, so as to embed the super-prior decoding network 604. Figure 6 yes Figure 4 A schematic diagram showing the specific expansion of the parts contained in each corresponding module.

[0068] like Figure 6 As shown, in a complete embodiment, firstly, the input image data x is obtained. The potential representation y of the original input image (i.e., image data x) is obtained through the image coding network 601. The potential representation y after quantization is obtained through the quantization module Q. The compressed binary bit stream is obtained through the arithmetic encoding module (AE in the figure). The binary bit stream is passed through the arithmetic decoding module (AD in the figure) to obtain the potential representation after arithmetic decoding. The arithmetic encoding module and the arithmetic decoding module are lossless compression. Read the potential representation after arithmetic decoding and pass it through the decoding network module to obtain the image reconstruction feature The quantized value of the potential representation After the context module (5×5mask 2N in the figure), we get The potential representation y is input into the super prior coding network (super prior coding module) to obtain the super prior potential representation z. z is quantized to obtain the quantized potential representation Quantized latent representation After AE, the binary bit stream of the super-prior potential representation is obtained. The binary bit stream of the super-prior potential representation is arithmetically decoded to obtain the super-prior features after arithmetically decoding. Input it into the super-super-prior decoding network (super-prior decoding module) to obtain the super-prior reconstruction features. The context features and super-prior reconstruction features are fused, and then linearly transformed to obtain the output potential features. Based on the Gaussian distribution parameter information, arithmetic encoding and arithmetic decoding are performed ( Figure 6 AE and AD on the left in the middle).

[0069] The fused attention module that fuses the channel attention mechanism and the window attention mechanism introduced in this application can be plug-and-play and can be nested into any image compression network (that is, an image encoding and decoding model that at least includes an image encoding network and an image decoding network), effectively improving the rate-distortion performance of deep learning image compression. Under the same settings, the use of this fused attention module can effectively improve the PSNR and MS-SSIM indicators of the reconstructed image under the same BPP conditions compared to not using it. Among them, BPP: Bits Per Pixel, the number of bits per pixel, the average number of bits required to encode the color information of each pixel. PSNR: Peak Signal-to-Noise Ratio, peak signal-to-noise ratio, is an objective indicator for measuring the quality of image reconstruction, defined as Among them MAX I It is the maximum value representing the color of the image, and MSE is the mean square error between the original image and the reconstructed image. The unit of PSNR is decibel (dB). MS-SSIM: Multi-Scale Structural Similarity, a method for measuring the similarity between two images based on multi-scale (the images are scaled from large to small according to certain rules), is defined as l stands for brightness, c stands for contrast, and s stands for structure.

[0070] In addition, based on this application, similar effects can be achieved by modifying the parameters in the fusion attention module, such as the number of convolutional layers, the number of channels, replacing activation functions, etc., and the weight ratio of the channel attention mechanism and the window attention mechanism can be adjusted to achieve the purpose of image compression. The specific implementation process will not be described in detail here.

[0071] The deep learning-based image compression method described in the embodiment of the present invention determines the image data to be processed; inputs the image data into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, and obtains the compressed target image output by the image encoding and decoding model; the fusion attention module is a processing module that integrates the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on the sample image and the image processing results corresponding to the sample image. It can better utilize the channel correlation information and the information of the window and the shift window for image compression, improve the rate-distortion performance of deep learning image compression, and effectively improve the image compression effect.

[0072] Corresponding to the above-mentioned image compression method based on deep learning, the present invention also provides an image compression device based on deep learning. Since the embodiment of the device is similar to the above-mentioned method embodiment, the description is relatively simple. For relevant parts, please refer to the description of the above-mentioned method embodiment. The embodiment of the image compression device based on deep learning described below is only illustrative. Please refer to Figure 7 As shown, it is a structural schematic diagram of an image compression device based on deep learning provided by an embodiment of the present invention.

[0073] The image compression device based on deep learning described in the present invention specifically includes the following parts:

[0074] An image data determination unit 701, used to determine image data to be processed;

[0075] The image compression processing unit 702 is used to input the image data into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing, and obtain a compressed target image output by the image encoding and decoding model; wherein the fusion attention module is a processing module that fuses the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on a sample image and the image processing results corresponding to the sample image.

[0076] Furthermore, before the image to be processed is input into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, it also includes: a first model construction and training module; the first model construction and training module is specifically used to:

[0077] Determine an initial image coding and decoding model; the initial image coding and decoding model includes an image coding network and an image decoding network; wherein the image coding network is used to encode the input image data to obtain a potential representation of the image data; the image decoding network is used to process the acquired potential representation after arithmetic decoding to obtain image reconstruction features;

[0078] The fusion attention module is respectively embedded into the corresponding target positions in the image encoding network and the image decoding network to obtain a first image encoding and decoding model to be trained; wherein the fusion attention module includes a channel attention module for obtaining a window attention feature map and a window attention module for obtaining a channel attention weight;

[0079] Obtain sample images; the sample images include sample images in a training set and sample images in a test set; train the first image encoding and decoding model based on the sample images in the training set, and verify the effect of the trained first image encoding and decoding model based on the sample images in the test set, to obtain a corresponding image encoding and decoding model pre-embedded with a fusion attention module.

[0080] Furthermore, the initial image coding and decoding model further includes a super-a priori coding network and a super-a priori decoding network; the super-a priori coding network is used to perform super-a priori coding processing on the input potential representation to obtain a corresponding super-a priori potential representation; the super-a priori decoding network is used to process the acquired arithmetic decoded super-a priori potential representation to obtain a super-a priori reconstruction feature;

[0081] After embedding the fused attention module into corresponding target positions in the image encoding network and the image decoding network respectively, the device also includes: a second model construction and training module; the second model construction and training module is specifically used to: embed the fused attention module into corresponding target positions in the super a priori encoding network and the super a priori decoding network respectively to obtain a second image encoding and decoding model to be trained; train the second image encoding and decoding model based on sample images in the training set, and verify the effect of the trained second image encoding and decoding model based on sample images in the test set to obtain a corresponding image encoding and decoding model pre-embedded with the fused attention module.

[0082] Furthermore, the initial image encoding and decoding model further includes a context module; the context module is used to extract corresponding context features based on the quantized potential representation;

[0083] After obtaining the context feature and the super a priori reconstruction feature, the device also includes: a feature fusion module, used to fuse the context feature and the super a priori reconstruction feature, and use a linear transformation module to process the fusion result to obtain Gaussian distribution parameter information corresponding to the latent representation; an arithmetic encoding and decoding module, used to perform arithmetical encoding and arithmetical decoding on the quantized latent representation based on the Gaussian distribution parameter information.

[0084] Furthermore, the image compression processing unit is specifically used for:

[0085] Input the image data into the image coding network pre-embedded with a fusion attention module for coding processing to obtain a potential representation of the image data; input the potential representation into the fusion attention module to obtain a window attention feature map and the channel attention weight; obtain a new potential representation based on the potential representation, the window attention feature map and the channel attention weight;

[0086] Quantizing the new latent representation based on a quantization module to obtain a quantized latent representation;

[0087] Based on the arithmetic coding module and the algorithmic decoding module, respectively, the quantized latent representation is subjected to arithmetic coding processing and algorithmic decoding processing to obtain an arithmetic-decoded latent representation;

[0088] Based on the image decoding network pre-embedded with a fusion attention module, corresponding decoding processing is performed on the potential representation after the arithmetic decoding to obtain image reconstruction features;

[0089] A compressed target image is obtained based on the image reconstruction feature.

[0090] Furthermore, the latent representation is input into the fusion attention module to obtain the window attention feature map and the channel attention weight, specifically including: inputting the latent representation into the channel attention module in the fusion attention module, and obtaining the channel attention feature map based on the convolutional layer and Gelu activation function in the channel attention module; obtaining the channel attention weight based on the sigmoid function in the channel attention module and the channel attention feature map; inputting the latent representation into the window attention module, and obtaining the window attention feature map based on the window attention mechanism submodule, Gelu activation function and shifted window attention mechanism submodule in the window attention module.

[0091] Furthermore, the new latent representation is obtained based on the latent representation, the window attention feature map and the channel attention weight, specifically including: multiplying the window attention feature map with the channel attention weight to obtain the residual of the latent representation; adding the residual to the latent representation to obtain a new latent representation.

[0092] The deep learning-based image compression device described in the embodiment of the present invention determines the image data to be processed; inputs the image data into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, and obtains the compressed target image output by the image encoding and decoding model; the fusion attention module is a processing module that integrates the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on the sample image and the image processing results corresponding to the sample image. It can better utilize the channel correlation information and the information of the window and the shift window for image compression, improve the rate-distortion performance of deep learning image compression, and effectively improve the image compression effect.

[0093] Corresponding to the above-mentioned deep learning-based image compression method, the present invention also provides an electronic device. Since the embodiment of the electronic device is similar to the above-mentioned method embodiment, the description is relatively simple. For relevant parts, please refer to the description of the above-mentioned method embodiment. The electronic device described below is only exemplary. Figure 8 As shown, it is a schematic diagram of the physical structure of an electronic device disclosed in an embodiment of the present invention. The electronic device may include: a processor 801, a memory 802 and a communication bus 803, wherein the processor 801 and the memory 802 complete mutual communication through the communication bus 803, and communicate with the outside through the communication interface 804. The processor 801 can call the logic instructions in the memory 802 to execute the image compression method based on deep learning, the method comprising: determining the image data to be processed; inputting the image data into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, and obtaining the compressed target image output by the image encoding and decoding model; wherein the fusion attention module is a processing module that fuses the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on the sample image and the image processing result corresponding to the sample image.

[0094] In addition, the logic instructions in the above-mentioned memory 802 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a memory chip, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0095] On the other hand, an embodiment of the present invention further provides a computer program product, the computer program product includes a computer program stored on a processor-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the image compression method based on deep learning provided by the above-mentioned method embodiments. The method includes: determining the image data to be processed; inputting the image data into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing, and obtaining a compressed target image output by the image encoding and decoding model; wherein the fusion attention module is a processing module that fuses a channel attention mechanism and a window attention mechanism; the image encoding and decoding model is a deep learning model trained based on a sample image and an image processing result corresponding to the sample image.

[0096] On the other hand, an embodiment of the present invention further provides a processor-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, it is implemented to execute the image compression method based on deep learning provided in the above embodiments. The method includes: determining the image data to be processed; inputting the image data into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing, and obtaining a compressed target image output by the image encoding and decoding model; wherein the fusion attention module is a processing module that fuses the channel attention mechanism and the window attention mechanism; the image encoding and decoding model is a deep learning model trained based on a sample image and the image processing result corresponding to the sample image.

[0097] The processor-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid-state drive (SSD)), etc.

[0098] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0099] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A deep learning-based image compression method, characterized in that: include: determining image data to be processed; The image data is input into an image encoding and decoding model pre-embedded with a fusion attention module for image encoding and decoding processing, and a compressed target image output by the image encoding and decoding model is obtained; wherein the fusion attention module is a processing module that fuses a channel attention mechanism and a window attention mechanism; and the image encoding and decoding model is a deep learning model trained based on a sample image and an image processing result corresponding to the sample image; Before the image to be processed is input into the image encoding and decoding model pre-embedded with the fusion attention module for image encoding and decoding processing, it also includes: Determine an initial image coding and decoding model; the initial image coding and decoding model includes an image coding network and an image decoding network; wherein the image coding network is used to encode the input image data to obtain a potential representation of the image data; the image decoding network is used to process the acquired potential representation after arithmetic decoding to obtain image reconstruction features; The fusion attention module is respectively embedded into the corresponding target positions in the image encoding network and the image decoding network to obtain a first image encoding and decoding model to be trained; wherein the fusion attention module includes a channel attention module for obtaining a window attention feature map and a window attention module for obtaining a channel attention weight; Obtain sample images; the sample images include sample images in a training set and sample images in a test set; train the first image encoding and decoding model based on the sample images in the training set, and verify the effect of the trained first image encoding and decoding model based on the sample images in the test set, to obtain a corresponding image encoding and decoding model pre-embedded with a fusion attention module.

2. The image compression method based on deep learning according to claim 1, characterized in that: The initial image coding and decoding model also includes a super-a priori coding network and a super-a priori decoding network; the super-a priori coding network is used to perform super-a priori coding processing on the input potential representation to obtain a corresponding super-a priori potential representation; the super-a priori decoding network is used to process the acquired arithmetic decoded super-a priori potential representation to obtain a super-a priori reconstruction feature; After embedding the fused attention module into corresponding target positions in the image encoding network and the image decoding network respectively, the method further includes: embedding the fused attention module into corresponding target positions in the super a priori encoding network and the super a priori decoding network respectively to obtain a second image encoding and decoding model to be trained; training the second image encoding and decoding model based on sample images in the training set, and verifying the effect of the trained second image encoding and decoding model based on sample images in the test set to obtain a corresponding image encoding and decoding model pre-embedded with the fused attention module.

3. The image compression method based on deep learning according to claim 2, characterized in that: The initial image encoding and decoding model also includes a context module; the context module is used to extract corresponding context features based on the quantized potential representation; After obtaining the context feature and the super a priori reconstruction feature, the method also includes: fusing the context feature and the super a priori reconstruction feature, and processing the fusion result using a linear transformation module to obtain Gaussian distribution parameter information corresponding to the latent representation; and arithmetically encoding and arithmetically decoding the quantized latent representation based on the Gaussian distribution parameter information.

4. The image compression method based on deep learning according to claim 1, characterized in that: Inputting the image data into an image encoding and decoding model pre-embedded in a fusion attention module for image encoding and decoding processing, and obtaining a compressed target image output by the image encoding and decoding model, specifically comprising: Input the image data into the image coding network pre-embedded with a fusion attention module for coding processing to obtain a potential representation of the image data; input the potential representation into the fusion attention module to obtain a window attention feature map and the channel attention weight; obtain a new potential representation based on the potential representation, the window attention feature map and the channel attention weight; Quantizing the new latent representation based on a quantization module to obtain a quantized latent representation; Based on the arithmetic coding module and the algorithmic decoding module, respectively, the quantized latent representation is subjected to arithmetic coding processing and algorithmic decoding processing to obtain an arithmetic-decoded latent representation; Based on the image decoding network pre-embedded with a fusion attention module, corresponding decoding processing is performed on the potential representation after the arithmetic decoding to obtain image reconstruction features; A compressed target image is obtained based on the image reconstruction feature.

5. The image compression method based on deep learning according to claim 4, characterized in that: The potential representation is input into the fusion attention module to obtain the window attention feature map and the channel attention weight, specifically including: inputting the potential representation into the channel attention module in the fusion attention module, and obtaining the channel attention feature map based on the convolutional layer and Gelu activation function in the channel attention module; obtaining the channel attention weight based on the sigmoid function in the channel attention module and the channel attention feature map; inputting the potential representation into the window attention module, and obtaining the window attention feature map based on the window attention mechanism submodule, Gelu activation function and shifted window attention mechanism submodule in the window attention module.

6. The image compression method based on deep learning according to claim 4, characterized in that: The method of obtaining a new latent representation based on the latent representation, the window attention feature map and the channel attention weight specifically includes: multiplying the window attention feature map with the channel attention weight to obtain a residual of the latent representation; and adding the residual to the latent representation to obtain a new latent representation.

7. An image compression device based on deep learning, characterized in that: include: An image data determination unit, used to determine image data to be processed; The image compression processing unit is used to input the image data into the image coding and decoding model pre-embedded in the fusion attention module for image coding and decoding processing, and obtain the compressed target image output by the image coding and decoding model; wherein the fusion attention module is a processing module that fuses the channel attention mechanism and the window attention mechanism; the image coding and decoding model is a deep learning model trained based on the sample image and the image processing result corresponding to the sample image; before the image to be processed is input into the image coding and decoding model pre-embedded in the fusion attention module for image coding and decoding processing, it also includes: a first model construction and training module; the first model construction and training module is specifically used to: Determine an initial image coding and decoding model; the initial image coding and decoding model includes an image coding network and an image decoding network; wherein the image coding network is used to encode the input image data to obtain a potential representation of the image data; the image decoding network is used to process the acquired potential representation after arithmetic decoding to obtain image reconstruction features; The fusion attention module is respectively embedded into the corresponding target positions in the image encoding network and the image decoding network to obtain a first image encoding and decoding model to be trained; wherein the fusion attention module includes a channel attention module for obtaining a window attention feature map and a window attention module for obtaining a channel attention weight; Obtain sample images; the sample images include sample images in a training set and sample images in a test set; train the first image encoding and decoding model based on the sample images in the training set, and verify the effect of the trained first image encoding and decoding model based on the sample images in the test set, to obtain a corresponding image encoding and decoding model pre-embedded with a fusion attention module.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the deep learning-based image compression method as described in any one of claims 1 to 6 are implemented.

9. A processor-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the deep learning-based image compression method as described in any one of claims 1 to 6 are implemented.