Image decoding method and device

Through deep learning-based image encoding and decoding technology, the image is extracted and reconstructed by n-level cascaded reversible modules and noise reduction modules, which solves the problem of low traditional image encoding quality and realizes high-quality image reconstruction.

CN120358359APending Publication Date: 2025-07-22HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410047695.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The traditional image encoding method has the problem of poor encoding effect and low quality of images reconstructed after encoding.

Method used

Using deep learning-based image encoding and decoding technology, n-level cascaded reversible modules and noise reduction modules are used for feature extraction and reconstruction. By decompressing and reverse processing of the encoded data of the target image, quantization noise is filtered to improve the quality of image reconstruction.

Benefits of technology

Even if there is a big difference between the target image and the training data set, it can effectively improve the quality of the reconstructed image after encoding, reduce quantization noise, and improve image reconstruction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358359A_ABST
    Figure CN120358359A_ABST
Patent Text Reader

Abstract

The invention provides an image decoding method and device, and relates to the technical field of image coding and decoding. Comprising the steps that coded data of a target image are acquired, the coded data of the target image are data acquired by compressing target features, and the target features are features acquired by performing forward processing on the target image through n-level cascaded reversible modules; performing decompression processing on the coded data of the target image to obtain reconstruction features of the target image; processing the reconstruction features of the target image through a reconstruction network composed of n stages of cascaded reconstruction units to obtain a reconstructed image of the target image; the ith-level reconstruction unit of the reconstruction network is composed of an (n-i + 1) th-level reversible module and a noise reduction module, and is used for carrying out the reverse processing of the input features of the ith-level reconstruction unit through the (n-i + 1) th-level reversible module, and filtering the quantization noise through the noise reduction module, so as to obtain the output features. Some embodiments of the invention are used for improving the quality of an image reconstructed after coding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Some embodiments of the present application relate to the field of image encoding and decoding technologies. More specifically, it relates to an image decoding method and apparatus. Background Art

[0002] With the development of the information age, in the field of the Internet, people obtain information by watching videos or images, and image transmission has also become an important communication method. With the increase in data volume and information volume, the physical space occupied by videos or images is also getting larger and larger; if unprocessed videos or images are directly transmitted over the network, it will occupy a large amount of network bandwidth and consume a large amount of traffic. Therefore, videos or images are compressed before transmission to reduce their volume. However, traditional image coding methods have problems such as poor coding effects and low quality of the reconstructed images after coding. Summary of the Invention

[0003] Exemplary embodiments of the present application provide an image decoding method and apparatus for improving the quality of the reconstructed images after coding.

[0004] Some embodiments of the present application provide the following technical solutions:

[0005] In a first aspect, some embodiments of the present application provide an image decoding method, including:

[0006] Obtaining encoded data of a target image, where the encoded data of the target image is data obtained by compressing a target feature, and the target feature is a feature obtained by performing forward processing on the target image through an n-stage cascaded reversible module. The compression processing includes quantization processing, and n is a positive integer;

[0007] Performing decompression processing on the encoded data of the target image to obtain the reconstructed feature of the target image;

[0008] Processing the reconstructed feature of the target image through a reconstruction network composed of an n-stage cascaded reconstruction unit to obtain the reconstructed image of the target image; the i-th reconstruction unit of the reconstruction network is composed of the (n - i + 1)-th reversible module and a noise reduction module, and is used to perform reverse processing on the input feature of the i-th reconstruction unit through the (n - i + 1)-th reversible module to obtain the initial feature of the i-th reconstruction unit, and filtering out the quantization noise in the initial feature of the i-th reconstruction unit through the noise reduction module to obtain the output feature of the i-th reconstruction unit. The quantization noise is the noise brought by the quantization processing, and i is a positive integer less than or equal to n.

[0009] In a second aspect, some embodiments of the present application provide an image decoding apparatus, including:

[0010] A memory configured to store a computer program;

[0011] A processor configured to, when calling the computer program, cause the image decoding device to implement the image decoding method described in the second aspect.

[0012] In a third aspect, some embodiments of the present application provide a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a computing device, the computing device is caused to implement the image decoding method described in the first aspect.

[0013] In a fourth aspect, some embodiments of the present application provide a computer program product. When the computer program product runs on a computer, the computer is caused to implement the image decoding method described in the first aspect.

[0014] As can be seen from the above technical solutions, the image decoding method provided by the embodiments of the present application, when obtaining the encoded data of the target image, first decompresses the encoded data of the target image to obtain the reconstruction features of the target image, and then processes the reconstruction features of the target image through a reconstruction network composed of n-level cascaded reconstruction units to obtain the reconstructed image of the target image. Since the i-th level reconstruction unit of the reconstruction network includes a noise reduction module, the noise reduction module can filter out the quantization noise brought by the quantization process in the features and reduce the difference between the target features and the reconstruction features. In addition, the encoded data of the target image is the data obtained by compressing the target features, the target features are the features obtained by forward processing the target image through n-level cascaded reversible modules, the i-th level reconstruction unit of the reconstruction network includes the (n - i + 1)-th level reversible module, and the input features of the i-th level reconstruction unit can be reversely processed through the (n - i + 1)-th level reversible module. Therefore, the embodiments of the present application can use the reversible module to extract features of the target image and reconstruct the target image according to the reconstruction features. Even if there are large differences between the target image and the training data set, the target features can be well extracted and the target image can be reconstructed according to the reconstruction features. Therefore, the embodiments of the present application can improve the quality of the reconstructed image after encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate some embodiments of the present application or the implementation manners in related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.

[0016] Figure 1 Shows a schematic structural diagram of an image encoder in some embodiments of the present application;

[0017] Figure 2 Shows a schematic structural diagram of a reversible module in some embodiments of the present application;

[0018] Figure 3 Shows a schematic structural diagram of a feature compression module in some embodiments of the present application;

[0019] Figure 4 Shows one of the schematic structural diagrams of an image decoder in some other embodiments of the present application;

[0020] Figure 5 Shows another schematic structural diagram of an image decoder in some embodiments of the present application;

[0021] Figure 6 Shows a schematic structural diagram of a noise reduction module in some embodiments of the present application;

[0022] Figure 7 Shows a schematic structural diagram of a noise reduction module in some other embodiments of the present application;

[0023] Figure 8 Shows a schematic structural diagram of a feature decompression module in some embodiments of the present application;

[0024] Figure 9 Shows a schematic diagram of an encoding and decoding framework based on a reversible neural network in some embodiments of the present application;

[0025] Figure 10 Shows one of the step flowcharts of an image decoding method in some embodiments of the present application;

[0026] Figure 11 Shows another step flowchart of an image decoding method in some embodiments of the present application;

[0027] Figure 12 Shows the third step flowchart of an image decoding method in some embodiments of the present application. Detailed implementation manners

[0028] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described embodiments. To make the purpose and embodiments of the present application clearer, the following will clearly and completely describe the exemplary embodiments of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0029] It is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.

[0030] The terms "comprising", "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components need not be limited to all the components clearly listed, but may include other components not clearly listed or inherent to such products or devices.

[0031] References in the specification to "some implementations", "some embodiments", etc. indicate that the described implementations or embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes that specific feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same implementation. Additionally, when a specific feature, structure, or characteristic is described in connection with an embodiment, it is considered within the knowledge of those skilled in the art to implement such feature, structure, or characteristic in connection with other implementations (whether explicitly described herein or not).

[0032] Traditional image coding methods have problems such as poor coding effect and low image quality after reconstruction. To improve the problems of poor coding effect and low image quality after reconstruction of traditional image coding methods, image coding and decoding technologies based on deep learning are developing rapidly. For example, the JPEG AI standard created by the Joint Photographic Experts Group (JPEG) is an extensible image coding standard based on machine learning. The general coding process of image coding and decoding technologies based on deep learning is as follows: First, the feature extraction network extracts features from the image to be coded, and then performs compression processing such as quantization on the features of the image to be coded extracted by the feature extraction network to obtain the coded data of the image to be coded. Correspondingly, the general decoding process of image coding and decoding technologies based on deep learning is as follows: First, the coded data of the image to be decoded is decompressed to obtain the features of the reconstructed image to be decoded, and then the image to be decoded is reconstructed through the image reconstruction network and the features of the image to be decoded. Among them, the feature extraction network and the image reconstruction network are generally network models obtained by stacking traditional convolutional neural networks and training on a specific data set. If the image to be coded and decoded has a large difference from the training data, it is difficult to extract features from this image with high quality or reconstruct this image based on the features. To further solve this problem, relevant technologies have proposed to perform feature extraction and feature-based image reconstruction through reversible neural networks in the process of image coding and decoding based on deep learning. Although performing feature extraction and feature-based image reconstruction through reversible neural networks can solve the problem that it is difficult to extract features from this image with high quality or reconstruct this image based on the features caused by the large difference between the image to be coded and decoded and the training data, there will be lossy compression processing such as quantization in the process of compressing the features of the image, and certain noise may be introduced during decompression. Therefore, the performance of image coding and decoding still needs to be further improved.

[0033] The images in the embodiments of the present application can be independent images or video frames in a video. Since a video can be regarded as a sequence of video frames composed of multiple images, when performing encoding and decoding on each video frame based on the image encoding method and the image decoding method, the encoding and decoding of the video can be performed based on the image encoding method and the image decoding method.

[0034] The embodiments of the present application relate to the technical field of image encoding and decoding. First, the image encoding and decoding framework provided by the embodiments of the present application will be described below.

[0035] Refer to Figure 1 As shown, the image encoder provided by some embodiments of the present application includes: a feature extraction module 11 and a feature compression module 12.

[0036] Among them, the feature extraction module 11 is a reversible neural network composed of n - stage cascaded reversible modules, which is used to extract features from the target image X to obtain the target feature F X , and the feature compression module 12 is used to perform compression processing on the target feature to obtain the encoded data B of the target image.

[0037] Refer to Figure 2 As shown, in some embodiments, the feature extraction module 11 is composed of 4 - stage cascaded reversible modules. Any stage of the reversible module includes: a channel recombination module (Channel Recombination Module) 111, a sixth convolutional layer 112, a first coupling layer (Coupling Layer) 113, a second coupling layer 114, and a third coupling layer 115 connected in series in sequence. That is, the reversible module is used to process the input feature I1 of the reversible module through the channel recombination module 111, the sixth convolutional layer 112, the first coupling layer 113, the second coupling layer 114, and the third coupling layer 115 in sequence to obtain the output feature O1 of the reversible module. Among them, the channel recombination module 111 is used to perform downsampling on the input feature I1 of the reversible module to reduce the number of feature channels, reduce the data volume, and extract features. The channel recombination module 111 can be composed of sub - modules such as a max - pooling layer, an average - pooling layer, and a convolutional layer. These sub - modules can perform downsampling on the input signal and perform recombination in the channel dimension. The coupling layer (the first coupling layer 113, the second coupling layer 114, and the third coupling layer 115) is used to fuse and interact the information between different channels or feature maps. The coupling layer can perform element - wise multiplication and addition operations in the channel dimension to increase the expression ability and flexibility of the model.

[0038] In some embodiments, the feature compression module 12 is a feature compression module based on a JPEG AI encoder.

[0039] Refer toFigure 3 As shown, in some embodiments, the feature compression module 12 includes: a residual unit 301. The residual unit 301 is used to calculate the difference between the target feature F X and the predicted feature μ to obtain a residual feature R. Wherein, the predicted feature μ is the predicted information of the target feature F obtained by the prediction fusion network based on the context feature output by the context network and the reconstructed latent feature of the hyperprior decoding network. The input of the context network is the target feature F X that has been encoded X of the elements X .

[0040] Referring to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a gain unit 302. The gain unit 302 is used to perform a gain process on the residual feature R to obtain a gain-processed residual feature Rg

[0041] Referring to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a first quantization unit 303. The first quantization unit 303 is used to perform a quantization process on the gain-processed residual feature Rg to obtain a quantized residual feature r

[0042] Referring to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a hyperprior encoding network 304. The hyperprior encoding network 304 is used to perform hyperprior encoding on the target feature F X to obtain a hyper-latent feature Z, and extract additional auxiliary information during the hyperprior encoding process, so that an accurate probability model for entropy encoding can be obtained based on the encoding result

[0043] Referring to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a second quantization unit 305. The second quantization unit 305 is used to perform a quantization process on the hyper-latent feature Z to obtain a quantized hyper-latent feature z

[0044] Referring to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a decomposed entropy model 306. The decomposed entropy model 306 is used to obtain a cumulative distribution function cdf (Cumulative Distribution Function, CDF)

[0045] Referring to Figure 3As shown, in some embodiments, the feature compression module 12 further includes: a first lossless encoder 307. The first lossless encoder 307 is configured to perform lossless encoding on the quantized hyper-latent feature z according to the cumulative distribution function cdf to obtain first encoded data b1.

[0046] Refer to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a lossless decoder 308. The lossless decoder 308 is configured to decode the first encoded data b1 according to the cumulative distribution function cdf to obtain the reconstructed hyper-latent feature

[0047] Refer to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a hyperprior scale encoding network 309. The hyperprior scale encoding network 309 is configured to process the reconstructed hyper-latent feature to obtain the variance N(0,σ) of a Gaussian distribution with a probability distribution mean of 0.

[0048] Refer to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a second lossless encoder 310. The second lossless encoder 310 is configured to perform lossless encoding on the quantized residual feature r based on the variance N(0,σ) to obtain second encoded data b2.

[0049] Refer to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a hyperprior decoding network 311. The hyperprior decoding network 311 is configured to perform decoding processing on the reconstructed hyper-latent feature to obtain the reconstructed target feature

[0050] Refer to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: an inverse gain unit 312. The inverse gain unit 312 is configured to perform inverse gain processing on the quantized residual feature r to obtain the inverse gain processed residual feature rig.

[0051] Refer to Figure 3 As shown, in some embodiments, the feature compression module 12 further includes: a context network 314. The context network 314 is configured to perform context feature extraction on the fused latent feature y to obtain the context feature cont of the fused latent feature y.

[0052] Refer to Figure 3As shown, in some embodiments, the feature compression module 12 further includes: a prediction fusion network 315. The prediction fusion network 315 is configured to obtain a predicted feature μ of the next element of the target feature F based on the context feature cont and the reconstructed prior feature and perform encoding of the next element based on the predicted feature μ. X

[0053] Referring to Figure 4 As shown, the image decoder for implementing some embodiments of the present application includes: a feature decompression module 41 and a reconstruction network 42 composed of n cascaded reconstruction units. The i-th reconstruction unit of the reconstruction network 42 is composed of the (n - i + 1)-th reversible module and a noise reduction module of the feature extraction module of the image encoder. The feature decompression module 41 performs decompression processing on the encoded data B of the target image to obtain the reconstructed feature of the target image The reconstruction network 42 is configured to process the reconstructed feature of the target image to obtain the reconstructed image of the target image X The i-th reconstruction unit is configured to perform reverse processing on the input feature F of the i-th reconstruction unit through the (n - i + 1)-th reversible module to obtain the initial feature F of the i-th reconstruction unit i in and filter out the quantization noise in the initial feature F of the i-th reconstruction unit through the noise reduction module to obtain the output feature F of the i-th reconstruction unit i 0 where the quantization noise is the noise brought by the quantization process, and i is a positive integer less than or equal to n. i 0 i out

[0054] Referring to Figure 5 As shown, in some embodiments, the reconstruction network 42 is composed of 4 cascaded reconstruction units. The first reconstruction unit of the reconstruction network 42 is composed of the fourth reversible module and a noise reduction module of the feature extraction module of the image decoder. The second reconstruction unit of the reconstruction network 42 is composed of the third reversible module and a noise reduction module of the feature extraction module of the image decoder. The third reconstruction unit of the reconstruction network 42 is composed of the second reversible module and a noise reduction module of the feature extraction module of the image decoder. The fourth reconstruction unit of the reconstruction network 42 is composed of the first reversible module and a noise reduction module of the feature extraction module of the image decoder.

[0055] Referring to Figure 6 ​​​As shown, in some embodiments, the noise reduction module includes: a first activation function layer 61, a first convolutional layer 62, a second activation function layer 63, a second convolutional layer 64, a first adjustment layer 65, and a first addition fusion layer 66. The feature processing process of the noise reduction module includes the following steps 61 to 66: Step 61, the input feature F of the noise reduction module is activated by the first activation function layer 61 in to obtain a first feature F1; Step 62, the first feature F1 is convolved by the first convolutional layer 62 to obtain a second feature F2; Step 63, the second feature F2 is activated by the second activation function layer 63 to obtain a third feature F3; Step 64, the third feature F3 is convolved by the second convolutional layer 64 to obtain a fourth feature F4; Step 65, the first adjustment layer 65 adjusts the fourth feature F4 based on the first adjustment coefficient α to obtain a fifth feature F5; Step 66, the first addition fusion layer 66 adds and fuses the input feature F of the noise reduction module in and the fifth feature F5 to obtain the output feature F of the noise reduction module out .

[0056] Among them, the input feature of the noise reduction module of any level of reconstruction unit is the output feature of the reversible module of this reconstruction unit. For example: the input feature of the noise reduction module of the first-level reconstruction unit is the output feature of the reversible module (the nth-level reversible module) of the first-level reconstruction unit. For another example: the input feature of the noise reduction module of the third-level reconstruction unit is the output feature of the reversible module (the n-2th-level reversible module) of the third-level reconstruction unit.

[0057] In some embodiments, the first activation function 61 and / or the second activation function 63 is a rectified linear unit (ReLU).

[0058] In some embodiments, the first convolutional layer 62 and / or the second convolutional layer 64 is a convolutional layer with a kernel size of 3*3.

[0059] In some embodiments, the first adjustment coefficient α is 0.2. That is, the first adjustment layer 65 multiplies the feature values of each element of the fourth feature F4 by 0.2 to obtain the fifth feature F5.

[0060] It should also be noted that the first adjustment coefficient α can adjust the proportion of the input feature F in the output feature F of the noise reduction module out of the noise reduction module. When it is necessary to make the output feature F of the noise reduction module in take the input feature F as the main feature, the first adjustment coefficient α can be set smaller, and when it is necessary to make the output feature F of the noise reduction module out take the output feature with the input feature F in as the main feature, the first adjustment coefficient α can be set smaller, and when it is necessary to make the output feature F of the noise reduction moduleout When the output feature takes the fifth feature F5 as the main feature, the first adjustment coefficient α can be set relatively large.

[0061] Refer to Figure 7 As shown, in some other embodiments, the noise reduction module includes: a first dense connection module (DenseConnection Block, DCB) 71, a third convolutional layer 72, a fourth convolutional layer 73, a fifth convolutional layer 74, a second dense connection module 75, a second adjustment layer 76, and a second addition fusion layer 77. The feature processing process of the noise reduction module includes the following steps 71 to step 76: Step 71, process the input feature F of the noise reduction module through the first dense connection module 71 in to obtain the sixth feature F6; Step 72, perform convolutional processing on the sixth feature F6 through the third convolutional layer 72 to obtain the seventh feature F7; Step 73, perform convolutional processing on the seventh feature F7 through the fourth convolutional layer 73 to obtain the eighth feature F8; Step 74, perform convolutional processing on the eighth feature F8 through the fifth convolutional layer 74 to obtain the ninth feature F9; Step 75, process the sixth feature F6, the seventh feature F7, the eighth feature F8, and the ninth feature F9 through the second dense connection module 75 to obtain the tenth feature F 10 ; Step 76, the second adjustment layer 76 adjusts the tenth feature F 10 based on the second adjustment coefficient β to obtain the eleventh feature F 11 ; Step 76, the second addition fusion layer 77 performs addition fusion on the input feature F of the noise reduction module in and the eleventh feature F 11 to obtain the output feature F of the noise reduction module out . Among them, the input feature of the noise reduction module of any level of the reconstruction unit is the output feature of the reversible module of this reconstruction unit.

[0062] In some embodiments, the third convolutional layer 72 and / or the fifth convolutional layer 74 is a convolutional layer with a convolutional kernel size of 3*3.

[0063] In some embodiments, the fourth convolutional layer 73 is a convolutional layer with a convolutional kernel size of 1*1.

[0064] In some embodiments, the second adjustment coefficient is 0.2. That is, the second adjustment layer 76 multiplies the feature values of each element of the tenth feature F 10 by 0.2 to obtain the eleventh feature F 11 .

[0065] Similarly, the second adjustment coefficient β can adjust the input feature F in the output feature F of the noise reduction module out ​in proportion, when it is necessary to make the output feature F of the noise reduction module out The output feature takes the input feature F in as the main feature, the second adjustment coefficient β can be set smaller, and when it is necessary to make the output feature F of the noise reduction module out The output feature takes the eleventh feature F 11 as the main feature, the second adjustment coefficient β can be set larger.

[0066] Since the input of the dense connection module is the input features of all the previous layer structures in the network structure before the dense connection block, for example: the input of the second dense connection module 75 includes: the sixth feature F6 output by the first dense connection module 71 before the second dense connection module 75, the seventh feature F7 output by the third convolutional layer 72, the eighth feature F8 output by the fourth convolutional layer 73, and the ninth feature F9 output by the fifth convolutional layer 74. Therefore, the noise reduction module based on the dense connection module can make more full use of the global information to alleviate the problem of gradient disappearance to a certain extent. And since the input of the dense connection module includes not only the output features of the previous network structure, but also the output features of other previous network structures, the robustness of the noise reduction module can also be improved.

[0067] In some embodiments, the feature decompression module 41 is a feature decompression module based on a JPEG AI decoder.

[0068] Referring to Figure 8 shown, in some embodiments, the feature decompression module 41 includes: a decomposition entropy model 801. The decomposition entropy model 801 is used to obtain the cumulative distribution function cdf.

[0069] Referring to Figure 8 shown, in some embodiments, the feature decompression module 41 further includes: a first lossless decoder 802. The first lossless decoder 802 is used to decode the first encoded data b1 according to the cumulative distribution function cdf to obtain the reconstructed super latent feature

[0070] Referring to Figure 8 shown, in some embodiments, the feature decompression module 41 further includes: a hyper prior scale decoding network 803. The hyper prior scale decoding network 803 is used to process the reconstructed super latent feature to obtain the variance of the Gaussian distribution N(0,σ) with a probability distribution of 0 mean.

[0071] Referring to Figure 8As shown, in some embodiments, the feature decompression module 41 further includes: a second lossless decoder 804. The second lossless decoder 804 is configured to decode the second encoded data b2 according to the variance N(0,σ) to obtain the reconstructed quantized residual features

[0072] Refer to Figure 8 As shown, in some embodiments, the feature decompression module 41 further includes: an inverse gain unit 805. The inverse gain unit 805 is configured to perform inverse gain processing on the reconstructed quantized residual features to obtain the residual features after inverse gain processing

[0073] Refer to Figure 8 As shown, in some embodiments, the feature decompression module 41 further includes: a fusion module 806. The fusion module 806 is configured to perform additive fusion on the residual features after inverse gain processing and the reconstructed prediction features to obtain the reconstructed fusion latent features and send them to the reconstruction network 42 for processing to obtain the reconstructed image of the target image X

[0074] Refer to Figure 8 As shown, in some embodiments, the feature decompression module 41 further includes: a hyperprior decoding network 807. The hyperprior decoding network 807 is configured to perform decoding processing on the reconstructed hyper-latent features to obtain the reconstructed target features

[0075] Refer to Figure 8 As shown, in some embodiments, the feature decompression module 41 further includes: a prediction fusion network 808. The prediction fusion network 808 is configured to obtain the reconstructed prediction features according to the reconstructed target features and the context features cont output by the context model 809

[0076] Refer to Figure 8 As shown, in some embodiments, the feature decompression module 41 further includes: a context model 809. The context model 809 is configured to obtain the reconstructed prediction features from the output context features cont The context model 809 is further configured to perform context feature extraction on the reconstructed fusion latent features to obtain the context features cont of the reconstructed fusion latent features for the reconstructed fusion latent features The context feature cont decompresses the next element of the first encoded data b1.

[0077] First, the overall inventive concept of the image decoding method provided by the embodiments of the present application will be introduced.

[0078] The goal of image coding and decoding is to minimize the reconstruction distortion at a certain bit rate. This goal can be transformed into an optimized balanced cost of bit rate and distortion through the Lagrangian algorithm. That is, the following rate-distortion cost:

[0079] J = R + λD (1)

[0080] Among them, R is the bit rate, representing the ratio of the number of bits after encoding to the number of bits of the original image, D is the distortion, representing the mean square error (MSE) of the difference between the reconstructed image and the original image, λ is the Lagrangian coefficient, which is a parameter that balances the bit rate and distortion in the rate-distortion cost, and J is the rate-distortion cost.

[0081] The goal of image coding is to minimize the rate-distortion cost minJ.

[0082] In deep learning-based image coding and decoding, since directly quantifying features is non-differentiable, quantization is replaced by adding uniform noise. Therefore, the goal of minimizing the rate-distortion cost J can be transformed into the following formula (2) on the entire training image set:

[0083]

[0084] Among them, x ∼ p(x) is the distribution of the training image set, is the process of adding noise to the image features to replace quantization.

[0085] Refer to Figure 9 As shown, the image coding and decoding framework using a reversible neural network includes: an encoding-end network reversible network 91 and a decoding-end reversible network 92. When training the image coding and decoding framework using a reversible neural network, the image coding process includes: first, the encoding-end network reversible network 91 extracts features from the input image x to obtain image features y, and then, uniform noise Δy is added to the image to replace quantization to obtain the quantized features of the input image x The image decoding process includes: the decoding-end network reversible network 92 reconstructs the input image x according to the quantized features of the input image x to obtain the reconstructed input image The reconstruction distortion D is the difference between the input image x and the reconstructed input image The mean square error. Among them, the feature extraction of the input image x by the reversible network 91 at the encoding end can be expressed by the following formula (3), and the reversible network 92 at the decoding end reconstructs the input image x according to the quantization feature of the input image x The process of reconstructing the input image x can be expressed by the following formula (4):

[0086] y = f(x) (3)

[0087]

[0088] Under the image encoding and decoding framework based on the reversible network, the above formula (2) can be further deduced as follows:

[0089]

[0090] Further using the Taylor expansion, the reconstruction process can be converted to:

[0091]

[0092] Among them, is the Jacobian matrix of the reversible network 92 at the decoding end.

[0093] By ignoring the high-order terms, formula (5) can be further transformed into:

[0094]

[0095] According to the theorem of matrix and vector norms, the above formula (7) can be further transformed into:

[0096]

[0097] From the above formula (8), it can be obtained that the distortion is determined by the Jacobian matrix of the decoding network and the quantization noise.

[0098] Considering that the code rate R can be approximated by entropy, formula (2) can be transformed into:

[0099]

[0100] It can be seen from the above formula (9) that the objective function of image encoding and decoding is determined by the code rate, the Jacobian matrix of the decoding network, and the quantization noise. And during the decoding process, the quantization noise can be modeled as:

[0101]

[0102] Therefore, reducing the quantization noise Δy during the image reconstruction process can improve the reconstruction quality and thus reduce the distortion.

[0103] Based on the above content, the embodiments of the present application provide an image decoding method, referring to Figure 10As shown, the image decoding method includes the following steps S101 to S103:

[0104] S101. Obtain the encoded data of the target image.

[0105] Among them, the encoded data of the target image is data obtained by compressing the target feature, and the target feature is a feature obtained by performing forward processing on the target image through an n-stage cascaded reversible module. The compression processing includes quantization processing, and n is a positive integer.

[0106] Represent the forward processing of the i-th reversible module as fi i (), represent the target image as X, the target feature as B, and the compression processing as enc(). Then the process of encoding the target image to obtain the encoded data of the target image can be expressed by the following formula:

[0107] B = enc(f n (f n-1 (…(f1(X))))) (11)

[0108] Obtaining the encoded data of the target image can be receiving the encoded data of the target image sent by the media resource server, or reading the encoded data of the target image from the local memory. The embodiments of the present application do not limit this, as long as the encoded data of the target image can be obtained.

[0109] S102. Perform decompression processing on the encoded data of the target image to obtain the reconstructed feature of the target image.

[0110] Represent the target feature as B and the reconstructed feature of the target image as Represent the decompression processing as dec(). Then the above step S102 can be expressed by the following formula:

[0111]

[0112] In some embodiments, it can be passed through Figure 8 The feature decompression module 41 in the shown decoding framework performs decompression processing on the encoded data of the target image to obtain the reconstructed feature of the target image.

[0113] S103. Process the reconstructed feature of the target image through a reconstruction network composed of an n-stage cascaded reconstruction unit to obtain the reconstructed image of the target image.

[0114] Among them, the i-th level reconstruction unit of the reconstruction network is composed of the (n - i + 1)-th level reversible module and the noise reduction module, and is used to reversely process the input features of the i-th level reconstruction unit through the (n - i + 1)-th level reversible module to obtain the initial features of the i-th level reconstruction unit, and filter out the quantization noise in the initial features of the i-th level reconstruction unit through the noise reduction module to obtain the output features of the i-th level reconstruction unit. The quantization noise is the noise brought by the quantization process, and i is a positive integer less than or equal to n.

[0115] The reconstructed features of the target image are represented as The reconstructed image of the target image is represented as The reverse processing of the i-th level reversible module is represented as f i -1 (), and the noise reduction processing of the noise reduction module of the i-th level reconstruction unit is represented as g i (), then the above step S103 can be expressed as the following formula:

[0116]

[0117] After obtaining the encoded data of the target image, the image decoding method provided by the embodiment of the present application first performs decompression processing on the encoded data of the target image to obtain the reconstructed features of the target image, and then processes the reconstructed features of the target image through a reconstruction network composed of n levels of cascaded reconstruction units to obtain the reconstructed image of the target image. Since the i-th level reconstruction unit of the reconstruction network includes a noise reduction module, the noise reduction module can filter out the quantization noise brought by the quantization process and reduce the difference between the target features and the reconstructed features. In addition, the encoded data of the target image is the data obtained by compressing the target features, the target features are the features obtained by forward processing the target image through n levels of cascaded reversible modules, the i-th level reconstruction unit of the reconstruction network includes the (n - i + 1)-th level reversible module, and the input features of the i-th level reconstruction unit can be reversely processed through the (n - i + 1)-th level reversible module. Therefore, the embodiment of the present application can use the reversible module to extract features from the target image and reconstruct the target image according to the reconstructed features. Even if there are large differences between the target image and the training data set, it can also well extract features from the target features and reconstruct the target image according to the reconstructed features. Therefore, the embodiment of the present application can improve the quality of the reconstructed image after encoding.

[0118] As an extension and refinement of the above embodiment, the embodiment of the present application provides another image decoding method. Refer to Figure 11 As shown, this image decoding method includes the following steps:

[0119] S1101. Obtain the encoded data of the target image.

[0120] Among them, the encoded data of the target image is data obtained by compressing the target features, and the target features are features obtained by forward processing the target image through an n-level cascaded reversible module. The compression processing includes quantization processing, and n is a positive integer.

[0121] S1102. Decompress the encoded data of the target image to obtain the reconstructed features of the target image.

[0122] S1103. Perform reverse processing on the reconstructed features of the target image through the nth-level reversible module to obtain the initial features of the first-level reconstruction unit.

[0123] Represent the reconstructed features of the target image as The initial features of the first-level reconstruction unit are represented as F a , and the reverse processing of the nth-level reversible module is represented as Then the above step S1103 can be expressed as the following formula:

[0124]

[0125] S1104. The first activation function layer performs activation processing on the input features of the noise reduction module through the first activation function to obtain the first feature.

[0126] In some embodiments, the first activation function is a rectified linear unit function.

[0127] Represent the input features of the noise reduction module as F in , the first feature is represented as F1, and the activation processing of the first activation function layer is represented as ReLU1(). Then the above step S1104 can be expressed as the following formula:

[0128] F1 = ReLU1(F in ) (15)

[0129] S1105. The first convolutional layer performs convolutional processing on the first feature to obtain the second feature.

[0130] In some embodiments, the first convolutional layer is a convolutional layer with a kernel size of 3*3.

[0131] Represent the first feature as F1, the second feature as F2, and the convolutional processing of the first convolutional layer as conv1 3*3 (). Then the above step S1105 can be expressed as the following formula:

[0132] F2 = conv1 3*3 (F1) (16)

[0133] S1106. The second activation function layer activates the second feature through a second activation function to obtain a third feature.

[0134] Denote the second feature as F2, the third feature as F3, and the activation process of the second activation function layer as ReLU2(). Then the above step S1106 can be expressed as the following formula:

[0135] F3 = ReLU2(F2) (17)

[0136] S1107. The second convolutional layer performs a convolutional process on the third feature to obtain a fourth feature.

[0137] In some embodiments, the second convolutional layer is a convolutional layer with a convolutional kernel size of 3*3.

[0138] Denote the third feature as F3, the fourth feature as F4, and the convolutional process of the second convolutional layer as conv2 3*3 (). Then the above step S1107 can be expressed as the following formula:

[0139] F4 = conv2 3*3 (F3) (18)

[0140] S1108. The first adjustment layer adjusts the fourth feature based on a first adjustment coefficient to obtain a fifth feature.

[0141] In some embodiments, the first adjustment coefficient is 0.2.

[0142] Denote the fourth feature as F4, the fifth feature as F5, and the first adjustment coefficient as α. Then the above step S1108 can be expressed as the following formula:

[0143] F5 = F4 * α (19)

[0144] S1109. The first addition fusion layer performs an addition fusion on the input feature of the noise reduction module and the fifth feature to obtain the output feature of the noise reduction module.

[0145] Denote the output feature of the noise reduction module as F in and the output feature of the noise reduction module as F out , and the fifth feature as F5. Then the above step S1109 can be expressed as the following formula:

[0146] F out = F in + F5 (20)

[0147] So far, the processing of features by the first-level reconstruction unit of the reconstruction network is completed. Since the n-level cascaded reconstruction units of the reconstruction network are cascaded, the output features of the previous-level reconstruction unit are used as the input features for each subsequent-level reconstruction unit, and the processing similar to the above steps S1103 to S1109 is cyclically executed. The difference is that the reversible modules of different reconstruction units are different (the reversible module of the i-th level reconstruction unit is the (n - i + 1)-th level reversible module), and the input features of the noise reduction modules of each reconstruction unit are the output features of the reversible module belonging to the same reconstruction unit. After all the reconstruction units of the reconstruction network have completed the processing of features, the following step S1110 is executed:

[0148] S1110. Determine the output features of the n-th level reconstruction unit as the reconstructed image of the target image.

[0149] As an extension and refinement of the above embodiments, the embodiments of the present application provide another image decoding method. Referring to Figure 12 as shown, this image decoding method includes the following steps:

[0150] S1201. Obtain the encoded data of the target image.

[0151] Among them, the encoded data of the target image is the data obtained by compressing the target features. The target features are the features obtained by forward processing the target image through n-level cascaded reversible modules. The compression processing includes quantization processing, and n is a positive integer.

[0152] S1202. Decompress the encoded data of the target image to obtain the reconstructed features of the target image.

[0153] S1203. Perform reverse processing on the reconstructed features of the target image through the n-th level reversible module to obtain the initial features of the first-level reconstruction unit.

[0154] Similarly, the above step S1203 can also be expressed as the above formula (14). To avoid repetition, it will not be described again here.

[0155] S1204. The first dense connection module processes the input features of the noise reduction module to obtain the sixth feature.

[0156] Represent the input features of the noise reduction module as F in , represent the sixth feature as F6, and represent the processing of features by the first dense connection module as f dense1 , then the above step S1204 can be expressed as the following formula:

[0157] F6 = f dense1 (F in ) (21)

[0158] S1205. The third convolutional layer performs convolutional processing on the sixth feature to obtain a seventh feature.

[0159] In some embodiments, the third convolutional layer is a convolutional layer with a kernel size of 1*1.

[0160] Denote the sixth feature as F6, the seventh feature as F7, and the convolutional processing of the third convolutional layer as conv3 1*1 , then the above step S1205 can be expressed as:

[0161] F7 = conv3 1*1 (F6) (22)

[0162] S1206. The fourth convolutional layer performs convolutional processing on the seventh feature to obtain an eighth feature.

[0163] In some embodiments, the fourth convolutional layer is a convolutional layer with a kernel size of 3*3.

[0164] Denote the seventh feature as F7, the eighth feature as F8, and the convolutional processing of the fourth convolutional layer as conv4 3*3 , then the above step S1206 can be expressed as:

[0165] F8 = conv4 3*3 (F7) (23)

[0166] S1207. The fifth convolutional layer performs convolutional processing on the eighth feature to obtain a ninth feature.

[0167] In some embodiments, the fourth convolutional layer is a convolutional layer with a kernel size of 1*1.

[0168] Denote the eighth feature as F8, the ninth feature as F9, and the convolutional processing of the fifth convolutional layer as conv5 1*1 , then the above step S1207 can be expressed as:

[0169] F9 = conv5 1*1 (F8) (24)

[0170] S1208. The second dense connection module is used to process the sixth feature, the seventh feature, the eighth feature, and the ninth feature to obtain a tenth feature.

[0171] Denote the sixth feature as F6, the seventh feature as F7, the eighth feature as F8, the ninth feature as F9, and the tenth feature as F 10 , and the processing of the second dense connection module on the features is denoted as f dense2, then the above step S1208 can be expressed as the following formula:

[0172] F 10 = f dense2 (F6,F7,F8,F9,F 10 ) (25)

[0173] S1209. The second adjustment layer adjusts the tenth feature based on the second adjustment coefficient to obtain the eleventh feature.

[0174] In some embodiments, the second adjustment coefficient is 0.2.

[0175] Represent the tenth feature as F 10 , represent the eleventh feature as F 11 , represent the second adjustment coefficient as β, then the above step S1209 can be expressed as the following formula:

[0176] F 11 = F 10 *β (26)

[0177] S1210. The second addition fusion layer adds and fuses the input feature of the noise reduction module and the fifth feature to obtain the output feature of the noise reduction module.

[0178] Represent the output feature of the noise reduction module as F in , represent the output feature of the noise reduction module as F out , represent the eleventh feature as F 11 , then the above step S1210 can be expressed as the following formula:

[0179] F out = F in + F 11 (27)

[0180] So far, the processing of the features by the first-level reconstruction unit of the reconstruction network is completed. Since the n-level cascaded reconstruction units of the reconstruction network are cascaded, each subsequent reconstruction unit will use the output feature of the previous reconstruction unit as the input feature and loop to execute processing similar to the above steps S1203 to S1210. The difference is that the reversible modules of different reconstruction units are different (the reversible module of the i-th reconstruction unit is the (n - i + 1)-th reversible module), and the input features of the noise reduction modules of each reconstruction unit are all the output features of the reversible module belonging to the same reconstruction unit. After all the reconstruction units of the reconstruction network have completed the feature processing, execute the following step S1211:

[0181] S1211. Determine the output feature of the n-th reconstruction unit as the reconstructed image of the target image.

[0182] In some embodiments, the image decoding method provided by the embodiments of the present application further includes: training each network model in the image encoder and the image decoder based on a training data set.

[0183] Wherein, the training data set includes: a plurality of sample images. The process of training each network model in the image encoder and the image decoder based on the training data set may include: first inputting the sample images into the image encoder for encoding to obtain the encoded data of the images, then decoding the encoded data through the image decoder to obtain the reconstructed images of the sample images, then calculating the loss value according to the sample images and the reconstructed images, and adjusting the parameters of the network models in the image encoder and / or the image decoder according to the loss value until the image encoder and the image decoder converge.

[0184] In some embodiments, the data set provided by flicker_2W may be used as the training data set.

[0185] In some embodiments, the data set provided by kodak may also be used as a test data set to test the image encoding and decoding framework.

[0186] In some embodiments, when training each network model in the image encoder and the image decoder based on the training data set, the batch size may be set to 24.

[0187] In some embodiments, some embodiments of the present application provide an image decoding device, which includes:

[0188] A memory configured to store a computer program;

[0189] A processor configured to, when calling the computer program, enable the image decoding device to implement the image decoding method described in any of the above embodiments.

[0190] In some embodiments, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is enabled to implement the image decoding method described in any of the above embodiments or the image decoding method described in any of the above embodiments.

[0191] In some embodiments, some embodiments of the present application provide a computer program product, which, when running on a computer, enables the computer to implement the image decoding method described in any of the above embodiments or the image decoding method described in any of the above embodiments.

[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0193] For the sake of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. An image decoding method, characterized in that, Including: Obtaining encoded data of a target image, where the encoded data of the target image is data obtained by compressing a target feature, the target feature is a feature obtained by performing forward processing on the target image through an n-level cascaded reversible module, the compression processing includes quantization processing, and n is a positive integer; Performing decompression processing on the encoded data of the target image to obtain a reconstructed feature of the target image; Processing the reconstructed feature of the target image through a reconstruction network composed of an n-level cascaded reconstruction unit to obtain a reconstructed image of the target image; the i-th reconstruction unit of the reconstruction network is composed of the (n - i + 1)-th reversible module and a noise reduction module, and is used to perform reverse processing on the input feature of the i-th reconstruction unit through the (n - i + 1)-th reversible module to obtain an initial feature of the i-th reconstruction unit, and filtering quantization noise in the initial feature of the i-th reconstruction unit through the noise reduction module to obtain an output feature of the i-th reconstruction unit, where the quantization noise is the noise brought by the quantization processing, and i is a positive integer less than or equal to n.

2. The method according to claim 1, wherein The noise reduction module includes: a first activation function layer, a first convolutional layer, a second activation function layer, a second convolutional layer, a first adjustment layer, and a first addition fusion layer; The first activation function layer is used to perform activation processing on the input feature of the noise reduction module through a first activation function to obtain a first feature; The first convolutional layer is used to perform convolutional processing on the first feature to obtain a second feature; The second activation function layer is used to perform activation processing on the second feature through a second activation function to obtain a third feature; The second convolutional layer is used to perform convolutional processing on the third feature to obtain a fourth feature; The first adjustment layer is used to adjust the fourth feature based on a first adjustment coefficient to obtain a fifth feature; The first addition fusion layer is used to perform addition fusion on the input feature of the noise reduction module and the fifth feature to obtain the output feature of the noise reduction module.

3. The method according to claim 2, wherein The first activation function and / or the second activation function is a rectified linear unit function.

4. The method according to claim 2, wherein The first convolutional layer and / or the second convolutional layer is a convolutional layer with a convolutional kernel size of 3*3.

5. The method according to claim 2, wherein The first adjustment coefficient is 0.

2.

6. The method according to claim 1, wherein The noise reduction module includes: a first dense connection module, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a second dense connection module, a second adjustment layer, and a second addition fusion layer; The first dense connection module is used to process the input feature of the noise reduction module to obtain a sixth feature; The third convolutional layer is used to perform convolutional processing on the sixth feature to obtain a seventh feature; The fourth convolutional layer is used to perform convolutional processing on the seventh feature to obtain an eighth feature; The fifth convolutional layer is used to perform convolutional processing on the eighth feature to obtain a ninth feature; The second dense connection module is used to process the sixth feature, the seventh feature, the eighth feature, and the ninth feature to obtain a tenth feature; The second adjustment layer is used to adjust the tenth feature based on a second adjustment coefficient to obtain an eleventh feature; The second addition fusion layer is used to perform addition fusion on the input feature of the noise reduction module and the eleventh feature to obtain the output feature of the noise reduction module.

7. The method according to claim 6, characterized in that, The third convolutional layer and / or the fifth convolutional layer is a convolutional layer with a convolutional kernel size of 3*3.

8. The method according to claim 6, wherein The fourth convolutional layer is a convolutional layer with a convolutional kernel size of 1*1.

9. The method according to claim 6, characterized in that, The second adjustment coefficient is 0.

2.

10. The method according to any one of claims 1-9, characterized in that, The reversible module is composed of a channel rearrangement module, a sixth convolutional layer, a first coupling layer, a second coupling layer, and a third coupling layer connected in series in sequence.

11. An image decoding device, characterized in that, Comprising: A memory configured to store a computer program; A processor configured to, when calling the computer program, cause the image decoding device to implement the image decoding method according to any one of claims 1-10.