Image encoding method, image decoding method and device

By splitting the reversible neural network into two reversible block sequences and combining multi-stage channel downsampling, the problems of low encoding efficiency and poor reconstruction quality in traditional image encoding methods are solved, and more efficient image encoding and normal image reconstruction are achieved.

CN120358360APending Publication Date: 2025-07-22HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410050682.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The traditional image encoding method has the problem of low encoding efficiency and low image quality after encoding.

Method used

Using a first reversible block sequence composed of n cascaded reversible modules and a second reversible block sequence composed of m cascaded reversible modules, combined with the first and second channel downsampling modules, the target image is subjected to multi-stage channel downsampling processing to reduce channel redundancy in the features.

Benefits of technology

It improves the encoding efficiency of the image encoder, reduces channel redundancy in the features, and ensures normal reconstruction of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358360A_ABST
    Figure CN120358360A_ABST
Patent Text Reader

Abstract

The invention provides an image coding method and device and an image decoding method and device, and relates to the technical field of image coding and decoding. Comprising the steps that feature extraction is carried out on a target image through a first reversible block sequence to obtain a first feature, the first reversible block sequence is composed of n cascaded reversible modules, and n is a positive integer; performing channel down-sampling processing on the first feature through a first channel down-sampling module to obtain a second feature; the second feature is processed through a second reversible block sequence to obtain a third feature, the second reversible block sequence is composed of m cascaded reversible modules, and m is a positive integer; performing channel down-sampling processing on the third feature through a second channel down-sampling module to obtain an image feature of the target image; and encoding the image features of the target image to obtain encoded data of the target image. Some embodiments of the invention are used for improving the coding efficiency of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Some embodiments of the present application relate to the field of image encoding and decoding technologies. More specifically, it relates to an image encoding method, an image decoding method, and a device. Background Art

[0002] With the development of the information age, in the field of the Internet, people obtain information by watching videos or images, and image transmission has also become an important communication method. With the increase in the amount of data and information, the physical space occupied by videos or images is also getting larger and larger; if unprocessed videos or images are directly transmitted over the network, it will occupy a large amount of network bandwidth and consume a large amount of traffic. Therefore, videos or images are compressed before transmission to reduce their volume. However, traditional image encoding methods have the problem of low encoding efficiency. Summary of the Invention

[0003] Exemplary embodiments of the present application provide an image encoding method, an image decoding method, and a device for improving the encoding efficiency of images.

[0004] Some embodiments of the present application provide the following technical solutions:

[0005] In a first aspect, some embodiments of the present application provide an image encoding method, including:

[0006] Performing feature extraction on a target image through a first reversible block sequence to obtain a first feature, where the first reversible block sequence is composed of n cascaded reversible modules, and n is a positive integer;

[0007] Performing channel downsampling processing on the first feature through a first channel downsampling module to obtain a second feature;

[0008] Processing the second feature through a second reversible block sequence to obtain a third feature, where the second reversible block sequence is composed of m cascaded reversible modules, and m is a positive integer;

[0009] Performing channel downsampling processing on the third feature through a second channel downsampling module to obtain an image feature of the target image;

[0010] Encoding the image feature of the target image to obtain encoded data of the target image.

[0011] In a second aspect, some embodiments of the present application provide an image decoding method, including:

[0012] Obtain the encoded data of the target image, where the encoded data of the target image is the encoded data obtained by encoding the image features of the target image, and the image features of the target image are the features obtained by processing the target image through a first reversible block sequence, a first channel downsampling module, a second reversible block sequence, and a second channel downsampling module in sequence; the first reversible block sequence is composed of n cascaded reversible modules, and the second reversible block sequence is composed of m cascaded reversible modules, where n and m are positive integers;

[0013] Decode the encoded data of the target image to obtain the reconstructed image features of the target image;

[0014] Perform channel upsampling on the reconstructed image features through a first channel upsampling module to obtain a reconstructed third feature;

[0015] Perform reverse processing on the reconstructed third feature through the second reversible block sequence to obtain a reconstructed second feature;

[0016] Perform channel upsampling on the reconstructed second feature through a second channel upsampling module to obtain a reconstructed first feature;

[0017] Perform reverse processing on the reconstructed first feature through the first reversible block sequence to obtain the reconstructed image of the target image.

[0018] In a third aspect, some embodiments of the present application provide an image encoding device, including:

[0019] A memory configured to store a computer program;

[0020] A processor configured to, when calling the computer program, cause the image encoding device to implement the image encoding method described in the first aspect.

[0021] In a fourth aspect, some embodiments of the present application provide an image decoding device, including:

[0022] A memory configured to store a computer program;

[0023] A processor configured to, when calling the computer program, cause the image decoding device to implement the image decoding method described in the second aspect.

[0024] In a fifth aspect, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is caused to implement the image encoding method described in the first aspect or the image decoding method described in the second aspect.

[0025] In a sixth aspect, some embodiments of the present application provide a computer program product. When the computer program product runs on a computer, the computer is enabled to implement the image encoding method described in the first aspect or the image decoding method described in the second aspect.

[0026] As can be seen from the above technical solutions, when encoding a target image, the image encoding method provided by the embodiments of the present application first extracts features of the target image through a first reversible block sequence composed of n cascaded reversible modules to obtain a first feature. Secondly, the first feature is subjected to channel downsampling processing through a first channel downsampling module to obtain a second feature. Then, the second feature is processed through a second reversible block sequence composed of m cascaded reversible modules to obtain a third feature. Next, the third feature is subjected to channel downsampling processing through a second channel downsampling module to obtain the image feature of the target image. Finally, the image feature of the target image is encoded to obtain the encoded data of the target image. Since the image encoding method provided by the embodiments of the present application splits the reversible neural network into two reversible block sequences and performs channel downsampling in multiple stages through two channel downsampling modules, the embodiments of the present application can reduce the channel redundancy in the features, thereby improving the encoding efficiency of the image encoder. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate some embodiments of the present application or the implementation manners in related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0028] Figure 1 Shows a schematic structural diagram of an image encoder in related technologies;

[0029] Figure 2 Shows a schematic structural diagram of an image decoder in related technologies;

[0030] Figure 3 Shows a schematic structural diagram of an image encoder in some embodiments of the present application;

[0031] Figure 4 Shows a schematic structural diagram of a reversible module in some other embodiments of the present application;

[0032] Figure 5 Shows a schematic structural diagram of a feature encoding module in some embodiments of the present application;

[0033] Figure 6 Shows a schematic structural diagram of an image decoder in some embodiments of the present application;

[0034] Figure 7 The structural schematic diagram of an image encoder in some other embodiments of the present application is shown;

[0035] Figure 8 The structural schematic diagram of a channel upsampling module in some embodiments of the present application is shown;

[0036] Figure 9 The structural schematic diagram of a feature decoding module in some embodiments of the present application is shown;

[0037] Figure 10 The flowchart of the steps of an image encoding method in some embodiments of the present application is shown;

[0038] Figure 11 One of the flowcharts of the steps of an image decoding method in some embodiments of the present application is shown;

[0039] Figure 12 Another flowchart of the steps of an image decoding method in some embodiments of the present application is shown. Detailed implementation manners

[0040] It should be noted that the brief description of the terms in the present application is only for facilitating the understanding of the following described implementation manners. To make the purpose and implementation manners of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application in conjunction with the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0041] It is not intended to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.

[0042] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclude inclusion. For example, a product or device comprising a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.

[0043] The mention of "some implementation manners", "some embodiments", etc. in the specification indicates that the described implementation manners or embodiments may include specific features, structures, or characteristics, but not necessarily every embodiment includes the specific feature, structure, or characteristic. In addition, such phrases do not necessarily refer to the same implementation manner. Additionally, when describing a specific feature, structure, or characteristic in connection with an embodiment, it is considered that implementing such feature, structure, or characteristic in connection with other implementation manners (whether or not clearly described herein) is within the knowledge of those skilled in the art.

[0044] Traditional image coding methods have problems such as poor coding effect and low quality of the reconstructed image after coding. To improve the problems of poor coding effect and low quality of the reconstructed image after coding in traditional image coding methods, image coding and decoding technologies based on deep learning are developing rapidly. For example, the JPEG AI standard created by the Joint Photographic Experts Group (JPEG) is an extensible image coding standard based on machine learning. The general coding process of image coding and decoding technology based on deep learning is as follows: First, the feature extraction network extracts features from the image to be coded, and then compresses the features of the image to be coded extracted by the feature extraction network, such as quantization, to obtain the coding data of the image to be coded. Correspondingly, the general decoding process of image coding and decoding technology based on deep learning is as follows: First, decompress the coding data of the image to be decoded to obtain the features of the reconstructed image to be decoded, and then reconstruct the image to be decoded through the image reconstruction network and the features of the image to be decoded. Among them, the feature extraction network and the image reconstruction network are generally network models obtained by stacking traditional convolutional neural networks and training on a specific data set. If the image to be coded and decoded has a large difference from the training data, it is difficult to extract features of this image with high quality or reconstruct this image based on the features. To further solve this problem, related technologies propose to perform feature extraction and feature-based image reconstruction through reversible neural networks in the process of image coding and decoding based on deep learning. The total number of features before and after the reversible transformation of the reversible neural network remains unchanged. While the spatial resolution decreases, the channel resolution will increase. To improve the compression efficiency, channel compression processing is required.

[0045] Referring to Figure 1 As shown, the process of image coding in related technologies includes: First, the reversible neural network 11 at the coding end extracts the image feature F of the target image X X , and then the channel downsampling module 12 performs channel downsampling processing on the image feature F of the target image X X to obtain the channel downsampling feature with reduced number of channels Finally, the feature coding module 13 performs compression processing on the channel downsampling feature to obtain the coding data B of the target image X . Correspondingly, referring to Figure 2 As shown, the process of image decoding in related technologies includes: First, the feature decoding module 21 performs decompression processing on the coding data B of the target image X to obtain the reconstructed channel downsampling feature Then, the channel upsampling module 22 performs channel upsampling processing on the reconstructed channel downsampling feature to obtain the reconstructed image feature Finally, the reversible neural network 23 at the decoding end utilizes the reconstructed image features to perform image reconstruction to obtain the reconstructed image of the target image As shown above, when the related technology performs image encoding, only one channel downsampling module is used to compress the channels. However, this processing method will cause channel redundancy in the encoder and affect the encoding effect. In addition, when the related technology performs image decoding, only one channel upsampling module is used to decompress the channels. Therefore, it is difficult to obtain a good reconstruction effect, and the image encoding and decoding performance still needs to be further improved.

[0046] The image in the embodiment of the present application can be an independent image or a video frame in a video. Since a video can be regarded as a sequence of video frames composed of multiple images, when encoding and decoding each video frame based on the image encoding method and the image decoding method, the video can be encoded and decoded based on the image encoding method and the image decoding method.

[0047] The embodiment of the present application relates to the technical field of image encoding and decoding. First, the image encoding and decoding framework provided by the embodiment of the present application will be described below.

[0048] Referring to Figure 3 As shown, the image encoder provided by some embodiments of the present application includes: a first reversible block sequence 31, a first channel downsampling module 32, a second reversible block sequence 33, a second channel downsampling module 34, and a feature encoding module 35.

[0049] Among them, the first reversible block sequence 31 is composed of n cascaded reversible modules, and is used to extract features from the target image X to obtain the first image feature F1, where n is a positive integer. The first channel downsampling module 32 is used to perform channel downsampling processing on the first feature F1 to obtain the second feature F2. The second reversible block sequence 33 is composed of m cascaded reversible modules, and is used to process the second feature F2 to obtain the third feature F3, where m is a positive integer. The second channel downsampling module 34 is used to perform channel downsampling processing on the third feature F3 to obtain the image feature F of the target image X X . The feature encoding module 35 is used to encode the image feature F of the target image X X to obtain the encoded data B of the target image X X .

[0050] Referring to Figure 4As shown, in some embodiments, the first reversible block sequence 31 is composed of 3 cascaded reversible modules (the first-level reversible module, the second-level reversible module, and the third-level reversible module), and the second reversible block sequence 33 is composed of 1 reversible module (the fourth-level reversible module). Any reversible module includes: a channel recombination module (Downsampling Channel Recombination Module) 41, a convolutional layer 42, a first coupling layer (Coupling Layer) 43, a second coupling layer 44, and a third coupling layer 45 connected in series in sequence. That is, the reversible module is used to process the input feature I1 of the reversible module through the channel recombination module 41, the convolutional layer 42, the first coupling layer 43, the second coupling layer 44, and the third coupling layer 45 in sequence to obtain the output feature O1 of the reversible module. Among them, the channel recombination module 41 is used to perform a downsampling operation on the input signal I1 to reduce the number of feature channels, reduce the amount of data, and extract features. The channel recombination module 41 can be composed of sub-modules such as a max pooling layer, an average pooling layer, and a convolutional layer. These sub-modules can perform downsampling on the input signal and perform recombination in the channel dimension. The coupling layers (the first coupling layer 43, the second coupling layer 44, and the third coupling layer 45) are used to fuse and interact the information between different channels or feature maps. The coupling layer can perform element-wise multiplication and addition operations in the channel dimension to increase the expression ability and flexibility of the model.

[0051] In some embodiments, when the reversible module performs feature processing, the ratio of the number of channels of the input feature and the output feature of the channel recombination module 41 is 4, the ratio of the number of channels of the input feature and the output feature of the first channel downsampling module 32 is 1 / 2, and the ratio of the number of channels of the input feature and the output feature of the second channel downsampling module 34 is 1 / 3. That is, if the feature tensor of the target image is 3*H*W, then the feature tensor of the output feature of the first-level reversible module is 12*H*W, the feature tensor of the output feature of the second-level reversible module is 48*H*W, the feature tensor of the output feature (the first feature) of the third-level reversible module is 196*H*W, the feature tensor of the output feature (the second feature) of the first channel downsampling module 32 is 96*H*W, the feature tensor of the output feature (the third feature) of the fourth-level reversible module is 384*H*W, and the feature tensor of the output feature (the image feature of the target image) of the second channel downsampling module 34 is 128*H*W.

[0052] In some embodiments, the feature encoding module 35 is a feature encoding module based on a JPEG AI encoder.

[0053] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 includes: a residual unit 501. The residual unit 301 is used to calculate the target feature FX The difference from the predicted feature μ to obtain the residual feature R. Among them, the predicted feature μ is the predicted information of the target feature F obtained by the prediction fusion network based on the context feature output by the context network and the reconstructed latent feature of the hyperprior decoding network. The input of the context network is the target feature F that has been encoded X for the target feature F X The predicted information of. The input of the context network is the elements of the target feature F that has been encoded X of.

[0054] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a gain unit 502. The gain unit 502 is used to perform a gain process on the residual feature R to obtain the gain-processed residual feature Rg.

[0055] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a first quantization unit 503. The first quantization unit 503 is used to perform a quantization process on the gain-processed residual feature Rg to obtain the quantized residual feature r.

[0056] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a hyperprior encoding network 504. The hyperprior encoding network 504 is used to perform hyperprior encoding on the target feature F X to obtain a hyper-latent feature Z, and extract additional auxiliary information during the hyperprior encoding process, so that an accurate probability model for entropy encoding can be obtained based on the encoding result.

[0057] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a second quantization unit 505. The second quantization unit 505 is used to perform a quantization process on the hyper-latent feature Z to obtain the quantized hyper-latent feature z.

[0058] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a decomposition entropy model 506. The decomposition entropy model 506 is used to obtain a cumulative distribution function cdf (Cumulative Distribution Function, CDF).

[0059] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a first lossless encoder 507. The first lossless encoder 507 is used to perform lossless encoding on the quantized hyper-latent feature z according to the cumulative distribution function cdf to obtain the first encoded data b1.

[0060] Referring toFigure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a lossless decoder 508. The lossless decoder 508 is configured to decode the first encoded data b1 according to the cumulative distribution function cdf to obtain the reconstructed super latent feature

[0061] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a super prior scale encoding network 509. The super prior scale encoding network 509 is configured to process the reconstructed super latent feature to obtain the variance N(0,σ) of the Gaussian distribution with a probability distribution mean of 0.

[0062] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a second lossless encoder 510. The second lossless encoder 510 is configured to perform lossless encoding on the quantized residual feature r based on the variance N(0,σ) to obtain the second encoded data b2.

[0063] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a super prior decoding network 511. The super prior decoding network 511 is configured to perform decoding processing on the reconstructed super latent feature to obtain the reconstructed target feature

[0064] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: an inverse gain unit 512. The inverse gain unit 512 is configured to perform inverse gain processing on the quantized residual feature r to obtain the inversely gain-processed residual feature rig.

[0065] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a context network 514. The context network 514 is configured to perform context feature extraction on the fused latent feature y to obtain the context feature cont of the fused latent feature y.

[0066] Referring to Figure 5 As shown, in some embodiments, the feature encoding module 35 further includes: a prediction fusion network 515. The prediction fusion network 515 is configured to obtain the predicted feature μ of the next element of the target feature F based on the context feature cont and the reconstructed prior feature X for encoding the next element based on the predicted feature μ.

[0067] Referring to Figure 6As shown, the image decoder provided by some embodiments of the present application includes: a feature decoding module 61, a first channel upsampling module 62, a second reversible block sequence 63, a second channel upsampling module 64, and a first reversible block sequence 65. Among them, the feature decoding module 61 is used to decode the encoded data B of the target image X of the encoded data of the target image X for decoding to obtain the reconstructed image features of the target image X The first channel upsampling module 62 is used to perform channel upsampling processing on the reconstructed image features to obtain the reconstructed third feature The second reversible block sequence 63 is composed of m cascaded reversible modules and is used to perform reverse processing on the reconstructed third feature to obtain the reconstructed second feature m is a positive integer. The second channel upsampling module 64 is used to perform channel upsampling processing on the reconstructed second feature to obtain the reconstructed first feature The first reversible block sequence 65 is composed of n cascaded reversible modules and performs reverse processing on the reconstructed first feature to obtain the reconstructed image of the target image X

[0068] Referring to Figure 7 As shown, in some embodiments, the first reversible block sequence is composed of 3 cascaded reversible modules, and the second reversible block sequence is composed of 1 reversible module. That is Figure 6 As shown, the image decoder includes: a feature decoding module 61, a first channel upsampling module 62, a fourth-level reversible module 71, a second channel upsampling module 64, a third-level reversible module 72, a second-level reversible module 73, and a first-level reversible module 74. The structure of each reversible module (the fourth-level reversible module 71, the third-level reversible module 72, the second-level reversible module 73, and the first-level reversible module 74) can be referred to Figure 4 As shown, to avoid repetition, it will not be described again here

[0069] In some embodiments, when the reversible module performs reverse feature processing, the ratio of the number of channels of the input feature and the output feature of the channel recombination module 41 is 1 / 4, the ratio of the number of channels of the input feature and the output feature of the first channel upsampling module 62 is 3, and the ratio of the number of channels of the input feature and the output feature of the second channel upsampling module 64 is 2. That is, if the feature tensor of the reconstructed image feature of the target image is 128*H*W, then the feature tensor of the output feature (reconstructed third feature) of the first channel upsampling module 62 is 384*H*W, the feature tensor of the output feature (reconstructed second feature) of the 4th-level reversible module is 128*H*W, the feature tensor of the output feature (reconstructed first feature) of the second channel upsampling module 64 is 256*H*W, the feature tensor of the output feature of the 3rd-level reversible module is 64*H*W, the feature tensor of the output feature of the 2nd-level reversible module is 16*H*W, and the feature tensor of the output feature (reconstructed image of the target image) of the 1st-level reversible module is 3*H*W.

[0070] Referring to Figure 8 As shown, in some embodiments, the first channel upsampling module 62 and / or the second channel upsampling module 64 includes: a first convolutional layer 81, an activation function layer 82, a second convolutional layer 83, a third convolutional layer 84, and an addition fusion layer 85. Among them, the first convolutional layer 81 is used to perform convolutional processing on the input feature I2 (reconstructed image feature or reconstructed second feature ) of the channel upsampling module to obtain a first convolutional feature F conv1 . The activation function layer 82 is used to process the first convolutional feature F conv1 to obtain an activation feature F ReLU . The second convolutional layer 83 is used to perform convolutional processing on the activation feature F ReLU to obtain a second convolutional feature F conv2 . The third convolutional layer 84 is used to perform convolutional processing on the input feature I2 of the channel upsampling module to obtain a third convolutional feature F conv3 . The addition fusion layer 85 is used to perform addition fusion on the second convolutional feature F conv2 and the third convolutional feature F conv3 to obtain the output feature O2 of the channel upsampling module.

[0071] In some embodiments, the first convolutional layer 81, the second convolutional layer 83, and the third convolutional layer 84 are all convolutional layers with a convolutional kernel size of 3*3.

[0072] In some embodiments, the ratio of the number of channels of the input features to the number of channels of the output features of the first convolutional layer 81 and the third convolutional layer 84 is N, where N is greater than 1; the ratio of the number of channels of the input features to the number of channels of the output features of the second convolutional layer 83 is 1. That is, if the number of channels of the input features of the upsampling module is C, then the number of channels of the output features of the first convolutional layer 81, the second convolutional layer 83, and the third convolutional layer are all N*C.

[0073] In some embodiments, the activation function layer 82 is used to perform activation processing on the first convolutional feature F conv1 to obtain the activation feature F ReLU . The ReLU function is characterized in that when the input value is greater than 0, the output value is the input value itself; when the input value is less than or equal to 0, the output value is 0. The introduction of this function can solve the problem of gradient disappearance existing in traditional sigmoid or tanh activation functions in deep neural networks, thereby accelerating the training speed and improving the accuracy of the model.

[0074] In some embodiments, the feature decoding module 61 is a feature decoding module based on a JPEG AI decoder.

[0075] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 includes: a decomposition entropy model 901. The decomposition entropy model 901 is used to obtain the cumulative distribution function cdf.

[0076] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a first lossless decoder 902. The first lossless decoder 902 is used to decode the first encoded data b1 according to the cumulative distribution function cdf to obtain the reconstructed super latent feature

[0077] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a super prior scale decoding network 903. The super prior scale decoding network 903 is used to process the reconstructed super latent feature to obtain the variance of the Gaussian distribution N(0,σ) with a probability distribution of 0 mean.

[0078] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a second lossless decoder 904. The second lossless decoder 904 is used to decode the second encoded data b2 according to the variance N(0,σ) to obtain the reconstructed quantization residual feature

[0079] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: an inverse gain unit 905. The inverse gain unit 905 is configured to perform inverse gain processing on the reconstructed quantized residual feature to obtain the residual feature after inverse gain processing

[0080] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a fusion module 906. The fusion module 906 is configured to perform addition fusion on the residual feature after inverse gain processing and the reconstructed prediction feature to obtain the reconstructed fusion latent feature and sequentially send it to the first-channel upsampling module 62, the second reversible block sequence 63, the second-channel upsampling module 64, and the first reversible block sequence 65 to obtain the reconstructed image of the target image X

[0081] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a hyperprior decoding network 907. The hyperprior decoding network 907 is configured to perform decoding processing on the reconstructed hyper-latent feature to obtain the reconstructed target feature

[0082] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a prediction fusion network 908. The prediction fusion network 908 is configured to obtain the reconstructed prediction feature according to the reconstructed target feature and the context feature cont output by the context model 909

[0083] Referring to Figure 9 as shown, in some embodiments, the feature decoding module 61 further includes: a context model 909. The context model 909 is configured to obtain the reconstructed prediction feature from the output context feature cont The context model 909 is further configured to perform context feature extraction on the reconstructed fusion latent feature to obtain the context feature cont of the reconstructed fusion latent feature so as to decompress the next element of the first encoded data b1 through the context feature cont of the reconstructed fusion latent feature

[0084] The embodiments of the present application further provide an image encoding method. Referring to Figure 10 ​As shown, the image encoding method includes the following steps S101 to S105:

[0085] S101. Extract features from the target image through the first reversible block sequence to obtain the first feature.

[0086] Among them, the first reversible block sequence is composed of n cascaded reversible modules, and n is a positive integer.

[0087] In some embodiments, n = 3. That is, the first reversible block sequence is composed of 3 cascaded reversible modules.

[0088] Represent the forward processing of the i-th reversible module in the first reversible block sequence as fi i (), the target image is represented as X, and the first feature is represented as F1. Then the above step S101 can be expressed as the following formula:

[0089] F1 = f n (f n-1 (…(f1(X)))) (1)

[0090] S102. Perform channel downsampling on the first feature through the first channel downsampling module to obtain the second feature.

[0091] In some embodiments, the first channel downsampling module may perform channel downsampling on the first feature in an average pooling manner. If the downsampling factor is N, then the feature values of N channels of the first feature are averaged as the feature values of the channels of the second feature.

[0092] In some embodiments, the downsampling factor of the first channel downsampling module for performing channel downsampling on the first feature is 2. That is, if the feature tensor of the first feature is 64C*H*W, then the feature tensor of the second feature is 32C*H*W.

[0093] Represent the downsampling process of the first channel downsampling module as ds1(), the first feature as F1, and the second feature as F2. Then the above step S102 can be expressed as the following formula:

[0094] F2 = ds1(F1) (2)

[0095] S103. Process the second feature through the second reversible block sequence to obtain the third feature.

[0096] Among them, the second reversible block sequence is composed of m cascaded reversible modules, and m is a positive integer.

[0097] In some embodiments, m = 1. That is, the second reversible block sequence is composed of 1 reversible module.

[0098] Denote the forward processing of the j-th reversible module in the second reversible block sequence as f j (), the second feature is denoted as F2, and the third feature is denoted as F3. Then the above step S103 can be expressed as the following formula:

[0099] F3 = f m (f m-1 (…(f1(F2)))) (3)

[0100] S104. Perform channel downsampling on the third feature through the second channel downsampling module to obtain the image feature of the target image.

[0101] In some embodiments, the second channel downsampling module may perform channel downsampling on the first feature by using average pooling.

[0102] In some embodiments, the downsampling multiple of the second channel downsampling module for performing channel downsampling on the third feature is 3. That is, if the feature tensor of the third feature is 128C*H*W, then the feature tensor of the image feature of the target image is 128 / 3C*H*W.

[0103] Denote the downsampling process of the first channel downsampling module as ds2(), the third feature as F3, and the image feature of the target image as F X , then the above step S104 can be expressed as the following formula:

[0104] F X = ds2(F3) (4)

[0105] S105. Encode the image feature of the target image to obtain the encoded data of the target image.

[0106] Denote the encoding process as enc(), the image feature of the target image as F X , and the encoded data of the target image as B X , then the above step S105 can be expressed as the following formula:

[0107] B X = enc(F X ) (5)

[0108] In some embodiments, the feature encoding module 35 in the encoder shown in Figure 5 can be used to encode the image feature F of the target image X to obtain the encoded data B of the target image X .

[0109] When the image encoding method provided by the embodiment of the present application encodes a target image, first, feature extraction is performed on the target image through a first reversible block sequence composed of n cascaded reversible modules to obtain a first feature. Secondly, channel downsampling processing is performed on the first feature through a first channel downsampling module to obtain a second feature. Then, the second feature is processed through a second reversible block sequence composed of m cascaded reversible modules to obtain a third feature. Next, channel downsampling processing is performed on the third feature through a second channel downsampling module to obtain the image feature of the target image. Finally, the image feature of the target image is encoded to obtain the encoded data of the target image. Since the image encoding method provided by the embodiment of the present application splits the reversible neural network into two reversible block sequences and performs channel downsampling in multiple stages through two channel downsampling modules, the embodiment of the present application can reduce the channel redundancy in the features, thereby improving the encoding efficiency of the image encoder.

[0110] The embodiment of the present application also provides an image decoding method. Referring to Figure 11 as shown, the image decoding method includes the following steps S111 to S116:

[0111] S111. Obtain the encoded data of the target image.

[0112] Among them, the encoded data of the target image is the encoded data obtained by encoding the image feature of the target image, and the image feature of the target image is the feature obtained by processing the target image through a first reversible block sequence, a first channel downsampling module, a second reversible block sequence, and a second channel downsampling module in sequence; the first reversible block sequence is composed of n cascaded reversible modules, and the second reversible block sequence is composed of m cascaded reversible modules, where n and m are positive integers.

[0113] The implementation manner of encoding the target image to obtain the encoded data of the target image can refer to the above steps S101 to S105. To avoid repetition, it will not be described again here.

[0114] In some embodiments, obtaining the encoded data of the target image may be receiving the encoded data of the target image sent by a media resource server, or reading the encoded data of the target image from a local memory. The embodiment of the present application does not limit this, as long as the encoded data of the target image can be obtained.

[0115] S112. Decode the encoded data of the target image to obtain the reconstructed image feature of the target image.

[0116] Represent the image feature as B X , and represent the reconstructed feature of the target image as If the decoding process is represented as dec(), then the above step S112 can be expressed as the following formula:

[0117]

[0118] In some embodiments, it is possible to Figure 9 decode the encoded data of the target image through the feature decoding module 61 in the decoder shown in the figure to obtain the reconstructed features of the target image.

[0119] S113. Perform channel upsampling processing on the reconstructed image features through the first channel upsampling module to obtain the reconstructed third feature.

[0120] In some embodiments, the upsampling multiple of the first channel upsampling module for performing channel upsampling processing on the reconstructed image features is 3. That is, if the feature tensor of the reconstructed image features is 128 / 3C*H*W, then the feature tensor of the reconstructed third feature is 128C*H*W.

[0121] Represent the upsampling process of the first channel upsampling module as us1(), and the reconstructed third feature as The reconstructed image features are represented as Then the above step S113 can be expressed as the following formula:

[0122]

[0123] S114. Perform reverse processing on the reconstructed third feature through the second reversible block sequence to obtain the reconstructed second feature.

[0124] Represent the reconstructed third feature as The reconstructed second feature is represented as The reverse processing of the i-th level reversible module is represented as f i -1 (), then the above step S114 can be expressed as the following formula:

[0125]

[0126] S115. Perform channel upsampling processing on the reconstructed second feature through the second channel upsampling module to obtain the reconstructed first feature.

[0127] In some embodiments, the upsampling multiple of the second channel upsampling module for performing channel upsampling processing on the reconstructed second feature is 2. That is, if the feature tensor of the reconstructed second feature is 32C*H*W, then the feature tensor of the reconstructed first feature is 64C*H*W.

[0128] Represent the upsampling process of the first channel upsampling module as us2(), and the reconstructed second feature as Reconstruct the first feature representation as Then the above step S115 can be expressed as the following formula:

[0129]

[0130] S116. Perform reverse processing on the reconstructed first feature through the first reversible block sequence to obtain the reconstructed image of the target image.

[0131] Express the reconstructed first feature as Express the reconstructed image of the target image as The reverse processing of the i-th level reversible module of the first reversible block sequence is expressed as f i -1 (), then the above step S116 can be expressed as the following formula:

[0132]

[0133] The image decoding method provided by the embodiments of the present application obtains the encoded data of the target image. First, decode the encoded data of the target image to obtain the reconstructed image features of the target image. Secondly, perform channel upsampling processing on the reconstructed image features through the first channel upsampling module to obtain the reconstructed third feature. Then, perform reverse processing on the reconstructed third feature through the second reversible block sequence composed of m cascaded reversible modules to obtain the reconstructed second feature. Then, perform channel upsampling processing on the reconstructed second feature through the second channel upsampling module to obtain the reconstructed first feature. Finally, perform reverse processing on the reconstructed first feature through the first reversible block sequence to obtain the reconstructed image of the target image. Since the image decoding method provided by the embodiments of the present application splits the reversible neural network into two reversible block sequences and performs channel upsampling in multiple stages through two channel upsampling modules, the embodiments of the present application can decode the encoded data obtained by the image decoding method provided by the above embodiments to obtain the reconstructed image of the target image, thereby reducing the channel redundancy in the features and improving the encoding efficiency of the image encoder while ensuring the normal reconstruction of the image.

[0134] As an extension and refinement of the above embodiments, the embodiments of the present application provide another image decoding method. Refer to Figure 12 As shown, the image decoding method includes the following steps:

[0135] S1201. Obtain the encoded data of the target image.

[0136] Among them, the encoded data of the target image is the encoded data obtained by encoding the image features of the target image, and the image features of the target image are the features obtained by processing the target image through a first reversible block sequence, a first channel downsampling module, a second reversible block sequence, and a second channel downsampling module in sequence; the first reversible block sequence is composed of n cascaded reversible modules, and the second reversible block sequence is composed of m cascaded reversible modules, where n and m are positive integers.

[0137] S1202. Decode the encoded data of the target image to obtain the reconstructed image features of the target image.

[0138] S1203. The first convolutional layer performs convolutional processing on the reconstructed image features of the target image to obtain first convolutional features.

[0139] Represent the convolutional processing of the first convolutional layer as conv1 3*3 (), and represent the reconstructed image features of the target image as Represent the first convolutional features as F conv1 , then the above step S1203 can be expressed as the following formula:

[0140]

[0141] S1204. The activation function layer performs activation processing on the first convolutional features based on a preset activation function to obtain first activation features.

[0142] Represent the first convolutional features as F conv1 , represent the first activation features as F ReLU1 , and represent the activation processing of the activation function layer as ReLU(), then the above step S1204 can be expressed as the following formula:

[0143] F ReLU1 = ReLU(F conv1 ) (12)

[0144] S1205. The second convolutional layer performs convolutional processing on the first activation features to obtain second convolutional features.

[0145] Represent the convolutional processing of the first convolutional layer as conv2 3*3 (), represent the first activation features as F ReLU1 , represent the second convolutional features as F conv2 , then the above step S1205 can be expressed as the following formula:

[0146] F conv2 = conv2 3*3 (F ReLU1 ) (13)

[0147] S1206. The third convolutional layer performs convolutional processing on the reconstructed image features of the target image to obtain third convolutional features.

[0148] Represent the convolutional processing of the third convolutional layer as conv3 3*3 (), and represent the reconstructed image features of the target image as Represent the third convolutional features as F conv3 , then the above step S1206 can be expressed as the following formula:

[0149]

[0150] S1207. The addition fusion layer adds and fuses the second convolutional features and the third convolutional features to obtain reconstructed third features.

[0151] Represent the second convolutional features as F conv2 , and represent the third convolutional features as F conv3 , and represent the reconstructed third features as The above step S1207 can be expressed as the following formula:

[0152]

[0153] S1208. Perform reverse processing on the reconstructed third features through the second invertible block sequence to obtain reconstructed second features.

[0154] S1209. The first convolutional layer performs convolutional processing on the reconstructed second features to obtain fourth convolutional features.

[0155] Represent the convolutional processing of the first convolutional layer as conv1 3*3 (), and represent the reconstructed second features as Represent the fourth convolutional features as F conv4 , then the above step S1208 can be expressed as the following formula:

[0156]

[0157] S1210. The activation function layer performs activation processing on the fourth convolutional features based on a preset activation function to obtain second activation features.

[0158] Represent the fourth convolutional features as F conv4 , and represent the activation features as F ReLU2 , and represent the activation processing of the activation function layer as ReLU(), then the above step S1210 can be expressed as the following formula:

[0159] F ReLU2 =ReLU(F conv4 ) (17)

[0160] S1211. The second convolutional layer performs convolutional processing on the second activation feature to obtain a fifth convolutional feature.

[0161] Represent the convolutional processing of the first convolutional layer as conv2 3*3 (), represent the second activation feature as F ReLU2 , and represent the fifth convolutional feature as F conv5 , then the above step S1205 can be expressed as the following formula:

[0162] F conv5 = conv2 3*3 (F ReLU2 ) (18)

[0163] S1212. The third convolutional layer performs convolutional processing on the reconstructed second feature to obtain a sixth convolutional feature.

[0164] Represent the convolutional processing of the third convolutional layer as conv3 3*3 (), and represent the reconstructed second feature as Represent the sixth convolutional feature as F conv6 , then the above step S1212 can be expressed as the following formula:

[0165]

[0166] S1213. The addition fusion layer performs addition fusion on the fifth convolutional feature and the sixth convolutional feature to obtain a reconstructed first feature.

[0167] Represent the fifth convolutional feature as F conv5 , represent the sixth convolutional feature as F conv6 , and represent the reconstructed first feature as The above step S1213 can be expressed as the following formula:

[0168]

[0169] S1214. Perform reverse processing on the reconstructed first feature through the first invertible block sequence to obtain the reconstructed image of the target image.

[0170] In some embodiments, the image decoding method provided by the embodiments of the present application further includes: training each network model in the image encoder and the image decoder based on a training data set.

[0171] Among them, the training data set includes: a plurality of sample images. The process of training each network model in the image encoder and the image decoder based on the training data set may include: first, inputting the sample images into the image encoder for encoding to obtain the encoded data of the images, then decoding the encoded data through the image decoder to obtain the reconstructed images of the sample images, then calculating the loss value according to the sample images and the reconstructed images, and adjusting the parameters of the network models in the image encoder and / or the image decoder according to the loss value until the image encoder and the image decoder converge.

[0172] In some embodiments, the data set provided by flicker_2W can be used as the training data set.

[0173] In some embodiments, the data set provided by kodak can also be used as the test data set to test the image encoding and decoding framework.

[0174] In some embodiments, when training each network model in the image encoder and the image decoder based on the training data set, the batch size can be set to 24.

[0175] In some embodiments, some embodiments of the present application provide an image decoding device, which includes:

[0176] A memory configured to store a computer program;

[0177] A processor configured to, when calling the computer program, enable the image decoding device to implement the image decoding method described in any of the above embodiments.

[0178] In some embodiments, some embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is enabled to implement the image decoding method described in any of the above embodiments or the image decoding method described in any of the above embodiments.

[0179] In some embodiments, some embodiments of the present application provide a computer program product. When the computer program product runs on a computer, the computer is enabled to implement the image decoding method described in any of the above embodiments or the image decoding method described in any of the above embodiments.

[0180] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0181] For the sake of explanation, the above description has been presented in connection with specific embodiments. However, the above exemplary discussions are not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. According to the above teachings, many modifications and variations can be obtained. The selection and description of the above embodiments are for the purpose of better explaining the principles and actual applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. An image encoding method, characterized in that, Comprising: Performing feature extraction on a target image through a first reversible block sequence to obtain first features, where the first reversible block sequence consists of n cascaded reversible modules, and n is a positive integer; Performing channel downsampling processing on the first features through a first channel downsampling module to obtain second features; Processing the second features through a second reversible block sequence to obtain third features, where the second reversible block sequence consists of m cascaded reversible modules, and m is a positive integer; Performing channel downsampling processing on the third features through a second channel downsampling module to obtain image features of the target image; Encoding the image features of the target image to obtain encoded data of the target image.

2. The method according to claim 1, wherein The first reversible block sequence consists of 3 cascaded reversible modules.

3. The method according to claim 1, characterized in that, The performing channel downsampling processing on the first features through a first channel downsampling module to obtain second features includes: Performing channel downsampling processing on the first features in a channel downsampling manner using average pooling to obtain second features.

4. The method according to claim 1, characterized in that, The second reversible block sequence consists of 1 reversible module.

5. The method according to claim 1, wherein The performing channel downsampling processing on the third features through a second channel downsampling module to obtain fourth features includes: Performing channel downsampling processing on the third features in a channel downsampling manner using average pooling to obtain fourth features.

6. The method according to any one of claims 1-5, characterized in that, The reversible module consists of a channel rearrangement module, a convolutional layer, a first coupling layer, a second coupling layer, and a third coupling layer connected in series in sequence.

7. An image decoding method, characterized in that, Comprising: Obtaining encoded data of a target image, where the encoded data of the target image is encoded data obtained by encoding the image features of the target image, and the image features of the target image are features obtained by processing the target image through a first reversible block sequence, a first channel downsampling module, a second reversible block sequence, and a second channel downsampling module in sequence; The first reversible block sequence consists of n cascaded reversible modules, and the second reversible block sequence consists of m cascaded reversible modules, where n and m are positive integers; Decoding the encoded data of the target image to obtain reconstructed image features of the target image; Performing channel upsampling processing on the reconstructed image features through a first channel upsampling module to obtain reconstructed third features; Performing reverse processing on the reconstructed third features through the second reversible block sequence to obtain reconstructed second features; Performing channel upsampling processing on the reconstructed second features through a second channel upsampling module to obtain reconstructed first features; Performing reverse processing on the reconstructed first features through the first reversible block sequence to obtain a reconstructed image of the target image.

8. The method according to claim 7, characterized in that, The first channel upsampling module and / or the second channel upsampling module includes: A first convolutional layer for performing convolutional processing on the input features of the channel upsampling module to obtain first convolutional features; An activation function layer for performing activation processing on the first convolutional features based on a preset activation function to obtain activation features; A second convolutional layer for performing convolutional processing on the activation features to obtain second convolutional features; The third convolutional layer is used to perform convolutional processing on the input features of the channel upsampling module to obtain third convolutional features; The addition fusion layer is used to add and fuse the second convolutional features and the third convolutional features to obtain the output features of the channel upsampling module.

9. The method according to claim 8, wherein The first convolutional layer, the second convolutional layer, and the third convolutional layer are all convolutional layers with a convolutional kernel size of 3*3; The ratio of the number of channels of the input features to the number of channels of the output features of the first convolutional layer and the third convolutional layer is N, and N is greater than 1; The ratio of the number of channels of the input features to the number of channels of the output features of the second convolutional layer is 1.

10. The method according to claim 8, wherein The activation function layer is used to perform activation processing on the first convolutional features based on the rectified linear activation function to obtain the activation features.

11. An image encoding device, characterized in that, Comprising: A memory configured to store a computer program; A processor configured to cause the image encoding device to implement the image encoding method according to any one of claims 1-6 when the computer program is called.

12. An image decoding device, characterized in that, Comprising: A memory configured to store a computer program; A processor configured to cause the image decoding device to implement the image decoding method according to any one of claims 7-10 when the computer program is called.