Image encoding method, decoding method and device
By dividing the image into multiple image blocks and using neural network models to obtain the size information of the three-dimensional feature blocks for encoding, the problem of low image encoding and decoding efficiency is solved, and high compression ratio and clear image reconstruction are achieved.
Patent Information
- Application Number
- CN202110486600.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-22
- Filing Date
- 2021-04-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-04-30
AI Technical Summary
In the prior art, image encoding and decoding efficiency is low, and it is difficult to meet the users' growing image transmission and storage needs.
By dividing the image to be processed into multiple image blocks and inputting them into the neural network model, the size information of the three-dimensional feature blocks is obtained for encoding and generating an encoded code stream to improve the encoding and decoding efficiency.
It realizes a high compression ratio encoding code stream, ensuring that the decoding end can accurately reconstruct clear images, and improves image encoding and decoding efficiency.
Smart Images

Figure CN114125446B_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202010575457.0 and the application title "An Image Encoding and Decoding Method and Device", which was filed with the Chinese Patent Office on June 22, 2020, and the entire content is incorporated herein by reference. Technical Field
[0002] This application relates to the field of video or image compression technology based on artificial intelligence (AI), and in particular, to an image encoding method, a decoding method, and a device. Background Art
[0003] With the development of mobile communication technology and Internet technology, there are more and more multimedia contents mainly based on images. As an important carrier of information, images have many characteristics such as being intuitive, easy to understand, vivid, etc., and are an efficient information dissemination means, which are widely used in various fields such as life entertainment, road monitoring, media, and medical care.
[0004] By encoding an image, under the condition of meeting a certain quality, the image or the information contained in the image can be represented with fewer bits, thereby reducing the bandwidth resources required for transmitting the image and reducing the storage resources required for storing the image. Usually, technologies such as transformation, quantization, and entropy encoding can be used to eliminate the redundant information of the image, realize the encoding of the image, store and transmit high-quality images with a lower amount of data, and effectively transmit information.
[0005] Therefore, in order to meet the growing image transmission and storage needs of users, how to improve the image encoding and decoding efficiency has become an urgent technical problem to be solved. Summary of the Invention
[0006] This application provides an image encoding method, a decoding method, and a device to improve the image encoding and decoding efficiency.
[0007] In a first aspect, an embodiment of this application provides an image encoding method, which may include: obtaining an image to be processed. Dividing the image to be processed into at least two image blocks. Inputting the at least two image blocks into a first neural network model to obtain at least two three-dimensional feature blocks output by the first neural network model. The at least two image blocks and the at least two three-dimensional feature blocks may correspond one by one. Encoding the at least two three-dimensional feature blocks and encoding the size information of the at least two three-dimensional feature blocks to obtain an encoded bitstream.
[0008] Wherein, the size information of the at least two three-dimensional feature blocks may include: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, or the size of the three-dimensional feature blocks; the size information of the at least two image blocks and the parameters of the first neural network model are used to determine the size of the at least two three-dimensional feature blocks.
[0009] In this implementation manner, by encoding the size information of at least two three-dimensional feature blocks and the at least two three-dimensional feature blocks, an encoded bitstream is obtained. The size information of the at least two three-dimensional feature blocks is transmitted to the decoding end through the encoded bitstream, so that the decoding end can perform decoding based on the size information of the at least two three-dimensional feature blocks to obtain a reconstructed image. Since the size information of the three-dimensional feature blocks is added during the image encoding and decoding process, the encoded bitstream at the encoding end can have a high compression ratio, and the decoding end can decode the bitstream with a high compression ratio to obtain a reconstructed image, and it can ensure that the reconstructed image is relatively clear, so the image encoding and decoding efficiency can be improved.
[0010] Feature extraction is performed through the first neural network module to obtain an encoded bitstream. Since the neural network model has a relatively deep image modeling and expression ability, the encoded bitstream can have a high compression ratio to improve the image encoding efficiency.
[0011] In a possible design, the size information of the at least two image blocks may include the height and width of the image blocks. Or the size information of the at least two image blocks includes the partitioning method applied to the image to be processed and used to partition the at least two image blocks.
[0012] In this implementation manner, by transmitting the height and width of the at least two image blocks, or the partitioning method of the at least two image blocks to the decoding end, when the size information of the image blocks at the encoding end changes, the decoding end can accurately obtain the size information of the image blocks, and then accurately obtain the size information of the three-dimensional feature blocks, and accurately decode the encoded data of the three-dimensional feature blocks based on the size information of the three-dimensional feature blocks to obtain a reconstructed image.
[0013] In a possible design, the parameters of the first neural network model may include at least one of the number of channels of the convolutional kernel or the scaling step size.
[0014] In this implementation manner, by transmitting the parameters of the first neural network model to the decoding end, when the parameters of the first neural network model at the encoding end change, the decoding end can accurately obtain the parameters of the first neural network model, and then accurately obtain the size information of the three-dimensional feature blocks, and accurately decode the encoded data of the three-dimensional feature blocks based on the size information of the three-dimensional feature blocks to obtain a reconstructed image.
[0015] In a possible design, the sizes of the at least two three-dimensional feature blocks may include the length, width, and height of the at least two three-dimensional feature blocks, and the size information of the at least two image blocks may include the width and height of the at least two image blocks. The corresponding relationship between the length, width, and height of any one of the at least two three-dimensional feature blocks and the width and height of the corresponding one of the at least two image blocks is as follows:
[0016] M × N × R = W / S × H / S × C
[0017] Wherein, M is the length of any one of the three-dimensional feature blocks, N is the width of any one of the three-dimensional feature blocks, R is the height of any one of the three-dimensional feature blocks, W is the width of any one of the image blocks, H is the height of any one of the image blocks, C is the number of channels of the convolutional kernel, and S is the scaling step size.
[0018] In a possible design, the sizes of the at least two three-dimensional feature blocks may include the length, width, and height of the three-dimensional feature blocks.
[0019] In a possible design, any one of the at least two three-dimensional feature blocks includes multiple feature values. Encoding the at least two three-dimensional feature blocks may include: estimating the probability distribution of the feature values (which may also be referred to as three-dimensional feature values) of any one of the at least two three-dimensional feature blocks according to the size of any one of the at least two three-dimensional feature blocks, to obtain the probability distribution vector of the feature values. Entropy encoding any one of the three-dimensional feature blocks according to the probability distribution vector of the feature values.
[0020] In a possible design, estimating the probability distribution of the feature values of any one of the three-dimensional feature blocks according to the size of any one of the three-dimensional feature blocks, to obtain the probability distribution vector of the feature values, may include: determining the context information of the feature values according to the size of any one of the three-dimensional feature blocks, and the context information may include the encoded feature values in the neighborhood of the feature values determined according to the size of any one of the three-dimensional feature blocks. Inputting the context information of the feature values into the second neural network model. Obtaining the probability distribution vector of the feature values output by the second neural network model.
[0021] In this implementation manner, estimating the probability distribution of the feature values of the three-dimensional feature blocks through a neural network model, and performing entropy encoding on the three-dimensional feature blocks according to the probability distribution vector of the feature values to obtain an encoded bitstream, can achieve a relatively high compression ratio while ensuring that the reconstructed image is relatively clear and preserving more detailed textures.
[0022] In a possible design, encoding the at least two three-dimensional feature blocks may include: performing quantization processing on the at least two three-dimensional feature blocks to obtain at least two three-dimensional quantized feature blocks, and any one of the at least two three-dimensional quantized blocks includes multiple quantized feature values. Estimating the probability distribution of the quantized feature values (which may also be referred to as three-dimensional quantized feature values) of any one of the at least two three-dimensional quantized feature blocks according to the size of any one of the at least two three-dimensional quantized feature blocks, to obtain the probability distribution vector of the quantized feature values. The size of any one of the at least two three-dimensional quantized feature blocks is the same as the size of the corresponding any one of the three-dimensional feature blocks. Entropy encoding any one of the at least two three-dimensional quantized feature blocks according to the probability distribution vector of the quantized feature values.
[0023] In a possible design, according to the size of any one of the three-dimensional quantization feature blocks, a probability distribution estimation is performed on the quantization feature value of any one of the three-dimensional quantization feature blocks to obtain a probability distribution vector of the quantization feature value, which may include: determining context information of the quantization feature value according to the size of any one of the three-dimensional quantization feature blocks, where the context information includes the encoded feature values in the neighborhood of the quantization feature value determined according to the size of any one of the three-dimensional quantization feature blocks. Inputting the context information of the quantization feature value into a third neural network model. Obtaining the probability distribution vector of the quantization feature value output by the third neural network model.
[0024] In a second aspect, an embodiment of the present application provides an image decoding method, which may include: obtaining a bitstream to be decoded. The bitstream to be decoded may include encoded data of at least two three-dimensional feature blocks and size information of the at least two three-dimensional feature blocks. Obtaining at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks. Reconstructing at least two image blocks of the image to be processed according to the at least two three-dimensional feature blocks.
[0025] In a possible design, the size information of the at least two three-dimensional feature blocks may include the sizes of the at least two three-dimensional feature blocks, and the sizes of the at least two three-dimensional feature blocks may include the height and width of the image blocks.
[0026] In a possible design, the size information of the at least two three-dimensional feature blocks may include: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, where the size information of the at least two image blocks and the parameters of the first neural network model are used to determine the sizes of the at least two three-dimensional feature blocks.
[0027] In a possible design, the size information of the at least two image blocks may include the height and width of the at least two image blocks.
[0028] In a possible design, the parameters of the first neural network model may include at least one of the number of channels of the convolutional kernel or the scaling step size. The size information of the at least two image blocks may include the width and height of the at least two image blocks;
[0029] The corresponding relationship between the length, width, and height of any one of the at least two three-dimensional feature blocks and the width and height of the corresponding one of the at least two image blocks is as follows:
[0030] M×N×R=W / S×H / S×C;
[0031] Wherein, M is the length of any one of the three-dimensional feature blocks, N is the width of any one of the three-dimensional feature blocks, R is the height of any one of the three-dimensional feature blocks, W is the width of any one of the image blocks, H is the height of any one of the image blocks, C is the number of channels of the convolutional kernel, and S is the scaling step size.
[0032] In a possible design, the at least two three-dimensional feature blocks may include: multiple eigenvalues of any one of the at least two three-dimensional feature blocks, and size information of the at least two three-dimensional feature blocks.
[0033] In a possible design, obtaining the at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks may include: estimating the probability distribution of the eigenvalues of any one of the three-dimensional feature blocks according to the size information of any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues. Entropy decoding the encoded data of any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalues to obtain the eigenvalues.
[0034] In a possible design, obtaining the at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks may include: estimating the probability distribution of the eigenvalues of any one of the three-dimensional feature blocks according to the size information of any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues. Entropy decoding the encoded data of any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalues to obtain the quantized eigenvalues of the eigenvalues. Performing inverse quantization processing on the quantized eigenvalues to obtain the eigenvalues.
[0035] In a possible design, estimating the probability distribution of the eigenvalues of any one of the three-dimensional feature blocks according to the size information of any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues may include: determining context information of the eigenvalues according to the size information of any one of the three-dimensional feature blocks, where the context information includes the decoded eigenvalues in the neighborhood of the eigenvalues determined according to the size information of any one of the three-dimensional feature blocks. Inputting the context information of the eigenvalues into a second neural network model. Obtaining the probability distribution vector of the eigenvalues output by the second neural network model.
[0036] In a possible design, reconstructing at least two image blocks of an image to be processed according to the at least two three-dimensional feature blocks includes: inputting the at least two three-dimensional feature blocks into a third neural network model to obtain reconstructed image blocks of the at least two image blocks output by the third neural network model.
[0037] In a third aspect, an embodiment of the present application provides an image encoding device, which may include functional modules for implementing the functions of the method described in the first aspect or any possible design of the first aspect. For example, the image encoding device may include an acquisition module for acquiring an image to be processed, a processing module for dividing the image to be processed into at least two image blocks, inputting the at least two image blocks into a first neural network model to obtain at least two three-dimensional feature blocks output by the first neural network model, where the at least two image blocks and the at least two three-dimensional feature blocks may correspond one by one, encoding the at least two three-dimensional feature blocks, and encoding the size information of the at least two three-dimensional feature blocks to obtain an encoded bitstream.
[0038] Among them, the size information of the at least two three-dimensional feature blocks may include: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, or the size of the three-dimensional feature blocks; the size information of the at least two image blocks and the parameters of the first neural network model are used to determine the size of the at least two three-dimensional feature blocks.
[0039] In a possible design, the size information of the at least two image blocks may include the height and width of the image blocks. Or the size information of the at least two image blocks includes the division method applied to the image to be processed and used to divide the at least two image blocks.
[0040] In a possible design, the parameters of the first neural network model may include at least one of the number of channels of the convolution kernel or the scaling step.
[0041] In a possible design, the size of the at least two three-dimensional feature blocks may include the length, width, and height of the at least two three-dimensional feature blocks, and the size information of the at least two image blocks may include the width and height of the at least two image blocks. The correspondence between the length, width, and height of any one of the at least two three-dimensional feature blocks and the width and height of the corresponding one of the at least two image blocks is as follows:
[0042] M×N×R=W / S×H / S×C.
[0043] Wherein, M is the length of any one of the three-dimensional feature blocks, N is the width of any one of the three-dimensional feature blocks, R is the height of any one of the three-dimensional feature blocks, W is the width of any one of the image blocks, H is the height of any one of the image blocks, C is the number of channels of the convolution kernel, and S is the scaling step.
[0044] In a possible design, the size information of the at least two three-dimensional feature blocks may include the length, width, and height of the three-dimensional feature blocks.
[0045] In a possible design, the processing module is configured to: estimate the probability distribution of the eigenvalue of any one of at least two three-dimensional feature blocks according to the size of the any one three-dimensional feature block, so as to obtain the probability distribution vector of the eigenvalue. Perform entropy coding on the any one three-dimensional feature block according to the probability distribution vector of the eigenvalue.
[0046] In a possible design, the processing module is configured to: determine the context information of the eigenvalue according to the size of the any one three-dimensional feature block, where the context information may include the eigenvalues that have been encoded in the neighborhood of the eigenvalue determined according to the size of the any one three-dimensional feature block. Input the context information of the eigenvalue into the second neural network model. Obtain the probability distribution vector of the eigenvalue output by the second neural network model.
[0047] In a possible design, the processing module is configured to: perform quantization processing on the at least two three-dimensional feature blocks to obtain at least two three-dimensional quantized feature blocks, and any one of the at least two three-dimensional quantized blocks includes a plurality of quantized eigenvalues. Estimate the probability distribution of the quantized eigenvalues of the any one three-dimensional quantized feature block according to the size of the any one three-dimensional quantized feature block, so as to obtain the probability distribution vector of the quantized eigenvalues. The size of the any one three-dimensional quantized feature block is the same as the size of the corresponding any one three-dimensional feature block. Perform entropy coding on the any one three-dimensional quantized feature block according to the probability distribution vector of the quantized eigenvalues.
[0048] In a possible design, the processing module is configured to: determine the context information of the quantized eigenvalue according to the size of the any one three-dimensional quantized feature block, where the context information includes the eigenvalues that have been encoded in the neighborhood of the quantized eigenvalue determined according to the size of the any one three-dimensional quantized feature block. Input the context information of the quantized eigenvalue into the third neural network model. Obtain the probability distribution vector of the quantized eigenvalue output by the third neural network model.
[0049] In a fourth aspect, an embodiment of the present application provides an image decoding device, which may include functional modules for implementing the method described in the second aspect or any possible design of the second aspect. For example, the image decoding device may include: an acquisition module configured to acquire a bitstream to be decoded, where the bitstream to be decoded may include the encoded data of at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks. A processing module configured to obtain at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks. Reconstruct at least two image blocks of the image to be processed according to the at least two three-dimensional feature blocks.
[0050] In a possible design, the size information of the at least two three-dimensional feature blocks may include the sizes of the at least two three-dimensional feature blocks, and the sizes of the at least two three-dimensional feature blocks may include the height and width of the image block.
[0051] In a possible design, the size information of the at least two three-dimensional feature blocks may include: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, where the size information of the at least two image blocks and the parameters of the first neural network model are used to determine the size of the at least two three-dimensional feature blocks.
[0052] In a possible design, the size information of the at least two image blocks may include the height and width of the at least two image blocks.
[0053] In a possible design, the parameters of the first neural network model may include at least one of the number of channels of the convolutional kernel or the scaling step. The size information of the at least two image blocks may include the width and height of the at least two image blocks;
[0054] The correspondence between the length, width, and height of any one of the at least two three-dimensional feature blocks and the width and height of the corresponding one of the at least two image blocks is as follows:
[0055] M×N×R = W / S×H / S×C;
[0056] Wherein, M is the length of any one of the three-dimensional feature blocks, N is the width of any one of the three-dimensional feature blocks, R is the height of any one of the three-dimensional feature blocks, W is the width of any one of the image blocks, H is the height of any one of the image blocks, C is the number of channels of the convolutional kernel, and S is the scaling step.
[0057] In a possible design, the processing module is configured to: estimate the probability distribution of the eigenvalue of any one of the three-dimensional feature blocks according to the size information of the three-dimensional feature block, to obtain the probability distribution vector of the eigenvalue. Entropy decode the encoded data of any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalue, to obtain the eigenvalue.
[0058] In a possible design, the processing module is configured to: estimate the probability distribution of the eigenvalue of any one of the three-dimensional feature blocks according to the size information of the three-dimensional feature block, to obtain the probability distribution vector of the eigenvalue. Entropy decode the encoded data of any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalue, to obtain the quantized eigenvalue of the eigenvalue. Perform an inverse quantization process on the quantized eigenvalue, to obtain the eigenvalue.
[0059] In a possible design, the processing module is configured to: determine context information of the feature value according to the size information of any one of the three-dimensional feature blocks, where the context information includes the decoded feature values within the neighborhood of the feature value determined according to the size information of any one of the three-dimensional feature blocks. Input the context information of the feature value into the second neural network model. Obtain the probability distribution vector of the feature value output by the second neural network model.
[0060] In a possible design, the processing module is configured to: input the at least two three-dimensional feature blocks into a third neural network model, and obtain the reconstructed image blocks of the at least two image blocks output by the third neural network model.
[0061] The method described in the first aspect of the embodiments of the present application can be executed by the device described in the third aspect of the embodiments of the present application. Other features and implementation manners of the method described in the first aspect of the embodiments of the present application directly depend on the functionality and implementation manners of the device described in the third aspect of the embodiments of the present application.
[0062] The method described in the second aspect of the embodiments of the present application can be executed by the device described in the fourth aspect of the embodiments of the present application. Other features and implementation manners of the method described in the second aspect of the embodiments of the present application directly depend on the functionality and implementation manners of the device described in the fourth aspect of the embodiments of the present application.
[0063] In a fifth aspect, an image encoding device provided by the embodiments of the present application may include: a non-volatile memory and a processor coupled to each other, and the processor calls program code stored in the memory to execute the method described in the first aspect or any possible design of the first aspect.
[0064] In a sixth aspect, an image decoding device provided by the embodiments of the present application may include: a non-volatile memory and a processor coupled to each other, and the processor calls program code stored in the memory to execute the method described in the second aspect or any possible design of the second aspect.
[0065] In a seventh aspect, an image encoding device provided by the embodiments of the present application includes: an encoder, and the encoder is configured to execute the method described in the first aspect or any possible design of the first aspect.
[0066] In an eighth aspect, an image decoding device provided by the embodiments of the present application includes: a decoder, and the decoder is configured to execute the method described in the second aspect or any possible design of the second aspect.
[0067] In a ninth aspect, a computer-readable storage medium provided by the embodiments of the present application includes an encoded bitstream obtained according to the method described in the first aspect or any possible design of the first aspect.
[0068] In a tenth aspect, an embodiment of the present application provides a computer-readable storage medium storing instructions that, when executed, cause one or more processors to encode video data. The instructions cause the one or more processors to execute the method in the first or second aspect or any possible design of the first or second aspect.
[0069] In an eleventh aspect, an embodiment of the present application provides a computer program product including program code that, when running, executes the method in the first or second aspect or any possible design of the first or second aspect.
[0070] In a twelfth aspect, an embodiment of the present application provides an encoder including a processing circuit for executing the method as described in the first aspect or any possible design of the first aspect.
[0071] In a thirteenth aspect, an embodiment of the present application provides a decoder including a processing circuit for executing the method as described in the second aspect or any possible design of the second aspect.
[0072] In a fourteenth aspect, an embodiment of the present application provides a decoder including: one or more processors; a non-transitory computer-readable storage medium coupled to the processors and storing a program executed by the processors, wherein when the program is executed by the processors, the decoder executes the method as described in the second aspect or any possible design of the second aspect.
[0073] In a fifteenth aspect, an embodiment of the present application provides an encoder including: one or more processors; a non-transitory computer-readable storage medium coupled to the processors and storing a program executed by the processors, wherein when the program is executed by the processors, the encoder executes the method as described in the first aspect or any possible design of the first aspect.
[0074] In a sixteenth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium including program code that, when executed by a computer device, is used to execute the method in the first or second aspect or any possible design of the first or second aspect.
[0075] In a seventeenth aspect, an embodiment of the present application provides a non-transitory storage medium, characterized by including a bitstream encoded using the method as described in the first aspect or any possible design of the first aspect.
[0076] In an eighteenth aspect, an embodiment of the present application provides an encoded bitstream of a video signal, the encoded bitstream including a plurality of syntax elements, the plurality of syntax elements including size information of a three-dimensional feature block.
[0077] In a nineteenth aspect, an embodiment of the present application provides a non-transitory storage medium, characterized by including an encoded bitstream decoded by an image decoding device, the bitstream being generated by dividing frames of a video signal or an image signal into a plurality of image blocks, the encoded bitstream including a plurality of syntax elements, and the plurality of syntax elements including size information of three-dimensional feature blocks.
[0078] In the image encoding method, decoding method, and device according to the embodiments of the present application, by dividing an image to be processed into a plurality of image blocks, inputting the plurality of image blocks into a first neural network model, obtaining a plurality of three-dimensional feature blocks output by the first neural network model, encoding the size information of the plurality of three-dimensional feature blocks and the plurality of three-dimensional feature blocks to obtain an encoded bitstream, and transmitting the size information of the plurality of three-dimensional feature blocks from an encoding end to a decoding end, so that the decoding end can perform decoding according to the size information of the plurality of three-dimensional feature blocks to obtain a reconstructed image. Since the size information of the plurality of three-dimensional feature blocks is added during the image encoding and decoding process, the encoded bitstream at the encoding end can have a high compression ratio, the decoding end can decode the bitstream with a high compression ratio to obtain a reconstructed image, and the relative clarity of the reconstructed image can be ensured. Therefore, the image encoding and decoding efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] Figure 1A A block diagram of an example of a video decoding system for implementing the embodiments of the present application, where the system uses a neural network to encode or decode video images;
[0080] Figure 1B A block diagram of another example of a video decoding system for implementing the embodiments of the present application, where the video encoder and / or video decoder uses a neural network to encode or decode video images;
[0081] Figure 2 A block diagram of an example of a video encoder for implementing the embodiments of the present application, where the video encoder 20 uses a neural network to encode video images;
[0082] Figure 3 A block diagram of an example of a video decoder for implementing the embodiments of the present application, where the video decoder 30 uses a neural network to decode video images;
[0083] Figure 4 A schematic block diagram of a video decoding device for implementing the embodiments of the present application;
[0084] Figure 5 A schematic block diagram of a video decoding device for implementing the embodiments of the present application;
[0085] Figures 6a - 6c A training schematic diagram of the neural network according to the embodiments of the present application;
[0086] Figures 7a - 7e The convolutional neural network shown is only an example diagram of a convolutional neural network;
[0087] Figure 8 It is a flowchart of image compression based on deep learning;
[0088] Figure 9 It is a schematic diagram of the processing process of the image coding method according to the embodiment of the present application;
[0089] Figure 10 It is a flowchart of the image coding method according to the embodiment of the present application;
[0090] Figure 11 It is a schematic diagram of the image block division according to the embodiment of the present application;
[0091] Figure 12 It is a schematic diagram of the processing process of the image decoding method according to the embodiment of the present application;
[0092] Figure 13 It is a flowchart of the image decoding method according to the embodiment of the present application;
[0093] Figure 14 It is a schematic diagram of the processing process of the image coding method according to the embodiment of the present application;
[0094] Figure 15 It is a flowchart of the image coding method according to the embodiment of the present application;
[0095] Figure 16 It is a schematic diagram of the processing process of the image decoding method according to the embodiment of the present application;
[0096] Figure 17 It is a flowchart of the image decoding method according to the embodiment of the present application;
[0097] Figure 18 It is a schematic flowchart of an end-to-end image coding method according to the embodiment of the present application;
[0098] Figure 19 It is a schematic block diagram of the image coding device according to the embodiment of the present application;
[0099] Figure 20 It is a schematic block diagram of the image decoding device according to the embodiment of the present application;
[0100] Figure 21 It is a schematic block diagram of the image coding device according to the embodiment of the present application;
[0101] Figure 22 It is a schematic block diagram of the image decoding device according to the embodiment of the present application. Detailed implementation manners
[0102] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application.
[0103] The terms "first", "second", etc. involved in the embodiments of the present application are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0104] It should be understood that in the embodiments of the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0105] The embodiments of the present application provide an AI-based video image compression technology, especially a neural network-based video compression technology, to improve the traditional hybrid video codec system.
[0106] Video encoding generally refers to processing an image sequence that forms a video or video sequence. In the field of video encoding, the terms "picture", "frame", or "image" can be used as synonyms. Video encoding (or generally referred to as encoding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and generally includes processing (e.g., compressing) the original video image to reduce the amount of data required to represent the video image (thus enabling more efficient storage and / or transmission). Video decoding is performed on the destination side and generally includes performing inverse processing relative to the encoder to reconstruct the video image. The "encoding" of the video image (or generally referred to as an image) involved in the embodiments should be understood as the "encoding" or "decoding" of the video image or video sequence. The encoding part and the decoding part are also collectively referred to as codec (encoding and decoding, CODEC).
[0107] In the case of lossless video coding, the original video image can be reconstructed, that is, the reconstructed video image has the same quality as the original video image (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, further compression is performed through quantization, etc., to reduce the amount of data required to represent the video image, and the decoder side cannot fully reconstruct the video image, that is, the quality of the reconstructed video image is lower or worse than that of the original video image.
[0108] Several video coding standards belong to "lossy hybrid video coding and decoding" (that is, combining spatial and temporal prediction in the pixel domain with 2D transform coding for applying quantization in the transform domain). Each image in a video sequence is usually divided into a set of non-overlapping blocks, and encoding is usually performed at the block level. In other words, the encoder usually processes and encodes the video at the block (video block) level. For example, prediction blocks are generated through spatial (intra-frame) prediction and temporal (inter-frame) prediction; the prediction blocks are subtracted from the current block (the currently processed / block to be processed) to obtain a residual block; the residual block is transformed and quantized in the transform domain to reduce the amount of data to be transmitted (compressed), and the decoder side applies the inverse processing part relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. Additionally, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same predictions (such as intra-frame prediction and inter-frame prediction) and / or reconstruct pixels for processing subsequent blocks.
[0109] In the following embodiments of the decoding system 10, the encoder 20 and the decoder 30 are described according to Figures 1A to 3 this.
[0110] Figure 1A FIG. is a schematic block diagram of an exemplary decoding system 10, for example, a video decoding system 10 (or simply referred to as the decoding system 10) that can utilize the technology of the present application. The video encoder 20 (or simply referred to as the encoder 20) and the video decoder 30 (or simply referred to as the decoder 30) in the video decoding system 10 represent devices that can be used to execute various techniques according to the various examples described in the present application.
[0111] As Figure 1A shown, the decoding system 10 includes a source device 12, and the source device 12 is used to provide encoded image data 21 such as encoded images to a destination device 14 for decoding the encoded image data 21.
[0112] The source device 12 includes an encoder 20, and additionally, optionally, may include an image source 16, a pre-processor (or pre-processing unit) 18 such as an image pre-processor, and a communication interface (or communication unit) 22.
[0113] The image source 16 can include or can be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer animation images or any type of device for obtaining and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images and / or any combination thereof (e.g., augmented reality (AR) images). The image source can be any type of memory or storage that stores any of the above images.
[0114] To distinguish the processing performed by the pre-processor (or pre-processing unit) 18, the image (or image data) 17 can also be referred to as the original image (or original image data) 17.
[0115] The pre-processor 18 is used to receive (original) image data 17 and pre-process the image data 17 to obtain pre-processed image (or pre-processed image data) 19. For example, the pre-processing performed by the pre-processor 18 can include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It can be understood that the pre-processing unit 18 can be an optional component.
[0116] The video encoder (or encoder) 20 is used to receive the pre-processed image data 19 and provide encoded image data 21 (which will be further described below according to Figure 2 etc.).
[0117] The communication interface 22 in the source device 12 can be used to: receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.
[0118] The destination device 14 includes a decoder 30, and additionally, optionally, can include a communication interface (or communication unit) 28, a post-processor (or post-processing unit) 32, and a display device 34.
[0119] The communication interface 28 in the destination device 14 is used to directly receive the encoded image data 21 (or any other processed version) from the source device 12 or from any other source device such as a storage device. For example, the storage device is an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0120] The communication interfaces 22 and 28 can be used to send or receive encoded image data (or encoded data) 21 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, etc., or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network and public network, or any combination of any type thereof.
[0121] For example, the communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a packet, and / or use any type of transport encoding or processing to process the encoded image data for transmission over the communication link or communication network.
[0122] The communication interface 28 corresponds to the communication interface 22. For example, it can be used to receive the transmitted data and process the transmitted data using any type of corresponding transport decoding or processing and / or de-encapsulation to obtain the encoded image data 21.
[0123] Both the communication interface 22 and the communication interface 28 can be configured as Figure 1A a unidirectional communication interface as indicated by the arrow of the corresponding communication channel 13 pointing from the source device 12 to the destination device 14 in
[0124] or a bidirectional communication interface, and can be used to send and receive messages, etc., to establish a connection, confirm, and exchange any other information related to the communication link and / or data transmission such as the transmission of encoded image data, etc. Figure 3 etc. as further described below.
[0125] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (which will be further described below according to
[0126] The display device 34 is configured to receive the post - processed image data 33 to display an image to a user, viewer, etc. The display device 34 may be or include any type of display for presenting the reconstructed image, for example, an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro - LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display screen.
[0127] The decoding system 10 further includes a training engine 25. The training engine 25 is configured to train the encoder 20 (especially the feature extraction module in the encoder 20) or the decoder 30 (especially the feature decoding module in the decoder 30) to process an input image or an image region or an image patch to generate a three - dimensional feature patch of the input image or the image region or the image patch.
[0128] In the embodiments of the present application, the training data includes: training image patches and corresponding training three - dimensional feature patches, for example, Figures 6a to 6c the images or image regions or image patches shown; where
[0129] The training data can be stored in a database (not shown). The training engine 25 trains a target model (for example: it can be a neural network for feature extraction, etc.) based on the training data. It should be noted that the embodiments of the present application do not limit the source of the training data. For example, the training data can be obtained from the cloud or other places for model training.
[0130] The image patches of the image to be processed are input into the target model after relevant pre - processing, so as to obtain three - dimensional feature patches. The target model in the embodiments of the present application can specifically be a convolutional neural network, which will be described in detail below in combination with Figures 7a - 7e to detail the target model.
[0131] The target model trained by the training engine 25 can be applied to the decoding systems 10, 40, for example, applied to Figure 1AThe source device 12 shown (such as the encoder 20) or the destination device 14 (such as the decoder 30). The training engine 25 can train a target model in the cloud, and then the decoding system 10 downloads and uses the target model from the cloud; alternatively, the training engine 25 can train a target model in the cloud and use the target model, and the decoding system 10 directly obtains the processing result from the cloud. For example, the training engine 25 trains a target model with filtering function, the decoding system 10 downloads the target model from the cloud, and then the loop filter 220 in the encoder 20 or the loop filter 320 in the decoder 30 can filter the input reconstructed image or image block according to the target model to obtain the filtered image or image block. For another example, the training engine 25 trains a target model with filtering function, the decoding system 10 does not need to download the target model from the cloud, and the encoder 20 or the decoder 30 transmits the reconstructed image or image block to the cloud, and the cloud filters the reconstructed image or image block through the target model to obtain the filtered image or image block and transmits it to the encoder 20 or the decoder 30.
[0132] It should be noted that the training engine 25 can also be used to train the encoder 20 (especially the entropy coding module in the encoder 20) or the decoder 30 (especially the entropy decoding module in the decoder 30) to process the three-dimensional feature blocks of the input image or image region or image block to generate the probability distribution vector of the eigenvalues of the three-dimensional feature blocks, and perform entropy coding on the three-dimensional feature blocks according to the probability distribution vector of the eigenvalues to obtain the encoded bitstream. Correspondingly, the training data for the target model used for probability distribution estimation can include: training three-dimensional feature blocks and the probability distribution vectors corresponding to the training three-dimensional feature blocks.
[0133] Although Figure 1A The source device 12 and the destination device 14 are shown as independent devices, but the device embodiments can also include both the source device 12 and the destination device 14 or the functions of both the source device 12 and the destination device 14 at the same time, that is, include both the source device 12 or the corresponding function and the destination device 14 or the corresponding function at the same time. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function can be implemented using the same hardware and / or software or through separate hardware and / or software or any combination thereof.
[0134] According to the description, Figure 1A The presence and (accurate) division of different units or functions in the source device 12 and / or the destination device 14 shown may vary according to the actual device and application, which is obvious to those skilled in the art.
[0135] The encoder 20 (such as the video encoder 20) or the decoder 30 (such as the video decoder 30) or both can be implemented by, for example, Figure 1BThe processing circuitry shown is implemented, for example, by one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding-specific processors, or any combination thereof. The encoder 20 may be implemented by the processing circuitry 46 to include the various modules discussed with reference to Figure 2 the encoder 20 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by the processing circuitry 46 to include the various modules discussed with reference to Figure 3 the decoder 30 and / or any other decoder system or subsystem described herein. The processing circuitry 46 may be used to perform the various operations discussed below. As Figure 5 shown, if part of the technology is implemented in software, the device may store the instructions of the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of this application. One of the video encoder 20 and the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, as Figure 1B shown.
[0136] The source device 12 and the destination device 14 may include any of a variety of devices, including any type of handheld or fixed device, such as, for example, a laptop or notebook computer, a cell phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, etc., and may or may not use any type of operating system. In some cases, the source device 12 and the destination device 14 may be equipped with components for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0137] In some cases, Figure 1AThe illustrated video decoding system 10 is merely exemplary, and the techniques provided by this application are applicable to video coding settings (e.g., video encoding or video decoding), which may not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from a local memory, sent over a network, and so on. A video encoding device may encode data and store the data in a memory, and / or a video decoding device may retrieve data from the memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but merely encode data into a memory and / or retrieve and decode data from the memory.
[0138] Figure 1B is an illustration of an example of a video decoding system 40 that includes Figure 2 a video encoder 20 and / or Figure 3 a video decoder 30 according to an exemplary embodiment. The video decoding system 40 may include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video codec implemented by processing circuitry 46), an antenna 42, one or more processors 43, one or more memory memories 44, and / or a display device 45.
[0139] As Figure 1B shown, the imaging device 41, the antenna 42, the processing circuitry 46, the video encoder 20, the video decoder 30, the processor 43, the memory memory 44, and / or the display device 45 are capable of communicating with each other. In different examples, the video decoding system 40 may include only the video encoder 20 or only the video decoder 30.
[0140] In some instances, antenna 42 can be used to transmit or receive an encoded bitstream of video data. Additionally, in some instances, display device 45 can be used to present video data. Processing circuitry 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Video decoding system 40 can also include optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Additionally, memory 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting instance, memory 44 can be implemented by cache memory. In other instances, processing circuitry 46 can include memory (e.g., cache, etc.) for implementing an image buffer and the like.
[0141] In some instances, video encoder 20 implemented by logic circuitry can include an image buffer (e.g., implemented by processing circuitry 46 or memory 44) and a graphics processing unit (e.g., implemented by processing circuitry 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include video encoder 20 implemented by processing circuitry 46 to implement various modules discussed with reference to Figure 2 and / or any other encoder system or subsystem described herein. The logic circuitry can be used to perform the various operations discussed herein.
[0142] In some instances, video decoder 30 can be implemented by processing circuitry 46 in a similar manner to implement video decoder 30 discussed with reference to Figure 3 and / or any other decoder system or subsystem described herein. In some instances, video decoder 30 implemented by logic circuitry can include an image buffer (implemented by processing circuitry 46 or memory 44) and a graphics processing unit (e.g., implemented by processing circuitry 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include video decoder 30 implemented by processing circuitry 46 to implement various modules discussed with reference to Figure 3 and / or any other decoder system or subsystem described herein.
[0143] In some examples, antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data, indicators, index values, mode selection data, etc. related to the encoded video frames as discussed herein, such as data related to encoded partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators as discussed, and / or data defining the encoded partitions). The video decoding system 40 may also include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. A display device 45 is used to present video frames.
[0144] It should be understood that for the examples described herein with reference to the video encoder 20, the video decoder 30 may be used to perform the reverse process. Regarding signaling syntax elements, the video decoder 30 may be used to receive and parse such syntax elements and accordingly decode the relevant video data. In some examples, the video encoder 20 may entropy encode the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 may parse such syntax elements and accordingly decode the relevant video data.
[0145] For ease of description, embodiments of the present application are described with reference to the Versatile video coding (VVC) reference software or the High-Efficiency Video Coding (HEVC) developed by the Joint Collaboration Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG). Those of ordinary skill in the art understand that the embodiments of the present application are not limited to HEVC or VVC.
[0146] Encoder and encoding method
[0147] Figure 2 FIG. is a schematic block diagram of an example of a video encoder 20 for implementing the technology of the present application. In Figure 2In the example, video encoder 20 includes an input end (or input interface) 201, a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output end (or output interface) 272. The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a segmentation unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The illustrated video encoder 20 may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.
[0148] See Figure 2 , the inter prediction module / intra prediction module / loop filtering module / XXX includes (is) a trained target model (also referred to as a neural network), and the neural network is used to process an input image or an image region or an image block to generate a predicted value of the input image block. For example, the neural network for inter prediction / intra prediction / loop filtering is used to receive an input image or an image region or an image block, for example, Figures 6a to 6c the illustrated input image data, and generate a predicted value of the input image or an image region or an image block. The neural network for inter prediction / intra prediction / loop filtering / XXX will be described in detail below in conjunction with Figures 7a - 7e the relevant content.
[0149] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 form the forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 form the backward signal path of the encoder, where the backward signal path of the encoder 20 corresponds to the signal path of the decoder (see the decoder 30 in Figure 3 ). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer 230, the inter prediction unit 244, and the intra prediction unit 254 also form the "built-in decoder" of the video encoder 20.
[0150] Images and image segmentation (images and blocks)
[0151] The encoder 20 can be used to receive an image (or image data) 17 through an input terminal 201, etc., for example, an image in an image sequence forming a video or a video sequence. The received image or image data can also be a preprocessed image (or preprocessed image data) 19. For simplicity, the following description uses the image 17. The image 17 can also be referred to as the current image or the image to be encoded (especially when distinguishing the current image from other images in video coding, other images such as previously encoded and / or decoded images in the same video sequence, i.e., the video sequence that also includes the current image).
[0152] (Digital) images are or can be regarded as two-dimensional arrays or matrices composed of pixel points with intensity values. The pixel points in the array can also be called pixels (pixel or pel, short for picture element). The number of pixel points in the array or image in the horizontal and vertical directions (or axes) determines the size and / or resolution of the image. To represent color, usually three color components are used, that is, the image can be represented as or include three pixel point arrays. In the RBG format or color space, the image includes corresponding red, green, and blue pixel point arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space, such as YCbCr, including a luminance component indicated by Y (sometimes also denoted by L) and two chrominance components denoted by Cb and Cr. The luminance component Y represents the luminance or gray-level intensity (for example, the two are the same in a grayscale image), while the two chrominance components Cb and Cr represent the chrominance or color information components. Accordingly, an image in the YCbCr format includes a luminance pixel point array of luminance pixel point values (Y) and two chrominance pixel point arrays of chrominance values (Cb and Cr). An image in the RGB format can be converted or transformed into the YCbCr format, and vice versa, and this process is also called color transformation or conversion. If the image is black and white, then the image can only include a luminance pixel point array. Accordingly, the image can be, for example, a luminance pixel point array in a monochrome format or a luminance pixel point array and two corresponding chrominance pixel point arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0153] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit ( Figure 2(not shown in the figure) for splitting the image 17 into multiple (usually non-overlapping) image blocks 203. These blocks may also be referred to as root blocks, macro blocks (H.264 / AVC), or coding tree blocks (CTBs), or coding tree units (CTUs) in the H.265 / HEVC and VVC standards. The splitting unit can be used to use the same block size for all images in the video sequence and the corresponding grid with a defined block size, or to change the block size between images or subsets of images or groups of images, and split each image into corresponding blocks.
[0154] In other embodiments, the video encoder can be used to directly receive the blocks 203 of the image 17, for example, one, several, or all of the blocks that make up the image 17. The image blocks 203 can also be referred to as current image blocks or image blocks to be encoded.
[0155] Similar to the image 17, the image block 203 is also or can be considered as a two-dimensional array or matrix composed of pixel points with intensity values (pixel point values), but the image block 203 is smaller than the image 17. In other words, the block 203 can include an array of pixel points (e.g., the luminance array in the case of a monochrome image 17 or the luminance array or chrominance arrays in the case of a color image) or three arrays of pixel points (e.g., one luminance array and two chrominance arrays in the case of a color image 17) or any other number and / or type of arrays according to the adopted color format. The number of pixel points in the horizontal and vertical directions (or axes) of the block 203 defines the size of the block 203. Accordingly, the block can be an array of M×N (M columns × N rows) pixel points, or an array of M×N transform coefficients, etc.
[0156] In one embodiment, Figure 2 The shown video encoder 20 is used to encode the image 17 block by block, for example, perform encoding and prediction on each block 203.
[0157] In one embodiment, Figure 2 The shown video encoder 20 can also be used to split and / or encode the image using slices (also called video slices), where the image can be split or encoded using one or more slices (usually non-overlapping). Each slice can include one or more blocks (e.g., coding tree units CTUs) or one or more groups of blocks (e.g., coding blocks (tiles) in the H.265 / HEVC / VVC standards and bricks in the VVC standard).
[0158] In one embodiment, Figure 2The video encoder 20 shown can also be used to segment and / or encode an image using slices / coding tree units (also referred to as video coding tree units) and / or coding units (also referred to as video coding units), where the image can be segmented or encoded using one or more slices / coding tree units (usually non-overlapping), each slice / coding tree unit may include one or more blocks (e.g., CTUs) or one or more coding units, etc., and each coding unit can be in a shape such as a rectangle, and may include one or more whole or partial blocks (e.g., CTUs).
[0159] Residual calculation
[0160] The residual calculation unit 204 is used to calculate the residual block 205 based on the image block 203 and the prediction block 265 (the prediction block 265 is introduced in detail later) in the following manner: for example, subtract the pixel value of the prediction block 265 from the pixel value of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.
[0161] Transformation
[0162] The transformation processing unit 206 is used to perform a discrete cosine transform (DCT) or a discrete sine transform (DST), etc. on the pixel values of the residual block 205 to obtain the transform coefficients 207 in the transform domain. The transform coefficients 207 can also be referred to as transform residual coefficients, representing the residual block 205 in the transform domain.
[0163] The transformation processing unit 206 can be used to apply an integer approximation of DCT / DST, such as the transformation specified for H.265 / HEVC. Compared with the orthogonal DCT transform, this integer approximation is usually scaled by a certain factor. In order to maintain the norm of the residual block after forward and inverse transformation processing, other scaling factors are used as part of the transformation process. The scaling factors are usually selected according to certain constraints, such as the power of 2 for shift operations, the bit depth of the transform coefficients, the trade-off between accuracy and implementation cost, etc. For example, specific scaling factors are specified for the inverse transformation by the inverse transformation processing unit 212 on the encoder 20 side (and for the corresponding inverse transformation by, for example, the inverse transformation processing unit 312 on the decoder 30 side), and correspondingly, the corresponding scaling factors can be specified for the forward transformation by the transformation processing unit 206 on the encoder 20 side.
[0164] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) can be used to output transform parameters such as the type of one or more transforms. For example, it can be directly output or output after being encoded or compressed by the entropy encoding unit 270, such that the video decoder 30 can receive and use the transform parameters for decoding.
[0165] Quantization
[0166] The quantization unit 208 is used to quantize the transform coefficients 207 through, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209. The quantized transform coefficients 209 can also be referred to as quantized residual coefficients 209.
[0167] The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, during quantization, an n-bit transform coefficient can be rounded down to an m-bit transform coefficient, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different degrees of scaling can be applied to achieve finer or coarser quantization. A smaller quantization step corresponds to finer quantization, while a larger quantization step corresponds to coarser quantization. The appropriate quantization step can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index of a predefined set of appropriate quantization steps. For example, a smaller quantization parameter can correspond to fine quantization (smaller quantization step), and a larger quantization parameter can correspond to coarse quantization (larger quantization step), and vice versa. Quantization can include dividing by the quantization step, and the corresponding or inverse dequantization performed by the dequantization unit 210 etc. can include multiplying by the quantization step. Embodiments according to some standards such as HEVC can be used to determine the quantization step using the quantization parameter. Generally, the quantization step can be calculated using a fixed-point approximation of an equation involving division based on the quantization parameter. Other scaling factors can be introduced for quantization and dequantization to restore the norm of the residual block that may be modified due to the scaling used in the fixed-point approximation of the equations for the quantization step and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream. Quantization is a lossy operation, where the larger the quantization step, the greater the loss.
[0168] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) can be used to output the quantization parameter (QP), for example, directly output or output after being encoded or compressed by the entropy encoding unit 270, such that the video decoder 30 can receive and use the quantization parameter for decoding.
[0169] Dequantization
[0170] The inverse quantization unit 210 is used to perform inverse quantization on the quantization coefficients by the quantization unit 208 to obtain dequantized coefficients 211. For example, the inverse quantization scheme of the quantization scheme performed by the quantization unit 208 is performed according to or using the same quantization step as the quantization unit 208. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211, corresponding to the transform coefficients 207. However, due to the loss caused by quantization, the inverse quantization coefficients 211 are usually not exactly the same as the transform coefficients.
[0171] Inverse transform
[0172] The inverse transform processing unit 212 is used to perform the inverse transform of the transform performed by the transform processing unit 206, for example, inverse discrete cosine transform (DCT) or inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 213.
[0173] Reconstruction
[0174] The reconstruction unit 214 (e.g., adder 214) is used to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain a reconstructed block 215 in the pixel domain. For example, the pixel point values of the reconstructed residual block 213 and the pixel point values of the prediction block 265 are added.
[0175] Filtering
[0176] The loop filter unit 220 (or simply referred to as "loop filter" 220) is used to filter the reconstructed block 215 to obtain a filtered block 221, or is generally used to filter reconstructed pixel points to obtain filtered pixel point values. For example, the loop filter unit is used to smoothly perform pixel transitions or improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 220 is shown as a loop filter in Figure 2 it may be implemented as a post-loop filter in other configurations. The filtered block 221 may also be referred to as a filtered reconstructed block 221.
[0177] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be used to output loop filter parameters (such as SAO filtering parameters, ALF filtering parameters, or LMCS parameters), for example, directly output or output after entropy coding by the entropy coding unit 270, such that the decoder 30 can receive and use the same or different loop filter parameters for decoding.
[0178] Decoded picture buffer
[0179] The decoded picture buffer (DPB) 230 can be a reference picture memory that stores reference picture data for use by the video encoder 20 when encoding video data. The DPB 230 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The decoded picture buffer 230 can be used to store one or more filtered blocks 221. The decoded picture buffer 230 can also be used to store other previously filtered blocks, such as previously reconstructed and filtered blocks 221, of the same current picture or different pictures such as previous reconstructed pictures, and can provide the complete previously reconstructed i.e., decoded picture (and corresponding reference blocks and pixels) and / or a partially reconstructed current picture (and corresponding reference blocks and pixels), for example, for inter prediction. The decoded picture buffer 230 can also be used to store one or more unfiltered reconstructed blocks 215, or generally unfiltered reconstructed pixels, e.g., reconstructed blocks 215 that are not filtered by the loop filter unit 220, or reconstructed blocks or reconstructed pixels that have not undergone any other processing.
[0180] Mode Selection (Partitioning and Prediction)
[0181] The mode selection unit 260 includes a partitioning unit 262, an inter prediction unit 244, and an intra prediction unit 254, and is configured to receive or obtain raw picture data such as the raw block 203 (the current block 203 of the current picture 17) and reconstructed picture data from the decoded picture buffer 230 or other buffers (e.g., a column buffer, not shown in the figure), e.g., filtered and / or unfiltered reconstructed pixels or reconstructed blocks of the same (current) picture and / or one or more previously decoded pictures. The reconstructed picture data is used as reference picture data required for prediction such as inter prediction or intra prediction to obtain a predicted block 265 or a predicted value 265.
[0182] The mode selection unit 260 can be used to determine or select a partitioning for the current block prediction mode (including no partitioning) and prediction mode (e.g., intra or inter prediction mode), and generate a corresponding predicted block 265 to calculate the residual block 205 and reconstruct the reconstructed block 215.
[0183] In one embodiment, the mode selection unit 260 may be used to select a partitioning and prediction mode (e.g., from the prediction modes supported or available to the mode selection unit 260), where the prediction mode provides the best match or the smallest residual (the smallest residual means better compression in transmission or storage), or provides the smallest signaling overhead (the smallest signaling overhead means better compression in transmission or storage), or considers or balances both of the above. The mode selection unit 260 may be used to determine the partitioning and prediction mode according to rate distortion Optimization (RDO), i.e., select the prediction mode that provides the smallest rate distortion optimization. Terms such as "best", "lowest", "optimal", etc. in this article do not necessarily refer to "the best", "the lowest", "the most optimal" in general, but may also refer to cases that meet the termination or selection criteria. For example, values exceeding or falling below a threshold or other limitations may result in a "sub-optimal choice", but will reduce complexity and processing time.
[0184] In other words, the partitioning unit 262 may be used to partition the images in the video sequence into a sequence of coding tree units (CTUs), and the CTU 203 may be further partitioned into smaller block parts or sub-blocks (again forming blocks), for example, by iteratively using quad-tree partitioning (QT), binary-tree partitioning (BT), triple-tree partitioning (TT), or any combination thereof, and used to perform prediction on, for example, each of the block parts or sub-blocks, where mode selection includes selecting the tree structure for partitioning the block 203 and selecting the prediction mode applied to each of the block parts or sub-blocks.
[0185] The partitioning (e.g., performed by the partitioning unit 262) and prediction processing (e.g., performed by the inter-frame prediction unit 244 and the intra-frame prediction unit 254) performed by the video encoder 20 will be described in detail below.
[0186] Partitioning
[0187] The splitting unit 262 may split (or partition) a coding tree unit 203 into smaller parts, such as small blocks in the shape of a square or a rectangle. For an image with an array of three pixel points, a CTU consists of an N×N block of luminance pixel points and two corresponding chrominance pixel point blocks. The maximum allowed size of the luminance block in the currently under-developed Versatile Video Coding (VVC) standard is specified as 128×128, but it may be specified as a value different from 128×128 in the future, such as 256×256. The CTUs of an image may be grouped / concentrated into slices / coding tree blocks, coding blocks, or tiles. A coding block covers a rectangular area of an image, and a coding block may be divided into one or more tiles. A tile consists of multiple CTU rows within a coding block. A coding block that is not divided into multiple tiles may be called a tile. However, a tile is a proper subset of a coding block and thus is not called a coding block. VVC supports two coding tree block modes, namely the raster scan slice / coding tree block mode and the rectangular slice mode. In the raster scan coding tree block mode, a slice / coding tree block contains a sequence of coding blocks in the raster scan of the coding blocks of an image. In the rectangular slice mode, a slice contains multiple tiles of an image, and these tiles together form a rectangular area of the image. The tiles within a rectangular slice are arranged in the tile raster scan order of the slice. These smaller blocks (which may also be called sub-blocks) may be further split into even smaller parts. This is also called tree splitting or hierarchical tree splitting, where a root block at the root tree level 0 (hierarchical level 0, depth 0), etc., can be recursively split into two or more blocks at the next lower tree level, such as nodes at tree level 1 (hierarchical level 1, depth 1). These blocks may in turn be split into two or more blocks at the next lower level, such as tree level 2 (hierarchical level 2, depth 2), etc., until the splitting ends (because an end criterion is met, such as reaching the maximum tree depth or the minimum block size). Blocks that are not further split are also called leaf blocks or leaf nodes of the tree. A tree split into two parts is called a binary-tree (BT), a tree split into three parts is called a ternary-tree (TT), and a tree split into four parts is called a quad-tree (QT).
[0188] For example, a coding tree unit (CTU) may be or include a CTB of luma pixel points, two corresponding CTBs of chroma pixel points of an image having three pixel point arrays, or a CTB of pixel points of a monochrome image or a CTB of pixel points of an image encoded using three independent color planes and a syntax structure (for encoding pixel points). Accordingly, a coding tree block (CTB) may be a block of N×N pixel points, where N may be set to a certain value such that a component is divided into CTBs, which is segmentation. A coding unit (CU) may be or include an encoded block of luma pixel points, two corresponding encoded blocks of chroma pixel points of an image having three pixel point arrays, or an encoded block of pixel points of a monochrome image or an encoded block of pixel points of an image encoded using three independent color planes and a syntax structure (for encoding pixel points). Accordingly, an encoded block (CB) may be a block of M×N pixel points, where M and N may be set to a certain value such that a CTB is divided into encoded blocks, which is segmentation.
[0189] For example, in an embodiment, according to HEVC, a coding tree unit (CTU) may be divided into multiple CUs by using a quadtree structure represented as a coding tree. A decision on whether to use inter (temporal) prediction or intra (spatial) prediction to encode an image region is made at the leaf CU level. Each leaf CU may be further divided into one, two, or four PUs according to the PU partition type. The same prediction process is used within one PU, and relevant information is transmitted to a decoder in units of PUs. After obtaining a residual block by applying a prediction process according to the PU partition type, a leaf CU may be segmented into transform units (TUs) according to another quadtree structure similar to the coding tree used for CUs.
[0190] For example, in an embodiment, according to the latest video coding standard currently under development (referred to as Versatile Video Coding (VVC)), a combined quadtree using nested multi-type trees (such as binary trees and ternary trees) is used to divide the segmentation structure for segmenting coding tree units. In the coding tree structure within a coding tree unit, a CU can be square or rectangular. For example, a coding tree unit (CTU) is first divided by a quadtree structure. The quadtree leaf nodes are further divided by a multi-type tree structure. The multi-type tree structure has four division types: vertical binary tree division (SPLIT_BT_VER), horizontal binary tree division (SPLIT_BT_HOR), vertical ternary tree division (SPLIT_TT_VER), and horizontal ternary tree division (SPLIT_TT_HOR). The multi-type tree leaf nodes are called coding units (CUs), unless the CU is too large for the maximum transform length. Such segments are used for prediction and transform processing without any further segmentation. In most cases, this means that the CU, PU, and TU have the same block size in the coding block structure of the quadtree nested multi-type tree. This exception occurs when the maximum supported transform length is less than the width or height of the color component of the CU. VVC has developed a unique signaling mechanism for the segmentation division information in the coding structure with a quadtree nested multi-type tree. In the signaling mechanism, the coding tree unit (CTU), as the root of the quadtree, is first divided by the quadtree structure. Then each quadtree leaf node (when large enough to be) is further divided into a multi-type tree structure. In the multi-type tree structure, it is indicated by the first identifier (mtt_split_cu_flag) whether the node is further divided. When the node is further divided, the division direction is first indicated by the second identifier (mtt_split_cu_vertical_flag), and then it is indicated by the third identifier (mtt_split_cu_binary_flag) whether the division is a binary tree division or a ternary tree division. According to the values of mtt_split_cu_vertical_flag and mtt_split_cu_binary_flag, the decoder can derive the multi-type tree division mode (MttSplitMode) of the CU based on predefined rules or tables. It should be noted that for a certain design, such as the 64×64 luma block and 32×32 chroma pipeline design in a VVC hardware decoder, TT division is not allowed when the width or height of the luma coding block is greater than 64. TT division is also not allowed when the width or height of the chroma coding block is greater than 32. The pipeline design divides the image into multiple virtual pipeline data units (VPDUs), and each VPDU is defined as a non-overlapping unit in the image. In the hardware decoder, consecutive VPDUs are processed simultaneously in multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is necessary to keep the VPDU small.In most hardware decoders, the VPDU size can be set to the maximum transform block (TB) size. However, in VVC, the splitting of the ternary tree (TT) and binary tree (BT) may increase the VPDU size.
[0191] In addition, it should be noted that when a part of the tree node block extends beyond the bottom or the right image boundary, the tree node block is forced to be divided until all pixel points of each coded CU are within the image boundary.
[0192] For example, the intra sub - partitions (ISP) tool can vertically or horizontally divide the luma intra - prediction block into two or four sub - parts according to the block size.
[0193] In one example, the mode selection unit 260 of the video encoder 20 can be used to perform any combination of the segmentation techniques described above.
[0194] As described above, the video encoder 20 is used to determine or select the best or optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes may include, for example, intra - prediction modes and / or inter - prediction modes.
[0195] Intra - prediction
[0196] The set of intra - prediction modes may include 35 different intra - prediction modes. For example, non - directional modes such as the DC (or mean) mode and the planar mode, or directional modes defined as in HEVC. Or it may include 67 different intra - prediction modes. For example, non - directional modes such as the DC (or mean) mode and the planar mode, or directional modes defined in VVC. For example, several traditional angular intra - prediction modes are adaptively replaced by wide - angle intra - prediction modes for non - square blocks defined in VVC. Also, for example, to avoid the division operation in DC prediction, only the longer side is used to calculate the average value of non - square blocks. And the intra - prediction result of the planar mode can also be modified using the position - dependent intra - prediction combination (PDPC) method.
[0197] The intra - prediction unit 254 is used to generate an intra - prediction block 265 using the reconstructed pixel points of adjacent blocks in the same current image according to the intra - prediction mode in the set of intra - prediction modes.
[0198] The intra prediction unit 254 (or generally the mode selection unit 260) is also used to output intra prediction parameters (or generally information indicating the selected intra prediction mode of the block) in the form of a syntax element 266 to be sent to the entropy coding unit 270 to be included in the coded picture data 21, so that the video decoder 30 can perform operations such as receiving and using the prediction parameters for decoding.
[0199] Inter - prediction
[0200] In a possible implementation, the set of inter - prediction modes depends on the available reference images (i.e., for example, at least a part of the previously decoded images stored in the DBP 230 as described above) and other inter - prediction parameters, such as depending on whether the entire reference image or only a part of the reference image, such as a search window area near the region of the current block, is used to search for the best - matching reference block, and / or for example depending on whether pixel interpolation such as half - pixel, quarter - pixel, and / or sixteenth - pixel interpolation is performed.
[0201] In addition to the above - mentioned prediction modes, a skip mode and / or a direct mode can also be adopted.
[0202] For example, for extended merge prediction, the merge candidate list for this mode consists of the following five candidate types in order: spatial MVP from spatially adjacent CUs, temporal MVP from collocated CUs, history-based MVP from the FIFO table, pairwise average MVP, and zero MV. A decoder side motion vector refinement (DMVR) based on bilateral matching can be used to increase the accuracy of the MV for the merge mode. The merge mode with MVD (MMVD) comes from the merge mode with motion vector differences. The MMVD flag is sent immediately after the skip flag and the merge flag to specify whether the CU uses the MMVD mode. A CU-level adaptive motion vector resolution (AMVR) scheme can be used. AMVR supports encoding the MVD of the CU with different precisions. The MVD of the current CU is adaptively selected according to the prediction mode of the current CU. When the CU is encoded in the merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. The inter and intra prediction signals are weighted and averaged to obtain the CIIP prediction. For affine motion compensation prediction, the affine motion field of the block is described by the motion information of the motion vector with 2 control points (4 parameters) or 3 control points (6 parameters). The subblock-based temporal motion vector prediction (SbTMVP) is similar to the temporal motion vector prediction (TMVP) in HEVC, but predicts the motion vector of the sub-CU within the current CU. The bi-directional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces computations, especially in terms of the number of multiplications and the size of the multipliers. In the triangle partitioning mode, the CU is evenly divided into two triangular parts in two partitioning ways: diagonal partitioning and anti-diagonal partitioning. In addition, the bi-directional prediction mode is extended based on simple averaging to support the weighted average of two prediction signals.
[0203] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in Figure 2(not shown in the figure). The motion estimation unit can be used to receive or obtain the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or more previously reconstructed blocks, for example, the reconstructed blocks of one or more other / different previously decoded images 231, to perform motion estimation. For example, the video sequence can include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 can be part of the image sequence forming the video sequence or form the image sequence.
[0204] For example, the encoder 20 can be used to select a reference block from multiple reference blocks of the same or different images among multiple other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also called a motion vector (MV).
[0205] The motion compensation unit is used to obtain, for example, receive, the inter-frame prediction parameter, and perform inter-frame prediction according to or using the inter-frame prediction parameter to obtain the inter-frame prediction block 246. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block according to the motion / block vector determined by motion estimation, and may also include performing interpolation at sub-pixel accuracy. The interpolation filter can generate pixel points of other pixels from the pixel points of known pixels, thereby potentially increasing the number of candidate prediction blocks available for encoding the image block. Once the motion vector corresponding to the PU of the current image block is received, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.
[0206] The motion compensation unit can also generate syntax elements related to the block and the video slice for use by the video decoder 30 when decoding the image blocks of the video slice. In addition, or as an alternative to the slice and the corresponding syntax elements, coded block groups and / or coded blocks and the corresponding syntax elements can be generated or used.
[0207] Entropy coding
[0208] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CALVC), arithmetic coding scheme, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, to obtain coded image data 21 that can be output in the form of a coded bitstream 21, etc. through the output terminal 272, such that a video decoder 30, etc. can receive and use the parameters for decoding. The coded bitstream 21 can be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.
[0209] Other structural variants of the video encoder 20 can be used to encode a video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal in the case where some blocks or frames do not have a transform processing unit 206. In another implementation, the encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.
[0210] Decoder and decoding method
[0211] Figure 3 An exemplary video decoder 30 for implementing the technology of the present application is shown. The video decoder 30 is used to receive coded image data 21 (e.g., coded bitstream 21) encoded by, for example, the encoder 20, to obtain a decoded image 331. The coded image data or bitstream includes information for decoding the coded image data, such as data representing image blocks of a coded video slice (and / or coded group of blocks or coded block) and related syntax elements.
[0212] In Figure 3In the example, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (such as an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding process that is generally opposite to the encoding process described with reference to Figure 2 the video encoder 100.
[0213] See Figure 3 , the inter prediction module / intra prediction module / loop filter module / XXX includes (is) a trained target model (also called a neural network), and this neural network is used to process the input image or image region or image block to generate a predicted value of the input image block. For example, the neural network for inter prediction / intra prediction / loop filtering is used to receive the input image or image region or image block. For example, Figures 6a to 6c the input image data shown in the figure, and generate a predicted value of the input image or image region or image block. The neural network for inter prediction / intra prediction / loop filtering / XXX will be described in detail below in conjunction with Figures 7a - 7e .
[0214] As described for the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer DPB 230, the inter prediction unit 344, and the intra prediction unit 354 also form the "built-in decoder" of the video encoder 20. Correspondingly, the inverse quantization unit 310 may be functionally the same as the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally the same as the inverse transform processing unit 122, the reconstruction unit 314 may be functionally the same as the reconstruction unit 214, the loop filter 320 may be functionally the same as the loop filter 220, and the decoded picture buffer 330 may be functionally the same as the decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of the video encoder 20 are correspondingly applicable to the corresponding units and functions of the video decoder 30.
[0215] Entropy decoding
[0216] The entropy decoding unit 304 is used to parse the bitstream 21 (or generally the encoded image data 21) and perform entropy decoding on the encoded image data 21 to obtain the quantization coefficients 309 and / or the decoded encoded parameters ( Figure 3etc. (not shown in the figure), such as any one or all of inter-frame prediction parameters (e.g., reference image index and motion vector), intra-frame prediction parameters (e.g., intra-frame prediction mode or index), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 can be used to apply the decoding algorithm or scheme corresponding to the encoding scheme of the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 can also be used to provide inter-frame prediction parameters, intra-frame prediction parameters, and / or other syntax elements to the mode application unit 360, and provide other parameters to other units of the decoder 30. The video decoder 30 can receive syntax elements at the video slice and / or video block level. Additionally, or as an alternative to the slice and corresponding syntax elements, coded block groups and / or coded blocks and corresponding syntax elements can be received or used.
[0217] Inverse quantization
[0218] The inverse quantization unit 310 can be used to receive a quantization parameter (QP) (or generally information related to inverse quantization) and quantization coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse-quantize the decoded quantization coefficients 309 based on the quantization parameter to obtain inverse-quantized coefficients 311, which can also be referred to as transform coefficients 311. The inverse quantization process can include using the quantization parameter calculated by the video encoder 20 for each video block in the video slice to determine the degree of quantization, and also determine the degree of inverse quantization to be performed.
[0219] Inverse transform
[0220] The inverse transform processing unit 312 can be used to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the pixel domain. The reconstructed residual block 213 can also be referred to as a transform block 313. The transform can be an inverse transform, such as an inverse DCT, inverse DST, inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 can also be used to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304) to determine the transform to be applied to the dequantized coefficients 311.
[0221] Reconstruction
[0222] The reconstruction unit 314 (e.g., adder 314) is used to add the reconstructed residual block 313 to the prediction block 365 to obtain a reconstructed block 315 in the pixel domain, e.g., adding the pixel point values of the reconstructed residual block 313 and the pixel point values of the prediction block 365.
[0223] Filtering
[0224] The loop filter unit 320 (in or after the encoding loop) is used to filter the reconstructed block 315 to obtain a filtered block 321, so as to smoothly perform pixel transformation or improve video quality, etc. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chromascaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 320 is shown as a loop filter in Figure 3 it may be implemented as a post-loop filter in other configurations.
[0225] Decoded picture buffer
[0226] Subsequently, the decoded video block 321 in an image is stored in the decoded picture buffer 330, and the decoded picture buffer 330 stores the decoded picture 331 as a reference picture, and the reference picture is used for subsequent motion compensation of other pictures and / or output for display respectively.
[0227] The decoder 30 is used to output the decoded picture 311 through the output terminal 312, etc., for the user to display or view.
[0228] Prediction
[0229] The inter-frame prediction unit 344 can be functionally the same as the inter-frame prediction unit 244 (especially the motion compensation unit), and the intra-frame prediction unit 354 can be functionally the same as the inter-frame prediction unit 254, and determines the division or segmentation and performs prediction based on the segmentation and / or prediction parameters or corresponding information received from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304). The mode application unit 360 can be used to perform prediction (intra-frame or inter-frame prediction) for each block according to the reconstructed image, block or corresponding pixel points (filtered or unfiltered), and obtain the predicted block 365.
[0230] When encoding a video slice as an intra-coded (I) slice, the intra-frame prediction unit 354 in the mode application unit 360 is used to generate a predicted block 365 for the image block of the current video slice according to the indicated intra-frame prediction mode and the data of the previously decoded blocks from the current image. When the video image is encoded as an inter-coded (i.e., B or P) slice, the inter-frame prediction unit 344 (e.g., the motion compensation unit) in the mode application unit 360 is used to generate a predicted block 365 for the video block of the current video slice according to the motion vector and other syntax elements received from the entropy decoding unit 304. For inter-frame prediction, these predicted blocks can be generated from one of the reference images in one of the reference image lists. The video decoder 30 can use the default construction technique to construct reference frame list 0 and list 1 according to the reference images stored in the DPB 330. In addition to or as an alternative to a slice (e.g., a video slice), the same or similar process can be applied to embodiments of coded block groups (e.g., video coded block groups) and / or coded blocks (e.g., video coded blocks), for example, a video can be encoded using I, P, or B coded block groups and / or coded blocks.
[0231] The mode application unit 360 is used to determine the prediction information for the video block of the current video slice by parsing the motion vector and other syntax elements, and use the prediction information to generate a predicted block for the current video block being decoded. For example, the mode application unit 360 uses some received syntax elements to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for the video block of the encoded video slice, the inter-frame prediction slice type (e.g., B slice, P slice or GPB slice), the construction information for one or more reference image lists for the slice, the motion vector for each inter-frame coded video block of the slice, the inter-frame prediction state for each inter-frame coded video block of the slice, and other information to decode the video blocks within the current video slice. In addition to or as an alternative to a slice (e.g., a video slice), the same or similar process can be applied to embodiments of coded block groups (e.g., video coded block groups) and / or coded blocks (e.g., video coded blocks), for example, a video can be encoded using I, P, or B coded block groups and / or coded blocks.
[0232] In one embodiment, Figure 3The illustrated video encoder 30 may also be used to segment and / or decode an image using slices (also referred to as video slices), where the image may be segmented or decoded using one or more slices (usually non-overlapping). Each slice may include one or more blocks (e.g., CTUs) or one or more groups of blocks (e.g., coding tree units in the H.265 / HEVC / VVC standards and tiles in the VVC standard).
[0233] In one embodiment, Figure 3 The illustrated video decoder 30 may also be used to segment and / or decode an image using slice / coding tree unit groups (also referred to as video coding tree unit groups) and / or coding tree units (also referred to as video coding tree units), where the image may be segmented or decoded using one or more slice / coding tree unit groups (usually non-overlapping), each slice / coding tree unit group may include one or more blocks (e.g., CTUs) or one or more coding tree units, etc., where each coding tree unit may be in a shape such as a rectangle and may include one or more whole or partial blocks (e.g., CTUs).
[0234] Other variants of the video decoder 30 may be used to decode the encoded image data 21. For example, the decoder 30 may produce an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 may directly dequantize the residual signal without the inverse transform processing unit 312 for certain blocks or frames. In another implementation, the video decoder 30 may have a dequantization unit 310 and an inverse transform processing unit 312 combined into a single unit.
[0235] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clip or shift operations may be performed on the processing result of interpolation filtering, motion vector derivation, or loop filtering.
[0236] It should be noted that further operations can be performed on the derived motion vectors of the current block (including but not limited to the control point motion vectors in the affine mode, the sub-block motion vectors in the affine, planar, and ATMVP modes, the temporal motion vectors, etc.). For example, the value of the motion vector can be restricted to a predefined range according to the representation bits of the motion vector. If the representation bit of the motion vector is bitDepth, the range is from -2^(bitDepth - 1) to 2^(bitDepth - 1) - 1, where "^" represents the power. For example, if bitDepth is set to 16, the range is from -32768 to 32767; if bitDepth is set to 18, the range is from -131072 to 131071. For example, the value of the derived motion vector (such as the MV of 4 4×4 sub-blocks in an 8×8 block) is restricted such that the maximum difference between the integer parts of the 4 4×4 sub-block MVs does not exceed N pixels, for example, does not exceed 1 pixel. Two methods for restricting the motion vector according to bitDepth are provided here.
[0237] Although the above embodiments mainly describe video coding and decoding, it should be noted that the embodiments of the decoding system 10, the encoder 20, and the decoder 30, as well as other embodiments described herein, can also be used for still image processing or coding and decoding, that is, the processing or coding and decoding of a single image independent of any previous or consecutive images in video coding and decoding. Generally, if the image processing is limited to a single image 17, the inter-frame prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can equally be used for static image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, dequantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy coding 270, and entropy decoding 304.
[0238] Figure 4 Schematic diagram of the video decoding device 400 provided by the embodiments of the present application. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 can be a decoder, such as Figure 1A the video decoder 30 in Figure 1A or an encoder, such as
[0239] The video decoding device 400 includes: an input port 410 (or input port 410) for receiving data and a receiver unit (Rx) 420; a processor, logic unit, or central processing unit (CPU) 430 for processing data; for example, the processor 430 here can be a neural network processor 430; a transmitter unit (Tx) 440 and an output port 450 (or output port 450) for transmitting data; and a memory 460 for storing data. The video decoding device 400 may further include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the input port 410, the receiving unit 420, the transmitting unit 440, and the output port 450 for the exit or entry of optical or electrical signals.
[0240] The processor 430 is implemented by hardware and software. The processor 430 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the input port 410, the receiving unit 420, the transmitting unit 440, the output port 450, and the memory 460. The processor 430 includes a decoding module 470 (e.g., a decoding module 470 based on a neural network NN). The decoding module 470 implements the embodiments disclosed above. For example, the decoding module 470 performs, processes, prepares, or provides various encoding operations. Therefore, the decoding module 470 provides a substantial improvement to the functions of the video decoding device 400 and affects the switching of the video decoding device 400 to different states. Alternatively, the decoding module 470 is implemented by instructions stored in the memory 460 and executed by the processor 430.
[0241] The memory 460 includes one or more disks, tape drives, and solid-state drives and can be used as an overflow data storage device for storing such programs when a selected program is to be executed and for storing the instructions and data read during the execution of the program. The memory 460 can be volatile and / or non-volatile and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0242] Figure 5 A simplified block diagram of the apparatus 500 provided for an exemplary embodiment. The apparatus 500 can be used asFigure 1A either one or both of the source device 12 and the destination device 14 in
[0243] The processor 502 in the apparatus 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices, existing or to be developed in the future, capable of manipulating or processing information. Although a single processor such as the processor 502 shown in the figure may be used to implement the disclosed implementations, using more than one processor is faster and more efficient.
[0244] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessible by the processor 502 via the bus 512. The memory 504 may also include an operating system 508 and an application 510, and the application 510 includes at least one program that allows the processor 502 to execute the methods described herein. For example, the application 510 may include applications 1 to N, and also includes a video decoding application that executes the methods described herein.
[0245] The apparatus 500 may also include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that can be used to sense touch inputs. The display 518 may be coupled to the processor 502 via the bus 512.
[0246] Although the bus 512 in the apparatus 500 is described herein as a single bus, the bus 512 may include multiple buses. In addition, the auxiliary storage may be directly coupled to other components of the apparatus 500 or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, the apparatus 500 may have a variety of configurations.
[0247] Since the embodiments of this application relate to the application of neural networks, for ease of understanding, some nouns or terms used in the embodiments of this application are explained below, and these nouns or terms are also part of the invention content.
[0248] (1) Neural network
[0249] A neural network (Neural Network, NN) is a machine learning model. A neural network may be composed of neural units, and a neural unit may refer to an arithmetic unit that takes xs and an intercept 1 as inputs. The output of this arithmetic unit may be:
[0250]
[0251] Among them, s = 1, 2, …… n, where n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0252] (2) Deep Neural Network
[0253] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i + 1-th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer only performs the following simple operation on the input vector to obtain the output vector Since the DNN has many layers, the number of coefficients W and the offset vector is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the k-th neuron in the L - 1-th layer to the j-th neuron in the L-th layer is defined as It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0254] (3) Convolutional Neural Network
[0255] A convolutional neural network (CNN) is a deep neural network with a convolutional structure and a deep learning architecture. A deep learning architecture refers to multiple levels of learning at different abstraction levels through machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network in which each neuron can respond to the input image. A convolutional neural network contains a feature extractor composed of convolutional layers and pooling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input image or a convolutional feature plane (feature map).
[0256] A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolutional processing on the input signal. The convolutional layer can include many convolutional operators, which are also called kernels. Their role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolutional operation on the image, the weight matrix usually processes pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction of the input image, thus completing the work of extracting specific features from the image. The size of this weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as that of the input image. During the convolutional operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices with the same size (row × column), that is, multiple matrices of the same type, are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image, where the dimension can be understood as determined by the above-mentioned "multiple". Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and yet another weight matrix is used to blur the unwanted noise in the image, etc. These multiple weight matrices have the same size (row × column), and the size of the feature maps extracted by these multiple weight matrices of the same size is also the same. Then, the multiple feature maps of the same size extracted are merged to form the output of the convolutional operation. The weight values in these weight matrices need to be obtained through a large amount of training in practical applications. Each weight matrix formed by the weight values obtained through training can be used to extract information from the input image, so that the convolutional neural network can make correct predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layer often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network increases, the features extracted by the subsequent convolutional layers become more and more complex, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.
[0257] Since it is often necessary to reduce the number of training parameters, a pooling layer is often periodically introduced after the convolutional layer. It can be a pooling layer following a convolutional layer, or one or more pooling layers following multiple convolutional layers. In the process of image processing, the sole purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer can include an average pooling operator and / or a max pooling operator for sampling the input image to obtain a smaller-sized image. The average pooling operator can calculate the average value of pixel values in the image within a specific range as the result of average pooling. The max pooling operator can take the pixel with the maximum value within the specific range as the result of max pooling. Additionally, just as the size of the weight matrix in the convolutional layer should be related to the image size, the operators in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer can be smaller than the size of the image input to the pooling layer. Each pixel point in the image output by the pooling layer represents the average or maximum value of the corresponding sub-region of the image input to the pooling layer.
[0258] After being processed by the convolutional layer / pooling layer, the convolutional neural network is still not sufficient to output the required output information. As mentioned before, the convolutional layer / pooling layer only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network needs to use neural network layers to generate one or a set of outputs with the number of classes required. Therefore, the neural network layer can include multiple hidden layers, and the parameters contained in these multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include image recognition, image classification, image super-resolution reconstruction, etc.
[0259] Optionally, after the multiple hidden layers in the neural network layer, there is also an output layer of the entire convolutional neural network. This output layer has a loss function similar to categorical cross-entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, backpropagation will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network, that is, the error between the result output by the convolutional neural network through the output layer and the ideal result.
[0260] (4) Recurrent Neural Network
[0261] Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, it goes from the input layer to the hidden layer and then to the output layer, with full connections between layers, while there are no connections between individual nodes within each layer. Although this ordinary neural network has solved many problems, it is still powerless in many aspects. For example, when you want to predict the next word in a sentence, you generally need to use the previous words because the words before and after in a sentence are not independent. The reason why RNN is called a recurrent neural network is that the current output of a sequence is also related to the previous output. The specific manifestation is that the network will remember the previous information and apply it to the calculation of the current output, that is, the nodes within the hidden layer itself are no longer unconnected but connected, and the input to the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNNs can process sequential data of any length. The training of RNNs is the same as that of traditional CNNs or DNNs. The error backpropagation algorithm is also used, but there is one difference: that is, if the RNN is unfolded, the parameters such as W are shared; while the traditional neural network mentioned above is not like this. And in the use of the gradient descent algorithm, the output at each step depends not only on the network at the current step but also on the states of the networks in several previous steps. This learning algorithm is called Backpropagation Through Time (BPTT).
[0262] Since we already have convolutional neural networks, why do we still need recurrent neural networks? The reason is very simple. In convolutional neural networks, there is a prerequisite assumption that elements are independent of each other, and the input and output are also independent, such as a cat and a dog. However, in the real world, many elements are interconnected, such as the change of stocks over time, or for example, a person says: "I like traveling, and the favorite place is Yunnan. I must go there if I have the chance in the future." Here, fill in the blank. Humans should all know that the answer is "Yunnan". Because humans will make inferences based on the context, but how can we make machines do this? This is where RNNs come into being. RNNs aim to enable machines to have the ability to remember like humans. Therefore, the output of RNNs needs to depend on the current input information and historical memory information.
[0263] (5) Recursive residual convolutional neuronnetwork (RR-CNN)
[0264] (6) Artificial Neural Networks (ANN)
[0265] (7) Loss function
[0266] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0267] (8) Backpropagation algorithm
[0268] The convolutional neural network can use the backpropagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during the training process, so that the reconstruction error loss of the super-resolution model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial super-resolution model parameters are updated by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0269] (9) Generative adversarial network
[0270] Generative adversarial networks (GAN) is a deep learning model. There are at least two modules in this model: one module is the Generative Model, and the other module is the Discriminative Model. Through the mutual game learning of these two modules, better outputs can be generated. Both the Generative Model and the Discriminative Model can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of GAN is as follows: Taking the GAN for generating pictures as an example, assume there are two networks, G (Generator) and D (Discriminator). Among them, G is a network for generating pictures. It receives a random noise z and generates a picture through this noise, denoted as G(z); D is a discriminative network used to determine whether a picture is "real". Its input parameter is x, where x represents a picture, and the output D(x) represents the probability that x is a real picture. If it is 1, it means it is a 100% real picture. If it is 0, it means it is impossible to be a real picture. During the training process of this generative adversarial network, the goal of the generative network G is to generate as real pictures as possible to deceive the discriminative network D, while the goal of the discriminative network D is to try to distinguish the pictures generated by G from the real pictures. In this way, G and D constitute a dynamic "game" process, that is, the "adversary" in the "generative adversarial network". Finally, in the ideal state, the result of the game is that G can generate pictures G(z) that are "indistinguishable from the real ones", and D is difficult to determine whether the pictures generated by G are real, that is, D(G(z)) = 0.5. In this way, an excellent generative model G is obtained, which can be used to generate pictures.
[0271] The following will be combined with Figures 7a - 7e Describe in detail the target model (also known as a neural network) for feature extraction or probability distribution estimation. Figures 7a - 7e Schematic diagram of the example architecture 700 of the target model (such as a neural network for feature extraction, abbreviated as a feature extraction network). Using the input pixel matrix of the image block as the input of the neural network, the neural network processes the input data using the convolutional layer 220 and the network layer 230, and outputs a three-dimensional feature block using the output layer 240.
[0272] It should be noted that, as Figures 7a - 7e The convolutional neural network shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models, and this application does not make specific limitations on this.
[0273]
[0274]
[0275] Image compression is a fundamental task in the field of image processing. With the development of modern technologies, the construction of smart cities, the popularity of mobile device photography and short videos, and the increasing number of road monitoring systems, etc., image storage and transmission have become problems to be solved. Therefore, image compression is becoming increasingly important in saving transmission bandwidth and storage devices.
[0276] The traditional classic image compression standard JPEG (Joint Photographic Experts Group) was released in 1992 and is currently the most widely used image compression coding standard. Its technical solution is simple and royalty-free. In the nearly 30 years since the release of the JPEG standard, international standard organizations have developed international standards such as JPEG2000, H.264 intra-frame coding, and H.265 intra-frame coding that can be applied to image compression. The compression efficiency of these new standards far exceeds that of the JPEG standard. In the past 20-odd years, the JPEG standard has been widely applied to hardware devices such as video surveillance cameras, and the generated image files can be decoded and viewed by almost all devices, with a perfect ecosystem. Therefore, it will continue to be used for a relatively long time.
[0277] However, for image compression, these image compression methods cannot achieve high compression performance and inevitably generate unnatural information, such as ringing and compression artifacts, etc. The compressed files generated by JPEG encoding consume a large amount of storage resources on the server side. If there is a transcoding solution that runs on the server side and can perform lossless transcoding on JPEG-encoded files, it will be able to greatly save the storage resource overhead on the server side.
[0278] In recent years, neural networks have made great developments in tasks related to computer vision and have better performance compared to traditional methods. Due to their deeper image modeling and expression capabilities, through multi-layer convolutions for non-linear analysis and mapping transformation, and end-to-end training and learning, they can achieve a high compression ratio while ensuring that the reconstructed images are relatively clear and retain more detailed textures.
[0279] Figure 8Shown is an image compression method based on deep learning, also known as an image compression method based on a neural network. Generally, an image compression method based on a neural network includes the following parts: a feature extraction module, a feature quantization module, entropy encoding, entropy decoding, a feature inverse quantization module, and a feature decoding module. Among them, at the encoder end, the feature extraction module can obtain a three-dimensional feature map after extraction by stacking multiple layers of convolutions and combining a non-linear mapping activation function. The feature quantization module quantizes the floating-point feature values through a feature value quantization method to obtain the quantized feature values. The quantized feature values are subjected to lossless entropy encoding to obtain an encoded bitstream. At the decoder end, when receiving the bitstream after entropy encoding, first perform lossless entropy decoding to obtain three-dimensional quantized feature values. Through the feature decoding module, decode the features into a reconstructed image, thereby achieving the decoding purpose.
[0280] After the image to be compressed passes through the feature extraction module and the feature quantization module, a three-dimensional quantized feature block is obtained. When the entropy encoding module processes each feature value in the three-dimensional quantized feature block, it can use the already processed feature values in the neighborhood as context to estimate the probability distribution of the feature value, obtain the probability distribution of the feature value, and perform subsequent encoding according to the probability distribution to obtain the encoded bitstream.
[0281] Specifically, the method for estimating the probability distribution of the feature value is as follows: Method 1: The probability estimation network directly estimates the probability of the values within the dynamic value range of the feature value to obtain a probability vector, where the probability vector includes the probability values of each value within the value range, and the sum of all probability values is 1. Method 2: A probability distribution model can also be used to obtain the probability vector. For example, use a single Gaussian model (GSM) or a Gaussian mixture model (GMM) for modeling, use the probability estimation network to estimate the model parameters, and substitute the model parameters into the probability distribution model to obtain the probability vector.
[0282] Among them, entropy encoding includes two parts: context modeling and arithmetic coding. Before arithmetic coding, context modeling needs to be performed. According to the previously encoded feature values, the probability value of the current feature value is estimated through the probability estimation network to obtain the probability distribution of the feature value. Then, subsequent arithmetic coding is performed according to the corresponding probability model to output the encoded bitstream.
[0283] The process of entropy decoding is the opposite of the process of entropy encoding. When processing each code element in the bitstream, the already decoded feature values in the neighborhood of the current code element are input into the probability estimation network and the probability of the current code element to be decoded is estimated to obtain the probability distribution of the current code element to be decoded. Then, the probability distribution is used to decode the current code element to obtain the quantized feature value of the current code element. After decoding each code element, three-dimensional quantized feature values are obtained.
[0284] This application proposes an end-to-end image coding scheme, and also provides a method for obtaining the size information of three-dimensional feature blocks and a bitstream transmission scheme, which can effectively improve the encoding and decoding efficiency.
[0285] The main method is: transmit the size information of three-dimensional feature blocks or the size information of sub-images (image blocks) from the encoding end to the decoding end, and the decoding end can perform decoding according to the size information of the three-dimensional feature blocks, which can effectively improve the encoding and decoding efficiency.
[0286] This application relates to an end-to-end image coding scheme: perform block partitioning before feature analysis. Specify the bitstream transmission scheme for the size information of three-dimensional feature blocks.
[0287] To meet the growing image transmission and storage needs of users and improve the image encoding and decoding efficiency, an embodiment of this application provides an image coding method. This method transmits the size information of three-dimensional feature blocks from the encoding end to the decoding end, so that the decoding end can perform decoding based on the size information of the three-dimensional feature blocks to obtain a reconstructed image. Since the size information of three-dimensional feature blocks is added during the image encoding and decoding process, the encoding bitstream at the encoding end can have a high compression ratio, and the decoding end can decode the bitstream with a high compression ratio to obtain a reconstructed image, and can ensure that the reconstructed image is relatively clear, so the image encoding and decoding efficiency can be improved.
[0288] A high compression ratio can reduce the size of the bitstream, thereby reducing the bandwidth resources required for transmitting images and the storage resources required for storing images.
[0289] The following elaborates in detail the specific process of image encoding and decoding in the embodiments of this application with reference to the accompanying drawings.
[0290] Figure 9 It is a schematic diagram of the processing process of the image coding method in the embodiments of this application, as Figure 9 shown, perform block partitioning on the image to be processed to obtain a plurality of image blocks, perform feature extraction on the plurality of image blocks to obtain a plurality of three-dimensional feature blocks, and perform entropy coding on the size information of the plurality of three-dimensional feature blocks and the plurality of three-dimensional feature blocks to obtain an encoded bitstream.
[0291] Specifically, refer to Figure 10 , Figure 10 It is a flowchart of the image coding method in the embodiments of this application, Figure 10 The method shown can be executed by an encoding device or an encoder, hereinafter collectively referred to as the encoding end, as Figure 10 shown, the method may include:
[0292] Step 201, obtain the image to be processed.
[0293] The image to be processed can also be referred to as the image to be compressed.
[0294] Step 202: Divide the image to be processed into at least two image blocks.
[0295] Divide the image to be processed into multiple image blocks. The multiple image blocks can be of the same size. As an example, Figure 3 is a schematic diagram of the image block division according to the embodiment of the present application. As Figure 11 shown, the image to be processed is divided into 8 image blocks of the same size. The size of each image block can be represented by W and H, where W is the width of the image block and H is the height of the image block. W and H are positive integers respectively. The values of W and H can be any positive integers, and the values of W and H can be the same or different. For example, W = 4 and H = 4, that is, the size of the image block is 4×4. For another example, W = 4 and H = 8, that is, the size of the image block is 4×8. For yet another example, W = 8 and H = 4, that is, the size of the image block is 8×4. Of course, it can be understood that the size of the image block can also be 8×8, 16×16, 32×32, 32×32, 64×64, 128×128 or 256×256, etc. The embodiments of the present application will not list them one by one. The size of the image block can be a preset size.
[0296] Step 203: Input at least two image blocks into the first neural network model to obtain at least two three-dimensional feature blocks output by the first neural network model.
[0297] The first neural network model can be used to extract features from multiple image blocks to obtain multiple three-dimensional feature blocks. The multiple image blocks can correspond one by one to the multiple three-dimensional feature blocks.
[0298] The eigenvalue of any three-dimensional feature block can be a floating point number between 0 and 1.
[0299] As an example, the features of multiple image blocks can be extracted in a preset order to obtain the three-dimensional feature blocks corresponding to each image block respectively. For example, taking the above Figure 11 as a further example, the features of 8 image blocks of the same size can be extracted respectively to obtain the three-dimensional feature blocks corresponding to each of the 8 image blocks of the same size.
[0300] The size of any three-dimensional feature block among the multiple three-dimensional feature blocks can include the length, width and height of the three-dimensional feature block. M is the length of the three-dimensional feature block, N is the width of the three-dimensional feature block, and R is the height of the three-dimensional feature block.
[0301] An achievable way of feature extraction is to process the image patch through the first neural network model to obtain a three-dimensional feature patch. Taking an image patch as an example, the image patch can be input into the first neural network model to obtain the three-dimensional feature patch output by the first neural network model. Among them, the first neural network model can adopt any network structure, for example, a fully connected network, a convolutional neural network, or a recurrent neural network, etc. In some embodiments, the first neural network model can adopt a multi-layer deep neural network structure to achieve better feature extraction effects. As an example, taking the first neural network model as a convolutional neural network, the convolutional neural network includes multiple convolutional layers, a non-linear mapping activation function layer, and a pooling layer. The image patch is processed through the multiple convolutional layers, the non-linear mapping activation function layer, and the pooling layer to obtain a three-dimensional feature patch. Of course, it can be understood that the structure of the first neural network model can also be in other forms, and the embodiments of the present application will not list them one by one.
[0302] Before using the first neural network model to extract features from the image patch, the neural network model can also be trained through a training process to obtain the first neural network model. The training data used in the training process can be training image patches and the corresponding training three-dimensional feature patches.
[0303] Step 204: Encode at least two three-dimensional feature patches and encode the size information of at least two three-dimensional feature patches to obtain an encoded bitstream.
[0304] Among them, the size information of the at least two three-dimensional feature patches may include: the size information of the at least two image patches, or the size information of the at least two image patches and the parameters of the first neural network model, or the sizes of the at least two three-dimensional feature patches. The size information of the at least two image patches and the parameters of the first neural network model are used to determine the sizes of the at least two three-dimensional feature patches.
[0305] The size information of the at least two image patches is used to represent the sizes of the at least two image patches divided. An achievable way is that the size information of the at least two image patches may include the above-mentioned W and H. Based on the above-mentioned W and H, the size information of the three-dimensional feature patch can be determined. For example, M×N×R can be determined. Another achievable way is that the size information of the at least two image patches may include the division method applied to the image to be processed and used to divide the at least two image patches. For example, taking Figure 11 the illustrated embodiment as an example, the division method may be 2×4, that is, it means that the image to be processed is divided into 2×4 image patches. Based on this division method, W and H can be determined, and according to W and H, the size information of the three-dimensional feature patch can be determined. For example, M×N×R can be determined.
[0306] Among them, the size information (MxNxR) of the three-dimensional feature block can be determined according to the size information of the input image block:
[0307] In the first method, the correspondence between W and H, and M, N, and R can be: M×N×R = W / S×H / S×C, where C is the number of channels of the convolution kernel of the first neural network model, and S is the scaling step of the first neural network model. The values of S and C can be preset. In other words, the parameters of the first neural network model can be built into the encoding end and the decoding end without transmission.
[0308] Among them, S is a number greater than 0, which can be a decimal or an integer, and is not limited here. If a pooling layer is used in the network structure, then S is the step of the pooling layer, or a value obtained by multiplying the step by a coefficient a (the coefficient a can be a number greater than 1 or less than 1, such as 0.5 or 2); if no pooling layer is used, then S is 1.
[0309] The parameters of the first neural network model can include S and C. Different from the parameters of the first neural network model that can be built into the encoding end and the decoding end, the parameters of the first neural network model can be transmitted from the encoding end to the decoding end so that the decoding end can determine M×N×R based on the received parameters of the first neural network model and the size information of the at least two image blocks.
[0310] In the second method, the length of the three-dimensional feature block is M = Wb, the width is N = Hb, and the height is R = C, where b is the scaling factor of the neural network. The value of this scaling factor can be obtained according to the scaling factors of each layer of the network. For example, b is obtained by multiplying the scaling factors of each layer of the network. The value of b is a number greater than 0; C is the number of channels (channels) of the neural network, and the value is an integer greater than or equal to 1.
[0311] In the third method, the length of the three-dimensional feature block is M = Wb, the width is N = Hc, and the height is R = C, where b and c are the scaling factors of the neural network. The values of these scaling factors can be obtained according to the scaling factors of each layer of the network. For example, b or c is obtained by multiplying the scaling factors of each layer of the network. The values of b and c are different, and the value range is a number greater than 0; C is the number of channels (channels) of the neural network, and the value is an integer greater than or equal to 1.
[0312] The embodiments of the present application can encode the size information of at least two three-dimensional feature blocks and at least two three-dimensional feature blocks to obtain an encoded bitstream. The encoding end sends the encoded bitstream to the decoding end. The decoding end can obtain the size information of at least two three-dimensional feature blocks from the encoded bitstream, and then decode the code elements according to the size information of at least two three-dimensional feature blocks to obtain a reconstructed image.
[0313] Optionally, an implementable way to encode at least two three-dimensional feature blocks is to estimate the probability distribution of the feature values of any three-dimensional feature block according to the size of any three-dimensional feature block, obtain the probability distribution vector of the feature value, and perform entropy encoding on any three-dimensional feature block according to the probability distribution vector of the feature value. The probability distribution vector may include the probability values of each value within the value range, and the sum of all probability values is 1. Taking the encoded bitstream as a binary bitstream as an example, the probability distribution vector may include the probability value of 0 and the probability value of 1, and the sum of the two is 1.
[0314] An implementable way to perform entropy encoding on the size information of at least two three-dimensional feature blocks and at least two three-dimensional feature blocks is to perform entropy encoding on at least two three-dimensional feature blocks to obtain the encoded bitstream, and write the encoded image block partition information into the bitstream. For example, the size information of the encoded at least two three-dimensional feature blocks can be written into the bitstream through one or more syntax elements. For example, the size information of at least two three-dimensional feature blocks is written into the bitstream through the values of one or more syntax elements. Taking the size information of at least two three-dimensional feature blocks including M, N, and R as an example, M, N, and R are written into the bitstream respectively through the values of one or more syntax elements.
[0315] In some embodiments, a neural network model can be used for probability distribution estimation. Specifically, according to the size of any three-dimensional feature block, the context information of the feature value of the three-dimensional feature block can be determined. The context information may include the encoded feature values within the neighborhood of the feature value determined according to the size of the three-dimensional feature block. The context information of the feature value of the three-dimensional feature block can be respectively input into the second neural network model. Obtain the probability distribution vector of the feature value of the three-dimensional feature block output by the second neural network model.
[0316] The second neural network model can be a deep learning network, such as a Recurrent Neural Network (RNN) and a Convolutional Neural Network (PixelCNN), etc.
[0317] As an example, taking a three-dimensional feature block as an example, according to the size of the three-dimensional feature block, the context information of each feature value of the three-dimensional feature block can be determined. The context information of each feature value includes the encoded feature values within the neighborhood of each feature value. The context information of each feature value of the three-dimensional feature block is respectively input into the second neural network model. Obtain the probability distribution vector of each feature value output by the second neural network model respectively. Among them, the encoded feature values within the neighborhood of each feature value can be obtained according to the size of the three-dimensional feature block. The second neural network model can adopt any network structure, for example, a convolutional neural network, or a recurrent neural network, etc.
[0318] Before using the second neural network model to estimate the probability distribution of the image blocks, the neural network model can also be trained through a training process to obtain the second neural network model.
[0319] The training processes of the first neural network model and the second neural network model described above can be independent of each other or correlated with each other.
[0320] In this embodiment, by dividing the image to be processed into multiple image blocks, inputting the multiple image blocks into the first neural network model, obtaining multiple three-dimensional feature blocks output by the first neural network model, encoding the size information of the multiple three-dimensional feature blocks and the multiple three-dimensional feature blocks to obtain an encoded bitstream, and transmitting the size information of the multiple three-dimensional feature blocks from the encoding end to the decoding end, so that the decoding end can perform decoding according to the size information of the multiple three-dimensional feature blocks to obtain a reconstructed image. Since the size information of the multiple three-dimensional feature blocks is added in the image encoding and decoding process, the encoded bitstream at the encoding end can have a high compression ratio, the decoding end can decode the bitstream with a high compression ratio to obtain a reconstructed image, and the relative clarity of the reconstructed image can be ensured. Therefore, the image encoding and decoding efficiency can be improved.
[0321] Figure 12 It is a schematic diagram of the processing process of the image decoding method according to the embodiment of the present application. As Figure 12 shown, perform entropy decoding on the encoded bitstream to obtain three-dimensional feature values, perform feature decoding on the three-dimensional feature values to obtain reconstructed image blocks, and splice the reconstructed image blocks to obtain a reconstructed image.
[0322] Specifically, refer to Figure 13 , Figure 13 It is a flowchart of the image decoding method according to the embodiment of the present application. Figure 13 The method shown can be executed by a decoding device or a decoder, hereinafter collectively referred to as the decoding end. As Figure 13 shown, the method may include:
[0323] Step 501, obtain the bitstream to be decoded.
[0324] The bitstream to be decoded is the above-mentioned encoded bitstream. The bitstream to be decoded may include encoded data of at least two three-dimensional feature blocks and size information of at least two three-dimensional feature blocks. Among them, for the explanation of the size information of the at least two three-dimensional feature blocks, reference can be made to the relevant explanation in Figure 10 step 204 of the shown embodiment, which will not be elaborated here.
[0325] An implementation method is to parse the to-be-decoded bitstream, and the size information of at least two three-dimensional feature blocks obtained includes the size of the three-dimensional feature blocks. For example, by parsing one or more syntax elements in the to-be-decoded bitstream, the length (M), width (N), and height (R) of the three-dimensional feature blocks are obtained respectively.
[0326] Another implementation method is to parse the to-be-decoded bitstream, and the size information of at least two three-dimensional feature blocks obtained includes the size information of at least two image blocks. For example, the size information of at least two image blocks is W and H in the above embodiment. The decoding end can determine the size information of at least two three-dimensional feature blocks according to W and H.
[0327] Still another implementation method is to parse the to-be-decoded bitstream, and the size information of at least two three-dimensional feature blocks obtained includes the size information of at least two image blocks and the parameters of the first neural network model. For example, the size information of at least two image blocks is W and H in the above embodiment, and the parameters of the first neural network model include S and C in the above embodiment. The decoding end can determine the size information of the three-dimensional feature blocks according to W, H, S, and C.
[0328] Method 1: M×N×R = W / S×H / S×C.
[0329] Method 2: The length of the three-dimensional feature block is M = Wb, the width is N = Hb, and the height is R = C.
[0330] Method 3: The length of the three-dimensional feature block is M = Wb, the width is N = Hc, and the height is R = C.
[0331] Among them, the explanations of b, c, and C can be referred to Figure 10 the explanation in step 204 shown, which will not be elaborated here.
[0332] Step 502: Obtain at least two three-dimensional feature blocks according to the encoded data of at least two three-dimensional feature blocks and the size information of at least two three-dimensional feature blocks.
[0333] The entropy decoding can be performed on the encoded data of at least two three-dimensional feature blocks with the size information of at least two three-dimensional feature blocks to obtain at least two three-dimensional feature blocks. Taking the encoded data of one three-dimensional feature block as an example, the encoded data can include multiple to-be-decoded code elements. The entropy decoding is performed on the multiple to-be-decoded code elements (to-be-decoded feature values) according to the size information to obtain multiple feature values, and according to the multiple feature values and the size information, the three-dimensional feature block can be obtained.
[0334] According to the encoded data of at least two three-dimensional feature blocks and the size information of at least two three-dimensional feature blocks, a way to obtain at least two three-dimensional feature blocks can be as follows: According to the size information of any three-dimensional feature block, perform probability distribution estimation on the feature values of any three-dimensional feature block to obtain a probability distribution vector of the feature values. Then, perform entropy decoding on the encoded data of the three-dimensional feature block according to the probability distribution vector of the feature values to obtain the feature values of the three-dimensional feature block.
[0335] Optionally, a neural network can be used for probability distribution estimation. For example, according to the size information of any three-dimensional feature block, determine the context information of the feature values (feature values to be decoded) of the three-dimensional feature block, and input the context information into the third neural network model to obtain the probability distribution vector of the feature values (feature values to be decoded) of the three-dimensional feature block output by the third neural network model. The context information of the feature values (feature values to be decoded) of the three-dimensional feature block can include the decoded feature values within the neighborhood of the feature values (feature values to be decoded) of the three-dimensional feature block. The decoded feature values within the neighborhood of the feature values (feature values to be decoded) of the three-dimensional feature block can be determined according to the size information of the three-dimensional feature block. The third neural network model can adopt any network structure, such as a convolutional neural network or a recurrent neural network, etc.
[0336] Before using the third neural network model to perform probability distribution estimation on the image block, the neural network model can also be trained through a training process to obtain the third neural network model.
[0337] Step 503: Reconstruct at least two image blocks of the image to be processed according to at least two three-dimensional feature blocks.
[0338] According to at least two three-dimensional feature blocks, reconstructed image blocks can be obtained. That is, at least two image blocks of the image to be processed are reconstructed, and then a reconstructed image, that is, the reconstructed image to be processed, is obtained. The reconstructed image blocks can correspond to the image blocks in the above Figure 11 illustrated embodiments.
[0339] Optionally, a neural network can be used for feature decoding. For example, input the multiple feature values of the three-dimensional feature block obtained by entropy decoding and the size information of the three-dimensional feature block into the fourth neural network model to obtain a reconstructed image block, and obtain a reconstructed image according to the reconstructed image block. The fourth neural network model can adopt any network structure, such as a fully connected network, a convolutional neural network, or a recurrent neural network, etc.
[0340] Before using the fourth neural network model to perform probability distribution estimation on the image block, the neural network model can also be trained through a training process to obtain the fourth neural network model.
[0341] In this embodiment, by obtaining the size information of at least two three-dimensional feature blocks from the bitstream to be decoded, and based on the size information of the at least two three-dimensional feature blocks, decoding the encoded data of the at least two three-dimensional feature blocks to obtain a reconstructed image. Since the size information of the three-dimensional feature blocks is added during the image encoding and decoding process, the encoded bitstream at the encoding end can have a high compression ratio, and the decoding end can decode the bitstream with a high compression ratio to obtain a reconstructed image, and it can ensure that the reconstructed image is relatively clear. Therefore, the image encoding and decoding efficiency can be improved.
[0342] Different from the above Figures 9 to 13 shown embodiment, in the following embodiment, a feature quantization operation is added at the encoding end, and a feature inverse quantization operation is added at the decoding end. For specific explanations, please refer to the description of the following embodiment.
[0343] Figure 14 is a schematic diagram of the processing process of the image encoding method according to an embodiment of the present application. As Figure 14 shown, the image to be processed is divided into blocks to obtain a plurality of image blocks, feature extraction is performed on the plurality of image blocks to obtain a plurality of three-dimensional feature blocks, feature quantization is performed on the plurality of three-dimensional feature blocks to obtain a plurality of three-dimensional quantized feature blocks, and entropy encoding is performed on the size information of the plurality of three-dimensional feature blocks and the plurality of three-dimensional quantized feature blocks to obtain an encoded bitstream.
[0344] Specifically, referring to Figure 15 , Figure 15 is a flowchart of the image encoding method according to an embodiment of the present application. Figure 15 The method shown can be executed by an encoding device or an encoder, hereinafter collectively referred to as the encoding end. As Figure 15 shown, the method may include:
[0345] Step 701, obtain the image to be processed.
[0346] Step 702, divide the image to be processed into at least two image blocks.
[0347] Step 703, input the at least two image blocks into a first neural network model, and obtain at least two three-dimensional feature blocks output by the first neural network model.
[0348] Among them, for specific explanations of steps 701 to 703, please refer to Figure 10 steps 201 to 203 of the shown embodiment, which will not be elaborated here.
[0349] Step 704, perform quantization processing on the at least two three-dimensional feature blocks to obtain at least two three-dimensional quantized feature blocks.
[0350] The eigenvalues in at least two three-dimensional feature blocks can be quantized to obtain quantized eigenvalues, also known as quantization feature values. The three-dimensional quantization feature blocks include a plurality of quantized eigenvalues.
[0351] Step 705: Encode at least two three-dimensional quantization feature blocks and encode the size information of at least two three-dimensional feature blocks to obtain an encoded bitstream.
[0352] Among them, the size information of the at least two three-dimensional feature blocks includes: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, or the sizes of the at least two three-dimensional feature blocks. The size information of the at least two image blocks and the parameters of the first neural network model are used to determine the sizes of the at least two three-dimensional feature blocks. For a specific explanation of the size information of the at least two three-dimensional feature blocks, reference can be made to Figure 10 Step 204 of the illustrated embodiment, which will not be elaborated here.
[0353] In the embodiments of the present application, the size information of at least two three-dimensional feature blocks and at least two three-dimensional quantization feature blocks can be encoded to obtain an encoded bitstream. The encoding end sends the encoded bitstream to the decoding end, and the decoding end can obtain the size information of at least two three-dimensional feature blocks from the encoded bitstream, and then decode according to the size information of at least two three-dimensional feature blocks to obtain a reconstructed image.
[0354] It should be noted that the size of any three-dimensional quantization feature block can be the same as the size of the corresponding three-dimensional feature block.
[0355] Optionally, an implementable way to encode at least two three-dimensional quantization feature blocks is to estimate the probability distribution of the quantization feature values of any three-dimensional quantization feature block according to the size of the any three-dimensional quantization feature block to obtain a probability distribution vector of the quantization feature values, and perform entropy encoding on the three-dimensional quantization feature block according to the probability distribution vector of the quantization feature values. The probability distribution vector can include the probability values of each value within the value range, and the sum of all probability values is 1. Taking the encoded bitstream as a binary bitstream as an example, the probability distribution vector can include the probability value of 0 and the probability value of 1, and the sum of the two is 1.
[0356] An implementable way to encode the size information of at least two three-dimensional feature blocks is to perform entropy encoding on the three-dimensional quantization feature blocks to obtain an encoded bitstream, and write the encoded size information of the at least two three-dimensional feature blocks into the bitstream. For example, the encoded size information of the at least two three-dimensional feature blocks can be written into the bitstream through one or more syntax elements.
[0357] In some embodiments, a neural network model can be used for probability distribution estimation. The context information of the quantization feature values of any three-dimensional quantization feature block can be determined according to the size of the three-dimensional quantization feature block. The context information can include the quantization feature values that have been encoded within the neighborhood of the quantization feature values determined according to the size of the three-dimensional quantization feature block. The context information of the quantization feature values of the three-dimensional quantization feature block can be input into a second neural network model. The probability distribution vector of the quantization feature values of the three-dimensional quantization feature block output by the second neural network model is obtained.
[0358] As an example, taking a three-dimensional quantization feature block as an example, the context information of each quantization feature value of the three-dimensional quantization feature block can be determined according to the size of the three-dimensional quantization feature block. The context information of each quantization feature value includes the quantization feature values that have been encoded within the neighborhood of each quantization feature value. The context information of each quantization feature value of the three-dimensional quantization feature block can be separately input into the second neural network model. The probability distribution vector of each quantization feature value output by the second neural network model separately is obtained. Among them, the quantization feature values that have been encoded within the neighborhood of each quantization feature value can be determined according to the size of the three-dimensional quantization feature block.
[0359] In this embodiment, by dividing the image to be processed into multiple image blocks, inputting the multiple image blocks into the first neural network model, obtaining multiple three-dimensional feature blocks output by the first neural network model, quantizing the multiple three-dimensional feature blocks to obtain multiple three-dimensional quantization feature blocks, encoding the size information of the multiple three-dimensional feature blocks and the multiple three-dimensional quantization feature blocks to obtain an encoded bitstream, and the encoding end transmits the size information of the multiple three-dimensional feature blocks to the decoding end, so that the decoding end can perform decoding based on the size information of the multiple three-dimensional feature blocks to obtain a reconstructed image. Since the size information of multiple three-dimensional feature blocks is added during the image encoding and decoding process, the encoded bitstream at the encoding end can have a high compression ratio, the decoding end can decode the bitstream with a high compression ratio to obtain a reconstructed image, and it can ensure that the reconstructed image is relatively clear, so the image encoding and decoding efficiency can be improved.
[0360] Figure 16 It is a schematic diagram of the processing process of the image decoding method according to the embodiment of the present application. As Figure 16 shown, entropy decoding is performed on the encoded bitstream to obtain quantization feature values, inverse quantization is performed on the quantization feature values to obtain three-dimensional feature values, feature decoding is performed on the three-dimensional feature values to obtain reconstructed image blocks, and the reconstructed image blocks are spliced to obtain a reconstructed image.
[0361] Specifically, referring to Figure 17 , Figure 17 It is a flowchart of the image decoding method according to the embodiment of the present application. Figure 17The method shown can be executed by a decoding device or a decoder, hereinafter collectively referred to as the decoding end. As Figure 17 shown, the method may include:
[0362] Step 901, obtaining a bitstream to be decoded.
[0363] For the explanatory note of step 901, reference can be made to Figure 13 the explanatory note of step 501 in the embodiment shown, which will not be elaborated here.
[0364] Step 902, performing entropy decoding on the encoded data of at least two three-dimensional quantization feature blocks according to the size information of at least two three-dimensional feature blocks to obtain the quantization feature values of at least two three-dimensional quantization feature blocks, and performing inverse quantization processing on the quantization feature values of the at least two three-dimensional quantization feature blocks to obtain the feature values of at least two three-dimensional feature blocks.
[0365] Taking the encoded data of one three-dimensional quantization feature block as an example, the encoded data includes a plurality of code elements to be decoded. Entropy decoding is performed on the plurality of code elements to be decoded (quantization feature values to be decoded) according to the size information to obtain a plurality of quantization feature values, inverse quantization processing is performed on the plurality of quantization feature values to obtain a plurality of feature values, and according to the plurality of feature values and the size information, a three-dimensional feature block can be obtained.
[0366] The realizable manner of performing entropy decoding on the encoded data of at least two three-dimensional quantization feature blocks according to the size information of at least two three-dimensional feature blocks to obtain the quantization feature values of at least two three-dimensional quantization feature blocks can be: estimating the probability distribution of the code elements to be decoded (quantization feature values to be decoded of any three-dimensional quantization feature block) according to the size information of any three-dimensional feature block to obtain the probability distribution vector of the code elements to be decoded, and performing entropy decoding on the code elements to be decoded according to the probability distribution vector of the code elements to be decoded to obtain the quantization feature values of the three-dimensional quantization feature block. Then, inverse quantization processing can be performed on the quantization feature values to obtain the feature values of the three-dimensional feature block.
[0367] Optionally, a neural network can be used for probability distribution estimation. For example, according to the size information of any three-dimensional feature block, the context information of the code elements to be decoded can be determined, and the context information of the code elements to be decoded is input into the third neural network model to obtain the probability distribution vector of the code elements to be decoded output by the third neural network model. The context of the code elements to be decoded may include the decoded code elements in the neighborhood of the code elements to be decoded. The third neural network model can adopt any network structure, such as a convolutional neural network or a recurrent neural network, etc. The decoded code elements in the neighborhood of the code elements to be decoded can be determined according to the size information of the three-dimensional feature block.
[0368] Before using the third neural network model to estimate the probability distribution of image patches, the neural network model can also be trained through a training process to obtain the third neural network model.
[0369] Step 903: Reconstruct at least two image patches of the image to be processed according to the eigenvalue of at least two three-dimensional feature patches and the size information of at least two three-dimensional feature patches.
[0370] Perform feature decoding on the eigenvalues of at least two three-dimensional feature patches to obtain the reconstructed image patches, and obtain the reconstructed image (the reconstructed image of the image to be processed) according to the reconstructed image patches. The reconstructed image patches can correspond to the image patches in the above Figure 11 illustrated embodiments.
[0371] Optionally, a neural network can be used for feature decoding. For example, the eigenvalue obtained by inverse quantization and the size information of the three-dimensional feature patch can be input into the fourth neural network model to obtain the reconstructed image patches, and the reconstructed image can be obtained according to the reconstructed image patches. The fourth neural network model can adopt any network structure, such as a fully connected network, a convolutional neural network, or a recurrent neural network, etc.
[0372] Before using the fourth neural network model to estimate the probability distribution of image patches, the neural network model can also be trained through a training process to obtain the fourth neural network model.
[0373] In this embodiment, by obtaining the size information of at least two three-dimensional feature patches from the code stream to be decoded, and based on the size information of at least two three-dimensional feature patches, decoding the encoded data of at least two three-dimensional feature patches to obtain the reconstructed image. Since the size information of the three-dimensional feature patches is added in the image encoding and decoding process, the encoded code stream at the encoding end can have a high compression ratio, and the decoding end can decode the code stream with a high compression ratio to obtain the reconstructed image, and it can ensure that the reconstructed image is relatively clear, so the image encoding and decoding efficiency can be improved.
[0374] This embodiment relates to an end-to-end image coding scheme, such as Figure 18As shown, it mainly includes the following parts: a block partitioning module, a feature extraction module, a feature quantization module, entropy encoding, entropy decoding, a feature inverse quantization module, and a feature decoding module. At the encoding end, the block partitioning module divides the image to be compressed into multiple sub-image blocks, and the feature extraction module uses a neural network-based method to obtain the extracted three-dimensional feature map (or three-dimensional feature block). The feature quantization module quantizes the feature values of the three-dimensional feature block through eigenvalue quantization to obtain the quantized feature values. The quantized feature values are subjected to lossless entropy encoding to obtain the encoded bitstream. At the decoding end, when the encoded bitstream is received, first, lossless entropy decoding is performed to obtain the three-dimensional quantized feature values. Through the feature decoding module, the three-dimensional feature values are decoded into a reconstructed image, thus achieving the decoding purpose. In some cases, there is no feature quantization module at the encoding end, and correspondingly, there is no feature inverse quantization module at the decoding end.
[0375] A novel image encoding method according to an embodiment of the present application can further reduce the size of the compressed file and lower the storage cost of video image files on the server side.
[0376] The above has introduced the image encoding method and decoding method of the embodiments of the present application in detail with reference to the accompanying drawings. Next, in combination with Figures 19 to 22 the image encoding device and image decoding device of the embodiments of the present application will be introduced. It should be understood that Figures 19 to 22 the image encoding device in Figures 19 to 22 can execute the image encoding method of the embodiments of the present application,
[0377] Figure 19 is a schematic block diagram of the image encoding device according to an embodiment of the present application.
[0378] Figure 19 The image encoding device 10000 shown may include an acquisition module 10001 and a processing module 10002. The image encoding device 10000 can execute the image encoding method of the embodiments of the present application. Specifically, the image encoding device 10000 can execute Figure 9 or Figure 10 or Figure 14 or Figure 15 the image encoding methods in
[0379] Figure 20 is a schematic block diagram of the image decoding device according to an embodiment of the present application.
[0380] Figure 20The image decoding device 11000 shown includes an acquisition module 11001 and a processing module 11002. The image decoding device 11000 can execute the image decoding method of the embodiments of the present application. Specifically, the image decoding device 11000 can execute Figure 12 or Figure 13 or Figure 16 or Figure 17 the image decoding method in.
[0381] Figure 21 is a schematic diagram of the hardware structure of the image encoding device of the embodiments of the present application.
[0382] Figure 21 The image encoding device 12000 shown (the image encoding device 12000 can specifically be a computer device) includes a memory 12001, a memory 12002, a communication interface 12003, and a bus 12004. Among them, the memory 12001, the memory 12002, and the communication interface 12003 are communicatively connected to each other through the bus 12004.
[0383] The memory 12001 can be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 12001 can store a program. When the program stored in the memory 12001 is executed by the memory 12002, the memory 12002 is used to execute each step of the image encoding method of the embodiments of the present application.
[0384] The memory 12002 can adopt a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits, and is used to execute relevant programs to implement the image encoding method of the method embodiments of the present application.
[0385] The memory 12002 can also be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the image encoding method of the present application can be completed by the integrated logic circuit in the hardware in the memory 12002 or the instructions in software form.
[0386] The aforementioned memory 12002 can also be a general - purpose processor, a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application.
[0387] The general - purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read - only memory, programmable read - only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in memory 12001, and the processor 12002 reads the information in memory 12001 and combines its hardware to complete the functions required to be executed by the modules included in this image encoding device, or executes the image encoding method of the method embodiment of the present application.
[0388] The communication interface 12003 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the image encoding device 12000 and other devices or communication networks. For example, the image to be processed can be obtained through the communication interface 12003.
[0389] The bus 12004 can include a path for transmitting information between various components of the image encoding device 12000 (for example, memory 12001, memory 12002, communication interface 12003).
[0390] The acquisition module 10001 and the processing module 10002 in the aforementioned image encoding device 10000 are equivalent to the processor 12002 in the image encoding device 12000.
[0391] Figure 22 It is a schematic diagram of the hardware structure of the image decoding device for the embodiments of the present application. Figure 22 The shown image decoding device 13000 (this image decoding device 13000 can specifically be a computer device) includes a memory 13001, a processor 13002, a communication interface 13003, and a bus 13004. Among them, the memory 13001, the processor 13002, and the communication interface 13003 are communicatively connected to each other through the bus 13004.
[0392] The definitions and explanations of the various modules in the image decoding device 12000 in the above text also apply to the image decoding device 13000, and will not be described in detail here.
[0393] The above-mentioned memory 13001 can be used to store programs, and the processor 13002 is used to execute the programs stored in the memory 13001. When the programs stored in the memory 13001 are executed, the processor 13002 is used to execute the various steps of the image decoding method according to the embodiments of the present application.
[0394] The acquisition module 11001 and the processing module 11002 in the above-mentioned image decoding device 11000 are equivalent to the processor 13002 in the image decoding device 13000.
[0395] Those skilled in the art can appreciate that the functions described in connection with the various illustrative logical blocks, modules, and algorithm steps disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions described in the various illustrative logical blocks, modules, and steps can be stored or transmitted as one or more instructions or codes on a computer-readable medium and executed by a hardware-based processing unit. The computer-readable medium can include a computer-readable storage medium corresponding to a tangible medium, such as a data storage medium, or a communication medium including any medium that facilitates the transfer of a computer program from one place to another (e.g., according to a communication protocol). In this way, the computer-readable medium generally corresponds to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium, such as a signal or a carrier wave. The data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, codes, and / or data structures for implementing the techniques described in the present application. A computer program product can include a computer-readable medium.
[0396] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. As used herein, disk and optical disks include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), and Blu-ray disk, where disks typically reproduce data magnetically, while optical disks utilize lasers to optically reproduce data. Combinations of the above should also be included within the scope of computer-readable media.
[0397] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functions described for the various illustrative logical blocks, modules, and steps can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Moreover, the techniques can be fully implemented in one or more circuits or logic elements.
[0398] The techniques of this application can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs), or a group of ICs (e.g., a chip set). The various components, modules, or units described in this application are described to emphasize functional aspects of the devices for performing the disclosed techniques, but need not be implemented by different hardware units. In fact, as described above, the various units can be combined in a codec hardware unit with suitable software and / or firmware, or provided by interoperating hardware units including one or more processors as described above.
[0399] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the said claims.
Claims
1. An image encoding method, characterized in that, The method includes: Obtaining an image to be processed; Dividing the image to be processed into at least two image blocks; Inputting the at least two image blocks into a first neural network model to obtain at least two three-dimensional feature blocks output by the first neural network model; Encoding the at least two three-dimensional feature blocks and encoding the size information of the at least two three-dimensional feature blocks to obtain an encoded bitstream; Wherein, the size information of the at least two three-dimensional feature blocks includes: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, or the size of the three-dimensional feature blocks, and the size information of the at least two image blocks and the parameters of the first neural network model are used to determine the size of the at least two three-dimensional feature blocks.
2. The method according to claim 1, wherein The size information of the at least two image blocks includes the height and width of the image blocks, or the size information of the at least two image blocks includes the partitioning method applied to the image to be processed and used to divide the at least two image blocks.
3. The method according to claim 1, wherein The parameters of the first neural network model include at least one of the number of channels of the convolutional kernel or the scaling step.
4. The method according to claim 3, characterized in that, The size of the at least two three-dimensional feature blocks includes the length, width, and height of the at least two three-dimensional feature blocks, and the size information of the at least two image blocks includes the width and height of the at least two image blocks; The correspondence between the length, width, and height of any one of the at least two three-dimensional feature blocks and the width and height of the corresponding one of the at least two image blocks is as follows: M×N×R = W / S×H / S×C; Wherein, M is the length of any one of the three-dimensional feature blocks, N is the width of any one of the three-dimensional feature blocks, R is the height of any one of the three-dimensional feature blocks, W is the width of any one of the image blocks, H is the height of any one of the image blocks, C is the number of channels of the convolutional kernel, and S is the scaling step.
5. The method according to claim 1, characterized in that, The size of the at least two three-dimensional feature blocks includes the length, width, and height of the three-dimensional feature blocks.
6. The method according to any one of claims 1 to 5, characterized in that Any one of the at least two three-dimensional feature blocks includes a plurality of eigenvalues, and the encoding of the at least two three-dimensional feature blocks includes: Estimating the probability distribution of the eigenvalues of any one of the three-dimensional feature blocks according to the size of the any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues; Performing entropy encoding on any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalues.
7. The method according to claim 6, characterized in that The estimating the probability distribution of the eigenvalues of any one of the three-dimensional feature blocks according to the size of the any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues includes: Determining the context information of the eigenvalues according to the size of any one of the three-dimensional feature blocks, where the context information includes the encoded eigenvalues within the neighborhood of the eigenvalues determined according to the size of any one of the three-dimensional feature blocks; Inputting the context information of the eigenvalues into a second neural network model; Obtaining the probability distribution vector of the eigenvalues output by the second neural network model.
8. The method according to any one of claims 1 to 5, characterized in that The encoding of the at least two three-dimensional feature blocks includes: Quantize the at least two three-dimensional feature blocks to obtain at least two three-dimensional quantized feature blocks, and any one of the at least two three-dimensional quantized feature blocks includes a plurality of quantization feature values; Estimate the probability distribution of the quantization feature values of any one of the at least two three-dimensional quantized feature blocks according to the size of any one of the at least two three-dimensional quantized feature blocks to obtain a probability distribution vector of the quantization feature values; the size of any one of the at least two three-dimensional quantized feature blocks is the same as the size of the corresponding one of the at least two three-dimensional feature blocks; Perform entropy coding on any one of the at least two three-dimensional quantized feature blocks according to the probability distribution vector of the quantization feature values.
9. The method according to claim 8, wherein The estimating the probability distribution of the quantization feature values of any one of the at least two three-dimensional quantized feature blocks according to the size of any one of the at least two three-dimensional quantized feature blocks to obtain a probability distribution vector of the quantization feature values includes: Determine the context information of the quantization feature values according to the size of any one of the at least two three-dimensional quantized feature blocks, where the context information includes the quantized feature values that have been encoded in the neighborhood of the quantization feature values determined according to the size of any one of the at least two three-dimensional quantized feature blocks; Input the context information of the quantization feature values into a third neural network model; Obtain the probability distribution vector of the quantization feature values output by the third neural network model.
10. An image decoding method, characterized in that, The method includes: Obtain a bitstream to be decoded, where the bitstream to be decoded includes the encoded data of at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks; Obtain the at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks; Reconstruct at least two image blocks of the image to be processed according to the at least two three-dimensional feature blocks.
11. The method according to claim 10, characterized in that The size information of the at least two three-dimensional feature blocks includes the sizes of the at least two three-dimensional feature blocks, and the sizes of the at least two three-dimensional feature blocks include the length, width, and height of the three-dimensional feature blocks.
12. The method according to claim 10, wherein The size information of the at least two three-dimensional feature blocks includes: the size information of the at least two image blocks, or the size information of the at least two image blocks and the parameters of the first neural network model, where the size information of the at least two image blocks and the parameters of the first neural network model are used to determine the sizes of the at least two three-dimensional feature blocks.
13. The method according to claim 12, wherein The size information of the at least two image blocks includes the height and width of the at least two image blocks.
14. The method according to claim 12, characterized in that, The parameters of the first neural network model include at least one of the number of channels of the convolutional kernel or the scaling step; the size information of the at least two image blocks includes the width and height of the at least two image blocks; The corresponding relationship between the length, width, and height of any one of the at least two three-dimensional feature blocks and the width and height of the corresponding one of the at least two image blocks is as follows: M×N×R=W / S×H / S×C; where M is the length of any one of the at least two three-dimensional feature blocks, N is the width of any one of the at least two three-dimensional feature blocks, R is the height of any one of the at least two three-dimensional feature blocks, W is the width of any one of the at least two image blocks, H is the height of any one of the at least two image blocks, C is the number of channels of the convolutional kernel, and S is the scaling step.
15. The method according to any one of claims 10 to 14, characterized in that, The at least two three-dimensional feature blocks include: a plurality of eigenvalues of any one of the at least two three-dimensional feature blocks, and size information of the at least two three-dimensional feature blocks.
16. The method according to claim 15, wherein The obtaining of the at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks includes: Performing a probability distribution estimation on the eigenvalues of any one of the three-dimensional feature blocks according to the size information of the any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues; Performing entropy decoding on the encoded data of the any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalues to obtain the eigenvalues.
17. The method according to claim 15, wherein The obtaining of the at least two three-dimensional feature blocks according to the encoded data of the at least two three-dimensional feature blocks and the size information of the at least two three-dimensional feature blocks includes: Performing a probability distribution estimation on the eigenvalues of any one of the three-dimensional feature blocks according to the size information of the any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues; Performing entropy decoding on the encoded data of the any one of the three-dimensional feature blocks according to the probability distribution vector of the eigenvalues to obtain quantized eigenvalues of the eigenvalues; Performing an inverse quantization process on the quantized eigenvalues to obtain the eigenvalues.
18. The method according to claim 16 or 17, characterized in that, The performing of a probability distribution estimation on the eigenvalues of any one of the three-dimensional feature blocks according to the size information of the any one of the three-dimensional feature blocks to obtain a probability distribution vector of the eigenvalues includes: Determining context information of the eigenvalues according to the size information of the any one of the three-dimensional feature blocks, where the context information includes the decoded eigenvalues in the neighborhood of the eigenvalues determined according to the size information of the any one of the three-dimensional feature blocks; Inputting the context information of the eigenvalues into a second neural network model; Obtaining the probability distribution vector of the eigenvalues output by the second neural network model.
19. The method according to any one of claims 10 to 18, characterized in that, The reconstructing of at least two image blocks of a to-be-processed image according to the at least two three-dimensional feature blocks includes: Inputting the at least two three-dimensional feature blocks into a third neural network model to obtain reconstructed image blocks of the at least two image blocks output by the third neural network model.
20. An image encoding device, characterized in that, including: A non-volatile memory and a processor coupled to each other, and the processor calls program code stored in the memory to execute the method according to any one of claims 1 to 9.
21. An image decoding apparatus, characterized in that, including: A non-volatile memory and a processor coupled to each other, and the processor calls program code stored in the memory to execute the method according to any one of claims 10 to 19.
22. An image encoding device, characterized in that, including: an encoder, and the encoder is used to execute the method according to any one of claims 1 to 9.
23. An image decoding device, characterized in that, including: A decoder, and the decoder is used to execute the method according to any one of claims 10 to 19.
24. A computer-readable storage medium, characterized in that, including an encoded bitstream obtained according to the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image data extraction using neural networks
US20190384970A1