Enhanced layer encoding and decoding method and apparatus
Through the adaptive transform block division method and direct difference calculation method, the problem of low encoding efficiency in the existing video encoding technology is solved, and more efficient video encoding is achieved, especially the compression efficiency of the enhancement layer is improved.
Patent Information
- Application Number
- CN202011424168.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-12-08
AI Technical Summary
When existing video encoding technologies transmit video data in networks with limited bandwidth, it is difficult to effectively improve the encoding efficiency, especially in the H.265/HEVC standard, the encoding efficiency of brightness quantization parameters needs to be improved.
Adaptive transformation block division method is adopted to divide the residual blocks of the enhancement layer through iterative tree structure, determine the optimal division method through rate distortion optimization, reduce the processing flow of the encoder, and obtain the residual blocks of the enhancement layer by using the direct difference method to avoid using the same transformation block division method as the basic layer.
It improves the compression efficiency of the enhancement layer, reduces the encoder processing flow, improves the encoding efficiency, and is suitable for hierarchical video encoding in airspace and quality domains.
Smart Images

Figure CN114615500B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of video or image compression, and particularly to an enhancement layer encoding and decoding method and apparatus. Background Art
[0002] Video encoding (video encoding and decoding) is widely used in digital video applications, such as broadcast digital television, video transmission over the Internet and mobile networks, real-time session applications such as video chat and video conferencing, DVDs and Blu-ray discs, video content acquisition and editing systems, and security applications for portable cameras.
[0003] Even in the case of short films, a large amount of video data needs to be described, which may cause difficulties when the data is to be sent over a network with limited bandwidth capacity or transmitted in other ways. Therefore, video data is usually compressed first and then transmitted in modern telecommunication networks. Since memory resources may be limited, the size of the video may also be a problem when storing the video on a storage device. Video compression devices typically use software and / or hardware on the source side to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. Then, the compressed data is received by a video decompression device on the destination side. In the context of limited network resources and the growing demand for higher video quality, improved compression and decompression technologies are needed, which can increase the compression ratio with little impact on the image quality.
[0004] In the H.265 / HEVC standard, a coding unit (CU) contains a luminance quantization parameter (QP) and two chrominance quantization parameters, where the chrominance quantization parameters can be derived from the luminance quantization parameter. Therefore, improving the encoding efficiency of the luminance quantization parameter becomes a key technology for improving video encoding efficiency. Summary of the Invention
[0005] The present application provides an enhancement layer encoding and decoding method and apparatus, which can reduce the processing flow of the encoder, improve the encoding efficiency of the encoder, and by adopting an adaptive TU partitioning method, the TU size of the enhancement layer is no longer limited by the CU size, and the compression efficiency of the residual block can be more effectively improved.
[0006] In a first aspect, the present application provides an enhancement layer encoding and decoding method, including: The encoder obtains a reconstructed block of the base layer of the image block to be encoded. The encoder calculates the difference between the corresponding pixel points of the image block to be encoded and the reconstructed block of the base layer to obtain a residual block of the enhancement layer of the image block to be encoded. The encoder determines the transform block partitioning method of the residual block of the enhancement layer. The encoder performs a transform on the residual block of the enhancement layer according to the transform block partitioning method to obtain a bitstream of the residual block of the enhancement layer. The decoder obtains the bitstream of the residual block of the enhancement layer of the image block to be decoded. The decoder performs entropy decoding, inverse quantization, and inverse transform on the bitstream of the residual block of the enhancement layer to obtain a reconstructed residual block of the enhancement layer. The decoder obtains the reconstructed block of the base layer of the image block to be decoded. The decoder sums the corresponding pixel points of the reconstructed residual block of the enhancement layer and the reconstructed block of the base layer to obtain a reconstructed block of the enhancement layer.
[0007] Scalable video coding, also known as scalable video encoding, is an extended coding standard of current video coding standards (generally an extended standard of advanced video coding (AVC) (H.264), scalable video coding (SVC), or an extended standard of high efficiency video coding (HEVC) (H.265), scalable high efficiency video coding (SHVC)). The emergence of scalable video coding is mainly to solve the problems of packet loss and delay jitter caused by the real-time change of network bandwidth in real-time video transmission.
[0008] The basic structure in scalable video coding can be called a layer. Through spatial scalability (resolution scalability) of the original image block, scalable video coding technology can obtain bitstreams of different resolution layers. Resolution can refer to the size of the image block in pixels. The resolution of the lower layer is lower, and the resolution of the higher layer is not lower than that of the lower layer; or, through temporal scalability (frame rate scalability) of the original image block, bitstreams of different frame rate layers can be obtained. Frame rate can refer to the number of image frames contained in the video per unit time. The frame rate of the lower layer is lower, and the frame rate of the higher layer is not lower than that of the lower layer; or, through quality scalability of the original image block, bitstreams of different coding quality layers can be obtained. Coding quality can refer to the quality of the video. The image distortion degree of the lower layer is larger, and the image distortion degree of the higher layer is not higher than that of the lower layer.
[0009] Generally, the layer called the base layer is the bottom layer in scalable video coding. In spatial scalability, the base layer image blocks are encoded using the lowest resolution; in temporal scalability, the base layer image blocks are encoded using the lowest frame rate; in quality scalability, the base layer image blocks are encoded using the highest QP or the lowest bitrate. That is, the base layer is the layer with the lowest quality in scalable video coding. The layer called the enhancement layer is the layer above the base layer in scalable video coding, which can be divided into multiple enhancement layers from low to high. The lowest-layer enhancement layer obtains the encoded information based on the base layer, and its encoded resolution is higher than that of the base layer, or the frame rate is higher than that of the base layer, or the bitrate is larger than that of the base layer. The higher-layer enhancement layers can encode higher-quality image blocks based on the encoded information of the lower-layer enhancement layers.
[0010] For example, Figure 9 is an exemplary layer schematic diagram of the scalable video coding of this application. As Figure 9 shown, after the original image blocks are sent into the scalable encoder, they can be layered into base layer image blocks B and enhancement layer image blocks (E1 to En, n≥1) according to different coding configurations, and then encoded respectively to obtain a bitstream containing the base layer bitstream and the enhancement layer bitstream. The base layer bitstream is generally the bitstream obtained by encoding the lowest spatial image blocks, the lowest temporal image blocks, or the lowest quality image blocks. The enhancement layer bitstream is a bitstream obtained by taking the base layer as the basis and superimposing and encoding the high-level spatial, high-level temporal, or high-level quality image blocks. As the number of enhancement layers increases, the encoded spatial level, temporal level, or quality level will also become higher and higher. When the encoder transmits the bitstream to the decoder, it first ensures the transmission of the base layer bitstream. When there is network surplus, it gradually transmits the bitstreams of higher and higher levels. The decoder first receives and decodes the base layer bitstream, and then, according to the received enhancement layer bitstream, in the order from the lower level to the higher level, decodes the bitstreams with higher and higher levels of spatial, temporal, or quality levels layer by layer, and then superimposes the decoded information of the higher level on the reconstructed blocks of the lower level to obtain reconstructed blocks with higher resolution, higher frame rate, or higher quality.
[0011] The image block to be encoded can refer to the image block that the encoder is currently processing, and this image block can refer to the largest coding unit (LCU) in the entire frame image. The entire frame image can refer to any image frame in the image sequence included in the video being processed by the encoder. This image frame has not been divided and its size is the size of a complete image frame. In the H.265 standard, before video encoding, the original image frame is divided into multiple coding tree units (CTUs). A CTU is the largest coding unit for video encoding and can be divided into CUs of different sizes in a quadtree manner. As the largest coding unit, a CTU is also called an LCU; alternatively, this image block can also refer to the entire frame image; or this image block can also refer to the region of interest (ROI) in the entire frame image, that is, a specified image area in the image that needs to be processed.
[0012] As described above, each image in a video sequence is usually segmented into a set of non-overlapping blocks and is usually encoded at the block level. In other words, the encoder usually processes and encodes video at the block (image block) level. For example, prediction blocks are generated through spatial (intra-frame) prediction and temporal (inter-frame) prediction; the prediction blocks are subtracted from the image block (the currently processed / pending block) to obtain a residual block; the residual block is transformed and quantized in the transform domain to reduce the amount of data to be transmitted (compressed). The encoder also needs to perform inverse quantization and inverse transformation to obtain a reconstructed residual block, and then add the pixel values of the reconstructed residual block and the pixel values of the prediction block to obtain a reconstructed block. The reconstructed block of the base layer refers to the reconstructed block obtained by performing the above operations on the base layer image block obtained by hierarchically dividing the original image block. The encoder obtains the prediction block of the base layer based on the original image block (such as an LCU), then calculates the difference between the corresponding pixels in the original image block and the prediction block of the base layer to obtain the residual block of the base layer. After dividing the residual block of the base layer, it is transformed, quantized, and jointly entropy-encoded with the base layer coding control information, prediction information, motion information, etc. to obtain the bitstream of the base layer. The encoder performs inverse quantization and inverse transformation on the quantized quantization coefficients to obtain the reconstructed residual block of the base layer, and then sums the corresponding pixels in the prediction block of the base layer and the reconstructed residual block of the base layer to obtain the reconstructed block of the base layer.
[0013] The resolution of the enhancement layer image block is not lower than that of the base layer image block, or the coding quality of the enhancement layer image block is not lower than that of the base layer image block. As described above, the coding quality can refer to the quality of the video. The coding quality of the enhancement layer image block not being lower than that of the base layer image block can mean that the image distortion degree of the base layer is larger, while the image distortion degree of the enhancement layer is not higher than that of the base layer.
[0014] An image block includes a plurality of pixel points, which are arranged in an array. Each pixel point can uniquely identify its position in the image block using a row number and a column number. Assume that the sizes of the image block to be encoded and the reconstructed block of the base layer are both M×N, that is, the image block to be encoded and the reconstructed block of the base layer each contain M×N pixel points. a(i1, j1) represents the pixel point at the i1-th column and j1-th row in the image block to be encoded, where i1 = 1 to M and j1 = 1 to N. b(i2, j2) represents the pixel point at the i1-th column and j1-th row in the reconstructed block of the base layer, where i2 = 1 to M and j2 = 1 to N. The corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer refer to the pixel points whose row numbers and column numbers in their respective image blocks are equal, that is, i1 = i2 and j1 = j2. For example, if the size of the image block is 16×16, it contains 16×16 pixel points, with row numbers from 0 to 15 and column numbers from 0 to 15. The pixel point labeled a(0, 0) in the image block to be encoded and the pixel point labeled b(0, 0) in the reconstructed block of the base layer are corresponding pixel points, or the pixel point labeled a(6, 9) in the image block to be encoded and the pixel point labeled b(6, 9) in the reconstructed block of the base layer are corresponding pixel points, and so on. Taking the difference between corresponding pixel points can be to take the difference between the pixel values of the corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer. The pixel value can be the luminance value, chrominance value, etc. of the pixel point. This application does not make specific limitations on this.
[0015] In the SHVC standard, the encoder will predict the prediction block of the enhancement layer based on the reconstructed block of the base layer, and then take the difference between the corresponding pixel points in the reconstructed block of the base layer and the prediction block of the enhancement layer to obtain the residual block of the enhancement layer. This application can directly take the difference between the corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer to obtain the residual block of the enhancement layer, reducing the process of obtaining the prediction block of the enhancement layer, which can reduce the processing flow of the encoder and improve the encoding efficiency of the encoder.
[0016] In this application, the transform block partitioning method of the residual block of the enhancement layer is different from that of the residual block of the base layer. That is, when the encoder processes the residual block of the enhancement layer, the transform block partitioning method used is different from that used when processing the residual block of the base layer. For example, the residual block of the base layer is divided into three sub-blocks using the TT partitioning method, but the residual block of the enhancement layer is not divided using the TT partitioning method. The encoder can use the following adaptive method to determine the transform block partitioning method of the residual block of the enhancement layer:
[0017] In a possible implementation, the encoder may first perform iterative tree-structured partitioning on the first largest transform unit (LTU) of the residual block of the enhancement layer to obtain transform units (TUs) with multiple partitioning depths. The maximum partitioning depth among the multiple partitioning depths is equal to Dmax, where Dmax is a positive integer. Then, the partitioning method of the first TU is determined based on the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU. The partitioning depth of the first TU is i, and the depth of the second TU is i + 1, where 0 ≤ i ≤ Dmax - 1.
[0018] Corresponding to the size of the image block to be encoded, the size of the residual block of the enhancement layer can be the size of the entire frame image, or the size of the image block (such as CTU or CU) divided in the entire frame image, or the size of the ROI in the entire frame image. The LTU is the largest-sized image block in the image block for performing transformation processing. The size of the LTU can be the same as the size of the reconstruction block of the base layer, which can maximize the efficiency of transform coding while ensuring parallel processing of the enhancement layer between different reconstruction blocks. An LTU can be divided into multiple nodes based on the configured partitioning information, and each node can be further divided based on the configured partitioning information until all nodes are no longer divided. This process can be called iterative tree-structured partitioning. The tree-structured partitioning can include QT partitioning, BT partitioning, and / or TT partitioning, and can also include EQT partitioning. The present application does not specifically limit the partitioning method of the LTU. A node is divided once to obtain multiple nodes. The divided node is called the parent node, and the divided nodes are called child nodes. The depth of the root node is 0, and the depth of the child node is the depth of the parent node plus 1. Therefore, iterative tree-structured partitioning starts from the root node LTU of the residual block of the enhancement layer, and TUs with multiple depths can be obtained. The maximum depth (i.e., the depth of the smallest TU obtained by partitioning) among the multiple depths is Dmax. Assuming that the width and height of the LTU are both L, and the width and height of the smallest TU are both S, then the partitioning depth of the smallest TU where represents rounding down, so Dmax ≤ d.
[0019] Based on the TUs obtained by the above transformation partitioning, the encoder performs rate distortion optimization (RDO) processing to determine the partitioning method of the LTU. Taking the first TU with a partitioning depth of i and the second TU with a depth of i + 1 as an example, i starts from Dmax - 1 and decreases. The second TU is obtained by partitioning the first TU. The encoder first performs transformation and quantization on the first TU to obtain the quantization coefficients of the first TU (the quantization coefficients can be as Figure 10obtained through the quantization process in [reference], then pre-encode the quantization coefficients of the first TU (pre-encoding is a coding process to estimate the length of the coded word after encoding, or a process using a method approximate to encoding) to obtain the coded word length (i.e., the bitstream size) R of the first TU. Then, perform inverse quantization and inverse transformation on the quantization coefficients of the first TU to obtain the reconstructed block of the TU, calculate the sum of the squared errors between the first TU and the reconstructed block of the first TU to obtain the distortion value D of the first TU. Finally, obtain the loss estimation value C of the first TU according to the bitstream size R of the first TU and the distortion value D of the first TU.
[0020] The calculation formula for the distortion value D of the first TU is as follows:
[0021]
[0022] where, P rs (i,j) represents the original value of the residual pixel in the first TU at the coordinate point (i,j) within the range of this TU, and P rc (i,j) represents the reconstructed value of the residual pixel in the first TU at the coordinate point (i,j) within the range of this TU.
[0023] The calculation formula for the loss estimation value C of the first TU is as follows:
[0024] C = D + λR
[0025] where, λ represents a constant value related to the quantization coefficients of the current layer, which determines the pixel distortion situation of the current layer.
[0026] The encoder can calculate the loss estimation value of the second TU using the same method as above. After obtaining the loss estimation value of the first TU and the loss estimation values of the second TUs, the encoder compares the loss estimation value of the first TU with the sum of the loss estimation values of multiple second TUs. The first TU is divided into multiple second TUs. The division method corresponding to the smaller of the two is determined as the division method of the first TU. That is, if the loss estimation value of the first TU is greater than the sum of the loss estimation values of multiple second TUs, then the division method of dividing the first TU into multiple second TUs is determined as the division method of the first TU; if the loss estimation value of the first TU is less than or equal to the sum of the loss estimation values of multiple second TUs, then the first TU is no longer divided. After the encoder traverses all the TUs of the LTU through the above method, the division method of the LTU can be obtained. When the residual block of the enhancement layer contains only one LTU, the division method of the LTU is the division method of the residual block of the enhancement layer; when the residual block of the enhancement layer contains multiple LTUs, the division method of the LTU and the method of dividing the residual block of the enhancement layer into LTUs constitute the division method of the residual block of the enhancement layer.
[0027] In this application, the TU partitioning method for the residual blocks in the enhancement layer is no longer the same as that for the residual blocks in the base layer. Instead, an adaptive TU partitioning method is adopted to obtain a TU partitioning method suitable for the residual blocks in the enhancement layer, and then encoding is performed. When the TU partitioning of the residual blocks in the enhancement layer is independent of the CU partitioning method of the base layer, the TU size of the enhancement layer is no longer limited by the CU size, which can improve the flexibility of encoding.
[0028] The encoder can perform transformation, quantization, and entropy coding on the residual blocks in the enhancement layer according to the above transformation block partitioning method to obtain the bitstream of the residual blocks in the enhancement layer. The methods of transformation, quantization, and entropy coding can refer to the above description and will not be elaborated here.
[0029] After the decoder obtains the bitstream of the residual blocks in the enhancement layer, it performs entropy decoding, inverse quantization, and inverse transformation processing on it based on the syntax elements carried in the bitstream to obtain the reconstructed residual blocks in the enhancement layer. The decoder adds the pixel points with the same row numbers and column numbers in the two image blocks according to the row numbers and column numbers of the multiple pixel points included in the reconstructed residual blocks in the enhancement layer and the row numbers and column numbers of the multiple pixel points included in the reconstructed blocks in the base layer to obtain the reconstructed blocks in the enhancement layer.
[0030] This application can directly calculate the difference between the corresponding pixel points in the image block to be encoded and the reconstructed blocks in the base layer to obtain the residual blocks in the enhancement layer, reducing the process of obtaining the prediction blocks in the enhancement layer, which can reduce the processing flow of the encoder and improve the encoding efficiency of the encoder. In addition, the TU partitioning method for the residual blocks in the enhancement layer is no longer the same as that for the residual blocks in the base layer. Instead, an adaptive TU partitioning method is adopted to obtain a TU partitioning method suitable for the residual blocks in the enhancement layer, and then encoding is performed. When the TU partitioning of the residual blocks in the enhancement layer is independent of the CU partitioning method of the base layer, the TU size of the enhancement layer is no longer limited by the CU size, which can more effectively improve the compression efficiency of the residual blocks.
[0031] The above method can be applied to two hierarchical methods: spatial domain hierarchical and quality domain hierarchical. When the image blocks are hierarchically divided in the spatial domain to obtain the base layer image blocks and the enhancement layer image blocks, the resolution of the base layer image blocks is smaller than that of the enhancement layer image blocks. Therefore, before the encoder calculates the difference between the corresponding pixel points in the image block to be encoded and the reconstructed blocks in the base layer to obtain the residual blocks in the enhancement layer of the image block to be encoded, it can downsample the original image block to be encoded to obtain the image block to be encoded with the first resolution, and upsample the original reconstructed blocks in the base layer to obtain the reconstructed blocks in the base layer with the first resolution. That is, the encoder can downsample the original image block to be encoded and upsample the original reconstructed blocks in the base layer respectively to make the resolution of the image block to be encoded the same as that of the reconstructed blocks in the base layer.
[0032] Second aspect, the present application provides an enhanced layer encoding and decoding method, including: The encoder obtains the reconstructed block of the first enhanced layer of the image block to be encoded. The encoder calculates the difference between the corresponding pixel points of the image block to be encoded and the reconstructed block of the first enhanced layer to obtain the residual block of the second enhanced layer of the image block to be encoded. The encoder determines the transform block partitioning method of the residual block of the second enhanced layer. The encoder transforms the residual block of the second enhanced layer according to the transform block partitioning method to obtain the bitstream of the residual block of the second enhanced layer. The decoder obtains the bitstream of the residual block of the second enhanced layer of the image block to be decoded. The decoder performs entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the second enhanced layer to obtain the reconstructed residual block of the second enhanced layer. The decoder obtains the reconstructed block of the first enhanced layer of the image block to be decoded. The decoder sums the corresponding pixel points of the reconstructed residual block of the second enhanced layer and the reconstructed block of the first enhanced layer to obtain the reconstructed block of the second enhanced layer.
[0033] In the embodiments of the present application, after the image blocks are layered using spatial domain grading, temporal domain grading, or quality domain grading, in addition to the base layer, they are also divided into multiple enhanced layers. Compared with the first aspect, this embodiment is an encoding and decoding method for the second enhanced layer of the higher layer based on the lower layer enhanced layer (the first enhanced layer).
[0034] The difference from the method for obtaining the reconstructed block of the base layer is that: the encoding block can calculate the difference between the corresponding pixel points of the image block to be processed and the reconstructed block of the third layer of the image block to be processed to obtain the residual block of the first enhanced layer. The third layer is a layer lower than the first enhanced layer, which can be the base layer or an enhanced layer. Then, the residual block of the first enhanced layer is transformed and quantized to obtain the quantization coefficients of the residual block of the first enhanced layer, and then the quantization coefficients of the residual block of the first enhanced layer are inverse quantized and inverse transformed to obtain the reconstructed residual block of the first enhanced layer. Finally, the corresponding pixel points of the reconstructed block of the third layer and the reconstructed residual block of the first enhanced layer are summed to obtain the reconstructed block of the first enhanced layer.
[0035] In the embodiments of the present application, the transform block partitioning method of the residual block of the second enhanced layer is different from that of the residual block of the first enhanced layer. The encoder can also use the above RDO processing to determine the transform block partitioning method of the residual block of the second enhanced layer, which will not be elaborated here.
[0036] The present application no longer uses the same TU partitioning method for the residual block of the second enhanced layer as that for the residual block of the first enhanced layer and the residual block of the base layer. Instead, an adaptive TU partitioning method is adopted to obtain the TU partitioning method suitable for the residual block of the second enhanced layer, and then encoding is performed. When the TU partitioning of the residual block of the second enhanced layer is independent of the CU or TU partitioning methods of other layers, the TU size of the second enhanced layer is no longer limited by the CU size, or the TU size of the second enhanced layer is no longer limited by the TU size of the first enhanced layer, which can improve the flexibility of encoding.
[0037] The present application can directly calculate the difference between the corresponding pixel points in the to-be-coded image block and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer, reducing the process of obtaining the prediction block of the second enhancement layer, which can reduce the processing flow of the encoder and improve the encoding efficiency of the encoder. Additionally, for the residual block of the second enhancement layer, instead of using the same TU partitioning method as the residual block of the first enhancement layer, an adaptive TU partitioning method is adopted to obtain the TU partitioning method suitable for the residual block of the second enhancement layer, and then encoding is performed. When the TU partitioning of the residual block of the second enhancement layer is independent of the CU or TU partitioning method of the first enhancement layer, the TU size of the second enhancement layer is no longer limited by the CU size, or the TU size of the second enhancement layer is no longer limited by the TU size of the first enhancement layer, which can more effectively improve the compression efficiency of the residual block.
[0038] The above method can be applied to two hierarchical methods: spatial domain hierarchical and quality domain hierarchical. When the image block is hierarchically processed in the spatial domain to obtain the first enhancement layer image block and the second enhancement layer image block, the resolution of the first enhancement layer image block is lower than that of the second enhancement layer image block. Therefore, before calculating the difference between the corresponding pixel points in the to-be-coded image block and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the to-be-coded image block, the original to-be-coded image block can be downsampled to obtain the to-be-coded image block with the first resolution, and the reconstructed block of the original first enhancement layer can be upsampled to obtain the reconstructed block of the first enhancement layer with the first resolution. That is, the encoder can downsample the original to-be-coded image block and upsample the reconstructed block of the original first enhancement layer respectively, so that the resolution of the to-be-coded image block is the same as that of the reconstructed block of the first enhancement layer.
[0039] In a third aspect, the present application provides an encoding device, including: an acquisition module, configured to acquire the reconstructed block of the base layer of the to-be-coded image block; calculate the difference between the corresponding pixel points in the to-be-coded image block and the reconstructed block of the base layer to obtain the residual block of the enhancement layer of the to-be-coded image block, where the resolution of the enhancement layer image block is not lower than that of the base layer image block, or the encoding quality of the enhancement layer image block is not lower than that of the base layer image block; a determination module, configured to determine the transform block partitioning method of the residual block of the enhancement layer, where the transform block partitioning method of the residual block of the enhancement layer is different from that of the residual block of the base layer; and an encoding module, configured to perform transformation on the residual block of the enhancement layer according to the transform block partitioning method to obtain the bitstream of the residual block of the enhancement layer.
[0040] In a possible implementation, the determining module is specifically configured to perform iterative tree-structured partitioning on the first largest transformation unit (LTU) of the residual block of the enhancement layer to obtain multiple transformation units (TUs) with different split depths, where the maximum split depth among the multiple split depths is equal to Dmax, and Dmax is a positive integer; determine the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, where the split depth of the first TU is i, and the depth of the second TU is i + 1, and 0 ≤ i ≤ Dmax - 1.
[0041] In a possible implementation, the determining module is specifically configured to perform transformation and quantization on the TU to obtain the quantization coefficient of the TU, where the TU is the first TU or the second TU; perform precoding on the quantization coefficient of the TU to obtain the codeword length of the TU; perform inverse quantization and inverse transformation on the quantization coefficient of the TU to obtain the reconstructed block of the TU; calculate the sum of squared differences (SSD) between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; obtain the loss estimation value of the TU according to the codeword length and the distortion value of the TU; determine the smaller value between the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; and determine the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
[0042] In a possible implementation, the size of the LTU is the same as the size of the reconstructed block of the base layer; or, the size of the LTU is the same as the size of the reconstructed block of the original base layer after upsampling.
[0043] In a possible implementation, the image block to be encoded refers to the largest coding unit (LCU) in the entire frame of the image; or, the image block to be encoded refers to the entire frame of the image; or, the image block to be encoded refers to the region of interest (ROI) in the entire frame of the image.
[0044] In a possible implementation, the base layer and the enhancement layer are obtained by layer splitting based on resolution or coding quality.
[0045] In a possible implementation, when the base layer and the enhancement layer are obtained by layer splitting based on resolution, the obtaining module is further configured to perform downsampling on the original image block to be encoded to obtain the image block to be encoded with the first resolution; and perform upsampling on the reconstructed block of the original base layer to obtain the reconstructed block of the base layer with the first resolution.
[0046] In a possible implementation, the obtaining module is specifically configured to calculate the difference between corresponding pixel points in the to-be-processed image block and the prediction block of the to-be-processed image block to obtain the residual block of the to-be-processed image block; perform transformation and quantization on the residual block of the to-be-processed image block to obtain the quantization coefficients of the to-be-processed image block; perform inverse quantization and inverse transformation on the quantization coefficients of the to-be-processed image block to obtain the reconstructed residual block of the to-be-processed image block; and calculate the sum of corresponding pixel points in the prediction block of the to-be-processed image block and the reconstructed residual block of the to-be-processed image block to obtain the reconstructed block of the base layer.
[0047] In a fourth aspect, the present application provides an encoding device, including: an obtaining module, configured to obtain the reconstructed block of the first enhancement layer of the to-be-encoded image block; calculate the difference between corresponding pixel points in the to-be-encoded image block and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the to-be-encoded image block, where the resolution of the second enhancement layer image block is not lower than that of the first enhancement layer image block, or the encoding quality of the second enhancement layer image block is not lower than that of the first enhancement layer image block; a determining module, configured to determine the transform block partitioning method of the residual block of the second enhancement layer, where the transform block partitioning method of the residual block of the second enhancement layer is different from that of the residual block of the first layer; and an encoding module, configured to perform transformation on the residual block of the second enhancement layer according to the transform block partitioning method to obtain the bitstream of the residual block of the second enhancement layer.
[0048] In a possible implementation, the determining module is specifically configured to perform iterative tree-structured partitioning on the first largest transform unit (LTU) of the residual block of the second enhancement layer to obtain transform units (TUs) with multiple segmentation depths, where the maximum segmentation depth among the multiple segmentation depths is equal to Dmax, and Dmax is a positive integer; and determine the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, where the segmentation depth of the first TU is i, and the depth of the second TU is i + 1, and 0 ≤ i ≤ Dmax - 1.
[0049] In a possible implementation, the determining module is specifically configured to perform transformation and quantization on the TU to obtain quantization coefficients of the TU, where the TU is the first TU or the second TU; perform precoding on the quantization coefficients of the TU to obtain the codeword length of the TU; perform inverse quantization and inverse transformation on the quantization coefficients of the TU to obtain the reconstructed block of the TU; calculate the sum of squared differences (SSD) between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; obtain the loss estimation value of the TU according to the codeword length and the distortion value of the TU; determine the smaller value between the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; and determine the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
[0050] In a possible implementation, the size of the LTU is the same as the size of the reconstructed block of the first enhancement layer; or, the size of the LTU is the same as the size of the reconstructed block of the original first enhancement layer after upsampling.
[0051] In a possible implementation, the image block to be encoded refers to the largest coding unit (LCU) in the entire frame of the image; or, the image block to be encoded refers to the entire frame of the image; or, the image block to be encoded refers to the region of interest (ROI) in the entire frame of the image.
[0052] In a possible implementation, the first enhancement layer and the second enhancement layer are obtained by layer splitting in the quality domain or in the spatial domain.
[0053] In a possible implementation, when the first enhancement layer and the second enhancement layer are obtained by layer splitting in the spatial domain, the obtaining module is further configured to downsample the original image block to be encoded to obtain the image block to be encoded at a second resolution; and upsample the reconstructed block of the original first enhancement layer to obtain the reconstructed block of the first enhancement layer at the second resolution.
[0054] In a possible implementation, the obtaining module is specifically configured to calculate the difference between corresponding pixel points of the image block to be processed and the reconstructed block of the third layer of the image block to be processed to obtain the residual block of the first enhancement layer, where the resolution of the image block of the first enhancement layer is not lower than the resolution of the image block of the third layer, or the coding quality of the image block of the first enhancement layer is not lower than the coding quality of the image block of the third layer; perform transformation and quantization on the residual block of the first enhancement layer to obtain the quantization coefficients of the residual block of the first enhancement layer; perform inverse quantization and inverse transformation on the quantization coefficients of the residual block of the first enhancement layer to obtain the reconstructed residual block of the first enhancement layer; and calculate the sum of corresponding pixel points of the reconstructed block of the third layer and the reconstructed residual block of the first enhancement layer to obtain the reconstructed block of the first enhancement layer.
[0055] Fifth aspect, the present application provides a decoding device, including: an acquisition module, configured to acquire a bitstream of a residual block of an enhancement layer of a to-be-decoded image block; a decoding module, configured to perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the enhancement layer to obtain a reconstructed residual block of the enhancement layer; a reconstruction module, configured to acquire a reconstructed block of a base layer of the to-be-decoded image block, wherein the resolution of the enhancement layer image block is not lower than that of the base layer image block, or the coding quality of the enhancement layer image block is not lower than that of the base layer image block; and sum corresponding pixel points in the reconstructed residual block of the enhancement layer and the reconstructed block of the base layer to obtain a reconstructed block of the enhancement layer.
[0056] In a possible implementation manner, the to-be-decoded image block refers to the largest coding unit (LCU) in a whole frame of image; or, the to-be-decoded image block refers to a whole frame of image; or, the to-be-decoded image block refers to a region of interest (ROI) in a whole frame of image.
[0057] In a possible implementation manner, the base layer and the enhancement layer are obtained by layer splitting in a quality domain or in a spatial domain.
[0058] In a possible implementation manner, when the base layer and the enhancement layer are obtained by layer splitting in a spatial domain, the reconstruction module is further configured to perform upsampling on the reconstructed block of the original base layer to obtain a reconstructed block of the base layer with a third resolution, where the third resolution is the same as the resolution of the reconstructed residual block of the enhancement layer.
[0059] In a possible implementation manner, the reconstruction module is specifically configured to acquire a bitstream of a residual block of the base layer; perform entropy decoding on the bitstream of the residual block of the base layer to obtain decoded data of the residual block of the base layer; perform inverse quantization and inverse transformation on the decoded data of the residual block of the base layer to obtain a reconstructed residual block of the base layer; acquire a predicted block of the base layer according to the decoded data of the residual block of the base layer; and sum corresponding pixel points in the reconstructed residual block of the base layer and the predicted block of the base layer to obtain a reconstructed block of the base layer.
[0060] Sixth aspect, the present application provides a decoding device, including: an acquisition module, configured to acquire a bitstream of a residual block of a second enhancement layer of a to-be-decoded image block; a decoding module, configured to perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the second enhancement layer to obtain a reconstructed residual block of the second enhancement layer; a reconstruction module, configured to acquire a reconstructed block of a first enhancement layer of the to-be-decoded image block, where the resolution of the second enhancement layer image block is not lower than that of the first enhancement layer image block, or the coding quality of the second enhancement layer image block is not lower than that of the first enhancement layer image block; and sum corresponding pixel points in the reconstructed residual block of the second enhancement layer and the reconstructed block of the first enhancement layer to obtain the reconstructed block of the second enhancement layer.
[0061] In a possible implementation manner, the to-be-decoded image block refers to the largest coding unit LCU in a whole frame of image; or, the to-be-decoded image block refers to a whole frame of image; or, the to-be-decoded image block refers to a region of interest ROI in a whole frame of image.
[0062] In a possible implementation manner, the first enhancement layer and the second enhancement layer are obtained by layer division in a quality domain or in a spatial domain.
[0063] In a possible implementation manner, when the first enhancement layer and the second enhancement layer are obtained by layer division in a spatial domain, the reconstruction module is further configured to perform upsampling on the reconstructed block of the original first enhancement layer to obtain the reconstructed block of the first enhancement layer with a fourth resolution, where the fourth resolution is the same as the resolution of the reconstructed residual block of the second enhancement layer.
[0064] In a possible implementation manner, the reconstruction module is specifically configured to acquire a bitstream of a residual block of the first enhancement layer; perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the first enhancement layer to obtain a reconstructed residual block of the first enhancement layer; and sum corresponding pixel points in the reconstructed residual block of the first enhancement layer and the reconstructed block of a third layer to obtain the reconstructed block of the first enhancement layer, where the resolution of the first enhancement layer image block is not lower than that of the third layer image block, or the coding quality of the first enhancement layer image block is not lower than that of the third layer image block.
[0065] Seventh aspect, the present application provides an encoder, including: a processor and a memory; the processor is coupled to the memory, and computer-readable instructions are stored in the memory; the processor is configured to read the computer-readable instructions to enable the encoder to implement the method according to any one of the first to second aspects described above.
[0066] In an eighth aspect, the present application provides a decoder, comprising: a processor and a memory; the processor is coupled to the memory, and computer-readable instructions are stored in the memory; the processor is configured to read the computer-readable instructions to cause the encoder to implement the method according to any one of the first to second aspects described above.
[0067] In a ninth aspect, the present application provides a computer program product, comprising program code that, when executed on a computer or a processor, is configured to execute the method according to any one of the first to second aspects described above.
[0068] In a tenth aspect, the present application provides a computer-readable storage medium, characterized by comprising program code that, when executed by a computer device, is configured to execute the method according to any one of the first to second aspects described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1A It is an exemplary block diagram of a decoding system 10 according to an embodiment of the present application;
[0070] Figure 1B It is an exemplary block diagram of a video decoding system 40 according to an embodiment of the present application;
[0071] Figure 2 It is an exemplary block diagram of a video encoder 20 according to an embodiment of the present application;
[0072] Figure 3 It is an exemplary block diagram of a video decoder 30 according to an embodiment of the present application;
[0073] Figure 4 It is an exemplary block diagram of a video decoding device 400 according to an embodiment of the present application;
[0074] Figure 5 It is an exemplary block diagram of a device 500 according to an embodiment of the present application;
[0075] Figures 6a to 6g It is several exemplary schematic diagrams of a partitioning method according to an embodiment of the present application;
[0076] Figure 7 It is an exemplary schematic diagram of a QT-MTT partitioning method according to the present application;
[0077] Figure 8 It is an exemplary flowchart of an enhancement layer encoding method according to the present application;
[0078] Figure 9 It is an exemplary hierarchical schematic diagram of scalable video coding according to the present application;
[0079] Figure 10 It is an exemplary flowchart of an encoding method for an enhancement layer according to the present application;
[0080] Figure 11 An exemplary flowchart of the encoding method for the enhancement layer of the present application;
[0081] Figure 12a and 12b An exemplary flowchart of the encoding method for the enhancement layer of the present application;
[0082] Figure 13a and 13b An exemplary flowchart of the encoding method for the enhancement layer of the present application;
[0083] Figure 14a and 14b An exemplary flowchart of the encoding method for the enhancement layer of the present application;
[0084] Figure 15 An exemplary flowchart of the encoding method for the enhancement layer of the present application;
[0085] Figure 16 An exemplary structural schematic diagram of the encoding apparatus of the present application;
[0086] Figure 17 An exemplary structural schematic diagram of the decoding apparatus of the present application. Detailed implementation manners
[0087] The embodiments of the present application provide a video image compression technology, and specifically provide an encoding and decoding architecture based on hierarchical residual encoding to improve the traditional hybrid video encoding and decoding system.
[0088] Video encoding generally refers to processing an image sequence that forms a video or a video sequence. In the field of video encoding, the terms "picture", "frame", or "image" can be used as synonyms. Video encoding (or generally referred to as encoding) includes two parts: video encoding and video decoding. Video encoding is performed on the source side and generally includes processing (e.g., compressing) the original video image to reduce the amount of data required to represent the video image (thus enabling more efficient storage and / or transmission). Video decoding is performed on the destination side and generally includes performing inverse processing relative to the encoder to reconstruct the video image. The "encoding" of the video image (or generally referred to as an image) involved in the embodiments should be understood as the "encoding" or "decoding" of the video image or video sequence. The encoding part and the decoding part are also collectively referred to as encoding and decoding (encoding and decoding, CODEC).
[0089] In the case of lossless video coding, the original video image can be reconstructed, that is, the reconstructed video image has the same quality as the original video image (assuming no transmission loss or other data loss during storage or transmission). In the case of lossy video coding, further compression is performed through quantization, etc., to reduce the amount of data required to represent the video image, and the decoder side cannot fully reconstruct the video image, that is, the quality of the reconstructed video image is lower or worse than that of the original video image.
[0090] Several video coding standards belong to "lossy hybrid video coding and decoding" (that is, combining spatial and temporal prediction in the pixel domain with 2D transform coding for applying quantization in the transform domain). Each image in a video sequence is usually divided into a set of non-overlapping blocks, and encoding is usually performed at the block level. In other words, the encoder usually processes and encodes the video at the block (video block) level. For example, prediction blocks are generated through spatial (intra-frame) prediction and temporal (inter-frame) prediction; the prediction blocks are subtracted from the current block (the currently processed / block to be processed) to obtain a residual block; the residual block is transformed and quantized in the transform domain to reduce the amount of data to be transmitted (compressed), and the decoder side applies the inverse processing part relative to the encoder to the encoded or compressed block to reconstruct the current block for representation. Additionally, the encoder needs to repeat the processing steps of the decoder so that the encoder and the decoder generate the same predictions (such as intra-frame prediction and inter-frame prediction) and / or reconstruct pixels for processing, that is, for encoding subsequent blocks.
[0091] In the following embodiments of the decoding system 10, the encoder 20 and the decoder 30 are described according to Figures 1A to 3 this.
[0092] Figure 1A FIG. is an exemplary block diagram of the decoding system 10 according to an embodiment of the present application. For example, a video decoding system 10 (or simply referred to as the decoding system 10) that can utilize the technology of the present application. The video encoder 20 (or simply referred to as the encoder 20) and the video decoder 30 (or simply referred to as the decoder 30) in the video decoding system 10 represent devices that can be used to execute various techniques according to the various examples described in the present application.
[0093] As Figure 1A shown, the decoding system 10 includes a source device 12, and the source device 12 is used to provide encoded image data 21 such as encoded images to a destination device 14 for decoding the encoded image data 21.
[0094] The source device 12 includes an encoder 20, and additionally, optionally, may include an image source 16, a pre-processor (or pre-processing unit) 18 such as an image pre-processor, and a communication interface (or communication unit) 22.
[0095] The image source 16 may include or be any type of image capture device for capturing real-world images, etc., and / or any type of image generation device, such as a computer graphics processor for generating computer animation images or any type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images, and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory or storage that stores any of the above images.
[0096] To distinguish the processing performed by the preprocessor (or preprocessing unit) 18, the image (or image data) 17 may also be referred to as the raw image (or raw image data) 17.
[0097] The preprocessor 18 is used to receive the raw image data 17 and preprocess the raw image data 17 to obtain preprocessed image (or preprocessed image data) 19. For example, the preprocessing performed by the preprocessor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It can be understood that the preprocessing unit 18 may be an optional component.
[0098] The video encoder (or encoder) 20 is used to receive the preprocessed image data 19 and provide encoded image data 21 (which will be further described below according to Figure 2 etc.).
[0099] The communication interface 22 in the source device 12 can be used to: receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.
[0100] The destination device 14 includes a decoder 30, and additionally, optionally, may include a communication interface (or communication unit) 28, a postprocessor (or postprocessing unit) 32, and a display device 34.
[0101] The communication interface 28 in the destination device 14 is used to directly receive the encoded image data 21 (or any other processed version) from the source device 12 or from any other source device such as a storage device. For example, the storage device is an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0102] The communication interfaces 22 and 28 can be used to send or receive encoded image data (or encoded data) 21 via a direct communication link between the source device 12 and the destination device 14, such as a direct wired or wireless connection, etc., or via any type of network, such as a wired network, a wireless network, or any combination thereof, any type of private network and public network, or any combination of any type thereof.
[0103] For example, the communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a packet, and / or use any type of transport encoding or processing to process the encoded image data for transmission over the communication link or communication network.
[0104] The communication interface 28 corresponds to the communication interface 22. For example, it can be used to receive the transmitted data and process the transmitted data using any type of corresponding transport decoding or processing and / or de-encapsulation to obtain the encoded image data 21.
[0105] Both the communication interface 22 and the communication interface 28 can be configured as Figure 1A a unidirectional communication interface as indicated by the arrow of the corresponding communication channel 13 pointing from the source device 12 to the destination device 14 in
[0106] or a bidirectional communication interface, and can be used to send and receive messages, etc., to establish a connection, confirm and exchange any other information related to the communication link and / or data transmission such as the transmission of encoded image data, etc. Figure 3 etc. will be further described below.
[0107] The video decoder (or decoder) 30 is used to receive the encoded image data 21 and provide decoded image data (or decoded image data) 31 (which will be further described below according to
[0108] The display device 34 is configured to receive the post-processed image data 33 to display an image to a user, viewer, etc. The display device 34 may be or include any type of display for presenting the reconstructed image, e.g., an integrated or external display screen or monitor. For example, the display screen may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display screen.
[0109] Although Figure 1A the source device 12 and the destination device 14 are shown as separate devices, device embodiments may also include both the source device 12 and the destination device 14 or the functions of both the source device 12 and the destination device 14, i.e., include both the source device 12 or corresponding function and the destination device 14 or corresponding function. In these embodiments, the source device 12 or corresponding function and the destination device 14 or corresponding function may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0110] According to the description, Figure 1A the presence and (exact) partitioning of the different units or functions in the illustrated source device 12 and / or destination device 14 may vary depending on the actual device and application, which will be apparent to those skilled in the art.
[0111] The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30) or both may be implemented by a processing circuit as Figure 1B shown, e.g., one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video encoding dedicated processors, or any combination thereof. The encoder 20 may be implemented by the processing circuit 46 to include the various modules discussed with reference to Figure 2 the encoder 20 and / or any other encoder system or subsystem described herein. The decoder 30 may be implemented by the processing circuit 46 to include the various modules discussed with reference to Figure 3The various modules discussed for decoder 30 and / or any other decoder system or subsystem described herein. The processing circuitry 46 can be used to perform the various operations discussed below. As Figure 5 shown, if part of the technology is implemented in software, the device can store the instructions of the software in a suitable computer-readable storage medium and execute the instructions in hardware using one or more processors to implement the technology of the present invention. One of the video encoder 20 and the video decoder 30 can be integrated as part of a combined codec (encoder / decoder, CODEC) in a single device, as Figure 1B shown.
[0112] The source device 12 and the destination device 14 can include any of a variety of devices, including any type of handheld or fixed device, such as, for example, a laptop or notebook computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (e.g., a content service server or a content distribution server), a broadcast receiving device, a broadcast transmitting device, etc., and may or may not use any type of operating system. In some cases, the source device 12 and the destination device 14 can be equipped with components for wireless communication. Thus, the source device 12 and the destination device 14 can be wireless communication devices.
[0113] In some cases, Figure 1A the video decoding system 10 shown is merely exemplary, and the technology provided in this application can be applied to video coding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory, sent over a network, etc. The video encoding device can encode data and store the data in memory, and / or the video decoding device can retrieve data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but only encode data into memory and / or retrieve and decode data from memory.
[0114] Figure 1B is an exemplary block diagram of a video decoding system 40 according to an embodiment of this application. As Figure 1B shown, the video decoding system 40 can include an imaging device 41, a video encoder 20, a video decoder 30 (and / or a video codec implemented by the processing circuitry 46), an antenna 42, one or more processors 43, one or more memory memories 44, and / or a display device 45.
[0115] As Figure 1BAs shown, the imaging device 41, antenna 42, processing circuit 46, video encoder 20, video decoder 30, processor 43, memory 44, and / or display device 45 can communicate with each other. In different instances, the video decoding system 40 may include only the video encoder 20 or only the video decoder 30.
[0116] In some instances, the antenna 42 can be used to transmit or receive an encoded bitstream of video data. Additionally, in some instances, the display device 45 can be used to present video data. The processing circuit 46 can include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. The video decoding system 40 can also include an optional processor 43, which can similarly include application-specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, and the like. Additionally, the memory 44 can be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting instance, the memory 44 can be implemented by cache memory. In other instances, the processing circuit 46 can include a memory (e.g., a cache, etc.) for implementing an image buffer, etc.
[0117] In some instances, the video encoder 20 implemented by logic circuits can include an image buffer (e.g., implemented by the processing circuit 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing circuit 46). The graphics processing unit can be communicatively coupled to the image buffer. The graphics processing unit can include the video encoder 20 implemented by the processing circuit 46 to implement various modules discussed with reference to Figure 2 and / or any other encoder system or subsystem described herein. The logic circuits can be used to perform various operations discussed herein.
[0118] In some instances, the video decoder 30 can be implemented in a similar manner by the processing circuit 46 to implement with reference to Figure 3The various modules discussed for the video decoder 30 and / or any other decoder system or subsystem described herein. In some examples, the video decoder 30 implemented by logic circuitry may include an image buffer (implemented by the processing circuitry 46 or the memory 44) and a graphics processing unit (e.g., implemented by the processing circuitry 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include the video decoder 30 implemented by the processing circuitry 46 to implement with reference to Figure 3 and / or the various modules discussed for any other decoder system or subsystem described herein.
[0119] In some examples, the antenna 42 may be used to receive an encoded bitstream of video data. As discussed, the encoded bitstream may include data, indicators, index values, mode selection data, etc. discussed herein related to the coded video frames, e.g., data related to coded partitions (e.g., transform coefficients or quantized transform coefficients, optional indicators as discussed, and / or data defining the coded partitions). The video decoding system 40 may also include a video decoder 30 coupled to the antenna 42 and configured to decode the encoded bitstream. A display device 45 is used to present the video frames.
[0120] It should be understood that for the examples described with reference to the video encoder 20 in the embodiments of the present application, the video decoder 30 may be used to perform the reverse process. Regarding signaling syntax elements, the video decoder 30 may be used to receive and parse such syntax elements and accordingly decode the relevant video data. In some examples, the video encoder 20 may entropy encode the syntax elements into an encoded video bitstream. In such examples, the video decoder 30 may parse such syntax elements and accordingly decode the relevant video data.
[0121] For ease of description, embodiments of the present invention are described with reference to the Versatile Video Coding (VVC) reference software or the High-Efficiency Video Coding (HEVC) developed by the Video Coding Experts Group (VCEG) of the ITU-T and the Joint Collaboration Team on Video Coding (JCT-VC) of the ISO / IEC Moving Picture Experts Group (MPEG). Those of ordinary skill in the art understand that the embodiments of the present invention are not limited to HEVC or VVC.
[0122] Encoder and encoding method
[0123] Figure 2Exemplary block diagram of video encoder 20 according to an embodiment of the present application. As Figure 2 shown, video encoder 20 includes an input end (or input interface) 201, a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a loop filter 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output end (or output interface) 272. The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a segmentation unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). Figure 2 The video encoder 20 shown may also be referred to as a hybrid video encoder or a video encoder based on a hybrid video codec.
[0124] The residual calculation unit 204, the transformation processing unit 206, the quantization unit 208, and the mode selection unit 260 form the forward signal path of encoder 20, while the inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 form the backward signal path of the encoder, where the backward signal path of encoder 20 corresponds to the signal path of the decoder (see Figure 3 decoder 30 in). The inverse quantization unit 210, the inverse transformation processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer 230, the inter prediction unit 244, and the intra prediction unit 254 also form the "built-in decoder" of video encoder 20.
[0125] Images and image segmentation (images and blocks)
[0126] Encoder 20 can be used to receive an image (or image data) 17 through the input end 201, etc., for example, an image in an image sequence forming a video or video sequence. The received image or image data can also be a preprocessed image (or preprocessed image data) 19. For simplicity, image 17 is used in the following description. Image 17 can also be referred to as the current image or the image to be encoded (especially when distinguishing the current image from other images in video coding, other images such as previously encoded and / or decoded images in the same video sequence, i.e., the video sequence that also includes the current image).
[0127] (Digital) images are or can be regarded as two-dimensional arrays or matrices composed of pixel points with intensity values. The pixel points in the array can also be called pixels (pixel or pel, short for picture element). The number of pixel points in the horizontal and vertical directions (or axes) of the array or image determines the size and / or resolution of the image. To represent colors, usually three color components are used, that is, the image can be represented as or include three pixel point arrays. In the RGB format or color space, the image includes corresponding red, green, and blue pixel point arrays. However, in video coding, each pixel is usually represented in a luminance / chrominance format or color space, such as YCbCr, including a luminance component indicated by Y (sometimes also represented by L) and two chrominance components represented by Cb and Cr. The luminance component Y represents the luminance or gray-level intensity (for example, they are the same in a grayscale image), while the two chrominance components Cb and Cr represent the chrominance or color information components. Accordingly, an image in the YCbCr format includes a luminance pixel point array of luminance pixel point values (Y) and two chrominance pixel point arrays of chrominance values (Cb and Cr). An image in the RGB format can be converted or transformed into the YCbCr format, and vice versa, and this process is also called color transformation or conversion. If the image is black and white, then the image can only include a luminance pixel point array. Accordingly, the image can be, for example, a luminance pixel point array in a monochrome format or a luminance pixel point array and two corresponding chrominance pixel point arrays in 4:2:0, 4:2:2, and 4:4:4 color formats.
[0128] In one embodiment, an embodiment of the video encoder 20 may include an image segmentation unit ( Figure 2 not shown in the figure) for segmenting the image 17 into a plurality of (usually non-overlapping) image blocks 203. These blocks may also be called root blocks, macroblocks (H.264 / AVC), or coding tree blocks (CTB), or coding tree units (CTU) in the H.265 / HEVC and VVC standards. The segmentation unit can be used to use the same block size for all images in the video sequence and a corresponding grid defining the block size, or to change the block size between images or subsets of images or groups of images, and segment each image into corresponding blocks.
[0129] In other embodiments, the video encoder can be used to directly receive the blocks 203 of the image 17, for example, one, several, or all of the blocks constituting the image 17. The image block 203 can also be called the current image block or the image block to be encoded.
[0130] Similar to Image 17, Image Block 203 is also or can be considered as a two-dimensional array or matrix composed of pixel points with intensity values (pixel values), but Image Block 203 is smaller than Image 17. In other words, Block 203 can include an array of pixel points (e.g., the luminance array in the case of a monochrome image 17 or the luminance array or chrominance arrays in the case of a color image) or three arrays of pixel points (e.g., one luminance array and two chrominance arrays in the case of a color image 17) or any other number and / or type of arrays according to the color format adopted. The number of pixel points in the horizontal and vertical directions (or axes) of Block 203 defines the size of Block 203. Accordingly, the block can be an array of M×N (M columns × N rows) pixel points, or an array of M×N transform coefficients, etc.
[0131] In one embodiment, Figure 2 The illustrated video encoder 20 is used to encode Image 17 block by block. For example, encoding and prediction are performed on each block 203.
[0132] In one embodiment, Figure 2 The illustrated video encoder 20 can also be used to segment and / or encode an image using slices (also called video slices), where the image can be segmented or encoded using one or more slices (usually non-overlapping). Each slice can include one or more blocks (e.g., Coding Tree Unit CTU) or one or more groups of blocks (e.g., coding blocks (tile) in H.265 / HEVC / VVC standards and bricks in VVC standards).
[0133] In one embodiment, Figure 2 The illustrated video encoder 20 can also be used to segment and / or encode an image using slice / coding block group (also called video coding block group) and / or coding block (also called video coding block), where the image can be segmented or encoded using one or more slice / coding block groups (usually non-overlapping), each slice / coding block group can include one or more blocks (e.g., CTU) or one or more coding blocks, etc., where each coding block can be in a shape such as a rectangle and can include one or more complete or partial blocks (e.g., CTU).
[0134] Residual calculation
[0135] The residual calculation unit 204 is used to calculate the residual block 205 based on the image block 203 and the prediction block 265 (the prediction block 265 is introduced in detail later) in the following way: for example, subtract the pixel values of the prediction block 265 from the pixel values of the image block 203 pixel by pixel (pixel by pixel) to obtain the residual block 205 in the pixel domain.
[0136] Transformation
[0137] The transform processing unit 206 is used to perform a discrete cosine transform (DCT), a discrete sine transform (DST), etc. on the pixel values of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 can also be referred to as transform residual coefficients, representing the residual block 205 in the transform domain.
[0138] The transform processing unit 206 can be used to apply an integer approximation of DCT / DST, such as the transform specified for H.265 / HEVC. Compared with the orthogonal DCT transform, this integer approximation is usually scaled by a certain factor. To maintain the norm of the residual block after forward and inverse transform processing, other scaling factors are used as part of the transform process. The scaling factors are usually selected according to certain constraints, such as the power of 2 for shift operations, the bit depth of the transform coefficients, the trade-off between accuracy and implementation cost, etc. For example, at the encoder 20 side, a specific scaling factor is specified for the inverse transform by the inverse transform processing unit 212 (and at the decoder 30 side, a corresponding inverse transform is performed by, for example, the inverse transform processing unit 312), and correspondingly, at the encoder 20 side, a corresponding scaling factor can be specified for the forward transform by the transform processing unit 206.
[0139] In one embodiment, the video encoder 20 (correspondingly, the transform processing unit 206) can be used to output transform parameters such as the type of one or more transforms, for example, directly output or output after being encoded or compressed by the entropy encoding unit 270, such that the video decoder 30 can receive and use the transform parameters for decoding.
[0140] Quantization
[0141] The quantization unit 208 is used to quantize the transform coefficients 207 through, for example, scalar quantization or vector quantization to obtain quantized transform coefficients 209. The quantized transform coefficients 209 can also be referred to as quantized residual coefficients 209.
[0142] The quantization process can reduce the bit depth associated with some or all of the transform coefficients 207. For example, during quantization, an n-bit transform coefficient can be rounded down to an m-bit transform coefficient, where n is greater than m. The degree of quantization can be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different degrees of scaling can be applied to achieve finer or coarser quantization. A smaller quantization step corresponds to finer quantization, while a larger quantization step corresponds to coarser quantization. The appropriate quantization step can be indicated by the quantization parameter (QP). For example, the quantization parameter can be an index of a predefined set of appropriate quantization steps. For example, a smaller quantization parameter can correspond to fine quantization (smaller quantization step), and a larger quantization parameter can correspond to coarse quantization (larger quantization step), and vice versa. Quantization can include dividing by the quantization step, and the corresponding or inverse dequantization performed by the dequantization unit 210 and the like can include multiplying by the quantization step. Embodiments according to some standards such as HEVC can be used to determine the quantization step using the quantization parameter. Generally, the quantization step can be calculated using a fixed-point approximation of an equation containing division according to the quantization parameter. Other scaling factors can be introduced for quantization and dequantization to recover the norm of the residual block that may have been modified due to the scaling used in the fixed-point approximation of the equations for the quantization step and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization can be combined. Alternatively, a custom quantization table can be used and indicated from the encoder to the decoder in the bitstream. Quantization is a lossy operation, and the larger the quantization step, the greater the loss.
[0143] In one embodiment, the video encoder 20 (correspondingly, the quantization unit 208) can be used to output the quantization parameter (QP), for example, directly output or output after being encoded or compressed by the entropy encoding unit 270, such that the video decoder 30 can receive and use the quantization parameter for decoding.
[0144] Dequantization
[0145] The dequantization unit 210 is used to perform the inverse quantization of the quantization unit 208 on the quantization coefficients to obtain the dequantized coefficients 211. For example, the inverse quantization scheme of the quantization scheme performed by the quantization unit 208 is performed according to or using the same quantization step as the quantization unit 208. The dequantized coefficients 211 can also be referred to as dequantized residual coefficients 211, corresponding to the transform coefficients 207. However, due to the loss caused by quantization, the dequantized coefficients 211 are usually not exactly the same as the transform coefficients.
[0146] Inverse transform
[0147] The inverse transformation processing unit 212 is used to perform the inverse transformation of the transformation performed by the transformation processing unit 206, for example, inverse discrete cosine transform (DCT) or inverse discrete sine transform (DST), to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the pixel domain. The reconstructed residual block 213 may also be referred to as the transformed block 213.
[0148] Reconstruction
[0149] The reconstruction unit 214 (e.g., adder 214) is used to add the transformed block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain a reconstructed block 215 in the pixel domain, for example, adding the pixel point values of the reconstructed residual block 213 and the pixel point values of the prediction block 265.
[0150] Filtering
[0151] The loop filter unit 220 (or simply referred to as "loop filter" 220) is used to filter the reconstructed block 215 to obtain a filtered block 221, or is generally used to filter the reconstructed pixel points to obtain filtered pixel point values. For example, the loop filter unit is used to smoothly perform pixel transitions or improve video quality. The loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in-loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub-block edges, such as affine sub-block edges, ATMVP sub-block edges, sub-block transform (SBT) edges, and intra sub-partition (ISP) edges. Although the loop filter unit 220 is shown as a loop filter in Figure 2 it may be implemented as a post-loop filter in other configurations. The filtered block 221 may also be referred to as the filtered reconstructed block 221.
[0152] In one embodiment, the video encoder 20 (correspondingly, the loop filter unit 220) may be used to output loop filter parameters (such as SAO filter parameters, ALF filter parameters, or LMCS parameters), for example, directly output or output after being entropy encoded by the entropy encoding unit 270, such that the decoder 30 can receive and use the same or different loop filter parameters for decoding.
[0153] Decoded picture buffer
[0154] The decoded picture buffer (DPB) 230 may be a reference image memory that stores reference image data for use by the video encoder 20 when encoding video data. The DPB 230 may be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of storage devices. The decoded picture buffer 230 may be used to store one or more filtered blocks 221. The decoded picture buffer 230 may also be used to store other previously filtered blocks, such as previously reconstructed and filtered blocks 221, of the same current image or different images such as previous reconstructed blocks, and may provide the complete previously reconstructed, i.e., decoded, image (and corresponding reference blocks and pixels) and / or a partially reconstructed current image (and corresponding reference blocks and pixels), for example, for inter-frame prediction. The decoded picture buffer 230 may also be used to store one or more unfiltered reconstructed blocks 215, or generally store unfiltered reconstructed pixels, for example, reconstructed blocks 215 that have not been filtered by the loop filter unit 220, or reconstructed blocks or reconstructed pixels that have not undergone any other processing.
[0155] Mode selection (segmentation and prediction)
[0156] The mode selection unit 260 includes a segmentation unit 262, an inter-frame prediction unit 244, and an intra-frame prediction unit 254, and is used to receive or obtain original image data such as the original block 203 (the current block 203 of the current image 17) and reconstructed block data from the decoded picture buffer 230 or other buffers (such as a column buffer, not shown in the figure), for example, filtered and / or unfiltered reconstructed pixels or reconstructed blocks of the same (current) image and / or one or more previously decoded images. The reconstructed block data is used as reference image data required for prediction such as inter-frame prediction or intra-frame prediction to obtain a predicted block 265 or a predicted value 265.
[0157] The mode selection unit 260 can be used to determine or select a segmentation for the prediction mode (e.g., intra-frame or inter-frame prediction mode) of the current block (including non-segmentation), generate a corresponding prediction block 265, to calculate the residual block 205 and reconstruct the reconstructed block 215.
[0158] In one embodiment, the mode selection unit 260 can be used to select a segmentation and a prediction mode (e.g., from the prediction modes supported or available by the mode selection unit 260), where the prediction mode provides the best match or the smallest residual (the smallest residual means better compression in transmission or storage), or provides the smallest signaling overhead (the smallest signaling overhead means better compression in transmission or storage), or considers or balances both of the above at the same time. The mode selection unit 260 can be used to determine the segmentation and the prediction mode according to rate distortion Optimization (RDO), that is, select the prediction mode that provides the smallest rate distortion optimization. The terms "best", "lowest", "optimal", etc. in this article do not necessarily refer to "the best", "the lowest", "the optimal" in general, but can also refer to the situation that meets the termination or selection criteria. For example, values exceeding or below a threshold or other limitations may result in a "sub-optimal choice", but will reduce the complexity and processing time.
[0159] In other words, the segmentation unit 262 can be used to segment the images in the video sequence into a sequence of coding tree units (CTUs). The CTU 203 can be further segmented into smaller block parts or sub-blocks (forming blocks again). For example, by iteratively using quadtree (QT) partitioning, binary tree (BT) partitioning, or triple-tree (TT) partitioning or any combination thereof, and used to perform prediction on each of the block parts or sub-blocks. Wherein the mode selection includes selecting the tree structure for segmenting the block 203 and selecting the prediction mode applied to each of the block parts or sub-blocks.
[0160] The segmentation (e.g., performed by the segmentation unit 262) and the prediction processing (e.g., performed by the inter-frame prediction unit 244 and the intra-frame prediction unit 254) performed by the video encoder 20 will be described in detail below.
[0161] Segmentation
[0162] The segmentation unit 262 can segment (or divide) a coding tree unit 203 into smaller parts, such as small blocks in the shape of a square or a rectangle. For an image with an array of three pixel points, a CTU consists of an N×N block of luminance pixel points and two corresponding chrominance pixel point blocks.
[0163] The H.265 / HEVC video coding standard divides a frame of image into non-overlapping CTUs. The size of a CTU can be set to 64×64 (the size of a CTU can also be set to other values, such as in the JVET reference software JEM, the CTU size is increased to 128×128 or 256×256). A 64×64 CTU contains a rectangular pixel lattice with 64 columns and 64 pixels in each column. Each pixel contains a luminance component and / or a chrominance component.
[0164] H.265 uses a QT-based CTU partitioning method. The CTU is used as the root node of the QT, and according to the partitioning method of the QT, the CTU is recursively partitioned into several leaf nodes, as shown in Figure 6a the following. A node corresponds to an image region. If a node is not partitioned, it is called a leaf node, and the corresponding image region is a CU. If a node is further partitioned, the corresponding image region of the node can be divided into four regions of the same size (its length and width are each half of the divided region), and each region corresponds to a node. It is necessary to determine whether these nodes will be further partitioned respectively. Whether a node is partitioned is indicated by the partitioning flag bit split_cu_flag corresponding to the node in the bitstream. When a node A is partitioned once, 4 nodes Bi are obtained, where i = 0 to 3, and Bi is called the child node of A, and A is called the parent node of Bi. The QT level (qtDepth) of the root node is 0, and the QT level of a node is the QT level of its parent node plus 1. For the sake of concise expression, in the following text, the size and shape of a node refer to the size and shape of the image region corresponding to the node.
[0165] For example, for a 64×64 CTU (its QT level is 0), according to its corresponding split_cu_flag, if it is determined not to be partitioned, it becomes a 64×64 CU; if it is determined to be partitioned into 4 32×32 nodes (QT level is 1). Each of the 4 32×32 nodes can be determined to be further partitioned or not according to its corresponding split_cu_flag. If one of the 32×32 nodes is further partitioned, 4 16×16 nodes (QT level is 2) are generated. And so on, until all nodes are no longer partitioned. In this way, a CTU is partitioned into a group of CUs. The minimum size of a CU is identified in the sequence parameter set (SPS). For example, 8×8 is the minimum size of a CU. In the above recursive partitioning process, if the size of a node is equal to the minimum CU size, this node is defaulted to not be partitioned, and at the same time, it is not necessary to include its partitioning flag bit in the bitstream.
[0166] After parsing a node as a leaf node, the leaf node is a CU, and the encoding information corresponding to the CU is further parsed (including information such as the prediction mode and transform coefficients of the CU, for example, the coding_unit() syntax structure in H.265), and then decoding processes such as prediction, inverse quantization, inverse transform, and loop filtering are performed on the CU according to this encoding information to obtain the reconstructed block corresponding to the CU. QT partitioning enables the CTU to be partitioned into a set of CUs of appropriate sizes according to the local characteristics of the image. For example, smooth regions are partitioned into larger-sized CUs, while texture-rich regions are partitioned into smaller-sized CUs.
[0167] Based on QT partitioning, the H.266 / VVC standard adds BT partitioning and TT partitioning methods.
[0168] The BT partitioning method divides 1 node into 2 sub-nodes. There are two specific BT partitioning methods:
[0169] 1) Horizontal bisection: The region corresponding to the node is divided into two regions of the same size, upper and lower, that is, the width remains unchanged, and the height becomes half of the region before partitioning. Each region after partitioning corresponds to a sub-node. For example, as Figure 6b shown.
[0170] 2) Vertical bisection: The region corresponding to the node is divided into two regions of the same size, left and right, that is, the height remains unchanged, and the width becomes half of the region before partitioning. Each region after partitioning corresponds to a sub-node. For example, as Figure 6c shown.
[0171] The TT partitioning method divides 1 node into 3 sub-nodes. There are two specific TT partitioning methods:
[0172] 1) Horizontal trisection: The region corresponding to the node is divided into three regions, upper, middle, and lower. The heights of the upper, middle, and lower regions are 1 / 4, 1 / 2, and 1 / 4 of the region before partitioning respectively. Each region after partitioning corresponds to a sub-node. For example, as Figure 6d shown.
[0173] 2) Vertical trisection: The region corresponding to the node is divided into three regions, left, middle, and right. The widths of the left, middle, and right regions are 1 / 4, 1 / 2, and 1 / 4 of the region before partitioning respectively. Each region after partitioning corresponds to a sub-node. For example, as Figure 6e shown.
[0174] In H.266, the QT cascaded BT / TT partitioning method is used, which is simply referred to as the QT-MTT partitioning method. That is, the CTU is partitioned by QT to generate 4 QT nodes. The QT nodes can continue to be partitioned into 4 QT nodes using the QT partitioning method until the QT nodes no longer use the QT partitioning method for further partitioning and serve as QT leaf nodes, or the QT nodes do not use the QT partitioning method for partitioning and serve as QT leaf nodes. The QT leaf nodes serve as the root nodes of the MTT. The nodes in the MTT can be partitioned into child nodes using one of the four partitioning methods: horizontal bisection, vertical bisection, horizontal trisection, and vertical trisection, or no longer be partitioned and serve as 1 MTT leaf node. The MTT leaf node is a CU.
[0175] For example, Figure 7 is an exemplary schematic diagram of the QT-MTT partitioning method of this application. As Figure 7 shown, using QT-MTT to partition a CTU into 16 CUs from a to p. Each endpoint in the tree diagram represents a node. 4 lines connected to 1 node represent QT partitioning, 2 lines connected to 1 node represent BT partitioning, and 3 lines connected to 1 node represent TT partitioning. The solid line represents QT partitioning, the dashed line represents the first layer of MTT partitioning, and the dotted line represents the second layer of MTT partitioning. a to p are 16 MTT leaf nodes, and each MTT leaf node corresponds to a CU. Based on the above tree diagram, the partitioning diagram of the CTU in Figure 7 can be obtained.
[0176] In the QT-MTT partitioning method, each CU has a QT level (or QT depth) (quad-tree depth, QT depth) and an MTT level (MTT depth) (multi-type tree depth, MTT depth). The QT level represents the QT level of the QT leaf node to which the CU belongs, and the MTT level represents the MTT level of the MTT leaf node to which the CU belongs. For example, Figure 7 in, the QT levels of a, b, c, d, e, f, g, i, j are 1, and the MTT levels are 2; the QT level of h is 1, and the MTT level is 1; the QT levels of n, o, p are 2, and the MTT levels are 0; the QT levels of l, m are 2, and the MTT levels are 1. If the CTU is only partitioned into one CU, then the QT level of this CU is 0, and the MTT level is 0.
[0177] The AVS3 standard adopts another QT-MTT partitioning method, namely the partitioning method of QT cascading BT / EQT. That is, the AVS3 standard replaces the TT partitioning in H.266 with the extended quad-tree (EQT) partitioning. Specifically, through QT partitioning, a CTU generates 4 QT nodes. A QT node can continue to be partitioned into 4 QT nodes using the QT partitioning method until the QT node no longer uses the QT partitioning method for further partitioning and becomes a QT leaf node, or the QT node does not use the QT partitioning method for partitioning and becomes a QT leaf node. The QT leaf node serves as the root node of the MTT. Nodes in the MTT can be partitioned into child nodes using one of the four partitioning methods: horizontal bisection, vertical bisection, horizontal quartering, and vertical quartering, or no longer be partitioned and become 1 MTT leaf node. The MTT leaf node is a CU.
[0178] The EQT partitioning method divides 1 node into 4 child nodes. There are two specific EQT partitioning methods:
[0179] 1) Horizontal quartering: The area corresponding to the node is divided into four areas: upper, upper-middle left, upper-middle right, and lower. The heights of the upper, upper-middle left, upper-middle right, and lower areas are 1 / 4, 1 / 2, 1 / 2, and 1 / 4 of the area before partitioning respectively. The widths of the upper-middle left and upper-middle right are both 1 / 2 of the area before partitioning. Each area after partitioning corresponds to a child node. For example, as Figure 6f shown.
[0180] 2) Vertical quartering: The area corresponding to the node is divided into four areas: left, upper-middle, lower-middle, and right. The widths of the left, upper-middle, lower-middle, and right areas are 1 / 4, 1 / 2, 1 / 2, and 1 / 4 of the area before partitioning respectively. The heights of the upper-middle and lower-middle are both 1 / 2 of the area before partitioning. Each area after partitioning corresponds to a child node. For example, as Figure 6g shown.
[0181] In the H.265 / HEVC standard, for an image in YUV4:2:0 format, a CTU contains one luma block and two chroma blocks. The luma block and chroma blocks can be partitioned in the same way, which is called the joint luma-chroma coding tree. In VVC, if the current frame is an I-frame, when a CTU is a node of a preset size (such as 64×64) in an intra-coded frame (I-frame), the luma block contained in this node is partitioned into a set of coding units containing only luma blocks through the luma coding tree, and the chroma blocks contained in this node are partitioned into a set of coding units containing only chroma blocks through the chroma coding tree; the partitioning of the luma coding tree and the chroma coding tree are independent of each other. This use of independent coding trees for luma blocks and chroma blocks is called separate trees. In H.265, a CU contains luma pixels and chroma pixels; in standards such as H.266 and AVS3, in addition to CUs that contain both luma pixels and chroma pixels, there are also luma CUs that contain only luma pixels and chroma CUs that contain only chroma pixels.
[0182] As described above, the video encoder 20 is used to determine or select the best or optimal prediction mode from a (predetermined) set of prediction modes. The set of prediction modes may include, for example, intra prediction modes and / or inter prediction modes.
[0183] Intra prediction
[0184] The set of intra prediction modes may include 35 different intra prediction modes. For example, non-directional modes such as the DC (or mean) mode and the planar mode, or directional modes defined as in HEVC. Or it may include 67 different intra prediction modes. For example, non-directional modes such as the DC (or mean) mode and the planar mode, or directional modes defined in VVC. For example, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks defined in VVC. Also, for example, to avoid the division operation in DC prediction, only the longer side is used to calculate the average value of non-square blocks. And the intra prediction result of the planar mode can also be modified using the position-dependent intra prediction combination (PDPC) method.
[0185] The intra prediction unit 254 is used to generate an intra prediction block 265 using the reconstructed pixel points of adjacent blocks of the same current image according to the intra prediction mode in the set of intra prediction modes.
[0186] The intra prediction unit 254 (or generally the mode selection unit 260) is also used to output intra prediction parameters (or generally information indicating the selected intra prediction mode of the block) in the form of syntax elements 266 to the entropy coding unit 270 to be included in the encoded image data 21 so that the video decoder 30 can perform operations such as receiving and using the prediction parameters for decoding.
[0187] Inter - frame prediction
[0188] In a possible implementation, the set of inter - frame prediction modes depends on the available reference images (i.e., for example, at least part of the previously decoded images stored in the DBP 230 as described above) and other inter - frame prediction parameters, such as depending on whether the entire reference image or only a part of the reference image, such as a search window region near the area of the current block, is used to search for the best - matching reference block, and / or for example depending on whether pixel interpolation of half - pixel, quarter - pixel, and / or sixteenth - pixel interpolation is performed.
[0189] In addition to the above - mentioned prediction modes, a skip mode and / or a direct mode can also be adopted.
[0190] For example, in extended merge prediction, the merge candidate list for this mode consists of the following five candidate types in order: spatial MVP from spatially adjacent CUs, temporal MVP from collocated CUs, history-based MVP from the FIFO table, pairwise average MVP, and zero MV. Decoder side motion vector refinement (DMVR) based on bilateral matching can be used to increase the accuracy of the MVs in the merge mode. The merge mode with MVD (MMVD) comes from the merge mode with motion vector differences. The MMVD flag is sent immediately after the skip flag and the merge flag to specify whether the CU uses the MMVD mode. An adaptive motion vector resolution (AMVR) scheme for the CU level can be used. AMVR supports encoding the MVD of the CU with different precisions. The MVD of the current CU is adaptively selected according to the prediction mode of the current CU. When the CU is encoded in the merge mode, the combined inter / intra prediction (CIIP) mode can be applied to the current CU. The inter and intra prediction signals are weighted and averaged to obtain the CIIP prediction. For affine motion compensation prediction, the affine motion field of the block is described by the motion information of the motion vectors of 2 control points (4 parameters) or 3 control points (6 parameters). Subblock-based temporal motion vector prediction (SbTMVP), similar to the temporal motion vector prediction (TMVP) in HEVC, predicts the motion vectors of the sub-CUs within the current CU. Bi-directional optical flow (BDOF), formerly known as BIO, is a simplified version that reduces calculations, especially in terms of the number of multiplications and the size of the multipliers. In the triangular split mode, the CU is evenly divided into two triangular parts in two ways: diagonal division and anti-diagonal division. In addition, the bi-directional prediction mode is extended based on simple averaging to support weighted averaging of two prediction signals.
[0191] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (both in Figure 2(not shown in the figure). The motion estimation unit can be used to receive or obtain the image block 203 (the current image block 203 of the current image 17) and the decoded image 231, or at least one or more previously reconstructed blocks, for example, the reconstructed blocks of one or more other / different previously decoded images 231, to perform motion estimation. For example, the video sequence can include the current image and the previously decoded image 231, or in other words, the current image and the previously decoded image 231 can be part of the image sequence forming the video sequence or form the image sequence.
[0192] For example, the encoder 20 can be used to select a reference block from multiple reference blocks of the same or different images in multiple other images, and provide the reference image (or reference image index) and / or the offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-frame prediction parameter to the motion estimation unit. This offset is also called a motion vector (MV).
[0193] The motion compensation unit is used to obtain, for example, receive, the inter-frame prediction parameter, and perform inter-frame prediction according to or using the inter-frame prediction parameter to obtain the inter-frame prediction block 246. The motion compensation performed by the motion compensation unit may include extracting or generating a prediction block according to the motion / block vector determined by motion estimation, and may also include performing interpolation at sub-pixel accuracy. The interpolation filter can generate pixel points of other pixels from the pixel points of known pixels, thereby potentially increasing the number of candidate prediction blocks available for encoding the image block. Once the motion vector corresponding to the PU of the current image block is received, the motion compensation unit can locate the prediction block pointed to by the motion vector in one of the reference image lists.
[0194] The motion compensation unit can also generate syntax elements related to the block and the video slice for use by the video decoder 30 when decoding the image block of the video slice. In addition, or as an alternative to the slice and the corresponding syntax elements, coding block groups and / or coding blocks and the corresponding syntax elements can be generated or used.
[0195] Entropy coding
[0196] The entropy coding unit 270 is used to apply an entropy coding algorithm or scheme (e.g., variable length coding (VLC) scheme, context adaptive VLC (CALVC), arithmetic coding scheme, binarization algorithm, context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques) to the quantized residual coefficients 209, inter-frame prediction parameters, intra-frame prediction parameters, loop filter parameters, and / or other syntax elements, to obtain coded image data 21 that can be output in the form of a coded bitstream 21, etc. through the output terminal 272, such that a video decoder 30, etc. can receive and use the parameters for decoding. The coded bitstream 21 can be transmitted to the video decoder 30 or stored in a memory for later transmission or retrieval by the video decoder 30.
[0197] Other structural variants of the video encoder 20 can be used to encode a video stream. For example, a non-transform-based encoder 20 can directly quantize the residual signal in the case where some blocks or frames do not have a transform processing unit 206. In another implementation, the encoder 20 can have a quantization unit 208 and an inverse quantization unit 210 combined into a single unit.
[0198] Decoder and decoding method
[0199] Figure 3 This is an exemplary block diagram of the video decoder 30 according to an embodiment of the present application. The video decoder 30 is used to receive coded image data 21 (e.g., coded bitstream 21) encoded by, for example, the encoder 20, to obtain a decoded image 331. The coded image data or bitstream includes information for decoding the coded image data, such as data representing image blocks of a coded video slice (and / or coded group of blocks or coded block) and related syntax elements.
[0200] In Figure 3In the example of, the decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (such as an adder 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. The inter prediction unit 344 may be or include a motion compensation unit. In some examples, the video decoder 30 may perform a decoding process that is generally the reverse of the encoding process described with reference to Figure 2 the video encoder 100.
[0201] As described for the encoder 20, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer DPB 230, the inter prediction unit 344, and the intra prediction unit 354 also constitute the "built-in decoder" of the video encoder 20. Accordingly, the inverse quantization unit 310 may be functionally the same as the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally the same as the inverse transform processing unit 122, the reconstruction unit 314 may be functionally the same as the reconstruction unit 214, the loop filter 320 may be functionally the same as the loop filter 220, and the decoded picture buffer 330 may be functionally the same as the decoded picture buffer 230. Therefore, the explanations of the corresponding units and functions of the video encoder 20 correspondingly apply to the corresponding units and functions of the video decoder 30.
[0202] Entropy Decoding
[0203] The entropy decoding unit 304 is used to parse the bitstream 21 (or generally the encoded picture data 21) and perform entropy decoding on the encoded picture data 21 to obtain quantization coefficients 309 and / or decoded encoded parameters ( Figure 3 not shown in), such as any one or all of inter prediction parameters (such as reference picture indices and motion vectors), intra prediction parameters (such as intra prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be used to apply a decoding algorithm or scheme corresponding to the encoding scheme of the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may also be used to provide inter prediction parameters, intra prediction parameters, and / or other syntax elements to the mode application unit 360, and other parameters to other units of the decoder 30. The video decoder 30 may receive video slice and / or video block-level syntax elements. Additionally, or as an alternative to slices and corresponding syntax elements, encoded block groups and / or encoded blocks and corresponding syntax elements may be received or used.
[0204] Inverse Quantization
[0205] The inverse quantization unit 310 can be used to receive a quantization parameter (QP) (or generally information related to inverse quantization) and quantized coefficients from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304), and inverse-quantize the decoded quantized coefficients 309 based on the quantization parameter to obtain inverse-quantized coefficients 311, which may also be referred to as transform coefficients 311. The inverse quantization process may include determining the degree of quantization using the quantization parameter calculated by the video encoder 20 for each video block in the video slice, and also determining the degree of inverse quantization to be performed.
[0206] Inverse transform
[0207] The inverse transform processing unit 312 can be used to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain a reconstructed residual block 213 in the pixel domain. The reconstructed residual block 213 may also be referred to as a transform block 313. The transform can be an inverse transform, such as an inverse DCT, inverse DST, inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 can also be used to receive transform parameters or corresponding information from the encoded image data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304) to determine the transform to be applied to the dequantized coefficients 311.
[0208] Reconstruction
[0209] The reconstruction unit 314 (e.g., adder 314) is used to add the reconstructed residual block 313 to the prediction block 365 to obtain a reconstructed block 315 in the pixel domain, e.g., adding the pixel point values of the reconstructed residual block 313 and the pixel point values of the prediction block 365.
[0210] Filtering
[0211] The loop filter unit 320 (in or after the encoding loop) is used to filter the reconstructed block 315 to obtain the filtered block 321, so as to smoothly perform pixel transformation or improve video quality, etc. The loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample - adaptive offset (SAO) filter, or one or more other filters, such as an adaptive loop filter (ALF), a noise suppression filter (NSF), or any combination. For example, the loop filter unit 220 may include a deblocking filter, an SAO filter, and an ALF filter. The order of the filtering process may be a deblocking filter, an SAO filter, and an ALF filter. For another example, a process called luma mapping with chroma scaling (LMCS) (i.e., an adaptive in - loop shaper) is added. This process is performed before deblocking. For another example, the deblocking filtering process may also be applied to internal sub - block edges, such as affine sub - block edges, ATMVP sub - block edges, sub - block transform (SBT) edges, and intra sub - partition (ISP) edges. Although the loop filter unit 320 is shown as a loop filter in Figure 3 it may be implemented as a post - loop filter in other configurations.
[0212] Decoded picture buffer
[0213] Subsequently, the decoded video block 321 in an image is stored in the decoded picture buffer 330, and the decoded picture buffer 330 stores the decoded picture 331 as a reference picture, and the reference picture is used for subsequent motion compensation of other pictures and / or separately output for display.
[0214] The decoder 30 is used to output the decoded picture 311 through the output terminal 312, etc., for the user to display or view.
[0215] Prediction
[0216] The inter - prediction unit 344 may be functionally the same as the inter - prediction unit 244 (especially the motion compensation unit), the intra - prediction unit 354 may be functionally the same as the intra - prediction unit 254, and determines the division or segmentation and performs prediction based on the segmentation and / or prediction parameters or corresponding information received from the encoded picture data 21 (e.g., parsed and / or decoded by the entropy decoding unit 304). The mode application unit 360 may be used to perform the prediction (intra - or inter - prediction) of each block according to the reconstructed block, block, or corresponding sample points (filtered or unfiltered) to obtain the predicted block 365.
[0217] When encoding a video slice as an intra-coded (I) slice, the intra prediction unit 354 in the mode application unit 360 is used to generate a prediction block 365 for an image block of the current video slice based on the indicated intra prediction mode and data from previously decoded blocks of the current image. When the video image is encoded as an inter-coded (i.e., B or P) slice, the inter prediction unit 344 (e.g., a motion compensation unit) in the mode application unit 360 is used to generate a prediction block 365 for a video block of the current video slice based on a motion vector and other syntax elements received from the entropy decoding unit 304. For inter prediction, these prediction blocks can be generated from one of the reference images in one of the reference image lists. The video decoder 30 can construct reference frame lists 0 and 1 using a default construction technique based on the reference images stored in the DPB 330. In addition to or as an alternative to a slice (e.g., a video slice), the same or a similar process can be applied to embodiments of a coding block group (e.g., a video coding block group) and / or a coding block (e.g., a video coding block), e.g., a video can be encoded using I, P, or B coding block groups and / or coding blocks.
[0218] The mode application unit 360 is used to determine prediction information for a video block of the current video slice by parsing the motion vector and other syntax elements, and use the prediction information to generate a prediction block for the current video block being decoded. For example, the mode application unit 360 uses some received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for a video block of an encoded video slice, an inter prediction slice type (e.g., a B slice, a P slice, or a GPB slice), construction information for one or more reference image lists for the slice, a motion vector for each inter-coded video block of the slice, an inter prediction state for each inter-coded video block of the slice, and other information to decode a video block within the current video slice. In addition to or as an alternative to a slice (e.g., a video slice), the same or a similar process can be applied to embodiments of a coding block group (e.g., a video coding block group) and / or a coding block (e.g., a video coding block), e.g., a video can be encoded using I, P, or B coding block groups and / or coding blocks.
[0219] In one embodiment, Figure 3 The illustrated video encoder 30 can also be used to segment and / or decode an image using slices (also referred to as video slices), where the image can be segmented or decoded using one or more slices (usually non-overlapping). Each slice can include one or more blocks (e.g., CTUs) or one or more block groups (e.g., coding blocks in the H.265 / HEVC / VVC standards and tiles in the VVC standard).
[0220] In one embodiment, Figure 3The video decoder 30 shown can also be used to segment and / or decode an image using slices / coding block groups (also referred to as video coding block groups) and / or coding blocks (also referred to as video coding blocks), where the image can be segmented or decoded using one or more slices / coding block groups (usually non-overlapping), each slice / coding block group may include one or more blocks (e.g., CTUs) or one or more coding blocks, etc., and each coding block can be in a shape such as a rectangle and may include one or more whole or partial blocks (e.g., CTUs).
[0221] Other variants of the video decoder 30 can be used to decode the encoded image data 21. For example, the decoder 30 can generate an output video stream without the loop filter unit 320. For example, a non-transform-based decoder 30 can directly dequantize the residual signal without the inverse transform processing unit 312 for some blocks or frames. In another implementation, the video decoder 30 can have a dequantization unit 310 and an inverse transform processing unit 312 combined into a single unit.
[0222] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step can be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations can be performed on the processing results of interpolation filtering, motion vector derivation, or loop filtering, such as clip or shift operations.
[0223] It should be noted that further operations can be performed on the derived motion vectors of the current block (including but not limited to the control point motion vectors in the affine mode, the sub-block motion vectors in the affine, planar, ATMVP modes, the temporal motion vectors, etc.). For example, the value of the motion vector is restricted to a predefined range according to the representation bits of the motion vector. If the representation bits of the motion vector are bitDepth, the range is from -2^(bitDepth - 1) to 2^(bitDepth - 1) - 1, where "^" represents the power. For example, if bitDepth is set to 16, the range is from -32768 to 32767; if bitDepth is set to 18, the range is from -131072 to 131071. For example, the values of the derived motion vectors (e.g., the MVs of 4 4×4 sub-blocks in an 8×8 block) are restricted such that the maximum difference between the integer parts of the 4 4×4 sub-block MVs does not exceed N pixels, for example, does not exceed 1 pixel. Two methods for restricting the motion vector according to bitDepth are provided here.
[0224] Although the above embodiments mainly describe video encoding and decoding, it should be noted that embodiments of the decoding system 10, the encoder 20, and the decoder 30, as well as other embodiments described herein, can also be used for still image processing or encoding / decoding, that is, the processing or encoding / decoding of a single image independent of any previous or consecutive images in video encoding / decoding. Generally, if the image processing is limited to a single image 17, the inter-frame prediction units 244 (encoder) and 344 (decoder) may not be available. All other functions (also referred to as tools or techniques) of the video encoder 20 and the video decoder 30 can equally be used for still image processing, such as residual calculation 204 / 304, transformation 206, quantization 208, dequantization 210 / 310, (inverse) transformation 212 / 312, segmentation 262 / 362, intra-frame prediction 254 / 354, and / or loop filtering 220 / 320, entropy encoding 270, and entropy decoding 304.
[0225] Figure 4 Exemplary block diagram of the video decoding device 400 according to an embodiment of the present application. The video decoding device 400 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 400 can be a decoder, such as Figure 1A the video decoder 30 in Figure 1A or an encoder, such as
[0226] the video encoder 20 in
[0227] The processor 430 is implemented by hardware and software. The processor 430 can be implemented as one or more processor chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 communicates with the input port 410, the receiving unit 420, the transmitting unit 440, the output port 450, and the memory 460. The processor 430 includes a decoding module 470. The decoding module 470 implements the embodiments disclosed above. For example, the decoding module 470 performs, processes, prepares, or provides various encoding operations. Thus, the decoding module 470 provides a substantial improvement to the functions of the video decoding device 400 and affects the switching of the video decoding device 400 to different states. Alternatively, the decoding module 470 is implemented by instructions stored in the memory 460 and executed by the processor 430.
[0228] The memory 460 includes one or more disks, tape drives, and solid-state drives, and can be used as an overflow data storage device for storing such programs when a selected program is to be executed, and for storing instructions and data read during program execution. The memory 460 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0229] Figure 5 Exemplary block diagram of the apparatus 500 according to an embodiment of the present application. The apparatus 500 can be used as Figure 1A either or both of the source device 12 and the destination device 14 in
[0230] The processor 502 in the apparatus 500 can be a central processing unit. Alternatively, the processor 502 can be any other type of device or devices that can manipulate or process information, existing or to be developed in the future. Although the disclosed implementations can be implemented using a single processor such as the processor 502 shown in the figure, using more than one processor is faster and more efficient.
[0231] In one implementation, the memory 504 in the apparatus 500 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that the processor 502 accesses via the bus 512. The memory 504 may further include an operating system 508 and application programs 510, and the application programs 510 include at least one program that allows the processor 502 to execute the methods described herein. For example, the application programs 510 may include Applications 1 to N, and further include a video decoding application that executes the methods described herein.
[0232] The apparatus 500 may further include one or more output devices, such as a display 518. In one example, the display 518 may be a touch-sensitive display that combines a display with a touch-sensitive element that can be used to sense touch inputs. The display 518 may be coupled to the processor 502 via the bus 512.
[0233] Although the bus 512 in the apparatus 500 is described herein as a single bus, the bus 512 may include multiple buses. In addition, the secondary storage may be directly coupled to other components of the apparatus 500 or accessed via a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Therefore, the apparatus 500 may have various configurations.
[0234] Figure 8 This is an exemplary flowchart of the enhancement layer encoding method of the present application. The process 800 may be executed by the video encoder 20 (or encoder) and the video decoder 30 (or decoder). The process 800 is described as a series of steps or operations, and it should be understood that the process 800 may be executed in various orders and / or occur simultaneously, not limited to Figure 8 the execution order shown.
[0235] In step 801, the encoder obtains the reconstructed block of the base layer of the image block to be encoded.
[0236] Scalable video coding, also known as scalable video encoding, is an extended coding standard for current video coding standards (generally the extended standard of advanced video coding (AVC) (H.264), scalable video coding (SVC), or the extended standard of high efficiency video coding (HEVC) (H.265), scalable high efficiency video coding (SHVC)). The emergence of scalable video coding is mainly to solve the problems of packet loss and delay jitter caused by the real-time change of network bandwidth in real-time video transmission.
[0237] The basic structure in scalable video coding can be called a layer. Through spatial scalability (resolution scalability) of the original image blocks, scalable video coding technology can obtain bitstreams of different resolution layers. Resolution can refer to the size of the image block in pixels. The resolution of the lower layer is lower, and the resolution of the higher layer is not lower than that of the lower layer; or, through temporal scalability (frame rate scalability) of the original image blocks, bitstreams of different frame rate layers can be obtained. Frame rate can refer to the number of image frames contained in the video per unit time. The frame rate of the lower layer is lower, and the frame rate of the higher layer is not lower than that of the lower layer; or, through quality scalability of the original image blocks, bitstreams of different coding quality layers can be obtained. Coding quality can refer to the quality of the video. The degree of image distortion of the lower layer is larger, and the degree of image distortion of the higher layer is not higher than that of the lower layer.
[0238] Generally, the layer called the base layer is the bottom layer in scalable video coding. In spatial scalability, the base layer image blocks are encoded using the lowest resolution; in temporal scalability, the base layer image blocks are encoded using the lowest frame rate; in quality scalability, the base layer image blocks are encoded using the highest QP or the lowest bit rate. That is, the base layer is the layer with the lowest quality in scalable video coding. The layer called the enhancement layer is the layer above the base layer in scalable video coding, which can be divided into multiple enhancement layers from low to high. The lowest layer enhancement layer obtains the coding information based on the base layer, and its coding resolution is higher than that of the base layer, or the frame rate is higher than that of the base layer, or the bit rate is larger than that of the base layer. Higher layer enhancement layers can encode higher quality image blocks based on the coding information of lower layer enhancement layers.
[0239] For example, Figure 9 is an exemplary layer schematic diagram of the scalable video coding of this application, as Figure 9As shown, after the original image block is fed into the scalable encoder, it can be layered into a base layer image block B and enhancement layer image blocks (E1 to En, n≥1) according to different coding configurations, and then encoded separately to obtain a bitstream containing the base layer bitstream and the enhancement layer bitstream. The base layer bitstream is generally the bitstream obtained by encoding the lowest spatial domain image block, the lowest temporal domain image block, or the lowest quality image block. The enhancement layer bitstream is based on the base layer and is obtained by superposing and encoding high-level spatial domain, high-level temporal domain, or high-level quality image blocks. As the number of enhancement layers increases, the encoded spatial domain level, temporal domain level, or quality level will also become higher and higher. When the encoder transmits the bitstream to the decoder, it first ensures the transmission of the base layer bitstream. When there is network bandwidth available, it gradually transmits bitstreams of higher and higher levels. The decoder first receives and decodes the base layer bitstream, and then, according to the received enhancement layer bitstream, decodes bitstreams with higher and higher levels of spatial domain, temporal domain, or quality in order from the lower level to the higher level, and then superimposes the decoded information of the higher level on the reconstructed block of the lower level to obtain a reconstructed block with higher resolution, higher frame rate, or higher quality.
[0240] The image block to be encoded can refer to the image block that the encoder is currently processing, and this image block can refer to the largest coding unit (LCU) in the entire frame image. The entire frame image can refer to any image frame in the image sequence contained in the video that the encoder is processing. This image frame has not been divided and its size is the size of a complete image frame. In the H.265 standard, before video encoding, the original image frame is divided into multiple coding tree units (CTUs). A CTU is the largest coding unit for video encoding and can be divided into CUs of different sizes in a quadtree manner. Since the CTU is the largest coding unit, it is also called the LCU; alternatively, this image block can also refer to the entire frame image; or, this image block can also refer to the region of interest (ROI) in the entire frame image, that is, a specified image area that needs to be processed in the image.
[0241] As described above, each image in the video sequence is usually divided into a set of non-overlapping blocks and is usually encoded at the block level. In other words, the encoder usually processes and encodes the video at the block (image block) level. For example, it generates a predicted block through spatial (intra-frame) prediction and temporal (inter-frame) prediction; subtracts the predicted block from the image block (the currently processed / to-be-processed block) to obtain a residual block; transforms and quantizes the residual block in the transform domain to reduce the amount of data to be transmitted (compressed). The encoder also needs to perform inverse quantization and inverse transformation to obtain the reconstructed residual block, and then add the pixel values of the reconstructed residual block and the pixel values of the predicted block to obtain the reconstructed block. The reconstructed block of the base layer refers to the reconstructed block obtained by performing the above operations on the base layer image block obtained by layer-dividing the original image block. For example,Figure 10 FIG. 0 is an exemplary flowchart of an encoding method for an enhancement layer of the present application. As shown in Figure 10 FIG. 1, the encoder obtains a prediction block of the base layer according to an original image block (e.g., LCU), then calculates the difference between corresponding pixel points in the original image block and the prediction block of the base layer to obtain a residual block of the base layer. After dividing the residual block of the base layer, transform and quantization are performed, and entropy encoding is jointly performed with base layer encoding control information, prediction information, motion information, etc. to obtain a bitstream of the base layer. The encoder performs inverse quantization and inverse transform on the quantized quantization coefficients to obtain a reconstructed residual block of the base layer, and then sums the corresponding pixel points in the prediction block of the base layer and the reconstructed residual block of the base layer to obtain a reconstructed block of the base layer.
[0242] Step 802: The encoder calculates the difference between corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer to obtain a residual block of the enhancement layer of the image block to be encoded.
[0243] The resolution of the enhancement layer image block is not lower than that of the base layer image block, or the encoding quality of the enhancement layer image block is not lower than that of the base layer image block. As described above, the encoding quality may refer to the quality of the video. The encoding quality of the enhancement layer image block not being lower than that of the base layer image block may mean that the image distortion degree of the base layer is relatively large, while the image distortion degree of the enhancement layer is not higher than that of the base layer.
[0244] An image block includes a plurality of pixel points, and these pixel points are arranged in an array. Each pixel point can be uniquely identified by a row number and a column number for its position in the image block. Assume that the sizes of the image block to be encoded and the reconstructed block of the base layer are both M×N, that is, the image block to be encoded and the reconstructed block of the base layer respectively contain M×N pixel points. a(i1, j1) represents the pixel point located in the i1-th column and the j1-th row in the image block to be encoded, where i1 = 1~M and j1 = 1~N. b(i2, j2) represents the pixel point located in the i1-th column and the j1-th row in the reconstructed block of the base layer, where i2 = 1~M and j2 = 1~N. The corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer refer to the pixel points whose row numbers and column numbers in their respective image blocks are equal, that is, i1 = i2 and j1 = j2. For example, if the size of the image block is 16×16, it contains 16×16 pixel points, with row numbers from 0 to 15 and column numbers from 0 to 15. The pixel point labeled a(0, 0) in the image block to be encoded and the pixel point labeled b(0, 0) in the reconstructed block of the base layer are corresponding pixel points, or the pixel point labeled a(6, 9) in the image block to be encoded and the pixel point labeled b(6, 9) in the reconstructed block of the base layer are corresponding pixel points, and so on. Calculating the difference between corresponding pixel points may be to calculate the difference between the pixel values of the corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer. The pixel value may be the luminance value, chrominance value, etc. of the pixel point. The present application does not make specific limitations on this.
[0245] As Figure 10 described, in the SHVC standard, the encoder predicts the prediction blocks of the enhancement layer based on the reconstructed blocks of the base layer, and then calculates the difference between the corresponding pixel points in the reconstructed blocks of the base layer and the prediction blocks of the enhancement layer to obtain the residual blocks of the enhancement layer. However, in step 802, the difference can be directly calculated between the corresponding pixel points in the block to be encoded and the reconstructed blocks of the base layer to obtain the residual blocks of the enhancement layer, reducing the process of obtaining the prediction blocks of the enhancement layer, which can reduce the processing flow of the encoder and improve the encoding efficiency of the encoder.
[0246] Step 803: The encoder determines the transform block partitioning method for the residual blocks of the enhancement layer.
[0247] The transform block partitioning method for the residual blocks of the enhancement layer is different from that for the residual blocks of the base layer. That is, when the encoder processes the residual blocks of the enhancement layer, the transform block partitioning method used is different from that used when processing the residual blocks of the base layer. For example, the residual blocks of the base layer are partitioned into three sub-blocks using the TT partitioning method, but the residual blocks of the enhancement layer are not partitioned using the TT partitioning method. The encoder can determine the transform block partitioning method for the residual blocks of the enhancement layer in the following adaptive manner:
[0248] In a possible implementation, the encoder can first perform iterative tree-structured partitioning on the first largest transform unit (LTU) of the residual blocks of the enhancement layer to obtain transform units (TUs) with multiple segmentation depths. The maximum segmentation depth among the multiple segmentation depths is equal to Dmax, where Dmax is a positive integer. Then, the partitioning method of the first TU is determined based on the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs included in the first TU. The segmentation depth of the first TU is i, and the depth of the second TU is i + 1, where 0 ≤ i ≤ Dmax - 1.
[0249] Corresponding to the size of the image block to be encoded, the size of the residual block in the enhancement layer can be the size of the entire frame image, or the size of the image blocks (such as CTUs or CUs) divided in the entire frame image, or the size of the ROI in the entire frame image. The LTU is the image block with the largest size for transform processing in the image block. The size of the LTU can be the same as the size of the reconstruction block in the base layer, which can maximize the efficiency of transform coding while ensuring parallel processing of the enhancement layer between different reconstruction blocks. An LTU can be divided into multiple nodes based on the configured partitioning information, and each node can be further divided based on the configured partitioning information until all nodes are no longer divided. This process can be called iterative tree-structured partitioning. The tree-structured partitioning can include QT partitioning, BT partitioning, and / or TT partitioning, and can also include EQT partitioning. The present application does not specifically limit the partitioning method of the LTU. A node is divided once to obtain multiple nodes. The divided node is called the parent node, and the divided nodes are called child nodes. The splitting depth (depth) of the root node is 0, and the depth of the child node is the depth of the parent node plus 1. Therefore, the residual block in the enhancement layer starts iterative tree-structured partitioning from the root node LTU, and multiple TUs with different depths can be obtained. For example, as Figure 7 shown. The maximum depth among the multiple depths (i.e., the depth of the smallest TU obtained by partitioning) is Dmax. Assume that the width and height of the LTU are both L, and the width and height of the smallest TU are both S. Then the splitting depth of the smallest TU where represents rounding down. Therefore, Dmax ≤ d.
[0250] Based on the TU obtained by the above transform partitioning, the encoder performs rate distortion optimization (RDO) processing to determine the partitioning method of the LTU. Taking the first TU with a splitting depth of i and the second TU with a depth of i + 1 as an example, i starts from Dmax - 1 and decreases. The second TU is obtained by partitioning the first TU. The encoder first performs transform and quantization on the first TU to obtain the quantization coefficients of the first TU (the quantization coefficients can be obtained through quantization processing as in Figure 10 ), then pre-codes the quantization coefficients of the first TU (pre-coding is a coding process to estimate the length of the coded codeword, or a process using a method similar to coding) to obtain the bitstream size R of the first TU, then inverse quantizes and inverse-transforms the quantization coefficients of the first TU to obtain the reconstruction block of the TU, calculates the sum of squared errors between the first TU and the reconstruction block of the first TU to obtain the distortion value D of the first TU, and finally obtains the loss estimation value C of the first TU according to the bitstream size R of the first TU and the distortion value D of the first TU.
[0251] The calculation formula for the distortion value D of the first TU is as follows:
[0252]
[0253] Among them, P rs (i,j) represents the original value of the residual pixel in the first TU at the coordinate point (i,j) within the range of this TU, and P rc (i,j) represents the reconstructed value of the residual pixel in the first TU at the coordinate point (i,j) within the range of this TU.
[0254] The calculation formula for the loss estimation value C of the first TU is as follows:
[0255] C = D + λR
[0256] Among them, λ represents a constant value related to the quantization coefficient of the current layer, which determines the pixel distortion situation of the current layer.
[0257] The encoder can calculate the loss estimation value of the second TU by using the same method as above. After obtaining the loss estimation value of the first TU and the loss estimation values of multiple second TUs, the encoder compares the loss estimation value of the first TU with the sum of the loss estimation values of multiple second TUs. The first TU is divided into multiple second TUs, and the partitioning method corresponding to the smaller of the two is determined as the partitioning method of the first TU. That is, if the loss estimation value of the first TU is greater than the sum of the loss estimation values of multiple second TUs, then the method of dividing the first TU into multiple second TUs is determined as the partitioning method of the first TU; if the loss estimation value of the first TU is less than or equal to the sum of the loss estimation values of multiple second TUs, then the first TU is no longer divided. After the encoder traverses all the TUs of the LTU through the above method, the partitioning method of the LTU can be obtained. When the residual block of the enhancement layer contains only one LTU, the partitioning method of the LTU is the partitioning method of the residual block of the enhancement layer; when the residual block of the enhancement layer contains multiple LTUs, the partitioning method of the LTU and the method of dividing the residual block of the enhancement layer into LTUs constitute the partitioning method of the residual block of the enhancement layer.
[0258] This application no longer uses the same TU partitioning method as the residual block of the base layer for the residual block of the enhancement layer, but adopts an adaptive TU partitioning method to obtain a TU partitioning method suitable for the residual block of the enhancement layer, and then performs encoding. When the TU partitioning of the residual block of the enhancement layer is independent of the CU partitioning method of the base layer, the size of the TU of the enhancement layer is no longer limited by the size of the CU, which can improve the flexibility of encoding.
[0259] Step 804: The encoder transforms the residual block of the enhancement layer according to the transform block partitioning method to obtain the bitstream of the residual block of the enhancement layer.
[0260] The encoder can perform transformation, quantization, and entropy coding on the residual blocks of the enhancement layer according to the above transformation block partitioning method to obtain the bitstream of the residual blocks of the enhancement layer. The methods of transformation, quantization, and entropy coding can refer to the above description and will not be elaborated here.
[0261] Step 805: The decoder obtains the bitstream of the residual blocks of the enhancement layer of the image block to be decoded.
[0262] The encoder and the decoder can transmit the bitstream (including the bitstream of the base layer and the bitstream of the enhancement layer) through a wired or wireless link between them, which will not be elaborated here.
[0263] Step 806: The decoder performs entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual blocks of the enhancement layer to obtain the reconstructed residual blocks of the enhancement layer.
[0264] After the decoder obtains the bitstream of the residual blocks of the enhancement layer, it performs entropy decoding, inverse quantization, and inverse transformation processing on it based on the syntax elements carried in the bitstream to obtain the reconstructed residual blocks of the enhancement layer. This process can refer to the above description. The encoder also performs the above operations in the process of obtaining the reconstructed blocks, so it will not be elaborated here.
[0265] Step 807: The decoder obtains the reconstructed blocks of the base layer of the image block to be decoded.
[0266] The decoding end can refer to the description of step 801 and use the same method as the encoder to obtain the reconstructed blocks of the base layer, which will not be elaborated here.
[0267] Step 808: The decoder sums the corresponding pixel points in the reconstructed residual blocks of the enhancement layer and the reconstructed blocks of the base layer to obtain the reconstructed blocks of the enhancement layer.
[0268] The decoder can obtain the reconstructed blocks of the enhancement layer by adding the pixel points with the same row numbers and column numbers in the two image blocks according to the row numbers and column numbers of the multiple pixel points included in the reconstructed residual blocks of the enhancement layer and the row numbers and column numbers of the multiple pixel points included in the reconstructed blocks of the base layer.
[0269] This application can directly calculate the difference between the corresponding pixel points in the image block to be encoded and the reconstructed blocks of the base layer to obtain the residual blocks of the enhancement layer, reducing the process of obtaining the prediction blocks of the enhancement layer, which can reduce the processing flow of the encoder and improve the encoding efficiency of the encoder. In addition, the same TU partitioning method as that of the residual blocks of the base layer is no longer used for the residual blocks of the enhancement layer. Instead, an adaptive TU partitioning method is adopted to obtain the TU partitioning method suitable for the residual blocks of the enhancement layer and then perform encoding. When the TU partitioning of the residual blocks of the enhancement layer is independent of the CU partitioning method of the base layer, the TU size of the enhancement layer is no longer limited by the CU size, which can more effectively improve the compression efficiency of the residual blocks.
[0270] Figure 8The method shown can be applied to both spatial domain grading and quality domain grading. When image blocks are graded in the spatial domain to obtain base layer image blocks and enhancement layer image blocks, the resolution of the base layer image blocks is less than that of the enhancement layer image blocks. Therefore, before the encoder calculates the difference between the corresponding pixel points in the image block to be encoded and the reconstructed blocks of the base layer to obtain the residual block of the enhancement layer of the image block to be encoded, the original image block to be encoded can be downsampled to obtain the image block to be encoded with the first resolution, and the reconstructed blocks of the original base layer can be upsampled to obtain the reconstructed blocks of the base layer with the first resolution. That is, the encoder can downsample the original image block to be encoded and upsample the reconstructed blocks of the original base layer respectively, so that the resolution of the image block to be encoded is the same as that of the reconstructed blocks of the base layer.
[0271] Figure 11 FIG. 1100 is an exemplary flowchart of the encoding method for the enhancement layer of the present application. Process 1100 may be executed by video encoder 20 (or encoder) and video decoder 30 (or decoder). Process 1100 is described as a series of steps or operations. It should be understood that process 1100 may be executed in various orders and / or occur simultaneously, not limited to Figure 11 the execution order shown.
[0272] Step 1101, the encoder obtains the reconstructed block of the first enhancement layer of the image block to be encoded.
[0273] As Figure 9 shown, after the image blocks are graded by spatial domain grading, temporal domain grading or quality domain grading, in addition to the base layer, they are also divided into multiple enhancement layers. Figure 8 The method shown is an encoding and decoding method based on the base layer and the enhancement layer one level higher than the base layer. This embodiment is an encoding and decoding method based on the lower level enhancement layer (the first enhancement layer) and the second enhancement layer one level higher.
[0274] The difference from the method of obtaining the reconstructed blocks of the base layer is that: the encoding block can calculate the difference between the corresponding pixel points in the image block to be processed and the reconstructed blocks of the third layer of the image block to be processed to obtain the residual block of the first enhancement layer. The third layer is a level lower than the first enhancement layer, and it can be the base layer or an enhancement layer. Then, the residual block of the first enhancement layer is transformed and quantized to obtain the quantization coefficients of the residual block of the first enhancement layer. Then, the quantization coefficients of the residual block of the first enhancement layer are inverse quantized and inverse transformed to obtain the reconstructed residual block of the first enhancement layer. Finally, the corresponding pixel points in the reconstructed blocks of the third layer and the reconstructed residual block of the first enhancement layer are summed to obtain the reconstructed block of the first enhancement layer.
[0275] Step 1102, the encoder calculates the difference between the corresponding pixel points in the image block to be encoded and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the image block to be encoded.
[0276] Step 1103: The encoder determines the transform block partitioning method for the residual blocks of the second enhancement layer.
[0277] The transform block partitioning method for the residual blocks of the second enhancement layer is different from that of the first enhancement layer.
[0278] The encoder can also use RDO processing to determine the transform block partitioning method for the residual blocks of the second enhancement layer, which will not be elaborated here.
[0279] In this application, the same TU partitioning method as that of the residual blocks of the first enhancement layer and the basic layer is no longer used for the residual blocks of the second enhancement layer. Instead, an adaptive TU partitioning method is adopted to obtain the TU partitioning method suitable for the residual blocks of the second enhancement layer, and then encoding is performed. When the TU partitioning of the residual blocks of the second enhancement layer is independent of the CU or TU partitioning methods of other layers, the TU size of the second enhancement layer is no longer restricted by the CU size, or the TU size of the second enhancement layer is no longer restricted by the TU size of the first enhancement layer, which can improve the flexibility of encoding.
[0280] Step 1104: The encoder transforms the residual blocks of the second enhancement layer according to the transform block partitioning method to obtain the bitstream of the residual blocks of the second enhancement layer.
[0281] Step 1105: The decoder obtains the bitstream of the residual blocks of the second enhancement layer of the image block to be decoded.
[0282] Step 1106: The decoder performs entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual blocks of the second enhancement layer to obtain the reconstructed residual blocks of the second enhancement layer.
[0283] Step 1107: The decoder obtains the reconstructed blocks of the first enhancement layer of the image block to be decoded.
[0284] Step 1108: The decoder sums the corresponding pixel points in the reconstructed residual blocks of the second enhancement layer and the reconstructed blocks of the first enhancement layer to obtain the reconstructed blocks of the second enhancement layer.
[0285] The present application can directly calculate the difference between the corresponding pixel points in the image block to be coded and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer, reducing the process of obtaining the prediction block of the second enhancement layer, which can reduce the processing flow of the encoder and improve the coding efficiency of the encoder. Additionally, for the residual block of the second enhancement layer, instead of using the same TU partitioning method as the residual block of the first enhancement layer, an adaptive TU partitioning method is adopted to obtain the TU partitioning method suitable for the residual block of the second enhancement layer, and then encoding is performed. When the TU partitioning of the residual block of the second enhancement layer is independent of the CU or TU partitioning method of the first enhancement layer, the TU size of the second enhancement layer is no longer restricted by the CU size, or the TU size of the second enhancement layer is no longer restricted by the TU size of the first enhancement layer, which can more effectively improve the compression efficiency of the residual block.
[0286] Figure 11 The method shown can be applicable to two hierarchical methods: spatial domain hierarchical and quality domain hierarchical. When the image block is hierarchically divided into the first enhancement layer image block and the second enhancement layer image block using spatial domain hierarchical, the resolution of the first enhancement layer image block is smaller than that of the second enhancement layer image block. Therefore, before the encoder calculates the difference between the corresponding pixel points in the image block to be coded and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the image block to be coded, the original image block to be coded can be downsampled to obtain the image block to be coded with the first resolution, and the reconstructed block of the original first enhancement layer can be upsampled to obtain the reconstructed block of the first enhancement layer with the first resolution. That is, the encoder can downsample the original image block to be coded and upsample the reconstructed block of the original first enhancement layer respectively, so that the resolution of the image block to be coded is the same as that of the reconstructed block of the first enhancement layer.
[0287] The following uses several specific embodiments to Figure 8 and Figure 11 describe the embodiments shown.
[0288] Embodiment 1
[0289] Figure 12a and 12b are an exemplary flowchart of the coding method for the enhancement layer of the present application. As shown in Figure 12a and 12b , in this embodiment, quality domain hierarchical is used to hierarchically divide the image block, and the image block is an LCU.
[0290] Coding end:
[0291] Step 1: The encoder obtains the original image frame, encodes the basic layer image block according to the LCU, and obtains the reconstructed block of the basic layer of the LCU and the basic layer bitstream.
[0292] In this step, the encoder may use an encoder compliant with the H.264 / H.265 standard or other non-standard video encoders for encoding, and this application does not make any limitations in this regard. In this embodiment, the encoder is described by taking the H.265 standard as an example.
[0293] The original image frame is segmented according to the size of the LCU. For each LCU, CU division is performed to obtain the best prediction unit (PU), and intra-frame prediction or inter-frame prediction of the image block is performed to obtain the prediction block of the base layer. The residual quadtree (RQT) division process can be performed within the CU to obtain the best TU division. According to this TU division method, the residual blocks of the base layer can be transformed, quantized, and jointly entropy-encoded together with the LCU control information, prediction information, motion information, etc., so as to obtain the base layer bitstream. In the above process of obtaining the best PU and best TU divisions, RDO processing can be used to achieve the best coding compression ratio.
[0294] The quantized quantization coefficients are obtained as the reconstructed residual blocks of the base layer after the inverse quantization and inverse transformation processes, and added to the corresponding pixel points in the prediction blocks of the base layer to obtain the reconstructed blocks of the base layer. In this embodiment, the above process can be performed for each CU in the LCU, and the reconstructed blocks of the base layer of the entire LCU can be obtained.
[0295] Optionally, before performing step two, the encoder can perform further post-processing (such as loop filtering) on the reconstructed blocks of the base layer of the LCU, and then use them as the reconstructed blocks of the base layer in the LCU. There is no limitation on whether to perform the post-processing process. If the post-processing process is performed, the information indicating the post-processing needs to be encoded into the base layer bitstream.
[0296] Step two: Subtract the corresponding pixel points of the image block corresponding to the LCU from the reconstructed blocks of the base layer to obtain the residual blocks of the enhancement layer of the LCU.
[0297] Step three: Send the residual blocks of the enhancement layer of the LCU to the enhancement layer encoder for residual encoding.
[0298] In this step, after obtaining the residual blocks of the enhancement layer of the LCU, the LCU can be sent to the enhancement layer encoder for processing. It is pre-specified that the maximum transform unit (LTU) is the LCU of the enhancement layer at the current corresponding position, and TU adaptive division is performed to obtain TUs of different sizes. The specific method can use the RDO method based on the residual quadtree (RQT). The method is described as follows:
[0299] (1) Set the width and height of the LTU to be both L, which is a square of L×L. Predesignate the maximum splitting depth Dmax of the RQT splitting, where Dmax > 0, and the LTU is a TU with a splitting depth equal to 0. The designation of Dmax needs to consider the minimum TU size that can perform the transformation. The value of Dmax does not exceed the depth corresponding to the minimum TU size that can perform the transformation relative to the LTU size. For example, if the width and height values of the minimum TU that can perform the transformation process are S, then the splitting depth of the minimum TU where represents rounding down, that is, the value of Dmax must be less than or equal to d.
[0300] (2) Starting from the LTU, perform a quadtree split according to the tic-tac-toe grid, splitting it into four equal-sized TUs. The size of this TU is one-fourth of the LCU size, and the depth is 1. Determine whether the current TU depth is equal to the maximum splitting depth Dmax. If it is equal, no deeper split is performed; if not, perform the quadtree split process on each TU at this splitting depth, splitting it into smaller-sized TUs, and the splitting depth is based on the current depth value plus 1. This process continues until the quadtree is split to TUs with a splitting depth equal to Dmax.
[0301] (3) Perform RDO selection to obtain the best TU splitting method. Starting from the TU with a depth of Dmax - 1, perform transformation and quantization on the residual blocks of the enhancement layer at the TU position, obtain the quantized coefficients for precoding, and get the codeword length size R; on the other hand, perform inverse quantization and inverse transformation on the quantized coefficients to obtain the reconstructed residual blocks, calculate the sum of squared errors (SSD) between the reconstructed blocks and the residual blocks within the TU to obtain the distortion value Dr. According to R and Dr, the loss estimation value C1 can be obtained, which represents the loss estimation caused by not performing a quadtree split on this TU. Split the TU into four TUs with a splitting depth of Dmax by quadtree, and calculate in the same way as the splitting depth of Dmax - 1 for each TU to obtain the sum of the loss estimation values of the four TUs, denoted as C2, which represents the loss estimation caused by performing a quadtree split on this TU. Compare the values of C1 and C2, and select the smaller loss estimation value C3 = min(C1, C2) as the best loss value of this TU with a splitting depth of Dmax - 1, and use the splitting method corresponding to the smaller loss estimation value. According to this process, extrapolate and calculate for TUs with smaller splitting depths, calculate the best loss values for whether this TU performs a quadtree split or not respectively, compare to obtain the smaller loss value as the best loss of the TU, and perform the split according to the splitting method corresponding to the smaller loss value. Until it is decided whether the LTU should perform a quadtree split, the best TU splitting method of this LTU can be obtained.
[0302] After obtaining the optimal TU partitioning method for the LTU, according to this partitioning, the residual blocks of the enhancement layer in the LTU can be transformed and quantized in each TU block to obtain quantization coefficients, and entropy coding is performed together with the LTU partitioning information, thereby obtaining the enhancement layer bitstream of the LTU. On the other hand, through the inverse quantization and inverse transformation processes of the quantization coefficients, the coded distortion reconstruction block of the enhancement layer residual blocks of the LTU can be obtained, and this image is superimposed on the reconstruction block of the base layer in the LCU at the corresponding position of the LTU, and the reconstruction block corresponding to the enhancement layer of the LTU can be obtained.
[0303] Here, the enhancement layer bitstream of the LTU can be encoded using the syntax element organization and encoding method of the TU part bitstream in the existing coding standard.
[0304] In addition, for the TU adaptive partitioning method in the LTU, it is not limited to using the RQT method. Other methods that can adaptively partition the TU according to the enhancement layer residual blocks can also be applied here, and the corresponding coding methods for incorporating the TU partitioning information into the enhancement layer bitstream can also be different. After partitioning, the sizes of the TU blocks in the LTU can be different.
[0305] Step 4: According to Steps 1 to 3, when all the LCUs of the entire source image are completed, the reconstruction blocks of the base layer and the enhancement layer of the entire image, as well as the base layer bitstream and the enhancement layer bitstream of the entire image, can be obtained. These reconstruction blocks can be put into the reference frame list for subsequent image frames in the time domain to refer to. After obtaining the reconstruction blocks of the base layer or the enhancement layer, further post-processing (such as loop filtering) can also be performed. After processing, it is then put into the reference frame list, which is not limited here. If post-processing is performed, the indication information of performing post-processing and then putting it into the reference frame list needs to be encoded into the corresponding bitstream.
[0306] Here, the information indicating that the residual blocks of the enhancement layer are no longer partitioned according to the partitioning method of the base layer LCU or CU can also be encoded into the bitstream for the decoding end to perform enhancement layer processing after decoding.
[0307] Step 5: Repeat Steps 1 to 4 to encode the subsequent frame images of the image sequence until the entire image sequence encoding is completed.
[0308] Decoding end:
[0309] Step 1: Obtain the base layer bitstream and send it to the base layer decoder for decoding to obtain the base layer reconstruction block.
[0310] In this step, the base layer decoder uses a decoder that can decode the base layer bitstream corresponding to the base layer bitstream encoding rule. Here, a decoder compliant with the H.264 / H.265 standard can be used, or other non-standard video decoders can be used for decoding, which is not limited here. In this embodiment, the base layer decoder is described by taking a decoder compliant with the H.265 encoding standard as an example.
[0311] After the bitstream is sent into the decoder for entropy decoding, on the one hand, the coding block partition information, transform block partition information, control information, prediction information, and motion information of the LCU can be obtained. Through these information, intra prediction and inter prediction are performed to obtain the prediction block of the LCU. On the other hand, the residual information can be obtained. According to the transform block partition information obtained by decoding, through the inverse quantization and inverse transform processes, the residual image of the LCU can be obtained. By superimposing the prediction block and the residual image, the base layer reconstruction block corresponding to the LCU can be obtained. Repeat this process until all the LCUs of the entire frame image are decoded, and the reconstruction block of the entire image can be obtained.
[0312] If the information that needs to perform post-processing (such as loop filtering) in the reconstruction block is obtained by decoding, the corresponding post-processing process is performed after the reconstruction block is obtained to make it correspond to the process in the encoding.
[0313] After this step is completed, the base layer reconstruction block is obtained.
[0314] Step 2: Obtain the enhancement layer bitstream, send it into the enhancement layer decoder for decoding, obtain the reconstruction residual block of the enhancement layer, and further obtain the reconstruction block of the enhancement layer.
[0315] In this step, after the bitstream is sent into the decoder for entropy decoding, the partition information of the LTU and the residual information of the LTU are obtained. According to the partition of the LTU, the residual information is inversely quantized and inversely transformed, and the reconstruction residual block corresponding to the LTU can be obtained. Repeat this process until all the LTUs of the entire frame image are decoded, and the enhancement layer reconstruction residual block can be obtained. By adding the base layer reconstruction block and the enhancement layer reconstruction residual block, the reconstruction block is obtained.
[0316] If there is information in the bitstream indicating that the enhancement layer no longer follows the partition method of the LCU or CU of the base layer, this information can be obtained by entropy decoding first, and then Step 2 is executed.
[0317] In this step, if the information that needs to perform post-processing (such as loop filtering) in the reconstruction block is obtained by decoding, the corresponding post-processing process is performed after the reconstruction block to make it correspond to the process in the encoding.
[0318] After this step is completed, the obtained reconstruction block is the enhancement layer reconstruction block.
[0319] Step 3: Put the basic layer reconstruction blocks and enhancement layer reconstruction blocks into the reference frame list for subsequent frame images in the temporal domain to refer to.
[0320] If there is indication information in the bitstream to perform post-processing (such as loop filtering) on the reconstruction blocks and then put them into the reference frame list, this information can be obtained through entropy decoding. Perform post-processing on the basic layer reconstruction blocks and enhancement layer reconstruction blocks, and then put them into the reference frame list to make it correspond to the process in encoding.
[0321] Step 4: Repeat Steps 1 to 3 until the entire image sequence is decoded.
[0322] In this embodiment, the enhancement layer no longer performs the prediction process. Instead, it directly performs adaptive partitioning and transform coding on the residual blocks of the enhancement layer. Compared with the coding method that requires prediction, it can improve the coding speed more, reduce the pipelining process of the hardware, and save the hardware processing cost. For the residual blocks of the enhancement layer, the TU partitioning method in the basic layer is no longer used. Instead, a more suitable TU partitioning method is adaptively selected for coding. Therefore, the compression efficiency of the residual blocks can be improved more effectively. In addition, after the LCU in the basic layer is processed, the obtained LCU residual blocks can be directly sent to the enhancement layer for subsequent separate processing, while the basic layer can simultaneously process the next LCU in parallel. This method is more friendly to the hardware implementation of the pipelining process and can make the processing process more efficient.
[0323] Optionally, the above basic layer and enhancement layer can be processed according to the size of the entire frame image, rather than according to the size of the LCU in the above embodiment.
[0324] Embodiment 2
[0325] Figure 13a and 13b is an exemplary flowchart of the encoding method for the enhancement layer of the present application. As Figure 13a and 13b shown, this embodiment uses quality-domain grading to layer the image blocks to obtain at least two enhancement layers, and the image blocks are LCUs. This embodiment is an extension of Embodiment 1, that is, the encoding and decoding processing method when there are more than one enhancement layer.
[0326] Encoding end:
[0327] Step 1: The same as Step 1 of the encoding end in Embodiment 1
[0328] Step 2: The same as Step 2 of the encoding end in Embodiment 1
[0329] Step 3: The same as Step 3 of the encoding end in Embodiment 1.
[0330] Optionally, before performing Step Four, the encoder can also perform further post-processing (such as loop filtering) on the reconstructed blocks of the base layer of the LCU, and then use them as the reconstructed blocks of the first enhancement layer in the LCU. There is no limitation on whether to perform the post-processing process. If the post-processing process is performed, information indicating the post-processing needs to be added to the first enhancement layer bitstream.
[0331] Step Four: Subtract the corresponding pixel points of the image block corresponding to the LCU from the reconstructed blocks of the first enhancement layer to obtain the residual block of the second enhancement layer of the LCU.
[0332] Step Five: Send the residual block of the second enhancement layer obtained in Step Four into the second enhancement layer encoder for residual coding. The coding steps are the same as those in Step Three. The difference is that the coded distortion reconstructed block of the residual block of the second enhancement layer is superimposed on the reconstructed blocks of the first enhancement layer to obtain the reconstructed block of the second enhancement layer corresponding to the LCU.
[0333] Here, further post-processing (such as loop filtering) can also be performed on the superimposed image, and then used as the reconstructed block of the second enhancement layer in the LCU. There is no limitation on whether to perform the post-processing process. If the post-processing process is performed, information indicating the post-processing needs to be added to the second enhancement layer bitstream.
[0334] Step Six: Repeat Step Four and Step Five. Subtract the image block of the LCU from the pixels of the reconstructed blocks of the current previous layer of the enhancement layer to obtain the residual block of the next layer of the enhancement layer corresponding to the LCU, and send it into the enhancement layer encoder of the corresponding layer for residual coding to obtain the reconstructed block of the high-layer enhancement layer and the high-layer enhancement layer bitstream.
[0335] Step Seven: Repeat Steps One to Six to complete all LCUs of the entire source image, and the reconstructed blocks of the base layer of the entire image and all enhancement layer reconstructed blocks, as well as the base layer bitstream of the entire image and all enhancement layer bitstreams can be obtained. These reconstructed blocks can be put into the reference frame list for subsequent frame images in the time domain to refer to. After obtaining the reconstructed blocks of the base layer or enhancement layer, further post-processing (such as loop filtering) can also be performed, and then put into the reference frame list. There is no limitation here. If post-processing is performed, the indication information of performing post-processing and then putting it into the reference frame list needs to be encoded into the corresponding bitstream.
[0336] Here, information indicating that the enhancement layer no longer follows the LCU or CU partitioning method of the base layer, or information indicating that the high-layer enhancement layer no longer follows the LCU of the base layer or the LTU partitioning method of the low-layer enhancement layer can also be encoded into the bitstream for the decoding end to perform enhancement layer processing after decoding.
[0337] Step Eight: Repeat Steps One to Seven to encode the subsequent frame images of the image sequence until the entire image sequence encoding is completed.
[0338] Decoder side:
[0339] Step 1: Obtain the base layer bitstream, send it to the base layer decoder for decoding, and obtain the base layer reconstructed blocks.
[0340] In this step, the base layer decoder uses a decoder that can decode the base layer bitstream corresponding to the base layer bitstream encoding rule. Here, a decoder that conforms to the H.264 / H.265 standard can be used, or other non-standard video decoders can be used for decoding, which is not limited here. In this embodiment, the base layer decoder is described by taking a decoder that conforms to the H.265 encoding standard as an example.
[0341] After the bitstream is sent to the decoder for entropy decoding, on the one hand, the coding block partition information, transform block partition information, and control information, prediction information, and motion information of the LCU can be obtained. Through these information, intra-frame prediction and inter-frame prediction are performed to obtain the predicted block of the LCU; on the other hand, the residual information can be obtained. According to the transform block partition information obtained by decoding, through the inverse quantization and inverse transform processes, the residual image of the LCU can be obtained. By superimposing the predicted block and the residual image, the corresponding base layer reconstructed block of the LCU can be obtained. Repeat this process until all LCUs of the entire frame image are decoded, and the reconstructed blocks of the entire image can be obtained.
[0342] If the information that needs to perform post-processing (such as loop filtering) in the reconstructed block is obtained by decoding, then after obtaining the reconstructed block, the corresponding post-processing process is performed to make it correspond to the process in the encoding.
[0343] After completing this step, the base layer reconstructed blocks are obtained.
[0344] Step 2: Obtain the first enhancement layer bitstream, send it to the first enhancement layer decoder for decoding, obtain the reconstructed residual blocks of the first enhancement layer, and further obtain the reconstructed blocks of the first enhancement layer.
[0345] In this step, after the bitstream is sent to the decoder for entropy decoding, the partition information of the LTU and the residual information of the LTU are obtained. According to the partition of the LTU, the residual information is inverse quantized and inverse transformed to obtain the corresponding reconstructed residual blocks of the LTU. Repeat this process until all LTUs of the entire frame image are decoded, and the reconstructed residual blocks of the first enhancement layer can be obtained. Add the base layer reconstructed blocks and the reconstructed residual blocks of the first enhancement layer to obtain the reconstructed blocks.
[0346] If there is information in the bitstream indicating that the first enhancement layer no longer follows the partition method of the base layer's LCU or CU, this information can be obtained by entropy decoding first, and then step 2 is executed.
[0347] In this step, if information that post-processing (such as loop filtering) needs to be performed on the reconstructed block is decoded, the corresponding post-processing process is performed after the reconstructed block to make it corresponding to the process in encoding.
[0348] After this step is completed, the obtained reconstructed block is the reconstructed block of the first enhancement layer.
[0349] Step 3: Obtain the second enhancement layer bitstream, send it to the enhancement layer decoder for decoding to obtain the second enhancement layer reconstructed residual block, and further obtain the second enhancement layer reconstructed block.
[0350] This step is similar to Step 2, the difference being that the reconstructed block is obtained by adding the reconstructed block of the first enhancement layer and the reconstructed residual block of the second enhancement layer.
[0351] If there is information in the bitstream indicating that the enhancement layer no longer follows the partitioning method of the LCU or CU of the base layer, or information indicating that the high-level enhancement layer no longer follows the partitioning method of the LCU of the base layer or the LTU of the low-level enhancement layer, this information can be obtained through entropy decoding, and then Step 3 is executed.
[0352] In this step, if information that post-processing (such as loop filtering) needs to be performed on the reconstructed block is decoded, the indicated post-processing process is performed after the reconstructed block to make it corresponding to the process in encoding.
[0353] After this step is completed, the obtained reconstructed block is the reconstructed block of the second enhancement layer.
[0354] Step 4: Repeat Step 3 to obtain the bitstream of a higher-level enhancement layer, send it to the enhancement layer decoder for decoding to obtain the reconstructed residual block of the high-level enhancement layer, and further obtain the reconstructed block of the high-level enhancement layer. Until the decoding of all enhancement layer bitstreams of the current frame is completed.
[0355] Step 5: Put the reconstructed block of the base layer and the reconstructed blocks of all enhancement layers into the reference frame list for subsequent frames in the temporal domain to refer to.
[0356] If there is indication information in the bitstream to perform post-processing (such as loop filtering) on the reconstructed block and then put it into the reference frame list, this information can be obtained through entropy decoding, perform post-processing on the reconstructed block of the base layer and the reconstructed blocks of the enhancement layers, and then put them into the reference frame list to make it corresponding to the process in encoding.
[0357] Step 6: Repeat Steps 1 to 5 until the decoding of the entire image sequence is completed.
[0358] Embodiment 3
[0359] Figure 14a and 14b is an exemplary flowchart of the encoding method for the enhancement layer of the present application, as Figure 14a and14b As shown, in this embodiment, spatial domain grading is used to layer the image blocks to obtain at least two enhancement layers, and the image blocks are LCUs.
[0360] Encoding end:
[0361] Step 1: The encoder downsamples the original image according to Resolution 1 (image width and height are W1×H1) to generate the downsampled image 1. Here, Resolution 1 is less than the resolution of the original image block, and the sampling method can be any method, such as linear interpolation, bicubic interpolation, etc., which is not limited here.
[0362] Step 2: Send the downsampled image 1 into the base layer encoder for encoding, and the encoding steps are the same as Step 1 in Embodiment 1.
[0363] Step 3: Downsample the original image according to Image Resolution 2 (image width and height are W2×H2) to generate the downsampled image 2, where Resolution 2 is higher than Resolution 1. The scaling ratio of the downsampled image 2 to image 1 is Sw1 = W2 / W1 in width and Sh1 = H2 / H1 in height.
[0364] Step 4: Upsample the reconstructed block in the base layer of the LCU to obtain a reconstructed block with the same size as the LTU of the first enhancement layer, and then obtain the residual block of the first enhancement layer corresponding to the LTU.
[0365] Set the size of the LTU of the first enhancement layer to be the size of the LCU of the base layer multiplied by the scaling ratio Sw1×Sh1 to obtain an LTU of size Lw1×Lh1, where Lw1 = L×Sw1 and Lh1 = L×Sh1. Upsample the reconstructed block of the base layer corresponding to the LCU obtained in Step 2 to the size of Lw1×Lh1 to obtain a reconstructed block of the base layer with the same size as the LTU of the first enhancement layer. Take the LTU source image at the same position in the downsampled image 2. Subtract the above source image from the above reconstructed block of the base layer to obtain the residual block in the LTU of the first enhancement layer. Here, the upsampling method and the downsampling method in Step 1 can correspond or not, which is not limited here.
[0366] Step 5: Send the residual block in the LTU of the first enhancement layer into the first enhancement layer encoder for residual encoding. The encoding steps are the same as Step 3 in Embodiment 1.
[0367] Step 6: Downsample the original image according to Image Resolution 3 (image width and height are W3×H3) to generate the downsampled image 3, where Resolution 3 is higher than Resolution 2. The scaling ratio of the downsampled image 3 to image 2 is Sw2 = W3 / W2 in width and Sh2 = H3 / H2 in height.
[0368] Step 7: Upsample the LTU reconstruction blocks of the first enhancement layer to obtain reconstruction blocks with the same size as the LTUs of the second enhancement layer, and then obtain the residual blocks corresponding to the LTUs of the second enhancement layer.
[0369] This step is similar to Step 4. Set the size of the LTUs of the second enhancement layer to be the size of the LTUs of the first enhancement layer multiplied by the scaling ratios Sw2×Sh2. Upsample the reconstruction blocks in the LTUs of the first enhancement layer to obtain reconstruction blocks with the same size as the LTUs of the second enhancement layer. Also, take the LTU source images at the same positions in the downsampled image 3. After subtracting the above source images from the reconstruction blocks, the residual blocks in the LTUs of the second enhancement layer can be obtained.
[0370] Step 8: Send the residual blocks in the LTUs of the second enhancement layer into the second enhancement layer encoder for residual encoding. The encoding steps are the same as those in Step 3 of Embodiment 1.
[0371] Step 9: Repeat Step 6 to Step 8 until the reconstruction blocks corresponding to the LTUs of all enhancement layers and the enhancement layer bitstreams are obtained.
[0372] Step 10: Repeat Steps 1 to 9. After finishing all the LTUs of the downsampled images input for each layer, the base layer reconstruction blocks and all enhancement layer reconstruction blocks of the entire image can be obtained, as well as the base layer bitstream and all enhancement layer bitstreams. These reconstruction blocks can be put into the reference frame list for subsequent frame images in the time domain to refer to. After obtaining the base layer or enhancement layer reconstruction blocks, further post-processing (such as loop filtering) can be performed. After the processing, they are put into the reference frame list again, which is not limited here. If post-processing is performed, the indication information of putting the post-processed result into the reference frame list needs to be encoded into the corresponding bitstream.
[0373] Here, the resolution information of the downsampled images input for each layer is encoded into the bitstream for all layers to be processed during decoding. In addition, the information of the upsampling method for each layer can also be encoded into the bitstream.
[0374] Step 11: Repeat Steps 1 to 10 to encode the subsequent frame images of the image sequence until the entire image sequence encoding is completed.
[0375] Decoder side:
[0376] Step 1: The same as Step 1 of the decoder side in Embodiment 1
[0377] Step 2: Obtain the first enhancement layer bitstream of the image block being processed, send it into the enhancement layer decoder for decoding, and obtain the first enhancement layer reconstruction residual blocks. This step is the same as the process of obtaining the first enhancement layer reconstruction residual blocks in Step 2 of Embodiment 1.
[0378] Step 3: Based on the resolution of the first enhanced layer residual block, upsample the basic layer reconstruction block to the same resolution as the first enhanced layer residual block. Add the two together to obtain the reconstruction block of the first enhanced layer.
[0379] If there is information about the upsampling method in the bitstream, the information can be obtained by entropy decoding first, and the basic layer reconstruction block is upsampled according to the obtained upsampling information in this step.
[0380] Step 4: Obtain the bitstream of the second enhanced layer of the image block being processed, and send it to the enhanced layer decoder for decoding to obtain the second enhanced layer reconstruction residual block. This step is the same as the process of obtaining the first enhanced layer residual block.
[0381] Step 5: Based on the resolution of the second enhanced layer residual block, upsample the first enhanced layer reconstruction block to the same resolution as the second enhanced layer residual block. Add the two together to obtain the reconstruction block of the second enhanced layer.
[0382] If there is information about the upsampling method in the bitstream, the information can be obtained by entropy decoding first, and the first enhanced layer reconstruction block is upsampled according to the upsampling information in this step.
[0383] Step 6: Repeat steps 4 to 5 to obtain the bitstream of a higher-level enhanced layer, send it to the enhanced layer decoder for decoding to obtain the high-level enhanced layer reconstruction residual block, and add it to the upsampled image of the low-level enhanced layer reconstruction block to obtain the high-level enhanced layer reconstruction block. Repeat until the decoding of all enhanced layer bitstreams of the current frame is completed.
[0384] Step 7: The same as step 3 of the decoding end in Embodiment 1.
[0385] Step 8: Repeat steps 1 to 7 until the entire image sequence is decoded.
[0386] Embodiment 4
[0387] Figure 15 It is an exemplary flowchart of the encoding method for the enhanced layer of the present application. As Figure 15 shown, in this embodiment, quality-domain grading is used to layer the image blocks. In the basic layer, the entire image frame is processed, and in the enhanced layer, a partial area (such as the ROI in the entire image) is taken from the entire image for processing. The difference between this embodiment and Embodiment 1 is that a partial area is taken from the original image frame, and a corresponding partial area is taken from the reconstruction block of the basic layer. The difference between the corresponding pixel points in these two partial areas is obtained to get the residual block of this partial area, which is used as the residual block of the enhanced layer and sent to the enhanced layer encoder.
[0388] Encoding end:
[0389] Step 1: The same as step 1 of the encoding end in Embodiment 1
[0390] Step 2: Subtract a partial area (such as ROI) of the original image frame from the corresponding partial area in the reconstruction block of the base layer to obtain the residual block of the enhancement layer for this partial area.
[0391] It should be noted that the number of the above partial areas can be one or more, and the present application does not make specific limitations in this regard. If multiple partial areas are included, the following steps are respectively executed for each partial area.
[0392] Step 3: The same as Step 3 of the encoding end in Embodiment 1, the difference is that after the quantization coefficient undergoes inverse quantization and inverse transformation, the reconstructed residual block of the enhancement layer for the partial area is obtained, and the reconstructed residual block of the enhancement layer for this partial area is superimposed on the corresponding reconstructed block of the base layer of the partial area to obtain the reconstructed block of the enhancement layer for the partial area. The position and size information of the above partial area can be encoded into the bitstream.
[0393] Step 4: The same as Step 4 of the encoding end in Embodiment 1.
[0394] Step 5: The same as Step 5 of the encoding end in Embodiment 1.
[0395] Decoding end:
[0396] Step 1: The same as Step 1 of the decoding end in Embodiment 1.
[0397] Step 2: The same as Step 2 of the decoding end in Embodiment 1, the difference is that after obtaining the reconstructed residual block of the enhancement layer for the partial area, the reconstructed residual block of the enhancement layer is superimposed on the corresponding reconstructed block of the base layer of the partial area to obtain the reconstructed block of the enhancement layer for this partial area.
[0398] If the position and size information of the above partial area is stored in the bitstream, first decode to obtain this position and size information, then find this partial area in the reconstructed block of the base layer according to this position and size information, and then superimpose the reconstructed residual block of the enhancement layer of the partial area.
[0399] Step 3: The same as Step 3 of the decoding end in Embodiment 1.
[0400] Step 4: The same as Step 4 of the decoding end in Embodiment 1.
[0401] Embodiment 5
[0402] In this embodiment, a hybrid grading of quality domain grading and spatial domain grading is used to layer the image block to obtain at least two enhancement layers, and the image block is an LCU.
[0403] Based on the hierarchical structures of Embodiment 2 and Embodiment 3, it can be known that whether it is quality-domain grading or spatial-domain grading, the enhancement layer encoders used in both for the enhancement layer residual blocks are the same, both being residual block coding based on adaptive TU partitioning, and the coding process has been described in detail in Embodiments 1 to 4.
[0404] The difference between the two is as follows: In spatial-domain grading, the reconstruction block of the low-layer enhancement layer needs to go through an upsampling process to obtain a reconstruction block corresponding in size to the LTU of the high-layer enhancement layer, so as to obtain a residual block by taking the difference between the source image block and the corresponding pixels in the upsampled reconstruction block, and then send it to the high-layer enhancement layer encoder; on the other hand, the upsampled reconstruction block is superimposed with the reconstructed residual block of the high-layer enhancement layer to obtain the reconstructed block of the high-layer enhancement layer. Therefore, during the grading process, either spatial-domain grading or quality-domain grading can be used for any layer, and the two grading methods can be used in a mixed and interspersed manner. For example, the first enhancement layer uses spatial-domain grading, the second enhancement layer uses quality-domain grading, and the third enhancement layer uses spatial-domain grading again, and so on.
[0405] Correspondingly, a decoding method based on the above mixed grading bitstream can be obtained. Based on Embodiment 2 and Embodiment 3, it can be seen that whether it is quality-domain grading or spatial-domain grading, the decoding methods of the enhancement layer bitstreams are the same, both being the process of directly decoding the TU partitioning information and the residual information, and then processing the residual information into the reconstructed residual block of the enhancement layer based on the TU partitioning information. Therefore, for the bitstream that is a mixture of quality-domain grading bitstream and spatial-domain grading, it can also be correctly decoded according to the corresponding grading method of that layer.
[0406] Figure 16 This is a schematic structural diagram of an exemplary encoding device of the present application. As Figure 16 shown, the device of this embodiment can correspond to the video encoder 20. The device may include: an acquisition module 1601, a determination module 1602, and an encoding module 1603. Among them,
[0407] In a possible implementation manner, the acquisition module 1601 is configured to acquire the reconstructed block of the base layer of the image block to be encoded; take the difference between the corresponding pixel points of the image block to be encoded and the reconstructed block of the base layer to obtain the residual block of the enhancement layer of the image block to be encoded, where the resolution of the enhancement layer image block is not lower than that of the base layer image block, or the coding quality of the enhancement layer image block is not lower than that of the base layer image block; the determination module 1602 is configured to determine the transform block partitioning method of the residual block of the enhancement layer, where the transform block partitioning method of the residual block of the enhancement layer is different from that of the residual block of the base layer; the encoding module 1603 is configured to perform a transform on the residual block of the enhancement layer according to the transform block partitioning method to obtain the bitstream of the residual block of the enhancement layer.
[0408] In a possible implementation, the determining module 1602 is specifically configured to perform iterative tree-structured partitioning on the first largest transformation unit LTU of the residual block of the enhancement layer to obtain multiple transformation units TU with different segmentation depths, where the maximum segmentation depth among the multiple segmentation depths is equal to Dmax, and Dmax is a positive integer; determine the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, the segmentation depth of the first TU is i, and the depth of the second TU is i + 1, where 0 ≤ i ≤ Dmax - 1.
[0409] In a possible implementation, the determining module 1602 is specifically configured to perform transformation and quantization on the TU to obtain the quantization coefficient of the TU, where the TU is the first TU or the second TU; perform precoding on the quantization coefficient of the TU to obtain the codeword length of the TU; perform inverse quantization and inverse transformation on the quantization coefficient of the TU to obtain the reconstructed block of the TU; calculate the sum of squared differences SSD between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; obtain the loss estimation value of the TU according to the codeword length and the distortion value of the TU; determine the smaller value among the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; and determine the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
[0410] In a possible implementation, the size of the LTU is the same as the size of the reconstructed block of the base layer; or, the size of the LTU is the same as the size of the reconstructed block of the original base layer after upsampling.
[0411] In a possible implementation, the image block to be encoded refers to the largest coding unit LCU in the entire frame of the image; or, the image block to be encoded refers to the entire frame of the image; or, the image block to be encoded refers to the region of interest ROI in the entire frame of the image.
[0412] In a possible implementation, the base layer and the enhancement layer are obtained by layer division based on resolution or coding quality.
[0413] In a possible implementation, when the base layer and the enhancement layer are obtained by layer division based on resolution, the obtaining module 1601 is further configured to downsample the original image block to be encoded to obtain the image block to be encoded with the first resolution; and upsample the reconstructed block of the original base layer to obtain the reconstructed block of the base layer with the first resolution.
[0414] In a possible implementation, the obtaining module 1601 is specifically configured to calculate the difference between corresponding pixel points in the to-be-processed image block and the prediction block of the to-be-processed image block to obtain the residual block of the to-be-processed image block; perform transformation and quantization on the residual block of the to-be-processed image block to obtain the quantization coefficients of the to-be-processed image block; perform inverse quantization and inverse transformation on the quantization coefficients of the to-be-processed image block to obtain the reconstructed residual block of the to-be-processed image block; and calculate the sum of corresponding pixel points in the prediction block of the to-be-processed image block and the reconstructed residual block of the to-be-processed image block to obtain the reconstructed block of the base layer.
[0415] In a possible implementation, the obtaining module 1601 is configured to obtain the reconstructed block of the first enhancement layer of the to-be-encoded image block; calculate the difference between corresponding pixel points in the to-be-encoded image block and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the to-be-encoded image block, where the resolution of the second enhancement layer image block is not lower than that of the first enhancement layer image block, or the coding quality of the second enhancement layer image block is not lower than that of the first enhancement layer image block; the determining module 1602 is configured to determine the transformation block partitioning method of the residual block of the second enhancement layer, where the transformation block partitioning method of the residual block of the second enhancement layer is different from that of the residual block of the first layer; and the encoding module 1603 is configured to perform transformation on the residual block of the second enhancement layer according to the transformation block partitioning method to obtain the bitstream of the residual block of the second enhancement layer.
[0416] In a possible implementation, the determining module 1602 is specifically configured to perform iterative tree-structured partitioning on the first largest transform unit (LTU) of the residual block of the second enhancement layer to obtain transform units (TUs) with multiple partitioning depths, where the maximum partitioning depth among the multiple partitioning depths is equal to Dmax, and Dmax is a positive integer; and determine the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, where the partitioning depth of the first TU is i, and the depth of the second TU is i + 1, and 0 ≤ i ≤ Dmax - 1.
[0417] In a possible implementation, the determining module 1602 is specifically configured to perform transformation and quantization on the TU to obtain quantization coefficients of the TU, where the TU is the first TU or the second TU; perform precoding on the quantization coefficients of the TU to obtain the codeword length of the TU; perform inverse quantization and inverse transformation on the quantization coefficients of the TU to obtain the reconstructed block of the TU; calculate the sum of squared differences SSD between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; obtain the loss estimation value of the TU according to the codeword length and the distortion value of the TU; determine the smaller value among the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; and determine the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
[0418] In a possible implementation, the size of the LTU is the same as the size of the reconstructed block of the first enhancement layer; or, the size of the LTU is the same as the size after upsampling the reconstructed block of the original first enhancement layer.
[0419] In a possible implementation, the image block to be encoded refers to the largest coding unit LCU in the entire frame image; or, the image block to be encoded refers to the entire frame image; or, the image block to be encoded refers to the region of interest ROI in the entire frame image.
[0420] In a possible implementation, the first enhancement layer and the second enhancement layer are obtained by layer separation in the quality domain or in the spatial domain.
[0421] In a possible implementation, when the first enhancement layer and the second enhancement layer are obtained by layer separation in the spatial domain, the obtaining module 1601 is further configured to downsample the original image block to be encoded to obtain the image block to be encoded with the second resolution; and upsample the reconstructed block of the original first enhancement layer to obtain the reconstructed block of the first enhancement layer with the second resolution.
[0422] In a possible implementation, the obtaining module 1601 is specifically configured to calculate the difference between corresponding pixel points in the to-be-processed image block and the reconstruction block of the third layer of the to-be-processed image block to obtain the residual block of the first enhancement layer, where the resolution of the image block of the first enhancement layer is not lower than that of the image block of the third layer, or the coding quality of the image block of the first enhancement layer is not lower than that of the image block of the third layer; perform transformation and quantization on the residual block of the first enhancement layer to obtain the quantization coefficients of the residual block of the first enhancement layer; perform inverse quantization and inverse transformation on the quantization coefficients of the residual block of the first enhancement layer to obtain the reconstructed residual block of the first enhancement layer; and sum the corresponding pixel points in the reconstruction block of the third layer and the reconstructed residual block of the first enhancement layer to obtain the reconstruction block of the first enhancement layer.
[0423] The device of this embodiment can be used to execute Figure 8 or Figure 11 the technical solutions implemented by the encoder in the method embodiment shown, and the implementation principles and technical effects are similar, which will not be elaborated here.
[0424] Figure 17 This is a schematic structural diagram of an exemplary decoding device of the present application. As Figure 17 shown, the device of this embodiment can correspond to the video decoder 30. The device may include: an obtaining module 1701, a decoding module 1702, and a reconstruction module 1703. Among them,
[0425] In a possible implementation, the obtaining module 1701 is configured to obtain the bitstream of the residual block of the enhancement layer of the to-be-decoded image block; the decoding module 1702 is configured to perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the enhancement layer to obtain the reconstructed residual block of the enhancement layer; the reconstruction module 1703 is configured to obtain the reconstruction block of the base layer of the to-be-decoded image block, where the resolution of the image block of the enhancement layer is not lower than that of the image block of the base layer, or the coding quality of the image block of the enhancement layer is not lower than that of the image block of the base layer; and sum the corresponding pixel points in the reconstructed residual block of the enhancement layer and the reconstruction block of the base layer to obtain the reconstruction block of the enhancement layer.
[0426] In a possible implementation, the to-be-decoded image block refers to the largest coding unit LCU in the entire frame of image; or, the to-be-decoded image block refers to the entire frame of image; or, the to-be-decoded image block refers to the region of interest ROI in the entire frame of image.
[0427] In a possible implementation, the base layer and the enhancement layer are obtained by layer division in the quality domain or layer division in the spatial domain.
[0428] In a possible implementation, when the base layer and the enhancement layer are obtained by spatial layer division, the reconstruction module 1703 is further configured to upsample the reconstruction blocks of the original base layer to obtain the reconstruction blocks of the base layer with a third resolution, where the third resolution is the same as the resolution of the reconstruction residual blocks of the enhancement layer.
[0429] In a possible implementation, the reconstruction module 1703 is specifically configured to obtain the bitstream of the residual blocks of the base layer; perform entropy decoding on the bitstream of the residual blocks of the base layer to obtain the decoded data of the residual blocks of the base layer; perform inverse quantization and inverse transformation on the decoded data of the residual blocks of the base layer to obtain the reconstructed residual blocks of the base layer; obtain the prediction blocks of the base layer according to the decoded data of the residual blocks of the base layer; and sum the corresponding pixel points in the reconstructed residual blocks of the base layer and the prediction blocks of the base layer to obtain the reconstructed blocks of the base layer.
[0430] In a possible implementation, the acquisition module 1701 is configured to obtain the bitstream of the residual blocks of the second enhancement layer of the image block to be decoded; the decoding module 1702 is configured to perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual blocks of the second enhancement layer to obtain the reconstructed residual blocks of the second enhancement layer; the reconstruction module 1703 is configured to obtain the reconstructed blocks of the first enhancement layer of the image block to be decoded, where the resolution of the second enhancement layer image block is not lower than that of the first enhancement layer image block, or the coding quality of the second enhancement layer image block is not lower than that of the first enhancement layer image block; and sum the corresponding pixel points in the reconstructed residual blocks of the second enhancement layer and the reconstructed blocks of the first enhancement layer to obtain the reconstructed blocks of the second enhancement layer.
[0431] In a possible implementation, the image block to be decoded refers to the largest coding unit (LCU) in the entire frame of image; or, the image block to be decoded refers to the entire frame of image; or, the image block to be decoded refers to the region of interest (ROI) in the entire frame of image.
[0432] In a possible implementation, the first enhancement layer and the second enhancement layer are obtained by layer division in the quality domain or spatial layer division.
[0433] In a possible implementation, when the first enhancement layer and the second enhancement layer are obtained by spatial layer division, the reconstruction module 1703 is further configured to upsample the reconstruction blocks of the original first enhancement layer to obtain the reconstruction blocks of the first enhancement layer with a fourth resolution, where the fourth resolution is the same as the resolution of the reconstruction residual blocks of the second enhancement layer.
[0434] In a possible implementation manner, the reconstruction module 1703 is specifically configured to obtain the bitstream of the residual block of the first enhancement layer; perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the first enhancement layer to obtain the reconstructed residual block of the first enhancement layer; sum the corresponding pixel points in the reconstructed residual block of the first enhancement layer and the reconstructed block of the third layer to obtain the reconstructed block of the first enhancement layer, where the resolution of the image block of the first enhancement layer is not lower than that of the image block of the third layer, or the coding quality of the image block of the first enhancement layer is not lower than that of the image block of the third layer.
[0435] The device in this embodiment can be used to execute Figure 8 or Figure 11 the technical solutions implemented by the decoder in the method embodiments shown. The implementation principles and technical effects are similar and will not be elaborated here.
[0436] In the implementation process, each step of the above method embodiment can be completed by the integrated logic circuit in hardware in the processor or instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by a hardware encoding processor, or executed and completed by a combination of hardware and software modules in the encoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.
[0437] The memory mentioned in the above embodiments may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include but not be limited to these and any other suitable types of memory.
[0438] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0439] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0440] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0441] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0442] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0443] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0444] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An enhanced layer encoding method, characterized in that, Including: Obtaining a reconstructed block of a base layer of an image block to be encoded; Calculating the difference between corresponding pixel points in the image block to be encoded and the reconstructed block of the base layer to obtain a residual block of an enhancement layer of the image block to be encoded, where the resolution of the enhancement layer image block is not lower than that of the base layer image block, or the coding quality of the enhancement layer image block is not lower than that of the base layer image block; Using an adaptive transform unit partitioning method to determine a transform block partitioning manner of the residual block of the enhancement layer, where the transform block partitioning manner of the residual block of the enhancement layer is different from that of the residual block of the base layer, and the transform block partitioning manner of the residual block of the enhancement layer is independent of that of the residual block of the base layer; Transforming the residual block of the enhancement layer according to the transform block partitioning manner to obtain a bitstream of the residual block of the enhancement layer; Among them, obtaining the reconstructed block of the base layer of the image block to be encoded includes: performing prediction on the image block to be encoded, and the prediction includes inter-frame prediction.
2. The method according to claim 1, wherein Determining the transform block partitioning manner of the residual block of the enhancement layer includes: Performing iterative tree-structured partitioning on a first largest transform unit (LTU) of the residual block of the enhancement layer to obtain transform units (TUs) with multiple segmentation depths, where the maximum segmentation depth among the multiple segmentation depths is equal to Dmax, and Dmax is a positive integer; Determining the partitioning manner of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, where the segmentation depth of the first TU is i, and the depth of the second TU is i + 1, and 0 ≤ i ≤ Dmax - 1.
3. The method according to claim 2, wherein The determining the partitioning manner of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU includes: Performing transformation and quantization on the TU to obtain quantization coefficients of the TU, where the TU is the first TU or the second TU; Performing precoding on the quantization coefficients of the TU to obtain the codeword length of the TU; Performing inverse quantization and inverse transformation on the quantization coefficients of the TU to obtain a reconstructed block of the TU; Calculating the sum of squared differences (SSD) between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; Obtaining the loss estimation value of the TU according to the codeword length and the distortion value of the TU; Determining the smaller value between the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; Determining the partitioning manner corresponding to the smaller value as the partitioning manner of the first TU.
4. The method according to claim 2 or 3, characterized in that, The size of the LTU is the same as the size of the reconstructed block of the base layer; or, the size of the LTU is the same as the size of the upsampled original reconstructed block of the base layer.
5. The method according to any one of claims 1 to 3, characterized in that The image block to be encoded refers to the largest coding unit (LCU) in a whole frame of image; or, the image block to be encoded refers to a whole frame of image; or, the image block to be encoded refers to the region of interest (ROI) in a whole frame of image.
6. The method according to any one of claims 1-3, characterized in that, The base layer and the enhancement layer are obtained by layer division based on resolution or coding quality.
7. The method according to any one of claims 1 to 3, characterized in that, Before obtaining the residual block of the enhancement layer of the image block to be encoded by taking the difference between the corresponding pixel points of the image block to be encoded and the reconstructed block of the base layer when the base layer and the enhancement layer are obtained by hierarchical division based on resolution, the method further includes: Downsampling the original image block to be encoded to obtain the image block to be encoded at the first resolution; Upsampling the reconstructed block of the original base layer to obtain the reconstructed block of the base layer at the first resolution.
8. The method according to any one of claims 1 to 3, characterized in that, The obtaining of the reconstructed block of the base layer of the image block to be encoded includes: Taking the difference between the corresponding pixel points of the image block to be encoded and the predicted block of the image block to be encoded to obtain the residual block of the image block to be encoded; Performing transformation and quantization on the residual block of the image block to be encoded to obtain the quantization coefficients of the image block to be encoded; Performing inverse quantization and inverse transformation on the quantization coefficients of the image block to be encoded to obtain the reconstructed residual block of the image block to be encoded; Summing the corresponding pixel points of the predicted block of the image block to be encoded and the reconstructed residual block of the image block to be encoded to obtain the reconstructed block of the base layer.
9. An enhanced layer encoding method, characterized in that, Includes: Obtaining the reconstructed block of the first enhancement layer of the image block to be encoded; Taking the difference between the corresponding pixel points of the image block to be encoded and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the image block to be encoded, where the resolution of the second enhancement layer image block is not lower than that of the first enhancement layer image block, or the coding quality of the second enhancement layer image block is not lower than that of the first enhancement layer image block; Adopting an adaptive transformation unit division method to determine the transformation block division method of the residual block of the second enhancement layer, where the transformation block division method of the residual block of the second enhancement layer is different from that of the residual block of the first enhancement layer and that of the residual block of the base layer, and the transformation block division method of the residual block of the second enhancement layer is independent of that of the residual block of the first enhancement layer and that of the residual block of the base layer, and the resolution of the first enhancement layer image block is not lower than that of the base layer image block, or the coding quality of the first enhancement layer image block is not lower than that of the base layer image block; Performing transformation on the residual block of the second enhancement layer according to the transformation block division method to obtain the bitstream of the residual block of the second enhancement layer.
10. The method according to claim 9, wherein The determining of the transformation block division method of the residual block of the second enhancement layer includes: Performing iterative tree-structured division on the first largest transformation unit (LTU) of the residual block of the second enhancement layer to obtain transformation units (TUs) with multiple segmentation depths, where the maximum segmentation depth among the multiple segmentation depths is equal to Dmax, and Dmax is a positive integer; Determining the division method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, where the segmentation depth of the first TU is i, and the depth of the second TU is i + 1, and 0 ≤ i ≤ Dmax - 1.
11. The method according to claim 10, characterized in that, Determining the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU includes: Performing transformation and quantization on the TU to obtain quantization coefficients of the TU, where the TU is the first TU or the second TU; Performing precoding on the quantization coefficients of the TU to obtain the codeword length of the TU; Performing inverse quantization and inverse transformation on the quantization coefficients of the TU to obtain the reconstructed block of the TU; Calculating the sum of squared differences SSD between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; Obtaining the loss estimation value of the TU according to the codeword length of the TU and the distortion value of the TU; Determining the smaller value between the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; Determining the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
12. The method according to claim 10 or 11, characterized in that, The size of the LTU is the same as the size of the reconstructed block of the first enhancement layer; or, the size of the LTU is the same as the size after upsampling the reconstructed block of the original first enhancement layer.
13. The method according to any one of claims 9-11, characterized in that, The image block to be encoded refers to the largest coding unit LCU in the entire frame of the image; or, the image block to be encoded refers to the entire frame of the image; or, the image block to be encoded refers to the region of interest ROI in the entire frame of the image.
14. The method according to any one of claims 9-11, characterized in that, The first enhancement layer and the second enhancement layer are obtained by layer division in the quality domain or the spatial domain.
15. The method according to any one of claims 9-11, characterized in that, When the first enhancement layer and the second enhancement layer are obtained by layer division in the spatial domain, before obtaining the residual block of the second enhancement layer of the image block to be encoded by taking the difference between the corresponding pixel points in the image block to be encoded and the reconstructed block of the first enhancement layer, it further includes: Performing downsampling on the original image block to be encoded to obtain the image block to be encoded with the second resolution; Performing upsampling on the reconstructed block of the original first enhancement layer to obtain the reconstructed block of the first enhancement layer with the second resolution.
16. The method according to any one of claims 9-11, characterized in that, Obtaining the reconstructed block of the first enhancement layer of the image block to be encoded includes: Taking the difference between the corresponding pixel points in the image block to be encoded and the reconstructed block of the third layer of the image block to be encoded to obtain the residual block of the first enhancement layer, where the resolution of the first enhancement layer image block is not lower than the resolution of the third layer image block, or the coding quality of the first enhancement layer image block is not lower than the coding quality of the third layer image block; Performing transformation and quantization on the residual block of the first enhancement layer to obtain quantization coefficients of the residual block of the first enhancement layer; Performing inverse quantization and inverse transformation on the quantization coefficients of the residual block of the first enhancement layer to obtain the reconstructed residual block of the first enhancement layer; Taking the sum of the corresponding pixel points in the reconstructed block of the third layer and the reconstructed residual block of the first enhancement layer to obtain the reconstructed block of the first enhancement layer.
17. An enhanced layer decoding method, characterized in that, Includes: Obtaining the bitstream of the residual block of the enhancement layer of the image block to be decoded; Performing entropy decoding, inverse quantization and inverse transformation on the bitstream of the residual block of the enhancement layer to obtain the reconstructed residual block of the enhancement layer; Obtain the reconstructed block of the base layer of the image block to be decoded. The resolution of the enhanced layer image block is not lower than that of the base layer image block, or the coding quality of the enhanced layer image block is not lower than that of the base layer image block. The transform block partitioning method of the residual block in the enhanced layer is different from that of the residual block in the base layer. The transform block partitioning method of the residual block in the enhanced layer is determined by using an adaptive transform unit partitioning method, and the transform block partitioning method of the residual block in the enhanced layer is independent of the transform block partitioning method of the residual block in the base layer; Sum the corresponding pixel points in the reconstructed residual block of the enhanced layer and the reconstructed block of the base layer to obtain the reconstructed block of the enhanced layer; Among them, obtaining the reconstructed block of the base layer of the image block to be decoded includes: performing prediction on the image block to be decoded, and the prediction includes inter-frame prediction.
18. The method according to claim 17, wherein The image block to be decoded refers to the largest coding unit (LCU) in the entire frame image; or, the image block to be decoded refers to the entire frame image; or, the image block to be decoded refers to the region of interest (ROI) in the entire frame image.
19. The method according to claim 17 or 18, characterized in that, The base layer and the enhanced layer are obtained by layer splitting in the quality domain or the spatial domain.
20. The method according to claim 17 or 18, characterized in that, When the base layer and the enhanced layer are obtained by layer splitting in the spatial domain, before summing the corresponding pixel points in the reconstructed residual block of the enhanced layer and the reconstructed block of the base layer to obtain the reconstructed block of the enhanced layer, it further includes: Upsample the reconstructed block of the original base layer to obtain the reconstructed block of the base layer with the third resolution, and the third resolution is the same as the resolution of the reconstructed residual block of the enhanced layer.
21. The method according to claim 17 or 18, characterized in that, Obtaining the reconstructed block of the base layer of the image block to be decoded includes: Obtain the bitstream of the residual block of the base layer; Perform entropy decoding on the bitstream of the residual block of the base layer to obtain the decoded data of the residual block of the base layer; Perform inverse quantization and inverse transformation on the decoded data of the residual block of the base layer to obtain the reconstructed residual block of the base layer; Obtain the predicted block of the base layer according to the decoded data of the residual block of the base layer; Sum the corresponding pixel points in the reconstructed residual block of the base layer and the predicted block of the base layer to obtain the reconstructed block of the base layer.
22. A method for enhanced layer decoding, characterized in that, Include: Obtain the bitstream of the residual block of the second enhanced layer of the image block to be decoded; Perform entropy decoding, inverse quantization and inverse transformation on the bitstream of the residual block of the second enhanced layer to obtain the reconstructed residual block of the second enhanced layer; Obtain the reconstructed block of the first enhanced layer of the image block to be decoded. The resolution of the second enhanced layer image block is not lower than that of the first enhanced layer image block, or the coding quality of the second enhanced layer image block is not lower than that of the first enhanced layer image block; Sum the corresponding pixel points in the reconstruction residual block of the second enhancement layer and the reconstruction block of the first enhancement layer to obtain the reconstruction block of the second enhancement layer. The transform block partitioning method of the residual block of the second enhancement layer is different from that of the residual block of the first enhancement layer and the transform block partitioning method of the residual block of the base layer. The transform block partitioning method of the residual block of the second enhancement layer is determined by using an adaptive transform unit partitioning method. The transform block partitioning method of the residual block of the second enhancement layer is independent of the transform block partitioning method of the residual block of the first enhancement layer and the transform block partitioning method of the residual block of the base layer. The resolution of the image block of the first enhancement layer is not lower than that of the image block of the base layer, or the coding quality of the image block of the first enhancement layer is not lower than that of the image block of the base layer.
23. The method according to claim 22, wherein The image block to be decoded refers to the largest coding unit (LCU) in the entire frame image; or, the image block to be decoded refers to the entire frame image; or, the image block to be decoded refers to the region of interest (ROI) in the entire frame image.
24. The method according to claim 22 or 23, characterized in that, The first enhancement layer and the second enhancement layer are obtained by layer splitting in the quality domain or the spatial domain.
25. The method according to claim 22 or 23, characterized in that, When the first enhancement layer and the second enhancement layer are obtained by layer splitting in the spatial domain, before summing the corresponding pixel points in the reconstruction residual block of the second enhancement layer and the reconstruction block of the first enhancement layer to obtain the reconstruction block of the second enhancement layer, it further includes: Upsample the reconstruction block of the original first enhancement layer to obtain the reconstruction block of the first enhancement layer with the fourth resolution, where the fourth resolution is the same as the resolution of the reconstruction residual block of the second enhancement layer.
26. The method according to claim 22 or 23, characterized in that The obtaining of the reconstruction block of the first enhancement layer of the image block to be decoded includes: Obtain the bitstream of the residual block of the first enhancement layer; Perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the first enhancement layer to obtain the reconstruction residual block of the first enhancement layer; Sum the corresponding pixel points in the reconstruction residual block of the first enhancement layer and the reconstruction block of the third layer to obtain the reconstruction block of the first enhancement layer. The resolution of the image block of the first enhancement layer is not lower than that of the image block of the third layer, or the coding quality of the image block of the first enhancement layer is not lower than that of the image block of the third layer.
27. A coding device, characterized in that, It includes: An acquisition module, configured to acquire the reconstruction block of the base layer of the image block to be encoded; Subtract the corresponding pixel points in the image block to be encoded and the reconstruction block of the base layer to obtain the residual block of the enhancement layer of the image block to be encoded. The resolution of the image block of the enhancement layer is not lower than that of the image block of the base layer, or the coding quality of the image block of the enhancement layer is not lower than that of the image block of the base layer; A determination module, configured to use an adaptive transform unit partitioning method to determine the transform block partitioning method of the residual block of the enhancement layer. The transform block partitioning method of the residual block of the enhancement layer is different from that of the residual block of the base layer, and the transform block partitioning method of the residual block of the enhancement layer is independent of the transform block partitioning method of the residual block of the base layer; An encoding module, configured to transform the residual blocks of the enhancement layer according to the transform block partitioning method to obtain the bitstream of the residual blocks of the enhancement layer; The obtaining module is specifically configured to perform prediction on the image block to be encoded, and the prediction includes inter-frame prediction.
28. The device according to claim 27, characterized in that, The determining module is specifically configured to perform iterative tree-structured partitioning on the first largest transform unit (LTU) of the residual blocks of the enhancement layer to obtain transform units (TUs) with multiple segmentation depths, where the maximum segmentation depth among the multiple segmentation depths is equal to Dmax, and Dmax is a positive integer; determine the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU, the segmentation depth of the first TU is i, and the depth of the second TU is i + 1, 0 ≤ i ≤ Dmax - 1.
29. The device according to claim 28, characterized in that, The determining module is specifically configured to perform transformation and quantization on the TU to obtain the quantization coefficients of the TU, where the TU is the first TU or the second TU; perform precoding on the quantization coefficients of the TU to obtain the codeword length of the TU; perform inverse quantization and inverse transformation on the quantization coefficients of the TU to obtain the reconstructed block of the TU; calculate the sum of squared differences (SSD) between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; obtain the loss estimation value of the TU according to the codeword length and the distortion value of the TU; determine the smaller value among the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; and determine the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
30. The device according to claim 28 or 29, characterized in that, The size of the LTU is the same as the size of the reconstructed block of the base layer; or, the size of the LTU is the same as the size of the reconstructed block of the original base layer after upsampling.
31. The device according to any one of claims 27-29, characterized in that, The image block to be encoded refers to the largest coding unit (LCU) in the entire frame of the image; or, the image block to be encoded refers to the entire frame of the image; or, the image block to be encoded refers to the region of interest (ROI) in the entire frame of the image.
32. The device according to any one of claims 27 - 29, characterized in that, The base layer and the enhancement layer are obtained by layer division based on resolution or coding quality.
33. The device according to any one of claims 27-29, characterized in that, When the base layer and the enhancement layer are obtained by layer division based on resolution, the obtaining module is further configured to downsample the original image block to be encoded to obtain the image block to be encoded with the first resolution; and upsample the reconstructed block of the original base layer to obtain the reconstructed block of the base layer with the first resolution.
34. The device according to any one of claims 27 - 29, characterized in that, The obtaining module is specifically configured to calculate the difference between corresponding pixel points in the image block to be encoded and the predicted block of the image block to be encoded to obtain the residual block of the image block to be encoded; perform transformation and quantization on the residual block of the image block to be encoded to obtain the quantization coefficients of the image block to be encoded; perform inverse quantization and inverse transformation on the quantization coefficients of the image block to be encoded to obtain the reconstructed residual block of the image block to be encoded; and calculate the sum of corresponding pixel points in the predicted block of the image block to be encoded and the reconstructed residual block of the image block to be encoded to obtain the reconstructed block of the base layer.
35. A coding device, characterized in that, Including: An obtaining module, configured to obtain the reconstructed block of the first enhancement layer of the image block to be encoded; Calculate the difference between the corresponding pixel points in the image block to be encoded and the reconstructed block of the first enhancement layer to obtain the residual block of the second enhancement layer of the image block to be encoded. The resolution of the second enhancement layer image block is not lower than that of the first enhancement layer image block, or the coding quality of the second enhancement layer image block is not lower than that of the first enhancement layer image block. A determination module, configured to determine the transformation block partitioning method of the residual block of the second enhancement layer by using an adaptive transformation unit partitioning method. The transformation block partitioning method of the residual block of the second enhancement layer is different from the transformation block partitioning method of the residual block of the first enhancement layer and the transformation block partitioning method of the residual block of the base layer. The transformation block partitioning method of the residual block of the second enhancement layer is independent of the transformation block partitioning method of the residual block of the first enhancement layer and the transformation block partitioning method of the residual block of the base layer. The resolution of the first enhancement layer image block is not lower than that of the base layer image block, or the coding quality of the first enhancement layer image block is not lower than that of the base layer image block. An encoding module, configured to perform transformation on the residual block of the second enhancement layer according to the transformation block partitioning method to obtain the bitstream of the residual block of the second enhancement layer.
36. The device according to claim 35, characterized in that, Specifically, the determination module is configured to perform iterative tree-structured partitioning on the first largest transformation unit (LTU) of the residual block of the second enhancement layer to obtain transformation units (TUs) with multiple segmentation depths. The maximum segmentation depth among the multiple segmentation depths is equal to Dmax, where Dmax is a positive integer. Determine the partitioning method of the first TU according to the loss estimation value of the first TU and the sum of the loss estimation values of multiple second TUs included in the first TU. The segmentation depth of the first TU is i, and the depth of the second TU is i + 1, where 0 ≤ i ≤ Dmax - 1.
37. The device according to claim 36, characterized in that, Specifically, the determination module is configured to perform transformation and quantization on the TU to obtain the quantization coefficient of the TU, where the TU is the first TU or the second TU; perform precoding on the quantization coefficient of the TU to obtain the codeword length of the TU; perform inverse quantization and inverse transformation on the quantization coefficient of the TU to obtain the reconstructed block of the TU; calculate the sum of squared differences (SSD) between the TU and the reconstructed block of the TU to obtain the distortion value of the TU; obtain the loss estimation value of the TU according to the codeword length and the distortion value of the TU; determine the smaller value between the loss estimation value of the first TU and the sum of the loss estimation values of the multiple second TUs; and determine the partitioning method corresponding to the smaller value as the partitioning method of the first TU.
38. The device according to claim 36 or 37, characterized in that The size of the LTU is the same as the size of the reconstructed block of the first enhancement layer; or, the size of the LTU is the same as the size of the upsampled reconstructed block of the original first enhancement layer.
39. The device according to any one of claims 35 - 37, characterized in that, The image block to be encoded refers to the largest coding unit (LCU) in the entire frame image; or, the image block to be encoded refers to the entire frame image; or, the image block to be encoded refers to the region of interest (ROI) in the entire frame image.
40. The device according to any one of claims 35-37, characterized in that, The first enhancement layer and the second enhancement layer are obtained by layer division in the quality domain or in the spatial domain.
41. The device according to any one of claims 35 - 37, characterized in that, When the first enhancement layer and the second enhancement layer are obtained by layer division in the spatial domain, the obtaining module is further configured to downsample the original image block to be encoded to obtain the image block to be encoded with a second resolution; and upsample the reconstructed block of the original first enhancement layer to obtain the reconstructed block of the first enhancement layer with the second resolution.
42. The apparatus according to any one of claims 35 - 37, characterized in that, The obtaining module is specifically configured to calculate the difference between corresponding pixel points of the image block to be encoded and the reconstructed block of the third layer of the image block to be encoded to obtain the residual block of the first enhancement layer, where the resolution of the image block of the first enhancement layer is not lower than the resolution of the image block of the third layer, or the coding quality of the image block of the first enhancement layer is not lower than the coding quality of the image block of the third layer; perform transformation and quantization on the residual block of the first enhancement layer to obtain the quantization coefficients of the residual block of the first enhancement layer; perform inverse quantization and inverse transformation on the quantization coefficients of the residual block of the first enhancement layer to obtain the reconstructed residual block of the first enhancement layer; and calculate the sum of corresponding pixel points of the reconstructed block of the third layer and the reconstructed residual block of the first enhancement layer to obtain the reconstructed block of the first enhancement layer.
43. A decoding device, characterized in that, Comprising: An obtaining module, configured to obtain the bitstream of the residual block of the enhancement layer of the image block to be decoded; A decoding module, configured to perform entropy decoding, inverse quantization and inverse transformation on the bitstream of the residual block of the enhancement layer to obtain the reconstructed residual block of the enhancement layer; A reconstruction module, configured to obtain the reconstructed block of the base layer of the image block to be decoded, where the resolution of the image block of the enhancement layer is not lower than the resolution of the image block of the base layer, or the coding quality of the image block of the enhancement layer is not lower than the coding quality of the image block of the base layer; calculate the sum of corresponding pixel points of the reconstructed residual block of the enhancement layer and the reconstructed block of the base layer to obtain the reconstructed block of the enhancement layer, where the transform block partitioning method of the residual block of the enhancement layer is different from the transform block partitioning method of the residual block of the base layer, the transform block partitioning method of the residual block of the enhancement layer is determined by using an adaptive transform unit partitioning method, and the transform block partitioning method of the residual block of the enhancement layer is independent of the transform block partitioning method of the residual block of the base layer; The reconstruction module is specifically configured to perform prediction on the image block to be decoded, and the prediction includes inter-frame prediction.
44. The device according to claim 43, characterized in that, The image block to be decoded refers to the largest coding unit (LCU) in the entire frame of image; or, the image block to be decoded refers to the entire frame of image; or, the image block to be decoded refers to the region of interest (ROI) in the entire frame of image.
45. The device according to claim 43 or 44, characterized in that, The base layer and the enhancement layer are obtained by layer division in the quality domain or in the spatial domain.
46. The device according to claim 43 or 44, characterized in that, When the base layer and the enhancement layer are obtained by layer division in the spatial domain, the reconstruction module is further configured to upsample the reconstructed block of the original base layer to obtain the reconstructed block of the base layer with a third resolution, where the third resolution is the same as the resolution of the reconstructed residual block of the enhancement layer.
47. The device according to claim 43 or 44, characterized in that The reconstruction module is specifically configured to obtain the bitstream of the residual block of the base layer; perform entropy decoding on the bitstream of the residual block of the base layer to obtain the decoded data of the residual block of the base layer; perform inverse quantization and inverse transformation on the decoded data of the residual block of the base layer to obtain the reconstructed residual block of the base layer; obtain the prediction block of the base layer according to the decoded data of the residual block of the base layer; sum the corresponding pixel points in the reconstructed residual block of the base layer and the prediction block of the base layer to obtain the reconstructed block of the base layer.
48. A decoding device, characterized in that, It includes: An acquisition module, configured to obtain the bitstream of the residual block of the second enhancement layer of the image block to be decoded; A decoding module, configured to perform entropy decoding, inverse quantization and inverse transformation on the bitstream of the residual block of the second enhancement layer to obtain the reconstructed residual block of the second enhancement layer; A reconstruction module, configured to obtain the reconstructed block of the first enhancement layer of the image block to be decoded, the resolution of the second enhancement layer image block is not lower than the resolution of the first enhancement layer image block, or the coding quality of the second enhancement layer image block is not lower than the coding quality of the first enhancement layer image block; sum the corresponding pixel points in the reconstructed residual block of the second enhancement layer and the reconstructed block of the first enhancement layer to obtain the reconstructed block of the second enhancement layer, the transform block partitioning method of the residual block of the second enhancement layer is different from the transform block partitioning method of the residual block of the first enhancement layer and the transform block partitioning method of the residual block of the base layer, the transform block partitioning method of the residual block of the second enhancement layer is determined by using an adaptive transform unit partitioning method, the transform block partitioning method of the residual block of the second enhancement layer is independent of the transform block partitioning method of the residual block of the first enhancement layer and the transform block partitioning method of the residual block of the base layer, the resolution of the first enhancement layer image block is not lower than the resolution of the base layer image block, or the coding quality of the first enhancement layer image block is not lower than the coding quality of the base layer image block.
49. The device according to claim 48, characterized in that, The image block to be decoded refers to the largest coding unit (LCU) in the entire frame image; or, the image block to be decoded refers to the entire frame image; or, the image block to be decoded refers to the region of interest (ROI) in the entire frame image. The device according to claim 48 or 49, characterized in that, The first enhancement layer and the second enhancement layer are obtained by layer division in the quality domain or the spatial domain.
51. The device according to claim 48 or 49, characterized in that, When the first enhancement layer and the second enhancement layer are obtained by layer division in the spatial domain, the reconstruction module is further configured to perform upsampling on the reconstructed block of the original first enhancement layer to obtain the reconstructed block of the first enhancement layer with the fourth resolution, and the fourth resolution is the same as the resolution of the reconstructed residual block of the second enhancement layer.
52. The device according to claim 48 or 49, characterized in that, The reconstruction module is specifically configured to obtain the bitstream of the residual block of the first enhancement layer; perform entropy decoding, inverse quantization, and inverse transformation on the bitstream of the residual block of the first enhancement layer to obtain the reconstructed residual block of the first enhancement layer; sum the corresponding pixel points in the reconstructed residual block of the first enhancement layer and the reconstructed block of the third layer to obtain the reconstructed block of the first enhancement layer, where the resolution of the image block of the first enhancement layer is not lower than that of the image block of the third layer, or the coding quality of the image block of the first enhancement layer is not lower than that of the image block of the third layer.
53. An encoder, characterized in that, Comprising: a processor and a memory; the processor is coupled to the memory, and computer-readable instructions are stored in the memory; the processor is configured to read the computer-readable instructions to enable the encoder to implement the method according to any one of claims 1-16.
54. A decoder, characterized in that, Comprising: a processor and a memory; the processor is coupled to the memory, and computer-readable instructions are stored in the memory; the processor is configured to read the computer-readable instructions to enable the decoder to implement the method according to any one of claims 17-26.
55. A computer program product, characterized in that, Comprising program code which, when executed on a computer or a processor, is used to execute the method according to any one of claims 1-26.
56. A computer-readable storage medium, characterized in that, Comprising program code which, when executed by a computer device, is used to execute the method according to any one of claims 1-26.
Citation Information
Patent Citations
HEVC (High Efficiency Video Coding)-oriented quality scalable inter-layer prediction coding
CN103281531A
Conversion unit diversion method and device
CN104902276A
Method and devices for encoding a sequence of images into a scalable video bit-stream, and decoding a corresponding scalable video bit-stream
WO2013128010A2
Cited By
Encoding and decoding method and apparatus for enhancement layer
WO2022121770A1