Video processing method, video processing device, encoder, decoder, medium, and computer program

The implementation of a history-based motion vector prediction list for coding tree units addresses the challenge of further compression efficiency in video coding, enhancing encoding and decoding performance by initializing and updating the list for improved CTU processing.

JP7727618B2Active Publication Date: 2025-08-21HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022212116
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-08-10
Filing Date
2022-12-28
Publication Date
2025-08-21
Estimated Expiration
2039-08-12

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in achieving further compression efficiency beyond High Efficiency Video Coding (HEVC) without sacrificing picture quality, particularly in handling motion vector prediction for coding tree units.

Method used

Implementing a history-based motion vector prediction (HMVP) list initialization for current coding tree units (CTUs) at the start of processing, allowing for improved coding and decoding efficiency, especially in wavefront parallel processing mode, and updating the HMVP list based on processed CTUs.

Benefits of technology

Enhances encoding and decoding efficiency by initializing the HMVP list for CTUs, enabling simultaneous processing of CTU rows in a picture frame, thereby improving compression efficiency beyond HEVC standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007727618000002
    Figure 0007727618000002
  • Figure 0007727618000003
    Figure 0007727618000003
  • Figure 0007727618000004
    Figure 0007727618000004
Patent Text Reader

Abstract

A video processing method and corresponding apparatus are provided for improving coding efficiency. A video processing method includes the steps of: initializing an HMVP list for a current CTU row when the current CTU is the start CTU of the current CTU row; and processing the current CTU row based on the HMVP list, thereby improving coding and decoding efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD Embodiments of the present application (disclosure) relate generally to the field of video coding, and more particularly to video processing methods, video processing devices, encoders, decoders, media, and computer programs. [Background technology]

[0002] Video coding (video encoding and video decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat, video conferencing, DVDs and Blu-ray® discs, video content acquisition and editing systems, and camcorders for security applications.

[0003] Since the development of the block-based hybrid video coding technique in the H.261 standard in 1990, new video coding techniques and tools have been developed to form the basis of new video coding standards. Additional video coding standards include MPEG-1 video, MPEG-2 video, ITU-T H.262 / MPEG-2, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), ITU-T H.266 / Versatile Video Coding (VVC), and extensions to these standards, such as scalability and / or three-dimensional (3D) extensions. As video production and use become increasingly ubiquitous, video traffic places a significant burden on communication networks and data storage; therefore, one of the goals of many video coding standards has been to achieve a reduction in bit rate compared to their predecessors without sacrificing picture quality. Even the latest High Efficiency Video Coding (HEVC) can compress video about twice as much as AVC without sacrificing quality, so there is a need to further compress video compared to HEVC. Summary of the Invention [Means for solving the problem]

[0004] SUMMARY OF THE INVENTION The embodiments of the present application provide a video processing method and corresponding apparatus for improving coding efficiency.

[0005] These and other objects are achieved by the subject matter of the independent claims. Further implementations are evident from the independent claims, the detailed description and the figures.

[0006] A first aspect of the present invention provides a video processing method, the video processing method including the steps of: initializing a history-based motion vector prediction (HMVP) list for a current coding tree unit (CTU) row when the current CTU is a start CTU of the current CTU row; and processing the current CTU row based on the HMVP list. The start CTU, sometimes called a starting CTU, is the first CTU of the CTUs in the same CTU row being processed.

[0007] It can be seen that the HMVP list for the current CTU row is initialized at the start of processing the current CTU row, and the process of the current CTU row does not need to be based on the HMVP list of the previous CTU row, thereby improving coding and decoding efficiency.

[0008] Referring to the first aspect, in a first possible implementation method of the first aspect, the number of candidate motion vectors in the initialized HMVP list is zero.

[0009] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a second possible implementation method of the first aspect, the current CTU row belongs to a picture area consisting of multiple CTU rows, and the current CTU row is any one of the multiple CTU rows, for example, the first (e.g., top) CTU row, the second CTU row, ... and the last (e.g., bottom) CTU row of the picture area.

[0010]

[0023] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a third possible implementation method of the first aspect, the method further includes a step of initializing an HMVP list for each of a plurality of CTU rows, excluding a current CTU row, where the HMVP lists for the plurality of CTU rows are the same or different. In other words, the embodiment may additionally initialize the HMVP lists for all other CTU rows in the picture area, i.e., initialize the HMVP lists for all CTU rows in the picture area.

[0011] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a fourth possible implementation method of the first aspect, the step of processing the current CTU row based on the HMVP list includes the steps of processing a current CTU in the current CTU row, updating the initialized HMVP list based on the processed current CTU, and processing a second CTU in the current CTU row based on the updated HMVP list.

[0012] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a fifth possible implementation method of the first aspect, the HMVP list is updated according to the processed CTU of the current CTU row.

[0013] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a sixth possible implementation method of the first aspect, the HMVP list for the current CTU row is initialized as follows, i.e., to empty the HMVP list for the current CTU row.

[0014] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a seventh possible implementation method of the first aspect, the step of processing the current CTU row based on the HMVP list includes a step of processing the current CTU row based on an HMVP list from a second CTU in the current CTU row, where the second CTU is adjacent to the starting CTU.

[0015] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in an eighth possible implementation method of the first aspect, multiple CTU rows are processed in wavefront parallel processing (WPP) mode.

[0016] It can be seen that, since the HMVP list for the current CTU row is initialized at the start of processing the current CTU row, when combined with WPP mode, the CTU rows of a picture frame or picture area can be processed simultaneously, thereby further improving encoding and decoding efficiency.

[0017] Referring to the first aspect or any one of the above-mentioned implementation methods of the first aspect, in a ninth possible implementation method of the first aspect, the current CTU row begins to be processed (or processing of the current CTU row begins) when a specific CTU of the previous CTU row is processed.

[0018] Referring to the first aspect or any one of the aforementioned implementation methods of the first aspect, in a tenth possible implementation method of the first aspect, the previous CTU row is a CTU row that is directly adjacent to the current CTU row and is above or above the current CTU row.

[0019] Referring to the ninth implementation method of the first aspect or the tenth implementation method of the first aspect, in an eleventh possible implementation form of the first aspect, a particular CTU in the previous CTU row is the second CTU in the previous CTU row; or, a particular CTU in the previous CTU row is the first CTU in the previous CTU row.

[0020] A second aspect of the present invention provides a video processing device including: an initialization unit configured to initialize a history-based motion vector prediction (HMVP) list for a current coding tree unit (CTU) row when the CTU is the start CTU of the current CTU row; and a processing unit configured to process the current CTU row based on the HMVP list.

[0021] Referring to the second aspect of the present invention, in a first possible implementation method of the second aspect, the number of candidate motion vectors in the initialized HMVP list is zero.

[0022] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in a second possible implementation method of the second aspect, the current CTU row belongs to a picture area consisting of multiple CTU rows, and the current CTU row is any one of the multiple CTU rows.

[0023] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in a third possible implementation method of the second aspect, the initialization unit is further configured to initialize an HMVP list for each of a plurality of CTU rows, excluding the current CTU row, wherein the HMVP lists for the plurality of CTU rows are the same or different.

[0024] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in a fourth possible implementation method of the second aspect, the processing unit is further configured to process a current CTU of a current CTU row, update the initialized HMVP list based on the processed current CTU, and process a second CTU of the current CTU row based on the updated HMVP list.

[0025] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in a fifth possible implementation method of the second aspect, the HMVP list is updated according to the processed CTU of the current CTU row.

[0026] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in a sixth possible implementation method of the second aspect, the initialization unit is further configured to initialize the HMVP list for the current CTU row as follows, i.e., to empty the HMVP list for the current CTU row.

[0027] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in a seventh possible implementation method of the second aspect, the processing unit is further configured to process the current CTU row based on the HMVP list as follows: processing the current CTU row based on the HMVP list from a second CTU in the current CTU row, where the second CTU is adjacent to the start CTU.

[0028] Referring to the second aspect or any one of the above-mentioned implementation methods of the second aspect, in an eighth possible implementation method of the second aspect, multiple CTU rows are processed in wavefront parallel processing (WPP) mode.

[0029] Referring to the second aspect, or any one of the aforementioned implementation methods of the second aspect, in a ninth possible implementation method of the second aspect, the current CTU row begins to be processed (or processing of the current CTU row begins) when a specific CTU of the previous CTU row is processed.

[0030] Referring to the second aspect or any one of the aforementioned implementation methods of the second aspect, in a tenth possible implementation method of the second aspect, the previous CTU row is a CTU row that is directly adjacent to and above the current CTU row.

[0031] With reference to the ninth implementation method of the second aspect or the tenth implementation method of the second aspect, in an eleventh possible implementation method of the second aspect, a specific CTU in the previous CTU row is the second CTU in the previous CTU row; or a specific CTU in the previous CTU row is the first CTU in the previous CTU row.

[0032] A third aspect of the present invention provides a coding method implemented by a decoding device, the coding method including the steps of: constructing / initializing an HMVP list for a current CTU row; and processing CTUs in the current CTU row based on the constructed / initialized HMVP list. Referring to the third aspect, in a first possible implementation method of the third aspect, the HMVP list for the current CTU row is constructed / initialized as follows: to empty the HMVP list for the current CTU row, and / or to set a default value for the HMVP list for the current CTU row, and / or to construct / initialize the HMVP list for the current CTU row based on the HMVP list of the CTU of the previous CTU row.

[0033] Referring to the first possible implementation method of the third aspect, in a second possible implementation method of the third aspect, the step of setting default values ​​for the HMVP list for the current CTU row includes a step of populating the MVs of the HMVP list as MVs of a uni-predictive method, where the MVs of the uni-predictive method are either zero motion vectors or are not zero motion vectors, and the reference pictures include the first reference picture in the L0 list, and / or a step of populating the MVs of the HMVP list as MVs of a bi-predictive method, where the MVs of the bi-predictive method are either zero motion vectors or are not zero motion vectors, and the reference pictures include the first reference picture in the L0 list and the first reference picture in the L1 list.

[0034] Referring to the first possible implementation method of the third aspect, in the third possible implementation method of the third aspect, each co-located picture can store a temporal HMVP list for each CTU row or for the entire picture, and the step of setting a default value for the HMVP list for the current CTU row includes a step of initializing / constructing the HMVP list for the current CTU row based on the temporal HMVP list.

[0035] Referring to the first possible implementation method of the third aspect, in a fourth possible implementation method of the third aspect, the previous CTU row is a CTU row that is directly adjacent to and above the current CTU row.

[0036] Referring to the fourth possible implementation method of the third aspect, in a fifth possible implementation method of the third aspect, the CTU in the previous CTU row is the second CTU in the previous CTU row.

[0037] Referring to the fourth possible implementation method of the third aspect, in a fifth possible implementation method of the third aspect, the CTU in the previous CTU row is the first CTU in the previous CTU row.

[0038] A fourth aspect of the present invention provides a coding method implemented by an encoding device, the coding method including the steps of: constructing / initializing an HMVP list for a current CTU row; and processing CTUs in the current CTU row based on the constructed / initialized HMVP list.

[0039] Referring to the fourth aspect, in a first possible implementation method of the fourth aspect, the HMVP list for the current CTU row is constructed / initialized as follows: to empty the HMVP list for the current CTU row; and / or to set a default value for the HMVP list for the current CTU row, and / or to construct / initialize the HMVP list for the current CTU row based on the HMVP list of the CTUs in the previous CTU row.

[0040] Referring to the first possible implementation method of the fourth aspect, in a second possible implementation method of the fourth aspect, the step of setting default values ​​for the HMVP list for the current CTU row includes a step of populating the MVs of the HMVP list as MVs of a uni-predictive method, where the MVs of the uni-predictive method are either zero motion vectors or are not zero motion vectors, and the reference pictures include the first reference picture in the L0 list; and / or a step of populating the MVs of the HMVP list as MVs of a bi-predictive method, where the MVs of the bi-predictive method are either zero motion vectors or are not zero motion vectors, and the reference pictures include the first reference picture in the L0 list and the first reference picture in the L1 list.

[0041] Referring to the first possible implementation method of the fourth aspect, in a third possible implementation method of the fourth aspect, each co-located picture can store a temporal HMVP list for each CTU row or for the entire picture, and the step of setting a default value for the HMVP list for the current CTU row includes a step of initializing / constructing the HMVP list for the current CTU row based on the temporal HMVP list.

[0042] Referring to the first possible implementation method of the fourth aspect, in the fourth possible implementation method of the fourth aspect, the previous CTU row is a CTU row that is directly adjacent to the current CTU row and is above or above the current CTU row.

[0043] Referring to the fourth possible implementation method of the fourth aspect, in a fifth possible implementation method of the fourth aspect, the CTU in the previous CTU row is the second CTU in the previous CTU row.

[0044] Referring to the fourth possible implementation method of the fourth aspect, in a sixth possible implementation method of the fourth aspect, the CTU in the previous CTU row is the first CTU in the previous CTU row.

[0045] A fifth aspect of the present invention provides an encoder including a processing circuit for performing a method according to the first aspect or any one of the implementation methods of the first aspect, or according to the third aspect or any one of the implementation methods of the third aspect, or according to the fourth aspect or any one of the implementation methods of the fourth aspect. For example, the encoder may include an initialization circuit configured to initialize a history-based motion vector prediction (HMVP) list for a current coding tree unit (CTU) row when the current CTU is a start CTU of the current CTU row, and a processing circuit configured to process the current CTU row based on the HMVP list.

[0046] A sixth aspect of the present invention provides a decoder, the decoder including a processing circuit for performing a method according to the first aspect or any one of the implementation methods of the first aspect, according to the third aspect or any one of the implementation methods of the third aspect, or according to the fourth aspect or any one of the implementation methods of the fourth aspect. For example, the decoder may include an initialization circuit configured to initialize a history-based motion vector prediction (HMVP) list for a current coding tree unit (CTU) row when the current CTU is a start CTU of the current CTU row, and a processing circuit configured to process the current CTU row based on the HMVP list.

[0047] A seventh aspect of the present invention provides a computer program product comprising program code for carrying out a method according to the first aspect or any one of the methods of implementing the first aspect, or according to the third aspect or any one of the methods of implementing the third aspect, or according to the fourth aspect or any one of the methods of implementing the fourth aspect.

[0048] An eighth aspect of the present invention provides a computer-readable storage medium having stored thereon computer instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to the first aspect or any one of the methods for implementing the first aspect, or according to the third aspect or any one of the methods for implementing the third aspect, or according to the fourth aspect or any one of the methods for implementing the fourth aspect.

[0049] A ninth aspect of the present invention provides a decoder comprising: one or more processors; and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processor, configuring the decoder to perform a method according to the first aspect or any one of the methods for implementing the first aspect, or according to the third aspect or any one of the methods for implementing the third aspect, or according to the fourth aspect or any one of the methods for implementing the fourth aspect.

[0050] A tenth aspect of the present invention provides an encoder comprising: one or more processors; and a non-transitory computer-readable storage medium coupled to the processors and having stored thereon programming for execution by the processors, the programming, when executed by the processor, configuring the encoder to perform a method according to the first aspect or any one of the methods for implementing the first aspect, or according to the third aspect or any one of the methods for implementing the third aspect, or according to the fourth aspect or any one of the methods for implementing the fourth aspect.

[0051] In the following, embodiments of the invention will be explained in more detail with reference to the accompanying figures and drawings. [Brief explanation of the drawings]

[0052] [Figure 1A]1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Figure 1B] FIG. 2 is a block diagram illustrating another example of a video coding system configured to implement embodiments of the present invention. [Figure 2] 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present invention; [Figure 3] 1 is a block diagram illustrating one exemplary structure of a video decoder configured to implement embodiments of the present invention. [Figure 4] FIG. 1 is a block diagram illustrating an example of an encoding device or a decoding device. [Figure 5] FIG. 10 is a block diagram showing another example of an encoding device or a decoding device. [Figure 6] A diagram showing the locations of spatially adjacent blocks used in merging and AMVP candidate list construction. [Figure 7] 1 is a decoding flow chart of the HMVP method. [Figure 8] FIG. 10 is a block diagram showing the WPP processing sequence. [Figure 9] 4 is a flow diagram illustrating an exemplary operation of a video decoder according to one embodiment. [Figure 10] 1 is a flow diagram illustrating an exemplary operation according to one embodiment. [Figure 11] FIG. 1 is a block diagram illustrating an example of a video processing device. DETAILED DESCRIPTION OF THE INVENTION

[0053] Hereinafter, unless there is any specific note regarding the differences between identical reference signs, those identical reference signs refer to identical or at least functionally equivalent features.

[0054] In the following description, reference is made to the accompanying drawings, which form a part of this disclosure and which show, by way of illustration, specific aspects of embodiments of the present invention or in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other ways and may include structural or logical changes not shown in the drawings. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0055] For example, it is understood that disclosure regarding a described method may also be valid for a corresponding device or system configured to perform the method, and vice versa. For example, when one or more particular method steps are described, a corresponding device may include one or more units, e.g., functional units, for performing the described one or more method steps, even if such one or more units are not explicitly described or shown in a figure (e.g., one unit performs one or more steps, or multiple units each perform one or more of the multiple steps). On the other hand, for example, when a particular apparatus is described based on one or more units, e.g., functional units, a corresponding method may include one step for performing the functionality of the one or more units, even if such one or more steps are not explicitly described or shown in a figure (e.g., one step performs the functionality of one or more units, or multiple steps each perform the functionality of one or more of the multiple units). Furthermore, it is understood that features of various exemplary embodiments and / or aspects described herein may be combined with each other, unless expressly specified otherwise.

[0056] Video coding generally refers to the processing of a sequence of pictures that form a video or video sequence. Instead of the term "picture," the terms "frame" or "image" are sometimes used synonymously in the field of video coding. Video coding, as used in this application (or this disclosure), refers to either video encoding or video decoding. Video encoding generally occurs at the source side, including processing the original video picture (e.g., by compression) to reduce the amount of data needed to represent the video picture (for more efficient storage and / or transmission). Video decoding occurs at the destination side, and generally involves the reverse processing compared to the encoder to reconstruct the video picture. It should be understood that embodiments referring to "coding" a video picture (or, as described later, a picture in general) relate to either "encoding" or "decoding" a video sequence. The combination of the encoding and decoding parts is also referred to as a CODEC (coding and decoding).

[0057] In the case of lossless video coding, the original video pictures are reconstructable, i.e., the reconstructed video pictures have the same quality as the original video pictures (assuming there are no transmission losses or other data losses during storage or transmission). In the case of lossy video coding, further compression, e.g., by quantization, is performed to reduce the amount of data representing the video pictures, but these video pictures cannot be perfectly reconstructed at the decoder, i.e., the quality of the reconstructed video pictures is lower or worse than the quality of the original video pictures.

[0058] Since H.261, several video coding standards belong to the group of "lossy hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding to apply quantization in the transform domain). Each picture in a video sequence is generally divided into a set of non-overlapping blocks, and coding is generally performed at the block level. In other words, at an encoder, video is generally processed (i.e., encoded) at the block (video block) level, for example, by using spatial (intra-picture) and temporal (inter-picture) prediction to generate a predictive block, subtracting the predictive block from a current block (the block currently being / to be processed) to obtain a residual block, transforming the residual block, and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression). At a decoder, in contrast to the encoder, an inverse process is partially applied to the coded or compressed block to reconstruct the current block for representation. Additionally, the encoder replicates the decoder processing loop so that both generate the same predictions (e.g., intra-prediction and inter-prediction) and / or reconstructions for processing, i.e., coding, subsequent blocks.

[0059] As used herein, the term "block" may refer to a portion of a picture or a frame. For ease of explanation, embodiments of the present invention are described herein with reference to High Efficiency Video Coding (HEVC) or the reference software for Versatile Video Coding (VVC) developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Motion Picture Experts Group (MPEG) Joint Collaboration Team for Video Coding (JCT-VC). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC. Embodiments of the present invention may refer to CUs, PUs, and TUs. In HEVC, a CTU is divided into CUs by using a quadtree structure, denoted as a coding tree. The decision as to whether to code a picture area using inter-picture (temporal) prediction or intra-picture (spatial) prediction is made at the CU level. Each CU may be further divided into one, two, or four PUs according to the PU partition type. Within a PU, the same prediction process is applied, and related information is transmitted to the decoder on a PU-by-PU basis. After obtaining the residual block by applying the prediction process based on the PU partition type, the CU can be divided into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU. In the latest developments in video compression technology, quadtree and binary tree (QTBT) partitioned frames are used to divide coding blocks. In the QTBT block structure, CUs can have either square or rectangular shapes. For example, coding tree units (CTUs) are first divided by a quadtree structure. The quadtree leaf nodes are further divided by a binary tree structure. The binary tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transform processing without further division. This means that CUs, PUs, and TUs have the same block size in the QTBT coding block structure. In parallel, multi-partitioning, such as ternary tree partitioning, has also been proposed for use with the QTBT block structure.

[0060] In the following, embodiments of the encoder 20, the decoder 30 and the coding system 10 will be described based on FIGS.

[0061] 1A is a conceptual or schematic block diagram illustrating one exemplary coding system 10, e.g., a video coding system 10 that may utilize the techniques of this application (this disclosure). An encoder 20 (e.g., video encoder 20) and a decoder 30 (e.g., video decoder 30) of video coding system 10 represent examples of devices that may be configured to perform techniques according to various examples described herein. As shown in FIG. 1A, coding system 10 includes a source device 12 configured to provide encoded data 13, e.g., encoded pictures 13, to, e.g., a destination device 14, for decoding the encoded data 13.

[0062] Source device 12 includes an encoder 20 and may additionally, i.e., optionally, include a picture source 16, a pre-processing unit 18, eg, a picture pre-processing unit 18, and a communication interface or unit 22.

[0063] Picture source 16 may include or be, for example, any kind of picture capture device for capturing real-world pictures, and / or any kind of picture or comment (in the case of screen content coding, some text on the screen is also considered part of the picture or image to be coded) generation device, such as a computer graphics processor for generating computer-animated pictures, or any kind of device for obtaining and / or providing real-world pictures, computer-animated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures).

[0064] A (digital) picture can be considered as a two-dimensional array or matrix of samples with intensity values. The samples in the array are sometimes called pixels (short for picture element) or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. To represent color, three color components are generally employed, i.e., a picture can be represented by or contain three sample arrays. In an RBG format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each pixel is generally represented in a luma / chroma format or color space, e.g., YCbCr, which contains a luma component denoted by Y (sometimes L is used instead) and two chroma components denoted by Cb and Cr. The luma (or luma for short) component Y represents brightness or density intensity (e.g., as in a grayscale picture), and the two chroma (or chroma for short) components Cb and Cr represent color or color information components. Thus, a picture in YCbCr format includes a luma sample array of luma sample values ​​(Y) and two chroma sample arrays of chroma values ​​(Cb and Cr). A picture in RGB format can be converted or transformed to YCbCr format, and vice versa; this process is also called color transformation or color conversion. If a picture is monochrome, the picture may include only a luma sample array.

[0065] In monochrome sampling, there is only one sample array, nominally considered the luma array.

[0066] In 4:2:0 sampling, each of the two chroma arrays has half the height and half the width of the luma array.

[0067] In 4:2:2 sampling, each of the two chroma arrays has the same height and half the width of the luma array.

[0068] For 4:4:4 sampling, depending on the value of separate_colour_plane_flag the following applies: - If separate_colour_plane_flag is equal to 0, each of the two chroma arrays has the same height and width as the luma array. - Otherwise (separate_colour_plane_flag equals 1), the three colour planes are processed separately as a monochrome sampled picture.

[0069] Picture source 16 (e.g., video source 16) may be, for example, a camera for capturing a picture, a memory, e.g., a picture memory containing or storing previously captured or generated pictures, and / or any kind of interface (internal or external) for acquiring or receiving pictures. The camera may be, for example, a local camera or an integrated camera, e.g., integrated within the source device, and the memory may be a local memory or an integrated memory, e.g., integrated within the source device. The interface may be, for example, an external interface for receiving pictures from an external video source, e.g., a camera, an external memory, or an external picture generation device, e.g., an external computer graphics processor, computer, or server, or other external picture capture device. The interface may be any kind of interface, e.g., a wired or wireless interface, an optical interface, according to any specific or standardized interface protocol. The interface for acquiring picture data 17 may be the same interface as communication interface 22 or may be part of communication interface 22.

[0070] To distinguish from pre-processing unit 18 and the processing performed by pre-processing unit 18, pictures or picture data 17 (eg, video data 16) are sometimes referred to as raw pictures or raw picture data 17.

[0071] The pre-processing unit 18 is configured to receive (raw) picture data 17 and perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. The pre-processing performed by the pre-processing unit 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It may be understood that the pre-processing unit 18 may be an optional component.

[0072] The encoder 20 (e.g., video encoder 20) is configured to receive pre-processed picture data 19 and provide encoded picture data 21 (further details are described below, e.g., based on Figure 2 or Figure 4).

[0073] The communications interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and transmit it to another device, e.g., the destination device 14 or any other device, for storage or direct reconstruction, or to process the encoded picture data 21 before storing the encoded data 13 and / or before transmitting the encoded data 13 to another device, e.g., the destination device 14 or any other device, for decoding or storage, respectively.

[0074] Destination device 14 includes a decoder 30 (eg, a video decoder 30) and may additionally, i.e., optionally, include a communication interface or unit 28, a post-processing unit 32, and a display device 34.

[0075] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 or the encoded data 13, for example, directly from the source device 12 or from any other source, for example, a storage device, for example, an encoded picture data storage device.

[0076] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g., a direct wired or wireless connection, or via any type of network, e.g., a wired or wireless network, or any type of combination thereof, or any type of private and public network, or any type of combination thereof.

[0077] The communications interface 22 may be configured to package the encoded picture data 21 into an appropriate format, eg, packets, for transmission over a communications link or network.

[0078] The counterpart of the communication interface 22 , the communication interface 28 , may be configured to unpackage the encoded data 13 , for example, to obtain the encoded picture data 21 .

[0079] Both communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces, as indicated by the arrow for encoded picture data 13 in FIG. 1A pointing from source device 12 to destination device 14, or as bidirectional communication interfaces, and may be configured, for example, to send and receive messages to set up a connection and acknowledge and exchange any other information related to the communication link and / or data transmission, e.g., the encoded picture data transmission.

[0080] The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded pictures 31 (further details are described below, for example, based on Figure 3 or Figure 5).

[0081] Post-processor 32 of destination device 14 is configured to post-process decoded picture data 31 (also referred to as reconstructed picture data), e.g., decoded picture 31, to obtain post-processed picture data 33, e.g., post-processed picture 33. The post-processing performed by post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare decoded picture data 31 for, e.g., display, by, e.g., display device 34.

[0082] Display device 34 of destination device 14 is configured to receive the post-processed picture data 33 for displaying the picture, e.g., to a user or viewer. Display device 34 may be or comprise any type of display for presenting the reconstructed picture, e.g., an integrated or external display or monitor. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0083] 1A depicts source device 12 and destination device 14 as separate devices, an embodiment of the devices may comprise both or both functionality, source device 12 or corresponding functionality, and destination device 14 or corresponding functionality. In such an embodiment, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0084] As will be apparent to those skilled in the art based on the description, the presence and (exact) separation of functionality of different units or functionality within source device 12 and / or destination device 14, as shown in FIG. 1A, may vary depending on the actual device and application.

[0085] Encoder 20 (e.g., video encoder 20) and decoder 30 (e.g., video decoder 30) may each be implemented as any of a variety of suitable circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, or any combination thereof. Where these techniques are implemented partially in software, a device may store instructions for the software in a suitable, non-transitory computer-readable storage medium and may execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Any of the foregoing (including hardware, software, a combination of hardware and software, etc.) may be considered to be one or more processors. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) within the respective device.

[0086] Source device 12 may be referred to as a video encoding device or video encoding apparatus. Destination device 14 may be referred to as a video decoding device or video decoding apparatus. Source device 12 and destination device 14 may be examples of video coding devices or video coding apparatus.

[0087] Source device 12 and destination device 14 may comprise any of a wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content distribution server), a broadcast receiver device, a broadcast transmitter device, etc., and may use no operating system or any type of operating system.

[0088] In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices.

[0089] 1A is merely an example, and the techniques of the present application may be applied to video coding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between encoding and decoding devices. In other examples, data may be retrieved from local memory, streamed over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve and decode data from memory. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data to and / or retrieve and decode data from memory.

[0090] For each of the above examples described with reference to video encoder 20, it should be understood that video decoder 30 may be configured to perform the inverse process. With respect to syntax element signaling, video decoder 30 is configured to receive and parse such syntax elements and decode associated video data accordingly. In some examples, video encoder 20 may entropy encode one or more syntax elements into an encoded video bitstream. In such examples, video decoder 30 may parse such syntax elements and decode associated video data accordingly.

[0091] 1B is an illustrative diagram of another example video coding system 40 including the encoder 20 of FIG. 2 and / or the decoder 30 of FIG. 3, according to one exemplary embodiment. System 40 may implement techniques according to various examples described herein. In the illustrated implementation, video coding system 40 may include an imaging device 41, a video encoder 100, a video decoder 30 (and / or a video coder implemented via logic 47 of a processing unit 46), an antenna 42, one or more processors 43, one or more memory stores 44, and / or a display device 45.

[0092] As illustrated, imaging device 41, antenna 42, processing unit 46, logic circuitry 47, video encoder 20, video decoder 30, processor 43, memory store 44, and / or display device 45 may be capable of communicating with one another. As described above, while shown with both video encoder 20 and video decoder 30, video coding system 40 may include only video encoder 20 or only video decoder 30 in various examples.

[0093] As shown, in some examples, video coding system 40 may include antenna 42. Antenna 42 may be configured to transmit or receive, for example, an encoded bitstream of video data. Further, in some examples, video coding system 40 may include display device 45. Display device 45 may be configured to present the video data. As shown, in some examples, logic circuitry 47 may be implemented via processing unit 46. Processing unit 46 may include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. Video coding system 40 may also include an optional processor 43, which may also include application specific integrated circuit (ASIC) logic, a graphics processor, a general-purpose processor, etc. In some examples, logic circuitry 47 may be implemented via hardware, video coding dedicated hardware, etc., and processor 43 may be implemented via general-purpose software, an operating system, etc. Additionally, memory store 44 may be any type of memory, such as volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, memory store 44 may be implemented by cache memory. In some examples, logic circuitry 47 may access memory store 44 (e.g., for implementing an image buffer). In other examples, logic circuitry 47 and / or processing unit 46 may include a memory store (e.g., a cache, etc.) for implementing an image buffer, etc.

[0094] In some examples, video encoder 100 implemented via logic circuitry may include an image buffer (e.g., via either processing unit 46 or memory store 44) and a graphics processing unit (e.g., via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video encoder 100 implemented via logic circuitry 47 to implement various modules such as those described with respect to FIG. 2 and / or any other encoder system or subsystem described herein. The logic circuitry may be configured to perform various operations as described herein.

[0095] Video decoder 30 may be implemented in a manner similar to that implemented via logic circuitry 47 to implement various modules as described with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. In some examples, video decoder 30 may be implemented via logic circuitry and may include an image buffer (e.g., via either processing unit 420 or memory store 44) and a graphics processing unit (e.g., via processing unit 46). The graphics processing unit may be communicatively coupled to the image buffer. The graphics processing unit may include video decoder 30 as implemented via logic circuitry 47 to implement various modules as described with respect to FIG. 3 and / or any other decoder system or subsystem described herein.

[0096] In some examples, antenna 42 of video coding system 40 may be configured to receive an encoded bitstream of video data. As described, the encoded bitstream may include data, indicators, index values, mode selection data, etc. associated with encoding video frames as described herein, such as data related to coding partitions (e.g., transform coefficients or quantized transform coefficients, any indicators (as described), and / or data defining the coding partitions). Video coding system 40 may also include video decoder 30 coupled to antenna 42 and configured to decode the encoded bitstream. Display device 45 is configured to present the video frames.

[0097] Encoder and encoding method FIG. 2 illustrates a schematic / conceptual block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a buffer 216, a loop filter unit 220, a decoded picture buffer (DPB) 230, a prediction processing unit 260, and an entropy coding unit 270. The prediction processing unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a mode selection unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 illustrated in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder using a hybrid video codec.

[0098] For example, the residual calculation unit 204, the transform processing unit 206, the quantization unit 208, the prediction processing unit 260, and the entropy coding unit 270 form a forward signal path of the encoder 20, while, for example, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, and the prediction processing unit 260 form a backward signal path of the encoder, which corresponds to the signal path of the decoder (see decoder 30 in Figure 3).

[0099] Encoder 20 is configured to receive, for example, via input 202, block 203 of picture 20 or picture 201, e.g., a picture of a video or a sequence of pictures forming a video sequence. Picture block 203 may also be referred to as a current picture block or a picture block to be coded, and picture 201 may also be referred to as a current picture or a picture to be coded (particularly in video coding, to distinguish the current picture from other pictures, e.g., pictures that have been coded and / or decoded previously in the same video sequence, i.e., a video sequence that also includes the current picture).

[0100] Split An embodiment of encoder 20 may include a division unit (not shown in FIG. 2) configured to divide picture 201 into a number of blocks, which are generally non-overlapping blocks, such as block 203. The division unit may be configured to use the same block size for all pictures of a video sequence and a corresponding grid defining the block sizes, or to vary the block size between pictures or subsets or groups of pictures, and divide each picture into corresponding blocks.

[0101] In one example, prediction processing unit 260 of video encoder 20 may be configured to perform any combination of the partitioning techniques described above.

[0102] Like picture 201, block 203 may be or may be considered to be a two-dimensional array or matrix of samples having intensity values ​​(sample values), although again of smaller dimensions than picture 201. In other words, block 203 may include, for example, one sample array (e.g., one luma array for monochrome picture 201), or three sample arrays (e.g., one luma array and two chroma arrays for color picture 201), or any other number and / or type of arrays, depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 define the size of block 203.

[0103] The encoder 20 shown in FIG. 2 is configured to code a picture 201 block-wise, eg, coding and prediction is performed block 203-by-block.

[0104] Residual calculation The residual calculation unit 204 is configured to calculate the residual block 205 based on the picture block 203 and the prediction block 265 (further details regarding the prediction block 265 are described below), for example, by subtracting sample values ​​of the prediction block 265 from sample values ​​of the picture block 203 on a sample-by-sample (pixel-by-pixel) basis to obtain the residual block 205 in the sample domain.

[0105] conversion The transform processing unit 206 is configured to apply a transform, for example, a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values ​​of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207, sometimes referred to as transform residual coefficients, represent the residual block 205 in the transform domain.

[0106] The transform processing unit 206 may be configured to apply an integer approximation of a DCT / DST, such as a transform specified for HEVC / H.265. Compared to an orthogonal DCT transform, such an integer approximation is generally scaled by a factor. To maintain the standard of the residual blocks processed by the forward and inverse transforms, an additional scaling factor is applied as part of the transform process. The scaling factor is generally selected based on certain constraints, such as the scaling factor being a power of two for shift operations, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. A particular scaling factor may be specified for the inverse transform at the decoder 30, e.g., by the inverse transform processing unit 212 (and a corresponding inverse transform at the encoder 20, e.g., by the inverse transform processing unit 212), and a corresponding scaling factor may be specified for the forward transform at the encoder 20, e.g., by the transform processing unit 206.

[0107] Quantization The quantization unit 208 is configured to quantize the transform coefficients 207 to obtain quantized transform coefficients 209, for example, by applying scalar quantization or vector quantization. The quantized transform coefficients 209 are sometimes referred to as quantized residual coefficients 209. The quantization process may reduce a bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be truncated to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting a quantization parameter (QP). For example, in the case of scalar quantization, different scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, whereas a larger quantization step size corresponds to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may, for example, be an index to a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to fine quantization (small quantization step size), a large quantization parameter may correspond to coarse quantization (large quantization step size), or vice versa. Quantization may involve division by the quantization step size, and corresponding or inverse dequantization, e.g., by inverse quantization 210, may involve multiplication by the quantization step size. Some standards, e.g., HEVC, embodiments may be configured to use the quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of an equation, including division. To restore the reference of the residual block, additional scaling factors may be introduced for quantization and dequantization, and the reference of the residual block may be modified by the scaling used in the fixed-point approximation of the equation for the quantization step size and the quantization parameter. In one exemplary implementation, the scaling of the inverse transform and dequantization may be combined.Alternatively, customized quantization tables may be used and signaled, for example, in the bitstream, from the encoder to the decoder. Quantization is a lossy operation, and the loss increases with increasing quantization step size.

[0108] Inverse quantization unit 210 is configured to apply the inverse quantization of quantization unit 208 to the quantized coefficients to obtain dequantized coefficients 211, e.g., by applying the inverse of the quantization scheme applied by quantization unit 208, e.g., based on or using the same quantization step size as quantization unit 208. Dequantized coefficients 211, sometimes referred to as dequantized residual coefficients 211, generally correspond to transform coefficients 207, although they are not identical to the transform coefficients due to losses due to quantization.

[0109] The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, for example, an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST), to obtain an inverse transform block 213 in the sample domain. The inverse transform block 213 may also be referred to as an inverse transformed dequantized block 213 or an inverse transformed residual block 213.

[0110] The reconstruction unit 214 (e.g., adder 214) is configured to add the inverse transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265, for example, by adding the sample values ​​of the reconstructed residual block 213 and the sample values ​​of the prediction block 265, to obtain a reconstructed block 215 in the sample domain.

[0111] An optional buffer unit 216 (or "buffer" 216 for short), e.g., a line buffer 216, is configured to buffer or store the reconstructed blocks 215 and their respective sample values, e.g., for intra-prediction. In further embodiments, the encoder may be configured to use the unfiltered reconstructed blocks and / or their respective sample values ​​stored in the buffer unit 216 for any kind of estimation and / or prediction, e.g., intra-prediction.

[0112] Embodiments of encoder 20 may be configured, for example, such that buffer unit 216 is used to store reconstructed blocks 215 not only for intra prediction 254 but also for loop filter unit 220 (not shown in FIG. 2), and / or such that buffer unit 216 and decoded picture buffer unit 230 form one buffer. Further embodiments may be configured to use filtered blocks 221 and / or blocks or samples from decoded picture buffer 230 (both not shown in FIG. 2) as input or basis for intra prediction 254.

[0113] Loop filter unit 220 (or “loop filter” 220 for short) is configured to filter reconstructed block 215 to obtain filtered block 221, e.g., to smooth pixel transitions or otherwise improve video quality. Loop filter unit 220 is intended to represent one or more loop filters, such as a deblocking filter, a sample-adaptive offset (SAO) filter, or other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening or smoothing filter, or a collaborative filter. Although loop filter unit 220 is illustrated in FIG. 2 as being an in-loop filter, in other configurations, loop filter unit 220 may be implemented as a post-loop filter. Filtered block 221 may also be referred to as filtered reconstructed block 221. Decoded picture buffer 230 may store the reconstructed coding block after loop filter unit 220 performs a filtering operation on the reconstructed coding block.

[0114] Embodiments of encoder 20 (respectively, loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information), e.g., directly or entropy coded via entropy coding unit 270 or any other entropy coding unit, such that decoder 30 can receive and apply the same loop filter parameters for decoding.

[0115] The decoded picture buffer (DPB) 230 may be a reference picture memory that stores reference picture data for use in encoding video data by the video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The DPB 230 and the buffer 216 may be provided by the same memory device or separate memory devices. In some examples, the decoded picture buffer (DPB) 230 is configured to store the filtered block 221. The decoded picture buffer 230 may further be configured to store other previously filtered blocks, e.g., previously reconstructed filtered block 221, of the same current picture or of a different picture, e.g., a previously reconstructed picture, and may provide a complete previously reconstructed, i.e., decoded, picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples), e.g., for inter-prediction. In some examples, if the reconstructed block 215 is reconstructed without in-loop filtering, the decoded picture buffer (DPB) 230 is configured to store the reconstructed block 215.

[0116] The prediction processing unit 260, sometimes referred to as the block prediction processing unit 260, is configured to receive or obtain the block 203 (the current block 203 of the current picture 201) and reconstructed picture data, e.g., reference samples of the same (current) picture, from the buffer 216 and / or reference picture data 231 from one or more previously decoded pictures from the decoded picture buffer 230, and to process such data for prediction, i.e., to provide a prediction block 265, which may be an inter-predicted block 245 or an intra-predicted block 255.

[0117] The mode selection unit 262 may be configured to select a prediction mode (e.g., intra prediction mode or inter prediction mode) and / or a corresponding prediction block 245 or 255 to be used as the prediction block 265 for calculation of the residual block 205 and for reconstruction of the reconstructed block 215.

[0118] Embodiments of mode selection unit 262 may be configured to select (e.g., from prediction modes supported by prediction processing unit 260) the prediction mode that provides the best match, i.e., the smallest residual (which means better compression for transmission or storage), or the smallest signaling overhead (which means better compression for transmission or storage), or that considers both or balances both. Mode selection unit 262 may be configured to determine the prediction mode based on rate distortion optimization (RDO), i.e., select the prediction mode that provides the smallest rate distortion optimization, or which associated rate distortion at least satisfies a prediction mode selection criterion.

[0119] The prediction processing (eg, by prediction processing unit 260) and mode selection (eg, by mode select unit 262) performed by one exemplary encoder 20 are described in more detail below.

[0120] As mentioned above, encoder 20 is configured to determine or select a best or optimal prediction mode from a (predetermined) set of prediction modes, which may include, for example, intra-prediction modes and / or inter-prediction modes.

[0121] The set of intra-prediction modes may include 35 different intra-prediction modes, e.g., omni-directional modes such as DC (or average) mode and planar mode, or directional modes as defined, e.g., in H.265, or may include 67 different intra-prediction modes, e.g., omni-directional modes such as DC (or average) mode and planar mode, or directional modes as defined, e.g., in the currently under development H.266.

[0122] The set of inter prediction modes (or possible inter prediction modes) depends on the available reference pictures (i.e., previous, at least partially decoded pictures stored, for example, in DBP 230) and other inter prediction parameters, such as whether the entire reference picture of the reference picture is used to search for the best matching reference block, or whether only a portion of the reference picture, for example, a search window area around the area of ​​the current block, is used, and / or whether pixel interpolation, for example, half / semi-pel and / or 1 / 4-pel interpolation, is or is not applied.

[0123] In addition to the above prediction modes, skip mode and / or direct mode may be applied.

[0124] The prediction processing unit 260 may be further configured to divide the block 203 into smaller block partitions or sub-blocks, for example, using quadtree partitioning (QT), binary tree partitioning (BT), or ternary tree partitioning (TT), or any combination thereof, iteratively, and to perform predictions for each of the block partitions or sub-blocks, for example, where the mode selection includes selecting a tree structure of the divided block 203 and a prediction mode to be applied to each of the block partitions or sub-blocks.

[0125] The inter prediction unit 244 may include a motion estimation (ME) unit (not shown in FIG. 2) and a motion compensation (MC) unit (not shown in FIG. 2). The motion estimation unit is configured to receive or obtain a picture block 203 (current picture block 203 of current picture 201) and a decoded picture 231, or at least one or more previously reconstructed blocks, e.g., reconstructed blocks of one or more other / different previously decoded pictures 231, for motion estimation. For example, a video sequence may include the current picture and the previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of or form a sequence of pictures that form a video sequence.

[0126] The encoder 20 may be configured to, for example, select one reference block from multiple reference blocks of the same or different pictures among multiple other pictures, and provide the reference picture (or reference picture index, ...) and / or an offset (spatial offset) between the position (x coordinate, y coordinate) of the reference block and the position of the current block to the motion estimation unit (not shown in FIG. 2) as an inter-prediction parameter. This offset is also called a motion vector (MV).

[0127] The motion compensation unit is configured to obtain, e.g., receive, inter prediction parameters and perform inter prediction based on or using the inter prediction parameters to obtain inter prediction block 245. Motion compensation performed by the motion compensation unit (not shown in FIG. 2) may require interpolation, possibly to sub-pixel accuracy, to fetch or generate a predictive block based on motion / block vectors determined by motion estimation. Interpolation filtering may generate additional pixel samples from known pixel samples, thereby potentially increasing the number of candidate predictive blocks that can be used to code the picture block. Upon receiving a motion vector for the PU of the current picture block, motion compensation unit 246 may locate the predictive block to which the motion vector points within one of the reference picture lists. Motion compensation unit 246 may also generate syntax elements associated with the block and the video slice for use by video decoder 30 in decoding picture blocks of the video slice.

[0128] The intra prediction unit 254 is configured to obtain, for example, receive, the picture block 203 (current picture block) and one or more previously reconstructed blocks of the same picture, for example, reconstructed neighboring blocks, for intra estimation. The encoder 20 may be configured, for example, to select an intra prediction mode from a plurality of (predetermined) intra prediction modes.

[0129] An embodiment of the encoder 20 may be configured to select an intra prediction mode based on an optimization criterion, for example, minimum residual (e.g., the intra prediction mode that provides the predicted block 255 that is most similar to the current picture block 203) or minimum rate distortion.

[0130] The intra prediction unit 254 is further configured to determine the intra prediction block 255 based on the intra prediction parameters, e.g., the selected intra prediction mode. In either case, after selecting the intra prediction mode for the block, the intra prediction unit 254 is also configured to provide the intra prediction parameters, i.e., information indicating the selected intra prediction mode for the block, to the entropy coding unit 270. In one example, the intra prediction unit 254 may be configured to perform any combination of the intra prediction techniques described below.

[0131] The entropy encoding unit 270 is configured to apply an entropy encoding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context adaptive VLC scheme (CALVC), an arithmetic coding scheme, a context adaptive binary arithmetic coding (CABAC), a syntax-based context-adaptive binary arithmetic coding (SBAC), a probability interval partitioning entropy (PIPE) coding, or other entropy encoding methodology or technique) to the quantized residual coefficients 209, the inter-prediction parameters, the intra-prediction parameters, and / or the loop filter parameters, individually or jointly (or not at all), to obtain encoded picture data 21, which may be output via an output 272, for example, in the form of an encoded bitstream 21. Encoded bitstream 21 may be transmitted to video decoder 30 or may be archived for later transmission or retrieval by video decoder 30. Entropy encoding unit 270 may further be configured to entropy encode other syntax elements for the current video slice being coded.

[0132] Other structural variations of the video encoder 20 may be used to encode the video stream. For example, a non-transform-based encoder 20 may directly quantize the residual signal for certain blocks or frames, without the transform processing unit 206. In other implementations, the encoder 20 may combine the quantization unit 208 and the inverse quantization unit 210 into a single unit.

[0133] 3 shows an example video decoder 30 configured to implement the techniques of the present application. Video decoder 30 is configured to receive coded picture data (e.g., coded bitstream) 21, e.g., coded by encoder 100, to obtain decoded picture 131. During the decoding process, video decoder 30 receives video data, e.g., a coded video bitstream and associated syntax elements representing picture blocks of coded video slices, from video encoder 100.

[0134] 3, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., adder 314), a buffer 316, a loop filter 320, a decoded picture buffer 330, and a prediction processing unit 360. Prediction processing unit 360 may include an inter prediction unit 344, an intra prediction unit 354, and a mode selection unit 362. Video decoder 30, in some examples, may perform a decoding pass that is generally inverse to the encoding pass described with respect to video encoder 100 from FIG.

[0135] Entropy decoding unit 304 is configured to perform entropy decoding on coded picture data 21, e.g., to obtain quantized coefficients 309 and / or decoded coding parameters (not shown in FIG. 3), e.g., (decoded) any or all of inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements. Entropy decoding unit 304 is further configured to forward the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to prediction processing unit 360. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level.

[0136] The inverse quantization unit 310 may be functionally identical to the inverse quantization unit 110, the inverse transform processing unit 312 may be functionally identical to the inverse transform processing unit 112, the reconstruction unit 314 may be functionally identical to the reconstruction unit 114, the buffer 316 may be functionally identical to the buffer 116, the loop filter 320 may be functionally identical to the loop filter 120, and the decoded picture buffer 330 may be functionally identical to the decoded picture buffer 130.

[0137] Prediction processing unit 360 may include an inter prediction unit 344 and an intra prediction unit 354, which may be similar in functionality to inter prediction unit 144 and intra prediction unit 354, respectively. Prediction processing unit 360 is generally configured to perform block prediction and / or obtain a prediction block 365 from coded data 21, and to receive or obtain (explicitly or implicitly) information regarding prediction relationship parameters and / or a selected prediction mode, for example, from entropy decoding unit 304.

[0138] When a video slice is coded as an intra-coded (I) slice, intra prediction unit 354 of prediction processing unit 360 is configured to generate a predictive block 365 for a picture block of the current video slice based on the signaled intra prediction mode and data from a block decoded previously to the current frame or picture. When a video frame is coded as an inter-coded (i.e., B or P) slice, inter prediction unit 344 (e.g., a motion compensation unit) of prediction processing unit 360 is configured to create a predictive block 365 for a video block of the current video slice based on a motion vector and other syntax elements received from entropy decoding unit 304. For inter prediction, the predictive block may be created from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct the reference frame lists, i.e., List 0 and List 1, using a default construction technique based on the reference pictures stored in DPB 330.

[0139] Prediction processing unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing the motion vectors and other syntax elements, and use the prediction information to create a predictive block for the current video block being decoded. For example, prediction processing unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra-prediction or inter-prediction) used to code the video blocks of the video slice, the inter-prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more of the reference picture lists for the slice, the motion vectors for each inter-coded video block of the slice, the inter-prediction state for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice.

[0140] Inverse quantization unit 310 is configured to inverse quantize, i.e., dequantize, the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 304. The inverse quantization process may include using quantization parameters calculated by video encoder 100 for each video block in a video slice to determine the degree of quantization, and similarly, the degree of inverse quantization to be applied.

[0141] Inverse transform processing unit 312 is configured to apply an inverse transform, eg, an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to produce residual blocks in the pixel domain.

[0142] The reconstruction unit 314 (e.g., adder 314) is configured to add the inverse transform block 313 (i.e., the reconstructed residual block 313) to the prediction block 365, for example, by adding the sample values ​​of the reconstructed residual block 313 and the sample values ​​of the prediction block 365, to obtain a reconstructed block 315 in the sample domain.

[0143] Loop filter unit 320 is configured to filter reconstructed block 315 (either during or after the coding loop) to obtain filtered block 321, e.g., to smooth pixel transitions or otherwise improve video quality. In one example, loop filter unit 320 may be configured to perform any combination of the filtering techniques described below. Loop filter unit 320 is intended to represent one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or other filters, e.g., a bilateral filter, or an adaptive loop filter (ALF), or a sharpening or smoothing filter, or a collaborative filter. Although loop filter unit 320 is shown in FIG. 3 as being an in-loop filter, in other configurations, loop filter unit 320 may be implemented as a post-loop filter.

[0144] The decoded video blocks 321 in a given frame or picture are then stored in a decoded picture buffer 330, which stores reference pictures used for subsequent motion compensation.

[0145] The decoder 30 is configured to output the decoded pictures 311 for presentation to or viewing by a user, for example via an output 312.

[0146] Other variations of the video decoder 30 may be used to decode the compressed bitstream. For example, the decoder 30 may create an output video stream without the loop filtering unit 320. For example, a non-transform-based decoder 30 may directly inverse quantize the residual signal without the inverse transform processing unit 312 for certain blocks or frames. In other implementations, the video decoder 30 may combine the inverse quantization unit 310 and the inverse transform processing unit 312 into a single unit.

[0147] 4 is a schematic diagram of a video coding device 400 according to one embodiment of the present disclosure. Video coding device 400 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, video coding device 400 may be a decoder, such as video decoder 30 of FIG. 1A, or an encoder, such as video encoder 20 of FIG. 1A. In one embodiment, video coding device 400 may be one or more components of video decoder 30 of FIG. 1A or video encoder 20 of FIG. 1A, as described above.

[0148] Video coding device 400 includes an ingress port 410 and a receiver unit (Rx) 420 for receiving data, a processor, logic unit, or central processing unit (CPU) 430 for processing data, a transmitter unit (Tx) 440 and an egress port 450 for transmitting data, and a memory 460 for storing data. Video coding device 400 may also include optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to ingress port 410, receiver unit 420, transmitter unit 440, and egress port 450 for the egress or ingress of optical or electrical signals.

[0149] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 430 is in communication with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the disclosed embodiments described above. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. The inclusion of the coding module 470 thus significantly improves the functionality of the video coding device 400 and transforms the video coding device 400 into different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.

[0150] Memory 460 may include one or more disks, tape drives, and solid state drives, and may be used as overflow data storage devices for storing programs when such programs are selected for execution, and for storing instructions and data read during program execution. Memory 460 may be volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random access memory (SRAM).

[0151] 5 is a simplified block diagram of an apparatus 500 that may be used as either or both of the source device 310 and the destination device 320 from FIG. 1 , according to one exemplary embodiment. The apparatus 500 may implement the techniques of the present application described above. The apparatus 500 may be in the form of a computing system including multiple computing devices, or in the form of a single computing device, such as a mobile phone, tablet computer, laptop computer, notebook computer, desktop computer, etc.

[0152] The processor 502 in the device 500 may be a central processing unit. Alternatively, the processor 502 may be any other type of device or devices, now existing or later developed, capable of manipulating or processing information. While the disclosed implementations may be practiced with a single processor, e.g., processor 502, as shown, advantages in speed and efficiency can be achieved using two or more processors.

[0153] The memory 504 in the device 500, in one implementation, may be a read-only memory (ROM) device or a random-access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 accessed by the processor 502 using a bus 512. The memory 504 may further include an operating system 508 and application programs 510, which include at least one program that allows the processor 502 to perform the methods described herein. For example, the application programs 510 may include applications 1 through N, which further include a video coding application that performs the methods described herein. The device 500 may include additional memory in the form of a secondary storage device 514, which may be, for example, a memory card used with a mobile computing device. Because video communication sessions may contain a significant amount of information, these sessions may be stored in whole or in part in the secondary storage device 514 and loaded into the memory 504 as needed for processing.

[0154] The device 500 may include one or more output devices, such as a display 518. The display 518, in one example, may be a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512. In addition to, or instead of, the display 518, other output devices may be provided that allow a user to program or otherwise use the device 500. When the output device is or includes a display, the display may be implemented in various ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, a plasma display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.

[0155] The device 500 may include or be in communication with an image sensing device 520, for example, a camera or any other now existing or later developed image sensing device 520 that can sense images, such as an image of a user operating the device 500. The image sensing device 520 may be positioned such that the image sensing device 520 faces the user operating the device 500. In one example, the position and optical axis of the image sensing device 520 may be configured such that the field of view is directly adjacent to the display 518 and includes an area from which the display 518 is viewable.

[0156] The device 500 may include or be in communication with a voice sensing device 522, such as a microphone or any other now existing or later developed voice sensing device that can sense sound in the vicinity of the device 500. The voice sensing device 522 may be positioned such that the voice sensing device 522 faces a user operating the device 500 and may be configured to receive sound, e.g., speech or other utterances, made by the user while the user is operating the device 500.

[0157] While FIG. 5 depicts the processor 502 and memory 504 of device 500 as integrated into a single unit, other configurations may be utilized. The operations of processor 502 may be distributed across multiple machines (each machine having one or more of the processors) that may be coupled directly or across a local area network or other network. Memory 504 may be distributed across multiple machines, such as a network-based memory or memory in multiple machines that perform the operations of device 500. While shown here as a single bus, bus 512 of device 500 may consist of multiple buses. Furthermore, secondary storage 514 may be directly coupled to other components of device 500 or may be accessed over a network and may include a single integrated unit, such as a memory card, or may include multiple units, such as multiple memory cards. Device 500 may therefore be implemented in a wide variety of configurations.

[0158] In VVC, the motion vector of an inter-coded block can be signaled in two ways: Advanced motion vector prediction (AMVP) mode or merge mode. In the AMVP mode, the difference between the actual motion vector and the motion vector prediction (MVP), a reference index pointing to the AMVP candidate list, and an MVP index are signaled, where the reference index points to the reference picture from which the reference block is copied for motion compensation. In the merge mode, a merge index pointing to the merge candidate list is signaled, and all motion information related to the merge candidate is inherited.

[0159] For both the AMVP candidate list and the merge candidate list, they are derived from coded blocks that are temporally or spatially adjacent. More specifically, the merge candidate list is constructed by sequentially examining the following four types of merge MVP candidates: 1. As shown in FIG. 6, spatial merge candidates can be determined from five spatially adjacent blocks, namely, blocks A0 and A1 located in the lower left corner, blocks B0 and B1 located in the upper right corner, and block B2 located in the upper left corner. 2. Temporal MVP (TMVP) merge candidates. 3. Combined bi-predictive merging candidates. 4. Zero motion vector merging candidates.

[0160] The merge candidate list construction process terminates when the number of available merge candidates reaches the signaled maximum allowed merge candidates (e.g., 5 under typical test conditions). Note that the maximum allowed merge candidates may be different under different conditions.

[0161] Similarly, for the AMVP candidate list, three types of MVP candidates are examined in turn: 1. At most two spatial MVP candidates, one of which is determined from blocks B0, B1, and B2 as shown in Figure 6, and the other of which is determined from blocks A0 and A1 as shown in Figure 6. 2. Temporal MVP (TMVP) candidate. 3. Zero MVP candidates.

[0162] The history-based motion vector prediction (HMVP) method was introduced by JVET-K0104 (accessible at http: / / phenix.it-sudparis.eu / jvet / ), an input document to the Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, where an HMVP candidate is defined as the motion information of a previously coded block. A table with multiple HMVP candidates is maintained during the encoding / decoding process. This table is emptied when a new slice is encountered. Whenever an inter-coded block exists, the associated motion information is added to the last entry of the table as a new HMVP candidate. The overall coding flow is shown in Figure 7, including:

[0163] Step 701. Load the HMVP candidates into a table.

[0164] Step 702. Decode the block with the HMVP candidates in the loaded table.

[0165] Step 703. When decoding a block, update the table with the decoded motion information.

[0166] Steps 701 to 703 may be performed cyclically.

[0167] The HMVP candidates may be used in the merge candidate list construction process. All HMVP candidates from the last entry to the first entry in the table are inserted after the TMVP candidate. Pruning may be applied to the HMVP candidates. When the total number of available merge candidates reaches the signaled maximum allowed merge candidates, the merge candidate list construction process terminates.

[0168] The act of pruning refers to identifying identical motion estimator candidates in the list and removing one of the identical candidates from the list.

[0169] Similarly, HMVP candidates may be used in the AMVP candidate list construction process. The motion vectors of the last K HMVP candidates in the table are inserted after the TMVP candidate. In some implementations, only HMVP candidates with the same reference picture as the AMVP target reference picture are used to construct the AMVP candidate list. Pruning may be applied to HMVP candidates.

[0170] To improve processing efficiency, a process called wavefront parallel processing (WPP) is introduced, where WPP mode allows rows of CTUs to be processed simultaneously. In WPP mode, each CTU row is processed relative to its preceding (directly adjacent) CTU row by using a delay of two consecutive CTUs. For example, referring to FIG. 8, a picture frame or picture area consists of multiple CTU rows, and each thread (row) includes 11 CTUs: thread 1 includes CTU0 to CTU10, thread 2 includes CTU11 to CTU21, thread 3 includes CTU22 to CTU32, and thread 4 includes CTU33 to 43... Therefore, in WPP mode, when the encoding / decoding process of CTU1 in thread 1 is completed, the encoding / decoding process of CTU11 in thread 2 may start; similarly, when the encoding / decoding process of CTU12 in thread 2 is completed, the encoding / decoding process of CTU22 in thread 3 may start; when the encoding / decoding process of CTU23 in thread 3 is completed, the encoding / decoding process of CTU33 in thread 4 may start; and when the encoding / decoding process of CTU34 in thread 4 is completed, the encoding / decoding process of CTU44 in thread 5 may start.

[0171] However, when combining WPP with HMVP, as described above, an HMVP list is maintained and updated after processing each coding block, and thus continues to be updated until the last CTU in the CTU row, and one HMVP list is maintained, and therefore, wavefront parallel processing is not possible because thread N needs to wait for the last CTU in the above CTU row to be processed.

[0172] 9 is a flow diagram illustrating one exemplary operation of a video decoder, such as the video decoder 30 of FIG. 3, according to one embodiment of the present application. One or more structural elements of the video decoder 30, including the inter prediction unit 344, may be configured to perform the techniques of FIG. 9. In the example of FIG. 9, the video decoder 30 may perform the following steps:

[0173] 901. At the start of processing a CTU row, a step is taken to build / initialize the HMVP list for the CTU row.

[0174] When the CTU to be processed is the first CTU (start CTU) in a CTU row, an HMVP list for the CTU row is constructed or initialized, and thus the first CTU in the CTU row can be processed based on the HMVP list for the CTU row.

[0175] When the method is an encoding method, the HMVP list for the CTU row may be constructed or initialized by the inter prediction unit 344 of Figure 3. Alternatively, when the method is a decoding method, the HMVP list for the CTU row may be constructed or initialized by the inter prediction unit 244 of Figure 2.

[0176] In one implementation, all CTU rows with different HMVP lists may be maintained for a picture frame. In another implementation, all CTU rows with different HMVP lists may be maintained for a picture area, where the picture area consists of multiple CTU rows and the picture may be a VVC slice, tile, or brick.

[0177] Where a brick is a rectangular region of CTU rows within a particular tile in a picture, the tile may be divided into multiple bricks, each of which consists of one or more CTU rows within the tile. A tile that is not divided into multiple bricks is also called a brick. However, a brick that is a true subset of a tile is not called a tile.

[0178] Note that maintaining a different HMVP list for every CTU row simply means that a particular HMVP list may be maintained for a CTU row, but the candidates in the different HMVP lists may be the same, e.g., all candidates in one HMVP list are the same as candidates in other HMVP lists; the candidates in one HMVP list may not have redundancy; or the candidates in different HMVP lists may have overlaps, e.g., some of the candidates in one HMVP list are the same as some of the candidates in other HMVP lists, and some of the candidates in one HMVP list do not have the same candidate in other HMVP lists; or the candidates in different HMVP lists may be completely different, e.g., none of the candidates in one HMVP list have the same candidate in other HMVP lists. Note that when all CTUs in a CTU row have been processed, the HMVP list maintained for that CTU row may be released, thus reducing storage requirements.

[0179] This disclosure provided the following methods for building / initializing an HMVP list.

[0180] Method 1: At the start of processing a CTU row, the corresponding HMVP list is either emptied or set to a default value, which is a predetermined candidate known to both the encoder and the decoder.

[0181] For example, the corresponding HMVP list is populated with the following default MVs: a) a MV from a uni-prediction method, where the MV may be a zero motion vector and the reference picture may include the first reference picture in the L0 list; and / or b) MVs from a bi-predictive method, where the MVs may be zero motion vectors and the reference pictures may include the first reference picture in the L0 list and the first reference picture in the L1 list; and / or c) MVs of previously processed pictures according to the picture processing order, more specifically, MVs belonging to previously processed pictures and within the spatial neighborhood of the current block when the current block position is overlaid on the previous picture; and / or d) MVs of the temporal HMVP list, where each co-located picture may store a temporal HMVP list for each CTU row or for the entire picture, and thus the temporal HMVP list may be used to build / initialize the HMVP list for the current CTU row.

[0182] Method 2: At the start of processing the current CTU row, the corresponding HMVP list is constructed / initialized based on the HMVP list of the second CTU of the previous CTU row, where the previous CTU row is the CTU row that is directly adjacent to and above the current CTU row.

[0183] Using Figure 8 as an example, when the current CTU row is the CTU row of thread 2, the previous CTU row is the CTU row of thread 1, and the second CTU in the previous row is CTU1; when the current CTU row is the CTU row of thread 3, the previous CTU row is the CTU row of thread 2, and the second CTU in the previous row is CTU12; when the current CTU row is the CTU row of thread 4, the previous CTU row is the CTU row of thread 3, and the second CTU in the previous row is CTU23; when the current CTU row is the CTU row of thread 5, the previous CTU row is the CTU row of thread 4, and the second CTU in the previous row is CTU34; when the current CTU row is the CTU row of thread 6, the previous CTU row is the CTU row of thread 5, and the second CTU in the previous row is CTU45.

[0184] Method 3: At the start of processing the current CTU row, the corresponding HMVP list is constructed / initialized based on the HMVP list of the first CTU in the previous CTU row, which is the CTU row directly adjacent to and above the current CTU row.

[0185] Using Figure 8 as an example, when the current CTU row is the CTU row of thread 2, the previous CTU row is the CTU row of thread 1, and the first CTU in the previous row is CTU0; when the current CTU row is the CTU row of thread 3, the previous CTU row is the CTU row of thread 2, and the first CTU in the previous row is CTU11; when the current CTU row is the CTU row of thread 4, the previous CTU row is the CTU row of thread 3, and the first CTU in the previous row is CTU22; when the current CTU row is the CTU row of thread 5, the previous CTU row is the CTU row of thread 4, and the first CTU in the previous row is CTU33; when the current CTU row is the CTU row of thread 6, the previous CTU row is the CTU row of thread 5, and the first CTU in the previous row is CTU44.

[0186] According to methods 1 to 3, the processing of the current CTU row does not need to wait for the processing of the CTU row before the current CTU row to be completed, and thus may improve the processing efficiency of the current picture frame.

[0187] 902. Process the CTUs in the CTU row based on the constructed / initialized HMVP list.

[0188] The processing of the CTU may be an inter-prediction process performed during the decoding process, i.e., the processing of the CTU may be implemented by the inter-prediction unit 344 of Figure 3. Alternatively, the processing of the CTU may be an inter-prediction process performed during the encoding process, i.e., the processing of the CTU may be implemented by the inter-prediction unit 244 of Figure 2.

[0189] It should be noted that the above method for constructing / initializing the HMVP list can also be used for normal HMVP processing without wavefronts, e.g., HMVP processing without WPP. As a result, the HMVP processing is identical regardless of the application of WPP, which reduces the need for additional logic implementation.

[0190] It should be noted that the process of FIG. 9 may be an encoding process implemented by an encoder, such as the video encoder 20 of FIG. 2, according to one embodiment of the present application.

[0191] Furthermore, it should be noted that the above-described method for combining wavefront and HMVP-based prediction can also be used for intra prediction, i.e., the history intra mode may be used and the history table for each CTU row is initialized to default values.

[0192] For example, the initialization of the HMVP list for each CTU row in intra prediction may be performed in a default intra mode, such as planar mode, DC mode, vertical mode, horizontal mode, mode 2 mode, VDIA mode, and DIA mode.

[0193] Figure 10 is a flow diagram illustrating one exemplary operation of a video decoder or video encoder, such as the video decoder 30 of Figure 3 and the video encoder 20 of Figure 2 according to one embodiment of the present application. One or more structural elements of the video decoder 30 / encoder 20, including the inter prediction unit 344 / inter prediction unit 244, may be configured to perform the techniques of Figure 10. In the example of Figure 10, the video decoder 30 / video encoder 20 may perform the following steps:

[0194] Step 1010: When the current CTU is the start CTU of the current CTU row, initialize the HMVP list for the current CTU row.

[0195] It should be noted that the current CTU row may be any CTU row in a picture frame consisting of multiple CTU rows, or a picture area (which may be a portion of a picture frame) may consist of multiple CTU rows, and the current CTU row may be any one of the multiple CTU rows.

[0196] Whether the current CTU is the starting CTU (or the first CTU) of the current CTU row can be determined based on the index of the current CTU. For example, as disclosed in FIG. 8, each CTU has a unique index, and therefore, based on the index of the current CTU, it can be determined whether the current CTU is the first CTU of the current CTU row. For example, CTUs with indexes 0, 11, 22, 33, 44, or 55... are the first CTUs of the CTU row, respectively. Alternatively, using FIG. 8 as an example, each CTU row includes 11 CTUs, that is, the width of each CTU row is 11. Therefore, the index of the CTU can be divided by the width of the CTU row to determine whether the remainder is 0. If the remainder is 0, the corresponding CTU is the first CTU of the CTU row; otherwise, if the remainder is not 0, the corresponding CTU is not the first CTU of the CTU row. That is, if a CTU's index % width of the CTU row = 0, the CTU is the first CTU in the CTU row; otherwise, if a CTU's index % width of the CTU row ≠ 0, the CTU is not the first CTU in the CTU row. Note that when the process of the CTU row is from right to left, whether a CTU is the starting CTU of the CTU row can be determined in a similar manner.

[0197] After the initialization of the HMVP list, the number of candidate motion vectors in the initialized HMVP list is zero.

[0198] Initialization can be performed by emptying the HMVP list for the current CTU row, i.e., by removing the contents of the HMVP list for the current CTU row, in other words, by setting the number of candidates in the HMVP list for the current CTU row to zero.

[0199] In another implementation method, the method may further include the following step: initializing an HMVP list for each of a plurality of CTU rows, excluding the current CTU row, wherein the HMVP lists for the plurality of CTU rows are the same or different.

[0200] Initialization can be performed by setting default values ​​for the HMVP list for the current CTU row, or by initializing the HMVP list for the current CTU row based on the HMVP list of the CTUs in the previous CTU row, as described above.

[0201] Step 1020 processes the current CTU row based on the HMVP list.

[0202] The process may be an inter-prediction process, whereby a predictive block may be obtained. Reconstruction may be performed based on the predictive block to obtain a reconstructed block, and finally, a decoded picture may be obtained based on the reconstructed block. Details of these processes are described above.

[0203] As shown in FIG. 8, the current picture frame includes multiple CTU rows to improve coding / decoding efficiency, and the multiple CTU rows can be processed in wavefront parallel processing (WPP) mode. That is, when a specific CTU in the previous CTU row is processed, the current CTU row begins to be processed (or processing of the current CTU row begins), where the previous CTU row is the CTU row directly adjacent to and above the current CTU row, and the specific CTU in the previous CTU row is the second CTU in the previous CTU row; or the specific CTU in the previous CTU row is the first CTU in the previous CTU row. Using FIG. 8 as an example, when the current CTU row is thread 3, the previous CTU row is thread 2, and the specific CTU in the previous CTU row may be CTU12. That is, when CTU12 is processed, the decoder / encoder begins to process the CTU row in thread 3, that is, the decoder / encoder begins to process CTU22. Using Figure 8 as another example, when the current CTU row is thread 4, the previous CTU row is thread 3, and a particular CTU in the previous CTU row may be CTU23, i.e., when CTU23 is processed, the decoder / encoder begins processing the CTU row of thread 4, i.e., the decoder / encoder begins processing CTU33.

[0204] In one implementation method, the step of processing the current CTU row based on the HMVP list may include the steps of processing a current CTU in the current CTU row, updating the initialized HMVP list based on the processed current CTU, and processing a second CTU in the current CTU row based on the updated HMVP list.

[0205] FIG. 11 is a block diagram illustrating an example of a video processing device 1100 configured to implement an embodiment of the present invention, which may be an encoder 20 or a decoder 30 as shown in FIG. 11, and which includes:

[0206] an initialization unit 1110 configured to initialize an HMVP list for a current CTU row when the current CTU is the starting CTU (first CTU) of the current CTU row;

[0207] For details of the initialization performed by the initialization unit 1110, see step 1010.

[0208] A processing unit 1120 configured to process the current CTU row based on the HMVP list.

[0209] For details of the processing performed by the processing unit 1120, refer to step 1020.

[0210] The process may be an inter-prediction process, whereby a predictive block may be obtained. Reconstruction may be performed based on the predictive block to obtain a reconstructed block, and finally, a decoded picture may be obtained based on the reconstructed block. Details of these processes are described above.

[0211] As shown in FIG. 8, the current picture frame includes multiple CTU rows, and in order to improve coding / decoding efficiency, multiple CTU rows can be processed in WPP mode. That is, when a specific CTU in the previous CTU row is processed, the current CTU row begins to be processed, where the previous CTU row is a CTU row directly adjacent to and above the current CTU row, and the specific CTU in the previous CTU row is the second CTU in the previous CTU row; or the specific CTU in the previous CTU row is the first CTU in the previous CTU row. Using FIG. 8 as an example, when the current CTU row is thread 3, the previous CTU row may be thread 2, and the specific CTU in the previous CTU row may be CTU12. That is, when CTU12 is processed, the decoder / encoder begins to process the CTU row in thread 3, that is, the decoder / encoder begins to process CTU22. Using Figure 8 as another example, when the current CTU row is thread 4, the previous CTU row is thread 3, and a particular CTU in the previous CTU row may be CTU23, i.e., when CTU23 is processed, the decoder / encoder begins processing the CTU row of thread 4, i.e., the decoder / encoder begins processing CTU33.

[0212] The present disclosure further discloses an encoder including processing circuitry for performing the video processing method or coding method of the present disclosure.

[0213] The present disclosure further discloses a decoder including processing circuitry for performing the video processing method or coding method of the present disclosure.

[0214] The present disclosure further discloses a computer program product comprising program code for performing the video processing method or coding method of the present disclosure.

[0215] The present disclosure further discloses a computer-readable storage medium having stored thereon computer instructions that, when executed by one or more processors, cause the one or more processors to perform a video processing method or coding method of the present disclosure. The computer-readable storage medium may be non-transitory or transitory.

[0216] The present disclosure further discloses a decoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processors, configuring the decoder to perform a video processing method or coding method of the present disclosure.

[0217] The present disclosure further discloses an encoder including one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, the programming, when executed by the processors, configuring the encoder to perform a video processing method or coding method of the present disclosure.

[0218] The initialization process for the HMVP list is described in the VVC (Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, Multipurpose Video Coding (Draft 6)) generic slice data syntax, and section 7.3.8.1 of VVC states the following:

[0219] [Table 1]

[0220] Here, j%BrickWidth[SliceBrickIdx[i]])==0 means that the CTU with index j is the starting CTU of the CTU row, and NumHmvpCand=0 means that the quantity of candidates in the HMVP list is set to 0, in other words, the HMVP list is emptied.

[0221] The update process for the HMVP list is described in Section 8.5.2.16 of VVC (Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11, Multipurpose Video Coding (Draft 6)), which states: The inputs to this process are: Luma motion vectors mvL0 and mvL1 with 1 / 16 fractional sample precision, Reference indices refIdxL0 and refIdxL1, Prediction list usage flags predFlagL0 and predFlagL1, Biprediction weight index bcwIdx. The MVP candidate hMvpCand consists of luma motion vectors mvL0 and mvL1, reference indices refIdxL0 and refIdxL1, prediction list usage flags predFlagL0 and predFlagL1, and bi-prediction weight index bcwIdx. The candidate list HmvpCandList is modified using the candidate hMvpCands by the following ordered steps: The variable identicalCandExist is set equal to FALSE and the variable removeIdx is set equal to 0. If NumHmvpCand is greater than 0, then for each index hMvpIdx with hMvpIdx=0..NumHmvpCand-1, the following steps are applied until identicalCandExist is equal to TRUE: When hMvpCand is equal to HmvpCandList[hMvpIdx], identicalCandExist is set equal to TRUE and removeIdx is set equal to hMvpIdx. The candidate list HmvpCandList is updated as follows: If identicalCandExist is equal to TRUE or NumHmvpCand is equal to 5, the following applies: For each index i, with i=(removeIdx+1)..(NumHmvpCand-1), HmvpCandList[i-1] is set equal to HmvpCandList[i]. HmvpCandList[NumHmvpCand-1] is set equal to mvCand. Otherwise (identicalCandExist equals FALSE and NumHmvpCand is less than 5), the following applies: HmvpCandList[NumHmvpCand++] is set equal to mvCand.

[0222] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates the transfer of computer programs between each other, for example, according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) communication media, such as signal waves or carrier waves. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0223] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair wire, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair wire, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of medium. Note, however, that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory, tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically replicate data magnetically and discs replicate data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0224] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein, may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided in dedicated hardware and / or software modules configured to encode and decode, or may be incorporated into a combined codec. Also, these techniques may be implemented entirely in one or more circuit or logic elements.

[0225] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). Various components, modules, or units are described in this disclosure to highlight functional aspects of devices configured to perform the disclosed techniques, but do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined into a codec hardware unit or may be provided by a collection of interoperable hardware units, including one or more processors as described above, along with suitable software and / or firmware. [Explanation of symbols]

[0226] 10 Coding Systems 12 Source Devices 13 Encoded Picture 14 Destination Device 16 Picture Source 17 Picture Data 18 Pre-treatment Unit 19 Preprocessed Picture Data 20 Encoder 21 Encoded Picture Data 22 Communication Interface 28 Communication Interface 30 Decoder 31 Decoded Pictures 32 Post Processors 33 Post-processed pictures 34 Display Devices 40 Video Coding System 41 Imaging Device 42 Antenna 43 processors 44 Memory Store 45 Display Devices 46 Processing Unit 47 Logic Circuits 110 Inverse Quantization Unit 112 Inverse Transformation Processing Unit 114 Reconstruction Unit 116 buffers 120 Loop Filter 130 Decoded Picture Buffer 131 decoded pictures 144 Inter Prediction Units 154 intra prediction units 201 Pictures 202 Input 203 Picture Block 204 Residual Calculation Unit 205 Residual Blocks 206 Conversion Processing Unit 207 Conversion Factor 208 quantization units 209 Quantization Coefficients 210 Inverse Quantization Unit 211 Dequantization Factors 212 Inverse Transformation Processing Unit 213 Inverse Transform Block 214 Reconstruction Unit 215 reconstructed blocks 216 buffers 220 Loop Filter 221 Filtered Blocks 230 Decoded Picture Buffer 231 decoded pictures 244 Inter Prediction Units 245 Inter-Predicted Blocks 254 intra prediction units 255 intra predicted blocks 260 Prediction Processing Unit 262 Mode Selection Unit 265 predicted blocks 270 Entropy Coding Unit 272 output 304 Entropy Decoding Unit 309 Quantization Coefficients 310 Inverse Quantization Unit 311 Decoded Picture 312 Inverse Transformation Processing Unit 313 Inverse Transform Block 314 Reconstruction Unit 315 reconstructed blocks 316 buffers 320 Loop Filter 321 Filtered Blocks 330 Decoded Picture Buffer 344 Inter Prediction Unit 354 Intra Prediction Units 360 Prediction Processing Unit 362 Mode Selection Unit 365 predicted blocks 400 Video Coding Device 410 Inlet Port 420 Tx / Rx 430 processor 440 Tx / Rx 450 outlet port 460 memory 470 Coding Module 500 devices 502 processor 504 memory 506 Data 508 Operating Systems 510 Application Program 512 Bus 514 Secondary storage device 518 Display 520 Image sensing device 522 Voice Sensing Devices 1100 Video Processing Unit 1110 Initialization Unit 1120 Processing Unit

Claims

1. 1. A video processing method comprising: determining whether a current coding tree unit (CTU) is a start CTU of a current CTU row of a picture area, the picture area consisting of a plurality of CTU rows; initializing a history-based motion vector prediction (HMVP) list for a current CTU row when the current CTU is the starting CTU of the current CTU row; processing the current CTU row based on the HMVP list; initializing an HMVP list for each of the plurality of CTU rows, excluding the current CTU row; A method comprising:

2. The method of claim 1 , wherein the number of candidate motion vectors in the initialized HMVP list is zero.

3. The method of claim 1 or 2, wherein the current CTU row is any one of the plurality of CTU rows.

4. The step of processing the current CTU row based on the HMVP list includes: applying a prediction process to the current CTU in the current CTU row; updating the initialized HMVP list based on the prediction result of the current CTU to obtain an updated HMVP list; applying a prediction process to a second CTU in the current CTU row based on the updated HMVP list; 4. The method of claim 1, comprising:

5. The method of claim 4 , further comprising updating the updated HMVP list based on the processed second CTU.

6. The HMVP list for the current CTU row is as follows: Empty the HMVP list for the current CTU row. The method according to any one of claims 1 to 5, wherein the first and second inputs are initialized.

7. A computer program for causing a computer to carry out the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a program recorded thereon, the program causing a computer to perform the method of any one of claims 1 to 6.

9. A decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor; wherein the programming, when executed by the processor, configures the decoder to perform the method of any one of claims 1 to 6. decoder.

10. 1. An encoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor; wherein the programming, when executed by the processor, configures the encoder to perform the method of any one of claims 1 to 6. Encoder.

11. an encoder configured to obtain a bitstream by performing the method according to any one of claims 1 to 6; a storage configured to store the bitstream; and a transmitter configured to transmit the bitstream. Encoding device.

Citation Information

Patent Citations

  • Coded signal demultiplexer / Multiplexer, coded signal demultiplexing / Multiplexing method and medium for recording coded signal demultiplexing / Multiplexing program

    JP2002223441A

  • Method and apparatus for history-based motion vector prediction

    WO2020018297A1