Intra-mode JVET compilation method

By identifying and instantiating a subset of JVET intra-predicted compilation patterns and using appropriate encoding methods, the problems of intra-mode compilation burden and bandwidth in the prior art are solved, and more efficient compilation efficiency and video quality are achieved.

CN115174913BActive Publication Date: 2025-05-09ARRIS ENTERPRISES LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210748079.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-07-24
Filing Date
2018-07-24
Publication Date
2025-05-09
Estimated Expiration
2038-07-24

AI Technical Summary

Technical Problem

Intra-mode compilation in existing JVET standards has problems with compilation burden and high bandwidth, especially when dealing with MPM lists and selected mode sets.

Method used

The balancing of intra prediction modes is achieved by defining a unique set of intra prediction compilation modes, identifying and instantiating a subset of MPM intra prediction compilation modes, a subset of selected intra prediction compilation modes, and a subset of unselected intra prediction compilation modes. A subset of the MPM intra-predicted compilation modes is encoded using truncated unary binarization and the selected set of modes is encoded with a 4-bit fixed-length code.

Benefits of technology

Reduces the compilation burden and bandwidth related to intra-mode compilation, and improves compilation efficiency and video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115174913B_ABST
    Figure CN115174913B_ABST
Patent Text Reader

Abstract

The present invention relates to an intra-frame mode JVET coding method. A method for JVET segmentation of a video coding block, wherein an MPM set includes a set other than 6 intra-frame prediction coding modes, and can be encoded using truncated unary binarization, 16 selected intra-frame prediction coding modes can be encoded using 4-bit fixed-length codes, and the remaining unselected coding modes can be encoded using truncated binary coding, and a JVET coding tree unit can be encoded into a root node in a quadtree plus binary tree (QTBT) structure, which can have a quadtree branching from the root node and a binary tree branching from a leaf node of each quadtree using asymmetric binary partitioning to separate the coding unit represented by the quadtree leaf node into a child node, which represents the child node as a leaf node in the binary tree branching from the quadtree leaf node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese invention patent application “Intra-mode JVET coding method” with PCT application number PCT / US2018 / 043438, international application date July 24, 2018, and Chinese application number 201880049860.0, which entered the Chinese national phase on January 23, 2020.

[0002] Priority declaration

[0003] This application claims priority under 35 U.S.C. § 119(e) from earlier-filed U.S. Provisional Application Serial No. 62 / 536,072, filed on July 24, 2017, the entire contents of which are incorporated herein by reference. Technical Field

[0004] The present disclosure relates to the field of video coding, and more particularly to efficient intra mode coding. Background Art

[0005] Technical improvements in evolving video coding standards illustrate the trend of increasing coding efficiency to achieve higher bit rates, higher resolutions, and better video quality. The Joint Video Exploration Group is developing a new video coding scheme called JVET. Similar to other video coding schemes like HEVC (High Efficiency Video Coding), JVET is a block-based hybrid space-time prediction coding scheme. However, relative to HEVC, JVET includes many modifications to the bitstream structure, syntax, constraints, and mappings for generating decoded pictures. JVET has been implemented in the Joint Exploration Model (JEM) encoder and decoder.

[0006] There are a total of 67 intra prediction modes described in the current JVET standard, including planar, DC mode, and 65 angular intra modes. In order to efficiently code these 67 modes, all intra modes are subdivided into three sets, including 6 most probable mode (MPM) sets, 16 selection mode sets, and 45 non-selection mode sets.

[0007] Six MPMs are derived from the modes of available neighboring blocks, derived intra modes, and the default intra mode. Figure 1aThe intra modes of the five neighboring blocks of the current block are depicted in . They are Left (L), Top (A), Bottom Left (BL), Top Right (AR), and Top Left (AL), and they are used to form the MPM list of the current block. The initial MPM list is formed by inserting the five neighboring intra modes as well as the plane mode and the DC mode into the MPM list. A pruning process is used to remove repeated modes so that only unique modes can be included in the MPM list. The order in which the initial modes are included is: Left, Top, Plane, DC, Bottom Left, Top Right, and then Top Left.

[0008] If the MPM list is incomplete, the derived modes are added; these intra modes are obtained by adding -1 or +1 to the angle modes already included in the MPM list. If the MPM list is still incomplete, the default modes are added in the following order: vertical, horizontal, mode 2, and diagonal mode. As a result of this process, a unique list of 6 MPM modes is generated.

[0009] For entropy compilation of 6 MPMs, Figure 1b The truncated unary binarization shown in is currently used. The first three bins of the MPM pattern are compiled by a context that depends on the MPM pattern associated with the bin currently being signaled. The MPM patterns are classified into one of three categories: (a) primarily horizontal patterns (i.e., the number of MPM patterns is less than or equal to the number of patterns in the diagonal direction), (b) primarily vertical patterns (i.e., the number of MPM patterns is greater than the number of patterns in the diagonal direction), and (c) non-angular (DC and planar) classes. Therefore, based on this classification, three contexts are used to signal the MPM index.

[0010] The compilation for selecting the remaining 61 non-MPMs is performed as follows. The 61 non-MPMs are first divided into two sets: a selected mode set and an unselected mode set. The selected mode set contains 16 modes, and the rest (45 modes) are assigned to the unselected mode set. The mode set to which the current mode belongs is indicated in the bitstream by a flag. If the mode to be indicated is in the selected mode set, the selected mode is signaled with a 4-bit fixed length code, and if the mode to be indicated is from the unselected set, the selected mode is signaled with a truncated binary code. By way of example, the following selected mode set is generated by subsampling the 61 non-MPM modes:

[0011] Selected mode set = {0, 4, 8, 12, 16, 20...60}

[0012] Unselected pattern set = {1, 2, 3, 5, 6, 7, 9, 10...59}

[0013] In the following Figure 1b The current JVET intra mode compilation is summarized in .

[0014] like Figure 1b As shown in , the last two entries of the MPM list require six bins, which is the same number of bins assigned to the 16 selected modes. For the last two modes on the MPM list, this design has no advantage in terms of compilation performance. Moreover, because the first three bins of the MPM pattern are compiled using context-based entropy compilation, the complexity of compiling the six bins of the MPM pattern is higher than the complexity of compiling the six bins of the selected pattern.

[0015] What is needed is a system and method for reducing the coding burden and bandwidth associated with intra-mode coding. Summary of the invention

[0016] The present disclosure provides a method for video coding for JVET intra prediction, including defining a set of unique intra prediction coding modes, which may be 67 modes in some embodiments, and identifying and instantiating a subset of unique MPM intra prediction coding modes from the set of unique intra prediction coding modes in a memory, which may be 5 or less of 7 or more in some embodiments. The method also provides that a subset of unique selected intra prediction coding modes from a set of unique intra prediction coding modes other than the subset of the unique MPM intra prediction coding modes is identified and instantiated in a memory, which may include 16 coding modes, and a subset of unique unselected intra prediction coding modes from a set of unique intra prediction coding modes other than the subset of the unique MPM intra prediction coding modes and other than the subset of the unique selected intra prediction coding modes is identified and instantiated in a memory, constituting a balance of intra prediction modes. Then, the subset of unique MPM intra prediction coding modes is encoded using truncated unary binarization.

[0017] The present disclosure also provides a video coding system for JVET intra-frame prediction. In some embodiments, the system may include the following steps: instantiating a set of 67 unique intra-frame prediction coding modes in a memory; instantiating a subset of unique MPM intra-frame prediction coding modes from the set of unique intra-frame prediction coding modes in a memory; instantiating a subset of 16 unique selected intra-frame prediction modes from the set of unique intra-frame prediction coding modes other than the subset of unique MPM intra-frame prediction coding modes in a storage; instantiating a subset of unique unselected intra-frame prediction coding modes from the set of unique intra-frame prediction coding modes other than the subset of unique MPM intra-frame prediction coding modes and other than the subset of unique selected intra-frame prediction coding modes in a memory; encoding the subset of unique MPM intra-frame prediction coding modes using truncated unary binarization; and encoding the subset of 16 unique selected intra-frame prediction coding modes using 4 bits of a fixed-length code. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The invention is explained in further detail with the help of the accompanying drawings, in which:

[0019] Figure 1a Depicts the current coding block and associated neighboring blocks.

[0020] Figure 1b Table depicting the current JVET compilation for intra mode prediction.

[0021] Figure 1c Describes dividing a frame into multiple coding tree units (CTUs).

[0022] Figure 2 Depicted is an exemplary partitioning of a CTU into coding units (CUs) using quadtree partitioning and symmetric binary partitioning methods.

[0023] Figure 3 Depiction Figure 2 The partitioned quadtree plus binary tree (QTBT) representation.

[0024] Figure 4 Depicts four possible types of asymmetric binary partitioning of a CU into two smaller CUs.

[0025] Figure 5 Depicted are exemplary partitionings of a CTU into CUs using quadtree partitioning, symmetric binary partitioning, and asymmetric binary partitioning.

[0026] Figure 6 Depiction Figure 5 The segmented QTBT representation.

[0027] Figure 7Depicting a simplified block diagram for CU coding in a JVET encoder.

[0028] Figure 8 Depicts the 67 possible intra prediction modes for the luma component in JVET.

[0029] Fig. 9 Depicting a simplified block diagram for CU coding in a JVET encoder.

[0030] Fig.10 An embodiment of a method of CU coding in a JVET encoder is described.

[0031] Fig.11 Depicting a simplified block diagram for CU coding in a JVET encoder.

[0032] Fig.12 Depicts a simplified block diagram for CU decoding in a JVET decoder.

[0033] Fig.13 Depicting an alternative simplified block diagram of JVET coding for intra mode prediction.

[0034] Fig.14 Table depicting alternative JVET compilation for intra mode prediction.

[0035] Fig.15 Embodiments of a computer system adapted for and / or configured to process a method of CU compilation are depicted.

[0036] Fig.16 An embodiment of an encoder / decoder system for CU coding / decoding in a JVET encoder / decoder is depicted. DETAILED DESCRIPTION

[0037] FIG. 1 depicts the partitioning of a frame into a plurality of coding tree units (CTUs) 100. A frame may be an image in a video sequence. A frame may include a matrix or a collection of matrices having pixel values ​​representing intensity measures in an image. Thus, a collection of these matrices may generate a video sequence. Pixel values ​​may be defined to represent color and brightness in full-color video coding, where pixels are partitioned into three channels. For example, in a YCbCr color space, a pixel may have a brightness value Y representing grayscale intensity in an image and two chrominance values, Cb and Cr, representing how different the color is from gray to blue and red. In other embodiments, pixel values ​​may be represented by values ​​in different color spaces or models. The resolution of a video may determine the number of pixels in a frame. Higher resolutions may mean more pixels and better image clarity, but may also result in higher bandwidth, storage, and transmission requirements.

[0038] JVET can be used to encode and decode frames of a video sequence. JVET is a video coding scheme being developed by the Joint Video Exploration Group. Versions of JVET have been implemented in the JEM (Joint Exploration Model) encoder and decoder. Similar to other video coding schemes like HEVC (High Efficiency Video Coding), JVET is a block-based hybrid space-time prediction coding scheme. During coding with JVET, the frame is first divided into square blocks called CTU 100, as shown in Figure 1. For example, CTU 100 can be a block of 128×128 pixels.

[0039] Figure 2 Depicts an exemplary partitioning of a CTU 100 into CUs 102. Each CTU 100 in a frame may be partitioned into one or more CUs (coding units) 102. CUs 102 may be used for prediction and transformation as described below. Unlike HEVC, in JVET, CUs 102 may be rectangular or square and may be coded without further partitioning into prediction units or transform units. CUs 102 may be as large as their root CTUs 100, or may be smaller subdivisions of the root CTUs 100 as small as 4×4 blocks.

[0040] In JVET, CTU 100 may be partitioned into CU 102 according to a quadtree plus binary tree (QTBT) scheme, where CTU 100 may be recursively separated into square blocks according to a quadtree, and then the square blocks may be recursively separated horizontally or vertically according to a binary tree. Parameters such as CTU size, minimum size for quadtree and binary tree leaf nodes, maximum size for binary tree root node, and maximum depth for binary tree may be set to control separation according to QTBT.

[0041] In some embodiments, JVET may restrict binary partitioning in the binary tree portion of the QTBT to symmetric partitioning, where a block may be divided into two halves vertically or horizontally along a midline.

[0042] By way of non-limiting example, Figure 2 A CTU 100 is shown partitioned into CUs 102 by solid lines indicating quadtree splitting and dashed lines indicating symmetric binary tree splitting. As shown, binary splitting allows for symmetric horizontal and vertical splitting to define the structure of a CTU and its subdivision into CUs.

[0043] Figure 3 Show Figure 2The quadtree root node represents CTU 100, and each child node in the quadtree portion represents one of the four square blocks separated from the parent square block. The square blocks represented by the quadtree leaf nodes can then be symmetrically divided zero or more times using a binary tree, where the quadtree leaf node is the root node of the binary tree. At each level of the binary tree portion, the blocks can be divided symmetrically, vertically, or horizontally. A flag set to "0" indicates that the block is separated horizontally symmetrically, while a flag set to "1" indicates that the block is separated vertically symmetrically.

[0044] In other embodiments, JVET may allow symmetric binary partitioning or asymmetric binary partitioning in the binary tree portion of the QTBT. Asymmetric motion partitioning (AMP) is allowed in different contexts in HEVC when partitioning prediction units (PUs). However, for partitioning of CU 102 in JVET according to the QTBT structure, asymmetric binary partitioning may result in improved partitioning relative to symmetric binary partitioning when the relevant region of the CU 102 is not positioned on either side of a midline passing through the center of the CU. By way of non-limiting example, when the CU 102 depicts one object close to the center of the CU and another object at the side of the CU 102, the CU 102 may be partitioned asymmetrically to place each object in a separate smaller CU 102 of a different size.

[0045] Figure 4 Four possible types of asymmetric binary partitions are depicted in which a CU 102 is separated into two smaller CUs 102 along a line passing through the length or height of the CU 102 such that one of the smaller CUs 102 is 25% of the size of the parent CU 102 and the other is 75% of the size of the parent CU 102 . Figure 4 The four types of asymmetric binary splits shown in allow a CU 102 to be split along a line that starts 25% of the way from the left side of the CU 102, 25% of the way from the right side of the CPU 102, 25% of the way from the top of the CPU 102, or 25% of the way from the bottom of the CPU 102. In alternative embodiments, the asymmetric split line at which the CU 102 is split may be positioned at any other location so that the CU 102 is not symmetrically divided into two halves.

[0046] Figure 5 A non-limiting example of a CTU 100 partitioned into a CU 102 using a scheme that allows both symmetric binary partitioning and asymmetric binary partitioning in the binary tree portion of a QTBT is depicted. Figure 5 The dashed line shows the asymmetric binary split line, where the Figure 4 The parent CU 102 is separated using one of the split types shown in .

[0047] Figure 6 Show Figure 5 The segmented QTBT representation of Figure 6 , the two solid lines extending from the nodes indicate a symmetric partition in the binary tree portion of the QTBT, while the two dashed lines extending from the nodes indicate an asymmetric partition in the binary tree portion.

[0048] Syntax may be encoded in the bitstream that indicates how the CTU 100 is partitioned into the CUs 102. By way of non-limiting example, syntax may be encoded in the bitstream that indicates which nodes are separated by quadtree partitioning, which nodes are separated by symmetric binary partitioning, and which nodes are separated by asymmetric binary partitioning. Similarly, syntax may be encoded in the bitstream for nodes separated by asymmetric binary partitioning that indicates which type of asymmetric binary partitioning is used, such as Figure 4 One of the four types shown in .

[0049] In some embodiments, the use of asymmetric partitioning can be limited to partitioning CU 102 at the leaf nodes of the quadtree portion of the QTBT. In these embodiments, the CU 102 at the child node separated from the parent node using quadtree partitioning in the quadtree portion can be the final CU 102, or they can be further separated using quadtree partitioning, symmetric binary partitioning, or asymmetric binary partitioning. The child nodes in the binary tree portion separated using symmetric binary partitioning can be the final CU 102, or they can be further recursively separated one or more times using only symmetric binary partitioning. In the case where further separation is not allowed, the child nodes in the binary tree portion separated from the QT leaf node using asymmetric binary partitioning can be the final CU 102.

[0050] In these embodiments, limiting the use of asymmetric partitioning to separating quad leaf nodes can reduce search complexity and / or limit overhead bits. Because only quad leaf nodes can be separated by asymmetric partitioning, the use of asymmetric partitioning can directly indicate the end of the QT partial branch without the need for other syntax or further signaling. Similarly, because the asymmetric partitioned nodes cannot be further separated, the use of asymmetric partitioning on a node can also directly indicate that its asymmetric partitioned child node is the final CU 102 without the need for other syntax or further signaling.

[0051] In alternative embodiments, such as when limiting search complexity and / or limiting the number of overhead bits becomes insignificant, asymmetric partitioning may be used to separate nodes generated by quadtree partitioning, symmetric binary partitioning, and / or asymmetric binary partitioning.

[0052] After quadtree splitting and binary tree splitting using any of the above QTBT structures, the blocks represented by the leaf nodes of the QTBT represent the final CU 102 to be coded, such as coding using inter prediction or intra prediction. For slices or full frames coded with inter prediction, different partitioning structures can be used for luma and chroma components. For example, for inter slices, the CU 102 can have coding blocks (CBs) for different color components, such as one luma CB and two chroma CBs. For slices or full frames coded by intra prediction, the partitioning structures for luma and chroma components can be the same.

[0053] In an alternative embodiment, JVET may use a two-level coding block structure as an alternative or extension to the above-mentioned QTBT partitioning. In the two-level coding block structure, the CTU 100 may first be partitioned into basic units (BUs) at a high level. The BUs may then be partitioned into multiple operation units (OUs) at a lower level.

[0054] In an embodiment employing a two-level coding block structure, at a high level, the CTU 100 may be partitioned into BUs according to one of the above-described QTBT structures or according to a quadtree (QT) structure such as that used in HEVC in which a block can be separated into only four equally sized sub-blocks. By way of non-limiting example, the CTU 100 may be partitioned into BUs according to the above-described QTBT structures. Figure 5-Figure 6 The described QTBT structure partitions CTU 102 into BUs so that leaf nodes in the quadtree portion can be separated using quadtree partitioning, symmetric binary partitioning, or asymmetric binary partitioning. In this example, the final leaf node of the QTBT can be a BU instead of a CU.

[0055] At a lower level in the two-level coding block structure, each BU partitioned from the CTU 100 may be further partitioned into one or more OUs. In some embodiments, when the BU is square, it may be separated into OUs using quadtree partitioning or binary partitioning, such as symmetric or asymmetric binary partitioning. However, when the BU is not square, it may only be separated into OUs using binary partitioning. Limiting the partitioning types that can be used for non-square BUs may limit the number of bits used to signal the partitioning type used to generate the BU.

[0056] Although the following discussion describes coding CU 102, in an embodiment using a two-level coding block structure, BUs and OUs may be coded instead of CU 102. By way of non-limiting example, a BU may be used for a higher level coding operation such as intra prediction or inter prediction, while a smaller OU may be used for a lower level coding operation such as transform and generating transform coefficients. Thus, the syntax for a BU to be coded indicates whether it is coded by intra prediction or inter prediction, or information identifying a specific intra prediction mode or motion vector is used to encode the BU. Similarly, the syntax of an OU may identify a specific transform operation or quantized transform coefficients used to code the OU.

[0057] Figure 7 Depicted is a simplified block diagram for CU coding in a JVET encoder. The main stages of video coding include segmentation to identify CU 102 as described above, followed by encoding CU 102 using prediction at 704 or 706, generating residual CU 710 at 708, transforming at 712, quantizing at 716, and entropy coding at 720. Figure 7 The encoder and encoding process illustrated in FIG. 1 also includes a decoding process, which will be described in more detail below.

[0058] Considering the current CU 102, the encoder can use intra prediction in space at 704 or inter prediction in time at 706 to obtain the predicted CU 702. The basic idea of ​​predictive coding is to send a difference or residual signal between the original signal and the prediction of the original signal. At the receiver side, the original signal can be reconstructed by adding the residual and the prediction, as will be described below. Because the difference signal has a lower correlation than the original signal, fewer bits are required for its transmission.

[0059] A slice coded entirely with an intra-predicted CU 102, such as an entire picture or a portion of a picture, may be an I slice, which may be decoded without reference to other slices, and as such, may be a possible point at which decoding may begin. A slice coded with at least some inter-predicted CUs may be a predicted (P) or bi-predicted (B) slice that may be decoded based on one or more reference pictures. A P slice may use intra-prediction and inter-prediction with previously coded slices. For example, by using inter-prediction, a P slice may be further compressed compared to an I slice, but previously coded slices may need to be coded to code them. A B slice may be coded using intra-prediction or inter-prediction that applies interpolated predictions from two different frames, using data from previous and / or subsequent slices, thereby increasing the accuracy of the motion estimation process. In some cases, P slices and B slices may also be coded or alternatively coded using intra-block copying, where data from other portions of the same slice is used.

[0060] As will be discussed below, intra prediction or inter prediction may be performed based on the reconstructed CU 734 from a previously coded CU 102 , such as a neighboring CU 102 or a CU 102 in a reference picture.

[0061] When the CU 102 is spatially coded using intra prediction at 704, an intra prediction mode may be found that best predicts pixel values ​​of the CU 102 based on samples from neighboring CUs 102 in the picture.

[0062] When encoding the luma component of a CU, the encoder can generate a list of candidate intra prediction modes. While HEVC has 35 possible intra prediction modes for luma components, in JVET, there are 67 possible intra prediction modes for luma components. These include a planar mode that uses a three-dimensional plane of values ​​generated from neighboring pixels, a DC mode that uses values ​​averaged from neighboring pixels, and an in-plane mode that uses values ​​copied from neighboring pixels along the indicated direction. Figure 8 65 directional modes are shown in FIG.

[0063] When generating a list of candidate intra prediction modes for the luma component of a CU, the number of candidate modes on the list may depend on the size of the CU. The candidate list may include: a subset of HEVC's 35 modes with the lowest SATD (Sum of Absolute Transform Difference) cost; a new directional mode added to the JVET adjacent to the candidate found from the HEVC mode; and a mode from a set of six most probable modes (MPMs) for CU 102 identified based on the intra prediction modes for previously coded neighboring blocks and the default mode list.

[0064] When coding the chroma components of a CU, a list of candidate intra prediction modes may also be generated. The candidate mode list may include a mode generated using a cross-component linear model projection from luma samples, an intra prediction mode found for a luma CB in a specific co-located position in a chroma block, and a chroma prediction mode previously found for a neighboring block. The encoder may find the candidate mode with the lowest rate-distortion cost in the list and use these intra prediction modes when coding the luma and chroma components of the CU. Syntax may be coded in the bitstream indicating the intra prediction mode used to code each CU 102.

[0065] After the best intra prediction modes for CU 102 have been selected, the encoder can use those modes to generate the predicted CU 402. When the selected mode is a directional mode, a 4-tap filter can be used to improve directional accuracy. A boundary prediction filter, such as a 2-tap or 3-tap filter, can be used to adjust the columns or rows at the top or left of the prediction block.

[0066] The predicted CU 702 may be further smoothed using a position-dependent intra prediction combining (PDPC) process that uses unfiltered samples of neighboring blocks to adjust the predicted CU 702 generated based on filtered samples of neighboring blocks, or adaptive reference sample smoothing using a 3-tap or 5-tap low-pass filter to process reference samples.

[0067] When CU 102 is temporally coded by inter prediction at 706, a set of motion vectors (MVs) can be found that point to samples in the reference picture that best predict the pixel values ​​of CU 102. Inter prediction exploits temporal redundancy between slices by representing the displacement of a block of pixels in a slice. The displacement is determined based on the pixel values ​​in the previous or next slice through a process called motion compensation. The motion vector and associated reference index indicating the displacement of the pixel relative to a particular reference picture can be provided to the decoder in the bitstream along with the residual between the original pixel and the motion compensated pixel. The decoder can use the residual and signal the motion vector and reference index to reconstruct the pixel block in the reconstructed slice.

[0068] In JVET, the motion vector precision can be stored at 1 / 16 pixel, and the difference between the motion vector and the predicted motion vector of the CU can be coded with quarter-pixel resolution or integer-pixel resolution.

[0069] In JVET, motion vectors may be found for multiple sub-CUs within CU 102 using techniques such as advanced temporal motion vector prediction (ATMVP), spatio-temporal motion vector prediction (STMVP), affine motion compensated prediction, pattern matching motion vector derivation (PMMVD), and / or bidirectional optical flow (BIO).

[0070] Using ATMVP, the encoder can find the temporal vector of CU 102, which points to the corresponding block in the reference picture. The temporal vector can be found based on the motion vector and reference picture found for the previously coded neighboring CU 102. Using the reference block to be pointed to by the temporal vector for the entire CU 102, the motion vector can be found for each sub-CU within the entire CU 102.

[0071] STMVP may find the motion vector of the sub-CU by scaling and averaging the motion vectors found for neighboring blocks previously coded using inter-prediction and the temporal vector.

[0072] Affine motion compensated prediction can be used to predict a field of motion vectors for each sub-CU in a block based on the two control motion vectors found for the corners of the block. For example, a motion vector for a sub-CU can be derived based on the corner motion vectors found for each 4x4 block within CU 102.

[0073] PMMVD can use bidirectional matching or template matching to find the initial motion vector for the current CU 102. Bidirectional matching can look at the current CU 102 and reference blocks in two different reference pictures along the motion trajectory, while template matching can look at the corresponding blocks in the current CU 102 and the reference picture identified by the template. The initial motion vector found for CU 102 can then be refined for each sub-CU separately.

[0074] BIO may be used when inter prediction is performed with bi-directional prediction based on earlier and later reference pictures, and allows finding a motion vector for a sub-CU based on the gradient of the difference between the two reference pictures.

[0075] In some cases, local illumination compensation (LIC) may be used at the CU level to find values ​​for the scaling factor parameters and the offset parameters based on samples of neighboring current CU 102 and corresponding samples of neighboring reference blocks identified by the candidate motion vector. In JVET, the LIC parameters may be modified and signaled at the CU level.

[0076] For some of the above methods, the motion vectors found for each sub-CU of the CU can be signaled to the decoder at the CU level. For other methods such as PMMVD and BIO, motion information is not signaled in the bitstream to save overhead, and the decoder can derive the motion vectors through the same process.

[0077] After motion vectors have been found for CU 102, the encoder can use those motion vectors to generate a predicted CU 702. In some cases, when motion vectors have been found for a single sub-CU, overlapped block motion compensation (OBMC) can be used when generating the predicted CU 702 by combining those motion vectors with motion vectors previously found for one or more neighboring sub-CUs.

[0078] When bidirectional prediction is used, JVET can use decoder-side motion vector modification (DMVR) to find motion vectors. DMVR allows motion vectors to be found based on two motion vectors found for bidirectional prediction using a bidirectional template matching process. In DMVR, a weighted combination of the prediction CU 702 generated with each of the two motion vectors can be found, and the two motion vectors can be modified by replacing them with a new motion vector that best points to the combined prediction CU 702. The two modified motion vectors can be used to generate the final prediction CU 702.

[0079] At 708 , once the prediction CU 702 is found, either by intra prediction at 704 or inter prediction at 706 , the encoder may subtract the prediction CU 702 from the current CU 102 to find the residual CU 710 , as described above.

[0080] The encoder may use one or more transform operations at 712 to convert the residual CU 710 into transform coefficients 714 that express the residual CU 710 in the transform domain, such as using a discrete cosine block transform (DCT transform) to transform the data to the transform domain. Compared to HEVC, JVET allows more types of transform operations, including DCT-II, DST-VII, DST-VII, DCT-VIII, DST-I, and DCT-V operations. The allowed transform operations may be grouped into subsets, and indications of which subsets and which specific operations in those subsets are used may be signaled by the encoder. In some cases, large block size transforms may be used to zero high frequency transform coefficients in CUs 102 that are larger than a certain size, so that only low frequency transform coefficients are maintained for those CUs 102.

[0081] In some cases, after the forward kernel transform, a mode-dependent non-separable secondary transform (MDNSST) may be applied to the low-frequency transform coefficients 714. The MDNSST operation may use a Hypercube-Givens transform (HyGT) based on the rotation data. When used, the encoder may signal an index value identifying a specific MDNSST operation.

[0082] At 716, the encoder may quantize the transform coefficients 714 into quantized transform coefficients 716. The quantization of each coefficient may be calculated by dividing the value of the coefficient by the quantization step size, which is derived from the quantization parameter (QP). In some embodiments, Qstep is defined as 2 (QP-4) / 6 . Quantization can assist in data compression because high precision transform coefficients 714 can be converted into quantized transform coefficients 716 having a limited number of possible values. Thus, quantization of the transform coefficients can limit the amount of bits generated and transmitted by the transform process. However, although quantization is a lossy operation, and quantization losses cannot be recovered, the quantization process presents a trade-off between the quality of the reconstructed sequence and the amount of information required to represent the sequence. For example, a lower QP value can result in better quality decoded video, although a higher amount of data may be required for representation and transmission. Conversely, a high QP value will result in a lower quality reconstructed video sequence, but with lower data and bandwidth requirements.

[0083] JVET may utilize a variance-based adaptive quantization technique that allows each CU 102 to use a different quantization parameter for its compilation or process (rather than using the same frame QP in the compilation of each CU 102 of a frame). The variance-based adaptive quantization technique may adaptively reduce the quantization parameter for certain blocks while increasing the quantization parameter in other blocks. To select a specific QP for a CU 102, the variance of the CU is calculated. In short, if the variance of the CU is higher than the average variance of the frame, a QP higher than the QP of the frame may be set for the CU 102. If the CU 102 exhibits a variance lower than the average variance of the frame, a lower QP may be assigned.

[0084] At 720, the encoder can find the final compressed bits 722 by entropy coding the quantized transform coefficients 718. Entropy coding is intended to remove statistical redundancy of the information to be sent. In JVET, the quantized transform coefficients 718 can be coded using CABAC (context adaptive binary arithmetic coding), which uses a probability metric to remove statistical redundancy. For CU 102 with non-zero quantized transform coefficients 718, the quantized transform coefficients 718 can be converted to binary. Each bit ("bin") of the binary representation can then be encoded using a context model. CU 102 can be decomposed into three regions, each region having its own set of context models for pixels within the region.

[0085] Multiple scan operations may be performed to encode a bin. In the operation of encoding the first three bins (bin0, bin1, and bin2), the index value indicating the context model to be used for the bin may be found by finding the sum of the bin positions in up to five previously coded adjacent quantization transform systems identified by the template.

[0086] The context model can be based on the probability of a bin value being "0" or "1". As the values ​​are compiled, the probabilities in the context model can be updated based on the actual number of "0" and "1" values ​​encountered. While HEVC uses a fixed table to reinitialize the context model for each new picture, in JVET, the probabilities of the context model for a new inter-predicted picture can be initialized based on the context model developed for a previously compiled inter-predicted picture.

[0087] The encoder may generate a bitstream containing entropy coded bits 722 of the residual CU 710, prediction information such as a selected intra prediction mode or motion vector, an indicator of how the CU 102 is partitioned from the CTU 100 according to the QTBT structure, and / or other information about the encoded video. The bitstream may be decoded by a decoder as described below.

[0088] In addition to using the quantized transform coefficients 718 to find the final compression bits 722, the encoder can also use the quantized transform coefficients 718 to generate a reconstructed CU 734 by following the same decoding process that the decoder will use to generate the reconstructed CU 734. Therefore, once the encoder calculates and quantizes the transform coefficients, the quantized transform coefficients 718 can be sent to the decoding loop in the encoder. After quantizing the transform coefficients of the CU, the decoding loop allows the encoder to generate the same reconstructed CU 734 as the CU 734 generated by the decoder during the decoding process. Therefore, the encoder can use the same reconstructed CU 734 that the decoder will use for neighboring CUs 102 or reference pictures when performing intra-frame prediction or inter-frame prediction for a new CU 102. The reconstructed CU 102, the reconstructed slice, or the fully reconstructed frame can be used as a reference for further prediction stages.

[0089] At the decode loop of the encoder (see below for the same operation in the decoder) that obtains the pixel values ​​of the reconstructed image, a dequantization process may be performed. To dequantize a frame, for example, the quantized value of each pixel of the frame is multiplied by a quantization step size, such as (Qstep) described above, to obtain the reconstructed dequantized transform coefficients 726. For example, in the encoder Figure 7 In the decoding process shown in , the quantized transform coefficients 718 of the residual CU 710 may be dequantized at 724 to find dequantized transform coefficients 726. If an MDNSST operation is performed during encoding, the operation may be reversed after dequantization.

[0090] At 728, the dequantized transform coefficients 726 may be inversely transformed to find a reconstructed residual CU 730, such as by applying a DCT to these values ​​to obtain a reconstructed image. At 732, the reconstructed residual CU 730 may be added to the corresponding prediction CU 702 found by intra prediction at 704 or by inter prediction at 706 to facilitate finding a reconstructed CU 734.

[0091] At 736, one or more filters may be applied to the reconstructed data during the decoding process (in the encoder, or as described below, in the decoder) at the picture level or CU level. For example, the encoder may apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). The encoder's decoding process may implement filters to estimate optimal filter parameters that can address potential artifacts in the reconstructed image and send them to the decoder. Such improvements increase the objective and subjective quality of the reconstructed video. In deblocking filtering, pixels near sub-CU boundaries may be modified, while in SAO, pixels in the CTU 100 may be modified using edge offsets or band offset classifications. JVET's ALF may use a filter with a circularly symmetric shape for each 2x2 block. An indication of the size and identity of the filter used for each 2x2 block may be signaled.

[0092] If the reconstructed pictures are reference pictures, they may be stored in a reference buffer 738 for inter prediction of the future CU 102 at 706 .

[0093] During the above steps, JVET allows the use of content adaptive cropping operations to adjust color values ​​to fit the upper and lower cropping boundaries. The cropping boundaries can be changed for each slice, and the parameters identifying the boundaries can be signaled in the bitstream.

[0094] Fig. 9 A simplified block diagram depicting CU coding in a JVET decoder. The JVET decoder may receive a bitstream containing information about an encoded CU 102. The bitstream may indicate how the CU 102 of a picture is partitioned from the CTU 100 according to the QTBT structure. As non-limiting examples, the bitstream may identify how the CU 102 is partitioned from each CTU 100 in the QTBT using quadtree partitioning, symmetric binary partitioning, and / or asymmetric binary partitioning. The bitstream may also indicate prediction information for the CU 102, such as an intra-prediction mode or a motion vector, and bits 902 representing an entropy-coded residual CU.

[0095] At 904, the decoder may use the CABAC context model signaled by the encoder in the bitstream to decode the entropy coded bits 902. The decoder may update the probabilities of the context model using parameters signaled by the encoder in the same manner as they were updated during the encoding process.

[0096] After the entropy encoding is reversed at 904 to find the quantized transform coefficients 906, the decoder may dequantize them at 908 to find the dequantized transform coefficients 910. If an MDNSST operation was performed during encoding, the operation may be reversed by the decoder after dequantization.

[0097] At 912, the dequantized transform coefficients 910 may be inversely transformed to find a reconstructed residual CU 914. At 916, the reconstructed residual CU 914 may be added to a corresponding prediction CU 926 found at 922 using intra prediction or at 924 using inter prediction to facilitate finding a reconstructed CU 918.

[0098] At 920, one or more filters may be applied to the reconstructed data at the picture level or CU level. For example, the decoder may apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). As described above, an in-loop filter located in the decoder loop of the encoder may be used to estimate optimal filter parameters to increase the objective and subjective quality of the frame. These parameters are sent to the decoder to filter the reconstructed frame at 920 to match the filtered reconstructed frame in the encoder.

[0099] After the reconstructed picture has been generated by finding the reconstructed CU 918 and applying the signaled filter, the decoder may output the reconstructed picture as output video 928. If the reconstructed picture is used as a reference picture, it may be stored in a reference buffer 930 for inter-prediction of the future CU 102 at 924.

[0100] Fig.10 An embodiment of a method of CU compilation 1000 in a JVET decoder is depicted. Fig.10 In the embodiment shown in , in step 1002, a coded bitstream 902 may be received, then in step 1004, a CABAC context model associated with the coded bitstream 902 may be determined, and then in step 1006, the determined CABAC context model may be used to decode the coded bitstream 902.

[0101] In step 1008 , quantized transform coefficients 906 associated with the encoded bitstream 902 may be determined, and then dequantized transform coefficients 910 may be determined from the quantized transform coefficients 906 in step 1010 .

[0102] In step 1012, it may be determined whether an MDNSST operation is performed during encoding and / or whether the bitstream 902 includes an indication that the bitstream 902 applies the MDNSST operation. If it is determined that an MDNSST operation is performed during the encoding process or the bitstream 902 includes an indication that the MDNSST operation is applied to the bitstream 902, an inverse MDNSST operation 1014 may be implemented before performing an inverse transform operation 912 on the bitstream 902 in step 1016. Alternatively, the operation 912 may be performed on the bitstream 902 in step 1016 without applying the inverse MDNSST operation in step 1014. The inverse transform operation 912 in step 1016 may determine and / or construct a reconstructed residual CU 914.

[0103] In step 1018, the reconstructed residual CU 914 from step 1016 may be combined with a prediction CU 918. The prediction CU 918 may be one of the intra prediction CU 922 determined in step 1020 and the inter prediction unit 924 determined in step 1022.

[0104] In step 1024, any one or more filters 920 may be applied to the reconstructed CU 914 and output in step 1026. In some embodiments, no filter 920 may be applied in step 1024.

[0105] In some embodiments, in step 1028 , the reconstructed CU 918 may be stored in a reference buffer 930 .

[0106] Fig.11 A simplified block diagram 1100 for CU coding in a JVET encoder is depicted. In step 1102, a JVET coding tree unit may be represented as a root node in a quadtree plus binary tree (QTBT) structure. In some embodiments, the QTBT may have a quadtree branching from the root node and / or a binary tree branching from one or more leaf nodes of the quadtree. The representation from step 1102 may proceed to steps 1104, 1106, or 1108.

[0107] In step 1104, an asymmetric binary split may be employed to separate the represented quadtree node into two blocks of unequal sizes. In some embodiments, the separated blocks may be represented in a binary tree branching from a quadtree node that is a leaf node capable of representing the final coding unit. In some embodiments, further separation is not allowed in the binary tree branching from the quadtree node that is a leaf node representing the final coding unit. In some embodiments, the asymmetric split may separate the coding unit into blocks of unequal sizes, with the first block representing 25% of the quadtree node and the second block representing 75% of the quadtree node.

[0108] In step 1106, quadtree partitioning may be employed to partition the quadtree node represented into four square blocks of equal size. In some embodiments, the separated blocks may be represented as quadtree nodes representing the final coding unit, or may be represented as child nodes that may be partitioned again by quadtree partitioning, symmetric binary partitioning, or asymmetric binary partitioning.

[0109] In step 1108, the quadtree node represented may be separated into two blocks of equal size using quadtree partitioning. In some embodiments, the separated blocks may be represented as quadtree nodes representing final coding units, or may be represented as child nodes that may be further partitioned by quadtree partitioning, symmetric binary partitioning, or asymmetric binary partitioning.

[0110] In step 1110, the child nodes from step 1106 or step 1108 may be represented as child nodes configured to be encoded. In some embodiments, the child nodes may be represented by leaf nodes of a binary tree using JVET.

[0111] In step 1112, the compilation unit from step 1104 or 1110 may be encoded using JVET.

[0112] Fig.12 A simplified block diagram 1200 is depicted for CU decoding in a JVET decoder. Fig.12 In the embodiment depicted in , in step 1202, a bitstream indicating how to partition a coding tree unit into coding units according to a QTBT structure may be received. The bitstream may indicate how to separate quadtree nodes by at least one of quadtree partitioning, symmetric binary partitioning, or asymmetric binary partitioning.

[0113] In step 1204, the coding unit represented by the leaf node of the QTBT structure may be identified. In some embodiments, the coding unit may indicate whether the node uses asymmetric binary partitioning to separate the leaf node from the quadtree. In some embodiments, the coding unit may indicate that the node represents the final coding unit to be decoded.

[0114] In step 1206, the identified coding unit may be decoded using JVET.

[0115] Fig.13 An alternative simplified block diagram of JVET compilation for intra mode prediction 1300 is depicted. Fig.13In the embodiment depicted in , in step 1302, a set of MPMs may be identified and instantiated in memory, then in step 1304, a set of 16 selected patterns may be identified and instantiated in memory, and in step 1304, a balance of 67 patterns may be defined and instantiated in memory. In some embodiments, the set of MPMs may be reduced from a standard set of 6 MPMs. In some embodiments, the set of MPMs may include 5 unique patterns, the selected patterns may include 16 unique patterns, and the set of unselected patterns may include the remaining 46 unselected unique patterns. However, in alternative embodiments, the set of MPMs may include fewer unique patterns, the selected patterns may remain fixed at 16 unique patterns, and the size of the set of unselected unique patterns may be adjusted accordingly to accommodate a total of 67 patterns.

[0116] By way of non-limiting example, in some embodiments, where the set of MPMs includes 5 unique modes instead of six MPMs, the number of bins assigned to the MPM mode may therefore be equal to or less than five bins if truncated unary binarization is used and a new binarization for the 5 MPMs may be utilized. Therefore, in some embodiments, 16 selected modes among the 62 remaining intra modes may be generated by uniformly subsampling these 62 intra modes, and each mode may be encoded by a 4-bit fixed length code. By way of non-limiting example, if it is assumed that the remaining 62 modes are indexed as {0, 1, 2, ..., 61}, the 16 selected modes = {0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 60}. The remaining 46 unselected modes = {1, 2, 3, 5, 6, 7, 9, 10 . . . 59, 61}, where these 46 unselected modes can be encoded using truncated binary codes.

[0117] Fig.14 Description Fig.13 Table 1400 of an alternative JVET compilation for intra mode prediction. Fig.14 In the depicted embodiment, the intra-frame prediction mode 1402 is shown as including 5 MPMs, 16 selected modes, and 46 unselected modes, where the binary string 1404 for the MPM can be encoded using truncated unary binarization, the 16 selected modes can be encoded using a 4-bit fixed-length code, and the 46 unselected modes can be encoded using truncated binary coding.

[0118] exist Fig.13 In an alternative embodiment, six MPMs may be used, but Fig.14As shown in , only the first five MPMs on the MPM list are binarized and compiled using the current context-based method described in the current JVET. The sixth MPM on the MPM list is now considered as one of the 16 selected modes and is compiled with a 4-bit fixed-length code along with the other 15 selected modes.

[0119] As a non-limiting example, if the remaining 61 modes are indexed as {0, 1, 2, ..., 60}, the following 15 selected modes can be obtained by uniformly subsampling the remaining 61 intra modes: the selected mode set can be {0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, where the 15 selected modes plus the sixth MPM are subsampled by a 4-bit fixed-length code. The rows are compiled, such as in the following set: {sixth MPM, 0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, and the set below shows the balance of 46 unselected modes and is compiled into an unselected mode set = {1, 2, 3, 4, 6, 7, 8, 9, 11, 12…49, 51, 52, 53, 54, 56, 57, 58, 59} through truncated binary code.

[0120] exist Fig.13 In yet another alternative embodiment, only the first five MPMs on the MPM list may be binarized, such as Fig.14 , and compiled using the current context-based method described in the current JVET standard. In such an embodiment, the sixth MPM on the MPM list can be considered as one of the 16 selected modes and compiled with the other 15 selected modes using a 4-bit fixed length code. Therefore, the selection of the other 15 selected modes can be established using any known convenient and / or desired selection process. By way of non-limiting example, they can be selected around MPM modes, or around (content-based) statistical popular modes, or around trained or historical popular modes, or using other methods or processes.

[0121] Again, the selection of 5 MPMs is merely a non-limiting example, and in alternative embodiments, the set of MPMs may be further reduced to 4 or 3 MPMs or expanded to more than 6, where there are still 16 selected modes, and a balance of 67 (or other known, convenient, and / or desired total) intra coding modes are included in the set of unselected intra coding modes. That is, embodiments where the total number of intra coding modes is greater or less than 67 are contemplated, where the set of MPMs may contain any known convenient or desired number of MPMs, and the number of selected modes may be any known convenient and / or desired number.

[0122] The execution of the instruction sequences required to practice the embodiments may be performed by Fig.15 1500. In an embodiment, the execution of the sequence of instructions is performed by a single computer system 1500. According to other embodiments, two or more computer systems 1500 coupled by a communication link 1515 can coordinate the execution of the sequence of instructions. Although only one computer system 1500 will be described below, it should be understood that any number of computer systems 1500 can be used to practice the embodiments.

[0123] Now refer to Fig.15 Describing a computer system 1500 according to an embodiment, Fig.15 is a block diagram of the functional components of computer system 1500. As used herein, the term computer system 1500 is used broadly to describe any computing device that can store and independently run one or more programs.

[0124] Each computer system 1500 may include a communication interface 1514 coupled to the bus 1506. The communication interface 1514 provides two-way communication between the computer systems 1500. The communication interface 1514 of each computer system 1500 sends and receives electrical, electromagnetic or optical signals, which include data streams representing various types of signal information, such as instructions, messages and data. The communication link 1515 links one computer system 1500 to another computer system 1500. For example, the communication link 1515 can be a LAN, in which case the communication interface 1514 can be a LAN card, or the communication link 1515 can be a PSTN, in which case the communication interface 1514 can be an integrated services digital network (ISDN) card or a modem, or the communication link 1515 can be the Internet, in which case the communication interface 1514 can be a dial-up, cable or wireless modem.

[0125] The computer system 1500 may send and receive messages, data, and instructions, including programs, i.e., applications, code, through its respective communication links 1515 and communication interfaces 1514. The received program code may be executed by the respective processor 1507 when it is received and / or stored in the storage device 1510 or other associated non-volatile media for later execution.

[0126] In an embodiment, the computer system 1500 operates with a data storage system 1531, for example, a data storage system 1531 including a database 1532 that is easily accessible by the computer system 1500. The computer system 1500 communicates with the data storage system 1531 via a data interface 1533. The data interface 1533 coupled to the bus 1506 sends and receives electrical, electromagnetic, or optical signals, including data streams representing various types of signal information, such as instructions, messages, and data. In an embodiment, the functions of the data interface 1533 may be performed by the communication interface 1514.

[0127] The computer system 1500 includes a bus 1506 or other communication mechanism for transmitting instructions, messages and data, collectively, information, and one or more processors 1507 coupled to the bus 1506 for processing information. The computer system 1500 also includes a main memory 1508, such as a random access memory (RAM) or other dynamic storage device, coupled to the bus 1506 for storing dynamic data and instructions to be executed by the processor 1507. The main memory 1508 may also be used to store temporary data, i.e., variables or other intermediate information, during the execution of instructions by the processor 1507.

[0128] The computer system 1500 may further include a read only memory (ROM) 1509 or other static storage device coupled to the bus 1506 for storing static data and instructions for the processor 1507. A storage device 1510, such as a magnetic disk or optical disk, may also be provided and coupled to the bus 1506 for storing data and instructions for the processor 1507.

[0129] Computer system 1500 may be coupled to a display device 1511, such as but not limited to a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, via bus 1506 for displaying information to a user. Input device 1512, such as alphanumeric and other keys, is coupled to bus 1506 for communicating information and command selections to processor 1507.

[0130] According to one embodiment, the individual computer systems 1500 perform specific operations by their respective processors 1507 executing one or more sequences of one or more instructions contained in the main memory 1508. Such instructions may be read into the main memory 1508 from another computer-usable medium such as ROM 1509 or storage device 1510. Execution of the sequences of instructions contained in the main memory 1508 causes the processor 1507 to perform the processes described herein. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. Therefore, the embodiments are not limited to any specific combination of hardware circuitry and / or software.

[0131] As used herein, the term "computer-usable medium" refers to any medium that provides information or can be used by processor 1507. Such media can take many forms, including but not limited to non-volatile, volatile, and transmission media. Non-volatile media, that is, media that can retain information without power, include ROM 1509, CD ROM, magnetic tape, and magnetic disk. Volatile media, that is, media that cannot retain information without power, include main memory 1508. Transmission media include coaxial cables, copper wire, and optical fiber, including the wires that make up bus 1506. Transmission media can also take the form of carrier waves; that is, electromagnetic waves that can be modulated in frequency, amplitude, or phase to send information signals. In addition, transmission media can take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0132] In the foregoing description, the embodiments have been described with reference to the specific elements of the embodiments. However, it is apparent that various modifications and changes may be made to them without departing from the broader spirit and scope of the embodiments. For example, the reader will appreciate that the specific order and combination of process actions shown in the process flow charts described herein are merely illustrative, and different or additional process actions may be used, or different combinations or orders of process actions may be used to implement the embodiments. Therefore, the description and drawings should be considered illustrative rather than restrictive.

[0133] It should also be noted that the present invention can be implemented in various computer systems. The various techniques described herein can be implemented in hardware or software or a combination of the two. Preferably, the technology is implemented in a computer program executed on a programmable computer, and the programmable computer includes a processor, a storage medium readable by the processor (including volatile and non-volatile memory and / or storage element), at least one input device and at least one output device. The program code is applied to the data input using the input device to perform the above functions and generate output information. The output information is applied to one or more output devices. Each program is preferably implemented in a high-level process or object-oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly language or machine language. In any case, the language can be a compiled language or an interpreted language. Each such computer program is preferably stored on a storage medium or device (e.g., ROM or disk) readable by a general or special programmable computer, for configuring and operating the computer when the storage medium or device is read by the computer to perform the above process. The system can also be considered to be implemented as a computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predetermined manner. Furthermore, the storage element of the exemplary computing application may be a relational or sequential (flat file) type computing database capable of storing data in various combinations and configurations.

[0134] Fig.16 16 is a high-level view of a source device 1612 and a destination device 1610 that may incorporate features of the systems and devices described herein. Fig.16 As shown in , the example video coding system 1610 includes a source device 1612 and a destination device 1616, where, in this example, the source device 1612 generates encoded video data. Therefore, the source device 1612 can be referred to as a video coding device. The destination device 1616 can decode the encoded video data generated by the source device 1612. Therefore, the destination device 1616 can be referred to as a video decoding device. The source device 1612 and the destination device 1616 can be examples of video coding devices.

[0135] Destination device 1616 may receive the encoded video data from source device 1612 via channel 1616. Channel 1616 may include a type of medium or device capable of moving the encoded video data from source device 1612 to destination device 1616. In one example, channel 1616 may include a communication medium that enables source device 1612 to send the encoded video data directly to destination device 1616 in real time.

[0136] In this example, source device 1612 may modulate the encoded video data according to a communication standard such as a wireless communication protocol, and may send the modulated video data to destination device 1616. The communication medium may include a wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or other devices that facilitate communication from source device 1612 to destination device 1616. In another example, channel 1616 may correspond to a storage medium that stores the encoded video data generated by source device 1612.

[0137] exist Fig.16 In the example of , source device 1612 includes a video source 1618, a video encoder 1620, and an output interface 1622. In some cases, output interface 1628 may include a modulator / demodulator (modem) and / or a transmitter. In source device 1612, video source 1618 may include a source such as a video capture device, for example, a camera, a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources.

[0138] The video encoder 1620 can encode captured, pre-captured, or computer-generated video data. Input images can be received by the video encoder 1620 and stored in the input frame memory 1621. The general purpose processor 1623 can load the information from here and perform the encoding. The program for driving the general purpose processor can be obtained from a program such as Fig.16 The general purpose processor may use processing memory 1622 to perform encoding, and the output of the encoding information of the general purpose processor may be stored in a buffer such as output buffer 1626.

[0139] The video encoder 1620 may include a resampling module 1625 that may be configured to code (e.g., encode) video data with a scalable video coding scheme that defines at least one base layer and at least one enhancement layer. As part of the encoding process, the resampling module 1625 may resample at least some of the video data, where the resampling may be performed in an adaptive manner using a resampling filter.

[0140] The encoded video data, e.g., a compiled bitstream, may be sent directly to the destination device 1616 via the output interface 1628 of the source device 1612. Fig.1616, destination device 1616 includes input interface 1638, video decoder 1630, and display device 1632. In some cases, input interface 1628 may include a receiver and / or a modem. Input interface 1638 of destination device 1616 receives encoded video data via channel 1616. The encoded video data may include various syntax elements generated by video encoder 1620 to represent the video data. Such syntax elements may be included with the encoded video data that is sent over a communication medium, stored on a storage medium, or stored in a file server.

[0141] The encoded video data may also be stored on a storage medium or file server for later access by the destination device 1616 for decoding and / or playback. For example, the compiled bitstream may be temporarily stored in an input buffer 1631 and then loaded into a general purpose processor 1633. A program for driving the general purpose processor may be loaded from a storage device or memory. The general purpose processor may use a processing memory 1632 to perform decoding. The video decoder 1630 may also include a resampling module 1635 similar to the resampling module 1625 employed in the video encoder 1620.

[0142] Fig.16 A resampling module 1635 is depicted as being separate from the general purpose processor 1633, but one skilled in the art will appreciate that the resampling function may be performed by a program executed by a general purpose processor, and that the processing in the video encoder may be accomplished using one or more processors. The decoded image may be stored in an output frame buffer 1636 and then sent to an input interface 1638.

[0143] Display device 1638 can be integrated with destination device 1616 or can be external thereto. In some examples, destination device 1616 can include an integrated display device and can also be configured to interface with an external display device. In other examples, destination device 1616 can be a display device. Typically, display device 1638 displays the decoded video data to a user.

[0144] The video encoder 1620 and the video decoder 1630 may operate according to a video compression standard. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) are studying the potential need to standardize future video coding techniques with compression capabilities that significantly exceed the current High Efficiency Video Coding HEVC standard (including its current extensions and near-term extensions for screen content coding and high dynamic range coding). The groups are working together on this exploratory activity under a joint collaboration called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by their experts in the field. The latest results developed by JVET are described in the "Algorithmic Description of the Joint Exploration Test Model 5 (JEM 5)" of JVET-E1001-V2, written by J. Chen, E. Alshina, G. Sullivan, J. Ohm, and J. Boyce.

[0145] Additionally or alternatively, the video encoder 1620 and the video decoder 1630 may operate in accordance with other proprietary or industry standards that function with the disclosed JVET features. Thus, other standards such as the ITU-T H.264 standard, alternatively referred to as MPEG-4, Part 10, Advanced Video Coding (AVC), or extensions of these standards. Thus, while new developments are being made for JVET, the techniques of the present disclosure are not limited to any particular coding standard or technique. Other examples of video compression standards and techniques include MPEG-2, ITU-T H.263, and proprietary or open source compression formats and related formats.

[0146] The video encoder 1620 and the video decoder 1630 may be implemented in hardware, software, firmware, or any combination thereof. For example, the video encoder 1620 and the decoder 1630 may employ one or more processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. When the video encoder 1620 and the decoder 1630 are partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium, and may use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Each of the video encoder 1620 and the video decoder 1630 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in a respective device.

[0147] Aspects of the subject matter described herein may be described in the general context of computer-executable instructions, such as program modules, executed by a computer, such as the general purpose processors 1623 and 1633 described above. Typically, program modules include routines, programs, objects, components, data structures, etc., which perform specific tasks or implement specific abstract data types. Aspects of the subject matter described herein may also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices linked through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including memory storage devices.

[0148] Examples of memory include random access memory (RAM), read-only memory (ROM), or both. The memory can store instructions for performing the above-described techniques, such as source code or binary code. The memory can also be used to store variables or other intermediate information during the execution of instructions to be executed by a processor such as processors 1623 and 1633.

[0149] The storage device may also store instructions for performing the above-mentioned techniques, such as source code or binary code. The storage device may additionally store data used and manipulated by the computer processor. For example, the storage device in the video encoder 1620 or the video decoder 1630 may be a database accessed by the computer system 1623 or 1633. Other examples of storage devices include random access memory (RAM), read-only memory (ROM), hard drive, disk, optical disk, CD-ROM, DVD, flash memory, USB memory card, or any other medium that can be read by a computer.

[0150] The memory or storage device may be an example of a non-transitory computer-readable storage medium for use with or in conjunction with a video encoder and / or decoder. The non-transitory computer-readable storage medium contains instructions for controlling a computer system that is configured to perform the functions described by the specific embodiments. When executed by one or more computer processors, the instructions may be configured to perform the functions described in the specific embodiments.

[0151] In addition, it should be noted that some embodiments have been described as processes that can be depicted as flow charts or block diagrams. Although each can describe the operations as a sequential process, many operations can be performed in parallel or simultaneously. In addition, the order of the operations can be rearranged. The process may have other steps not included in the drawings.

[0152] Particular embodiments may be implemented in a non-transitory computer-readable storage medium for use with or in conjunction with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium contains instructions for controlling a computer system to perform the methods described by the particular embodiments. The computer system may include one or more computing devices. When executed by one or more computer processors, the instructions may be configured to perform the methods described in the particular embodiments.

[0153] As used in the specification herein and throughout the claims that follow, “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Also, as used in the specification herein and throughout the claims that follow, the meaning of “in” includes “in” and “on” unless the context clearly dictates otherwise.

[0154] Although the exemplary embodiments of the present invention have been described in detail with the above-mentioned structural features and / or method action specific language, it is to be understood that those skilled in the art will readily appreciate that many additional modifications can be made in the exemplary embodiments without substantially departing from the novel teachings and advantages of the present invention. In addition, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the above-mentioned specific features or actions. Therefore, these and all such modifications are intended to be included in the scope of the present invention explained in the breadth and scope of the appended claims.

Claims

1. A method for decoding video data from a bitstream, the method comprising: (a) receiving a bitstream indicating how to partition a coding tree unit into coding units; (b) determining a first set of possible modes selectable based on possible mode MPM indices for a current block of the video data, wherein one of the first set of MPMs selectable based on the MPM indices comprises a direct horizontal mode and another one of the first set of MPMs selectable based on the MPM indices comprises a direct vertical mode and another one of the first set of MPMs selectable based on the MPM indices comprises an angular mode, wherein the first set of MPMs comprises only five different modes; (c) deriving (i) an MPM flag and (ii) another index from a bitstream, wherein the MPM flag comprises a total of 1 bit, and at least one of the MPM flag and the another index indicates whether the intra mode used to predict the current block is one of the first set of MPMs; (d) when the at least one of the MPM flag and the another index is used to indicate that the intra mode used to predict the current block is one of the first set of MPMs selectable based on the MPM index, selecting the intra mode of the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs; (e) when the at least one of the MPM flag and the another index indicates that the intra-mode used to predict the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determine a second set of at least one mode and (ii) determine a third set of at least one mode; (f) wherein the first set, the second set, and the third set include different patterns, wherein the combination of the first set, the second set, and the third set includes 67 different patterns; (g) determining an intra-mode for the current block of the second set of the at least one mode based on a first combination of the MPM flag and the another index that does not include any of the MPMs selectable based on the MPM index included in the first set of possible modes; as well as (h) determining an intra-mode for the current block of a third set of the at least one mode based on a second combination of the MPM flag of the first set not including any of the MPMs selectable based on the MPM index included in the first set of possible modes and the other index.

2. A computer-readable storage medium storing instructions which, when executed, cause a processor to: (a) receiving a bitstream indicating how to partition a coding tree unit into coding units; (b) determining, for a current block of video data, a first set of possible modes selectable based on possible mode MPM indices, wherein one of the first set of MPMs selectable based on the MPM indices comprises a direct horizontal mode and another one of the first set of MPMs selectable based on the MPM indices comprises a direct vertical mode and another one of the first set of MPMs selectable based on the MPM indices comprises an angular mode, wherein the first set of MPMs comprises only five different modes; (c) deriving (i) an MPM flag and (ii) another index from a bitstream, wherein the MPM flag comprises a total of 1 bit, and at least one of the MPM flag and the another index indicates whether the intra mode used to predict the current block is one of the first set of MPMs; (d) when the at least one of the MPM flag and the another index is used to indicate that the intra mode used to predict the current block is one of the first set of MPMs selectable based on the MPM index, selecting the intra mode of the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs; (e) when the at least one of the MPM flag and the another index indicates that the intra-mode used to predict the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determine a second set of at least one mode and (ii) determine a third set of at least one mode; (f) wherein the first set, the second set, and the third set include different patterns, wherein the combination of the first set, the second set, and the third set includes 67 different patterns; (g) determining an intra-mode for the current block of the second set of the at least one mode based on a first combination of the MPM flag and the another index that does not include any of the MPMs selectable based on the MPM index included in the first set of possible modes; as well as (h) determining an intra-mode for the current block of a third set of the at least one mode based on a second combination of the MPM flag of the first set not including any of the MPMs selectable based on the MPM index included in the first set of possible modes and the other index.

3. A method for encoding video data by an encoder, the method comprising: (a) providing a bitstream indicating how to partition a coding tree unit into coding units; (b) the bitstream comprises data adapted to determine, for a current block of the video data, a first set of possible modes selectable based on a possible mode MPM index, wherein one of the first set of MPMs selectable based on the MPM index comprises a direct horizontal mode and another one of the first set of MPMs selectable based on the MPM index comprises a direct vertical mode and another one of the first set of MPMs selectable based on the MPM index comprises an angular mode, wherein the first set of MPMs comprises only five different modes; (c) the bitstream comprises data suitable for deriving (i) an MPM flag and (ii) another index from the bitstream, the MPM flag comprising a total of 1 bit, at least one of the MPM flag and the another index indicating whether the intra mode used to predict the current block is one of the first set of MPMs; (d) the bitstream comprises data adapted to select an intra mode of the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs when the at least one of the MPM flag and the another index is used to indicate that the intra mode used to predict the current block is one of the first set of MPMs selectable based on the MPM index; (e) the bitstream comprises data adapted to, when the at least one of the MPM flag and the another index indicates that the intra-mode used to predict the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determine the second set of at least one mode and (ii) determine the third set of at least one mode; (f) the bitstream comprises data suitable for wherein the first set, the second set and the third set comprise different patterns, wherein a combination of the first set, the second set and the third set comprises 67 different patterns; (g) the bitstream comprises data adapted to determine an intra-mode for the current block of the second set of the at least one modes based on a first combination of the MPM flag and the further index not including any of the MPMs of the first set selectable based on the MPM index included in the first set of possible modes; as well as (h) the bitstream comprises data suitable for determining an intra-mode of the current block for a third set of the at least one mode based on a second combination of the MPM flag and the another index of the first set not including any of the MPMs selectable based on the MPM index included in the first set of possible modes.

Citation Information

Patent Citations

  • Video decoder with enhanced cabac decoding

    CN103959782A

  • Residual quad tree (rqt) coding for video coding

    CN104081777A