Intra Mode JVET Compilation Method

By defining a unique set of intra prediction modes and using truncated unary binary and quadtree plus binary tree (QTBT) structures, the problem of high complexity in JVET intra mode compilation is solved, and more efficient compilation performance is achieved.

CN115174914BActive Publication Date: 2025-07-22ARRIS ENTERPRISES LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210748094.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-07-24
Filing Date
2018-07-24
Publication Date
2025-07-22
Estimated Expiration
2038-07-24

AI Technical Summary

Technical Problem

In the existing JVET intra-mode compilation scheme, the compilation burden and bandwidth are high, especially when the MPM list is incomplete, the compilation complexity of the last two modes is insufficient.

Method used

By defining a set of unique intra-predictive compilation patterns, identifying and instantiating a subset of unique MPM intra-predictive compilation patterns, encoding them using truncated unary binarization, and encoding 16 selected patterns using 4-bit fixed-length codes, combining quad-tree plus binary tree (QTBT) structures for segmentation, limiting the use of asymmetric segmentation to reduce search complexity.

Benefits of technology

It effectively reduces the complexity and bandwidth of intra-mode compilation, improves compilation efficiency, and optimizes compilation performance, especially when the MPM list is incomplete.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115174914B_ABST
    Figure CN115174914B_ABST
Patent Text Reader

Abstract

The present invention relates to an in-frame mode JVET coding method. A method for coding JVET segmented video coding blocks, wherein the MPM set includes a set excluding six in-frame prediction coding modes, and can be coded using truncated unary binarization, 16 selected in-frame prediction coding modes can be coded using 4-bit fixed-length codes, and the remaining unselected coding modes can be coded using truncated binary coding, and the JVET coding tree unit can be compiled into a root node in a quadtree plus binary tree (QTBT) structure, which can have a quadtree branching from the root node and a binary tree branching from the leaf nodes of each quadtree using asymmetric binary splitting to separate the coding unit represented by the quadtree leaf node into child nodes, which represent the child nodes as leaf nodes in the binary tree branching from the quadtree leaf node.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the Chinese patent application "Intra Mode JVET Compilation Method" with Chinese application number 201880049860.0, PCT application number PCT / US2018 / 043438, international filing date of July 24, 2018, and entered the Chinese national phase on January 23, 2020.

[0002] Priority Claim

[0003] This application claims priority under 35 U.S.C. § 119(e) to the earlier filed U.S. Provisional Application Serial No. 62 / 536,072, filed July 24, 2017, the entire content of which is incorporated herein by reference. Technical Field

[0004] The present disclosure relates to the field of video coding, and more particularly to efficient intra mode coding. Background Art

[0005] Technical improvements in evolving video coding standards illustrate the trend of increasing coding efficiency to achieve higher bitrates, higher resolutions, and better video quality. The Joint Video Exploration Team is developing a new video coding scheme called JVET. Similar to other video coding schemes such as HEVC (High Efficiency Video Coding), JVET is a block-based hybrid spatio-temporal predictive coding scheme. However, relative to HEVC, JVET includes many modifications to the bitstream structure, syntax, constraints, and mappings for generating decoded pictures. JVET has been implemented in the Joint Exploration Model (JEM) encoder and decoder.

[0006] There are a total of 67 intra prediction modes described in the current JVET standard, including planar, DC mode, and 65 directional angle intra modes. To efficiently code these 67 modes, all intra modes are subdivided into three sets, including a 6 Most Probable Modes (MPM) set, a 16 Selected Modes set, and a 45 Non-Selected Modes set.

[0007] Six MPMs are derived from the modes of available neighboring blocks, derived intra modes, and default intra modes. Figure 1aThe intra modes of 5 neighboring blocks depicting the current block are shown. They are left (L), above (A), bottom - left (BL), top - right (AR), and top - left (AL) respectively, and they are used to form the MPM list of the current block. The initial MPM list is formed by inserting the 5 neighboring intra modes, along with the planar mode and the DC mode, into the MPM list. A pruning process is used to remove duplicate modes so that only unique modes can be included in the MPM list. The order of including the initial modes is: left, above, planar, DC, bottom - left, top - right, and then top - left.

[0008] If the MPM list is incomplete, exported modes are added; these intra modes can be obtained by adding - 1 or + 1 to the angular modes already included in the MPM list. If the MPM list is still incomplete, default modes are added in the following order: vertical, horizontal, mode 2, and diagonal mode. As a result of this process, a unique list of 6 MPM modes is generated.

[0009] For the entropy coding of the 6 MPMs, Figure 1b the truncated unary binarization shown in is currently used. The first three bins of the MPM mode are coded depending on the context of the MPM mode related to the bin currently being signaled. The MPM modes are classified into one of three categories: (a) modes that are mainly horizontal (i.e., the number of MPM modes is less than or equal to the number of modes in the diagonal direction), (b) modes that are mainly vertical (i.e., the number of MPM modes is greater than the number of modes in the diagonal direction), and (c) non - angular (DC and planar) classes. Thus, based on this classification, three contexts are used to signal the MPM index.

[0010] The coding for selecting the remaining 61 non - MPMs is carried out as follows. First, the 61 non - MPMs are divided into two sets: the selected mode set and the unselected mode set. The selected mode set contains 16 modes, and the remaining (45 modes) are assigned to the unselected mode set. The mode set to which the current mode belongs is indicated in the bitstream by a flag. If the mode to be indicated is in the selected mode set, the selected mode is signaled using a 4 - bit fixed - length code, and if the mode to be indicated is from the unselected set, the selected mode is signaled using a truncated binary code. By way of example, the selected mode set is generated by subsampling the 61 non - MPM modes as follows:

[0011] Selected mode set = {0, 4, 8, 12, 16, 20…60}

[0012] Unselected mode set = {1, 2, 3, 5, 6, 7, 9, 10…59}

[0013] In the following Figure 1b the current JVET intra mode coding is summarized.

[0014] As Figure 1b shown, the last two entries of the MPM list require six bins, which is the same number of bins assigned to the 16 selected modes. For the last two modes on the MPM list, this design has no advantage in terms of compilation performance. Also, since the first three bins of the MPM mode are compiled using context-based entropy coding, the complexity of compiling the six bins of the MPM mode is higher than the complexity of compiling the six bins of the selected modes.

[0015] A system and method are needed for reducing the compilation burden and bandwidth associated with intra-mode compilation. SUMMARY OF THE INVENTION

[0016] The present disclosure provides a method for video coding for JVET intra prediction, including defining a set of unique intra prediction coding modes, which can be 67 modes in some embodiments, and identifying and instantiating in memory a subset of unique MPM intra prediction coding modes from the set of unique intra prediction coding modes, which can be 5 or fewer out of 7 or more in some embodiments. The method further provides identifying and instantiating in memory a subset of unique selected intra prediction coding modes from the set of unique intra prediction coding modes other than the subset of unique MPM intra prediction coding modes, which can include 16 coding modes in some embodiments, and identifying and instantiating in memory a subset of unique unselected intra prediction coding modes from the set of unique intra prediction coding modes other than the subset of unique MPM intra prediction coding modes and other than the subset of unique selected intra prediction coding modes, to form a balance of intra prediction modes. Then, the subset of unique MPM intra prediction coding modes is coded using truncated unary binarization.

[0017] The present disclosure also provides a video encoding system for JVET intra prediction. In some embodiments, the system may include the following steps: instantiating a set of 67 unique intra prediction encoding modes in a memory; instantiating a subset of unique MPM intra prediction encoding modes from the set of unique intra prediction encoding modes in the memory; instantiating a subset of 16 unique selected intra prediction modes from the set of unique intra prediction encoding modes except for the subset of unique MPM intra prediction encoding modes in a storage; instantiating a subset of unique unselected intra prediction encoding modes from the set of unique intra prediction encoding modes except for the subset of unique MPM intra prediction encoding modes and except for the subset of unique selected intra prediction encoding modes in the memory; encoding the subset of unique MPM intra prediction encoding modes using truncated unary binarization; and encoding the subset of 16 unique selected intra prediction encoding modes using 4 bits of a fixed length code. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] More details of the present invention are explained with the help of the drawings, wherein:

[0019] Figure 1a Depicts the current encoding block and associated neighboring blocks.

[0020] Figure 1b Depicts a table for the current JVET encoding for intra mode prediction.

[0021] Figure 1c Depicts the partitioning of a frame into multiple encoding tree units (CTUs).

[0022] Figure 2 Depicts an exemplary partitioning of a CTU into encoding units (CUs) using the method of quadtree splitting and symmetric binary splitting.

[0023] Figure 3 Depicts Figure 2 the quadtree plus binary tree (QTBT) representation of the partitioning.

[0024] Figure 4 Depicts four possible types of asymmetric binary splitting for splitting a CU into two smaller CUs.

[0025] Figure 5 Depicts an exemplary partitioning of a CTU into CUs using quadtree splitting, symmetric binary splitting, and asymmetric binary splitting.

[0026] Figure 6 Depicts Figure 5 the QTBT representation of the partitioning.

[0027] Figure 7Depict a simplified block diagram for CU compilation in the JVET encoder.

[0028] Figure 8 Depict 67 possible intra prediction modes for the luma component in JVET.

[0029] Figure 9 Depict a simplified block diagram for CU compilation in the JVET encoder.

[0030] Figure 10 Depict an embodiment of a method for CU compilation in the JVET encoder.

[0031] Figure 11 Depict a simplified block diagram for CU compilation in the JVET encoder.

[0032] Figure 12 Depict a simplified block diagram for CU decoding in the JVET decoder.

[0033] Figure 13 Depict an alternative simplified block diagram for JVET compilation for intra mode prediction.

[0034] Figure 14 Depict a table for alternative JVET compilation for intra mode prediction.

[0035] Figure 15 Depict an embodiment of a computer system suitable for and / or configured to process a method for CU compilation.

[0036] Figure 16 Depict an embodiment of an encoder / decoder system for CU compilation / decoding in the JVET encoder / decoder. Detailed Description

[0037] Figure 1 depicts a frame divided into multiple coding tree units (CTUs) 100. A frame can be an image in a video sequence. A frame can include a matrix or a set of matrices having pixel values representing intensity measures in the image. Thus, a set of these matrices can generate a video sequence. Pixel values can be defined to represent color and luminance in a full-color video encoding, where pixels are divided into three channels. For example, in the YCbCr color space, a pixel can have a luminance value Y representing the gray intensity in the image and two chrominance values, Cb and Cr, representing the degree of color difference from gray to blue and red. In other embodiments, pixel values can be represented by values in different color spaces or models. The resolution of the video can determine the number of pixels in the frame. A higher resolution may mean more pixels and better image clarity, but may also result in higher bandwidth, storage, and transmission requirements.

[0038] Frames of a video sequence can be encoded and decoded using JVET. JVET is a video coding solution being developed by the Joint Video Exploration Team. The versions of JVET have been implemented in the JEM (Joint Exploration Model) encoder and decoder. Similar to other video coding solutions such as HEVC (High Efficiency Video Coding), JVET is a block-based hybrid spatio-temporal prediction coding solution. During coding with JVET, first, a frame is partitioned into square blocks called CTUs 100, as shown in FIG. 1. For example, a CTU 100 can be a block of 128×128 pixels.

[0039] Figure 2 An exemplary partitioning depicting the splitting of CTU 100 into CUs 102 is shown. Each CTU 100 in a frame can be split into one or more CUs (Coding Units) 102. CUs 102 can be used for prediction and transformation as described below. Different from HEVC, in JVET, a CU 102 can be rectangular or square and can be coded without further splitting into prediction units or transformation units. A CU 102 can be as large as its root CTU 100 or can be a smaller subdivision of the root CTU 100 as small as a 4×4 block.

[0040] In JVET, a CTU 100 can be split into CUs 102 according to the Quadtree plus Binary Tree (QTBT) scheme, where a CTU 100 can be recursively separated into square blocks according to the quadtree, and then these square blocks can be recursively separated horizontally or vertically according to the binary tree. Parameters such as CTU size, minimum size for quadtree and binary tree leaf nodes, maximum size for binary tree root nodes, and maximum depth for the binary tree can be set to control the separation according to QTBT.

[0041] In some embodiments, JVET can limit the binary splitting in the binary tree part of the QTBT to symmetric splitting, where a block can be divided vertically or horizontally into two halves along the midline.

[0042] By way of non-limiting example, Figure 2 A CTU 100 split into CUs 102 is shown by solid lines indicating quadtree separation and dashed lines indicating symmetric binary tree separation. As shown, the binary separation allows symmetric horizontal and vertical separations to define the structure of the CTU and its subdivision into CUs.

[0043] Figure 3 Shows Figure 2The segmented QTBT representation. The quadtree root node represents the CTU 100, and each child node in the quadtree part represents one of the four square blocks separated from the parent square block. Then, the square block represented by the quadtree leaf node can be symmetrically divided zero or more times using a binary tree, where the quadtree leaf node is the root node of the binary tree. At each level of the binary tree part, the block can be divided symmetrically, vertically, or horizontally. A flag set to "0" indicates that the block is symmetrically separated horizontally, while a flag set to "1" indicates that the block is symmetrically separated vertically.

[0044] In other embodiments, JVET may allow symmetric binary splitting or asymmetric binary splitting in the binary tree part of the QTBT. When splitting the prediction unit (PU), asymmetric motion splitting (AMP) is allowed in different contexts in HEVC. However, for splitting the CU 102 in JVET according to the QTBT structure, when the relevant region of the CU 102 is not located on either side of the midline passing through the center of the CU, asymmetric binary splitting can result in an improved split compared to symmetric binary splitting. By way of non-limiting example, when the CU 102 depicts one object near the center of the CU and another object at the side of the CU 102, the CU 102 can be split asymmetrically to place each object in a separate smaller CU 102 of different sizes.

[0045] Figure 4 Depicts four possible types of asymmetric binary splitting, where the CU 102 is separated into two smaller CU 102s along a line passing through the length or height of the CU 102, such that one of the smaller CU 102s is 25% of the size of the parent CU 102 and the other is 75% of the size of the parent CU 102. Figure 4 The four types of asymmetric binary splitting shown in allow the CU102 to be separated along a line that is 25% of the route starting from the left side of the CU 102, 25% of the route starting from the right side of the CPU 102, 25% of the route starting from the top of the CPU102, or 25% of the route starting from the bottom of the CPU 102. In an alternative embodiment, the asymmetric splitting line at which the CU 102 is separated can be located at any other position such that the CU102 is not symmetrically divided into two halves.

[0046] Figure 5 Depicts a non-limiting example of a CTU 100 split into CU 102s using a scheme that allows both symmetric binary splitting and asymmetric binary splitting in the binary tree part of the QTBT. In Figure 5 the dashed line shows the asymmetric binary splitting line, where the parent CU 102 is separated using one of the splitting types shown in Figure 4 the one shown in.

[0047] Figure 6 shows Figure 5 the segmented QTBT representation. In Figure 6 it, two solid lines extending from a node indicate a symmetric split in the binary tree portion of the QTBT, while two dashed lines extending from a node indicate an asymmetric split in the binary tree portion.

[0048] Syntax can be compiled in the bitstream that indicates how to split CTU 100 into CUs 102. By way of non-limiting example, syntax can be compiled in the bitstream that indicates which nodes are separated by quadtree splitting, which nodes are separated by symmetric binary splitting, and which nodes are separated by asymmetric binary splitting. Similarly, syntax can be compiled in the bitstream for nodes separated by asymmetric binary splitting that indicates which type of asymmetric binary splitting is used, such as Figure 4 one of the four types shown in

[0049] In some embodiments, the use of asymmetric splitting can be limited to splitting CU 102 at the leaf nodes of the quadtree portion of the QTBT. In these embodiments, the CU 102 at the child nodes separated from a parent node using quadtree splitting in the quadtree portion can be the final CU 102, or they can be further separated using quadtree splitting, symmetric binary splitting, or asymmetric binary splitting. The child nodes in the binary tree portion separated using symmetric binary splitting can be the final CU 102, or they can be recursively separated further one or more times using only symmetric binary splitting. In cases where further separation is not allowed, the child nodes in the binary tree portion separated from the QT leaf node using asymmetric binary splitting can be the final CU 102.

[0050] In these embodiments, limiting the use of asymmetric splitting to separating quadtree leaf nodes can reduce search complexity and / or limit overhead bits. Since only quadtree leaf nodes can be separated by asymmetric splitting, the use of asymmetric splitting can directly indicate the end of a QT portion branch without additional syntax or further signaling. Similarly, since nodes with asymmetric splitting cannot be further separated, the use of asymmetric splitting on a node can also directly indicate that its children with asymmetric splitting are the final CUs 102 without additional syntax or further signaling.

[0051] In alternative embodiments, such as when limiting search complexity and / or the number of overhead bits becomes immaterial, asymmetric splitting can be used to separate nodes generated by quadtree splitting, symmetric binary splitting, and / or asymmetric binary splitting.

[0052] After performing quadtree splitting and binary tree splitting using any of the above QTBT structures, the blocks represented by the leaf nodes of the QTBT represent the final CUs 102 to be compiled, such as compilation using inter prediction or intra prediction. For slices or full frames compiled using inter prediction, different splitting structures can be used for the luminance and chrominance components. For example, for an inter slice, the CU 102 can have compiled blocks (CBs) for different color components, such as one luminance CB and two chrominance CBs. For slices or full frames compiled using intra prediction, the splitting structures for the luminance and chrominance components can be the same.

[0053] In an alternative embodiment, JVET can use a two-level compiled block structure as an alternative or extension to the above QTBT splitting. In the two-level compiled block structure, first, the CTU 100 can be split into basic units (BUs) at a higher level. Then, the BUs can be split into multiple operation units (OUs) at a lower level.

[0054] In an embodiment employing the two-level compiled block structure, at a higher level, the CTU 100 can be split into BUs according to one of the above QTBT structures or according to a quadtree (QT) structure such as the quadtree (QT) structure used in HEVC where a block can be split only into four equal-sized sub-blocks. By way of non-limiting example, the CTU 102 can be split into BUs according to the QTBT structure described above such that quadtree splitting, symmetric binary splitting, or asymmetric binary splitting can be used to split the leaf nodes in the quadtree part. In this example, the final leaf nodes of the QTBT can be BUs instead of CUs. Figures 5 - 6 At a lower level in the two-level compiled block structure, each BU split from the CTU 100 can be further split into one or more OUs. In some embodiments, when the BU is square, quadtree splitting or binary splitting, such as symmetric or asymmetric binary splitting, can be used to split it into OUs. However, when the BU is not square, only binary splitting can be used to split it into OUs. Limiting the type of splitting that can be used for non-square BUs can limit the number of bits used to signal the type of splitting used to generate the BUs.

[0055]

[0056] ​Although the following discussion describes compiling CU 102, in embodiments using a two-level compilation block structure, the BU and OU rather than CU 102 can be compiled. By way of non-limiting example, the BU can be used for higher-level compilation operations such as intra prediction or inter prediction, while the smaller OU can be used for lower-level compilation operations such as transformation and generation of transform coefficients. Thus, the syntax for the BU that indicates whether to compile via intra prediction or inter prediction, or information identifying a specific intra prediction mode or motion vector, is used to encode the BU. Similarly, the syntax for the OU can identify a specific transformation operation or quantized transform coefficients for compiling the OU.

[0057] Figure 7 A simplified block diagram depicting CU compilation in a JVET encoder is shown. The main stages of video compilation include splitting to identify CU 102 as described above, subsequently encoding CU 102 using prediction at 704 or 706, generating a residual CU 710 at 708, performing transformation at 712, quantization at 716, and entropy compilation at 720. Figure 7 The encoder and encoding process illustrated in also includes a decoding process, which will be described in more detail below.

[0058] Given the current CU 102, the encoder can obtain a predicted CU 702 by using intra prediction spatially at 704 or inter prediction temporally at 706. The basic idea of predictive compilation is to transmit the difference or residual signal between the original signal and the prediction of the original signal. At the receiver side, the original signal can be reconstructed by adding the residual and the prediction, as will be described below. Since the difference signal has lower correlation than the original signal, fewer bits are required for its transmission.

[0059] A slice that is completely compiled with intra prediction of CU 102, such as an entire picture or a portion of a picture, can be an I slice, which can be decoded without reference to other slices and, as such, can be a possible starting point for decoding. A slice compiled with at least some inter prediction of CU can be a predicted (P) or bi-predicted (B) slice that can be decoded based on one or more reference pictures. A P slice can use intra prediction and inter prediction with previously compiled slices. For example, by using inter prediction, a P slice can be further compressed compared to an I slice, but the previously compiled slices need to be compiled to encode them. A B slice can use intra prediction or inter prediction that applies interpolation prediction from two different frames, encoding using data from previous and / or subsequent slices, thereby increasing the accuracy of the motion estimation process. In some cases, intra block copy can also be used to encode or alternatively encode P and B slices, where data from other parts of the same slice is used.

[0060] As will be discussed below, intra prediction or inter prediction can be performed based on reconstructed CUs 734 from previously encoded CUs 102 (such as neighboring CUs 102 or CUs 102 in a reference picture).

[0061] When spatially encoding CU 102 using intra prediction at 704, an intra prediction mode can be found that best predicts the pixel values of CU 102 based on samples from neighboring CUs 102 in the picture.

[0062] When encoding the luminance component of a CU, the encoder can generate a list of candidate intra prediction modes. Although HEVC has 35 possible intra prediction modes for the luminance component, in JVET, there are 67 possible intra prediction modes for the luminance component. These include a planar mode using a three-dimensional plane of values generated from neighboring pixels, a DC mode using an average of values from neighboring pixels, and 65 directional modes shown in Figure 8 that use values copied from neighboring pixels along the indicated directions.

[0063] When generating a list of candidate intra prediction modes for the luminance component of a CU, the number of candidate modes on the list can depend on the size of the CU. The candidate list can include: a subset of the 35 HEVC modes with the lowest SATD (Sum of Absolute Transform Differences) cost; new directional modes added for JVET that are adjacent to candidates found from HEVC modes; and modes from among a set of six most probable modes (MPMs) for CU 102 identified based on the intra prediction modes for previously encoded neighboring blocks and a default mode list.

[0064] When encoding the chrominance component of a CU, a list of candidate intra prediction modes can also be generated. The candidate mode list can include modes generated using cross-component linear model projections from luminance samples, intra prediction modes found for luminance CBs at specific co-location positions in the chrominance block, and previously found chrominance prediction modes for neighboring blocks. The encoder can find the candidate mode with the lowest rate-distortion cost in the list and use these intra prediction modes when encoding the luminance and chrominance components of the CU. The syntax can be encoded in the bitstream indicating the intra prediction mode used to encode each CU 102.

[0065] After the best intra prediction modes for CU 102 have been selected, the encoder can use those modes to generate predicted CU 402. When the selected mode is a directional mode, a 4-tap filter can be used to improve the directional accuracy. Boundary prediction filters, such as 2-tap or 3-tap filters, can be used to adjust columns or rows at the top or left of the predicted block.

[0066] The prediction CU 702 can be further smoothed using a position-dependent intra prediction combination (PDPC) process that uses the unfiltered samples of neighboring blocks to adjust the predicted CU 702 generated based on the filtered samples of neighboring blocks, or adaptive reference sample smoothing using a 3-tap or 5-tap low-pass filter to process the reference samples.

[0067] When the CU 102 is temporally compiled by inter prediction at 706, a set of motion vectors (MVs) can be found that point to samples in a reference picture of the pixel values of the best predicted CU 102. Inter prediction exploits the temporal redundancy between slices by representing the displacement of pixel blocks in the slice. The displacement is determined by a process called motion compensation based on the pixel values in the previous or subsequent slice. The motion vectors indicating the displacement of the pixels relative to a particular reference picture and the associated reference index can be provided to the decoder in the bitstream together with the residual between the original pixels and the motion-compensated pixels. The decoder can use the residual and the signaled motion vectors and reference index to reconstruct the pixel block in the reconstructed slice.

[0068] In JVET, the motion vector precision can be stored at 1 / 16 pixel, and the difference between the motion vector and the predicted motion vector of the CU can be compiled at a quarter-pixel resolution or an integer-pixel resolution.

[0069] In JVET, techniques such as advanced temporal motion vector prediction (ATMVP), spatio-temporal motion vector prediction (STMVP), affine motion compensation prediction, pattern matching motion vector derivation (PMMVD), and / or bidirectional optical flow (BIO) can be used to find motion vectors for multiple sub-CUs within the CU 102.

[0070] Using ATMVP, the encoder can find the temporal vector of the CU 102 that points to the corresponding block in the reference picture. The temporal vector can be found based on the motion vectors and reference pictures found for previously compiled adjacent CUs 102. Using the reference block pointed to by the temporal vector for the entire CU 102, motion vectors can be found for each sub-CU within the entire CU 102.

[0071] STMVP can find the motion vectors of the sub-CUs by scaling and averaging the motion vectors and temporal vectors found for neighboring blocks previously compiled using inter prediction.

[0072] Affine motion compensation prediction can be used to predict the field of motion vectors for each sub-CU in a block based on two control motion vectors found for the top corners of the block. For example, the motion vectors of the sub-CUs can be derived based on the top corner motion vectors found for each 4x4 block within the CU 102.

[0073] PMMVD can use bidirectional matching or template matching to find the initial motion vectors of the current CU 102. Bidirectional matching can look at the current CU 102 and reference blocks in two different reference pictures along the motion trajectory, while template matching can look at the corresponding block in the current CU 102 and the reference pictures identified by the template. Then, the initial motion vectors found for the CU 102 can be corrected separately for each sub-CU.

[0074] BIO can be used when performing inter prediction with bi-directional prediction based on earlier and later reference pictures, and allows finding motion vectors for sub-CUs based on the gradient of the difference between the two reference pictures.

[0075] In some cases, local luminance compensation (LIC) can be used at the CU level to find the values for the scaling factor parameter and the offset parameter based on the samples adjacent to the current CU 102 and the corresponding samples adjacent to the reference block identified by the candidate motion vector. In JVET, the LIC parameters can be changed and signaled at the CU level.

[0076] For some of the above methods, the motion vectors found for each sub-CU of a CU can be signaled to the decoder at the CU level. For other methods such as PMMVD and BIO, the motion information is not signaled in the bitstream to save overhead, and the decoder can derive the motion vectors through the same process.

[0077] After the motion vectors for the CU 102 have been found, the encoder can use those motion vectors to generate the predicted CU 702. In some cases, when the motion vectors have been found for a single sub-CU, overlapping block motion compensation (OBMC) can be used when generating the predicted CU 702 by combining those motion vectors with the motion vectors previously found for one or more neighboring sub-CUs.

[0078] When using bi-directional prediction, JVET can use decoder-side motion vector refinement (DMVR) to find the motion vectors. DMVR allows finding the motion vectors based on a bi-directional template matching process using the two motion vectors found for bi-directional prediction. In DMVR, a weighted combination of the predicted CU 702 generated by each of the two motion vectors can be found, and the two motion vectors can be refined by replacing them with the new motion vector that best points to the combined predicted CU 702. The two refined motion vectors can be used to generate the final predicted CU 702.

[0079] At 708, as described above, once the predicted CU 702 is found through intra prediction at 704 or inter prediction at 706, the encoder can subtract the predicted CU 702 from the current CU 102 to find the residual CU 710.

[0080] The encoder may use one or more transform operations at 712 to convert the residual CU 710 into transform coefficients 714 that represent the residual CU 710 in the transform domain, such as using a discrete cosine block transform (DCT transform) to transform the data into the transform domain. Compared with HEVC, JVET allows more types of transform operations, including DCT-II, DST-VII, DST-VII, DCT-VIII, DST-I, and DCT-V operations. The allowed transform operations can be grouped into subsets, and the encoder may signal an indication of which subsets and which specific operations within those subsets are used. In some cases, a transform of a large block size may be used to zero out the high-frequency transform coefficients in CUs 102 larger than a particular size, such that only the low-frequency transform coefficients are retained for those CUs 102.

[0081] In some cases, after the forward kernel transform, a mode-dependent non-separable second-order transform (MDNSST) may be applied to the low-frequency transform coefficients 714. The MDNSST operation may use a Hypercube-Givens transform (HyGT) based on rotated data. When used, the encoder may signal an index value that identifies a particular MDNSST operation.

[0082] At 716, the encoder may quantize the transform coefficients 714 to quantized transform coefficients 716. The quantization of each coefficient may be calculated by dividing the value of the coefficient by a quantization step size, which is derived from the quantization parameter (QP). In some embodiments, Qstep is defined as 2 (QP-4) / 6 . Since the high-precision transform coefficients 714 can be converted to quantized transform coefficients 716 with a finite number of possible values, quantization can assist in data compression. Thus, the quantization of the transform coefficients can limit the amount of bits generated and sent by the transform process. However, although quantization is a lossy operation and the quantization loss cannot be recovered, the quantization process presents a trade-off between the quality of the reconstructed sequence and the amount of information required to represent the sequence. For example, a lower QP value may result in a better quality decoded video, although a higher amount of data may be required for representation and transmission. Conversely, a high QP value results in a lower quality reconstructed video sequence but lower data and bandwidth requirements.

[0083] JVET can utilize variance-based adaptive quantization techniques, which allow each CU 102 to use different quantization parameters for its encoding or process (instead of using the same frame QP for the encoding of each CU 102 in a frame). Variance-based adaptive quantization techniques can adaptively reduce the quantization parameters for some blocks while increasing the quantization parameters in other blocks. To select a specific QP for the CU102, the variance of the CU is calculated. In short, if the variance of the CU is higher than the average variance of the frame, a QP higher than the QP of the frame can be set for the CU102. If the CU 102 exhibits a variance lower than the average variance of the frame, a lower QP can be assigned.

[0084] At 720, the encoder can find the final compressed bits 722 by performing entropy encoding on the quantized transform coefficients 718. Entropy encoding aims to remove the statistical redundancy of the information to be sent. In JVET, CABAC (Context-Adaptive Binary Arithmetic Coding) can be used to encode the quantized transform coefficients 718, which uses probability metrics to remove statistical redundancy. For a CU 102 with non-zero quantized transform coefficients 718, the quantized transform coefficients 718 can be converted to binary. Then, each bit ("bin") of the binary representation can be encoded using a context model. The CU 102 can be decomposed into three regions, each region having its own set of context models for the pixels within that region.

[0085] Multiple scan operations can be performed to encode the bins. In the operation of encoding the first three bins (bin0, bin1, and bin2), the index value indicating the context model to be used for that bin can be found by finding the sum of the bin positions in up to five previously encoded neighboring quantized transform systems identified by a template.

[0086] The context model can be based on the probability that the value of the bin is "0" or "1". When encoding the values, the probabilities in the context model can be updated based on the actual number of "0" and "1" values encountered. While HEVC uses a fixed table to re-initialize the context model for each new picture, in JVET, the probabilities of the context model for a new inter-predicted picture can be initialized based on the context model developed for the previously encoded inter-predicted pictures.

[0087] The encoder can generate a bitstream that contains the entropy-encoded bits 722 of the residual CU 710, prediction information such as the selected intra-prediction mode or motion vectors, an indicator of how the CU 102 is split from the CTU 100 according to the QTBT structure, and / or other information related to the encoded video. The bitstream can be decoded by the decoder as described below.

[0088] In addition to using the quantized transform coefficients 718 to find the final compressed bits 722, the encoder can also use the quantized transform coefficients 718 to generate a reconstructed CU 734 by following the same decoding process that the decoder will use to generate the reconstructed CU 734. Thus, once the encoder computes and quantizes the transform coefficients, the quantized transform coefficients 718 can be sent to a decoding loop in the encoder. After quantizing the transform coefficients of a CU, the decoding loop allows the encoder to generate the same reconstructed CU 734 as the decoder generates during the decoding process. Thus, the encoder can use the same reconstructed CU 734 that would be used for neighboring CUs 102 or reference pictures when the decoder performs intra prediction or inter prediction for a new CU 102. The reconstructed CU 102, reconstructed slice, or fully reconstructed frame can be used as a reference for further prediction stages.

[0089] At the decoding loop of the encoder where the pixel values of the reconstructed image are obtained (see below for the same operation in the decoder), a dequantization process can be performed. To dequantize a frame, for example, the quantized value of each pixel of the frame is multiplied by the quantization step, e.g., the above (Qstep), to obtain the reconstructed dequantized transform coefficients 726. For example, in the decoding process shown in Figure 7 the encoder, the quantized transform coefficients 718 of the residual CU 710 can be dequantized at 724 to find the dequantized transform coefficients 726. If the MDNSST operation is performed during encoding, this operation can be reversed after dequantization.

[0090] At 728, the dequantized transform coefficients 726 can be inversely transformed to find the reconstructed residual CU 730, such as by applying the DCT to these values to obtain the reconstructed image. At 732, the reconstructed residual CU 730 can be added to the corresponding predicted CU 702 found by intra prediction at 704 or by inter prediction at 706 in order to find the reconstructed CU 734.

[0091] At 736, one or more filters can be applied to the reconstructed data at the picture level or CU level during the decoding process (in the encoder or, as described below, in the decoder). For example, the encoder can apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). The decoding process of the encoder may implement filters to estimate the best filter parameters that can solve potential artifacts in the reconstructed image and send them to the decoder. Such improvements increase the objective and subjective quality of the reconstructed video. In deblocking filtering, pixels near the sub-CU boundary can be modified, while in SAO, pixels in the CTU 100 can be modified using edge offsets or band offset classification. The ALF of JVET can use a filter with a circular symmetric shape for each 2x2 block. An indication of the size and identity of the filter for each 2x2 block can be signaled.

[0092] If the reconstructed pictures are reference pictures, they can be stored in the reference buffer 738 for inter prediction of future CUs 102 at 706.

[0093] During the above steps, JVET allows the use of content adaptive cropping operations to adjust color values to fit the upper and lower cropping bounds. The cropping bounds can vary for each slice, and parameters identifying the bounds can be signaled in the bitstream.

[0094] Figure 9 A simplified block diagram depicting CU compilation in the JVET decoder is shown. The JVET decoder can receive a bitstream containing information about the encoded CU 102. The bitstream can indicate how to partition the CUs 102 of a picture from the CTU 100 according to the QTBT structure. As a non-limiting example, the bitstream can use quadtree partitioning, symmetric binary partitioning, and / or asymmetric binary partitioning to identify how to partition the CUs 102 from each CTU 100 in the QTBT. The bitstream can also indicate prediction information for the CU 102, such as an intra prediction mode or a motion vector, and the bits 902 representing the entropy-coded residual CU.

[0095] At 904, the decoder can use the CABAC context model signaled by the encoder in the bitstream to decode the entropy-coded bits 902. The decoder can update the probabilities of the context model using the parameters signaled by the encoder in the same way as they were updated during the encoding process.

[0096] After inverting the entropy coding at 904 to find the quantized transform coefficients 906, the decoder can dequantize them at 908 to find the dequantized transform coefficients 910. If the MDNSST operation was performed during encoding, this operation can be inverted by the decoder after dequantization.

[0097] At 912, the dequantized transform coefficients 910 can be inverse-transformed to find the reconstructed residual CU 914. At 916, the reconstructed residual CU 914 can be added to the corresponding predicted CU 926 found using intra prediction at 922 or inter prediction at 924 to facilitate finding the reconstructed CU 918.

[0098] At 920, one or more filters can be applied to the reconstructed data at a picture level or a CU level. For example, the decoder can apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). As described above, the in-loop filter located in the decoding loop of the encoder can be used to estimate optimal filter parameters to increase the objective and subjective quality of the frame. These parameters are sent to the decoder to filter the reconstructed frame at 920 to match the filtered reconstructed frame in the encoder.

[0099] After the reconstructed picture has been generated by finding the reconstructed CU 918 and applying the signaled filters, the decoder can output the reconstructed picture as the output video 928. If the reconstructed picture is used as a reference picture, it can be stored in the reference buffer 930 for inter prediction of future CUs 102 at 924.

[0100] Figure 10 An embodiment of a method for CU compilation 1000 in a JVET decoder is depicted. In Figure 10 the embodiment shown, at step 1002, an encoded bitstream 902 can be received, and then at step 1004, the CABAC context model associated with the encoded bitstream 902 can be determined, and then the determined CABAC context model can be used to decode the encoded bitstream 902 at step 1006.

[0101] At step 1008, the quantized transform coefficients 906 associated with the encoded bitstream 902 can be determined, and then the dequantized transform coefficients 910 can be determined from the quantized transform coefficients 906 at step 1010.

[0102] In step 1012, it can be determined whether the MDNSST operation is performed during encoding and / or whether the bitstream 902 contains an indication to apply the MDNSST operation to the bitstream 902. If it is determined that the MDNSST operation is performed during the encoding process or the bitstream 902 contains an indication to apply the MDNSST operation to the bitstream 902, then the inverse MDNSST operation 1014 can be implemented before performing the inverse transform operation 912 on the bitstream 902 in step 1016. Alternatively, the operation 912 can be performed on the bitstream 902 in step 1016 without applying the inverse MDNSST operation in step 1014. The inverse transform operation 912 in step 1016 can determine and / or construct the reconstructed residual CU 914.

[0103] In step 1018, the reconstructed residual CU 914 from step 1016 can be combined with the predicted CU 918. The predicted CU 918 can be one of the intra-predicted CU 922 determined in step 1020 and the inter-predicted unit 924 determined in step 1022.

[0104] In step 1024, any one or more filters 920 can be applied to the reconstructed CU 914 and output in step 1026. In some embodiments, the filter 920 may not be applied in step 1024.

[0105] In some embodiments, in step 1028. The reconstructed CU 918 can be stored in the reference buffer 930.

[0106] Figure 11 A simplified block diagram 1100 depicting CU compilation for a JVET encoder is shown. In step 1102, the JVET compilation tree unit can be represented as a root node in a quadtree plus binary tree (QTBT) structure. In some embodiments, the QTBT can have a quadtree branching from the root node and / or a binary tree branching from the leaf nodes of one or more quadtrees. The representation from step 1102 can proceed to step 1104, 1106, or 1108.

[0107] In step 1104, the represented quadtree nodes can be separated into two blocks of unequal size using asymmetric binary splitting. In some embodiments, the separated blocks can be represented in a binary tree branching from a quadtree node that is a leaf node capable of representing the final compiled unit. In some embodiments, further separation is not allowed in a binary tree branching from a quadtree node that is a leaf node representing the final compiled unit. In some embodiments, the asymmetric split can separate the compiled unit into blocks of unequal size, with the first block representing 25% of the quadtree node and the second block representing 75% of the quadtree node.

[0108] In step 1106, quadtree splitting can be employed to split the represented quadtree node into four equally sized square blocks. In some embodiments, the separated blocks can be represented as quadtree nodes representing the final compilation units, or can be represented as child nodes that can be split again by quadtree splitting, symmetric binary splitting, or asymmetric binary splitting.

[0109] In step 1108, quadtree splitting can be used to separate the represented quadtree node into two equally sized blocks. In some embodiments, the separated blocks can be represented as quadtree nodes representing the final compilation units, or can be represented as child nodes that can be split again by quadtree splitting, symmetric binary splitting, or asymmetric binary splitting.

[0110] In step 1110, the child nodes from step 1106 or step 1108 can be represented as child nodes configured to be encoded. In some embodiments, the child nodes can be represented by the leaf nodes of a binary tree using JVET.

[0111] In step 1112, JVET can be used to encode the compilation units from step 1104 or 1110.

[0112] Figure 12 A simplified block diagram 1200 depicting CU decoding in a JVET decoder is shown. In Figure 12 the depicted embodiment, in step 1202, a bitstream indicating how to split the compilation tree unit into compilation units according to the QTBT structure can be received. The bitstream can indicate how to separate the quadtree nodes by at least one of quadtree splitting, symmetric binary splitting, or asymmetric binary splitting.

[0113] In step 1204, the compilation units represented by the leaf nodes of the QTBT structure can be identified. In some embodiments, the compilation unit can indicate whether the node uses asymmetric binary splitting to separate the leaf node from the quadtree. In some embodiments, the compilation unit can indicate that the node represents the final compilation unit to be decoded.

[0114] In step 1206, JVET can be used to decode the identified compilation units.

[0115] Figure 13 An alternative simplified block diagram depicting JVET compilation for intra mode prediction 1300 is shown. In Figure 13In the embodiments depicted, in step 1302, a set of MPMs can be identified and instantiated in memory, and then in step 1304, a set of 16 selected modes can be identified and instantiated in memory, and in step 1304, a balance of 67 modes can be defined and instantiated in storage. In some embodiments, the set of MPMs can be reduced from a standard set of 6 MPMs. In some embodiments, the set of MPMs can include 5 unique modes, the selected modes can include 16 unique modes, and the set of unselected modes can include the remaining 46 unselected unique modes. However, in alternative embodiments, the set of MPMs can include fewer unique modes, the selected modes can remain fixed at 16 unique modes, and the size of the set of unselected unique modes can be adjusted accordingly to accommodate a total of 67 modes.

[0116] By way of non-limiting example, in some embodiments, where the set of MPMs includes 5 unique modes instead of six MPMs, thus, if truncated unary binarization is used and a new binarization for 5 MPMs can be utilized, the number of bins assigned to the MPM modes can thus be equal to or less than five bins. Thus, in some embodiments, 16 selected modes among the 62 remaining intra modes can be generated by uniformly downsampling these 62 intra modes, and each mode can be encoded by a 4-bit fixed-length code. By way of non-limiting example, if it is assumed that the remaining 62 modes are indexed as {0, 1, 2, …, 61}, then the 16 selected modes = {0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 60}. The remaining 46 unselected modes = {1, 2, 3, 5, 6, 7, 9, 10 … 59, 61}, where these 46 unselected modes can be encoded using truncated binary codes.

[0117] Figure 14 Depicting according to Figure 13 Table 1400 of an alternative JVET compilation for intra mode prediction. In Figure 14 the embodiments depicted, the intra prediction mode 1402 is shown as including 5 MPMs, 16 selected modes, and 46 unselected modes, where a binary string 1404 for the MPMs can be encoded using truncated unary binarization, the 16 selected modes can be encoded using a 4-bit fixed-length code, and the 46 unselected modes can be encoded using truncated binary compilation.

[0118] In Figure 13 an alternative embodiment, 6 MPMs can be utilized, but as Figure 14As shown, only the first five MPMs on the MPM list are binarized and compiled using the current context-based method described in the current JVET. Now, the sixth MPM on the MPM list is considered one of the 16 selected modes and is compiled together with the other 15 selected modes using a 4-bit fixed-length code.

[0119] As a non-limiting example, if the remaining 61 modes are indexed as {0, 1, 2, …, 60}, the following 15 selected modes can be obtained by uniformly subsampling the remaining 61 intra modes: the set of selected modes can be {0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, where the 15 selected modes plus the sixth MPM are compiled using a 4-bit fixed-length code, as in the following set: {sixth MPM, 0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, and the balance of the 46 non-selected modes is shown in the following set and is compiled into the set of non-selected modes = {1, 2, 3, 4, 6, 7, 8, 9, 11, 12 … 49, 51, 52, 53, 54, 56, 57, 58, 59} using a truncated binary code.

[0120] In Figure 13 yet another alternative embodiment, only the first five MPMs on the MPM list can be binarized, as Figure 14 shown, and compiled using the current context-based method described in the current JVET standard. In such an embodiment, the sixth MPM on the MPM list can be considered one of the 16 selected modes and is compiled together with the other 15 selected modes using a 4-bit fixed-length code. Thus, any known convenient and / or desired selection process can be used to establish the selection of the other 15 selected modes. By non-limiting example, they can be selected around the MPM mode, or around (content-based) statistical popular modes, or around trained or historical popular modes, or using other methods or processes.

[0121] Again, selecting 5 MPMs is only a non-limiting example, and in alternative embodiments, the set of MPMs can be further reduced to 4 or 3 MPMs or extended to more than 6, where there are still 16 selected modes, and the balance of the 67 (or other known, convenient, and / or desired total) intra-compiled modes is included in the set of non-selected intra-compiled modes. That is, embodiments where the total number of intra-compiled modes is greater than or less than 67 can be envisioned as embodiments where the set of MPMs can contain any known convenient or desired number of MPMs, and the number of selected modes can be any known convenient and / or desired number.

[0122] The execution of the instruction sequence required for the practice embodiment can be performed by a computer system 1500 as shown in Figure 15 . In an embodiment, the execution of the instruction sequence is performed by a single computer system 1500. According to other embodiments, two or more computer systems 1500 coupled by a communication link 1515 can execute the instruction sequence in coordination with each other. Although only one computer system 1500 will be described below, it should be understood that any number of computer systems 1500 can be employed to practice the embodiment.

[0123] Reference will now be made to Figure 15 describe a computer system 1500 according to an embodiment, Figure 15 is a block diagram of the functional components of the computer system 1500. As used herein, the term computer system 1500 is used broadly to describe any computing device that can store and independently run one or more programs.

[0124] Each computer system 1500 can include a communication interface 1514 coupled to a bus 1506. The communication interface 1514 provides two-way communication between computer systems 1500. The communication interfaces 1514 of the respective computer systems 1500 send and receive electrical, electromagnetic, or optical signals, which include data streams representing various types of signal information, such as instructions, messages, and data. The communication link 1515 links one computer system 1500 to another computer system 1500. For example, the communication link 1515 can be a LAN, in which case the communication interface 1514 can be a LAN card, or the communication link 1515 can be a PSTN, in which case the communication interface 1514 can be an integrated services digital network (ISDN) card or a modem, or the communication link 1515 can be the Internet, in which case the communication interface 1514 can be a dial-up, cable, or wireless modem.

[0125] Computer systems 1500 can send and receive messages, data, and instructions, including programs, i.e., application programs, code, through their respective communication links 1515 and communication interfaces 1514. When received, the received program code can be executed by their respective processors 1507 and / or stored in a storage device 1510 or other associated non-volatile media for later execution.

[0126] In an embodiment, computer system 1500 operates in conjunction with a data storage system 1531, e.g., a data storage system 1531 that includes a database 1532, which is readily accessible by computer system 1500. Computer system 1500 communicates with data storage system 1531 via a data interface 1533. Data interface 1533, coupled to bus 1506, transmits and receives electrical, electromagnetic, or optical signals, which include a data stream representing various types of signal information, e.g., instructions, messages, and data. In an embodiment, the functionality of data interface 1533 may be performed by communication interface 1514.

[0127] Computer system 1500 includes: a bus 1506 or other communication mechanism for conveying instructions, messages, and data, collectively, information; and one or more processors 1507 coupled to bus 1506 for processing the information. Computer system 1500 also includes a main memory 1508, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 1506 for storing dynamic data and instructions to be executed by processor 1507. Main memory 1508 may also be used to store temporary data during execution of instructions by processor 1507, i.e., variables or other intermediate information.

[0128] Computer system 1500 may further include a read only memory (ROM) 1509 or other static storage device coupled to bus 1506 for storing static data and instructions for processor 1507. A storage device 1510, such as a magnetic disk or optical disk, may also be provided and coupled to bus 1506 for storing data and instructions for processor 1507.

[0129] Computer system 1500 may be coupled via bus 1506 to a display device 1511, such as but not limited to a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor, for displaying information to a user. An input device 1512, e.g., alphanumeric keys and other keys, is coupled to bus 1506 for communicating information and command selections to processor 1507.

[0130] According to one embodiment, a single computer system 1500 performs particular operations by executing one or more sequences of one or more instructions contained in main memory 1508 by their respective processors 1507. Such instructions may be read into main memory 1508 from another computer - available medium, such as ROM 1509 or storage device 1510. Execution of the instruction sequences contained in main memory 1508 causes processor 1507 to perform the processes described herein. In alternative embodiments, hard - wired circuitry may be used in place of or in combination with software instructions. Accordingly, embodiments are not limited to any specific combination of hardware circuitry and / or software.

[0131] As used herein, the term "computer-usable medium" refers to any medium that provides information or that can be used by processor 1507. Such a medium can take many forms, including but not limited to non-volatile, volatile, and transmission media. Non-volatile media, i.e., media that can retain information without the presence of power, include ROM 1509, CD ROMs, magnetic tapes, and magnetic disks. Volatile media, i.e., media that cannot retain information without the presence of power, include main memory 1508. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up bus 1506. Transmission media can also take the form of carrier waves; i.e., electromagnetic waves that can be modulated in terms of frequency, amplitude, or phase to send information signals. Additionally, transmission media can take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.

[0132] In the foregoing specification, embodiments have been described with reference to specific elements of the embodiments. However, it is apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the embodiments. For example, the reader will understand that the specific order and combination of process actions shown in the process flow diagrams described herein are merely illustrative, and that different or additional process actions can be used, or different combinations or orders of process actions can be used to implement the embodiments. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

[0133] It should also be noted that the present invention can be implemented in various computer systems. The various techniques described herein can be implemented in hardware or software or a combination of both. Preferably, the techniques are implemented in a computer program executed on a programmable computer, which includes a processor, a processor-readable storage medium (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The program code is applied to the data input using the input device to perform the above functions and generate output information. The output information is applied to one or more output devices. Each program is preferably implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly language or machine language. In any case, the language can be a compiled language or an interpreted language. Each such computer program is preferably stored on a storage medium or device readable by a general or special programmable computer (e.g., ROM or disk) for configuring and operating the computer when the storage medium or device is read by the computer to perform the above process. The system can also be considered to be implemented as a computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predetermined manner. In addition, the storage elements of an exemplary computing application can be a relational or sequential (flat file) type of computing database capable of storing data in various combinations and configurations.

[0134] Figure 16 is a high-level view of source device 1612 and destination device 1610 that can incorporate the features of the systems and devices described herein. As Figure 16 shown, example video compilation system 1610 includes source device 1612 and destination device 1616, where, in this example, source device 1612 generates encoded video data. Thus, source device 1612 can be referred to as a video encoding device. Destination device 1616 can decode the encoded video data generated by source device 1612. Thus, destination device 1616 can be referred to as a video decoding device. Source device 1612 and destination device 1616 can be examples of video compilation devices.

[0135] Destination device 1616 can receive the encoded video data from source device 1612 via channel 1616. Channel 1616 can include a type of medium or device capable of moving the encoded video data from source device 1612 to destination device 1616. In one example, channel 1616 can include a communication medium that enables source device 1612 to directly send the encoded video data to destination device 1616 in real time.

[0136] In this example, the source device 1612 can modulate the encoded video data according to a communication standard such as a wireless communication protocol, and can send the modulated video data to the destination device 1616. The communication medium can include a wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or other devices that facilitate communication from the source device 1612 to the destination device 1616. In another example, the channel 1616 can correspond to a storage medium that stores the encoded video data generated by the source device 1612.

[0137] In Figure 16 the example of, the source device 1612 includes a video source 1618, a video encoder 1620, and an output interface 1622. In some cases, the output interface 1628 can include a modulator / demodulator (modem) and / or a transmitter. In the source device 1612, the video source 1618 can include sources such as a video capture device, e.g., a camera, a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources.

[0138] The video encoder 1620 can encode the captured, pre-captured, or computer-generated video data. The input image can be received by the video encoder 1620 and stored in the input frame memory 1621. The general-purpose processor 1623 can load information therefrom and perform the encoding. The program for driving the general-purpose processor can be loaded from a storage device such as Figure 16 the example storage module depicted in. The general-purpose processor can use the processing memory 1622 to perform the encoding, and the output of the encoded information of the general-purpose processor can be stored in a buffer such as the output buffer 1626.

[0139] The video encoder 1620 can include a resampling module 1625 that can be configured to compile (e.g., encode) the video data in a scalable video compilation scheme that defines at least one base layer and at least one enhancement layer. As part of the encoding process, the resampling module 1625 can resample at least some of the video data, where the resampling can be performed in an adaptive manner using a resampling filter.

[0140] The encoded video data, e.g., the compiled bitstream, can be directly sent to the destination device 1616 via the output interface 1628 of the source device 1612. In Figure 16In the example, the destination device 1616 includes an input interface 1638, a video decoder 1630, and a display device 1632. In some cases, the input interface 1628 may include a receiver and / or a modem. The input interface 1638 of the destination device 1616 receives the encoded video data via the channel 1616. The encoded video data may include various syntax elements representing the video data generated by the video encoder 1620. Such syntax elements may be included together with the encoded video data transmitted on a communication medium, stored on a storage medium, or stored in a file server.

[0141] The encoded video data may also be stored on a storage medium or a file server for later access by the destination device 1616 for decoding and / or playback. For example, the compiled bitstream may be temporarily stored in the input buffer 1631 and then loaded into the general-purpose processor 1633. A program for driving the general-purpose processor may be loaded from a storage device or a memory. The general-purpose processor may use the processing memory 1632 to perform decoding. The video decoder 1630 may also include a resampling module 1635 similar to the resampling module 1625 employed in the video encoder 1620.

[0142] Figure 16 The resampling module 1635 is depicted as being separate from the general-purpose processor 1633, but those skilled in the art will understand that the resampling function may be performed by a program executed by the general-purpose processor, and the processing in the video encoder may be completed using one or more processors. The decoded image may be stored in the output frame buffer 1636 and then sent to the input interface 1638.

[0143] The display device 1638 may be integrated with the destination device 1616 or may be external to it. In some examples, the destination device 1616 may include an integrated display device and may also be configured to interface with an external display device. In other examples, the destination device 1616 may be a display device. Generally, the display device 1638 displays the decoded video data to the user.

[0144] Video encoders 1620 and video decoders 1630 may operate in accordance with video compression standards. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) are studying the potential need to standardize future video coding technologies with compression capabilities significantly exceeding those of the current High Efficiency Video Coding (HEVC) standard, including its current extensions and recent extensions for screen content coding and high dynamic range coding. The groups are jointly collaborating on this exploration activity, which is called the Joint Video Exploration Team (JVET), to evaluate compression technology designs proposed by their experts in the field. The latest results developed by JVET are described in "Algorithm Description of Joint Exploration Test Model 5 (JEM 5)" of JVET-E1001-V2, written by J. Chen, E. Alshina, G. Sullivan, J. Ohm, and J. Boyce.

[0145] Additionally or alternatively, video encoders 1620 and video decoders 1630 may operate in accordance with other proprietary or industry standards that operate with the disclosed JVET features. Thus, other standards such as the ITU-T H.264 standard, alternatively known as MPEG-4, Part 10, Advanced Video Coding (AVC), or extensions of these standards. Thus, although new developments are made for JVET, the techniques of the present disclosure are not limited to any particular coding standard or technology. Other examples of video compression standards and technologies include MPEG-2, ITU-T H.263, and proprietary or open source compression formats and related formats.

[0146] Video encoders 1620 and video decoders 1630 may be implemented in hardware, software, firmware, or any combination thereof. For example, video encoders 1620 and decoders 1630 may employ one or more processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. When video encoders 1620 and decoders 1630 are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and may use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Each of video encoders 1620 and video decoders 1630 may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined encoder / decoder (CODEC) in their respective devices.

[0147] Aspects of the subject matter described herein may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer, such as the general-purpose processors 1623 and 1633 described above. Generally, program modules include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement particular abstract data types. Aspects of the subject matter described herein may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.

[0148] Examples of memory include random access memory (RAM), read-only memory (ROM), or both. The memory may store instructions for performing the techniques described above, such as source code or binary code. The memory may also be used to store variables or other intermediate information during the execution of instructions to be executed by processors such as processors 1623 and 1633.

[0149] The storage device may also store instructions for performing the techniques described above, such as source code or binary code. The storage device may additionally store data used and manipulated by a computer processor. For example, the storage device in the video encoder 1620 or the video decoder 1630 may be a database accessed by the computer system 1623 or 1633. Other examples of storage devices include random access memory (RAM), read-only memory (ROM), hard disk drives, magnetic disks, optical disks, CD-ROMs, DVDs, flash memory, USB memory cards, or any other medium readable by a computer.

[0150] The memory or storage device may be an example of a non-transitory computer-readable storage medium for use by and / or in conjunction with a video encoder and / or decoder. A non-transitory computer-readable storage medium contains instructions for controlling a computer system configured to perform functions described by a particular embodiment. When executed by one or more computer processors, the instructions may be configured to perform the functions described in a particular embodiment.

[0151] In addition, it should be noted that some embodiments have been described as processes that may be depicted as flowcharts or block diagrams. Although each may describe operations as a sequential process, many operations may be performed in parallel or concurrently. Additionally, the order of the operations may be rearranged. A process may have other steps not included in the figures.

[0152] Certain embodiments can be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium contains instructions for controlling a computer system to execute the methods described by certain embodiments. The computer system may include one or more computing devices. When executed by one or more computer processors, the instructions can be configured to perform the methods described in certain embodiments.

[0153] As used in this specification and the entire following claims, unless the context clearly dictates otherwise, the words "a," "an," and "the" include plural references. Also, as used in this specification and the following claims, unless the context clearly dictates otherwise, the meaning of "in" includes "in" and "on."

[0154] Although the exemplary embodiments of the present invention have been described in detail in terms of the above-described structural features and / or method acts, it is to be understood that those skilled in the art will readily appreciate that many additional modifications can be made to the exemplary embodiments without substantially departing from the novel teachings and advantages of the present invention. Further, it is to be understood that the subject matter defined in the appended claims need not be limited to the above-described specific features or acts. Accordingly, all such modifications are intended to be included within the scope of the present invention as interpreted in accordance with the breadth and scope of the appended claims.

Claims

1. A method for decoding video data from a bitstream, the method comprising: (a) receiving a bitstream indicating how to partition a coded tree unit into coded units, wherein a plurality of the coded units are rectangular, and wherein a plurality of the coded units are square; (b) determining, for a current block of the video data, a first set of possible modes selectable based on a most probable mode (MPM) index, wherein one of the first set of MPMs selectable based on the MPM index includes a direct horizontal mode and another of the first set of MPMs selectable based on the MPM index includes a direct vertical mode and another of the first set of MPMs selectable based on the MPM index includes an angular mode, wherein the first set of MPMs includes only five different modes; (c) deriving from the bitstream (i) an MPM flag and (ii) another index, the MPM flag comprising a total of 1 bit, wherein at least one of the MPM flag and the another index indicates whether an intra mode for predicting the current block is one of the first set of MPMs; (d) when at least one of the MPM flag and the another index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index, selecting the intra mode of the current block based on the MPM index decoded from the bitstream for one of the first set of MPMs; (e) when at least one of the MPM flag and the another index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determining a second set of at least one mode and (ii) determining a third set of at least one mode; (f) wherein the first set, the second set, and the third set include different modes, and wherein a combination of the first set, the second set, and the third set includes 67 different modes; (g) determining the intra mode of the current block for the second set of at least one mode based on a first combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes; and (h) determining the intra mode of the current block for the third set of at least one mode based on a second combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes.

2. A computer-readable storage medium storing instructions that, when executed, cause a processor to perform: (a) receiving a bitstream indicating how to partition a coded tree unit into coded units, wherein a plurality of the coded units are rectangular, and wherein a plurality of the coded units are square; (b) Determine a first set of possible modes selectable based on a Most Probable Mode (MPM) index for a current block of video data, wherein one of the first set of MPMs selectable based on the MPM index includes a direct horizontal mode and another of the first set of MPMs selectable based on the MPM index includes a direct vertical mode and another of the first set of MPMs selectable based on the MPM index includes an angular mode, wherein the first set of MPMs includes only five different modes; (c) Derive from the bitstream (i) an MPM flag and (ii) another index, the MPM flag comprising a total of 1 bit, wherein at least one of the MPM flag and the another index indicates whether the intra mode for predicting the current block is one of the first set of MPMs; (d) When at least one of the MPM flag and the another index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index, select the intra mode of the current block based on the MPM index decoded from the bitstream for one of the first set of MPMs; (e) When at least one of the MPM flag and the another index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determine a second set of at least one mode and (ii) determine a third set of at least one mode; (f) wherein the first set, the second set, and the third set include different modes, and wherein the combination of the first set, the second set, and the third set includes 67 different modes; (g) Determine the intra mode of the current block for the second set of at least one mode based on a first combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes; and (h) Determine the intra mode of the current block for the third set of at least one mode based on a second combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes.

3. A method for encoding video data by an encoder, the method comprising: (a) Provide a bitstream indicating how to split a coding tree unit into coding units, wherein a plurality of the coding units are rectangular and a plurality of the coding units are square; (b) The bitstream contains data adapted to determine a first set of possible modes from which a possible mode can be selected based on a possible mode MPM index for the current block of the video data, wherein one of the first set of MPMs selectable based on the MPM index includes a direct horizontal mode and another of the first set of MPMs selectable based on the MPM index includes a direct vertical mode and another of the first set of MPMs selectable based on the MPM index includes an angular mode, wherein the first set of MPMs includes only five different modes; (c) The bitstream contains data adapted to derive from the bitstream (i) an MPM flag and (ii) another index, the MPM flag comprising a total of 1 bit, and at least one of the MPM flag and the another index indicates whether the intra mode for predicting the current block is one of the first set of MPMs; (d) The bitstream contains data adapted to select the intra mode of the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs when at least one of the MPM flag and the another index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index; (e) The bitstream contains data adapted to, when at least one of the MPM flag and the another index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determine a second set of at least one mode and (ii) determine a third set of at least one mode; (f) The bitstream contains data adapted to wherein the first set, the second set and the third set include different modes, and the combination of the first set, the second set and the third set includes 67 different modes; (g) The bitstream contains data adapted to determine the intra mode of the current block for the second set of at least one mode based on a first combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes; and (h) The bitstream contains data adapted to determine the intra mode of the current block for the third set of at least one mode based on a second combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes.

Citation Information

Patent Citations

  • Enhanced intra-prediction mode signaling for video coding using neighboring mode

    CN103597832A

  • Video decoder with enhanced cabac decoding

    CN103959782A