Intra Mode JVET Compilation Method
By optimizing the set and encoding method of intra prediction compilation modes, the complexity and bandwidth requirements of intra mode compilation in the JVET standard are reduced, the compilation efficiency is improved, and the problems of compilation burden and bandwidth in the prior art are solved.
Patent Information
- Application Number
- CN202210747863.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-07-24
- Filing Date
- 2018-07-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2038-07-24
AI Technical Summary
In the existing JVET standard, the intra-mode compilation method has problems with high compilation burden and bandwidth. Especially when compiling 67 intra-prediction modes, the compilation complexity of the MPM list is high, and the compilation performance of the last two modes is insufficient.
By defining a collection of unique intra-predicted compilation modes, identifying and instantiating a subset of unique MPM intra-predicted compilation modes, a subset of selected intra-predicted compilation modes, and a subset of unselected intra-predicted compilation modes, and optimizing the compilation process using truncated unary binarization and fixed-length codes.
Reduces the complexity and bandwidth requirements of intra-mode compilation, improves compilation efficiency, and optimizes compilation performance.
Smart Images

Figure CN115174912B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application "Intra Mode JVET Compilation Method" with Chinese application number 201880049860.0, PCT application number PCT / US2018 / 043438, international filing date July 24, 2018, and which entered the Chinese national phase on January 23, 2020.
[0002] Priority Claim
[0003] This application claims priority under 35 U.S.C. § 119(e) to the earlier filed U.S. Provisional Application Serial No. 62 / 536,072, filed July 24, 2017, the entire content of which is incorporated herein by reference. Technical Field
[0004] The present disclosure relates to the field of video coding, and more particularly to efficient intra mode coding. Background Art
[0005] Technical improvements in evolving video coding standards illustrate the trend of increasing coding efficiency to achieve higher bitrates, higher resolutions, and better video quality. The Joint Video Exploration Team is developing a new video coding scheme called JVET. Similar to other video coding schemes such as HEVC (High Efficiency Video Coding), JVET is a block-based hybrid spatio-temporal prediction coding scheme. However, relative to HEVC, JVET includes many modifications to the bitstream structure, syntax, constraints, and mappings for generating decoded pictures. JVET has been implemented in the Joint Exploration Model (JEM) encoder and decoder.
[0006] There are a total of 67 intra prediction modes described in the current JVET standard, including the planar, DC mode, and 65 directional angle intra modes. To efficiently code these 67 modes, all intra modes are subdivided into three sets, including a 6 Most Probable Modes (MPM) set, a 16 Selected Modes set, and a 45 Non-Selected Modes set.
[0007] Six MPMs are derived from the modes of available neighboring blocks, derived intra modes, and default intra modes. Figure 1aThe intra modes of the 5 neighboring blocks depicting the current block are shown. They are respectively left (L), above (A), bottom - left (BL), top - right (AR), and top - left (AL), and they are used to form the MPM list of the current block. The initial MPM list is formed by inserting the 5 neighboring intra modes, as well as the planar mode and the DC mode into the MPM list. A pruning process is used to remove duplicate modes so that only unique modes can be included in the MPM list. The order of including the initial modes is: left, above, planar, DC, bottom - left, top - right, and then top - left.
[0008] If the MPM list is incomplete, exported modes are added; these intra modes can be obtained by adding - 1 or + 1 to the angular modes already included in the MPM list. If the MPM list is still incomplete, default modes are added in the following order: vertical, horizontal, mode 2, and diagonal mode. As a result of this process, a unique list of 6 MPM modes is generated.
[0009] For the entropy coding of the 6 MPMs, Figure 1b the truncated unary binarization shown in is currently used. The first three bins of the MPM mode are coded depending on the context of the MPM mode related to the bin currently being signaled. The MPM modes are classified into one of three categories: (a) modes that are mainly horizontal (i.e., the number of MPM modes is less than or equal to the number of modes in the diagonal direction), (b) modes that are mainly vertical (i.e., the number of MPM modes is greater than the number of modes in the diagonal direction), and (c) non - angular (DC and planar) classes. Thus, based on this classification, three contexts are used to signal the MPM index.
[0010] The coding for selecting the remaining 61 non - MPMs is carried out as follows. First, the 61 non - MPMs are divided into two sets: the selected mode set and the unselected mode set. The selected mode set contains 16 modes, and the rest (45 modes) are assigned to the unselected mode set. The mode set to which the current mode belongs is indicated in the bitstream by a flag. If the mode to be indicated is in the selected mode set, the selected mode is signaled with a 4 - bit fixed - length code, and if the mode to be indicated is from the unselected set, the selected mode is signaled with a truncated binary code. By way of example, the following selected mode set is generated by subsampling the 61 non - MPM modes:
[0011] Selected mode set = {0, 4, 8, 12, 16, 20…60}
[0012] Unselected mode set = {1, 2, 3, 5, 6, 7, 9, 10…59}
[0013] In the following Figure 1b the current JVET intra mode coding is summarized.
[0014] As Figure 1b shown, the last two entries of the MPM list require six bins, which is the same number of bins assigned for the 16 selected modes. For the last two modes on the MPM list, this design has no advantage in terms of compilation performance. Also, since the first three bins of the MPM mode are compiled using context-based entropy coding, the complexity of compiling the six bins of the MPM mode is higher than the complexity of compiling the six bins of the selected modes.
[0015] There is a need for a system and method for reducing the compilation burden and bandwidth associated with intra-mode compilation. SUMMARY OF THE INVENTION
[0016] The present disclosure provides a method for video coding for JVET intra prediction, including defining a set of unique intra prediction coding modes, which may be 67 modes in some embodiments, and identifying and instantiating in memory a subset of unique MPM intra prediction coding modes from the set of unique intra prediction coding modes, which may be 5 or fewer out of 7 or more in some embodiments. The method further provides identifying and instantiating in memory a subset of unique selected intra prediction coding modes from the set of unique intra prediction coding modes other than the subset of unique MPM intra prediction coding modes, which may include 16 coding modes in some embodiments, and identifying and instantiating in memory a subset of unique unselected intra prediction coding modes from the set of unique intra prediction coding modes other than the subset of unique MPM intra prediction coding modes and other than the subset of unique selected intra prediction coding modes, to form a balance of intra prediction modes. Then, the subset of unique MPM intra prediction coding modes is coded using truncated unary binary coding.
[0017] The present disclosure also provides a video encoding system for JVET intra prediction. In some embodiments, the system may include the following steps: instantiating a set of 67 unique intra prediction encoding modes in a memory; instantiating a subset of unique MPM intra prediction encoding modes from the set of unique intra prediction encoding modes in the memory; instantiating a subset of 16 unique selected intra prediction modes from the set of unique intra prediction encoding modes except for the subset of unique MPM intra prediction encoding modes in a storage; instantiating a subset of unique unselected intra prediction encoding modes from the set of unique intra prediction encoding modes except for the subset of unique MPM intra prediction encoding modes and except for the subset of unique selected intra prediction encoding modes in the memory; encoding the subset of unique MPM intra prediction encoding modes using truncated unary binarization; and encoding the subset of 16 unique selected intra prediction encoding modes using 4 bits of a fixed length code. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Further details of the present invention are explained with the help of the drawings, wherein:
[0019] Figure 1a Depicts the current encoding block and associated neighboring blocks.
[0020] Figure 1b Depicts a table for the current JVET encoding for intra mode prediction.
[0021] Figure 1c Depicts dividing a frame into multiple encoding tree units (CTUs).
[0022] Figure 2 Depicts an exemplary division of a CTU into encoding units (CUs) using a method of quadtree splitting and symmetric binary splitting.
[0023] Figure 3 Depicts Figure 2 the quadtree plus binary tree (QTBT) representation of the division.
[0024] Figure 4 Depicts four possible types of asymmetric binary splitting for dividing a CU into two smaller CUs.
[0025] Figure 5 Depicts an exemplary division of a CTU into CUs using quadtree splitting, symmetric binary splitting, and asymmetric binary splitting.
[0026] Figure 6 Depicts Figure 5 the QTBT representation of the division.
[0027] Figure 7Depict a simplified block diagram for CU compilation in the JVET encoder.
[0028] Figure 8 Depict 67 possible intra prediction modes for the luminance component in JVET.
[0029] Figure 9 Depict a simplified block diagram for CU compilation in the JVET encoder.
[0030] Figure 10 Depict an embodiment of a method for CU compilation in the JVET encoder.
[0031] Figure 11 Depict a simplified block diagram for CU compilation in the JVET encoder.
[0032] Figure 12 Depict a simplified block diagram for CU decoding in the JVET decoder.
[0033] Figure 13 Depict an alternative simplified block diagram for JVET compilation for intra mode prediction.
[0034] Figure 14 Depict a table for alternative JVET compilation for intra mode prediction.
[0035] Figure 15 Depict an embodiment of a computer system suitable for and / or configured to process a method for CU compilation.
[0036] Figure 16 Depict an embodiment of an encoder / decoder system for CU compilation / decoding in the JVET encoder / decoder. Detailed Description
[0037] Figure 1 depicts a frame divided into multiple coding tree units (CTUs) 100. The frame can be an image in a video sequence. The frame can include a matrix or a set of matrices with pixel values representing intensity measures in the image. Thus, a set of these matrices can generate a video sequence. The pixel values can be defined to represent color and luminance in a full-color video encoding, where the pixel is divided into three channels. For example, in the YCbCr color space, a pixel can have a luminance value Y representing the gray intensity in the image and two chrominance values, Cb and Cr, representing the degree of color difference from gray to blue and red. In other embodiments, the pixel values can be represented by values in different color spaces or models. The resolution of the video can determine the number of pixels in the frame. A higher resolution may mean more pixels and better image clarity, but may also result in higher bandwidth, storage, and transmission requirements.
[0038] Frames of a video sequence can be encoded and decoded using JVET. JVET is a video coding solution being developed by the Joint Video Exploration Team. The versions of JVET have been implemented in the JEM (Joint Exploration Model) encoder and decoder. Similar to other video coding solutions such as HEVC (High Efficiency Video Coding), JVET is a block-based hybrid spatio-temporal prediction coding solution. During coding with JVET, first, a frame is partitioned into square blocks called CTUs 100, as shown in FIG. 1. For example, a CTU 100 can be a block of 128×128 pixels.
[0039] Figure 2 Depicts an exemplary partitioning of a CTU 100 into CUs 102. Each CTU 100 in a frame can be partitioned into one or more CUs (Coding Units) 102. The CUs 102 can be used for prediction and transformation as described below. Different from HEVC, in JVET, a CU 102 can be rectangular or square and can be coded without further partitioning into prediction units or transform units. A CU 102 can be as large as its root CTU 100 or can be a smaller subdivision of the root CTU 100 as small as a 4×4 block.
[0040] In JVET, a CTU 100 can be partitioned into CUs 102 according to the quadtree plus binary tree (QTBT) scheme, where a CTU 100 can be recursively split into square blocks according to the quadtree and then these square blocks can be recursively split horizontally or vertically according to the binary tree. Parameters such as the CTU size, the minimum size for quadtree and binary tree leaf nodes, the maximum size for the binary tree root node, and the maximum depth for the binary tree can be set to control the splitting according to QTBT.
[0041] In some embodiments, JVET can limit the binary splitting in the binary tree part of the QTBT to symmetric splitting, where a block can be divided vertically or horizontally into two halves along the midline.
[0042] By way of non-limiting example, Figure 2 Shows a CTU 100 partitioned into CUs 102 by solid lines indicating quadtree splitting and dashed lines indicating symmetric binary tree splitting. As shown, the binary splitting allows symmetric horizontal and vertical splitting to define the structure of the CTU and its subdivision into CUs.
[0043] Figure 3 Shows Figure 2The segmented QTBT representation. The quadtree root node represents the CTU 100, and each child node in the quadtree part represents one of the four square blocks separated from the parent square block. Then, the square block represented by the quadtree leaf node can be symmetrically divided zero or more times using a binary tree, where the quadtree leaf node is the root node of the binary tree. At each level of the binary tree part, the block can be divided symmetrically, vertically, or horizontally. A flag set to "0" indicates that the block is symmetrically separated horizontally, while a flag set to "1" indicates that the block is symmetrically separated vertically.
[0044] In other embodiments, JVET may allow symmetric binary partitioning or asymmetric binary partitioning in the binary tree part of the QTBT. When partitioning the prediction unit (PU), asymmetric motion partitioning (AMP) is allowed in different contexts in HEVC. However, for partitioning the CU 102 in JVET according to the QTBT structure, when the relevant region of the CU 102 is not located on either side of the midline passing through the center of the CU, asymmetric binary partitioning can result in improved partitioning compared to symmetric binary partitioning. By way of non-limiting example, when the CU 102 depicts one object near the center of the CU and another object at the side of the CU 102, the CU 102 can be asymmetrically partitioned to place each object in a separate smaller CU 102 of different sizes.
[0045] Figure 4 Depicts four possible types of asymmetric binary partitioning, where the CU 102 is separated into two smaller CUs 102 along a line passing through the length or height of the CU 102, such that one of the smaller CUs 102 is 25% of the size of the parent CU 102 and the other is 75% of the size of the parent CU 102. Figure 4 The four types of asymmetric binary partitioning shown in allow the CU102 to be separated along a line that is 25% of the route starting from the left side of the CU 102, 25% of the route starting from the right side of the CPU 102, 25% of the route starting from the top of the CPU102, or 25% of the route starting from the bottom of the CPU 102. In an alternative embodiment, the asymmetric partitioning line at which the CU 102 is separated can be located at any other position such that the CU102 is not symmetrically divided into two halves.
[0046] Figure 5 Depicts a non-limiting example of a CTU 100 divided into CUs 102 using a scheme that allows both symmetric binary partitioning and asymmetric binary partitioning in the binary tree part of the QTBT. In Figure 5 the dashed line shows the asymmetric binary partitioning line, where the parent CU 102 is separated using one of the partitioning types shown in Figure 4 one of the partitioning types shown in.
[0047] Figure 6 shows Figure 5 the partitioned QTBT representation. In Figure 6 it, two solid lines extending from a node indicate a symmetric partition in the binary tree part of the QTBT, while two dashed lines extending from a node indicate an asymmetric partition in the binary tree part.
[0048] Syntax can be compiled in the bitstream that indicates how to partition CTU 100 into CUs 102. By way of non-limiting example, syntax can be compiled in the bitstream that indicates which nodes are separated by quadtree partitioning, which nodes are separated by symmetric binary partitioning, and which nodes are separated by asymmetric binary partitioning. Similarly, syntax can be compiled in the bitstream for nodes separated by asymmetric binary partitioning that indicates which type of asymmetric binary partitioning is used, such as Figure 4 one of the four types shown in
[0049] In some embodiments, the use of asymmetric partitioning can be limited to partitioning CU 102 at the leaf nodes of the quadtree part of the QTBT. In these embodiments, the CU 102 at the child nodes separated from a parent node using quadtree partitioning in the quadtree part can be the final CU 102, or they can be further separated using quadtree partitioning, symmetric binary partitioning, or asymmetric binary partitioning. The child nodes in the binary tree part separated using symmetric binary partitioning can be the final CU 102, or they can be further recursively separated one or more times using only symmetric binary partitioning. In cases where no further separation is allowed, the child nodes in the binary tree part separated from a QT leaf node using asymmetric binary partitioning can be the final CU 102.
[0050] In these embodiments, limiting the use of asymmetric partitioning to separating quadtree leaf nodes can reduce search complexity and / or limit overhead bits. Since only quadtree leaf nodes can be separated by asymmetric partitioning, the use of asymmetric partitioning can directly indicate the end of a QT part branch without additional syntax or further signaling. Similarly, since the nodes with asymmetric partitioning cannot be further separated, using asymmetric partitioning on a node can also directly indicate that its child nodes with asymmetric partitioning are the final CUs 102 without additional syntax or further signaling.
[0051] In alternative embodiments, such as when limiting search complexity and / or the number of overhead bits becomes immaterial, asymmetric partitioning can be used to separate nodes generated by quadtree partitioning, symmetric binary partitioning, and / or asymmetric binary partitioning.
[0052] After performing quadtree splitting and binary tree splitting using any of the above QTBT structures, the blocks represented by the leaf nodes of the QTBT represent the final CUs 102 to be compiled, such as compilation using inter prediction or intra prediction. For slices or full frames compiled using inter prediction, different splitting structures can be used for the luminance and chrominance components. For example, for an inter slice, the CU 102 can have compiled blocks (CBs) for different color components, such as one luminance CB and two chrominance CBs. For slices or full frames compiled using intra prediction, the splitting structures for the luminance and chrominance components can be the same.
[0053] In an alternative embodiment, JVET can use a two-level compiled block structure as an alternative or extension to the above QTBT splitting. In the two-level compiled block structure, the CTU 100 can first be split into basic units (BUs) at a higher level. Then the BUs can be split into multiple operation units (OUs) at a lower level.
[0054] In an embodiment employing the two-level compiled block structure, at a higher level, the CTU 100 can be split into BUs according to one of the above QTBT structures or according to a quadtree (QT) structure such as the quadtree (QT) structure used in HEVC where a block can be split only into four equally sized sub-blocks. By way of non-limiting example, the CTU 102 can be split into BUs according to the QTBT structure described above with respect to Figures 5 - 6 such that the leaf nodes in the quadtree part can be split using quadtree splitting, symmetric binary splitting, or asymmetric binary splitting. In this example, the final leaf nodes of the QTBT can be BUs instead of CUs.
[0055] At a lower level in the two-level compiled block structure, each BU split from the CTU 100 can be further split into one or more OUs. In some embodiments, when the BU is square, quadtree splitting or binary splitting, such as symmetric or asymmetric binary splitting, can be used to split it into OUs. However, when the BU is not square, only binary splitting can be used to split it into OUs. Limiting the type of splitting that can be used for non-square BUs can limit the number of bits used to signal the type of splitting used to generate the BUs.
[0056] Although the following discussion describes compiling CU 102, in embodiments using a two-level compilation block structure, the BU and OU rather than CU 102 can be compiled. By way of non-limiting example, the BU can be used for higher-level compilation operations such as intra prediction or inter prediction, while the smaller OU can be used for lower-level compilation operations such as transform and generation of transform coefficients. Thus, the syntax for the BU that indicates whether compilation is performed by intra prediction or inter prediction, or information identifying a particular intra prediction mode or motion vector, is used to encode the BU. Similarly, the syntax for the OU can identify a particular transform operation or quantized transform coefficients for compiling the OU.
[0057] Figure 7 A simplified block diagram depicting CU compilation in a JVET encoder is shown. The main stages of video compilation include splitting to identify CU 102 as described above, subsequently encoding CU 102 using prediction at 704 or 706, generating a residual CU 710 at 708, performing a transform at 712, quantization at 716, and entropy compilation at 720. Figure 7 The encoder and encoding process illustrated in also includes a decoding process, which will be described in more detail below.
[0058] Given the current CU 102, the encoder can obtain a predicted CU 702 by using intra prediction spatially at 704 or inter prediction temporally at 706. The basic idea of predictive compilation is to transmit the difference or residual signal between the original signal and the prediction of the original signal. At the receiver side, the original signal can be reconstructed by adding the residual and the prediction, as will be described below. Since the difference signal has lower correlation than the original signal, fewer bits are required for its transmission.
[0059] A slice that is completely compiled with intra prediction of CU 102, such as an entire picture or a part of a picture, can be an I slice, which can be decoded without reference to other slices, and as such, it can be a possible starting point for decoding. A slice compiled with at least some inter prediction of CU can be a predicted (P) or bi-predicted (B) slice that can be decoded based on one or more reference pictures. A P slice can use both intra prediction and inter prediction with previously compiled slices. For example, by using inter prediction, a P slice can be further compressed compared to an I slice, but the previously compiled slices need to be compiled to be able to compile them. A B slice can use intra prediction or inter prediction that applies interpolation prediction from two different frames, encoding using data from previous and / or subsequent slices, thereby increasing the accuracy of the motion estimation process. In some cases, intra block copy can also be used to encode or alternatively encode P and B slices, where data from other parts of the same slice is used.
[0060] As will be discussed below, intra prediction or inter prediction can be performed based on reconstructed CUs 734 from previously encoded CUs 102, such as neighboring CUs 102 or CUs 102 in a reference picture.
[0061] When spatially encoding a CU 102 using intra prediction at 704, an intra prediction mode can be found that best predicts the pixel values of the CU 102 based on samples from neighboring CUs 102 in the picture.
[0062] When encoding the luminance component of a CU, the encoder can generate a list of candidate intra prediction modes. Although HEVC has 35 possible intra prediction modes for the luminance component, in JVET, there are 67 possible intra prediction modes for the luminance component. These include a planar mode using a three-dimensional plane of values generated from neighboring pixels, a DC mode using the average of values from neighboring pixels, and 65 directional modes shown in Figure 8 in which values are copied from neighboring pixels along the indicated directions.
[0063] When generating a list of candidate intra prediction modes for the luminance component of a CU, the number of candidate modes on the list can depend on the size of the CU. The candidate list can include: a subset of the 35 HEVC modes that have the lowest SATD (Sum of Absolute Transform Differences) cost; new directional modes added for JVET that are adjacent to candidates found from HEVC modes; and modes from among a set of six most probable modes (MPMs) for the CU 102 identified based on the intra prediction modes for previously encoded neighboring blocks and a default mode list.
[0064] When encoding the chrominance component of a CU, a list of candidate intra prediction modes can also be generated. The candidate mode list can include modes generated using projections from a cross-component linear model from luminance samples, intra prediction modes found for luminance CBs at specific co-location positions in the chrominance block, and previously found chrominance prediction modes for neighboring blocks. The encoder can find the candidate mode with the lowest rate-distortion cost in the list and use these intra prediction modes when encoding the luminance and chrominance components of the CU. The syntax can be encoded in the bitstream indicating the intra prediction mode used for encoding each CU 102.
[0065] After the best intra prediction modes for the CU 102 have been selected, the encoder can use those modes to generate a predicted CU 402. When the selected mode is a directional mode, a 4-tap filter can be used to improve the directional accuracy. A boundary prediction filter, such as a 2-tap or 3-tap filter, can be used to adjust columns or rows at the top or left of the predicted block.
[0066] The prediction CU 702 can be further smoothed using a Position-Dependent Intra Prediction Combination (PDPC) process that uses unfiltered samples of neighboring blocks to adjust the predicted CU 702 generated based on filtered samples of neighboring blocks, or an Adaptive Reference Sample Smoothing using a 3-tap or 5-tap low-pass filter to process the reference samples.
[0067] When the CU 102 is temporally predicted by inter prediction at 706, a set of motion vectors (MVs) can be found that point to samples in a reference picture of the pixel values of the best predicted CU 102. Inter prediction exploits the temporal redundancy between slices by representing the displacement of pixel blocks in a slice. The displacement is determined by a process called motion compensation based on the pixel values in a previous or subsequent slice. The motion vectors indicating the displacement of pixels relative to a particular reference picture and the associated reference index can be provided to the decoder in the bitstream together with the residual between the original pixels and the motion-compensated pixels. The decoder can use the residual and the signaled motion vectors and reference index to reconstruct the pixel block in the reconstructed slice.
[0068] In JVET, the motion vector precision can be stored at 1 / 16 pixel, and the difference between the motion vector and the predicted motion vector of the CU can be coded at a quarter-pixel resolution or an integer-pixel resolution.
[0069] In JVET, techniques such as Advanced Temporal Motion Vector Prediction (ATMVP), Spatio-Temporal Motion Vector Prediction (STMVP), Affine Motion Compensation Prediction, Pattern Matching Motion Vector Derivation (PMMVD), and / or Bidirectional Optical Flow (BIO) can be used to find motion vectors for multiple sub-CUs within the CU 102.
[0070] Using ATMVP, the encoder can find the temporal vector of the CU 102 that points to the corresponding block in the reference picture. The temporal vector can be found based on the motion vectors and reference pictures found for previously coded adjacent CUs 102. Using the reference block pointed to by the temporal vector for the entire CU 102, motion vectors can be found for each sub-CU within the entire CU 102.
[0071] STMVP can find the motion vectors of the sub-CUs by scaling and averaging the motion vectors and temporal vectors found for neighboring blocks previously coded using inter prediction.
[0072] Affine Motion Compensation Prediction can be used to predict the field of motion vectors for each sub-CU in a block based on two control motion vectors found for the top corners of the block. For example, the motion vectors of the sub-CUs can be derived based on the top corner motion vectors found for each 4x4 block within the CU 102.
[0073] The PMMVD can use bidirectional matching or template matching to find the initial motion vector of the current CU 102. Bidirectional matching can look at the current CU 102 and reference blocks in two different reference pictures along the motion trajectory, while template matching can look at the corresponding block in the current CU 102 and the reference pictures identified by the template. Then, the initial motion vectors found for the CU 102 can be corrected separately for each sub-CU.
[0074] BIO can be used when performing inter-frame prediction in bi-directional prediction based on earlier and later reference pictures, and allows finding motion vectors for sub-CUs based on the gradient of the difference between the two reference pictures.
[0075] In some cases, local luminance compensation (LIC) can be used at the CU level to find the values for the scaling factor parameter and the offset parameter based on the samples adjacent to the current CU 102 and the corresponding samples adjacent to the reference block identified by the candidate motion vector. In JVET, the LIC parameters can be changed and signaled at the CU level.
[0076] For some of the above methods, the motion vectors found for each sub-CU of the CU can be signaled to the decoder at the CU level. For other methods such as PMMVD and BIO, the motion information is not signaled in the bitstream to save overhead, and the decoder can derive the motion vectors through the same process.
[0077] After the motion vectors for the CU 102 have been found, the encoder can use those motion vectors to generate the predicted CU 702. In some cases, when the motion vectors have been found for a single sub-CU, overlapping block motion compensation (OBMC) can be used when generating the predicted CU 702 by combining those motion vectors with the motion vectors previously found for one or more neighboring sub-CUs.
[0078] When using bi-directional prediction, JVET can use decoder-side motion vector refinement (DMVR) to find the motion vectors. DMVR allows finding motion vectors based on two motion vectors found for bi-directional prediction using a bi-directional template matching process. In DMVR, a weighted combination of the predicted CU 702 generated by each of the two motion vectors can be found, and the two motion vectors can be refined by replacing them with a new motion vector that optimally points to the combined predicted CU 702. The two refined motion vectors can be used to generate the final predicted CU 702.
[0079] At 708, as described above, once the predicted CU 702 is found through intra-frame prediction at 704 or inter-frame prediction at 706, the encoder can subtract the predicted CU 702 from the current CU 102 to find the residual CU 710.
[0080] The encoder may use one or more transform operations at 712 to convert the residual CU 710 into transform coefficients 714 that represent the residual CU 710 in the transform domain, such as using a Discrete Cosine Block Transform (DCT transform) to transform the data into the transform domain. Compared to HEVC, JVET allows for more types of transform operations, including DCT-II, DST-VII, DST-VII, DCT-VIII, DST-I, and DCT-V operations. The allowed transform operations may be grouped into subsets, and the encoder may signal an indication of which subsets and which specific operations within those subsets are used. In some cases, a transform of a large block size may be used to zero out the high-frequency transform coefficients in CUs 102 larger than a particular size, such that only the low-frequency transform coefficients are retained for those CUs 102.
[0081] In some cases, after the forward kernel transform, a mode-dependent non-separable second-order transform (MDNSST) may be applied to the low-frequency transform coefficients 714. The MDNSST operation may use a Hypercube-Givens transform (HyGT) based on rotated data. When used, the encoder may signal an index value that identifies a particular MDNSST operation.
[0082] At 716, the encoder may quantize the transform coefficients 714 to quantized transform coefficients 716. The quantization of each coefficient may be calculated by dividing the value of the coefficient by a quantization step size, which is derived from a quantization parameter (QP). In some embodiments, Qstep is defined as 2 (QP-4) / 6 . Since the high-precision transform coefficients 714 can be converted to quantized transform coefficients 716 with a finite number of possible values, quantization can assist in data compression. Thus, the quantization of the transform coefficients can limit the amount of bits generated and sent by the transform process. However, although quantization is a lossy operation and the quantization loss cannot be recovered, the quantization process presents a trade-off between the quality of the reconstructed sequence and the amount of information required to represent the sequence. For example, a lower QP value may result in a better quality decoded video, although a higher amount of data may be required for representation and transmission. Conversely, a high QP value results in a lower quality reconstructed video sequence but lower data and bandwidth requirements.
[0083] JVET can utilize variance-based adaptive quantization techniques, which allow each CU 102 to use different quantization parameters for its encoding or process (instead of using the same frame QP for the encoding of each CU 102 in a frame). The variance-based adaptive quantization techniques can adaptively reduce the quantization parameters for some blocks while increasing the quantization parameters in other blocks. To select a specific QP for the CU102, the variance of the CU is calculated. In short, if the variance of the CU is higher than the average variance of the frame, a QP higher than the frame's QP can be set for the CU102. If the CU 102 exhibits a variance lower than the average variance of the frame, a lower QP can be assigned.
[0084] At 720, the encoder can find the final compressed bits 722 by performing entropy encoding on the quantized transform coefficients 718. Entropy encoding aims to remove the statistical redundancy of the information to be sent. In JVET, CABAC (Context-Adaptive Binary Arithmetic Coding) can be used to encode the quantized transform coefficients 718, which uses probability metrics to remove statistical redundancy. For a CU 102 with non-zero quantized transform coefficients 718, the quantized transform coefficients 718 can be converted to binary. Then, each bit ("bin") of the binary representation can be encoded using a context model. The CU 102 can be decomposed into three regions, each region having its own set of context models for the pixels within that region.
[0085] Multiple scan operations can be performed to encode the bins. In the operation of encoding the first three bins (bin0, bin1, and bin2), the index value indicating the context model to be used for that bin can be found by finding the sum of the bin positions in up to five previously encoded neighboring quantized transform systems identified by a template.
[0086] The context model can be based on the probability that the value of the bin is "0" or "1". When encoding the values, the probabilities in the context model can be updated based on the actual number of "0" and "1" values encountered. While HEVC uses a fixed table to re-initialize the context model for each new picture, in JVET, the probabilities of the context model for a new inter-predicted picture can be initialized based on the context model developed for the previously encoded inter-predicted pictures.
[0087] The encoder can generate a bitstream that contains the entropy-encoded bits 722 of the residual CU 710, prediction information such as the selected intra-prediction mode or motion vectors, an indicator of how the CU 102 is split from the CTU 100 according to the QTBT structure, and / or other information about the encoded video. The bitstream can be decoded by the decoder, as described below.
[0088] In addition to using the quantized transform coefficients 718 to find the final compressed bits 722, the encoder can also use the quantized transform coefficients 718 to generate a reconstructed CU 734 by following the same decoding process that the decoder will use to generate the reconstructed CU 734. Thus, once the encoder computes and quantizes the transform coefficients, the quantized transform coefficients 718 can be sent to the decoding loop in the encoder. After quantizing the transform coefficients of a CU, the decoding loop allows the encoder to generate the same reconstructed CU 734 that the decoder generates during the decoding process. Thus, the encoder can use the same reconstructed CU 734 that would be used for neighboring CUs 102 or reference pictures when the decoder performs intra prediction or inter prediction for the new CU 102. The reconstructed CU 102, reconstructed slice, or fully reconstructed frame can be used as a reference for further prediction stages.
[0089] At the decoding loop of the encoder where the pixel values of the reconstructed image are obtained (see below for the same operation in the decoder), a dequantization process can be performed. To dequantize a frame, for example, the quantized value of each pixel of the frame is multiplied by the quantization step, e.g., the above (Qstep), to obtain the reconstructed dequantized transform coefficients 726. For example, in the Figure 7 decoding process shown in the encoder, the quantized transform coefficients 718 of the residual CU 710 can be dequantized at 724 to find the dequantized transform coefficients 726. If the MDNSST operation is performed during encoding, this operation can be inverted after dequantization.
[0090] At 728, the dequantized transform coefficients 726 can be inverse-transformed to find the reconstructed residual CU 730, such as by applying the DCT to these values to obtain the reconstructed image. At 732, the reconstructed residual CU 730 can be added to the corresponding predicted CU 702 found by intra prediction at 704 or inter prediction at 706 in order to find the reconstructed CU 734.
[0091] At 736, one or more filters can be applied to the reconstructed data at the picture level or at the CU level during the decoding process (in the encoder or, as described below, in the decoder). For example, the encoder can apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). The decoding process of the encoder may implement filters to estimate the best filter parameters that can address potential artifacts in the reconstructed image and send them to the decoder. Such improvements increase the objective and subjective quality of the reconstructed video. In deblocking filtering, pixels near the sub-CU boundary can be modified, while in SAO, pixels in the CTU 100 can be modified using edge offsets or band offset classifications. The JVET's ALF can use a filter with a circular symmetric shape for each 2x2 block. An indication of the size and identity of the filter for each 2x2 block can be signaled.
[0092] If the reconstructed pictures are reference pictures, they can be stored in the reference buffer 738 for inter-prediction of future CUs 102 at 706.
[0093] During the above steps, JVET allows the use of content-adaptive cropping operations to adjust color values to fit the top and bottom cropping bounds. The cropping bounds can vary for each slice, and parameters identifying the bounds can be signaled in the bitstream.
[0094] Figure 9 A simplified block diagram depicting CU compilation in the JVET decoder is shown. The JVET decoder can receive a bitstream containing information about the encoded CU 102. The bitstream can indicate how to partition the CUs 102 of a picture from the CTU 100 according to the QTBT structure. As a non-limiting example, the bitstream can use quadtree partitioning, symmetric binary partitioning, and / or asymmetric binary partitioning to identify how to partition the CUs 102 from each CTU 100 in the QTBT. The bitstream can also indicate prediction information for the CU 102, such as an intra-prediction mode or a motion vector, and the bits 902 representing the entropy-coded residual CU.
[0095] At 904, the decoder can use the CABAC context model signaled by the encoder in the bitstream to decode the entropy-coded bits 902. The decoder can update the probabilities of the context model using the parameters signaled by the encoder in the same way as was done during the encoding process.
[0096] After reversing the entropy coding at 904 to find the quantized transform coefficients 906, the decoder can dequantize them at 908 to find the dequantized transform coefficients 910. If the MDNSST operation was performed during encoding, this operation can be reversed by the decoder after dequantization.
[0097] At 912, the dequantized transform coefficients 910 may be inverse-transformed to find the reconstructed residual CU 914. At 916, the reconstructed residual CU 914 may be added to the corresponding predicted CU 926 found using intra prediction at 922 or inter prediction at 924 to facilitate finding the reconstructed CU 918.
[0098] At 920, one or more filters may be applied to the reconstructed data at the picture level or CU level. For example, the decoder may apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). As described above, the in-loop filter located in the decoding loop of the encoder may be used to estimate optimal filter parameters to increase the objective and subjective quality of the frame. These parameters are sent to the decoder to filter the reconstructed frame at 920 to match the filtered reconstructed frame in the encoder.
[0099] After the reconstructed picture has been generated by finding the reconstructed CU 918 and applying the signaled filters, the decoder may output the reconstructed picture as the output video 928. If the reconstructed picture is used as a reference picture, it may be stored in the reference buffer 930 for inter prediction of future CUs 102 at 924.
[0100] Figure 10 An embodiment of a method for CU compilation 1000 in a JVET decoder is depicted. In Figure 10 the embodiment shown, at step 1002, an encoded bitstream 902 may be received, and then at step 1004, the CABAC context model associated with the encoded bitstream 902 may be determined, and then the determined CABAC context model may be used to decode the encoded bitstream 902 at step 1006.
[0101] At step 1008, the quantized transform coefficients 906 associated with the encoded bitstream 902 may be determined, and then the dequantized transform coefficients 910 may be determined from the quantized transform coefficients 906 at step 1010.
[0102] In step 1012, it can be determined whether the MDNSST operation is performed during encoding and / or whether the bitstream 902 contains an indication to apply the MDNSST operation to the bitstream 902. If it is determined that the MDNSST operation is performed during the encoding process or the bitstream 902 contains an indication to apply the MDNSST operation to the bitstream 902, then the inverse MDNSST operation 1014 can be implemented before performing the inverse transform operation 912 on the bitstream 902 in step 1016. Alternatively, the operation 912 can be performed on the bitstream 902 in step 1016 without applying the inverse MDNSST operation in step 1014. The inverse transform operation 912 in step 1016 can determine and / or construct the reconstructed residual CU 914.
[0103] In step 1018, the reconstructed residual CU 914 from step 1016 can be combined with the predicted CU 918. The predicted CU 918 can be one of the intra-predicted CU 922 determined in step 1020 and the inter-predicted unit 924 determined in step 1022.
[0104] In step 1024, any one or more filters 920 can be applied to the reconstructed CU 914 and output in step 1026. In some embodiments, the filter 920 may not be applied in step 1024.
[0105] In some embodiments, in step 1028. The reconstructed CU 918 can be stored in the reference buffer 930.
[0106] Figure 11 A simplified block diagram 1100 depicting CU compilation for the JVET encoder is shown. In step 1102, the JVET compilation tree unit can be represented as a root node in a quadtree plus binary tree (QTBT) structure. In some embodiments, the QTBT can have a quadtree branching from the root node and / or a binary tree branching from the leaf nodes of one or more quadtrees. The representation from step 1102 can proceed to step 1104, 1106, or 1108.
[0107] In step 1104, the represented quadtree node can be separated into two blocks of unequal size using asymmetric binary splitting. In some embodiments, the separated blocks can be represented in a binary tree branching from a quadtree node that is a leaf node capable of representing the final compiled unit. In some embodiments, further separation is not allowed in a binary tree branching from a quadtree node that is a leaf node representing the final compiled unit. In some embodiments, the asymmetric split can separate the compiled unit into blocks of unequal size, with the first block representing 25% of the quadtree node and the second block representing 75% of the quadtree node.
[0108] In step 1106, quadtree splitting may be employed to split the represented quadtree node into four equally sized square blocks. In some embodiments, the separated blocks may be represented as quadtree nodes representing the final compilation units, or may be represented as child nodes that can be split again by quadtree splitting, symmetric binary splitting, or asymmetric binary splitting.
[0109] In step 1108, quadtree splitting may be used to separate the represented quadtree node into two equally sized blocks. In some embodiments, the separated blocks may be represented as quadtree nodes representing the final compilation units, or may be represented as child nodes that can be split again by quadtree splitting, symmetric binary splitting, or asymmetric binary splitting.
[0110] In step 1110, the child nodes from step 1106 or step 1108 may be represented as child nodes configured to be encoded. In some embodiments, the child nodes may be represented by leaf nodes of a binary tree using JVET.
[0111] In step 1112, JVET may be used to encode the compilation units from step 1104 or 1110.
[0112] Figure 12 A simplified block diagram 1200 depicting CU decoding in a JVET decoder is shown. In Figure 12 the depicted embodiment, in step 1202, a bitstream may be received indicating how to split the compilation tree unit into compilation units according to the QTBT structure. The bitstream may indicate how to separate the quadtree node by at least one of quadtree splitting, symmetric binary splitting, or asymmetric binary splitting.
[0113] In step 1204, the compilation units represented by the leaf nodes of the QTBT structure may be identified. In some embodiments, the compilation unit may indicate whether the node uses asymmetric binary splitting to separate the leaf node from the quadtree. In some embodiments, the compilation unit may indicate that the node represents the final compilation unit to be decoded.
[0114] In step 1206, JVET may be used to decode the identified compilation units.
[0115] Figure 13 An alternative simplified block diagram depicting JVET compilation for intra mode prediction 1300 is shown. In Figure 13In the embodiments depicted, in step 1302, a set of MPMs can be identified and instantiated in memory, and then in step 1304, a set of 16 selected modes can be identified and instantiated in memory, and in step 1304, a balance of 67 modes can be defined and instantiated in storage. In some embodiments, the set of MPMs can be reduced from a standard set of 6 MPMs. In some embodiments, the set of MPMs can include 5 unique modes, the selected modes can include 16 unique modes, and the set of unselected modes can include the remaining 46 unselected unique modes. However, in alternative embodiments, the set of MPMs can include fewer unique modes, the selected modes can remain fixed at 16 unique modes, and the size of the set of unselected unique modes can be adjusted accordingly to accommodate a total of 67 modes.
[0116] By way of non-limiting example, in some embodiments, where the set of MPMs includes 5 unique modes instead of six MPMs, thus, if truncated unary binarization is used and a new binarization for 5 MPMs can be exploited, the number of bins assigned to the MPM modes can thus be equal to or less than five bins. Thus, in some embodiments, 16 selected modes among the 62 remaining intra modes can be generated by uniformly downsampling these 62 intra modes, and each mode can be encoded by a 4-bit fixed-length code. By way of non-limiting example, if it is assumed that the remaining 62 modes are indexed as {0, 1, 2, …, 61}, then the 16 selected modes = {0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 60}. The remaining 46 unselected modes = {1, 2, 3, 5, 6, 7, 9, 10 … 59, 61}, and these 46 unselected modes can be encoded using truncated binary codes.
[0117] Figure 14 Depicted according to Figure 13 Table 1400 of an alternative JVET compilation for intra mode prediction. In Figure 14 the embodiments depicted, the intra prediction mode 1402 is shown to include 5 MPMs, 16 selected modes, and 46 unselected modes, where a binary string 1404 for the MPMs can be encoded using truncated unary binarization, the 16 selected modes can be encoded using a 4-bit fixed-length code, and the 46 unselected modes can be encoded using truncated binary compilation.
[0118] In Figure 13 alternative embodiments, 6 MPMs can be exploited, but as Figure 14As shown, only the first five MPMs on the MPM list are binarized and compiled using the context-based method described in the current JVET. Now, the sixth MPM on the MPM list is regarded as one of the 16 selected modes and is compiled together with the other 15 selected modes using a fixed-length code of 4 bits.
[0119] As a non-limiting example, if the remaining 61 modes are indexed as {0, 1, 2, …, 60}, the following 15 selected modes can be obtained by uniformly subsampling the remaining 61 intra modes: the set of selected modes can be {0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, where the 15 selected modes plus the sixth MPM are compiled using a fixed-length code of 4 bits, as in the following set: {sixth MPM, 0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, and the balance of the 46 unselected modes is shown in the following set and is compiled into the set of unselected modes = {1, 2, 3, 4, 6, 7, 8, 9, 11, 12 … 49, 51, 52, 53, 54, 56, 57, 58, 59} using a truncated binary code.
[0120] In Figure 13 yet another alternative embodiment, only the first five MPMs on the MPM list can be binarized, as Figure 14 shown, and compiled using the context-based method described in the current JVET standard. In such an embodiment, the sixth MPM on the MPM list can be considered as one of the 16 selected modes and is compiled together with the other 15 selected modes using a fixed-length code of 4 bits. Thus, any known convenient and / or desired selection process can be used to establish the selection of the other 15 selected modes. By non-limiting example, they can be selected around the MPM mode, or around (content-based) statistically popular modes, or around trained or historical popular modes, or using other methods or processes.
[0121] Again, selecting 5 MPMs is only a non-limiting example, and in alternative embodiments, the set of MPMs can be further reduced to 4 or 3 MPMs or extended to more than 6, where there are still 16 selected modes, and the balance of 67 (or other known, convenient, and / or desired total) intra-compiled modes is included in the set of unselected intra-compiled modes. That is, embodiments where the total number of intra-compiled modes is greater than or less than 67 can be envisioned as embodiments where the set of MPMs can contain any known convenient or desired number of MPMs, and the number of selected modes can be any known convenient and / or desired number.
[0122] The execution of the instruction sequence required for the practice embodiment can be performed by a computer system 1500 as shown in Figure 15 . In an embodiment, the execution of the instruction sequence is performed by a single computer system 1500. According to other embodiments, two or more computer systems 1500 coupled by a communication link 1515 can execute the instruction sequence in coordination with each other. Although only one computer system 1500 will be described below, it should be understood that any number of computer systems 1500 can be employed to practice the embodiment.
[0123] Reference will now be made to Figure 15 describe a computer system 1500 according to an embodiment, Figure 15 which is a block diagram of the functional components of the computer system 1500. As used herein, the term computer system 1500 is used broadly to describe any computing device that can store and independently run one or more programs.
[0124] Each computer system 1500 can include a communication interface 1514 coupled to a bus 1506. The communication interface 1514 provides two-way communication between computer systems 1500. The communication interfaces 1514 of the respective computer systems 1500 send and receive electrical, electromagnetic, or optical signals, which include data streams representing various types of signal information, such as instructions, messages, and data. The communication link 1515 links one computer system 1500 to another computer system 1500. For example, the communication link 1515 can be a LAN, in which case the communication interface 1514 can be a LAN card, or the communication link 1515 can be a PSTN, in which case the communication interface 1514 can be an integrated services digital network (ISDN) card or a modem, or the communication link 1515 can be the Internet, in which case the communication interface 1514 can be a dial-up, cable, or wireless modem.
[0125] The computer systems 1500 can send and receive messages, data, and instructions, including programs, i.e., application programs, code, via their respective communication links 1515 and communication interfaces 1514. When received, the received program code can be executed by their respective processors 1507 and / or stored in a storage device 1510 or other associated non-volatile media for later execution.
[0126] In an embodiment, computer system 1500 operates in conjunction with a data storage system 1531, for example, a data storage system 1531 that includes a database 1532, which is readily accessible by computer system 1500. Computer system 1500 communicates with data storage system 1531 via a data interface 1533. Data interface 1533, coupled to bus 1506, transmits and receives electrical, electromagnetic, or optical signals, including a data stream representing various types of signal information, such as instructions, messages, and data. In an embodiment, the functionality of data interface 1533 may be performed by communication interface 1514.
[0127] Computer system 1500 includes: a bus 1506 or other communication mechanism for conveying instructions, messages, and data, collectively, information; and one or more processors 1507 coupled to bus 1506 to process the information. Computer system 1500 also includes a main memory 1508, such as random access memory (RAM) or other dynamic storage device, coupled to bus 1506 for storing dynamic data and instructions to be executed by processor 1507. Main memory 1508 may also be used to store temporary data, i.e., variables or other intermediate information, during the execution of instructions by processor 1507.
[0128] Computer system 1500 may further include a read-only memory (ROM) 1509 or other static storage device coupled to bus 1506 for storing static data and instructions for processor 1507. A storage device 1510, such as a magnetic disk or optical disk, may also be provided and coupled to bus 1506 for storing data and instructions for processor 1507.
[0129] Computer system 1500 may be coupled via bus 1506 to a display device 1511, such as, but not limited to, a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor, for displaying information to a user. An input device 1512, such as alphanumeric keys and other keys, is coupled to bus 1506 for transmitting information and command selections to processor 1507.
[0130] According to one embodiment, a single computer system 1500 performs particular operations by executing one or more sequences of one or more instructions contained in main memory 1508 by their respective processors 1507. Such instructions may be read into main memory 1508 from another computer-usable medium, such as ROM 1509 or storage device 1510. Execution of the instruction sequences contained in main memory 1508 causes processor 1507 to perform the processes described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Accordingly, embodiments are not limited to any specific combination of hardware circuitry and / or software.
[0131] As used herein, the term "computer-usable medium" refers to any medium that provides information or can be used by processor 1507. Such a medium can take many forms, including but not limited to non-volatile, volatile, and transmission media. Non-volatile media, i.e., media that can retain information without power, include ROM 1509, CD ROM, magnetic tape, and magnetic disks. Volatile media, i.e., media that cannot retain information without power, include main memory 1508. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up bus 1506. Transmission media can also take the form of a carrier wave; i.e., an electromagnetic wave that can be modulated in terms of frequency, amplitude, or phase to send information signals. Additionally, transmission media can take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0132] In the foregoing specification, embodiments have been described with reference to specific elements of the embodiments. However, it will be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the embodiments. For example, the reader will understand that the specific order and combination of process actions shown in the process flow diagrams described herein are merely illustrative, and different or additional process actions can be used, or different combinations or orders of process actions can be used to implement the embodiments. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0133] It should also be noted that the present invention can be implemented in various computer systems. The various techniques described herein can be implemented in hardware or software or a combination of both. Preferably, the techniques are implemented in a computer program executed on a programmable computer, which includes a processor, a processor-readable storage medium (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The program code is applied to the data input using the input device to perform the above functions and generate output information. The output information is applied to one or more output devices. Each program is preferably implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly language or machine language. In any case, the language can be a compiled language or an interpreted language. Each such computer program is preferably stored on a storage medium or device readable by a general or special programmable computer (e.g., ROM or disk) for configuring and operating the computer when the storage medium or device is read by the computer to perform the above process. The system can also be considered to be implemented as a computer-readable storage medium configured with a computer program, wherein the storage medium so configured causes the computer to operate in a specific and predetermined manner. In addition, the storage elements of an exemplary computing application can be a relational or sequential (flat file) type of computing database capable of storing data in various combinations and configurations.
[0134] Figure 16 is a high-level view of source device 1612 and destination device 1610 that can incorporate the features of the systems and devices described herein. As Figure 16 shown, example video compilation system 1610 includes source device 1612 and destination device 1616, where, in this example, source device 1612 generates encoded video data. Thus, source device 1612 can be referred to as a video encoding device. Destination device 1616 can decode the encoded video data generated by source device 1612. Thus, destination device 1616 can be referred to as a video decoding device. Source device 1612 and destination device 1616 can be examples of video compilation devices.
[0135] Destination device 1616 can receive the encoded video data from source device 1612 via channel 1616. Channel 1616 can include a type of medium or device capable of moving the encoded video data from source device 1612 to destination device 1616. In one example, channel 1616 can include a communication medium that enables source device 1612 to send the encoded video data directly and in real time to destination device 1616.
[0136] In this example, the source device 1612 can modulate the encoded video data according to a communication standard such as a wireless communication protocol and can send the modulated video data to the destination device 1616. The communication medium can include a wireless or wired communication medium such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or other devices that facilitate communication from the source device 1612 to the destination device 1616. In another example, the channel 1616 can correspond to a storage medium that stores the encoded video data generated by the source device 1612.
[0137] In Figure 16 the example of, the source device 1612 includes a video source 1618, a video encoder 1620, and an output interface 1622. In some cases, the output interface 1628 can include a modulator / demodulator (modem) and / or a transmitter. In the source device 1612, the video source 1618 can include sources such as a video capture device, e.g., a camera, a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources.
[0138] The video encoder 1620 can encode the captured, pre-captured, or computer-generated video data. The input image can be received by the video encoder 1620 and stored in the input frame memory 1621. The general-purpose processor 1623 can load information therefrom and perform the encoding. The program for driving the general-purpose processor can be loaded from a storage device such as Figure 16 the example storage module depicted in. The general-purpose processor can use the processing memory 1622 to perform the encoding, and the output of the encoded information of the general-purpose processor can be stored in a buffer such as the output buffer 1626.
[0139] The video encoder 1620 can include a resampling module 1625 that can be configured to compile (e.g., encode) the video data in a scalable video compilation scheme that defines at least one base layer and at least one enhancement layer. As part of the encoding process, the resampling module 1625 can resample at least some of the video data, where the resampling can be performed in an adaptive manner using a resampling filter.
[0140] The encoded video data, e.g., the compiled bitstream, can be sent directly to the destination device 1616 via the output interface 1628 of the source device 1612. In Figure 16In the example, the destination device 1616 includes an input interface 1638, a video decoder 1630, and a display device 1632. In some cases, the input interface 1628 may include a receiver and / or a modem. The input interface 1638 of the destination device 1616 receives the encoded video data via a channel 1616. The encoded video data may include various syntax elements representing the video data generated by the video encoder 1620. Such syntax elements may be included together with the encoded video data transmitted on a communication medium, stored on a storage medium, or stored in a file server.
[0141] The encoded video data may also be stored on a storage medium or a file server for later access by the destination device 1616 for decoding and / or playback. For example, the compiled bitstream may be temporarily stored in an input buffer 1631 and then loaded into a general-purpose processor 1633. A program for driving the general-purpose processor may be loaded from a storage device or a memory. The general-purpose processor may use a processing memory 1632 to perform the decoding. The video decoder 1630 may also include a resampling module 1635 similar to the resampling module 1625 employed in the video encoder 1620.
[0142] Figure 16 The resampling module 1635 is depicted as separate from the general-purpose processor 1633, but those skilled in the art will understand that the resampling function may be performed by a program executed by the general-purpose processor, and the processing in the video encoder may be completed using one or more processors. The decoded image may be stored in an output frame buffer 1636 and then sent to the input interface 1638.
[0143] The display device 1638 may be integrated with the destination device 1616 or may be external to it. In some examples, the destination device 1616 may include an integrated display device and may also be configured to interface with an external display device. In other examples, the destination device 1616 may be a display device. Generally, the display device 1638 displays the decoded video data to the user.
[0144] Video encoders 1620 and video decoders 1630 may operate in accordance with video compression standards. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) are studying the potential need to standardize future video compression technologies with significantly greater compression capabilities than the current High Efficiency Video Coding (HEVC) standard, including its current extensions and recent extensions for screen content coding and high dynamic range coding. The groups are jointly collaborating on this exploration activity, which is called the Joint Video Exploration Team (JVET), to evaluate compression technology designs proposed by their experts in the field. The latest results developed by JVET are described in the "Algorithm Description of Joint Exploration Test Model 5 (JEM 5)" of JVET-E1001-V2, written by J. Chen, E. Alshina, G. Sullivan, J. Ohm, and J. Boyce.
[0145] Additionally or alternatively, video encoders 1620 and video decoders 1630 may operate in accordance with other proprietary or industry standards that operate with the disclosed JVET features. Thus, other standards such as the ITU-T H.264 standard, alternatively known as MPEG-4, Part 10, Advanced Video Coding (AVC), or extensions of these standards. Thus, although new developments are made for JVET, the techniques of the present disclosure are not limited to any particular coding standard or technology. Other examples of video compression standards and technologies include MPEG-2, ITU-T H.263, and proprietary or open-source compression formats and related formats.
[0146] Video encoders 1620 and video decoders 1630 may be implemented in hardware, software, firmware, or any combination thereof. For example, video encoders 1620 and decoders 1630 may employ one or more processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. When video encoders 1620 and decoders 1630 are implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and may use one or more processors to execute the instructions in hardware to perform the techniques of the present disclosure. Each of video encoders 1620 and video decoders 1630 may be included in one or more encoders or decoders, any one of which may be integrated as part of a combined encoder / decoder (CODEC) in their respective devices.
[0147] Aspects of the subject matter described herein may be described in the general context of computer-executable instructions, such as program modules, being executed by computers such as the above-described general-purpose processors 1623 and 1633. Generally, program modules include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement particular abstract data types. Aspects of the subject matter described herein may also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including memory storage devices.
[0148] Examples of memory include random access memory (RAM), read-only memory (ROM), or both. The memory may store instructions for performing the above-described techniques, such as source code or binary code. The memory may also be used to store variables or other intermediate information during the execution of instructions to be executed by processors such as processors 1623 and 1633.
[0149] The storage device may also store instructions for performing the above-described techniques, such as source code or binary code. The storage device may additionally store data used and manipulated by a computer processor. For example, the storage device in video encoder 1620 or video decoder 1630 may be a database accessed by computer system 1623 or 1633. Other examples of storage devices include random access memory (RAM), read-only memory (ROM), hard disk drives, magnetic disks, optical disks, CD-ROMs, DVDs, flash memory, USB memory cards, or any other medium readable by a computer.
[0150] The memory or storage device may be an example of a non-transitory computer-readable storage medium used by or in conjunction with a video encoder and / or decoder. The non-transitory computer-readable storage medium contains instructions for controlling a computer system configured to perform the functions described by a particular embodiment. When executed by one or more computer processors, the instructions may be configured to perform the functions described in a particular embodiment.
[0151] In addition, it should be noted that some embodiments have been described as processes that may be depicted as flowcharts or block diagrams. Although each may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. Additionally, the order of the operations may be rearranged. A process may have other steps not included in the figures.
[0152] Certain embodiments can be implemented on a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium contains instructions for controlling a computer system to execute the methods described by certain embodiments. The computer system may include one or more computing devices. When executed by one or more computer processors, the instructions may be configured to perform the methods described in certain embodiments.
[0153] As used in this specification and in the claims that follow throughout this application, unless the context clearly dictates otherwise, the words "a," "an," and "the" include plural references. Also, as used in this specification and the claims that follow, unless the context clearly dictates otherwise, the meaning of "in" includes both "in" and "on."
[0154] Although the exemplary embodiments of the present invention have been described in detail in terms of the above-described structural features and / or method acts, it is to be understood that those skilled in the art will readily appreciate that many additional modifications can be made in the exemplary embodiments without materially departing from the novel teachings and advantages of the present invention. Furthermore, it is to be understood that the subject matter defined in the appended claims need not be limited to the specific features or acts described above. Accordingly, all such modifications are intended to be included within the scope of the present invention as interpreted in light of the breadth and scope of the appended claims.
Claims
1. A method for decoding video data from a bitstream, the method comprising: (a) receiving a bitstream indicating how to split a coded tree unit into coded units, wherein a plurality of the coded units are square, and wherein one of the coded units is split into four square coded units; (b) determining a first set of possible modes selectable based on a most probable mode (MPM) index for a current block of the video data, wherein one of the first set of MPMs selectable based on the MPM index includes a direct horizontal mode and another of the first set of MPMs selectable based on the MPM index includes a direct vertical mode and another of the first set of MPMs selectable based on the MPM index includes an angular mode, wherein the first set of MPMs includes only five different modes; (c) deriving from the bitstream (i) an MPM flag and (ii) another index, the MPM flag comprising a total of 1 bit, wherein at least one of the MPM flag and the another index indicates whether an intra mode for predicting the current block is one of the first set of MPMs; (d) when at least one of the MPM flag and the another index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index, selecting the intra mode of the current block based on the MPM index decoded from the bitstream for one of the first set of MPMs; (e) when at least one of the MPM flag and the another index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determining a second set of at least one mode and (ii) determining a third set of at least one mode; (f) wherein the first set, the second set, and the third set include different modes, and wherein a combination of the first set, the second set, and the third set includes 67 different modes; (g) determining the intra mode of the current block for the second set of at least one mode based on a first combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes; and (h) determining the intra mode of the current block for the third set of at least one mode based on a second combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes.
2. A computer-readable storage medium storing instructions that, when executed, cause a processor to perform: (a) receiving a bitstream indicating how to split a coded tree unit into coded units, wherein a plurality of the coded units are square, and wherein one of the coded units is split into four square coded units; (b) Determine a first set of possible modes selectable based on a probable mode MPM index for a current block of video data, wherein one of the first set of MPMs selectable based on the MPM index includes a direct horizontal mode and another of the first set of MPMs selectable based on the MPM index includes a direct vertical mode and another of the first set of MPMs selectable based on the MPM index includes an angular mode, wherein the first set of MPMs includes only five different modes; (c) Derive from a bitstream (i) an MPM flag and (ii) another index, the MPM flag including a total of 1 bit, wherein at least one of the MPM flag and the another index indicates whether an intra mode for predicting the current block is one of the first set of MPMs; (d) When at least one of the MPM flag and the another index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index, select the intra mode of the current block based on the MPM index decoded from the bitstream for one of the first set of MPMs; (e) When at least one of the MPM flag and the another index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determine a second set of at least one mode and (ii) determine a third set of at least one mode; (f) wherein the first set, the second set, and the third set include different modes, and wherein a combination of the first set, the second set, and the third set includes 67 different modes; (g) Determine the intra mode of the current block for the second set of at least one mode based on a first combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes; And (h) Determine the intra mode of the current block for the third set of at least one mode based on a second combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes.
3. A method for encoding video data by an encoder, the method comprising: (a) Provide a bitstream indicating how to split a coding tree unit into coding units, wherein a plurality of the coding units are square, and wherein one of the coding units is split into four square coding units; (b) The bitstream contains data suitable for determining a first set of possible modes selectable based on a possible mode MPM index for the current block of the video data, wherein one of the first set of MPMs selectable based on the MPM index includes a direct horizontal mode and another of the first set of MPMs selectable based on the MPM index includes a direct vertical mode and another of the first set of MPMs selectable based on the MPM index includes an angular mode, wherein the first set of MPMs includes only five different modes; (c) The bitstream contains data suitable for deriving (i) an MPM flag and (ii) another index from the bitstream, the MPM flag comprising a total of 1 bit, and at least one of the MPM flag and the another index indicating whether the intra mode for predicting the current block is one of the first set of MPMs; (d) The bitstream contains data suitable for selecting the intra mode of the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs when at least one of the MPM flag and the another index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index; (e) The bitstream contains data suitable for, when at least one of the MPM flag and the another index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, the MPM flag and the another index (i) determining a second set of at least one mode and (ii) determining a third set of at least one mode; (f) The bitstream contains data suitable for wherein the first set, the second set and the third set include different modes, and the combination of the first set, the second set and the third set includes 67 different modes; (g) The bitstream contains data suitable for determining the intra mode of the current block for the second set of at least one mode based on a first combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes; And (h) The bitstream contains data suitable for determining the intra mode of the current block for the third set of at least one mode based on a second combination of the MPM flag and the another index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of possible modes.
Citation Information
Patent Citations
Indicating intra-prediction mode selection for video coding using CABAC
CN103299628A
Residual coding for depth intra prediction modes
CN105580361A