Intra-mode JVET coding
The JVET intra-mode coding method optimizes coding efficiency by encoding subsets of intra-predictive modes using truncated unary binarization and fixed-length codes, addressing inefficiencies in the current JVET scheme.
Patent Information
- Application Number
- JP2025002715
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-24
- Filing Date
- 2025-01-08
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2038-07-24
AI Technical Summary
The current JVET intra-mode coding scheme in video coding standards faces inefficiencies in coding burden and bandwidth due to the use of truncated unary binarization for the last two modes of the MPM list and higher coding complexity for the first three bins of the MPM mode, which does not offer advantages in coding performance.
A video coding method and system that identifies and encodes subsets of unique intra-predictive coding modes using truncated unary binarization, including a subset of unique MPM modes, a subset of selected modes, and a subset of non-selected modes, with the latter two using 4-bit fixed-length codes.
Reduces coding burden and bandwidth by optimizing the encoding of intra-mode coding subsets, improving coding efficiency and performance.
Smart Images

Figure 0007823236000001 
Figure 0007823236000002 
Figure 0007823236000003
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to the field of video coding, and more particularly, to efficient intra-mode coding. [Background technology]
[0002] Technical improvements in evolving video coding standards show a trend toward increasing coding efficiency, enabling higher bitrates, higher resolutions, and better video quality. The Joint Video Exploration Team (JVET) is developing a new video coding scheme. Similar to other video coding schemes such as High Efficiency Video Coding (HEVC), JVET is a block-based hybrid spatial and temporal predictive coding scheme. However, compared to HEVC, JVET includes many changes to the bitstream structure, syntax, constraints, and mappings for generating decoded multiple images. JVET is implemented in the Joint Exploration Model (JEM) encoder and decoder.
[0003] The current JVET standard describes a total of 67 intra prediction modes, including a planar mode, a DC mode, and 65 directional angular intra modes. To efficiently code these 67 modes, all intra modes are subdivided into three sets: a set of six most probable modes (MPMs), a set of 16 selected modes, and a set of 45 non-selected modes.
[0004] The six MPMs are derived from the available neighboring block modes, derived intra modes, and default intra modes. The intra modes of the five neighboring blocks of the current block are shown in Figure 1a. These are left (L), top (A), bottom-left (BL), top-right (AR), and top-left (AL), and are used to form the MPM list for the current block. The initial MPM list is created by inserting the five neighboring intra modes, the planar mode, and the DC mode into the MPM list. A pruning process is used to remove duplicate modes, allowing only unique modes to be included in the MPM list. The order in which multiple initial modes are included is left, top, planar, DC, bottom-left, top-right, and top-left.
[0005] If the MPM list is not full, derived modes are added; these intra modes are derived by adding "-1" or "+1" to the angular modes already included in the MPM list. If the MPM list is not yet complete, multiple default modes are added in the following order: vertical, horizontal, mode 2, and diagonal modes. This process results in a unique list of six MPM modes.
[0006] For entropy coding of the six MPMs, the truncated unary binarization shown in Figure 1b is currently used. The first three bins of the MPM modes are coded with multiple contexts depending on the MPM mode associated with the currently signaled bin. MPM modes are classified into one of three categories: (a) predominantly horizontal (i.e., MPM mode numbers are smaller than diagonal mode numbers), (b) predominantly vertical (i.e., MPM mode numbers are larger than diagonal mode numbers), and (c) non-angular (DC and planar) classes. Thus, three contexts are used to signal the MPM index based on this classification.
[0007] Coding for selecting the remaining 61 non-MPMs is performed as follows: The 61 non-MPMs are first divided into two sets: a selected mode set and a non-selected mode set. The selected mode set contains 16 modes, and the remainder (45 modes) are assigned to the non-selected mode set. The mode set to which the current mode belongs is indicated by a flag in the bitstream. If the indicated mode is in the selected mode set, the selected mode is signaled with a 4-bit fixed-length code; if the indicated mode is from the non-selected mode set, the selected mode is signaled with a truncated binary code. As an example, the selected mode set is generated by subsampling the 61 non-MPM modes as follows:
[0008] SelectionModeSet={0,4,8,12,16,20…60} Non-selected mode set = {1, 2, 3, 5, 6, 7, 9, 10...59} The current JVET intra-mode coding is summarized in Figure 1b below.
[0009] As shown in Figure 1b, the last two entries of the MPM list require six bins, which is the same number of bins allocated to the 16 selection modes. This configuration does not offer any advantage in terms of coding performance for the last two modes of the MPM list. Also, because the first three bins of the MPM mode are coded using context-based entropy coding, the coding complexity of the six bins of the MPM mode is higher than that of the six bins of the selection mode.
[0010] What is needed is a system and method that reduces the coding burden and bandwidth associated with intra-mode coding. Summary of the Invention
[0011] This disclosure provides a video coding method for JVET intra prediction, the video coding method including: defining a set of unique intra-predictive coding modes, which in some embodiments may be 67 modes; identifying and instantiating in memory a subset of unique MPM intra-predictive coding modes from the set of unique intra-predictive coding modes, which in some embodiments may be no more than five of seven or more. The method also provides identifying and instantiating in memory a subset of selected unique intra-predictive coding modes, which in some embodiments may include 16 coding modes, from the set of unique intra-predictive coding modes other than the subset of unique MPM intra-predictive coding modes; and identifying and instantiating in memory a subset of non-selected unique intra-predictive coding modes, which are other than the subset of unique MPM intra-predictive coding modes and constitute a balance of the intra-prediction modes, from the set of unique intra-predictive coding modes other than the subset of unique MPM intra-predictive coding modes and other than the selected subset of unique intra-predictive coding modes. The subset of unique MPM intra-predictive coding modes is then coded using truncated unary binarization.
[0012] This disclosure also provides a video coding system for JVET intra prediction, which in some embodiments comprises instantiating in memory a set of 67 unique intra-predictive coding modes; instantiating in memory a subset of unique MPM intra-predictive coding modes from the set of unique intra-predictive coding modes; instantiating in memory a subset of 16 unique selected intra-predictive coding modes from the set of unique intra-predictive coding modes other than the subset of unique MPM intra-predictive coding modes; instantiating in memory a subset of unselected unique intra-predictive coding modes from the set of unique intra-predictive coding modes other than the subset of unique MPM intra-predictive coding modes and other than the subset of unique selected intra-predictive coding modes; encoding the subset of unique MPM intra-predictive coding modes using truncated unary binarization; and encoding the subset of 16 selected unique intra-predictive coding modes using a 4-bit fixed length code. [Brief explanation of the drawings]
[0013] Further details of the invention are explained using the accompanying drawings. [Figure 1a] Indicates neighboring blocks relative to the current coding block. [Figure 1b] 1 shows a table of the current JVET coding for intra-mode prediction. [Figure 1c] This shows the division of a frame into multiple coding tree units (CTUs). [Figure 2] 1 illustrates an exemplary partitioning of a CTU into multiple Coding Units (CUs) using quadtree partitioning and symmetric bisection. [Figure 3]The QTBT (quadtree plus binary tree) representation of the partition in Fig. 2 is shown. [Figure 4] We present four possible types of asymmetric bisection of a CU into two smaller CUs. [Figure 5] 10 illustrates an exemplary partitioning of a CTU into multiple CUs using quadtree partitioning, symmetric bisection, and asymmetric bisection. [Figure 6] Figure 5 shows the QTBT representation of the partition. [Figure 7] 1 shows a simplified block diagram of CU coding in the JVET encoder. [Figure 8] 67 possible intra prediction modes for the luma component of JVET are shown. [Figure 9] 1 shows a simplified block diagram of CU coding in the JVET encoder. [Figure 10] 1 illustrates an embodiment of a method for CU coding in a JVET encoder. [Figure 11] 1 shows a simplified block diagram of CU coding in the JVET encoder. [Figure 12] 1 shows a simplified block diagram of CU decoding in the JVET decoder. [Figure 13] 10 shows an alternative simplified block diagram of JVET coding for intra-mode prediction. [Figure 14] 10 shows a table of alternative JVET coding for intra-mode prediction. [Figure 15] 1 illustrates an embodiment of a computer system adapted and / or configured to process a method of CU coding. [Figure 16] 1 illustrates an embodiment of an encoding / decoding system for CU encoding / decoding in a JVET encoder / decoder. DETAILED DESCRIPTION OF THE INVENTION
[0014] FIG. 1 illustrates the division of a frame into coding tree units (CTUs) 100. A frame may be an image of a video sequence. A frame may include a matrix or set of matrices with pixel values representing intensity measurements within the image. A video sequence may thus be generated by these sets of matrices. The pixel values may be defined to represent color and brightness in full-color video coding, where pixels are divided into three channels. For example, in the YCbCr color space, pixels have a luminance value Y representing the intensity of the image's gray levels and two chrominance values Cb and Cr representing the color variations from gray to blue and red. In other embodiments, the pixel values may be represented by values from different color spaces or models. The resolution of the video determines the number of pixels in a frame. Higher resolutions result in more pixels and sharper images, but also require higher bandwidth, storage, and transmission requirements.
[0015] Frames of a video sequence can be encoded and decoded using JVET. JVET is a video coding scheme being developed by the Joint Video Exploration Team. Versions of JVET have been implemented in the Joint Exploration Model (JEM) decoder and decoder. Similar to other video coding schemes such as High Efficiency Video Coding (HEVC), JVET is a block-based hybrid spatial and temporal predictive coding scheme. In coding with JVET, a frame is first divided into square blocks called CTUs 100, as shown in FIG. 1. For example, the CTUs 100 may be blocks of 128x128 pixels.
[0016] 2 shows an example division of a CTU 100 into multiple CUs 102. Each CTU 100 in a frame may be divided into one or more Coding Units (CUs) 102. One or more CUs 102 may be used for prediction and transform, as described below. Unlike HEVC, in JVET, the CUs 102 may be rectangular or square and may be coded without further division into prediction units or transform units. The CUs 102 may be the same size as their root CTU 100 or may be smaller subdivisions of the root CTU 100, as small as a 4x4 block.
[0017] In JVET, a CTU 100 can be partitioned into multiple CUs 102 according to a QTBT (quadtree plus binary tree) scheme. In this scheme, the CTU 100 is recursively partitioned into multiple square blocks according to a quadtree, and these square blocks can be recursively partitioned horizontally or vertically according to a binary tree. Several parameters can be set to control the partitioning according to QTBT, such as the CTU size, the minimum size of leaf nodes in the quadtree and binary tree, the maximum size of leaf nodes in the binary tree, and the maximum depth of the binary tree.
[0018] In some embodiments, JVET can restrict the binary partitioning of the binary tree portion of QTBT to a symmetric partition, where blocks are split in half either vertically or horizontally along the midline.
[0019] 2 illustrates a CTU 100 partitioned into multiple CUs 102, with solid lines indicating a quadtree partition and dashed lines indicating a symmetric binary tree partition. As illustrated, the binary partitioning allows for symmetric horizontal and vertical partitioning to define the structure of the CTU and its subdivision into multiple CUs.
[0020] FIG. 3 shows a QTBT representation of the partition of FIG. 2. The quadtree root node represents CTU 100, with each child node of the quadtree section representing one of the four square blocks partitioned from the parent square block. The square blocks, represented by the quadtree leaf nodes, are partitioned symmetrically zero or more times using a binary tree, and the quadtree leaf nodes are root nodes of the binary tree. At each level of the binary tree section, the block may be partitioned symmetrically vertically or horizontally. A flag set to "0" indicates that the block is partitioned symmetrically horizontally, and a flag set to "1" indicates that the block is partitioned symmetrically vertically.
[0021] In other embodiments, the JVET may enable either symmetric or asymmetric bipartitioning in the binary tree portion of the QTBT. Asymmetric motion partitioning (AMP) is possible in a different context in HEVC when partitioning multiple prediction units (PUs). However, when partitioning multiple CUs 102 in the JVET according to the QTBT structure, asymmetric bipartitioning can result in improved partitioning relative to symmetric bipartitioning when multiple correlated areas of a CU 102 are not located on either side of a midline through the center of the CU 102. As a non-limiting example, if a CU 102 represents one object close to the center of the CU and another object on the side of the CU 102, the CU 102 may be asymmetrically partitioned to place each object into separate smaller CUs 102 of different sizes.
[0022] 4 illustrates four possible types of asymmetric bisections, in which a CU 102 is divided into two smaller CUs 102 along a line across the length or height of the CU 102, where one of the smaller CUs 102 is 25% the size of the parent CU 102 and the other is 75% the size of the parent CU 102. The four types of asymmetric bisections illustrated in FIG. 4 allow the CU 102 to be divided along a line 25% away from the left side of the CU 102, 25% away from the right side of the CU 102, 25% away from the top of the CU 102, or 25% away from the bottom of the CU 102. In another embodiment, the asymmetric bisection line along which the CU 102 is divided may be located in any other position such that the CU 102 is not symmetrically divided into halves.
[0023] 5 shows a non-limiting example of a CTU 100 that has been split into multiple CUs 102 using a scheme that allows for both symmetric and asymmetric bisection in the binary tree portion of the QTBT. In FIG. 5, the dashed lines indicate asymmetric bisection lines, and the parent CU 102 has been split using one of the multiple split types shown in FIG. 4.
[0024] Figure 6 shows a QTBT representation of the partitioning of Figure 5. In Figure 6, two solid lines extending from a node indicate a symmetric partitioning in the binary tree portion of the QTBT, and two dashed lines extending from a node indicate an asymmetric partitioning in the binary tree portion.
[0025] Syntax indicating how the CTU 100 is partitioned into multiple CUs 102 may be coded into the bitstream. As a non-limiting example, syntax may be coded into the bitstream to indicate which nodes were partitioned using quadtree partitioning, symmetric bisection, and asymmetric bisection. Similarly, syntax may be coded into the bitstream for nodes partitioned using asymmetric bisection to indicate which type of asymmetric bisection was used, such as one of the four types shown in FIG. 4.
[0026] In some embodiments, the use of asymmetric splitting can be limited to splitting CUs 102 at leaf nodes of a QTBT quadtree portion. In these embodiments, the CUs 102 of child nodes split from a parent node in a quadtree portion using quadtree splitting can be final CUs 102 or can be further split using quadtree splitting, symmetric bisection, or asymmetric bisection. The child nodes of a bisection split using symmetric bisection can be final CUs 102 or can be further split recursively one or more times using only symmetric bisection. The child nodes of a bisection split from a QT leaf node using asymmetric bisection can be final CUs 102 that are not further split.
[0027] In these embodiments, the use of asymmetric splitting can be limited to splitting quadtree leaf nodes, thereby reducing search complexity and / or limiting overhead bits. Because only quadtree leaf nodes are split in an asymmetric split, the use of asymmetric splitting can directly indicate the end of a branch in the QT portion without other syntax or further signaling.
[0028] Similarly, because asymmetrically split nodes cannot be split further, the use of asymmetric split on a node can also directly indicate that its asymmetrically split child nodes are final CUs 102 without other syntax or further signaling.
[0029] In alternative embodiments, such as when limiting search complexity and / or limiting the number of additional bits is less of an issue, asymmetric splitting may be used to split multiple nodes generated by quadtree splitting, symmetric bisection, and / or asymmetric bisection.
[0030] After quadtree and binary tree partitioning using any of the QTBT structures described above, the blocks represented by the leaf nodes of the QTBT represent the final CUs 102 to be coded, such as for coding using inter prediction or intra prediction. For slices or full frames coded using inter prediction, different partitioning structures can be used for the luma and chroma components. For example, for an inter slice, CU 102 can have coding blocks (CBs) for different color components, such as one luma CB and two chroma CBs. For slices or full frames coded using intra prediction, the partitioning structure for the luma and chroma components is the same.
[0031] In an alternative embodiment, the JVET may use a two-level coding block structure as an alternative or extension to the QTBT partitioning described above. In a two-level coding block structure, the CTU 100 may first be partitioned at a higher level into multiple base units (BUs). The multiple BUs may then be partitioned at a lower level into multiple operating units (OUs).
[0032] In embodiments employing a two-level coding block structure, at a high level, the CTU 100 may be partitioned into BUs according to one of the QTBT structures described above or according to a quadtree (QT) structure such as that used in HEVC. A block may only be partitioned into four equal-sized sub-blocks. As a non-limiting example, the CTU 102 may be partitioned into BUs according to the QTBT structure described above with respect to Figures 5-6. The leaf nodes of the quadtree portion may be partitioned using quadtree partitioning, symmetric binary tree partitioning, or asymmetric binary tree partitioning.
[0033] In this example, the final leaf nodes of the QTBT can be multiple BUs instead of multiple CUs. At the lower level of the two-level coding block structure, each BU split from CTU 100 may be further split into one or more OUs. In some embodiments, if a BU is square, it may be split into multiple OUs using quadtree splitting or bisection, such as symmetric or asymmetric bisection. However, if a BU is not square, it may be split into multiple OUs using only bisection. Limiting the types of splitting available for non-square BUs may limit the number of bits used to indicate the type of splitting used to generate the multiple BUs.
[0034] Although the following description describes coding of CU 102, in embodiments using a two-level coding block structure, BUs and OUs can be coded instead of CU 102. As a non-limiting example, BUs may be used for higher-level coding operations, such as intra-prediction or inter-prediction, while smaller OUs may be used for lower-level coding operations, such as transforming or generating transform coefficients. Thus, the syntax coded for BUs may indicate whether they are coded with intra-prediction or inter-prediction, or may indicate information identifying a specific intra-prediction mode or motion vector used to code the BUs. Similarly, the syntax for OUs may identify a specific transform operation or quantized transform coefficients used to code the OUs.
[0035] 7 shows a simplified block diagram of CU coding in the JVET encoder. The main stages of video coding, as described above, are partitioning to identify multiple CUs 102, then encoding the multiple CUs 102 using prediction at 704 or 706 to generate residual CUs 710, which are transformed at 712, quantized at 716, and entropy coded at 720.
[0036] The encoder and encoding process shown in FIG. 7 also include a decoding process, which will be described in more detail below. Given a current CU 102, the encoder may obtain a predicted CU 702 using either intra prediction spatially at 704 or inter prediction temporally at 706. The basic idea of predictive coding is to transmit a difference or residual signal between an original signal and a prediction of the original signal. At the receiver side, the original signal can be reconstructed by adding the residual and the prediction, as described below. Because the difference signal is less correlated than the original signal, fewer bits are required for transmission.
[0037] A slice coded entirely by intra-predicted CU 102, such as an entire image or a portion of an image, may be an "I" slice, decoded without reference to other slices, and may represent a potential starting point for decoding. Slices coded with at least some inter-predicted CUs may be predicted (P) slices or bi-predictive (B) slices, which may be decoded based on one or more reference images. P slices may use intra-prediction and inter-prediction with previously coded slices. For example, P slices can be more compressed than "I" slices using inter-prediction, but require coding of previously coded slices to encode them. B slices may use intra- or inter-prediction with interpolated prediction from two different frames to use data from previous and / or subsequent slices, thereby improving the accuracy of the motion estimation process. In some cases, P slices and B slices may be coded, or alternatively coded, using intra-block copying, in which data from other portions of the same slice is used.
[0038] As described below, intra prediction or inter prediction may be performed based on multiple CUs 734 reconstructed from multiple previously coded CUs 102, such as multiple neighboring CUs 102 or multiple CUs 102 in a reference image.
[0039] When CU 102 is spatially coded using intra prediction at 704, an intra prediction mode that best predicts pixels of CU 102 based on samples from neighboring CUs 102 in the image may be identified.
[0040] When coding the luma component of a CU, the encoder can generate a list of candidate intra-prediction modes. While HEVC had 35 possible intra-prediction modes for the luma component, JVET has 67 possible intra-prediction modes for the luma component. These include a planar mode, which uses a three-dimensional plane of values generated from neighboring pixels; a DC mode, which uses values averaged from neighboring pixels; and the 65 directional modes shown in Figure 8, which use values copied from neighboring pixels along indicated directions.
[0041] When generating a list of candidate intra-prediction modes for the luma component of a CU, the number of candidate modes on the list depends on the size of the CU. The candidate list may include a subset of the 35 modes of HEVC with the lowest Sum of Absolute Transform Difference (SATD) cost, new directional modes added to the JVET adjacent to the candidates identified from the HEVC modes, and modes from a set of six most probable modes (MPMs) for CU 102 identified based on a list of intra-prediction modes used for neighboring previously coded blocks and default modes.
[0042] When coding the chroma components of a CU, a list of candidate intra-prediction modes can be generated. The list of candidate modes includes modes generated by cross-component linear model projection from luma samples, intra-prediction modes identified in the luma CB at specific array positions within the chroma block, and chroma prediction modes previously identified in neighboring blocks. The encoder identifies the candidate modes on the list with the lowest rate distortion cost and uses those intra-prediction modes when coding the luma and chroma components of the CU. Syntax can be coded into the bitstream indicating the intra-prediction modes used to code each CU 102.
[0043] After the optimal intra-prediction mode for CU 102 is selected, the encoder may use that mode to generate predicted CU 402. If the selected mode is a directional mode, a 4-tap filter may be used to improve the accuracy of the directionality. The top or left columns or rows of the prediction block may be adjusted with a boundary prediction filter, such as a 2-tap or 3-tap filter.
[0044] The predicted CU702 may be further smoothed by a position dependent intra prediction combination (PDPC) process, which adjusts the predicted CU702 generated based on filtered samples of neighboring blocks using unfiltered samples of neighboring blocks, or by adaptive reference sample smoothing, which uses a 3-tap or 5-tap low-pass filter to process multiple reference samples.
[0045] Once CU 102 is temporally coded using inter prediction at 706, a set of motion vectors (MVs) can be identified that point to samples in reference pictures that best predict the pixels of CU 102. Inter prediction exploits temporal redundancy between slices by representing the displacement of blocks of pixels within a slice. The displacement is determined according to the values of pixels in a previous or subsequent slice through a process called motion compensation. The motion vectors and associated reference indices indicating the pixel displacement relative to a particular reference picture can be provided in a bitstream to a decoder along with the residual between the original and motion-compensated pixels. The decoder can reconstruct the blocks of pixels in the reconstructed slice using the signaled motion vectors and reference indices of the residual.
[0046] In JVET, motion vectors are stored with 1 / 16 pel precision, and the difference between the motion vector and the predicted motion vector of a CU can be coded with 1 / 4 pel resolution or integer pel resolution.
[0047] In JVET, multiple motion vectors may be determined for multiple sub-CUs within CU 102 using techniques such as advanced temporal motion vector prediction (ATMP), spatial-temporal motion vector prediction (STMP), affine motion compensation prediction, pattern matched motion vector derivation (PMMVD), and / or bi-directional optical flow (BIO).
[0048] The encoder may use ATMVP to determine a time vector for CU 102 that points to a corresponding block in a reference image. The time vector may be determined based on multiple reference images and multiple motion vectors determined for previously coded neighboring CUs 102. The reference block indicated by the time vector for the entire CU 102 may be used to determine a motion vector for each sub-CU within CU 102.
[0049] STMVP may determine a motion vector for a sub-CU by scaling and averaging, along with a temporal vector, multiple motion vectors determined for neighboring blocks previously coded with inter prediction.
[0050] Affine motion compensation prediction may be used to predict a field of motion vectors for each sub-CU within a block based on two control motion vectors identified at the top corners of the block. For example, motion vectors for sub-CUs may be derived based on the top corner motion vectors identified at each 4x4 block within CU 102.
[0051] PMMVD can use bilateral matching or template matching to identify an initial motion vector for the current CU 102. Bilateral matching identifies the current CU 102 and reference blocks in two different reference images along the motion trajectory, while template matching can search for corresponding blocks in the current CU 102 and reference images identified by a template.
[0052] The initial motion vector determined for CU 102 may then be refined for each sub-CU individually. BIO may be used when performing inter prediction in bi-prediction based on previous and next reference images to determine the motion vector of a sub-CU based on the gradient of the difference between two reference images.
[0053] In some situations, local illumination compensation (LIC) can be used at the CU level to determine values for scaling factor and offset parameters based on samples neighboring the current CU 102 and corresponding samples neighboring the reference block identified by the candidate motion vector. In JVET, LIC parameters may be modified and signaled at the CU level. In some of the methods described above, the motion vectors determined for each sub-CU of the CU may be signaled to the decoder at the CU level. For other methods, such as PMMVD and BIO, motion information is not signaled in the bitstream to save overhead, and the decoder may derive motion vectors in the same process.
[0054] After the motion vectors for CU 102 are identified, the encoder may use those motion vectors to generate the predictive CU 702. In some cases, if multiple motion vectors are identified for an individual sub-CU, overlapped block motion compensation (OBMC) may be used when combining those motion vectors with motion vectors previously identified for one or more neighboring sub-CUs to generate the predictive CU 702.
[0055] Using bi-prediction, JVET may identify multiple motion vectors using decoder-side motion vector refinement (DMVR). DMVR may use a bilateral template matching process to identify a motion vector based on two motion vectors identified in bi-prediction. DMVR identifies a weighted combination of multiple predictive CUs 702 generated using each of the two motion vectors, and the two motion vectors may be refined by replacing them with a new motion vector that best represents the combined predictive CU 702.
[0056] The two refined motion vectors can be used to generate the final predicted CU 702 . As described above, once a predicted CU 702 has been identified in either intra prediction at 704 or inter prediction at 706 , the encoder may subtract the predicted CU 702 from the current CU 102 at 708 to identify a residual CU 710 .
[0057] The encoder may convert 712 the residual CU 710 into transform coefficients 714 that represent the residual CU 710 in the transform domain using one or more transform operations, such as a discrete cosine block transform (DCT) to transform data into the transform domain. JVET allows for more types of transform operations than HEVC, such as DCT-II, DST-VII, DST-VII, DCT-VIII, DST-I, and DCT-V operations. The possible transform operations are grouped into subsets, and an indication of which subsets and which specific operations within those subsets were used may be signaled by the encoder. In some cases, a large block size transform is used to zero out high-frequency transform coefficients in CUs 102 that are larger than a certain size, so that only low-frequency transform coefficients are maintained for these CUs 102.
[0058] In some cases, a mode dependent non-separable secondary transform (MDNSST) may be applied to the low frequency transform coefficients 714 after the forward core transform. The MDNSST operation may use a Hypercube-Givens Transform (HyGT) based on the rotated data. When used, an index value identifying the particular MDNSST operation may be signaled from the encoder.
[0059] At 716, the encoder may quantize the transform coefficients 714 into quantized transform coefficients 716. The quantization of each coefficient may be calculated by dividing the value of the coefficient by a quantization step derived from a quantization parameter (QP). In some embodiments, Qstep is 2 (QP-4) / 6 Quantization can aid in data compression because the high precision transform coefficients 714 can be converted into quantized transform coefficients 716, which have a finite number of possible values.
[0060] Quantization of transform coefficients may therefore limit the amount of bits generated and transmitted by the transform process. However, quantization is a lossy operation, and quantization losses cannot be recovered. However, the quantization process represents a trade-off between the quality of the reconstructed sequence and the amount of information required to represent the sequence. For example, a lower QP value may improve the quality of the decoded video, but may require a larger amount of data to represent and transmit. In contrast, a higher QP value may reduce the quality of the reconstructed video sequence, but require less data and bandwidth.
[0061] JVET can utilize a variance-based adaptive quantization technique that allows each CU 102 to use a different quantization parameter for its coding process (instead of using the same frame QP in coding each CU 102 of a frame). The variance-based adaptive quantization technique adaptively lowers the quantization parameter for certain blocks and increases it for other blocks. To select a specific QP for a CU 102, the variance of the CU is calculated. That is, if the variance of the CU is higher than the average variance of the frame, a higher QP than the QP of the frame may be set for the CU 102. If the CU 102 exhibits a lower variance than the average variance of the frame, a lower QP may be assigned.
[0062] At 720, the encoder may identify a number of final compression bits 722 by entropy coding the number of quantized transform coefficients 718. Entropy coding aims to remove statistical redundancy in the transmitted information. In JVET, the quantized transform coefficients 718 may be coded using CABAC (Context Adaptive Binary Arithmetic Coding), which uses a probability measure to remove statistical redundancy. For CUs 102 with non-zero quantized transform coefficients 718, the quantized transform coefficients 718 may be converted to binary. Each bit ("bin") of the binary representation may then be coded using a context model. The CU 102 is divided into three regions, each with its own set of context models to use for the pixels within that region.
[0063] Multiple scan passes may be performed to encode multiple bins. During the pass that encodes the first three bins (bin0, bin1, bin2), an index value indicating which context model to use for a bin may be identified by summing that bin's position over up to five neighboring quantized transform coefficients 718 that were previously coded as identified by the template.
[0064] The context model can be based on the probability that a bin's value is "0" or "1." As values are coded, the probabilities of the context model can be updated based on the actual number of "0" and "1" values. While HEVC used a fixed table to reinitialize the context model for each new image, in JVET, the probabilities of the context models for multiple new inter-predicted images can be initialized based on the context models generated for previously coded inter-predicted images.
[0065] The encoder may generate a bitstream that includes entropy coded bits 722 for multiple residual CUs 710, prediction information such as a selected intra-prediction mode or motion vectors, an indicator of how multiple CUs 102 were split from the CTU 100 according to the QTBT structure, and / or other information about the coded video. The bitstream may be decoded at a decoder, as described below.
[0066] In addition to using the quantized transform coefficients 718 to identify the final compressed bits 722, the encoder may also use the quantized transform coefficients 718 to generate reconstructed CUs 734 by following the same decoding process that the decoder uses to generate the reconstructed CUs 734. Thus, once the transform coefficients are calculated and quantized by the encoder, the quantized transform coefficients 718 may be sent to a decoding loop within the encoder. After quantizing the transform coefficients of the CUs, the decoding loop can cause the encoder to generate the same reconstructed CUs 734 that the decoder generates in the decoding process. Thus, when performing intra- or inter-prediction of a new CU 102, the encoder can use the same reconstructed CUs 734 that the decoder uses for neighboring CUs 102 or reference images. The reconstructed CUs 102, reconstructed slices, or a fully reconstructed frame may serve as references for further prediction stages.
[0067] To obtain pixel values of a reconstructed image, an inverse quantization process may be performed in the decoding loop of the encoder (for the same operation of the decoder, see below). To inverse quantize a frame, for example, the quantized value of each pixel of the frame is multiplied by a quantization step, such as the Qstep described above, to obtain reconstructed inverse quantized transform coefficients 726. For example, in the decoding process shown in FIG. 7 in the encoder, the quantized transform coefficients 718 of the residual CU 710 may be inverse quantized at 724 to obtain the inverse quantized transform coefficients 726. If an MDNSST operation was performed in encoding, the operation may be reversed after inverse quantization.
[0068] At 728, the inverse quantized transform coefficients 726 may be inverse transformed to identify a reconstructed residual CU 730, such as by applying a DCT to the values to obtain a reconstructed image. At 732, the reconstructed residual CU 730 may be added to the corresponding predicted CU 702 identified in the intra prediction at 704 or the inter prediction at 706 to identify a reconstructed CU 734.
[0069] At 736, one or more filters may be applied to the reconstructed data during the decoding process (at the encoder or, as described below, at the decoder), either at the picture level or CU level. For example, the encoder may apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). The encoder's decoding process may implement filters that estimate optimal filter parameters that can address potential artifacts in the reconstructed image and send them to the decoder. Such improvements improve the objective and subjective quality of the reconstructed video.
[0070] In deblocking filtering, pixels near sub-CU boundaries are modified, while in SAO, pixels within the CTU 100 can be modified using either edge offset or band offset classification. JVET's ALF can use circularly symmetric filters for each 2x2 block. An indication of the size and identity of the filter used for each 2x2 block can be signaled.
[0071] If the reconstructed images are reference images, then at 706 they may be stored in a reference buffer 738 for inter-prediction of future CUs 102 . In the above steps, JVET can use a content adaptive clipping operation to adjust color values to fit between upper and lower clipping bounds. The clipping bounds can be changed for each slice, and parameters identifying the bounds can be signaled in the bitstream.
[0072] 9 shows a simplified block diagram of CU coding in a JVET decoder. The JVET decoder may receive a bitstream containing information about coded CUs 102. The bitstream may indicate how the CUs 102 of an image were split from the CTUs 100 according to a QTBT structure. As a non-limiting example, the bitstream may identify how the CUs 102 were split from each CTU 100 within the QTBT using quadtree splitting, symmetric bisection, and / or asymmetric bisection. The bitstream may also indicate prediction information for the CUs 102, such as intra-prediction modes or motion vectors, and bits 902 representing entropy-coded residual CUs.
[0073] At 904, a decoder may decode the entropy-coded bits 902 using the CABAC context model signaled in the bitstream by the encoder. The decoder may use the parameters signaled by the encoder to update the probabilities of the context model in the same manner as they were updated during encoding.
[0074] After reversing the entropy coding at 904 to determine quantized transform coefficients 906, the decoder may dequantize them at 908 to determine dequantized transform coefficients 910. If an MDNSST operation was performed in the encoding, that operation may be reversed by the decoder after dequantization.
[0075] At 912, the dequantized transform coefficients 910 may be inverse transformed to identify a reconstructed residual CU 914. At 916, the reconstructed residual CU 914 may be added to a corresponding predicted CU 926 identified in the intra prediction at 922 or the inter prediction at 924 to identify a reconstructed CU 918.
[0076] At 920, one or more filters may be applied to the reconstructed data, either at the picture level or the CU level. For example, the decoder may apply a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF). As described above, an in-loop filter in the encoder's decoding loop may be used to estimate optimal filter parameters to improve the objective and subjective quality of the frame. At 920, these parameters are sent to the decoder to filter the reconstructed frame to match the filtered reconstructed frame in the encoder.
[0077] After generating the reconstructed image by applying the signaled filters identifying the reconstructed CUs 918, the decoder may output the reconstructed image as output video 928. If the reconstructed images are used as reference images, they may be stored 924 in a reference buffer 930 for inter-prediction of future CUs 102.
[0078] Figure 10 shows an embodiment of a method for CU coding in a JVET decoder 1000. In the embodiment shown in Figure 10, an encoded bitstream 902 may be received in step 1002, a CABAC context model associated with the encoded bitstream 902 may be determined in step 1004, and then the encoded bitstream 902 may be decoded in step 1006 using the determined CABAC context model.
[0079] Step 1008 may determine a plurality of quantized transform coefficients 906 associated with the encoded bitstream 902 , and step 1010 may determine inverse quantized transform coefficients 910 from the plurality of quantized transform coefficients 906 .
[0080] At step 1012, it may be determined whether an MDNSST operation was performed during encoding and / or whether the bitstream 902 includes an indication that an MDNSST operation was applied to the bitstream 902. If it is determined that an MDNSST operation was performed during the encoding process or that the bitstream 902 includes an indication that an MDNSST operation was applied to the bitstream 902, an inverse MDNSST operation 1014 may be performed on the bitstream 902 at step 1016 before the inverse transform operation 912 is performed. Alternatively, if an inverse MDNSST operation was not applied at step 1014, the inverse transform operation 912 may be performed on the bitstream 902 at step 1016. The inverse transform operation 912 at step 1016 may determine and / or construct a reconstructed residual CU 914.
[0081] In step 1018, the reconstructed residual CU 914 from step 1016 may be combined with a predicted CU 918. The predicted CU 918 may be one of the intra-predicted CU 922 determined in step 1020 and the inter-prediction unit 924 determined in step 1022.
[0082] In step 1024, any one or more filters 920 may be applied to the reconstructed CU 914 and output in step 1026. In some embodiments, filters 920 may not be applied in step 1024.
[0083] In some embodiments, in step 1028, the reconstructed CU 918 may be stored in a reference buffer 930. 11 shows a simplified block diagram 1100 of CU coding in a JVET encoder. In step 1102, a JVET coding tree unit may be represented as the root node of a quadtree plus binary tree (QTBT) structure. In some embodiments, the QTBT may have a quadtree branching from the root node and / or a binary tree branching from one or more leaf nodes of the quadtree. The representation from step 1102 may proceed to steps 1104, 1106, or 1108.
[0084] In step 1104, an asymmetric bisection may be used to divide the represented quadtree node into two unequal-sized blocks. In some embodiments, the divided blocks may be represented in a binary tree branching from the quadtree node as leaf nodes representing final coding units. In some embodiments, the binary tree branching from the quadtree node as leaf nodes indicates final coding units for which further division is not permitted. In some embodiments, the asymmetric bisection divides the coding unit into unequal-sized blocks, with a first block representing 25% of the quadtree node and a second block representing 75% of the quadtree node.
[0085] In step 1106, quadtree partitioning may be used to divide the represented quadtree node into four square blocks of equal size. In some embodiments, the divided blocks may be represented as quadtree nodes representing multiple final coding units, or as multiple child nodes that are further divided by quadtree partitioning, symmetric binary tree partitioning, or asymmetric binary tree partitioning.
[0086] In step 1108, a quadtree split may be used to split the represented quadtree node into two blocks of equal size. In some embodiments, the split blocks may be represented as quadtree nodes representing multiple final coding units, or as multiple child nodes that are split again by a quadtree split, a symmetric binary tree split, or an asymmetric binary tree split.
[0087] In step 1110, the multiple child nodes from step 1106 or step 1108 may be represented as multiple child nodes configured to be encoded. In some embodiments, the multiple child nodes may be represented by multiple leaf nodes of a binary tree in the JVET.
[0088] In step 1112, multiple coding units from step 1104 or 1110 may be encoded using JVET. 12 shows a simplified block diagram 1200 of CU decoding in a JVET decoder. In the embodiment shown in FIG. 12, a bitstream may be received in step 1202 indicating how a coding tree unit has been split into multiple coding units according to a QTBT structure. The bitstream may indicate how the quadtree nodes are split in at least one of a quadtree split, a symmetric bisection, or an asymmetric bisection.
[0089] In step 1204, multiple coding units represented by multiple leaf nodes of the QTBT structure may be identified. In some embodiments, the multiple coding units may indicate whether the node was split from the quadtree leaf node using an asymmetric bisection. In some embodiments, the coding unit may indicate that the node represents a final coding unit to be decoded.
[0090] In step 1206, the identified one or more coding units may be decoded using JVET. FIG. 13 shows an alternative simplified block diagram 1300 of JVET coding for intra-mode prediction. In the embodiment shown in FIG. 13, a set of MPMs may be identified and instantiated in memory in step 1302, a set of 16 selected modes may be identified and instantiated in memory in step 1304, and a balance of 67 modes may be defined and instantiated in memory in step 1304. In some embodiments, the set of MPMs may be reduced from a standard set of six MPMs. In some embodiments, the set of MPMs may include five unique modes, the selected modes may include 16 eigenmodes, and the set of unselected modes may include the remaining 46 unselected eigenmodes. However, in alternative embodiments, the set of MPMs may include fewer unique modes, the selected modes may remain fixed at 16 unique modes, and the set size of the unselected eigenmodes may be appropriately adjusted to accommodate a total of 67 modes. As a non-limiting example, in some embodiments where the set of MPMs includes five unique modes instead of six MPMs, a truncated unary binarization is used, and if a new binarization for the five MPMs is utilized, the number of bins assigned to the MPM modes may be equal to or less than five bins.
[0091] Thus, in some embodiments, 16 modes selected from the 62 remaining intra modes are generated by uniformly subsampling these 62 intra modes, each coded with a 4-bit fixed-length code. As a non-limiting example, assuming the remaining 62 modes are indexed as {0, 1, 2, ..., 61}, then the 16 selected modes = {0, 4, 8, 12, 16, 20, 24, 28, 32, 36, 40, 44, 48, 52, 56, 60}. The remaining 46 unselected modes = {1, 2, 3, 5, 6, 7, 9, 10 ... 59, 61}, and these 46 unselected modes may be coded with a truncated binary code.
[0092] Figure 14 shows an alternative JVET coding table 1400 for intra-mode prediction according to Figure 13. In the embodiment shown in Figure 14, a plurality of intra-prediction modes 1402 is shown to include 5 MPMs, 16 selected modes, and 46 non-selected modes, where a plurality of bin strings 1404 for the MPMs may be coded using truncated unary binarization, the 16 selected modes may be coded using 4 bits of a fixed-length code, and the 46 non-selected modes may be coded using truncated binary coding.
[0093] In the alternative embodiment of Figure 13, six MPMs are available, but only the first five MPMs in the MPM list are binarized and coded in a manner based on the current context as described in the current JVET, as shown in Figure 14. The sixth MPM in the MPM list is considered one of 16 selection modes and is coded with a 4-bit fixed length code along with the other 15 selection modes.
[0094] As a non-limiting example, if the remaining 61 modes are indexed as {0, 1, 2, ..., 60}, then 15 selected modes may be obtained by evenly subsampling the remaining 61 intra modes as follows: The set of 15 selected modes may be {0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, where the 15 selected modes plus the 6th The MPMs are coded with 4 bits of a fixed length code, such as the set {6th MPM, 0, 5, 10, 14, 18, 22, 26, 30, 34, 38, 42, 46, 50, 55, 60}, and the balance of the 46 non-selected modes is coded with a truncated binary code, such as the set of non-selected modes = {1, 2, 3, 4, 6, 7, 8, 9, 11, 12...49, 51, 52, 53, 54, 56, 57, 58, 59}.
[0095] In a further alternative embodiment of Figure 13, only the first five MPMs in the MPM list are binarized and coded in a manner based on the current context as described in the current JVET standard, as shown in Figure 14. In such an embodiment, the sixth MPM in the MPM list is considered one of the 16 selection modes and is coded with a 4-bit fixed-length code along with the other 15 selection modes. Thus, the selection of the other 15 selected modes may be established using any known, convenient, and / or desired selection process. As non-limiting examples, they may be selected relative to the MPM mode, or relative to (content-based) statistically well-known modes, or relative to trained or historically well-known modes, or using other methods or processes.
[0096] Again, the selection of five MPMs is merely a non-limiting example, and in alternative embodiments, the set of MPMs may be further reduced to four or three MPMs, or expanded to more than six. There are still 16 selected modes, and the balance of the 67 (or other known, convenient, and / or desired total number) intra-coding modes is included in the set of non-selected intra-coding modes. That is, embodiments are contemplated in which the total number of intra-coding modes is greater or less than 67, in which the set of MPMs includes any known, convenient, or desired number of MPMs, and in which the amount of selected modes can be any known, convenient, and / or desired amount.
[0097] Execution of the sequences of instructions necessary to implement the embodiments may be performed by a computer system 1500 as shown in Figure 15. In one embodiment, execution of the sequences of instructions is performed by a single computer system 1500. According to other embodiments, multiple computer systems 1500 connected by communications links 1515 may execute the sequences of instructions in cooperation with one another. While the following provides a description of only one computer system 1500, it should be understood that any number of computer systems 1500 may be used to implement the embodiments.
[0098] A computer system 1500 according to one embodiment will now be described with reference to Figure 15, which is a block diagram of several functional components of computer system 1500. As used herein, the term computer system 1500 is used broadly to describe any computing device capable of storing and independently executing one or more programs.
[0099] Each computer system 1500 may include a communication interface 1514 coupled to bus 1506. The communication interface 1514 provides two-way communication between the multiple computer systems 1500. The communication interface 1514 of each computer system 1500 sends and receives electrical, electromagnetic, or optical signals containing data streams representing various types of signal information, such as commands, messages, and data. The communication link 1515 links one computer system 1500 with another computer system 1500. For example, the communication link 1515 may be a LAN, in which case the communication interface 1514 may be a LAN card, or the communication link 1515 may be a PSTN, in which case the communication interface 1514 may be an integrated services digital network (ISDN) card or modem, or the communication link 1515 may be the Internet, in which case the communication interface 1514 may be a dial-up, cable, or wireless modem.
[0100] Computer system 1500 may send and receive messages, data, and instructions, including programs, i.e., applications, code, via its corresponding communications link 1515 and communications interface 1514. The received program code may be executed by the respective processor 1507 upon receipt and / or stored in storage device 1510 or other associated non-volatile media for later execution.
[0101] In one embodiment, computer system 1500 operates in conjunction with data storage system 1531, for example, data storage system 1531 including database 1532 readily accessible by computer system 1500. Computer system 1500 communicates with data storage system 1531 via data interface 1533. Data interface 1533 coupled to bus 1506 sends and receives electrical, electromagnetic, or optical signals containing data streams representing various types of signal information, e.g., commands, messages, and data. In various embodiments, the functionality of data interface 1533 may be performed by communication interface 1514.
[0102] Computer system 1500 includes a bus 1506 or other communication mechanism for communicating instructions, messages, and data, collectively information, and one or more processors 1507 coupled to bus 1506 for processing information. Computer system 1500 also includes a main memory 1508, such as a random access memory (RAM) or other dynamic storage device coupled to bus 1506 for storing dynamic data and instructions executed by the one or more processors 1507. Main memory 1508 may also be used for storing temporary data, i.e., variables, or other intermediate information during execution of instructions by the one or more processors 1507.
[0103] Computer system 1500 may further include a read-only memory (ROM) 1509 or other static storage device coupled to bus 1506 for storing static data and instructions for the one or more processors 1507. A storage device 1510, such as a magnetic disk or optical disk, may also be provided and coupled to bus 1506 for storing data and instructions for the one or more processors 1507.
[0104] Computer system 1500 may be coupled via bus 1506 to a display device 1511, such as, but not limited to, a cathode ray tube (CRT) or liquid-crystal display (LCD) monitor, for displaying information to a user. An input device 1512, such as alphanumeric and other keys, is coupled to bus 1506 for communicating information and command selections to processor 1507.
[0105] According to one embodiment, each computer system 1500 performs particular operations by way of its corresponding one or more processors 1507 executing one or more sequences of one or more instructions contained in main memory 1508. Such instructions may be read into main memory 1508 from another computer-usable medium, such as ROM 1509 or storage device 1510. Execution of the sequences of instructions contained in main memory 1508 causes the one or more processors 1507 to perform the processes described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions. Thus, embodiments are not limited to any specific combination of hardware circuitry and / or software.
[0106] As used herein, the term "computer-usable medium" refers to any medium that can provide information or be used by one or more processors 1507. Such media can take many forms, including, but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media, i.e., media that can retain information without electrical power, include ROM 1509, CD-ROM, magnetic tape, and magnetic disks. Volatile media, i.e., media that cannot retain information without electrical power, includes main memory 1508. Transmission media include coaxial cables, copper wire, and fiber optics, including the wires that comprise bus 1506. Transmission media can also be in the form of carrier waves, i.e., electromagnetic waves that are modulated in frequency, amplitude, or phase to transmit information signals. Further, transmission media can take the form of acoustic or light waves, such as those generated during radio wave or infrared data communications.
[0107] In the foregoing specification, several embodiments have been described with reference to specific components thereof. However, it will be apparent that various changes and modifications can be made without departing from the broader spirit and scope of the embodiments. For example, it will be understood that the specific ordering and combination of process actions shown in the process flow diagrams described herein are merely exemplary, and that embodiments can be implemented using different or additional process actions, or using a different combination or ordering of process actions. Accordingly, the specification and drawings should be considered in an illustrative and not a restrictive sense.
[0108] It should also be noted that the present invention can be implemented in a variety of computer systems. The various techniques described herein may be embodied in hardware or software, or a combination of both. Preferably, these techniques are embodied in programmable computer programs running on a plurality of computers, each including a processor, a processor-readable storage medium (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device. The program code is applied to data entered using the input device to perform the functions described above and generate output information. The output information is applied to one or more output devices. Each program is preferably implemented in a general procedural or object-oriented programming language to communicate with a computer system. However, if desired, the program may be embodied in assembly or machine language. In either case, the language may be a compiled or interpreted language. Each such computer program is preferably stored on a general-purpose or special-purpose programmable computer-readable storage medium or device (e.g., ROM or magnetic disk) for configuring and operating the computer when the storage medium or device is read by the computer to perform the procedures described above. The system is also contemplated as embodied as a computer-readable storage medium configured to have a computer program thereon, the storage medium so configured causing the computer to operate in a particular predetermined manner. Furthermore, the storage element of an exemplary computing application may be a relational or sequential (flat file) type computing database capable of storing data in various combinations and configurations.
[0109] Figure 16 is a schematic diagram of a source device 1612 and a destination device 1610 incorporating features of the systems and devices described herein. As shown in Figure 16, the exemplary video coding system 1610 includes a source device 1612 and a destination device 1616, where in this example, the source device 1612 generates encoded video data. Accordingly, the source device 1612 may be referred to as a video encoder. The destination device 1616 may decode the encoded video data generated by the source device 1612. Accordingly, the destination device 1616 may be referred to as a video decoder. The source device 1612 and the destination device 1616 may be examples of video coding devices.
[0110] The destination device 1616 may receive the encoded video data from the source device 1612 over the channel 1616. The channel 1616 may comprise any type of medium or device that can move the encoded video data from the source device 1612 to the destination device 1616. In one example, the channel 1616 may comprise a communications medium that allows the source device 1612 to transmit the encoded video data directly to the destination device 1616 in real time.
[0111] In this example, source device 1612 may modulate the encoded video data according to a communication standard, such as a wireless communication protocol, and transmit the modulated video data to destination device 1616. The communication medium may include a wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or other equipment that enables communication from source device 1612 to destination device 1616. In another example, channel 1616 may correspond to a storage medium that stores the encoded video data generated by source device 1612.
[0112] 16, source device 1612 includes a video source 1618, a video encoder 1620, and an output interface 1622. In some cases, output interface 1628 may include a modulator / demodulator (modem) and / or a transmitter. In source device 1612, video source 1618 may include sources such as a video capture device such as a video camera, a video archive containing previously captured video data, a video feed interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources.
[0113] Video encoder 1620 may encode captured, pre-captured, or computer-generated video data. Input images may be received by video encoder 1620 and stored in input frame memory 1621. General-purpose processor 1623 may load information therefrom and perform the encoding. Programs for driving the general-purpose processor may be loaded from a storage device such as the exemplary memory module shown in FIG. 16. The general-purpose processor performs the encoding using processing memory 1622, and output of the encoded information by the general-purpose processor may be stored in a buffer, such as output buffer 1626.
[0114] The video encoder 1620 may include a resampling module 1625 configured to code (e.g., encode) the video data in a scalable video coding scheme that defines at least one base layer and at least one enhancement layer. The resampling module 1625 may resample at least some of the video data as part of the encoding process, and the resampling may be performed adaptively using a resampling filter.
[0115] The encoded video data, e.g., a coded bitstream, may be transmitted directly to a destination device 1616 via an output interface 1628 of the source device 1612. In the example of FIG. 16, the destination device 1616 includes an input interface 1638, a video decoder 1630, and a display device 1632. In some cases, the input interface 1628 may include a receiver and / or a modem. The input interface 1638 of the destination device 1616 receives the encoded video data via the channel 1616. The encoded video data may include various syntax elements generated by the video encoder 1620 that represent the video data. Such syntax elements may be included in the encoded video data that is transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0116] The encoded video data may also be stored on a storage medium or file server for later access by the destination device 1616 for decoding and / or playback. For example, the coded bitstream may be temporarily stored in an input buffer 1631 and then loaded into a general-purpose processor 1633. A program for driving the general-purpose processor may be loaded from a storage device or memory. The general-purpose processor may perform the decoding using process memory 1632. The video decoder 1630 may also include a resampling module 1635 similar to the resampling module 1625 used in the video encoder 1620.
[0117] 16 shows the resampling module 1635 separate from the general-purpose processor 1633, those skilled in the art will appreciate that the resampling function is performed by a program executed by the general-purpose processor, and that processing in a video encoder is accomplished using one or more processors. The decoded image or images may be stored in an output frame buffer 1636 and then sent to an input interface 1638.
[0118] The display device 1638 may be integrated with or external to the destination device 1616. In some examples, the destination device 1616 may include an integrated display device or be configured to interface with an external display device. In other examples, the destination device 1616 may be a display device. In general, the display device 1638 displays the decoded video data to a user.
[0119] The video encoder 1620 and the video decoder 1630 may operate according to a video compression standard. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11) are currently studying the potential need for standardization of future video coding techniques with compression capabilities significantly exceeding those of the current High Efficiency Video Coding (HEVC) standard (including its current and near-term extensions for screen content coding and high dynamic range coding). These groups are collaborating in this research effort in a joint collaboration known as the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by their experts in this field. A recent record of JVET development is described in "Algorithm Description of Joint Exploration Test Model 5 (JEM 5)," JVET-E1001-V2, by J. Chen, E. Alshina, G. Sullivan, J. Ohm, and J. Boyce.
[0120] Additionally or alternatively, the video encoder 1620 and the video decoder 1630 may operate according to other proprietary or industry standards that work with the disclosed JVET features. That is, other standards include the ITU-T H.264 standard, alternatively MPEG-4, Part 10, Advanced Video Coding (AVC), or extensions to such standards. Thus, the techniques of this disclosure, although newly developed for JVET, are not limited to any particular coding standard or technique. Other examples of video compression standards and techniques include MPEG-2, ITU-T H.263, and proprietary or open-source compression formats and related formats.
[0121] The video encoder 1620 and the video decoder 1630 may be embodied in hardware, software, firmware, or any combination thereof. For example, the video encoder 1620 and the video decoder 1630 may use one or more processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof.
[0122] If the video encoder 1620 and decoder 1630 are embodied partially in software, the device may store the software instructions on a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the video encoder 1620 and video decoder 1630 may be included in one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (codec) within the respective device.
[0123] Aspects of the subject matter described herein may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer, such as the general purpose processors 1623 and 1633 described above. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Aspects of the subject matter described herein may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including memory storage devices.
[0124] Examples of memory include random access memory (RAM), read-only memory (ROM), or both. The memory may store instructions, such as source code or binary code, for executing the techniques described above. The memory may also be used to store variables or other intermediate information during execution of instructions by a processor, such as processors 1623 and 1633.
[0125] The storage device may store instructions, such as source code or binary code, for executing the techniques described above. The storage device may also store data used and manipulated by the computer processor. For example, the storage device in the video encoder 1620 or the video decoder 1630 may be a database accessed by the computer system 1623 or 1633.
[0126] Other examples of storage devices include random access memory (RAM), read-only memory (ROM), a hard drive, a magnetic disk, an optical disk, a CD-ROM, a DVD, flash memory, a USB memory card, or any other medium that can be read by a computer.
[0127] The memory or storage device may be an example of a non-transitory computer-readable storage medium for use by or in connection with a video encoder and / or decoder. The non-transitory computer-readable storage medium includes instructions for controlling a computer system configured to perform functions described by certain embodiments. The instructions, when executed by one or more computer processors, may be configured to perform those described in certain embodiments.
[0128] Also, note that some embodiments are described as processes that are shown as flow diagrams or block diagrams. While each is described as sequentially processing multiple operations, many of the multiple operations can be performed in parallel or simultaneously. Furthermore, the order of multiple operations can be rearranged. A process may have additional steps not included in the figures.
[0129] Certain embodiments may be implemented in a non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, system, or machine. The computer-readable storage medium includes instructions for controlling a computer system to perform a method as described by the particular embodiment. The computer system may include one or more computing devices. When executed by one or more computer processors, the instructions may be configured to perform those described in the particular embodiment.
[0130] As used in this description and throughout the claims that follow, the words "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Also, as used in this description and throughout the claims that follow, the meaning of "in" includes "in" and "on," unless the context clearly dictates otherwise.
[0131] While exemplary embodiments of the present invention have been described in detail and language specific to the structural features and / or methodological operations set forth above, those skilled in the art will readily appreciate that many additional modifications are possible in the exemplary embodiments without materially departing from the novel teachings and advantages of the present invention. Furthermore, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or operations set forth above. Accordingly, these and all such modifications are intended to be included within the scope of the present invention, which is to be construed in breadth and scope in accordance with the appended claims.
Claims
1. 1. A method for decoding video data from a bitstream, comprising: (a) receiving the bitstream indicating how a coding tree unit has been partitioned into a plurality of coding units, wherein the plurality of coding units are rectangular; (b) determining a first set of most probable modes (MPMs) for the current block of video data, wherein the first set of MPMs is selectable based on an MPM index, one of the first set of MPMs selectable based on the MPM index includes a true horizontal mode, another of the first set of MPMs selectable based on the MPM index includes a true vertical mode, and another of the first set of MPMs selectable based on the MPM index includes an angular mode, and the first set of MPMs includes only five different modes; (c) deriving from the bitstream (i) an MPM flag including a total of one bit and (ii) another index, wherein at least one of the MPM flag and the another index indicates whether an intra mode for predicting the current block is one of the first set of MPMs; (d) if at least one of the MPM flag and the other index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index, selecting an intra mode for the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs; (e) if at least one of the MPM flag and the further index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, determining (i) at least one mode of a second set, and (ii) at least one mode of a third set, according to the MPM flag and the further index; (f) wherein the first set, the second set, and the third set include different modes, and the combination of the first set, the second set, and the third set includes 67 different modes; (g) determining an intra mode for the current block for at least one mode in the second set based on a first combination of the MPM flag and the other index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of MPMs; (h) determining an intra mode for the current block for at least one mode in the third set based on a second combination of the MPM flag and the other index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of MPMs.
2. A program that, when executed by a decoder, causes the decoder to perform each step of the method for decoding video data from a bitstream described in claim 1.
3. 1. A method for encoding video data by an encoder, comprising: (a) providing a bitstream indicating how a coding tree unit is divided into a plurality of coding units, the plurality of coding units being rectangular; (b) the bitstream includes data suitable for determining a first set of most probable modes (MPMs) for a current block of the video data, wherein the first set of MPMs is selectable based on an MPM index, one of the first set of MPMs selectable based on the MPM index includes a true horizontal mode, another of the first set of MPMs selectable based on the MPM index includes a true vertical mode, and another of the first set of MPMs selectable based on the MPM index includes an angular mode, and the first set of MPMs includes only five different modes; (c) the bitstream includes data suitable for deriving from the bitstream an MPM flag and another index including a total of one bit, where at least one of the MPM flag and the other index indicates whether an intra mode for predicting the current block is one of the first set of MPMs; and (d) the bitstream includes data suitable for selecting an intra mode for the current block based on the MPM index decoded from the bitstream of one of the first set of MPMs when at least one of the MPM flag and the other index is used to indicate that the intra mode for predicting the current block is one of the first set of MPMs selectable based on the MPM index; (e) the bitstream includes data suitable for determining, when at least one of the MPM flag and the further index indicates that the intra mode for predicting the current block is not one of the first set of MPMs, (i) determining at least one mode of a second set, and (ii) determining at least one mode of a third set, according to the MPM flag and the further index; (f) wherein the first set, the second set, and the third set include different modes, and the combination of the first set, the second set, and the third set includes 67 different modes; (g) the bitstream includes data suitable for determining an intra mode of the current block for at least one mode of the second set based on a first combination of the MPM flag and the other index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of MPMs; (h) the bitstream includes data suitable for determining an intra mode of the current block for at least one mode in the third set based on a second combination of the MPM flag and the other index that does not include any of the first set of MPMs selectable based on the MPM index included in the first set of MPMs.
Citation Information
Patent Citations
Method and apparatus for encoding / decoding an in-screen prediction mode using a candidate in-screen prediction mode.
JP2014528670A