Method, device, and computer program for video encoding or decoding

JP2024147698A5Pending Publication Date: 2025-05-12TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024113766
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-03-12
Filing Date
2024-07-17
Publication Date
2025-05-12

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method for encoding or decoding a video sequence capable of supporting a color difference format in VCC (Versatile Video Coding).SOLUTION: A method for performing encoding or decoding includes the steps of: when encoding or decoding a video sequence using a 4:4:4 color difference format, copying an affine motion vector of one 4×4 luma block using operation other than an averaging operation and associating the affine motion vector to a co-located 4×4 chroma block; and when encoding or decoding a video sequence using a 4:2:2 color difference format, associating each 4×4 chroma block with two 4×4 co-located luma blocks, thereby, causing an affine motion vector of one 4×4 chroma block to become an average of the motion vectors of the two co-located luma blocks.SELECTED DRAWING: Figure 19
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority under 35 U.S.C. §119 to U.S. Provisional Application No. 62 / 817,517, filed in the U.S. Patent and Trademark Office on March 12, 2019, the disclosure of which is incorporated by reference in its entirety.

[0002] [Technical field] Methods and apparatus according to embodiments relate to video processing, and more particularly to encoding or decoding video sequences that can support different chrominance formats (eg, 4:4:4, 4:2:2) in Versatile Video Coding (VCC). [Background technology]

[0003] Recently, the Video Coding Experts Group (VCEG) of the ITU Telecommunication Standardization Sector (ITU-T), a division of the International Telecommunication Union (ITU), and ISO / IEC MPEG (JTC 1 / SC 29 / WG 11), a standardization subcommittee of the Joint Technical Committee ISO / IEC JTC 1 of the International Organization for Standardization (ISO) and the International Electrotechnical Commission (IEC), published the H.265 / HEVC (High Efficiency Video Coding) standard (version 1) in 2013. This standard was updated to version 2 in 2014, version 3 in 2015, and version 4 in 2016.

[0004] After that, we are studying the potential need for standardization of future video coding technologies with compression capabilities significantly exceeding the HEVC standard (including its extensions). In October 2017, we issued a Joint Call for Proposals (CfP) for video compression with capabilities exceeding HEVC. By February 15, 2018, a total of 22 CfP responses for standard dynamic range (SDR), 12 CfP responses for high dynamic range (HDR), and 12 CfP responses for 360 video categories were submitted. In April 2018, all received CfP responses were evaluated at the 122 MPEG / 10th JVET (Joint Video Exploration Team - Joint Video Expert Team) meeting. After careful evaluation, JVET officially started the standardization of the next generation video coding beyond HEVC, i.e., the so-called Versatile Video Coding (VVC).

[0005] Here, the HEVC block partition structure is described. In HEVC, a coding tree unit (CTU) may be partitioned into coding units (CU) by using a quadtree structure, denoted as a coding tree, to adapt to various local characteristics. The decision of whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction may be made at the CU level. Each CU may be further partitioned into one, two or four prediction units (PU) according to a PU partition type. Inside one PU, the same prediction process may be applied, and related information may be transmitted to the decoder on a PU basis. After obtaining the residual block by applying a prediction process based on the PU partition type, the CU may be partitioned into transform units (TU) according to another quadtree structure, such as a coding tree for the CU. The feature of the HEVC structure is that it may include multiple partition concepts, including CU, PU and TU. In HEVC, a CU or TU can only be generally square in shape, while a PU may be square or rectangular in shape for an inter-prediction block. In HEVC, one coding block may be further divided into four square sub-blocks, and a transform may be performed on each sub-block, i.e., TU. Each TU can be further divided recursively (using quad-tree partitioning) into smaller TUs, which are called Residual Quad-Tree (RQT).

[0006] At picture boundaries, HEVC uses implicit quadtree partitioning so that blocks retain their quadtree partitioning until their size fits the picture boundary.

[0007] Here, a block partition structure using a quad-tree (QT) + binary tree (BT) is described. In HEVC, a CTU may be partitioned into CUs by using a quad-tree structure, denoted as a coding tree, to adapt to various local characteristics. The decision of whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction may be made at the CU level. Each CU may be further partitioned into one, two or four PUs according to a PU partition type. Inside one PU, the same prediction process may be applied, and related information may be transmitted to the decoder on a PU basis. After obtaining the residual block by applying a prediction process based on the PU partition type, the CU may be partitioned into transform units (TUs) according to another quad-tree structure, such as a coding tree for the CU. A feature of the HEVC structure is that it may include multiple partition concepts, including CU, PU, ​​and TU.

[0008] The QTBT structure removes the concept of multiple partition types, i.e., removes the separation of the concepts of CU, PU, ​​and TU, and supports more flexibility for the shape of CU partition. In the QTBT block structure, a CU can have either a square or rectangular shape. As shown in FIG. 1A, a coding tree unit (CTU) is first partitioned by a quadtree structure. Then, the quadtree leaf node may be further partitioned by a binary tree structure. There may be two partition types in the binary tree partition: symmetric horizontal partition and symmetric vertical partition. The binary tree leaf node is called a coding unit (CU), and its segmentation may be used for prediction and transformation processes without further partitioning. This means that CU, PU, ​​and TU have the same block size in the QTBT coding block structure. In JEM, a CU may be composed of coding blocks (CBs) of different color components in some cases. For example, one CU may include one luma CB and two chroma CBs for P slices and B slices of 4:2:0 chroma format, or may include a single component CB. For example, one CU may include only one luma CB or only two chroma CBs in the case of an I slice.

[0009] The following parameters are defined for the QTBT partitioning method: CTU size: Root node size of the quadtree, the same concept as HEVC MinQTSize: The minimum allowed quadtree leaf node size. MaxBTSize: The maximum allowed size of a binary tree root node. MaxBTDepth: The maximum allowed binary tree depth. MinBTSize: The minimum allowed size of a binary tree leaf node.

[0010] In one example of a QTBT partitioning structure, the CTU size may be set as 128x128 luma samples with two corresponding 64x64 blocks of chroma samples, MinQTSize may be set as 16x16, MaxBTSize may be set as 64x64, MinBTSize (both width and height) may be set as 4x4, and MaxBTDepth may be set as 4. Quad-tree partitioning may first be applied to the CTU to generate quad-tree leaf nodes. The quad-tree leaf nodes may have sizes from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the leaf quad-tree node is 128x128, it is not further split by the binary tree since the size exceeds MaxBTSize (i.e., 64x64). Otherwise, the leaf quad-tree node may be further split by the binary tree. Therefore, a quadtree leaf node may also be the root node of the binary tree and have the binary tree depth as 0. If the binary tree depth reaches MaxBTDepth (i.e., 4), no further splits are considered. If a binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal splits are considered. Similarly, if a binary tree node has a height equal to MinBTSize, no further vertical splits are considered. A binary tree leaf node may be further processed by the prediction and transformation process without further splits. In JEM, the maximum CTU size may be 256x256 luma samples.

[0011] FIG. 1A shows an example of block partitioning by using QTBT, and FIG. 1B shows the corresponding tree representation. Solid lines indicate quadtree partitioning, and dotted lines indicate binary tree partitioning. At each partition (i.e., non-leaf) node of the binary tree, a flag may be conveyed to indicate which partition type (i.e., horizontal or vertical) may be used, with 0 indicating horizontal partitioning and 1 indicating vertical partitioning. In quadtree partitioning, there is no need to indicate the partition type, since the quadtree partitioning splits the block both horizontally and vertically to generate four sub-blocks with equal size.

[0012] Furthermore, the QTBT scheme supports the flexibility of luma and chroma having separate QTBT structures. Currently, for P and B slices, the luma and chroma CTBs in one CTU share the same QTBT structure. However, for I slices, the luma CTB may be divided into CUs by the QTBT structure, and the chroma CTB may be divided into chroma CUs by another QTBT structure, i.e., the DualTree (DT) structure. This means that a CU in an I slice is composed of a coding block of a luma component or a coding block of two chroma components, and a CU in a P or B slice is composed of coding blocks of all three chroma components.

[0013] In HEVC, inter prediction for small blocks may be restricted to reduce memory access for motion compensation, bidirectional prediction is not supported for 4 × 8 and 8 × 4 blocks, and inter prediction is not supported for 4 × 4 blocks. The QTBT method implemented in JEM-7.0 may remove these restrictions.

[0014] Here, we explain block division using a ternary tree (TT). A multi-type-tree (MTT) structure has been proposed. MTT is a more flexible tree structure than QTBT. In MTT, horizontal and vertical central ternary trees are introduced in addition to quad-trees and binary trees, as shown in Figures 2A and 2B.

[0015] Some advantages of ternary tree partitioning include: It provides a complement to quadtree and binary tree partitioning. Ternary tree partitioning can capture objects located at block centers, while quadtrees and binary trees can be partitioned at block centers. No further transformations are needed, since the width and height of a ternary tree partition can be a power of two. The design of a two-level tree is primarily motivated by reduced complexity. Theoretically, the complexity of traversing a tree is T^D, where T is the number of partition types and D is the depth of the tree.

[0016] Now, the YUV format will be described. Different YUV formats, i.e., chrominance formats, are shown in Figure 3. Different chrominance formats define different downsampling grids for different color components.

[0017] Here, cross-component linear modeling (CCLM) will be described. In VTM, for the chrominance component of an intra PU, an encoder selects the best chrominance prediction mode from eight modes including planar, DC, horizontal, vertical, direct copy of intra prediction mode from luma component (DM), left and top cross-component linear mode (LT_CCLM), left cross-component linear mode (L_CCLM), and top cross-component linear mode (T_CCLM). Among these modes, LT_CCLM, L_CCLM, and T_CCLM can be classified as cross-component linear modes (CCLM). The difference between these three modes is that different regions of neighboring samples may be used to derive parameters α and β. In LT_CCLM, both left and top neighboring samples may be used to derive parameters α and β. In L_CCLM, only the left neighboring samples are used to derive the parameters α and β. In T_CCLM, only the upper neighboring samples are used to derive the parameters α and β.

[0018] A cross-component linear model (CCLM) prediction mode may be used to reduce redundancy in the chrominance components, and the chrominance samples may be predicted based on the reconstructed luma samples of the same CU by using the following linear model: pred c (i,j)=αrec L'(i,j)+β (Equation 1)

[0019] Here, pred c (i,j) represents the predicted chrominance sample in a CU, and rec L '(i,j) represents the downsampled reconstructed luma sample of the same CU. The parameters α and β may be derived by a linear formula, for example, the max-min method. This calculation process may be performed as part of the decoding process, not just as an encoder search operation, so no syntax is used to communicate the values ​​of α and β.

[0020] For the chrominance 4:2:0 format, CCLM prediction applies a 6-tap interpolation filter to obtain downsampled luma samples corresponding to the chrominance samples, as shown in Figure 3. Here, downsampled luma samples Rec'L[x,y] may be calculated from the reconstructed luma samples.

[0021] The downsampled luma samples may be used to find the maximum and minimum sample points. The two points (combined luma and chroma) (A, B) may be the minimum and maximum values ​​in a set of adjacent luma samples, as shown in Figure 4. Here, the linear model parameters α and β may be obtained according to the following formula:

number

[0022] Here, divisions may be avoided and replaced by multiplications and shifts. A Look-up Table (LUT) may be used to store pre-calculated values, and the absolute difference value between the maximum and minimum luminance samples may be used to specify an entry index of the LUT, and the size of the LUT may be 512.

[0023] In the T_CCLM mode, only the adjacent samples shown in FIGS. 5A and 5B (including 2*W samples) may be used to calculate the linear model coefficients.

[0024] In the L_CCLM mode, only the left adjacent samples (including 2*H samples) may be used to calculate the linear model coefficients, as shown in FIGS. 6A and 6B.

[0025] CCLM prediction modes also include prediction between two chrominance components, i.e., the Cr component may be predicted from the Cb component. Instead of using the reconstructed sample signal, CCLM Cb-to-Cr prediction may be applied to the residual domain. This may be implemented by adding a weighted reconstructed Cb residual to the original Cr intra prediction to form the final Cr prediction. pred cr *(i,j)=pred cr (i,j)+αres cb '(i,j) (Formula 4)

[0026] The CCLM luma-to-chroma prediction mode may be added as one more chroma intra prediction mode. On the encoder side, one more RD cost check may be added for the chroma component to select the chroma intra prediction mode. If a cb intra prediction mode other than the CCLM luma-to-chroma prediction mode can be used for the chroma component of a CU, the CCLM Cb-to-Cr prediction may be used to predict the Cr component.

[0027] Multiple Model CCLM (MMLM) is another extension of CCLM. As the name suggests, there may be more than one model in MMLM, for example, two models may be used. In MMLM, the neighboring luma and chroma samples of the current block may be classified into two groups, and each group may be used as a training set to derive a linear model (i.e., a specific α and β may be derived for a specific group). Furthermore, the samples of the current luma block may be classified based on the same rules for the classification of the neighboring luma samples.

[0028] FIG. 8 shows an example of classifying adjacent samples into two groups. The threshold may be calculated as the average value of adjacent reconstructed luma samples. Adjacent samples with Rec'L[x,y]<=Threshold may be classified into group 1, and adjacent samples with Rec'L[x,y]>Threshold may be classified into group 2.

number

[0029] Here, affine motion compensation prediction will be described. In HEVC, a translational motion model may be applied to motion compensation prediction (MCP). However, many types of motion may exist, such as zoom in / out, rotation, squinting motion, and other irregular motions. In VTM4, block-based affine transformation motion compensation prediction may be applied. As shown in FIG. 9, the affine motion field of a block may be described by motion information of two control point motion vectors (four parameters) or three control point motion vectors (six parameters).

[0030] For a four-parameter affine motion model, the motion vector at a sample position (x,y) within a block may be derived as follows:

number

[0031] For a six-parameter affine motion model, the motion vector at a sample position (x,y) within a block may be derived as follows:

number

[0032] Here, mv 0x ,mv 0y is the motion vector of the upper left corner control point, and mv 1x ,mv 1y is the motion vector of the control point in the upper right corner, and mv 2x ,mv 2y is the motion vector of the bottom-left corner control point.

[0033] To simplify the motion compensation prediction, a block-based affine transformation prediction may be applied. To derive the motion vector of each 4×4 luma sub-block, the motion vector of the center sample of each sub-block may be calculated according to the above formula and rounded to 1 / 16 fractional precision, as shown in FIG. 10. Then, a motion compensation interpolation filter may be applied to generate a prediction of each sub-block with the derived motion vector. The sub-block size of the chroma component may also be set to 4×4. The MV of a 4×4 chroma sub-block may be calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks. As done for translational motion inter prediction, there may be two affine motion inter prediction modes: affine merge mode and affine AMVP mode.

[0034] Here, we will describe affine merge prediction. AF_MERGE mode may be applied to CUs with width and height of 8 or more. In this mode, the CPMV of the current CU may be generated based on the motion information of spatially neighboring CUs. There may be up to five CPMVP candidates, and one index may be signaled to indicate the one to be used for the current CU. The following three types of CPVM candidates may be used to form the affine merge candidate list: (1) Inherited affine merge candidates extrapolated from the CPMVs of adjacent CUs (2) A constructed affine merge candidate CPMVP that can be derived using the translation MVs of adjacent CUs (3) Zero's music video.

[0035] In VTM4, there may be at most two inherited affine candidates, one from the left neighboring CU and one from the top neighboring CU, which may be derived from the affine motion model of the neighboring blocks. The candidate blocks are shown in FIG. 11. For the left predictor, the scanning order may be A0->A1, and for the top predictor, the scanning order may be B0->B1->B2. The first inherited candidate from each side may be selected. No pruning check needs to be performed between the two inherited candidates. If a neighboring affine CU is identified, its control point motion vector may be used to derive a CPMVP candidate in the affine merge list of the current CU. As shown in FIG. 12, if the neighboring bottom-left block A is coded in affine mode, the motion vectors v2, v3, and v4 of the top-left corner, top-right corner, and bottom-left corner of the CU containing block A may be obtained. If block A is coded with a four-parameter affine model, the two CPMVs of the current CU may be calculated according to v2 and v3. If block A may be coded with a six-parameter affine model, the three CPMVs of the current CU may be calculated according to v2, v3, and v4.

[0036] Constructing affine candidates means that candidates may be constructed by combining neighboring translational motion information of each control point. The motion information of a control point may be derived from the designated spatial and temporal neighbors shown in FIG. 13. CPMVk (k=1,2,3,4) represents the kth control point. For CPMV1, blocks B2->B3->A2 may be examined, and the MV of the first available block may be used. For CPMV2, blocks B1->B0 may be examined, and for CPMV3, blocks A1->A0 may be examined. If TMVP is available, it is used as CPMV4.

[0037] After the MVs of the four control points are obtained, an affine merge candidate may be constructed based on the motion information. The control point MV combinations {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, and {CPMV1, CPMV3} may be used to construct the candidate affine merges in order.

[0038] A combination of three CPMVs constructs a six-parameter affine merge candidate, and a combination of two CPMVs constructs a four-parameter affine merge candidate. To avoid the motion scaling process, the associated combination of control point MVs may be discarded if the reference indexes of the control points are different.

[0039] After the inheritance and construction affine merge candidates have been examined, if the list is not yet full, a MV of zero may be inserted at the end of the list.

[0040] Here, we will describe affine AMVP prediction. Affine AMVP mode can be applied to CUs with both width and height of 16 or more. A CU-level affine flag may be signaled in the bitstream to indicate whether affine AMVP mode should be used, and another flag may be signaled to indicate whether 4-parameter affine or 6-parameter affine should be used. In this mode, the difference between the CPMV of the current CU and these predictors CPMVP may be signaled in the bitstream. The size of the affine AVMP candidate list may be 2, and it may be generated by using the following four types of CPVM candidates in order: (1) Inheritance affine AMVP candidates that can be extrapolated from the CPMV of neighboring CUs (2) Construct affine AMVP candidate CPMVP that can be derived using the translation MVs of neighboring CUs (3) Translational MV from adjacent CU (4) Zero's music video

[0041] The testing order of inherited affine AMVP candidates may be the same as that of inherited affine merge candidates. The difference is that for AVMP candidates, affine CUs that have the same reference picture as the current block may be considered. When inserting an inherited affine motion predictor into the candidate list, no pruning process needs to be applied.

[0042] The constructed AMVP candidate may be derived from the specified spatial neighborhood shown in FIG. 13. The same check order as that performed in the affine merge candidate construction may be used. In addition, the reference picture indexes of the neighboring blocks may also be checked. The first block in the check order that has the same reference picture as the current CU and may be inter-coded may be used. If the current CU is coded in a 4-parameter affine mode and both mv0 and mv1 may be available, they may be added as one candidate in the affine AMVP list. If the current CU is coded in a 6-parameter affine mode and all three CPMVs are available, they may be added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set as unavailable.

[0043] If the affine AMVP list candidates are still less than two after the inherited affine AMVP candidates and the constructed AMVP candidates have been examined, then mv0, mv1, and mv2 are added in order as translation MVs to predict all control point MVs of the current CU, if available. Finally, if the affine AMVP list is still not full, MVs of zero may be used to fill the affine AMVP list.

[0044] Here, we will describe the storage of affine motion information. In VTM4, the CPMV of an affine CU may be stored in a separate buffer. The stored CPMV may be used to generate the inherited CPMVP in the affine merge mode and the affine AMVP mode for the last coded CU. The sub-block MV derived from the CPMV may be used for motion compensation, merging of translational MVs / MV derivation of AMVP lists, and deblocking.

[0045] In order to avoid picture line buffer for additional CPMV, affine motion data inheritance from CU from upper CTU may be treated differently from inheritance from normal neighboring CU. If a candidate CU for affine motion data inheritance is in upper CTU line, the lower left and lower right sub-block MVs in the line buffer may be used for affine MVP derivation instead of CPMV. In this way, CPMV may be stored in a local buffer. If the candidate CU is 6-parameter affine coded, the affine model may be reduced to a 4-parameter model. As shown in FIG. 14, along the upper CTU boundary, the motion vectors of the lower left and lower right sub-blocks of the CU may be used for affine inheritance of the CU in the lower CTU.

[0046] Despite the above-mentioned advances, in VTM-4.0, the MV of a 4×4 chrominance sub-block in an affine-coded coding block is calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks. However, for chrominance 4:4:4 and 4:2:2 formats, each 4×4 chrominance sub-block is associated with one or two 4×4 luma sub-blocks, and the current scheme of MV derivation for the 4×4 chrominance components has room for improvement to accommodate the cases of chrominance 4:4:4 and 4:2:2 formats. Summary of the Invention

[0047] According to an aspect of the present disclosure, a method for encoding or decoding a video sequence may include encoding or decoding the video sequence using a 4:4:4 chroma format or encoding or decoding the video sequence using a 4:2:2 chroma format, where if encoding or decoding the video sequence using the 4:4:4 chroma format, the method may further include copying an affine motion vector of one 4x4 chroma block using an operation other than an average operation, where if encoding or decoding the video sequence using the 4:2:2 chroma format, the method may further include associating each 4x4 chroma block with two 4x4 co-located chroma blocks such that the affine motion vector of one 4x4 chroma block is the average of the motion vectors of the two co-located chroma blocks.

[0048] According to an aspect of the present disclosure, the method may further include dividing the current 4x4 chroma block into four 2x2 sub-blocks, regardless of the chroma format, deriving a first affine motion vector of a co-located luma block for the top-left 2x2 chroma sub-block, deriving a second affine motion vector of a co-located luma block for the bottom-right 2x2 chroma block, and deriving an affine motion vector for the current 4x4 chroma block using an average of the first and second affine motion vectors.

[0049] According to an aspect of the present disclosure, the method may include adjusting an interpolation filter used for motion compensation between the luma and chroma components.

[0050] According to this aspect of the disclosure, if the video sequence is input using a 4:2:0 chrominance format, the method may further include applying an 8-tap interpolation filter to the luma and chrominance components.

[0051] According to an aspect of the present disclosure, the method may further include encoding the components Y, Cb, and Cr as three separate trees, each tree of the three separate trees encoding one component of the components Y, Cb, and Cr.

[0052] According to this aspect of the disclosure, encoding as three separate trees may be performed for an I-slice or I-tile group.

[0053] According to aspects of the present disclosure, the maximum allowable transform size may be the same for different color components.

[0054] According to this aspect of the disclosure, when encoding or decoding a video sequence using a 4:2:2 chrominance format, the maximum vertical size may be the same among different chrominance components, and the maximum horizontal transform size for the chrominance components may be half of the maximum horizontal transform size for the luma component.

[0055] According to aspects of the present disclosure, at least one of Position-Dependent Predictor combination (PDPC), Multiple Transform Selection (MTS), Non-Separable Secondary Transform (NSST), Intra-Sub Partitioning (ISP), and Multiple reference line (MRL) intra prediction may be applied to both the luma and chroma components.

[0056] According to this aspect of the disclosure, if multiple reference line (MRL) intra prediction is applied to both the luma and chroma components and the step of encoding or decoding the video sequence is performed using a 4:4:4 chroma format, the method may further include a step of selecting an Nth reference for intra prediction and using the same reference line without explicit signaling for the chroma component, if intra subdivision (ISP) is applied to both the luma and chroma components, the method may further include a step of applying intra subdivision (ISP) at the block level of the current block for components Y, Cb and Cr, and if different trees are used for different chroma components, the method may further include a step of implicitly deriving coding parameters for U and V components from the co-located Y component without signaling.

[0057] According to aspects of the disclosure, a device for encoding or decoding a video sequence may include at least one memory configured to store program code and at least one processor configured to read the program code and operate as instructed by the program code, the program code including first encoding or decoding code configured to cause the at least one processor to encode or decode a video sequence using a 4:4:4 chroma format or to encode or decode a video sequence using a 4:2:2 chroma format, the first encoding or decoding code configured to cause the at least one processor to encode or decode a video sequence using a 4:4:4 chroma format. When the first encoding or decoding code is configured to cause the at least one processor to encode or decode a video sequence using a 4:2:2 chroma format, the first encoding or decoding code may further include code configured to cause the at least one processor to copy an affine motion vector of one 4x4 chroma block using an operation other than an average operation, and when the first encoding or decoding code is configured to cause the at least one processor to encode or decode a video sequence using a 4:2:2 chroma format, the first encoding or decoding code may further include code configured to cause the at least one processor to associate each 4x4 chroma block with two 4x4 co-located chroma blocks such that the affine motion vector of one 4x4 chroma block is the average of the motion vectors of two co-located chroma blocks.

[0058] According to an aspect of the present disclosure, the first encoding or decoding code may further include code configured to cause at least one processor to split a current 4x4 chroma block into four 2x2 sub-blocks, derive a first affine motion vector of a co-located luma block for the top-left 2x2 chroma sub-block, derive a second affine motion vector of a co-located luma block for the bottom-right 2x2 chroma block, and derive an affine motion vector for the current 4x4 chroma block using an average of the first and second affine motion vectors.

[0059] According to aspects of the present disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to adjust an interpolation filter used for motion compensation between the luma and chroma components.

[0060] According to this aspect of the disclosure, when the first encoding or decoding code is configured to cause the at least one processor to encode or decode the video sequence using a 4:2:2 chrominance format, the first encoding or decoding code may further include code configured to cause the at least one processor to apply an 8-tap interpolation filter to the luma and chrominance components.

[0061] According to aspects of the present disclosure, the first encoding or decoding code may further include code configured to cause at least one processor to encode the components Y, Cb, and Cr as three separate trees, each tree of the three separate trees encoding one component of the components Y, Cb, and Cr.

[0062] According to this aspect of the disclosure, the encoding as three separate trees may be configured to be performed for an I-slice or I-tile group.

[0063] According to aspects of the present disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to allow the maximum transform size to be the same for different color components.

[0064] According to this aspect of the disclosure, when the first encoding or decoding code is configured to cause at least one processor to encode or decode a video sequence using a 4:2:2 chrominance format, the first encoding or decoding code may further include code configured to cause the at least one processor to set the maximum vertical size to be the same among different color components and to set the maximum horizontal transform size for the chrominance components to be half of the maximum horizontal transform size for the luma component.

[0065] According to aspects of the disclosure, the first encoding or decoding code may further include code configured to cause the at least one processor to apply at least one of Position-Dependent Predictor combination (PDPC), Multiple Transform Selection (MTS), Non-Separable Secondary Transform (NSST), Intra-Sub Partitioning (ISP), and Multiple reference line (MRL) intra prediction to both the luma and chroma components.

[0066] According to aspects of the disclosure, a non-transitory computer-readable medium may be provided having instructions stored thereon, the instructions including one or more instructions that, when executed by one or more processors of a device, cause the one or more processors to encode or decode a video sequence using a 4:4:4 chrominance format, or to encode or decode a video sequence using a 4:2:2 chrominance format, and if the instructions, when executed by the one or more processors of the device, cause the one or more processors to encode or decode a video sequence using a 4:4:4 chrominance format, the instructions, when executed by the one or more processors of the device, cause the one or more processors to The instructions, when executed by one or more processors of the device, further cause the processor to copy the affine motion vector of one of the 4x4 chrominance blocks using an operation other than an average operation, and if the instructions, when executed by one or more processors of the device, cause the one or more processors to encode or decode a video sequence using a 4:2:2 chrominance format, the instructions, when executed by the one or more processors of the device, further cause at least one processor to associate each 4x4 chrominance block with two co-located 4x4 chrominance blocks, such that the affine motion vector of the one of the 4x4 chrominance blocks is the average of the motion vectors of the two co-located chrominance blocks.

[0067] Although the above methods, devices, and non-transitory computer-readable medium are described separately, this description is not intended to suggest any limitation as to their scope of use or functionality, and in fact these methods, devices, and non-transitory computer-readable medium may be combined in other aspects of the present disclosure. [Brief description of the drawings]

[0068] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings. [Figure 1A] FIG. 2 is a diagram of a partitioned coding tree unit according to one embodiment. [Figure 1B]FIG. 2 is a diagram of a coding tree unit according to one embodiment. [Figure 2A] FIG. 2 is a diagram of a coding tree unit according to one embodiment. [Figure 2B] FIG. 2 is a diagram of a coding tree unit according to one embodiment. [Diagram 3] FIG. 2 is an illustration of different YUV formats according to one embodiment. [Figure 4] FIG. 2 is a diagram of different luminance values ​​according to one embodiment. [Figure 5A] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 5B] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 6A] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 6B] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 7A] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 7B] FIG. 1 is a sample diagram used in cross-component linear modeling according to one embodiment. [Figure 8] 1 is an example of classification using multi-model CCLM according to one embodiment. [Figure 9A] 1 is an example of an affine motion field for a block according to one embodiment. [Figure 9B] 1 is an example of an affine motion field for a block according to one embodiment. [Figure 10] 1 is an example of an affine motion vector field according to one embodiment. [Figure 11] 1 is an example of a candidate block for prediction according to one embodiment. [Figure 12] 1 is an example of a candidate block for prediction according to one embodiment. [Figure 13]1 is an example of a candidate block for prediction according to one embodiment. [Figure 14] 1 is an example of the use of motion vectors according to one embodiment. [Figure 15] 1 is a simplified block diagram of a communication system according to one embodiment. [Figure 16] FIG. 1 is a diagram of a streaming environment according to one embodiment. [Figure 17] FIG. 2 is a block diagram of a video decoder according to one embodiment. [Figure 18] FIG. 2 is a block diagram of a video encoder according to one embodiment. [Figure 19] 1 is a flowchart of an exemplary process for encoding or decoding a video sequence, according to one embodiment. [Figure 20] FIG. 1 is a diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0069] FIG. 15 shows a simplified block diagram of a communication system (400) according to one embodiment of the present disclosure. The communication system (400) may include at least two terminals (410-420) interconnected via a network (450). For one-way transmission of data, a first terminal (410) may encode video data at a local location for transmission to the other terminal (420) via the network (450). The second terminal (420) may receive the other terminal's encoded video data from the network (450), decode the encoded data, and display the reconstructed video data. One-way data transmission may be common in media provisioning applications, etc.

[0070] 15 illustrates a second pair of terminals (430, 440) provided to support bidirectional transmission of encoded video that may occur, for example, during a video conference. For bidirectional transmission of data, each terminal (430, 440) may encode video data captured at a local location for transmission over the network (450) to the other terminal device. Each terminal (430, 440) may also receive encoded video data transmitted by the other terminal, decode the encoded data, and display the reconstructed video data on a local display device.

[0071] In FIG. 15, the terminals (410-440) may be depicted as servers, personal computers, and smartphones, although the principles of the present disclosure are not so limited. The embodiments of the present disclosure may be applied to laptop computers, tablet computers, media players, and / or dedicated videoconferencing equipment. The network (450) represents any number of networks that convey encoded video data between the terminals (410-440), including, for example, wired and / or wireless communication networks. The communication network (450) may exchange data in circuit-switched and / or packet-switched channels. Exemplary networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For purposes of this description, the architecture and topology of the network (450) is not important to the operation of the present disclosure, unless otherwise described herein below.

[0072] 16 shows an arrangement of video encoders and decoders in a streaming environment as an example of an application of the disclosed subject matter. The disclosed subject matter is similarly applicable to other video-enabled applications including, for example, videoconferencing, digital TV, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.), etc.

[0073] The streaming system may include a capture subsystem (513), which may include, for example, a video source (501) (e.g., a digital camera) that generates an uncompressed video sample stream (502). The sample stream (502), depicted as a thick line to emphasize its higher amount of data compared to an encoded video bitstream, may be processed by an encoder (503) coupled to the camera (501). The encoder (503) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video bitstream (504), depicted as a thin line to emphasize its lower amount of data compared to the sample stream, may be stored in a streaming server (505) for future use. One or more streaming clients (506, 508) may access a streaming server (505) to obtain copies (507, 509) of the encoded video bitstream (504). The client (506) may include a video decoder (510) that decodes an input copy (507) of the encoded video bitstream and generates an output video sample stream (511) that can be rendered on a display (512) or other rendering device (not shown). In some streaming systems, the video bitstreams (504, 507, 509) may be encoded according to a particular video encoding / compression standard. Examples of these standards include H.265 HEVC. The developing video encoding standard is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0074] FIG. 17 may be a functional block diagram of a video decoder (510) according to one embodiment of the disclosure.

[0075] The receiver (610) may receive one or more coded video sequences to be decoded by the decoder (610), and in the same or other embodiments, may receive one coded video sequence at a time, with the decoding of each coded video sequence being independent of the other coded video sequences. The coded video sequences may be received from a channel (612), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (610) may receive the coded video data along with other data (e.g., coded audio data and / or auxiliary data streams), which may be forwarded to respective usage entities (not shown). The receiver (610) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (615) may be coupled between the receiver (610) and the entropy decoder / parser (620), hereinafter referred to as the parser (520). If the receiver 610 is receiving data from a store-and-forward device with sufficient bandwidth and controllability, or from an isochronous network, the buffer 515 may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer 615 may be needed and may be relatively large, and advantageously of adaptive size.

[0076] The video decoder (510) may include a parser (620) for recovering symbols (621) from the entropy coded video sequence. These categories of symbols include information used to manage the operation of the decoder (610), and potentially include information for controlling a rendering device such as a display (512), which may not be an integral part of the decoder, but may be coupled to the decoder, as shown in FIG. 17. The control information for the rendering device may be in the form of Supplemental Enhancement Information (SEI) (SEI message) or Video Usability Information (VUI) parameter set fragments (not shown). The parser (620) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow any video coding technique or standard, and may follow principles well known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (620) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on the at least one parameter corresponding to the group. The subgroups may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transformation unit (TU), a prediction unit (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter (QP) values, motion vectors, etc. from the coded video sequence.

[0077] The parser (620) may perform an entropy decoding / parsing operation on the video sequence received from the buffer (615) to generate symbols (621). The parser (620) may receive the encoded data and selectively decode particular symbols (621). Additionally, the parser (620) may determine whether a particular symbol (621) should be provided to a motion compensated prediction unit (653), a scaler / inverse transform unit (651), an intra prediction unit (652), or a loop filter (656).

[0078] The reconstruction of the symbols (621) may involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how may be controlled by subgroup control information parsed from the coded video sequence by the parser (620). The flow of such subgroup control information between the parser (620) and the following units is not shown for clarity.

[0079] In addition to the functional blocks described above, the decoder (510) may be conceptually subdivided into a number of functional units, as described below. In a practical implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0080] The first unit is a scalar / inverse transform unit (651), which receives the quantized transform coefficients as symbols (621) from the parser (620), along with control information (including which transform to use, block size, quantization factor, quantization scaling matrix, etc.). The scalar / inverse transform unit (651) may output a block containing sample values ​​that can be input to an aggregator (655).

[0081] In some cases, the output samples of the scaler / inverse transform (651) may relate to intra-coded blocks, i.e., blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed part of the current picture. Such prediction information may be provided by an intra-picture prediction unit (652). In some cases, the intra-picture prediction unit (652) uses surrounding already reconstructed information taken from the (partially reconstructed) current picture (656) to generate blocks of the same size and shape of the block being reconstructed. In some cases, the aggregator (655) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (652) to the output sample information provided by the scaler / inverse transform unit (651).

[0082] In other cases, the output samples of the scalar / inverse transform unit (651) may relate to an inter-coded and potentially motion-compensated block. In such cases, the motion compensation prediction unit (653) may access the reference picture memory (657) to retrieve samples used for prediction. After motion compensating the retrieved samples according to the symbols (621) associated with the block, these samples may be added by the aggregator (655) to the output of the scalar / inverse transform unit (in this case called residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory available to the motion compensation unit from which the motion compensation unit retrieves prediction samples may be controlled by a motion vector, e.g., in the form of a symbol (621) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, etc.

[0083] The output samples of the aggregator (655) may be subjected to various loop filtering techniques in a loop filter unit (656). The video compression techniques may include in-loop filter techniques, controlled by parameters contained in the coded video bitstream and made available to the loop filter unit (656) as symbols (621) from the parser (620), that may be responsive to meta-information obtained during decoding of previous portions of the coded pictures or coded video sequences (in decoding order), as well as to previously reconstructed loop filtered sample values.

[0084] The output of the loop filter unit (556) may be a sample stream that may be output to a rendering device (512) and also stored in a reference picture memory (656) for use in future inter-picture prediction.

[0085] Once a particular coded picture is fully reconstructed, it may be used as a reference picture for future prediction. For example, once a coded picture is fully reconstructed and identified as a reference picture (e.g., by the parser (620)), the current reference picture (656) may become part of the reference picture buffer (657), and a new current picture memory may be reallocated before beginning reconstruction of the subsequent coded picture.

[0086] The video decoder (510) may perform decoding operations according to a given video compression technique, which may be documented in a standard, such as H.265 HEVC. The coded video sequence may conform to the syntax specified by the video compression technique or standard being used, in the sense that the coded video sequence conforms to the syntax of the video compression technique or standard specified in the video compression technique or standard, and in particular those profile documents. Also, what is necessary for compliance is that the complexity of the coded video sequence is within a range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstructed sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level may be further limited through a Hypothetical Reference Decoder (HRD) specification and metadata about HRD buffer management conveyed in the coded video sequence.

[0087] In one embodiment, the receiver (610) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder (510) to properly decode the data and / or to more accurately recover the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0088] FIG. 18 may be a functional block diagram of a video encoder (503) according to one embodiment of the present disclosure.

[0089] The encoder (503) may receive video samples from a video source (501) (not part of the encoder), which may capture video images to be encoded by the encoder (503).

[0090] The video source (501) may provide a source video sequence to be encoded by the encoder (503) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media presentation system, the video source (501) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (503) may be a camera that captures local image information as a video sequence. The video data may be provided as a number of individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., being used. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0091] According to one embodiment, the encoder (503) may encode and compress pictures of a source video sequence into an encoded video sequence (743) in real-time or under any other time constraint required by an application. Achieving an appropriate encoding rate is one function of the controller (750). The controller (750) controls and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller may include rate control related parameters (picture skip, quantization, lambda values ​​for rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can readily recognize other functions of the controller (750) that may be associated with a video encoder (703) optimized for a particular system design.

[0092] Some video encoders operate in what one skilled in the art would readily recognize as a "coding loop." As a very simplified description, the coding loop may consist of a coding part of the encoder (730) (hereafter referred to as the "source coder"), which is responsible for generating symbols based on the input picture to be encoded and reference pictures, and a (local) decoder (733) embedded in the encoder (503). The decoder (733) recovers the symbols to generate sample data in the same way that the (remote) decoder would (so that any compression between the symbols and the encoded video bitstream is lossless in the video compression techniques contemplated in the disclosed subject matter). The reconstructed sample stream is input to a reference picture memory (734). Since the decoding of the symbol stream produces bit-wise accurate results independent of the location of the decoder (local or remote), the contents of the reference picture buffer are also bit-wise accurate between the local and remote encoders. In other words, the prediction part of the encoder "sees" exactly the same sample values ​​as the reference picture samples that the decoder "sees" when using the prediction during decoding. This basic principle of reference picture synchronization (including the resulting drift when synchronization cannot be maintained, for example, due to channel errors) is well known to those skilled in the art.

[0093] The operation of the "local" decoder (733) may be the same as the "remote" decoder (510), which has already been described in detail above in connection with Figure 16. However, with brief reference to Figure 16, because symbols are available and the encoding / decoding of symbols into a coded video sequence by the entropy coder (745) and parser (620) may be lossless, the entropy decoding portion of the decoder (510), including the channel (612), receiver (610), buffer memory (615), and parser (620), may not be fully implemented in the local decoder (733).

[0094] An observation that can be made at this point is that any decoder technique, other than analysis / entropy decoding, that exists in the decoder must necessarily exist in substantially the same functional form in the corresponding encoder. A description of the encoder techniques can be omitted, since they are the inverse of the decoder techniques, which are described generically. Only in certain areas is a more detailed description necessary, and is provided below.

[0095] As part of its operation, the source coder (730) may perform motion-compensated predictive coding, which predictively codes an input frame with reference to one or more previously coded frames from the video sequence designated as “reference frames.” In this manner, the coding engine (732) codes differences between pixel blocks of the input frame and pixel blocks of reference frames that may be selected as predictive references for the input frame.

[0096] The local video decoder (733) may decode the encoded video data of frames that may be designated as reference frames based on the symbols generated by the source coder (730). The operation of the encoding engine (732) may advantageously be a lossy process. If the encoded video data can be decoded in a video decoder (not shown in FIG. 17), the reconstructed video sequence may be a replica of the source video sequence, typically with some errors. The local video decoder (733) may replicate the decoding process that may be performed by the video decoder on the reference frames and store the reconstructed reference frames in a reference picture cache (734). In this way, the encoder (503) may locally store copies of reconstructed reference frames that have common content as the reconstructed reference frames (without transmission errors) obtained by the far-end video decoder.

[0097] The predictor (735) may perform a prediction search for the coding engine (732). That is, for a new frame to be coded, the predictor (735) may search the reference picture memory (734) for sample data (as candidate reference pixel blocks) or specific metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (735) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, the input picture determined by the search results obtained by the predictor (735) may have prediction references drawn from multiple reference pictures stored in the reference picture memory (734).

[0098] The controller (750) may manage the encoding operations of the video coder (730), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0099] The output of all the above functional units may be subjected to entropy coding in an entropy coder (745), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques known to those skilled in the art, such as, for example, Huffman coding, variable length coding, arithmetic coding, etc.

[0100] The transmitter (740) may buffer the encoded video sequence produced by the entropy coder (745) and prepare it for transmission over a communication channel (760), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (740) may merge the encoded video data from the video coder (730) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (not shown).

[0101] A controller (750) may manage the operation of the encoder (503). During encoding, the controller (750) may assign each coded picture a particular coding picture type, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following frame types:

[0102] An Intra picture (I-picture) may be one that can be coded and decoded without using other pictures in a sequence as a source of prediction. Some video codecs allow different types of Intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize these variations of I-pictures and their respective uses and characteristics.

[0103] A predictive picture (P-picture) may be encoded and decoded using intra- or inter-prediction, using at most one motion vector and reference index to predict the sample values ​​of each block.

[0104] Bidirectionally predicted pictures (B-pictures) may be encoded and decoded using intra- or inter-prediction, using up to two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predicted pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0105] In general, a source picture may be spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8 or 16x16 samples, respectively) and coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to the respective picture of the blocks. For example, blocks of an I-picture may be non-predictively coded or may be predictively coded with reference to already coded blocks of the same picture (spatial or intra prediction). Pixel blocks of a P-picture may be non-predictively coded via spatial or temporal prediction with reference to the reference picture coded one step before. Blocks of a B-picture may be non-predictively coded via spatial or temporal prediction with reference to the reference picture coded one or two steps before.

[0106] The video coder (503) may perform encoding operations according to a given video encoding technology or standard, such as H.265 HEVC. In its operations, the video coder (503) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the encoded video data may conform to a syntax specified by the video encoding technology or standard being used.

[0107] In one embodiment, the transmitter (740) may transmit additional data along with the coded video. The video coder (730) may include such data as part of the coded video sequence. The additional data may include other types of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, Supplementary Enhancement Information (SEI) messages, Visual Usability Information (VUI) parameter set fragments, etc.

[0108] This disclosure is directed to several block partitioning methods in which motion information is taken into account during tree partitioning for video coding. More specifically, the techniques of this disclosure relate to tree partitioning methods for flexible tree structures based on motion field information. The techniques proposed in this disclosure are applicable to both uniform and non-uniform derived motion fields.

[0109] The derived motion field of a block is defined as uniform if the derived motion field is available for all sub-blocks in the block and all motion vectors in the derived motion field are similar (e.g., the motion vectors share the same reference frame and the absolute differences between the motion vectors are all less than a certain threshold). The threshold may be signaled in the bitstream or may be predefined.

[0110] The derived motion field of a block is defined as non-uniform if a derived motion field is available for all sub-blocks within the block and the motion vectors in the derived motion field are not similar (e.g., if at least one motion vector references a reference frame that is not referenced by the other motion vector, or if at least one absolute difference between two motion vectors within the field is greater than a signaled or predetermined threshold).

[0111] Figure 19 is a flow chart of an example process (800) for encoding or decoding a video sequence. In some implementations, one or more of the processing blocks of Figure 19 may be performed by the decoder (510). In some implementations, one or more of the processing blocks of Figure 19 may be performed by another device or group of devices, such as the encoder (503), that is separate from or includes the decoder (510).

[0112] As shown in FIG. 19, the process (800) may include encoding or decoding (810) a video sequence using a 4:4:4 chrominance format or a 4:2:2 chrominance format.

[0113] If the process (800) includes a step of encoding or decoding a video sequence using a 4:4:4 chrominance format, the process (800) may further include a step (820) of copying an affine motion vector of one 4x4 chrominance block using an operation other than an average operation.

[0114] If the process (800) includes a step of encoding or decoding a video sequence using a 4:2:2 chrominance format, the process (800) may further include a step (830) of associating each 4x4 chrominance block with two 4x4 co-located chrominance blocks such that the affine motion vector of one 4x4 chrominance block is the average of the motion vectors of the two co-located chrominance blocks.

[0115] Although Figure 19 illustrates example blocks of process (800), in some implementations, process (800) may include more, fewer, different, or differently arranged blocks than those illustrated in Figure 19. Additionally or alternatively, two or more of the blocks of process (800) may be performed in parallel.

[0116] Additionally, the proposed methods may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium to perform one or more of the proposed methods.

[0117] The above techniques may be implemented as computer software using computer readable instructions and physically stored on one or more computer readable media. For example, Figure 20 illustrates a computer system (900) suitable for implementing certain embodiments of the disclosed subject matter.

[0118] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking, or similar mechanisms to generate code including instructions, which may be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., either directly or through an interpreter, microcode execution, etc.

[0119] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0120] 20 for computer system (900) are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components illustrated in the exemplary embodiment of computer system (900).

[0121] The computer system (900) may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users, for example, through tactile input (keystrokes, swipes, data glove movements, etc.), audio input (voice, clapping, etc.), visual input (gestures, etc.), and olfactory input (not shown). Human interface devices may also be used to capture certain media not necessarily directly associated with conscious human input, such as audio (e.g., speech, music, ambient sounds), images (scanned images, photographic images obtained from a still camera, etc.), and video (2D video, 3D video including stereoscopic pictures, etc.).

[0122] The input human interface devices may include one or more of a keyboard (901), a mouse (902), a trackpad (903), a touch screen (910), a data glove (904), a joystick (905), a microphone (906), a scanner (907), and a camera (908).

[0123] The computer system (900) may also include certain human interface output devices that may stimulate one or more of the senses of a human user, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touch screen (910), data gloves (904) or joystick (905), although there may be haptic feedback devices that do not function as input devices), audio output devices (speakers (909), headphones (not shown), etc.), visual output devices (screens (910) including cathode ray tube (CRT) screens, liquid-crystal display (LCD) screens, plasma screens or organic light-emitting diode (OLED) screens, each of which may or may not have touch screen input capability, each of which may or may not have haptic feedback capability, some of which may be capable of outputting two-dimensional visual output or three or more dimensional output through such means as stereoscopic output, virtual reality glasses (not shown), holographic displays and smoke tanks (not shown)), and printers (not shown).

[0124] The computer system (900) may also include human accessible storage devices and associated media, such as optical media, including CD / DVD ROM / RW (920) with CD / DVD or similar media (921), thumb drives (922), removable hard drives or solid state drives (923), legacy magnetic media, such as tape and floppy disks (not shown), specialized ROM / ASIC / PLD based devices, such as security dongles (not shown), etc.

[0125] Additionally, those skilled in the art should understand that the term "computer-readable medium" as used in connection with the subject matter disclosed herein does not include transmission media, carrier waves, or other non-transitory signals.

[0126] The computer system (900) may also include interfaces to one or more communications networks. The networks may be, for example, wireless, wired, optical. The networks may be local, wide area, metropolitan, vehicular and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet, wireless LANs, cellular networks (including global systems for mobile communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), Long-Term Evolution (LTE), etc.), TV wired or wireless wide area digital networks (including cable TV, satellite TV, and terrestrial broadcast TV), vehicular and industrial (including CANBus), etc. Certain networks typically require an external network interface adapter (such as a universal serial bus (USB) port of the computer system (900)) attached to a particular general-purpose data port or peripheral bus (949), while other network interface adapters are typically integrated into the core of the computer system (900) by being attached to a system bus (such as an Ethernet interface to a PC computer system or a cellular network to a smartphone computer system) as described below. Using any of these networks, the computer system (900) can communicate with other entities. Such communications may be one-way receive only (e.g., broadcast TV), one-way transmit only (e.g., CANbus to certain CANbus devices), or bidirectional, for example, to other computer systems using local or wide area digital networks. Specific protocols and protocol stacks may be used in each of these networks and network interfaces.

[0127] The above-mentioned human interface devices, human accessible storage devices and network interfaces may be attached to a core (940) of the computer system (900).

[0128] The core (940) may include one or more central processing units (CPUs) (941), graphics processing units (GPUs) (942), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (943), hardware accelerators for specific tasks (944), etc. These devices may be connected through a system bus (948), along with read only memory (ROM) (945), random access memory (RAM) (946), and internal mass storage (internal non-user accessible hard drive, solid state drive (SSD), etc.) (947). In some computer systems, the system bus (948) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (948) or through a peripheral bus (949). Peripheral bus architectures include peripheral component interconnect (PCI), USB, etc.

[0129] The CPU (941), GPU (942), FPGA (943) and accelerator (944) may execute certain instructions, which in combination may constitute the above computer code. The computer code may be stored in a ROM (945) or a RAM (946). Also, temporary data may be stored in the RAM (946), while persistent data may be stored, for example, in an internal mass storage device (947). A cache memory, which may be closely associated with one or more of the CPU (941), GPU (942), mass storage device (947), ROM (945), RAM (946), etc., may be used to enable fast storage and retrieval in any of the memory devices.

[0130] The computer-readable medium may bear computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0131] By way of example and not limitation, the architecture (900), and in particular a computer system having a core (940), may provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media may be media associated with a user-accessible mass storage device as described above, as well as specific storage of the core (940) of a non-transitory nature, such as mass storage device (947) internal to the core or ROM (945). Software implementing various embodiments of the present disclosure may be stored in such devices and executed by the core (940). The computer-readable media may include one or more memory devices or chips according to particular needs. The software may cause the core (940), and in particular the processor (including a CPU, GPU, FPGA, etc.) therein, to perform certain operations or certain portions of certain operations described herein, including defining data structures stored in RAM (946) and modifying such data structures according to operations defined by the software. Additionally or alternatively, the computer system may provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerator (944)), which may operate in place of or in conjunction with software to perform particular operations or portions of particular operations described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may include circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure includes any appropriate combination of hardware and software.

[0132] While this disclosure describes several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be recognized that those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. 1. A method for decoding a video sequence by a decoder, comprising: decoding the video sequence using one of a 4:4:4 chrominance format and a 4:2:2 chrominance format; if the video sequence is decoded using the 4:4:4 chrominance format, associating an affine motion vector of one 4x4 luma block with a co-located 4x4 chrominance block; relating the two 4x4 luma blocks to the one 4x4 chroma block such that, if the video sequence is decoded using the 4:2:2 chroma format, an affine motion vector of the one 4x4 chroma block is an average of the motion vectors of the two 4x4 luma blocks corresponding to the one 4x4 chroma block; The method includes:

2. A method for producing a chrominance image comprising the steps of: dividing a current 4×4 chrominance block into four 2×2 sub-blocks, regardless of chrominance format; deriving a first affine motion vector for a co-located luma block for the top-left 2×2 chroma sub-block; deriving a second affine motion vector for the co-located luma block for the bottom right 2×2 chroma block; deriving an affine motion vector for the current 4×4 chrominance block using an average of the first affine motion vector and the second affine motion vector; The method of claim 1 further comprising:

3. The method further comprising the step of decoding the components Y, Cb, and Cr as three separate trees, 3. The method of claim 1, wherein each tree of the three separate trees decodes one of the components Y, Cb and Cr.

4. The method described in claim 3, wherein the step of decoding as three separate trees is performed for an I slice or an I tile group.

5. A method according to any one of claims 1 to 4, wherein when decoding the video sequence using the 4:4:4 chrominance format, the maximum allowable transform size is the same for different color components.

6. A method according to any one of claims 1 to 5, wherein at least one of PDPC (Position-Dependent Predictor combination), MTS (Multiple Transform Selection), NSST (Non-Separable Secondary Transform), ISP (Intra-Sub Partitioning) and MRL (Multiple reference line) intra prediction is applied to both the luma component and the chroma component.

7. when the MRL intra prediction is applied to both the luma and chroma components and the step of decoding the video sequence is performed using the 4:4:4 chroma format, selecting an Nth reference for intra prediction and using the same reference line without explicit signaling for the chroma components; if the ISP is applied to both the luminance and chrominance components, applying the ISP at block level of the current block for components Y, Cb and Cr; if different trees are used for different color components, implicitly deriving coding parameters for U and V components from the co-located Y component without signaling. The method of claim 6 further comprising:

8. A device for decoding a video sequence, comprising: at least one memory configured to store program code; At least one processor configured to read said program code and to execute the method according to any one of claims 1 to 7; The device that contains 9. A computer program product for causing one or more processors to carry out a method according to any one of claims 1 to 7.

10. A method for an encoder encoding a video sequence, comprising the steps of: encoding the video sequence using one of a 4:4:4 chrominance format and a 4:2:2 chrominance format; if the video sequence is encoded using the 4:4:4 chrominance format, associating an affine motion vector of a 4x4 luma block with a co-located 4x4 chrominance block; if the video sequence is encoded using the 4:2:2 chrominance format, associating the two 4x4 luma blocks with the one 4x4 chrominance block such that an affine motion vector of the one 4x4 chrominance block is an average of the motion vectors of the two 4x4 luma blocks corresponding to the one 4x4 chrominance block; The method includes:

11. A method for processing visual media data, comprising the steps of: obtaining a visual media file; performing a conversion between the visual media file and a bitstream of visual media data; Including, the bitstream includes an encoded video sequence having a 4:4:4 chrominance format and a 4:2:2 chrominance format; The step of performing the conversion comprises: using an affine motion vector of a 4x4 luma block associated with a co-located 4x4 chroma block in the 4:4:4 chroma format; using an affine motion vector for a 4×4 chrominance block that is an average of motion vectors for two 4×4 luma blocks corresponding to the 4×4 chrominance block in the 4:2:2 chrominance format; A method comprising: