Combined internal and intra predictions

Combined inter- and intra-prediction tools enhance video encoding efficiency by leveraging both temporal and spatial redundancy, particularly in video regions with mixed objects, using advanced prediction models and decoder-side intra-mode derivation.

JP2026067920APending Publication Date: 2026-04-21INTERDIGITAL VC HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL VC HOLDINGS INC
Filing Date
2026-01-19
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing video coding systems face inefficiencies in encoding video regions where old and newly appearing objects are mixed, as conventional inter- and intra-prediction methods do not optimally leverage both temporal and spatial redundancy.

Method used

Implementing combined inter- and intra-prediction tools that combine intra-prediction with inter-prediction, using techniques such as affine motion models, alternative temporal motion vector prediction, and cross-component linear models to enhance encoding efficiency and reduce implementation complexity.

Benefits of technology

Improves encoding efficiency by accurately predicting video content with mixed objects, reducing redundancy and enhancing coding performance through decoder-side intra-mode derivation and adaptive filtering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026067920000001_ABST
    Figure 2026067920000001_ABST
Patent Text Reader

Abstract

We provide devices for combined inter- and intra-prediction. [Solution] The video encoding device receives an MMVD mode indicator indicating whether the MMVD mode is used to generate inter predictions for CUs, and for example, when the MMVD mode indicator indicates that the MMVD mode is not used to generate inter predictions for CUs, it receives a CIIP indicator and decides, for example, based on the MMVD mode indicator and / or the CIIP indicator, whether to use the triangular merge mode for CUs, and may disable the triangular merge mode for CUs on the condition that the CIIP indicator indicates that CIIP is applied to CUs, or that the MMVD mode indicator indicates that the MMVD mode is used to generate inter predictions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to combined inter- and intra-prediction, and more particularly to combined inter- and intra-prediction in video processing. [Background technology]

[0002] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 786,653, filed on 31 December 2018, which is incorporated herein by reference in its entirety.

[0003] Background technology Video coding systems are widely used to compress digital video signals, reducing storage needs and / or transmission bandwidth for the aforementioned signals. For example, block-based hybrid video coding systems are widely used and developed, along with various types of video coding systems such as block-based, wavelet-based, and object-based systems. [Overview of the project]

[0004] Systems, methods, and means are disclosed for combined inter and intra prediction. A video encoding device may receive an MMVD mode indicator indicating whether the MMVD (motion vector difference) mode is used to generate the inter prediction for the encoding unit. A video encoding device may receive a CIIP (combined inter merge / intra prediction) indicator, for example, when the MMVD mode indicator indicates that the MMVD mode is not used to generate the inter prediction for the encoding unit. A video encoding device may, for example, decide whether to use a triangular merge mode for the encoding unit based on the MMVD mode indicator and / or the CIIP indicator. A CIIP indicator may not be received when the MMVD mode indicator indicates that the MMVD mode is used for the encoding unit. An MMVD mode indicator may be received for each encoding unit.

[0005] A video encoding device may disable triangular merge mode for a coding unit if the CIIP indication indicates that CIIP is applied to the coding unit. A video encoding device may disable triangular merge mode for a coding unit if the MMVD mode indication indicates that MMVD mode is used to generate inter predictions. A video encoding device may enable triangular merge mode for a coding unit if the CIIP indication indicates that CIIP is not applied to the coding unit, and the MMVD mode indication indicates that MMVD mode is not used to generate inter predictions. A video encoding device may enable triangular merge mode for a coding unit without receiving the triangular merge flag. A video encoding device may infer whether to enable triangular merge mode for a coding unit based on one or more of the MMVD mode indication or CIIP indication. The CIIP indication may not be received when the MMVD mode indication indicates that MMVD mode is used to the coding unit. The MMVD mode indication may be received per coding unit base. [Brief explanation of the drawing]

[0006] [Figure 1] An illustrative diagram of a block-based video encoder is provided. [Figure 2] Figure 2(a) illustrates an example associated with block division in a multi-type tree structure and a quaternary division; Figure 2(b) illustrates an example associated with block division in a multi-type tree structure and a vertical bipartite division; Figure 2(c) illustrates an example associated with block division in a multi-type tree structure and a horizontal bipartite division; Figure 2(d) illustrates an example associated with block division in a multi-type tree structure and a vertical tripartite division; and Figure 2(e) illustrates an example associated with block division in a multi-type tree structure and a horizontal tripartite division. [Figure 3] An illustrative diagram of a block-based video decoder is provided. [Figure 4] Figure 4(a) illustrates an example of combined inter and intra predictions associated with the DC planar mode, Figure 4(b) illustrates an example of combined inter and intra predictions associated with the horizontal mode, and Figure 4(c) illustrates an example of combined inter and intra predictions associated with the vertical mode. [Figure 5] An example related to selective time-motion vector prediction is provided. [Figure 6] An example related to the modeling of affine motion fields is provided. [Figure 7] Figure 7A illustrates an example associated with diagonal triangle segmentation-based motion compensation prediction, and Figure 7B illustrates an example associated with inverse diagonal triangle segmentation-based motion compensation prediction. [Figure 8] An example related to generating a unidirectional predictive motion vector (MV) in triangular mode is provided. [Figure 9] An example related to intra-mode derivation is provided. [Figure 10] An example illustrating the relationship between the calculation of the gradient of a template sample for intra-mode derivation is provided. [Figure 11] This document illustrates examples of color difference prediction associated with combined inter- and intra-predictions using CCLM (cross-component linear model) color difference prediction samples. [Figure 12] This example illustrates the relationship between color difference prediction for combined inter- and intra-predictions by combining color difference inter-prediction samples and CCLM color difference prediction samples. [Figure 13A] This is a system diagram of an exemplary communication system in which one or more of the disclosed embodiments may be implemented. [Figure 13B]Figure 13A is a system diagram of an exemplary WTRU (Wireless Transmitter / Receiver Unit) that may be used in the communication system illustrated. [Figure 13C] Figure 13A is a system diagram of an exemplary RAN (Radio Access Network) and an exemplary CN (Core Network) that may be used within the communication system illustrated. [Figure 13D] Figure 13A is a system diagram of a further exemplary RAN and a further exemplary CN that may be used within the communication system illustrated. [Modes for carrying out the invention]

[0007] Encoding tools that can provide high encoding efficiency and moderate implementation complexity may include at least one of the following: affine motion models, alternative temporal motion vector prediction (ATMVP) or advanced temporal motion vector prediction, IMV (integer motion vector), GBi (generalized bi-prediction), BDOF (bi-directional optical flow), CIIP (combined inter-merge / intra prediction), MMVD (merge with motion vector difference), pairwise average merge candidate, triangular inter prediction for intercoding, CCLM (cross-component linear model), multi-line intra prediction, CPR (current picture referencing) for intra-prediction, EMT (enhanced multiple transform), dependent quantization for quantization and transform coding, and ALF (adaptive loop filtering) for in-loop filtering.

[0008] Figure 1 shows a block diagram of an exemplary block-based hybrid video encoder system 200. The input video signal 202 may be processed block by block. Extended block sizes (e.g., referred to as coding units or CUs) may be used to compress high-resolution (e.g., 1080p and / or higher) video signals. CUs may include sizes up to 128x128 pixels. Blocks may be partitioned based on quadtrees. A coding tree unit (CTU) may be partitioned into CUs based on quadtree / binary / ternary trees to adapt to changing local characteristics. CUs may or may not be partitioned into prediction units or PUs to which separate predictions may be applied. CUs may be used as the base unit for prediction and transformation without further partitioning. In a multi-type tree structure, a CTU (e.g., one) may be partitioned by a quadtree structure (e.g., it may be partitioned first). The leaf nodes of the quadtree (e.g., each quadtree leaf node) may be further partitioned by binary and ternary tree structures. As shown in Figure 2, there are one or more (e.g., five) division patterns. One or more of the following division patterns may be exemplary: quadrilateral division (e.g., (a)), horizontal bilateral division (e.g., (c)), vertical bilateral division (e.g., (b)), horizontal trilateral division (e.g., (e)), and vertical trilateral division (e.g., (d)).

[0009] Referring to FIG. 1, for an input video block (e.g., an MB (macroblock) or a CU), spatial prediction 260 or motion prediction 262 may be performed. Spatial prediction (e.g., or intra prediction) may use pixels from adjacent blocks that have already been encoded in the same video picture and / or slice that predicts the current video block. Spatial prediction may reduce the spatial redundancy inherent in the video signal. Motion prediction (e.g., also referred to as inter prediction or temporal prediction) may use pixels from an already encoded video picture that predicts the current video block. Motion prediction may reduce the temporal redundancy inherent in the video signal. The motion prediction signal for a given video block may be signaled by a motion vector indicating the amount and / or direction of motion between the current block and its reference block. If multiple reference pictures are supported, the reference picture index of the video block may be signaled to the decoder. The reference index may be used to identify from which reference picture in the reference picture store 464 the temporal prediction signal may come.

[0010] After spatial and / or motion prediction, mode determination 280 in the encoder may select a prediction mode based, for example, on rate distortion optimization. The prediction block may be subtracted from the current video block in 216. The prediction residual may be decorrelated using the transformation module 204 and the quantization module 206 to achieve the target bitrate. The quantized residual coefficients may be inversely quantized in 210 and inversely transformed in 212 to form the reconstructed residual. The reconstructed residual may be added back to the prediction block in 226 to form the reconstructed video block. In-loop filters, such as deblocking filters and / or adaptive loop filters, may be applied to the reconstructed video block in 266 before being placed in the reference picture store 264. Reference pictures in the reference picture store 264 may be used to encode future video blocks. The output video bitstream 220 may be formed. The coding mode (e.g., inter or intra), prediction mode information, motion information, and / or quantized residual coefficients are sent to the entropy coding unit 208, where they may be compressed and packed to form the bitstream 220.

[0011] FIG. 3 shows a general block diagram of an exemplary block-based video encoder. A video bitstream 302 may be received, unpacked, and / or entropy decoded at an entropy decoding unit 308. Encoding mode and / or prediction information may be sent to a spatial prediction unit 360 (e.g., if intra encoded) and / or a temporal prediction unit 362 (e.g., if inter encoded). Prediction blocks may be formed at the spatial prediction unit 360 and / or the temporal prediction unit 362. Residual transform coefficients may be sent to an inverse quantization unit 310 and an inverse transform unit 312 to reconstruct a residual block. The prediction block and the residual block may be added at 326. The reconstructed block may pass through in-loop filtering 366 and may be stored in a reference picture store 364. The reconstructed video in the reference picture store 364 may be used to drive a display device and / or to predict future video blocks.

[0012] One or more encoding modules, e.g., those associated with inter prediction, may be enhanced to improve inter encoding efficiency. One or more encoding tools may be described herein.

[0013] Combined inter and intra prediction may be performed.

[0014] As shown in Figures 1 and 3, inter-prediction and intra-prediction can be used to leverage the temporal and spatial redundancy present in the video signal. For example, the PU may leverage the correlation of the original video in either the temporal or spatial domain. Given the characteristics of inter-prediction and intra-prediction, the above scheme may not be optimal for certain video content. For example, for video regions where old and newly appearing objects are mixed, better encoding efficiency may be expected if there is a way to combine inter-prediction and intra-prediction together. Based on the above considerations, a combined inter-and intra-prediction tool may be used. The combined inter-and intra-prediction tool may combine intra-prediction with inter-prediction generated from merge mode. For example, for each CU encoded in merge mode, a flag may be signaled to indicate whether or not the combined inter-and intra-mode is applied. When the flag is true, additional syntax may be signaled to select an intra-mode from a predefined list of intra-mode candidates. For the luminance component, intra-mode candidates may include 4-frequently selected intra-modes, such as planar, DC, horizontal, and vertical. For the chrominance component, DM modes (e.g., indicating that the chrominance component reuses luminance intra-modes to generate predictive samples) may be applied without any signaling. Weights may be applied to combine inter-predictive samples and intra-predictive samples. One or more of the following may be applied:

[0015] Equal weights (e.g., 0.5) may be applied to CUs predicted by DC mode or planar mode and to CUs with a width or height of 4 or less.

[0016] If the CU is predicted by either the horizontal or vertical mode, and the CU is greater than four samples in width and height, then the CU may be divided horizontally or vertically (for example, dependent on the applied intra mode), and the division may occur within four equally sized regions. (W_intra i ,W_inter i A combination of weights shown as ), where i=0,···,3. For example, (W_intra0,W_inter0)=(0.75, 0.25), (W_intra1,W_inter1)=(0.625,0.375), (W_intra2,W_inter2)=(0.375,0.625), and (W_intra3,W_inter3)=(0.25,0.75), where (W_intra0,W_inter0) may correspond to the region closest to the reconstructed neighbor sample (e.g., the intra-reference sample), and (W_intra3,W_inter3) may correspond to the region furthest from the reconstructed neighbor sample. Figure 4 illustrates exemplary weights that can be applied to combine inter-predicted samples and intra-predicted samples for a combined inter and intra-prediction mode. For example, (a) illustrates exemplary weights that may be applied in DC and / or plane modes, (b) illustrates exemplary weights that may be applied in horizontal modes, and (c) illustrates exemplary weights that may be applied in vertical modes. Table 1 shows an exemplary coded unit syntax table after incorporating the syntax elements added to the combined inter and intra predictions.

[0017] [Table 1]

[0018] Referring to Table 1, one or more of the following may apply: mh_intra_flag[x0][y0] may indicate whether a combined inter and intra prediction is applied to the current coding unit. If mh_intra_flag[x0][y0] is not present, it may be inferred to be equal to 0. mh_intra_luma_mpm_flag[x0][y0] and mh_intra_luma_mpm_idx[x0][y0] may indicate an intra prediction mode for a luminance sample by, for example, calling the luminance intra prediction mode derivation process for an mh intra mode having, as input, a sample position (x0,y0), the width of the current coding block, the height of the current coding block, mh_intra_luma_mpm_flag[x0][y0], and mh_intra_luma_mpm_idx[x0][y0].

[0019] For example, combined inter and intra prediction modes may be enabled when the current CU is encoded by five spatially adjacent elements and two temporally adjacent elements, as is done for conventional merge modes, such as the merge mode in HEVC. Combined inter and intra prediction modes may be disabled for interCUs that signal MVs in the bitstream, or for interCUs encoded in other merge modes (e.g., affine merge, ATMVP, MMVD, and triangular prediction).

[0020] Subblock merge mode may occur.

[0021] A CU encoded by a merge mode (e.g., each CU) may have a set of motion parameters (e.g., one motion vector and one reference picture index) for each predicted direction (e.g., each predicted direction). One or more merge candidates that enable the derivation of motion information at the subblock level may be included in the merge mode. The category of subblock merge candidates may include ATMVP (alternative temporal motion vector prediction). ATMVP may be built on the same concept as the TMVP (temporal motion vector prediction) tool and may allow a CU to fetch subblock motion information from multiple small blocks from temporally adjacent pictures (e.g., reference pictures at the same location). The category of subblock merge candidates may include an affine merge mode that models the motion of subblocks inside a CU based on an affine model.

[0022] ATMVP (Attendant-Based Voter) may be performed.

[0023] In ATMVP, time motion vector prediction can be improved, for example, by allowing a block to derive motion information (e.g., multiple motion pieces, including motion vectors and reference indices) for subblocks within the current block. The motion information for each subblock (e.g., each subblock) can be derived from corresponding smaller blocks of temporally adjacent pictures of the current picture.

[0024] ATMVP can derive motion information for subblocks of a block. One or more of the following may apply: The corresponding block of the current block (for example, sometimes called a collocated block) may be identified in a selected temporal reference picture. The current block may be divided into subblocks, and the motion information for each subblock may be derived from the corresponding subblock in a collocated picture, for example, as shown in Figure 5.

[0025] Figure 5 illustrates an exemplary subblock motion information derivation 500. The current block may be divided into subblocks, and the motion information of each subblock may be derived from the corresponding subblock in a colocation picture, for example, as shown in Figure 5. The selected temporal reference picture may be called a collocated picture. One or more of the following may apply: Colocation blocks and colocation pictures may be identified by the motion information of the spatially adjacent blocks of the current block. Figure 5 illustrates an example associated with ATMVP. Referring to Figure 5, block A may be identified as the first available merge candidate in the merge candidate list of the current block. The corresponding motion vector of block A (e.g., MV A The reference index can be used to identify identical pictures and identical blocks. The position of an identical block in an identical picture is determined by the motion vector (MV) of block A. A This can be determined by adding the coordinates of the current block to the coordinates of the current block.

[0026] The current block may be divided into subblocks, and the motion information for each subblock may be derived from the corresponding subblock in the same location picture, for example, as shown in Figure 5. For example, the motion information for each subblock in the current block may be derived from the corresponding subblock in the same location block (for example, as shown by the red arrow in Figure 5). The motion information for the subblocks (for example, each subblock within the same location block) can be identified and converted into the motion vector and reference index of the corresponding subblock in the current block (for example, in a similar manner to TMVP in HEVC, where scaling of the time motion vector may be applied).

[0027] Affine models can be used to represent motion information. A video sequence may have one or more types of motion, such as translational motion, zoom in / out, rotation, perspective motion, or other irregular motion. Motion compensation predictions based on affine motion field modeling can be applied. Figure 6 illustrates an example of affine motion field modeling. As shown in Figure 6, the affine motion field of a block may be described by one or more, for example, three, control point motion vectors. Based on three control-point motion, the motion field of one affine block may be described as follows:

[0028]

number

[0029] Motion vector (v 0x ,v 0y ) may be the motion vector of the control point at the top left corner, and the motion vector (v 1x ,v 1y) can be the motion vector of the control point at the upper right corner. When one video block is encoded in affine mode, the motion field can be derived based on the granularity of a 4x4 block. To derive the motion vector for each 4x4 block, the motion vector of the center sample of each 4x4 subblock can be calculated according to (1). It can be expressed as an approximation with an accuracy of 1 / 16pel. The derived motion vector can be used in the motion compensation stage to generate the predicted signal for each subblock inside the current block.

[0030] Triangle inter prediction is sometimes performed. Figure 7 shows example triangle prediction partitions 700 and 702.

[0031] In some content (e.g., nature video content), the boundary between two moving objects may not be horizontal or vertical (e.g., perfectly horizontal or vertical), making it difficult to accurately approximate with rectangular blocks. Triangular prediction can be applied, for example, to enable triangular division for motion-compensated prediction. As shown in Figure 7, triangular prediction may divide the CU into one or more (e.g., two) triangular prediction units, for example, diagonally or inversely diagonally. Each triangular prediction unit (e.g., each triangular prediction unit within the CU) may be interpreted using its own unidirectional prediction motion vector and reference frame index, which may be derived from a list of unidirectional prediction candidates.

[0032] Figure 8 depicts an exemplary unidirectional prediction motion vector candidate derivation 800. The unidirectional prediction candidate list may include one or more (e.g., five) unidirectional prediction motion vector candidates. The unidirectional prediction motion vector candidates can be derived from similar (e.g., the same) spatial / temporal adjacent blocks as used for merge processing (e.g., the merge processing of HEVC). By way of example, the unidirectional prediction MV candidates can be derived from five spatially adjacent blocks and two temporally co-located blocks, as shown in FIG. 8. Referring to FIG. 8, the motion vectors of the seven adjacent blocks are collected, in order, as the L0 motion vector of the adjacent block, the L1 motion vector of the adjacent block, and the averaged motion vector of the L0 and L1 motion vectors of the adjacent block (if the adjacent block is bi-directionally predicted), and may be added to the unidirectional prediction MV candidates. If the number of MV candidates is less than five, a zero (0) motion vector is added to the MV candidate list.

[0033] Cross-component prediction may be performed for chroma intra prediction. There may be a correlation between the luminance component and the chroma components of certain video content (e.g., natural video content). The CCLM (cross-component linear model) prediction mode may be used for chroma intra prediction. In the CCLM prediction mode, chroma samples can be predicted from the reconstructed luminance samples of a block (e.g., the same block) by using a linear model, e.g., (2).

[0034] [Number]

[0035] Referring to (2), pred c (i,j) may indicate the prediction of a chroma sample in a block, pred L(i,j) may represent a luminance sample reconstructed from the same block at the same resolution as the chrominance block, which can be downsampled for 4:2:0 chrominance format content. Parameters α and β may represent the scaling parameter and offset of the linear model, respectively.

[0036] As described herein, for the luminance component, one or more intra-modes (e.g., used up to four times), including, for example, plane mode, DC mode, horizontal mode, and vertical mode, may be supported (e.g., by a combination of inter and intra-predictive modes). An encoding device (e.g., an encoder) may test multiple intra-modes. The encoding device may select from among the multiple intra-modes the one that provides the best performance (e.g., in terms of rate distortion trade-offs). The encoding device may signal the selected intra-mode (e.g., explicitly signaled to a decoder). For a combination of inter and intra-predictive modes, a non-negligible (e.g., significant) amount of bitrate may be spent encoding the intra-mode. For example, with increasing computing power, modern devices (e.g., even battery-powered devices such as wireless mobile devices equipped with decoders) may perform some sophisticated operations. The intra-modes used for the combination of inter and intra-predictive modes may be derived on the decoder side. If the decoder's derivation is accurate, intra-mode signaling may be skipped, potentially improving encoding efficiency.

[0037] Combined inter and intra predictions can be enabled when a CU is encoded by a merge mode. A merge mode may include using five spatially adjacent and two temporally adjacent elements for the merge mode (as shown, for example, in Figure 8). Combined inter and intra predictions may be disabled (e.g., always) for interCUs predicted by other merge modes (e.g., affine merge, ATMVP, MMVD, and triangular prediction). For example, combined inter and intra predictions for MMVD may be disabled because the MMVD mode is primarily selected in true bidirectional prediction scenarios (e.g., there are previous and subsequent predictions from reference lists L0 and L1). Motion-compensated predictions (e.g., inter predictions) may be accurate in predicting the current CU. Additional inter predictions may be unnecessary. Combined inter and intra predictions may be disabled for other merge modes. For one or more merge modes (e.g., ATMVP and triangle prediction modes), the movement of subblocks derived from spatially adjacent blocks in the same reference picture, or from blocks at the same temporal location in the reference picture, may not be accurate. In this case, enabling combinations of these merge modes with combined inter and intra predictions may be beneficial (for example, from the standpoint of coding performance).

[0038] The coding performance of combined inter and intra predictive modes may be improved. Intra-mode derivation may skip the overhead of signaling intra-modes by, for example, utilizing the computational power of the coding device (e.g., a decoder). The coding of chrominance may be improved for combined inter and intra predictive modes. The application of combined inter and intra predictive modes may be extended by one or more coding tools, including, for example, triangular inter predictive and / or subblock merge modes.

[0039] Combined inter and intra predictions may be performed, for example, by decoder intra-mode derivation. For example, in combined inter and intra predictions, a selected intra-mode of the luminance component (e.g., plane mode, DC mode, horizontal mode, or vertical mode) may be signaled to and / or from the encoding device (e.g., encoder or decoder). Signaling the selected intra-mode may add a non-negligible portion of the bitstream (e.g., output bit-steam) and / or reduce overall encoding performance. The intra-mode used for combined inter and intra predictions may be derived at the encoding device (e.g., decoder) to reduce overhead. In other words, the decoder may derive the intra-mode for combined inter and intra predictions (e.g., based on one or more adjacent reconstruction samples). For example, when combined inter and intra predictions are applied to a CU (for instance, instead of directly signaling the intra-mode in the bitstream), the intra-mode may be derived from adjacent reconstructed samples of the CU.

[0040] Figure 9 illustrates an exemplary intra-mode derivation 900 (for example, a decoder-side derivation of the intra-mode for combined inter and intra-predictions). As shown in Figure 9, the current CU to which the combined inter and intra-predictions are applied may have an N×N (width × height) size. A template may specify a set of reconstructed samples used to derive the intra-mode for the current CU. The shaded regions in Figure 9 may represent templates. The template size may be indicated as the number of samples in an L-shaped region extending above and to the left of the target block, e.g., L. For example, gradient analysis may be applied on template samples when there is a strong correlation (e.g., a template) between the samples of the current CU and adjacent blocks. Gradient analysis may be used to estimate the intra-mode to be used for the combined inter and intra-predictions for the current CU. In the example, as shown in (3), a 3×3 Sobel filter may be applied to calculate the horizontal and vertical gradients of the template samples.

[0041]

number

[0042] Two matrices (for example, M x and M y ) can be multiplied by a 3x3 window of the sample (for example, a sample placed in the center of a template sample). Two gradient values, G x and G y And the activity value Act h This can be calculated by summing the horizontal and vertical gradients of each template sample, for example, using (4).

[0043]

number

[0044] (4)

[0045]

number

[0046] and

[0047]

number

[0048] These may represent the horizontal and vertical slopes of the template sample at coordinate (i,j), respectively. Ω may represent the set of coordinates of the sample in the template. Figure 10 illustrates an exemplary slope calculation of a template sample for intra-mode derivation 1000. Given the slope and activity values ​​calculated in (4), the final intra-mode (e.g., the final intra-mode applied to the current CU) can be determined by (5).

[0049]

number

[0050] (5) See th g and th act This may refer to two predefined thresholds for gradient and activity values ​​that can be fixed and used by encoding devices (e.g., encoders and / or decoders). For example, the thresholds for gradient and activity values ​​may be determined by an encoding device (e.g., an encoder) and signaled to other encoding devices (e.g., decoders) in the bitstream.

[0051] In Figure 9, reconstructed samples forming an L-shaped region (e.g., reconstructed samples closest to the current CU) may be used as templates. In the example, templates with different shapes and / or sizes may be selected, which may provide different complexity / performance trade-offs. Choosing a large template size may cause the template samples to be farther away from the target block. The correlation between the template and the target block may be insufficient. A large template size may increase the complexity of encoding and decoding (e.g., given that more samples may be considered for gradient and activity calculations). A large template size may compromise reliable estimation when given noise. In the example, an adaptive template size may be used for intra-mode estimation. For example, the template size may be determined based on the CU size. The template size may be represented by the value of "L" in Figure 9. In the example, a first template size may be used for a certain CU size (e.g., less than 64 samples), and a second template size may be used for other CU sizes (e.g., CUs with 64 or more samples). The first template size can be 3 (for example, L=3 as shown in Figure 9). The second template size can be 5 (for example, L=5 as shown in Figure 9).

[0052] The decoder may select a luminance intra-mode for combined inter and intra-predictions. The luminance intra-mode for combined inter and intra-predictions may be selected, for example, by minimizing the difference between the inter-prediction sample and the intra-prediction sample. The luminance intra-mode for combined inter and intra-predictions may also be selected (for example, for each intra-prediction mode) by calculating the measured cost between the inter-prediction signal and the intra-prediction signal. One or more of the following cost measurements (e.g., template cost measurements) may be applied, such as the sum of absolute difference (SAD), the sum of square difference (SSD), and / or the sum of absolute transformed difference (SATD). Based on cost comparisons, the intra-prediction mode that yields the smallest template cost may be selected as the intra-prediction mode for the current CU (e.g., the best intra-prediction mode for the current CU). In the example, the intra-prediction mode associated with the lowest measured cost may be selected (e.g., by the decoder).

[0053] Combined inter and intra prediction may be performed by color difference CCLM. In combined inter and intra prediction, DM (direct mode) may be applied to one or more color difference components, for example, without signaling. In DM, the luminance component intra mode may be reused for one or more color difference components. Inter-channel correlation may exist between the luminance component and the color difference component. For example, the luminance component may have an inter-channel correlation with the color difference component. CCLM may be used for color difference intra prediction. For example, CCLM may be used for color difference intra prediction where the color difference sample is predicted from the corresponding luminance sample by subsampling applied based on a linear model. One or more CCLM model parameters may be derived from one or more adjacent luminance and / or color difference samples (e.g., casual adjacent luminance and color difference samples around the current CU). DM mode can be replaced with CCLM mode for intra-prediction of a color difference sample, for example, when combined inter and intra-prediction is applied to CU. In other words, CCLM mode can be applied to one or more color difference components (for example, instead of DM mode).

[0054] The CCLM mode can be applied to the chrominance component in one or more of the following ways: For example, one or more luminance prediction samples of CU may be generated by mixing prediction samples generated from intra-predictions (e.g., based on intra-modes signaled in the bitstream) and inter-predictions (e.g., based on MVs of adjacent blocks indicated by merge boxes). One or more luminance prediction samples and residual samples of the luminance component may be added together to generate (e.g., form) one or more reconstructed luminance samples. One or more reconstructed luminance samples may be subsampled. One or more reconstructed luminance samples may be used to generate one or more chrominance prediction samples based on the CCLM mode.

[0055] For example, one or more color difference prediction samples (e.g., CCLM prediction samples) can be combined with color difference interpretation samples to generate a final prediction sample for the color difference component. One or more color difference prediction samples are sometimes referred to as color difference interpretation samples.

[0056] Figure 11 illustrates an exemplary color difference prediction 1100 for combined inter and intra predictions using CCLM color difference prediction samples (e.g., directly). In a first exemplary method of applying the CCLM mode to the color difference component (e.g., as shown in Figure 11), the color difference LM prediction samples can be used directly.

[0057] Figure 12 illustrates an exemplary color difference prediction 1200 for inter and intra predictions combined by combining color difference inter prediction samples and CCLM color difference prediction samples. In a second exemplary method of applying the CCLM mode to the color difference component (for example, as shown in Figure 12), color difference intra prediction samples may be combined with color difference inter prediction samples.

[0058] DM and CCLM modes can be enabled for combined inter and intra-predictive color difference intra-predictions. A flag may be signaled, for example, when combined inter and intra-predictions are enabled for the current CU. The flag may indicate whether DM mode or CCLM mode is applied. One or more of the following may be applied: When the flag is set to true, DM mode may be applied to the color difference component, for example, so that the same intra-mode of the color difference component is reused to generate color difference intra-predictive samples. When the flag is set to false, CCLM mode may be applied to the intra-prediction of the color difference sample, for example, so that the color difference inter-predictive sample is generated from a subsampled luminance reconstruction sample based on a linear mode.

[0059] Combined inter and intra predictions may interact with other coding tools. Combined inter and intra predictions may be enabled when the CU is coded by a merge mode (e.g., conventional merge mode). Combined inter and intra prediction modes may interact with one or more other intercoding tools (e.g., affine merge, ATMVP, and triangular prediction). The interaction(s) (e.g., co-action of interactions) between the combined inter and intra prediction modes and one or more other intercoding tools may be enhanced. One or more of the following may apply:

[0060] Combined inter and intra predictions may be enabled for subblock merge mode. Motion information may be derived at the subblock level. For example, subblock derivation of motion information may be enabled through subblock merge mode, e.g., ATMVP and affine mode. In the example, combined inter and intra predictions may be disabled (e.g., for CUs predicted by subblock mode). Subblock motion information may be derived (e.g., fully derived) from one or more spatially adjacent elements (e.g., in affine merge mode). Subblock motion information may be derived from one or more temporally adjacent elements (e.g., in ATMVP mode). Subblock motion information may not be accurate enough to generate prediction samples for the current CU. Combined inter and intra predictions may be enabled for subblock merge mode. In the example, combined inter and intra predictions may be enabled for ATMVP mode (e.g., simply enabled) and disabled for affine merge mode (e.g., always disabled). The combined inter and intra prediction flags may be signaled, for example, to enable combined inter and intra prediction. The combined inter and intra prediction flags may be signaled after the subblock merge mode index. When the subblock merge index refers to an affine candidate, the combined inter and intra prediction flags may be skipped. When the subblock merge index refers to an affine candidate, the combined inter and intra prediction flags may be inferred to be false, for example, to disable combined inter and intra prediction. If the subblock merge index refers to an ATMVP candidate, the combined inter / intra flags may be signaled to indicate whether combined inter and intra prediction is enabled for the current CU.

[0061] Combined inter and intra predictions may be enabled for triangular modes (e.g., triangular merge modes). In triangular prediction, the PUs (e.g., each PU) of the current CU may be inter-predicted using MVs (e.g., motion vector candidates) from unidirectional prediction candidates for the current CU. In the unidirectional prediction candidate list, MV candidates may be derived from spatially and / or temporally adjacent blocks of the current CU, which may not be accurate enough to describe the motion of the CU. Combined inter and intra predictions may be used for triangular prediction modes, for example, to improve the efficiency of triangular inter predictions. Intra-prediction modes (e.g., single intra-prediction modes) may be signaled for one or more (e.g., two) PUs. Intra-prediction modes may be signaled separately for each PU, for example.

[0062] Triangle mode can be disabled for combined inter- and intra-predictive and / or MMVD modes. Triangle mode is sometimes referred to as triangle merge mode. Triangle merge mode can be enabled when CIIP is not applied and MMVD mode is not used. For example, a video encoding device (e.g., an encoder configured as shown in Figure 1) may decide whether to enable triangle merge mode. When the encoder decides to enable triangle merge mode, the encoder may decide not to apply CIIP and not to use MMVD mode. For example, a video encoding device may enable triangle merge mode without receiving a triangle merge mode indication (e.g., a triangle merge mode flag). A video encoding device (e.g., a decoder configured as shown in Figure 3) may receive an MMVD mode indication (e.g., an MMVD flag). For example, an encoder may send an MMVD mode indication to a video encoding device. The video encoding device may be a WTRU (e.g., WTRU102 shown in Figures 13A-13D). The video encoding device may include a WTRU (for example, WTRU102 shown in Figures 13A-13D). The MMVD mode indication may indicate whether the MMVD mode is used to generate inter-predictions for the CU. The MMVD mode indication may be received per encoding unit base. The video encoding device may receive a CIIP (combine inter-merge / intra prediction) indication (for example, a CIIP flag). For example, an encoder may send a CIIP indication to the video encoding device. The CIIP indication may not be received when the MMVD mode indication indicates that the MMVD mode is used for the CU. The CIIP indication may indicate that CIIP is applied to the CU.

[0063] For example, the triangle merge flag may be signaled after, for instance, the subblock merge flag, the combined inter / intra flag, and / or the MMVD flag. If one or more of the above flags are set to true, the triangle flag may not be signaled. The triangle merge flag may not be signaled when the CIIP indication indicates that CIIP applies to the CU. For example, an encoder may decide not to include the triangle merge flag in a message to a video encoding device (e.g., a decoder) when the CIIP indication indicates that CIIP does not apply to the CU. The triangle merge flag may not be signaled when the MMVD mode indication indicates that MMVD mode is used to generate inter predictions. For example, an encoder may decide not to include the triangle merge flag in a message to a video encoding device (e.g., a decoder) when the MMVD mode indication indicates that MMVD mode is used to generate inter predictions. When the triangle merge flag is not signaled, it may be inferred to be false (e.g., disabling triangle mode). For example, a video coding device may infer whether to enable triangular merge mode for a CU based on the MMVD mode indication and / or the CIIP indication. The video coding device may disable triangular merge mode for a CU, for example, when the MMVD mode indication indicates that MMVD mode is used to generate interpretations. The video coding device may also disable triangular merge mode for a CU, for example, when the CIIP indication indicates that CIIP is applied to the CU. If all three flags are set to false, the triangular flag may signal, for example, whether triangular mode is applied to the current CU.

[0064] Figure 13A illustrates an exemplary communication system 100 in which one or more disclosed embodiments may be implemented. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 100 may enable multiple wireless users to access the above content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods such as CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), FDMA (Frequency Division Multiple Access), OFDMA (Orthogonal Frequency Division Multiple Access), SC-FDMA (Single Carrier FDMA), ZT UW DTS-s OFDM (zero-tail unique-word DFT-Spread OFDM), UW-OFDM (unique word OFDM), resource block-filtered OFDM, and FBMC (filter bank multicarrier).

[0065] As shown in Figure 13A, the communication system 100 may include WTRUs (Wireless Transmitter / Receiver Units) 102a, 102b, 102c, 102d, RAN 104 / 113, CN 106 / 115, Public Switched Telephone Network (PSTN) 108, the Internet 110, and other networks 112, but the disclosed embodiments will be understood to anticipate several WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d may be any type of device configured to operate and / or communicate in a wireless environment. For example, WTRU102a, 102b, 102c, and 102d may all be referred to as “station” and / or “STA” and may be configured to transmit and / or receive wireless signals, and may include UEs (User Equipment), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, PDAs (Personal Digital Assistants), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, IoT (Internet of Things) devices, watches or other wearables, HMDs (Head-Mounted Displays), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, and devices operating on commercial and / or industrial wireless networks. Any of WTRU102a, 102b, 102c, and 102d may be interchangeable with UE.

[0066] Furthermore, the communication system 100 may also include base stations 114a and / or base stations 114b. Each of the base stations 114a and 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks, such as CN106 / 115, the Internet 110, and / or other networks 112. For example, base stations 114a and 114b may be a BTS (wireless base station equipment), Node-B, eNode B, home Node B, home eNode B, gNB, NR Node B, site controller, AP (access point), wireless router, etc. While base stations 114a and 114b are each depicted as single elements, it will be understood that base stations 114a and 114b may include any number of interconnected base stations and / or network elements.

[0067] Base station 114a may also be part of RAN 104 / 113, which may also include other base stations and / or network elements (not shown), such as BSC (Base Station Control Unit), RNC (Radio Network Control Unit), and relay nodes. Base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be called a cell (not shown). The frequencies mentioned may be permitted spectrum, unpermitted spectrum, or a combination of permitted and unpermitted spectrum. A cell may provide wireless service coverage to a particular geographical area, which may be relatively fixed or may change in the future. Furthermore, a cell may be divided into sectors. For example, a cell associated with base station 114a may be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, i.e., one for each sector of the cell. In one embodiment, the base station 114a may employ MIMO (multiple-input multiple output) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0068] Base stations 114a and 114b may communicate with one or more WTRUs 102a, 102b, 102c, and 102d via an air interface 116, which may be any suitable wireless communication link (e.g., RF (radio frequency), microwave, centimeter wave, micrometer wave, IR (infrared), UV (ultraviolet), visible light, etc.). The air interface 116 may be established using any suitable RAT (radio access technology).

[0069] More specifically, as described above, the communication system 100 may be a multiple access system and may employ one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, base stations 114a and WTRU 102a, 102b, 102c in RAN 104 / 113 may implement radio technologies such as UTRA (UMTS (Universal Mobile Telecommunications System) Terrestrial Radio Access), which may establish air interfaces 115 / 116 / 117 using WCDMA (wideband CDMA). WCDMA may include communication protocols such as HSPA (High-Speed ​​Packet Access) and / or HSPA+ (Evolved HSPA). HSPA may include HSDPA (High-Speed ​​DL (Downlink) Packet Access) and / or HSUPA (High-Speed ​​UL Packet Access).

[0070] In this embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may establish an air interface 116 using LTE (Long Term Evolution) and / or LTE-A (LTE-Advanced) and / or LTE-A Pro (LTE-Advanced Pro), and may implement radio technologies such as E-UTRA (Evolved UMTS Terrestrial Radio Access).

[0071] In one embodiment, base stations 114a and WTRUs 102a, 102b, and 102c may establish an air interface 116 using NR (New Radio), and may implement radio technologies such as NR radio access.

[0072] In some embodiments, base stations 114a and WTRUs 102a, 102b, and 102c may implement multiple radio access technologies. For example, base stations 114a and WTRUs 102a, 102b, and 102c may implement both LTE radio access and NR radio access, for example, using the principle of DC (dual connectivity). Thus, the air interfaces utilized by WTRUs 102a, 102b, and 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0073] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c may implement wireless technologies such as, for example, IEEE 802.11 (i.e., WiFi (Wireless Fidelity)), IEEE 802.16 (i.e., WiMAX (Worldwide Interoperability for Microwave Access)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, IS-2000 (Interim Standard 2000), IS-95 (Interim Standard 95), IS-856 (Interim Standard 856), GSM (Global System for Mobile communications), EDGE (Enhanced Data rates for GSM Evolution), and GERAN (GSM EDGE).

[0074] In Figure 13A, base station 114b may be, for example, a wireless router, home Node B, home eNode B, or access point, and any RAT suitable for facilitating wireless connectivity in localized areas such as businesses, homes, vehicles, campuses, industrial facilities, aerial walkways (for example, for use by drones), and roadways may be used. In one embodiment, base station 114b and WTRU 102c, 102d may implement wireless technology such as IEEE 802.11 to establish a WLAN (wireless local area network). In another embodiment, base station 114b and WTRU 102c, 102d may implement wireless technology such as IEEE 802.15 to establish a WPAN (wireless personal area network). In yet another embodiment, base stations 114b and WTRUs 102c, 102d may utilize cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish picocells or femtocells. As shown in Figure 13A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not be required to access the internet 110 via CNs 106 / 115.

[0075] RAN104 / 113 may be in communication with CN106 / 115 and may be any type of network configured to provide voice, data, applications, and / or VoIP (voice over internet protocol) services to one or more of WTRU102a, 102b, 102c, and 102d. The data may have various QoS (quality of service) requirements, such as different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, and mobility requirements. CN106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, internet connectivity, video distribution, and / or perform high-level security functions, such as user authentication. Although not shown in Figure 13A, it will be understood that RAN104 / 113 and / or CN106 / 115 may be in direct or indirect communication with other RANs employing the same or different RAT as RAN104 / 113. For example, in addition to being connected to RAN104 / 113, which may be utilizing NR radio technology, CN106 / 115 may also be in communication with another RAN (not shown) employing GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0076] Furthermore, CN106 / 115 may also serve as a gateway for WTRU102a, 102b, 102c, and 102d to access PSTN108, the Internet 110, and / or other networks 112. PSTN108 may include a circuit-switched telephone network providing POTS (plain old telephone service). The Internet 110 may include a global system consisting of interconnected computer networks and devices using common communication protocols, such as TCP (transmission control protocol) / IP (internet protocol) in the Internet protocol suite, UDP (user datagram protocol), and / or IP. Network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, network 112 may include another CN connected to one or more RANs that may employ the same RAT as RAN104 / 113 or a different RAT.

[0077] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode capabilities (for example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers to communicate with separate wireless networks via separate wireless links). For example, WTRU 102c, shown in Figure 13A, may be configured to communicate with base station 114a, which may employ cellular-based radio technology, and base station 114b, which may employ IEEE 802 radio technology.

[0078] Figure 13B is a system diagram illustrating an exemplary WTRU 102. As shown in Figure 13B, the WTRU 102 may include, among many others, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a GPS (Global Positioning System) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the above elements without regard to its embodiment.

[0079] The processor 118 could be a general-purpose processor, a dedicated processor, a conventional processor, a DSP (digital signal processor), multiple microprocessors, one or more microprocessors with a DSP core, a controller, a microcontroller, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) circuit, any other type of IC (integrated circuit), or a state machine. The processor 118 could perform signal coding, data processing, power control, input / output processing, and / or any other function that enables the WTRU 102 to operate in a wireless environment. The processor 118 could be coupled to a transceiver 120, which may be coupled to a transmit / receive element 122. Figure 13B depicts the processor 118 and the transceiver 120 as separate components, although it should be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0080] The transmit / receive element 122 may be configured to transmit or receive signals to or from a base station (e.g., base station 114a) via the air interface 116. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In another embodiment, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR signals, UV signals, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and optical signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0081] Although the transmit / receive element 122 is depicted as a single element in Figure 13B, the WTRU 102 may contain any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO technology. Thus, in one embodiment, the WTRU 102 may contain two or more transmit / receive elements 122 (e.g., multiple antennas) to transmit and receive wireless signals via the air interface 116.

[0082] The transceiver 120 may be configured to modulate a signal that is to be transmitted by the transmit / receive element 122, and to demodulate a signal that is to be received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multimode capabilities. Therefore, for example, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate by multiple RATs, such as NR and IEEE 802.11.

[0083] The processor 118 of the WTRU102 may be coupled with a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (for example, an LCD (liquid crystal display) display unit or an OLED (organic light-emitting diode) display unit) and may receive user input data. Furthermore, the processor 118 may output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 may access information and store data in any suitable type of memory, such as a non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include RAM (random-access memory), ROM (read-only memory), a hard disk, or any other type of memory storage device. The removable memory 132 may include a SIM (subscriber identity module) card, a Memory Stick, an SD (secure digital) memory card, etc. In another embodiment, the processor 118 may access information and store data in memory, for example, of a server or home computer (not shown), which is not physically located in the WTRU 102.

[0084] The processor 118 may receive power from the power supply 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power supply 134 may be any device suitable for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., NiCd (nickel-cadmium), NiZn (nickel-zinc), NiMH (nickel-metal hydride), Li-ion (lithium-ion), etc.), a solar cell, a fuel cell, etc.

[0085] Furthermore, the processor 118 may be coupled to a GPS chipset 136 which may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of the information from the GPS chipset 136, the WTRU 102 may determine its location based on receiving location information from base stations (e.g., base stations 114a, 114b) via the air interface 116 and / or based on the timing of signals received from two or more neighboring base stations. It will be understood that the WTRU 102 may acquire location information through any appropriate location-determination method, without regard to the manner in which it was configured.

[0086] Furthermore, the processor 118 may be coupled to other peripherals 138, which may include one or more software modules and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, e-compass, satellite transceiver, digital camera (for photos and / or video), USB (Universal Serial Bus) port, vibration device, television transceiver, hands-free headset, Bluetooth® module, FM (frequency modulated) radio unit, digital music player, media player, video game player module, internet browser, VR / AR (virtual reality and / or augmented reality) device, activity tracker, and the like. Peripheral device 138 may include one or more sensors, which may be one or more of the following: gyroscope, accelerometer, Hall effect sensor, magnetometer, compass sensor, proximity sensor, temperature sensor, time sensor, geolocation sensor, altimeter, light sensor, touch sensor, magnetometer, barometer, gesture sensor, biometric sensor, and / or humidity sensor.

[0087] WTRU102 may include a full-duplex radio in which some or all of the transmission and reception of signals (e.g., associated with a particular subframe with respect to both UL (e.g., for transmission) and downlink (e.g., for reception) may be in parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and / or substantially eliminate self-interference by either hardware (e.g., chokes) or signal processing by a processor (e.g., separate processors (not shown), or by processor 118). In an embodiment, WRTU102 may include a half-duplex radio in which some or all of the transmission and reception of signals (e.g., associated with a particular subframe with respect to either UL (e.g., for transmission) or downlink (e.g., for reception) may be in parallel and / or simultaneously.

[0088] Figure 13C is a system diagram illustrating RAN104 and CN106 according to different embodiments. As described above, RAN104 may employ E-UTRA's wireless technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. Furthermore, RAN104 may also be in communication with CN106.

[0089] RAN104 may include eNode-B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of eNode-B without regard to the embodiment. Each of eNode-B160a, 160b, and 160c may include one or more transceivers to communicate with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, eNode-B160a, 160b, and 160c may implement MIMO technology. Thus, for example, eNode-B160a may use multiple antennas to transmit and / or receive wireless signals to WTRU102a.

[0090] Each of the eNode-B160a, 160b, and 160c may be associated with a specific cell (not shown) and may be configured to handle decisions regarding radio resource management, handover decisions, user scheduling in UL and / or DL, etc. As shown in Figure 13C, the eNode-B160a, 160b, and 160c may communicate with each other via the X2 interface.

[0091] The CN106 shown in Figure 13C may include an MME (mobility management entity) 162, an SGW (serving gateway) 164, and a PDN (packet data network) gateway (or PGW) 166. While each of the above elements is depicted as part of CN106, it should be understood that any of the elements mentioned may be owned and / or operated by an entity other than the CN operator.

[0092] MME162 may be connected to each of the eNode-B162a, 162b, and 162c in RAN104 via the S1 interface and may act as a control node. For example, MME162 may be responsible for authenticating users of WTRU102a, 102b, and 102c, activating / deactivating bearers, and selecting a specific serving gateway during the initial connection of WTRU102a, 102b, and 102c. MME162 may provide control plane functionality for switching between RAN104 and other RANs (not shown) employing other radio technologies such as GSM and / or WCDMA.

[0093] SGW164 may be connected to each of the eNode B160a, 160b, and 160c in RAN104 via the S1 interface. Generally, SGW164 may route and forward user data packets to WTRU102a, 102b, and 102c. SGW164 may also perform other functions, such as fixing the user plane during eNode B handovers, triggering paging when DL data is available to WTRU102a, 102b, and 102c, and managing and storing the context of WTRU102a, 102b, and 102c.

[0094] SGW164 may be connected to PGW166, which may provide WTRU102a, 102b, and 102c with access to a packet-switched network, such as the Internet 110, in order to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices.

[0095] CN106 may facilitate communication with other networks. For example, CN106 may provide WTRU102a, 102b, and 102c with access to a circuit-switched network, such as PSTN108, to facilitate communication between WTRU102a, 102b, and 102c and conventional terrestrial communication line communication devices. For example, CN106 may include, or communicate with, an IP gateway (e.g., an IMS (IP Multimedia Subsystem) server) acting as an interface between CN106 and PSTN108. In addition, CN106 may provide WTRU102a, 102b, and 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0096] Although the WTRU is described as a wireless terminal in Figures 13A to 13D, in a typical embodiment, it is expected that the terminal may use a wired communication interface with a communication network (for example, temporarily or permanently).

[0097] In a typical configuration, the other network 112 may be a WLAN.

[0098] In the BSS (Basic Service Set) mode of infrastructure, a WLAN may have an AP (Access Point) for the BSS and one or more STAs (Stations) associated with the AP. The AP may have access to or interfaces to a DS (Distribution System), or another type of wired / wireless network that carries traffic entering and leaving the BSS. Traffic originating outside the BSS and destined for the STA may arrive through the AP and be delivered to the STA. Traffic originating from the STA and destined for destinations outside the BSS may be sent to the AP and delivered to their respective destinations. Traffic between STAs within the BSS may be sent through the AP, for example, if the originating STA sends traffic to the AP, and the AP delivers the traffic to the destination STA. Traffic between STAs within the BSS may be considered and / or said to be peer-to-peer traffic. Peer-to-peer traffic may be sent between the originating and destination STAs (for example, directly between them) via a DLS (direct link setup). In a typical embodiment, DLS may use 802.11e DLS or 802.11z TDLS (tunneled DLS). A WLAN using IBSS (Independent BSS) mode may not have APs, and STAs within or using IBSS (e.g., all STAs) may communicate directly with each other. The IBSS mode of communication may sometimes be referred to as the “ad-hoc” mode of communication in this specification.

[0099] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as the primary channel. The primary channel may have a fixed width (e.g., a 20 MHz bandwidth) or a width dynamically set by signaling. The primary channel may be the operating channel of the BSS and may be used by the STA to establish a connection with the AP. In one typical embodiment, CSMA / CA (Carrier Sensing Multiple Access / Collision Avoidance) may be implemented in the 802.11 system, for example. With respect to CSMA / CA, an STA, including the AP (e.g., any STA), may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, that particular STA may back off. A single STA (e.g., just one station) may transmit at any given time in a given BSS.

[0100] A high-throughput (HT) STA may use a 40MHz wide channel for communication by combining a 20MHz primary channel with adjacent or non-adjacent 20MHz channels, for example, to configure a 40MHz wide channel.

[0101] A VHT (Very High Throughput) STA may support channels with widths of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz. 40 MHz and / or 80 MHz channels may be constructed by combining contiguous 20 MHz channels. A 160 MHz channel may be constructed by combining eight contiguous 20 MHz channels, or by combining two non-contiguous 80 MHz channels, which may constitute an 80+80 configuration. For the 80+80 configuration, data may proceed to a segment parser that, after channel encoding, splits the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time-domain processing may be performed separately on each stream. Streams may be mapped onto two 80 MHz channels, and data may be transmitted by the transmitting STA. In the receiving STA receiver, the operation for the 80+80 configuration described above may be inverted, and the combined data may be sent to MAC (Media Access Control).

[0102] A sub-1GHz mode of operation is supported by 802.11af and 802.11ah. The operating bandwidth and carrier of the channel are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5MHz, 10MHz, and 20MHz bandwidths in the TVWS (TV White Space) spectrum, while 802.11ah supports 1MHz, 2MHz, 4MHz, 8MHz, and 16MHz bandwidths using the non-TVWS spectrum. In a typical embodiment, 802.11ah may support Meter Type Control / Machine-Type Communication, for example, MTC devices in macro coverage areas. MTC devices may have limited performance, e.g., limited performance including support for certain and / or limited bandwidths (e.g., support only). MTC devices may contain batteries with battery life exceeding a threshold (for example, to maintain a very long battery life).

[0103] A WLAN system that may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, may include a channel that may be designated as the primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, even if the AP and other STAs in the BSS support operating modes of 2MHz, 4MHz, 8MHz, 16MHz, and / or other channel bandwidths, the primary channel may be 1MHz wide for an STA (e.g., an MTC type device) that supports (e.g., only supports) the 1MHz mode. Carrier sensing and / or NAV (Network Allocation Vector) settings may depend on the state of the primary channel. If the primary channel is busy, for example, due to an STA (which only supports a 1MHz operating mode) transmitting to the AP, then the entire available frequency band may be considered busy, even if a large portion of the frequency band could remain idle and be available.

[0104] In the United States, the available frequency band that may be used by 802.11ah is from 902 MHz to 928 MHz. In South Korea, the available frequency band is from 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is from 916.5 MHz to 927.5 MHz. The total bandwidth available for 802.11ah is from 6 MHz to 26 MHz, depending on the country code.

[0105] Figure 13D is a system diagram illustrating RAN113 and CN115 according to different embodiments. As described above, RAN113 may employ NR radio technology to communicate with WTRU102a, 102b, and 102c via the air interface 116. Furthermore, RAN113 may also be in communication with CN115.

[0106] RAN113 may include gNB180a, 180b, and 180c, but it will be understood that RAN113 may include any number of gNBs without regard to the embodiment. Each of the gNB180a, 180b, and 180c may include one or more transceivers to communicate with WTRU102a, 102b, and 102c via the air interface 116. In one embodiment, the gNB180a, 180b, and 180c may implement MIMO technology. For example, the gNB180a and 108b may utilize beamforming to transmit and / or receive signals to the gNB180a, 180b, and 180c. Thus, for example, the gNB180a may use multiple antennas to transmit and / or receive wireless signals to the WTRU102a. In some embodiments, gNB180a, 180b, and 180c may implement carrier aggregation techniques. For example, gNB180a may transmit multiple component carriers to WTRU102a (not shown). A subset of these component carriers may be on an unallowed spectrum while the remaining component carriers may be on an allowed spectrum. In some embodiments, gNB180a, 180b, and 180c may implement Coordinated Multi-Point (CoMP) techniques. For example, WTRU102a may receive coordinated transmissions from gNB180a and gNB180b (and / or gNB180c).

[0107] WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using transmissions associated with scalable numerology. For example, OFDM symbol spacing and / or OFDM subcarrier spacing may vary for separate transmissions, separate cells, and / or separate portions of the spectrum of wireless transmissions. WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using subframes or TTI (Transmit Time Interval) of varying or scalable lengths (e.g., including varying numbers of OFDM symbols and / or persistent absolute times of varying lengths).

[0108] gNB180a, 180b, and 180c may be configured to communicate with WTRU102a, 102b, and 102c in standalone and / or non-standalone configurations. In a standalone configuration, WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c without accessing other RANs (e.g., eNode-B160a, 160b, and 160c). In a standalone configuration, WTRU102a, 102b, and 102c may use one or more of gNB180a, 180b, and 180c as a mobility anchor point. In a standalone configuration, WTRU102a, 102b, and 102c may communicate with gNB180a, 180b, and 180c using signals in unauthorized bandwidths. In non-standalone configurations, WTRU102a, 102b, and 102c may communicate with / connect to gNB180a, 180b, and 180c while simultaneously communicating with / connecting to another RAN, such as eNode-B160a, 160b, and 160c. For example, WTRU102a, 102b, and 102c may implement DC principles to communicate substantially simultaneously with one or more gNB180a, 180b, and 180c and one or more eNode-B160a, 160b, and 160c. In non-standalone configurations, eNode-B160a, 160b, and 160c may act as mobility anchors for WTRU102a, 102b, and 102c, and gNB180a, 180b, and 180c may provide additional coverage and / or throughput to service WTRU102a, 102b, and 102c.

[0109] Each of the gNB180a, 180b, and 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to UPF (User Plane Function) 184a and 184b, and routing of control plane information to AMF (Access and Mobility Management Function) 182a and 182b. As shown in Figure 13D, the gNB180a, 180b, and 180c may communicate with each other via the Xn interface.

[0110] The CN115 shown in Figure 13D may include at least one AMF182a, 182b, at least one UPF184a, 184b, at least one SMF (Session Management Function)183a, 183b, and possibly a DN (Data Network)185a, 185b. While each of the above elements is depicted as part of the CN115, it will be understood that any of the elements described may be owned and / or operated by an entity other than the CN operator.

[0111] AMF182a and 182b may be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N2 interface and may act as control nodes. For example, AMF182a and 182b may be responsible for authenticating users of WTRU102a, 102b, and 102c, supporting network slicing (e.g., handling sessions of separate PDUs with distinct requirements), selecting specific SMF183a and 183b, managing registration areas, terminating NAS signaling, and mobility management. Network slicing may be used by AMF182a and 182b to customize CN support for WTRU102a, 102b, and 102c based on the type of services utilized by WTRU102a, 102b, and 102c. For example, separate network slices may be established for separate use cases, such as services that rely on ultra-high reliability low latency (URLLC) access, services that rely on eMBB (enhanced massive mobile broadband) access, or services related to MTC (machine type communication) access. The AMF162 may provide control plane functions for switching between RAN113 and other RANs (not shown) employing other radio technologies such as LTE, LTE-A, LTE-A Pro and / or non-3GPP access technologies such as WiFi.

[0112] SMF183a and 183b may be connected to AMF182a and 182b in CN115 via the N11 interface. Furthermore, SMF183a and 183b may be connected to UPF184a and 184b in CN115 via the N4 interface. SMF183a and 183b may select and control UPF184a and 184b and configure the routing of traffic passing through them. SMF183a and 183b may perform other functions, such as managing and assigning IP addresses to UEs, managing PDU sessions, controlling policy enforcement and QoS, and providing downlink data notifications. PDU session types may be IP-based, non-IP-based, Ethernet-based, etc.

[0113] UPF184a and 184b may be connected to one or more of gNB180a, 180b, and 180c in RAN113 via the N3 interface, and may provide WTRU102a, 102b, and 102c with access to a packet-switched network, such as the Internet 110, to facilitate communication between WTRU102a, 102b, and 102c and IP-enabled devices. UPF184 and 184b may perform other functions, such as routing and forwarding packets, enforcing user plane policies, supporting sessions for multi-homed PDUs, handling user plane QoS, buffering downlink packets, and providing mobility anchoring.

[0114] CN115 may facilitate communication with other networks. For example, CN115 may include, or communicate with, an IP gateway (e.g., an IMS (IP Multimedia Subsystem) server) that acts as an interface between CN115 and PSTN108. In addition, CN115 may provide WTRU102a,102b,102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, WTRU102a,102b,102c may be connected to local DN185a,185b via UPF184a,184b through an N3 interface to UPF184a,184b and an N6 interface between UPF184a,184b and DN (Data Network) 185a,185b.

[0115] In view of Figures 13A-13D and the descriptions corresponding to Figures 13A-13D, with respect to one or more of the WTRU102a-d, base stations 114a-b, eNode-B160a-c, MME162, SGW164, PGW166, gNB180a-c, AMF182a-b, UPF184a-b, SMF183a-b, DN185a-b, and / or any other device(s) described herein, one or more of the functions described herein may be performed by one or more emulation devices (not shown). An emulation device may be one or more devices configured to emulate one or more of the functions described herein. For example, an emulation device may be used to test other devices and / or to simulate network and / or WTRU functions.

[0116] Emulation devices may be designed to implement one or more tests of other devices in a lab environment and / or an operator's network environment. For example, one or more emulation devices may perform one, more, or all functions while being implemented and / or deployed as part of a wired and / or wireless communication network, either entirely or partially, to test other devices in a communication network. One or more emulation devices may perform one, more, or all functions while being temporarily implemented and / or deployed as part of a wired and / or wireless communication network. Emulation devices may be directly coupled to another device for testing purposes and / or perform testing using over-the-air (OTA) wireless communication.

[0117] One or more emulation devices may perform one or more functions, including all of the above, while not implemented / deployed as part of a wired and / or wireless communication network. For example, an emulation device may be used in a testing scenario in a testing laboratory and / or a wired and / or wireless communication network that is not deployed (e.g., for testing) to implement testing of one or more components. One or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may include, for example, one or more antennas) may be used by an emulation device to transmit and / or receive data.

[0118] The processes and techniques described herein may be implemented in computer programs, software, and / or firmware embedded on computer-readable media for execution by a computer and / or processor. Examples of computer-readable media include, but are not limited to, electrical signals (transmitted via wired and / or wireless connections) and / or computer-readable recording media. Examples of computer-readable recording media include, but are not limited to, ROM (described), RAM (random access memory), registers, cache memory, semiconductor memory devices, magnetic media such as, for example, internal hard disks and removable disks, magneto-optical media, and / or optical media such as, for example, CD-ROM disks and / or DVDs (digital versatile disks). Processors associated with software may be used to implement radio frequency transceivers for use in WTRUs, terminals, base stations, RNCs, and / or any host computer. [Explanation of Symbols]

[0119] 302 video bitstream 308 Entropy Decode Unit 310 Inverse Quantization Unit 312 Inverse Conversion Unit 360 Spatial Prediction Unit 362-hour forecast unit 364 Reference Picture Store 366 Intra-loop filtering

Claims

1. It is determined that combined inter-merge / intra-prediction (CIIP) is enabled for the video block, and the video block includes luminance and saturation components, Using CIIP, predict the luminance component of the video block. Based on the predicted luminance components, the luminance components of the video block are reconstructed. Based on the reconstructed luminance component of the video block, predict the saturation component of the video block. The video block is decoded based on the luminance component and saturation component of the video block. Processors configured in this way A device for video decoding characterized by having the following features.

2. The aforementioned processor, An intra-prediction of the saturation component of the video block is performed based on the reconstructed luminance component of the video block, and the saturation component of the video block is predicted by combining the intra-prediction of the saturation component of the video block and the inter-prediction of the saturation component of the video block. The device according to claim 1, further characterized by being configured as follows.

3. The aforementioned processor, Based on the fact that CIIP is enabled for the video block, intra-prediction and inter-prediction of the luminance component of the video block are performed, and the luminance component of the video block is predicted by combining the intra-prediction and inter-prediction of the luminance component of the video block. Based on the format of the reconstructed luminance component and the saturation component of the video block, a subsampled luminance component is obtained, and the saturation component of the video block is predicted using the subsampled luminance component. The device according to claim 1, further characterized by being configured as follows.

4. The aforementioned processor, It is determined that a cross-component linear model (CCLM) is enabled for the video block, and the saturation component of the video block is predicted based on the determination that the CCLM is enabled for the video block. The device according to claim 1, further characterized by being configured as follows.

5. The aforementioned processor, A saturation prediction index is obtained, which indicates whether a direct mode (DM) or a cross-component linear model (CCLM) is enabled for the video block, and the saturation components of the video block are predicted according to the saturation prediction index. The device according to claim 1, further characterized by being configured as follows.

6. The aforementioned processor, Determine the template for the aforementioned video block, Determine the multiple gradients associated with the sample in the template, Based on the aforementioned multiple gradients, an intra-prediction mode associated with the CIIP is determined, and the luminance component of the video block is predicted using the intra-prediction mode associated with the CIIP. The device according to claim 1, further characterized by being configured as follows.

7. The device according to claim 1, wherein the processor is further configured to determine that a subblock merge mode is enabled for the video block, and the video block is further decoded based on the subblock merge mode.

8. The aforementioned processor, Obtain subblock merge indicators that show affine merge candidates or advanced temporal motion vector prediction (ATMVP) merge candidates. Based on the fact that the subblock merge indication indicates the ATMVP merge candidate, it is decided to obtain the CIIP indication, and the decision that the CIIP is enabled for the video block is based on the CIIP indication. The device according to claim 1, further characterized by being configured as follows.

9. It was decided to enable combined inter-merge / intra-prediction (CIIP) for the video block, and the video block includes luminance and saturation components. Using CIIP, predict the luminance component of the video block. Based on the predicted luminance components, the luminance components of the video block are reconstructed. Based on the reconstructed luminance component of the video block, predict the saturation component of the video block. The video block is encoded based on the luminance component and saturation component of the video block. Processors configured in this way A device for video encoding characterized by having the following features.

10. The aforementioned processor, An intra-prediction of the saturation component of the video block is performed based on the reconstructed luminance component of the video block, and the saturation component of the video block is predicted by combining the intra-prediction of the saturation component of the video block and the inter-prediction of the saturation component of the video block. The device according to claim 9, further characterized by being configured as follows.

11. The aforementioned processor, Based on the fact that CIIP is enabled for the video block, intra-prediction and inter-prediction of the luminance component of the video block are performed, and the luminance component of the video block is predicted by combining the intra-prediction and inter-prediction of the luminance component of the video block. Based on the format of the reconstructed luminance component and the saturation component of the video block, a subsampled luminance component is obtained, and the saturation component of the video block is predicted using the subsampled luminance component. The device according to claim 9, further characterized by being configured as follows.

12. The determination that combined inter-merge / intra-prediction (CIIP) is enabled for a video block, wherein the video block includes luminance and saturation components, Predicting the luminance component of the video block using CIIP, Reconstructing the luminance components of the video block based on the predicted luminance components, Predicting the saturation component of the video block based on the reconstructed luminance component of the video block, Decoding the video block based on the luminance component and saturation component of the video block. A method for video decoding, characterized by comprising:

13. An intra-prediction of the saturation component of the video block is performed based on the reconstructed luminance component of the video block, wherein the saturation component of the video block is predicted by combining the intra-prediction of the saturation component of the video block and the inter-prediction of the saturation component of the video block. The method according to 12, further comprising:

14. Based on the fact that CIIP is enabled for the video block, intra-prediction and inter-prediction of the luminance component of the video block are performed, wherein the luminance component of the video block is predicted by combining the intra-prediction and inter-prediction of the luminance component of the video block. Obtaining a subsampled luminance component based on the format of the reconstructed luminance component and the saturation component of the video block, wherein the saturation component of the video block is predicted using the subsampled luminance component. The method according to 12, further comprising:

15. The determination to enable a cross-component linear model (CCLM) for the video block, wherein the saturation component of the video block is predicted based on the determination that the CCLM is enabled for the video block. The method according to 12, further comprising:

16. The method involves obtaining a saturation prediction index, the saturation prediction index indicating whether a direct mode (DM) or cross-component linear model (CCLM) is enabled for the video block, and the saturation components of the video block are predicted according to the saturation prediction index. The method according to 12, further comprising:

17. Determining the template for the aforementioned video block, Determining multiple gradients associated with the sample in the template, The intra-prediction mode associated with the CIIP is determined based on the plurality of gradients, wherein the luminance component of the video block is predicted using the intra-prediction mode associated with the CIIP. The method according to 12, further comprising:

18. The method according to 12, further comprising determining that a subblock merge mode is enabled for the video block, the video block being decoded based on the subblock merge mode.

19. Obtain subblock merge indicators that show affine merge candidates or advanced temporal motion vector prediction (ATMVP) merge candidates, The decision to obtain a CIIP indication based on the fact that the subblock merge indication indicates the ATMVP merge candidate, wherein the decision to enable the CIIP for the video block is based on the CIIP indication. The method according to 12, further comprising:

20. The decision to enable combined inter-merge / intra-prediction (CIIP) for a video block, wherein the video block includes luminance and saturation components, Predicting the luminance component of the video block using CIIP, Reconstructing the luminance components of the video block based on the predicted luminance components, Predicting the saturation component of the video block based on the reconstructed luminance component of the video block, Encoding the video block based on the luminance component and saturation component of the video block. A method for video encoding characterized by comprising:

21. An intra-prediction of the saturation component of the video block is performed based on the reconstructed luminance component of the video block, wherein the saturation component of the video block is predicted by combining the intra-prediction of the saturation component of the video block and the inter-prediction of the saturation component of the video block. The method according to the 20th invention, further comprising:

22. Based on the fact that CIIP is enabled for the video block, intra-prediction and inter-prediction of the luminance component of the video block are performed, wherein the luminance component of the video block is predicted by combining the intra-prediction and inter-prediction of the luminance component of the video block. Obtaining a subsampled luminance component based on the format of the reconstructed luminance component and the saturation component of the video block, wherein the saturation component of the video block is predicted using the subsampled luminance component. The method according to the 20th invention, further comprising: