Combined Inter and Intra Prediction

The combined inter and intra prediction tool enhances video coding efficiency by applying weighted sample combinations and decoder-side intra-mode derivation, addressing suboptimal coding in mixed object regions.

JP7808670B2Active Publication Date: 2026-01-29INTERDIGITAL VC HOLDINGS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024200862
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-31
Filing Date
2024-11-18
Publication Date
2026-01-29
Estimated Expiration
2039-12-20

AI Technical Summary

Technical Problem

Existing video coding systems struggle to optimally combine inter-prediction and intra-prediction for video regions where old and new objects are mixed, leading to suboptimal coding efficiency.

Method used

Implement a combined inter and intra prediction tool that combines intra-prediction with inter-prediction by applying weights to inter- and intra-predicted samples, using affine motion models, cross-component linear models, and decoder-side intra-mode derivation to enhance coding efficiency.

Benefits of technology

Improves coding efficiency by reducing redundancy in video signals and optimizing prediction modes, particularly in regions with mixed objects, while minimizing bitstream overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007808670000009
    Figure 0007808670000009
  • Figure 0007808670000010
    Figure 0007808670000010
  • Figure 0007808670000011
    Figure 0007808670000011
Patent Text Reader

Abstract

To provide a device for combined inter and intra prediction.SOLUTION: A video coding device may receive a motion vector difference (MMVD) mode indication that indicates whether MMVD mode is used to generate inter prediction of a coding unit (CU). The video coding device may receive a combined inter merge / intra prediction (CIIP) indication, for example, when the MMVD mode indication indicates that MMVD mode is not used to generate the inter prediction of the CU, The video coding device may determine whether to use triangle merge mode for the CU, for example, on the basis of the MMVD mode indication and / or the CIIP indication. On a condition that the CIIP indication indicates that CIIP is applied for the CU or the MMVD mode indication indicates that MMVD mode is used to generate the inter prediction, the video coding device may disable the triangle merge mode for the CU.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to combined inter and intra prediction, and more particularly to combined inter and intra prediction in video processing. [Background technology]

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 786,653, filed December 31, 2018, which is incorporated by reference herein in its entirety.

[0003] Background technology Video coding systems are widely used to compress digital video signals to reduce the storage needs and / or transmission bandwidth of such signals. Various types of video coding systems, such as block-based, wavelet-based, and object-based systems, as well as block-based and hybrid video coding systems, are widely used and developed. Summary of the Invention

[0004] Systems, methods, and means are disclosed for combined inter and intra prediction. A video encoding device may receive a motion vector difference (MMVD) mode indication indicating whether an MMVD mode is used to generate an inter prediction for a coding unit. The video encoding device may receive a combined inter merge / intra prediction (CIIP) indication, for example, when the MMVD mode indication indicates that an MMVD mode is not used to generate an inter prediction for the coding unit. The video encoding device may determine whether to use a triangle merge mode for a coding unit based on, for example, the MMVD mode indication and / or the CIIP indication. The CIIP indication may not be received when the MMVD mode indication indicates that an MMVD mode is used for the coding unit. The MMVD mode indication may be received for each coding unit.

[0005] On condition that the CIIP indication indicates that CIIP applies to the coding unit, the video encoding device may disable triangle merge mode for the coding unit. On condition that the MMVD mode indication indicates that MMVD mode is used to generate inter prediction, the video encoding device may disable triangle merge mode for the coding unit. On condition that the CIIP indication indicates that CIIP does not apply to the coding unit and the MMVD mode indication indicates that MMVD mode is not used to generate inter prediction, the video encoding device may enable triangle merge mode for the coding unit. The video encoding device may enable triangle merge mode for the coding unit without receiving a triangle merge flag. The video encoding device may infer whether to enable triangle merge mode for the coding unit based on one or more of the MMVD mode indication or the CIIP indication. The CIIP indication may not be received when the MMVD mode indication indicates that MMVD mode is used for the coding unit. The MMVD mode indication may be received on a per coding unit basis. [Brief explanation of the drawings]

[0006] [Figure 1] 1 illustrates an example diagram of a block-based video encoder. [Figure 2] FIG. 2(a) illustrates an example associated with a multi-type tree structure and block division into four divisions, FIG. 2(b) illustrates an example associated with a multi-type tree structure and block division into vertical halves, FIG. 2(c) illustrates an example associated with a multi-type tree structure and block division into horizontal halves, FIG. 2(d) illustrates an example associated with a multi-type tree structure and block division into vertical thirds, and FIG. 2(e) illustrates an example associated with a multi-type tree structure and block division into horizontal thirds. [Figure 3] 1 illustrates an example diagram of a block-based video decoder. [Figure 4] Figure 4(a) illustrates an example associated with combined inter and intra prediction and DC plane mode (planar mode), Figure 4(b) illustrates an example associated with combined inter and intra prediction and horizontal mode, and Figure 4(c) illustrates an example associated with combined inter and intra prediction and vertical mode. [Figure 5] 1 illustrates an example associated with selective temporal motion vector prediction. [Figure 6] 1 illustrates an example associated with affine motion field modeling. [Figure 7] FIG. 7A illustrates an example associated with diagonal triangle partitioning-based motion compensated prediction, and FIG. 7B illustrates an example associated with inverse diagonal triangle partitioning-based motion compensated prediction. [Figure 8] 1 illustrates an example associated with generating unidirectional predicted MVs (motion vectors) in triangular mode. [Figure 9] 1 illustrates an example associated with intra-mode derivation. [Figure 10] 10 illustrates an example associated with gradient calculation of template samples for intra-mode derivation. [Figure 11] 1 illustrates an example associated with chrominance prediction for combined inter and intra prediction using a cross-component linear model (CCLM) chrominance prediction sample. [Figure 12] 10 illustrates an example associated with chrominance prediction for combined inter and intra prediction by combining chrominance inter predicted samples and CCLM chrominance predicted samples. [Figure 13A] FIG. 1 is a system diagram of an example communication system in which one or more disclosed aspects may be implemented. [Figure 13B]13B is a system diagram of an example WTRU (Wireless Transmit / Receive Unit) that may be used within the communication system illustrated in FIG. 13A. [Figure 13C] 13B is a system diagram of an exemplary RAN (Radio Access Network) and an exemplary CN (Core Network) that may be used within the communication system illustrated in FIG. 13A. [Figure 13D] FIG. 13B is a system diagram of a further exemplary RAN and a further exemplary CN that may be used within the communication system illustrated in FIG. 13A. DETAILED DESCRIPTION OF THE INVENTION

[0007] Coding tools that may provide high coding efficiency and moderate implementation complexity may include at least one of the following: affine motion model, selective temporal motion vector prediction (alternative temporal motion vector prediction) or advanced temporal motion vector prediction (ATMVP), integer motion vector (IMV), generalized bi-prediction (GBi), bi-directional optical flow (BDOF), combined inter merge / intra prediction (CIIP), merge with motion vector difference (MMVD), pairwise average merge candidate, triangular inter prediction for inter coding, cross-component linear model (CCLM), multi-line intra prediction, current picture referencing (CPR) for intra prediction, enhanced multiple transform (EMT), dependent quantization for quantization and transform coding, and adaptive loop filtering (ALF) for in-loop filtering.

[0008] FIG. 1 shows a block diagram of an exemplary block-based hybrid video encoder system 200. An input video signal 202 may be processed block by block. Extended block sizes (e.g., referred to as coding units or CUs) may be used to compress high-resolution (e.g., 1080p and / or higher) video signals. CUs may include sizes up to 128x128 pixels. Blocks may be divided based on a quadtree. A coding tree unit (CTU) may be divided into CUs based on a quadtree / binary tree / ternary tree to adapt to changing local characteristics. A CU may be divided into prediction units or PUs to which separate predictions may be applied, or may not be divided. A CU may be used as a basic unit for prediction and transformation without further division. In a multi-type tree structure, (e.g., one) CTU may be divided (e.g., first divided) by a quadtree structure. Leaf nodes of the quadtree (e.g., each quadtree leaf node) may be further divided by a binary tree structure and a ternary tree structure. 2, there are one or more (e.g., five) dividing types. One or more of the following may be exemplary dividing types: quarter (e.g., (a)), horizontal bisection (e.g., (c)), vertical bisection (e.g., (b)), horizontal trisection (e.g., (e)), and vertical trisection (e.g., (d)).

[0009] Referring to FIG. 1 , spatial prediction 260 or motion prediction 262 may be performed on an input video block (e.g., a macroblock (MB) or CU). Spatial prediction (e.g., or intra prediction) may use pixels from neighboring blocks already coded in the same video picture and / or slice to predict the current video block. Spatial prediction may reduce spatial redundancy inherent in a video signal. Motion prediction (e.g., referred to as inter prediction or temporal prediction) may use pixels from an already coded video picture to predict the current video block. Motion prediction may reduce temporal redundancy inherent in a video signal. A motion prediction signal for a given video block may be signaled by a motion vector, which indicates the amount and / or direction of motion between the current block and its reference block. If multiple reference pictures are supported, the reference picture index of the video block may be signaled to the decoder. The reference index may be used to identify which reference picture in the reference picture store 464 the temporal prediction signal may come from.

[0010] After spatial prediction and / or motion prediction, a mode decision 280 in the encoder may select a prediction mode based on, for example, rate-distortion optimization. The prediction block may be subtracted from the current video block at 216. The prediction residual may be decorrelated using the transform module 204 and the quantization module 206 to achieve a target bitrate. The quantized residual coefficients may be inverse quantized at 210 and inverse transformed at 212 to form a reconstructed residual. The reconstructed residual may be added back to the prediction block at 226 to form a reconstructed video block. For example, an in-loop filter, such as a deblocking filter and / or an adaptive loop filter, may be applied to the reconstructed video block at 266 before being placed in the reference picture store 264. Reference pictures in the reference picture store 264 may be used to encode future video blocks. An output video bitstream 220 may be formed. The coding mode (e.g., inter or intra), prediction mode information, motion information, and / or quantized residual coefficients may be sent to an entropy coding unit 208 to be compressed and packed to form a bitstream 220.

[0011] Figure 3 shows a general block diagram of an example block-based video encoder. A video bitstream 302 may be received, unpacked, and / or entropy decoded at an entropy decoding unit 308. Coding mode and / or prediction information may be sent to a spatial prediction unit 360 (e.g., if intra-coded) and / or a temporal prediction unit 362 (e.g., if inter-coded). A prediction block may be formed at the spatial prediction unit 360 and / or the temporal prediction unit 362. Residual transform coefficients may be sent to an inverse quantization unit 310 and an inverse transform unit 312 to reconstruct a residual block. The prediction block and the residual block may be added at 326. The reconstructed block may pass through in-loop filtering 366 and may be stored in a reference picture store 364. The reconstructed video in the reference picture store 364 may be used to drive a display device and / or to predict future video blocks.

[0012] One or more coding modules, for example, those associated with inter prediction, may be enhanced to improve inter coding efficiency. One or more coding tools may be described herein.

[0013] A combined inter and intra prediction may be performed.

[0014] As shown in FIGS. 1 and 3, inter-prediction and intra-prediction can be used to exploit temporal and spatial redundancy present in a video signal. In an example, a PU can exploit correlation in the original video in either the temporal or spatial domain. Considering the characteristics of inter-prediction and intra-prediction, the above scheme may not be optimal for certain video content. For example, for a video region where old and new objects are mixed, better coding efficiency may be expected if there is a way to combine inter-prediction and intra-prediction together. Based on the above considerations, a combined inter- and intra-prediction tool may be implemented. The combined inter- and intra-prediction tool may combine intra-prediction with inter-prediction generated from a merge mode. In an example, for each CU coded in merge mode, a flag may be signaled to indicate whether the combined inter- and intra-mode is applied or not. When the flag is true, additional syntax may be signaled, for example, to select an intra-mode from a predefined intra-mode candidate list. For the luma component, the intra mode candidates may include 4-frequently selected intra modes, e.g., planar, DC, horizontal, and vertical. For the chroma component, the DM mode (e.g., indicating that the chroma component reuses the luma intra mode to generate prediction samples) may be applied without any signaling. Weights may be applied to combine the inter- and intra-predicted samples. One or more of the following may be applied:

[0015] An equal weight (eg, 0.5) may be applied to CUs predicted by DC mode or planar mode and CUs with width or height less than or equal to 4.

[0016] If a CU is predicted by either horizontal or vertical mode, and the CU is larger than four samples in width and height, the CU may be divided horizontally or vertically (e.g., depending on the applied intra mode), and the division may be into four equally sized regions. (W_intra i ,Winter i ), where i=0, , 3. In an example, (W_intra0,W_inter0)=(0.75, 0.25), (W_intra1,W_inter1)=(0.625, 0.375), (W_intra2,W_inter2)=(0.375, 0.625), and (W_intra3,W_inter3)=(0.25, 0.75), where (W_intra0,W_inter0) may correspond to the region closest to the reconstructed neighboring sample (e.g., intra reference sample), and (W_intra3,W_inter3) may correspond to the region farthest from the reconstructed neighboring sample. Figure 4 illustrates example weights that may be applied to combine inter- and intra-predicted samples for a combined inter and intra prediction mode. For example, (a) illustrates example weights that may be applied in DC and / or plane mode, (b) illustrates example weights that may be applied in horizontal mode, and (c) illustrates example weights that may be applied in vertical mode. Table 1 shows an example coding unit syntax table after incorporating syntax elements added for combined inter and intra prediction.

[0017] [Table 1]

[0018] With reference to Table 1, one or more of the following may apply: mh_intra_flag[x0][y0] may indicate whether combined inter and intra prediction is applied to the current coding unit. When mh_intra_flag[x0][y0] is not present, it may be inferred to be equal to 0. mh_intra_luma_mpm_flag[x0][y0] and mh_intra_luma_mpm_idx[x0][y0] may indicate the intra prediction mode for luma samples, for example, by invoking a luma intra prediction mode derivation process for an mh intra mode with sample position (x0, y0), the width of the current coding block, the height of the current coding block, mh_intra_luma_mpm_flag[x0][y0], and mh_intra_luma_mpm_idx[x0][y0] as input.

[0019] For example, combined inter and intra prediction mode may be enabled when the current CU is coded with five spatial neighbors and two temporal neighbors, such as those used for a conventional merge mode, e.g., the merge mode of HEVC. For inter CUs that signal MV in the bitstream or that are coded in other merge modes (e.g., affine merge, ATMVP, MMVD, and triangular prediction), combined inter and intra prediction may be disabled.

[0020] A sub-block merge mode may be performed.

[0021] A CU (e.g., each CU) coded by a merge mode may have a set of motion parameters (e.g., one motion vector and one reference picture index) for a prediction direction (e.g., each prediction direction). One or more merge candidates that enable derivation of motion information at the sub-block level may be included in the merge mode. A category of sub-block merge candidates may include alternative temporal motion vector prediction (ATMVP). ATMVP may be built on the same concept of the temporal motion vector prediction (TMVP) tool and may enable a CU to fetch motion information of a sub-block from multiple small blocks from temporally neighboring pictures (e.g., co-located reference pictures). A category of sub-block merge candidates may include an affine merge mode, which may model the motion of sub-blocks inside a CU based on an affine model.

[0022] ATMVP may be performed.

[0023] In ATMVP, temporal motion vector prediction can be improved, for example, by allowing a block to derive motion information (e.g., multiple pieces of motion information including motion vectors and reference indices) for sub-blocks in the current block. The motion information for a sub-block (e.g., each sub-block) can be derived from corresponding small blocks in pictures that are temporally adjacent to the current picture.

[0024] ATMVP may derive motion information for sub-blocks of a block. One or more of the following may apply: The corresponding block of the current block (e.g., sometimes referred to as a collocated block) may be identified in a selected temporal reference picture. The current block may be divided into sub-blocks, and the motion information of each sub-block may be derived from a corresponding small block in the co-located picture, for example, as shown in Figure 5.

[0025] FIG. 5 illustrates an example sub-block motion information derivation 500. The current block may be divided into sub-blocks, and the motion information of each sub-block may be derived from a corresponding small block in a co-located picture, for example, as shown in FIG. 5. The selected temporal reference picture may be referred to as a co-located picture. One or more of the following may apply: The co-located block and the co-located picture may be identified by the motion information of the spatial neighboring blocks of the current block. FIG. 5 illustrates an example associated with ATMVP. With reference to FIG. 5, block A may be identified as the first available merge candidate in the merge candidate list of the current block. The corresponding motion vector (e.g., MV A ) and the reference index can be used to identify the co-located picture and the co-located block. The location of the co-located block in the co-located picture is determined by the motion vector (MV A ) to the coordinates of the current block.

[0026] The current block may be divided into sub-blocks, and motion information for each sub-block may be derived from a corresponding small block in a co-located picture, for example, as shown in Figure 5. In an example, motion information for each sub-block in the current block may be derived from a corresponding small block in a co-located block (e.g., as shown by the red arrows in Figure 5). The motion information of a small block (e.g., each small block in a co-located block) may be identified and converted into a motion vector and reference index of the corresponding sub-block in the current block (e.g., in a manner similar to TMVP in HEVC, where temporal motion vector scaling may be applied).

[0027] An affine model may be used to represent motion information. There may be one or more types of motion present in a video sequence, for example, one or more of the following: translational motion, zoom-in / out, rotational, perspective motion, or other irregular motion. Motion compensation prediction based on affine motion field modeling may be applied. FIG. 6 illustrates an example of affine motion field modeling 600. As shown in FIG. 6, the affine motion field of a block may be described by one or more, for example, three, control point motion vectors. Based on three control-point motion, the motion field of one affine block may be described as follows:

[0028]

number

[0029] Motion vector (v 0x ,v 0y ) can be the motion vector of the upper left corner control point, and the motion vector (v 1x ,v 1y) may be the motion vector of the control point of the upper right corner. When a video block is coded using affine mode, the motion field may be derived based on the granularity of a 4x4 block. To derive the motion vector for each 4x4 block, the motion vector of the center sample of each 4x4 sub-block may be calculated according to (1), which may be expressed in 1 / 16-pel accuracy. The derived motion vector may be used in a motion compensation stage to generate a prediction signal for each sub-block inside the current block.

[0030] Triangle inter prediction may be performed. Figure 7 depicts an exemplary triangle prediction partition 700, 702.

[0031] In some content (e.g., natural video content), the boundary between two moving objects may not be horizontal or vertical (e.g., perfectly horizontal or vertical) and may be difficult to accurately approximate by a rectangular block. Triangular prediction may be applied, for example, to enable triangular division for motion-compensated prediction. As shown in FIG. 7, triangular prediction may divide a CU into one or more (e.g., two) triangular prediction units, for example, in a diagonal or anti-diagonal direction. A triangular prediction unit (e.g., each triangular prediction unit in a CU) may be inter-predicted using its own unidirectional prediction motion vector and reference frame index, which may be derived from a unidirectional prediction candidate list.

[0032] FIG. 8 illustrates an exemplary unidirectionally predicted motion vector candidate derivation 800. The unidirectionally predicted candidate list may include one or more (e.g., five) unidirectionally predicted motion vector candidates. The unidirectionally predicted motion vector candidates may be derived from similar (e.g., the same) spatial / temporal neighboring blocks as those used for the merge process (e.g., the merge process of HEVC). In an example, the unidirectionally predicted MV candidate may be derived from five spatially neighboring blocks and two temporally co-located blocks, as shown in FIG. 8. Referring to FIG. 8, the motion vectors of seven neighboring blocks may be collected and added to the unidirectionally predicted MV candidate in the order of the L0 motion vector of the neighboring block, the L1 motion vector of the neighboring block, and the averaged motion vector of the L0 and L1 motion vectors of the neighboring blocks (e.g., if the neighboring blocks are bidirectionally predicted). If the number of MV candidates is less than five, a zero (0) motion vector is added to the MV candidate list.

[0033] Cross-component prediction may be performed for chrominance intra prediction. Correlation may exist between the luma component and the chrominance component of certain video content (e.g., natural video content). A cross-component linear model (CCLM) prediction mode may be used for chrominance intra prediction. In the CCLM prediction mode, chrominance samples may be predicted from reconstructed luma samples of a block (e.g., the same block) by using a linear model, for example, (2).

[0034]

number

[0035] (2) Refer to pred c (i,j) may denote the prediction of the chrominance sample in the block, pred L(i,j) may denote the reconstructed luma samples of the same block at the same resolution as the chroma block, which may be downsampled for 4:2:0 chroma format content. The parameters α and β may denote the scaling parameter and offset of the linear model, respectively.

[0036] As described herein, for the luma component, one or more (e.g., used up to four times) intra modes, including, for example, plane mode, DC mode, horizontal mode, and vertical mode, may be supported (e.g., by combined inter and intra prediction modes). An encoding device (e.g., an encoder) may test multiple intra modes. The encoding device may select, from the multiple intra modes, the intra mode that provides the best performance (e.g., in terms of rate-distortion tradeoff). The encoding device may signal the selected intra mode (e.g., explicitly signaled to a decoder). For combined inter and intra prediction modes, a non-negligible (e.g., significant) amount of bitrate may be spent on encoding the intra mode. In an example, with increasing computational power, modern devices (e.g., even battery-powered devices such as wireless mobile devices equipped with decoders) may perform some sophisticated operations. The intra mode used for combined inter and intra prediction may be derived on the decoder side. If the decoder-side derivation is accurate, the intra-mode signaling may be skipped, which may improve coding efficiency.

[0037] Combined inter and intra prediction may be enabled when a CU is coded by a merge mode. The merge mode may include using five spatial neighbors and two temporal neighbors for the merge mode (e.g., as shown in FIG. 8). For inter CUs predicted by other merge modes (e.g., affine merge, ATMVP, MMVD, and triangular prediction), combined inter and intra prediction may be disabled (e.g., always enabled). For example, combined inter and intra prediction may be disabled for MMVD because the MMVD mode is primarily selected in true bidirectional prediction scenarios (e.g., with forward and backward predictions from reference lists L0 and L1). Motion-compensated prediction (e.g., inter prediction) may be accurate in predicting the current CU. Additional inter prediction may be unnecessary. Combined inter and intra prediction may be disabled for other merge modes. For one or more merge modes (e.g., ATMVP and triangular prediction modes), the motion of sub-blocks derived from spatial neighbors in the same reference picture or from co-located blocks in a temporal reference picture may not be accurate. In the present case, it may be beneficial (e.g., from the perspective of coding performance) to allow for the combination of combined inter and intra prediction with these merge modes.

[0038] The coding performance of combined inter and intra prediction modes may be improved. Intra mode derivation may skip the overhead of signaling intra modes, for example, by utilizing the computational power of an encoding device (e.g., a decoder). Chroma coding may be improved for combined inter and intra prediction modes. The application of combined inter and intra prediction modes may be extended by one or more coding tools, including, for example, triangular inter prediction and / or sub-block merging modes.

[0039] Combined inter and intra prediction may be performed, for example, by decoder intra mode derivation. For example, in combined inter and intra prediction, a selected intra mode (e.g., plane mode, DC mode, horizontal mode, or vertical mode) of the luma component may be signaled to and / or from an encoding device (e.g., an encoder or decoder). Signaling the selected intra mode may add a non-negligible portion of the bitstream (e.g., output bit-steam) and / or reduce overall coding performance. The intra mode used for combined inter and intra prediction may be derived at the encoding device (e.g., a decoder), which may reduce overhead. In other words, the decoder may derive the intra mode for combined inter and intra prediction (e.g., based on one or more neighboring reconstructed samples). For example, when combined inter and intra prediction is applied to a CU (eg, instead of directly signaling the intra mode in the bitstream), the intra mode may be derived from neighboring reconstructed samples of the CU.

[0040] FIG. 9 illustrates an example intra mode derivation 900 (e.g., decoder-side derivation of intra modes for combined inter and intra prediction). As shown in FIG. 9, the current CU to which combined inter and intra prediction is applied may have a size of N×N (width×height). A template may specify a set of reconstructed samples used to derive the intra mode of the current CU. The shaded region in FIG. 9 may represent the template. The template size may be indicated as the number of samples, e.g., L, within an L-shaped region extending above and to the left of the target block. For example, gradient analysis may be applied on the template samples when there is a strong correlation between the samples of the current CU and neighboring blocks (e.g., the template). Gradient analysis may be used to estimate the intra mode to be used for combined inter and intra prediction of the current CU. In an example, as shown in (3), a 3×3 Sobel filter may be applied to calculate the horizontal and vertical gradients of the template samples.

[0041]

number

[0042] Two matrices (say M x and M y ) may be multiplied with a 3×3 window of samples (e.g., samples centered on the template samples). Two gradient values, G x and G y and the activity value Act h can be calculated by summing the horizontal and vertical gradients of each template sample, for example using (4).

[0043]

number

[0044] (4)

[0045]

number

[0046] and

[0047]

number

[0048] may denote the horizontal and vertical gradients of the template sample at coordinate (i,j), respectively. Ω may denote the set of coordinates of the samples in the template. FIG. 10 illustrates an example gradient calculation 1000 of a template sample for intra-mode derivation. Given the gradient and activity values ​​calculated in (4), the final intra-mode (e.g., the final intra-mode applied to the current CU) may be determined by (5).

[0049]

number

[0050] (5) g and th act may indicate two predefined thresholds for the gradient and activity values ​​that may be fixed and used in an encoding device (e.g., an encoder and / or a decoder). In an example, the thresholds for the gradient and activity values ​​may be determined by an encoding device (e.g., an encoder) and signaled in the bitstream to another encoding device (e.g., a decoder).

[0051] In FIG. 9 , reconstructed samples forming an L-shaped region (e.g., reconstructed samples closest to the current CU) may be used as a template. In an example, templates having different shapes and / or sizes may be selected, which may provide different complexity / performance tradeoffs. Selecting a large template size may result in the template samples being far away from the target block. Correlation between the template and the target block may be poor. A large template size may increase the complexity of encoding and decoding (e.g., given that more samples may be considered for gradient and activity calculation). A large template size may result in less reliable estimation when noise is present. In an example, an adaptive template size may be used for intra-mode estimation. For example, the template size may be determined based on the CU size. The template size may be represented by the value of “L” in FIG. 9 . In an example, a first template size may be used for a certain CU size (e.g., less than 64 samples), and a second template size may be used for other CU sizes (e.g., CUs with 64 samples or more). The first template size may be 3 (e.g., L=3 as shown in FIG. 9). The second template size may be 5 (e.g., L=5 as shown in FIG. 9).

[0052] The decoder may select a luma intra mode for combined inter and intra prediction. The luma intra mode for combined inter and intra prediction may be selected, for example, by minimizing the difference between the inter-predicted samples and the intra-predicted samples. The luma intra mode for combined inter and intra prediction may be selected by calculating a measured cost between the inter-predicted signal and the intra-predicted signal (e.g., for each intra prediction mode). One or more of the following cost measures (e.g., template cost measures) may be applied, such as, for example, sum of absolute difference (SAD), sum of square difference (SSD), and / or sum of absolute transformed difference (SATD). Based on the cost comparison, the intra prediction mode that results in the smallest template cost may be selected as the intra prediction mode for the current CU (e.g., the best intra prediction mode for the current CU). In an example, the intra prediction mode associated with the lowest measured cost may be selected (e.g., by the decoder).

[0053] Combined inter and intra prediction may be performed using chroma CCLM. In combined inter and intra prediction, direct mode (DM) may be applied to one or more chroma components, for example, without signaling. In DM, the luma component intra mode may be reused for one or more chroma components. Inter-channel correlation may exist between the luma component and the chroma component. For example, the luma component may have inter-channel correlation with the chroma component. The chroma component may have inter-channel correlation with the luma component. CCLM may be used for chroma intra prediction. For example, CCLM may be used for chroma intra prediction in which chroma samples are predicted from corresponding luma samples with subsampling applied based on a linear model. One or more CCLM model parameters may be derived from one or more neighboring luma and / or chroma samples (e.g., perfunctory neighboring luma and chroma samples around the current CU). The DM mode may be replaced with a CCLM mode for intra prediction of chrominance samples, for example, when combined inter and intra prediction is applied to a CU. In other words, the CCLM mode may be applied to one or more chrominance components (e.g., instead of the DM mode).

[0054] The CCLM mode may be applied to the chroma component in one or more of the following ways: In an example, one or more luma prediction samples of a CU may be generated by mixing prediction samples generated from intra prediction (e.g., based on an intra mode signaled in the bitstream) and inter prediction (e.g., based on the MV of a neighboring block indicated by a merge box). One or more luma prediction samples and residual samples of the luma component may be added together to generate (e.g., form), for example, one or more reconstructed luma samples. One or more reconstructed luma samples may be subsampled. One or more reconstructed luma samples may be used to generate one or more chroma prediction samples based on the CCLM mode.

[0055] In an example, one or more chrominance prediction samples (e.g., CCLM prediction samples) may be combined with the chrominance inter prediction samples to generate a final prediction sample for the chrominance component. The one or more chrominance prediction samples may be referred to as chrominance inter prediction samples.

[0056] 11 illustrates an exemplary chrominance prediction 1100 for combined inter and intra prediction using (e.g., directly using) CCLM chrominance prediction samples. In a first exemplary method of applying CCLM mode to chrominance components (e.g., as shown in FIG. 11), the chrominance LM prediction samples may be directly used.

[0057] 12 illustrates an example chrominance prediction 1200 for combined inter and intra prediction by combining chrominance inter predicted samples and CCLM chrominance predicted samples. In a second example method of applying CCLM mode to the chrominance component (e.g., as shown in FIG. 12), chrominance intra predicted samples may be combined with chrominance inter predicted samples.

[0058] DM and CCLM modes may be enabled for chroma intra prediction of combined inter and intra prediction. A flag may be signaled, for example, when combined inter and intra prediction is enabled for the current CU. The flag may indicate whether DM mode or CCLM mode is applied. One or more of the following may be applied: When the flag is set to true, DM mode may be applied to the chroma components, for example, such that the same intra mode of the chroma components is reused to generate chroma intra predicted samples. When the flag is set to false, CCLM mode may be applied to intra prediction of chroma samples, for example, chroma inter predicted samples will be generated based on a linear mode from subsampled luma reconstructed samples.

[0059] The combined inter and intra prediction may interact with other coding tools. The combined inter and intra prediction may be enabled when a CU is coded by a merge mode (e.g., a conventional merge mode). The combined inter and intra prediction mode may interact with one or more other inter coding tools (e.g., including affine merge, ATMVP, and triangular prediction). The interaction(s) (e.g., synergy of interactions) between the combined inter and intra prediction mode and one or more other inter coding tools may be improved. One or more of the following may apply:

[0060] Combined inter and intra prediction may be enabled for a sub-block merge mode. Motion information may be derived at the sub-block level. For example, sub-block derivation of motion information may be enabled through a sub-block merge mode, such as ATMVP and affine mode. In an example, combined inter and intra prediction may be disabled (e.g., for a CU predicted by a sub-block mode). Sub-block motion information may be derived (e.g., fully derived) from one or more spatial neighbors (e.g., in affine merge mode). Sub-block motion information may be derived from one or more temporal neighbors (e.g., in ATMVP mode). The sub-block motion information may not be accurate enough to generate a predicted sample for the current CU. Combined inter and intra prediction may be enabled for a sub-block merge mode. In an example, combined inter and intra prediction may be enabled (e.g., simply enabled) for ATMVP mode and disabled (e.g., always disabled) for affine merge mode. The combined inter and intra prediction flag may be signaled, for example, to enable combined inter and intra prediction. The combined inter and intra prediction flag may be signaled after the subblock merge mode index. When the subblock merge index references an affine candidate, the combined inter and intra prediction flag may be skipped. When the subblock merge index references an affine candidate, the combined inter and intra prediction flag may be inferred to be false, e.g., disabling combined inter and intra prediction. If the subblock merge index references an ATMVP candidate, the combined inter / intra flag may be signaled to indicate whether combined inter and intra prediction is enabled for the current CU.

[0061] Combined inter and intra prediction may be enabled for a triangular mode (e.g., triangular merge mode). In triangular prediction, a PU (e.g., each PU) of a current CU may be inter predicted using MVs (e.g., motion vector candidates) from unidirectional prediction candidates for the current CU. MV candidates in a unidirectional prediction candidate list may be derived from spatially and / or temporally neighboring blocks of the current CU, which may not be accurate enough to account for the motion of the CU. Combined inter and intra prediction may be, for example, for a triangular prediction mode that improves the efficiency of triangular inter prediction. An intra prediction mode (e.g., a single intra prediction mode) may be signaled for one or more (e.g., two) PUs. The intra prediction mode may be signaled separately for each PU, for example.

[0062] The triangular mode may be disabled for combined inter and intra prediction and / or MMVD mode. The triangular mode may be referred to as the triangular merge mode. The triangular merge mode may be enabled when CIIP is not applied and MMVD mode is not used. For example, a video encoding device (e.g., an encoder configured as shown in FIG. 1) may determine whether to enable the triangular merge mode. When the encoder determines to enable the triangular merge mode, the encoder may determine not to apply CIIP and not to use the MMVD mode. For example, the video encoding device may enable the triangular merge mode without receiving a triangular merge mode indication (e.g., a triangular merge mode flag). A video encoding device (e.g., a decoder configured as shown in FIG. 3) may receive an MMVD mode indication (e.g., an MMVD flag). For example, the encoder may send the MMVD mode indication to the video encoding device. The video encoding device may be a WTRU (e.g., such as WTRU 102 shown in FIGS. 13A-13D). The video encoding device may include a WTRU (e.g., such as WTRU 102 shown in FIGS. 13A-13D). The MMVD mode indication may indicate whether an MMVD mode is used to generate inter prediction for a CU. The MMVD mode indication may be received on a per coding unit basis. The video encoding device may receive a combine inter merge / intra prediction (CIIP) indication (e.g., a CIIP flag). For example, an encoder may send a CIIP indication to the video encoding device. The CIIP indication may not be received when the MMVD mode indication indicates that an MMVD mode is used for a CU. The CIIP indication may indicate that CIIP applies to the CU.

[0063] In an example, the triangle merge flag may be signaled, for example, after the sub-block merge flag, the combined inter / intra flag, and / or the MMVD flag. If one or more of the above flags are set to true, the triangle flag may not be signaled. When the CIIP indication indicates that CIIP applies to the CU, the triangle merge flag may not be signaled. For example, the encoder may be determined not to include the triangle merge flag in a message to a video encoding device (e.g., a decoder) when the CIIP indication indicates that CIIP does not apply to the CU. When the MMVD mode indication indicates that an MMVD mode is used to generate inter prediction, the triangle merge flag may not be signaled. For example, the encoder may be determined not to include the triangle merge flag in a message to a video encoding device (e.g., a decoder) when the MMVD mode indication indicates that an MMVD mode is used to generate inter prediction. When the triangle merge flag is not signaled, the triangle merge flag may be inferred to be false (e.g., disabling triangle mode). For example, a video encoding device may infer whether to enable triangle merge mode for a CU based on the MMVD mode indication and / or the CIIP indication. The video encoding device may disable triangle merge mode for a CU, for example, when the MMVD mode indication indicates that an MMVD mode is used to generate inter prediction. The video encoding device may disable triangle merge mode for a CU, for example, when the CIIP indication indicates that CIIP is applied to the CU. If the three flags are set to false, a triangle flag may be signaled to indicate, for example, whether triangle mode is applied to the current CU.

[0064] 13A illustrates an example communication system 100 in which one or more disclosed aspects may be implemented. The communication system 100 may be a multiple-access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communication system 100 may enable the multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communication system 100 may employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal frequency division multiple access (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail unique-word DFT-Spread OFDM (ZT UW DTS-s OFDM), unique word OFDM (UW-OFDM), resource block-filtered OFDM, filter bank multicarrier (FBMC), etc.

[0065] 13A, communications system 100 may include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, RANs 104 / 113, CNs 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, although it will be understood that the disclosed aspects contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of WTRUs 102a, 102b, 102c, 102d may be any type of device configured to operate and / or communicate in a wireless environment. By way of example, the WTRUs 102a, 102b, 102c, 102d may all be referred to as “stations” and / or “STAs,” may be configured to transmit and / or receive wireless signals, and may include user equipment (UE), mobile stations, fixed or mobile subscriber units, subscription-based units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, hotspots or Mi-Fi devices, Internet of Things (IoT) devices, watches or other wearables, head-mounted displays (HMDs), vehicles, drones, medical devices and applications (e.g., remote surgery), industrial devices and applications (e.g., robots and / or other wireless devices operating in the context of industrial and / or automated processing chains), consumer electronics devices, devices operating on commercial and / or industrial wireless networks, etc. Any of the WTRUs 102a, 102b, 102c, and 102d may be referred to interchangeably as a UE.

[0066] Additionally, the communications system 100 may include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b may be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communications networks, such as the CN 106 / 115, the Internet 110, and / or other networks 112. By way of example, the base stations 114a, 114b may be a base transceiver station (BTS), a Node-B, an eNodeB, a home Node B, a home eNodeB, a gNB, a NR NodeB, a site controller, an access point (AP), a wireless router, etc. While the base stations 114a, 114b are each depicted as a single element, it will be understood that the base stations 114a, 114b may include any number of interconnected base stations and / or network elements.

[0067] The base station 114a may be part of the RAN 104 / 113, which may further include other base stations and / or network elements (not shown), such as, for example, a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base station 114a and / or base station 114b may be configured to transmit and / or receive wireless signals on one or more carrier frequencies, which may be referred to as a cell (not shown). The frequencies may be licensed spectrum, unlicensed spectrum, or a combination of licensed and unlicensed spectrum. A cell may provide coverage for wireless services over a particular geographic area, which may be relatively fixed or may change over time. Furthermore, a cell may be divided into cell sectors. For example, the cell associated with the base station 114a may be divided into three sectors. Thus, in one aspect, the base station 114a may include three transceivers, i.e., one for each sector of the cell. In an aspect, the base station 114a may employ multiple-input multiple output (MIMO) technology and may utilize multiple transceivers for each sector of the cell. For example, beamforming may be used to transmit and / or receive signals in a desired spatial direction.

[0068] The base stations 114a, 114b may communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 may be established using any suitable radio access technology (RAT).

[0069] More specifically, as mentioned above, the communication system 100 may be a multiple-access system and may employ one or more channel access schemes, such as, for example, CDMA, TDMA, FDMA, OFDMA, SC-FDMA, etc. For example, the base station 114a in the RAN 104 / 113 and the WTRUs 102a, 102b, 102c may implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which may establish the air interface 115 / 116 / 117 using, for example, wideband CDMA (WCDMA). WCDMA may include communication protocols such as High-Speed ​​Packet Access (HSPA) and / or Evolved HSPA+ (HSPA+). HSPA may include High-Speed ​​Downlink Packet Access (HSDPA) and / or High-Speed ​​Ultra-Low Packet Access (HSUPA).

[0070] In an aspect, the base station 114a and the WTRUs 102a, 102b, 102c may establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro), and may implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA).

[0071] In an aspect, the base station 114a and the WTRUs 102a, 102b, 102c may establish the air interface 116 using New Radio (NR) and may implement a radio technology such as NR radio access.

[0072] In an aspect, the base station 114a and the WTRUs 102a, 102b, 102c may implement multiple radio access technologies. For example, the base station 114a and the WTRUs 102a, 102b, 102c may implement both LTE radio access and NR radio access, e.g., using dual connectivity (DC) principles. Thus, the air interface utilized by the WTRUs 102a, 102b, 102c may be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., eNBs and gNBs).

[0073] In other aspects, the base station 114a and the WTRUs 102a, 102b, 102c may implement a wireless technology such as, for example, IEEE 802.11 (i.e., Wireless Fidelity (WiFi)), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GERAN (GSM EDGE), etc.

[0074] 13A may be, for example, a wireless router, a Home Node B, a Home eNode B, or an access point and may utilize any suitable RAT to facilitate wireless connectivity in a local area such as, for example, a business, a home, a vehicle, a campus, an industrial facility, an air corridor (e.g., for use by drones), a roadway, etc. In one aspect, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as, for example, IEEE 802.11 to establish a wireless local area network (WLAN). In an aspect, the base station 114b and the WTRUs 102c, 102d may implement a radio technology such as, for example, IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another aspect, the base station 114b and the WTRUs 102c, 102d may utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or a femtocell. As shown in FIG. 13A, the base station 114b may have a direct connection to the Internet 110. Thus, the base station 114b may not be required to access the Internet 110 via the CN 106 / 115.

[0075] The RAN 104 / 113 may be in communication with the CN 106 / 115 and may be any type of network configured to provide voice, data, application, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data may have various quality of service (QoS) requirements, such as, for example, different throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, etc. The CN 106 / 115 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as, for example, user authentication. 13A, it will be understood that the RAN 104 / 113 and / or the CN 106 / 115 may be in direct or indirect communication with other RANs employing the same RAT as the RAN 104 / 113 or a different RAT. For example, the CN 106 / 115, in addition to being connected to the RAN 104 / 113, which may utilize NR radio technology, may also be in communication with another RAN (not shown) employing GSM, UMTS, CDMA2000, WiMAX, E-UTRA, or WiFi radio technology.

[0076] Additionally, the CN 106 / 115 may serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or other networks 112. The PSTN 108 may include a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 110 may include a global system of interconnected computer networks and devices that use common communication protocols, such as transmission control protocol (TCP), user datagram protocol (UDP), and / or IP in the TCP / IP Internet Protocol suite. The network 112 may include wired and / or wireless communication networks owned and / or operated by other service providers. For example, the network 112 may include another CN connected to one or more RANs, which may employ the same RAT as the RAN 104 / 113 or a different RAT.

[0077] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 may include multi-mode capabilities (e.g., the WTRUs 102a, 102b, 102c, 102d may include multiple transceivers for communicating with separate wireless networks via separate wireless links). For example, the WTRU 102c shown in FIG. 13A may be configured to communicate with a base station 114a, which may employ a cellular-based radio technology, and with a base station 114b, which may employ an IEEE 802.11 radio technology.

[0078] 13B is a system diagram illustrating an example WTRU 102. As shown in FIG. 13B, the WTRU 102 may include, among other things, a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a GPS (Global Positioning System) chipset 136, and / or other peripherals 138. It will be understood that the WTRU 102 may include any sub-combination of the above elements without departing from the spirit and scope of the present invention.

[0079] The processor 118 may be a general-purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any of other types of integrated circuits (ICs), a state machine, etc. The processor 118 may perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 may be coupled to a transceiver 120, which may be coupled to a transmit / receive element 122. While FIG. 13B depicts the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 may be integrated together in an electronic package or chip.

[0080] The transmit / receive element 122 may be configured to transmit or receive signals to a base station (e.g., base station 114a) over the air interface 116. For example, in one aspect, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In an aspect, the transmit / receive element 122 may be an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another aspect, the transmit / receive element 122 may be configured to transmit and / or receive both RF and light signals. It will be understood that the transmit / receive element 122 may be configured to transmit and / or receive any combination of wireless signals.

[0081] 13B as a single element, the WTRU 102 may include any number of transmit / receive elements 122. More specifically, the WTRU 102 may employ MIMO techniques. Thus, in one aspect, the WTRU 102 may include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.

[0082] The transceiver 120 may be configured to modulate signals to be transmitted by the transmit / receive element 122 and to demodulate signals received by the transmit / receive element 122. As mentioned above, the WTRU 102 may have multi-mode capabilities. Thus, for example, the transceiver 120 may include multiple transceivers to enable the WTRU 102 to communicate over multiple RATs, such as NR and IEEE 802.11.

[0083] The processor 118 of the WTRU 102 may be coupled to and may receive user input data through a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit). Further, the processor 118 may output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. Additionally, the processor 118 may access information and store data in any suitable type of memory, such as non-removable memory 130 and / or removable memory 132. The non-removable memory 130 may include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other aspects, the processor 118 may access information and store data in memory that is not physically located in the WTRU 102, such as in a server or host computer (not shown).

[0084] The processor 118 may receive power from the power source 134 and may be configured to distribute and / or control power to other components in the WTRU 102. The power source 134 may be any device suitable for providing power to the WTRU 102. For example, the power source 134 may include one or more dry batteries (e.g., NiCd (nickel cadmium), NiZn (nickel zinc), NiMH (nickel metal hydride), Li-ion (lithium ion), etc.), solar cells, fuel cells, etc.

[0085] Additionally, the processor 118 may be coupled to a GPS chipset 136 that may be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to or instead of information from the GPS chipset 136, the WTRU 102 may receive location information from a base station (e.g., base stations 114a, 114b) over the air interface 116 and / or determine its location based on the timing of signals being received from two or more neighboring base stations. It will be understood that the WTRU 102 may obtain location information through any suitable location-determination method without departing from the spirit or scope of the present invention.

[0086] Additionally, the processor 118 may be coupled to other peripherals 138, which may include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 138 may include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos and / or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulated FM (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, a virtual reality and / or augmented reality (VR / AR) device, an activity tracker, etc. The peripheral device 138 may include one or more sensors, which may be one or more of a gyroscope, an accelerometer, a Hall effect sensor, a magnetometer, a direction sensor, a proximity sensor, a temperature sensor, a time sensor, a geolocation sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.

[0087] The WTRU 102 may include a full-duplex radio where transmission and reception of some or all of the signals (e.g., associated with a particular subframe for both the UL (e.g., for transmission) and downlink (e.g., for reception)) may be parallel and / or simultaneous. The full-duplex radio may include an interference management unit to reduce and or substantially eliminate self-interference either through hardware (e.g., a choke) or signal processing by a processor (e.g., a separate processor (not shown) or by processor 118). In an aspect, the WTRU 102 may include a half-duplex radio where transmission and reception of some or all of the signals (e.g., associated with a particular subframe for either the UL (e.g., for transmission) or downlink (e.g., for reception)) may be half-duplex.

[0088] 13C is a system diagram illustrating the RAN 104 and the CN 106 according to an aspect. As mentioned above, the RAN 104 may employ E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. Additionally, the RAN 104 may be in communication with the CN 106.

[0089] The RAN 104 may include eNode-Bs 160a, 160b, and 160c, although it will be understood that the RAN 104 may include any number of eNode-Bs while remaining consistent with an aspect. The eNode-Bs 160a, 160b, and 160c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one aspect, the eNode-Bs 160a, 160b, and 160c may implement MIMO technology. Thus, for example, the eNode-B 160a may use multiple antennas to transmit and / or receive wireless signals to the WTRU 102a.

[0090] Each of the eNode-Bs 160a, 160b, 160c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, etc. As shown in FIG. 13C, the eNode-Bs 160a, 160b, 160c may communicate with each other via an X2 interface.

[0091] 13C may include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the above elements is depicted as part of the CN 106, it will be understood that any of the just-mentioned elements may be owned and / or operated by an entity other than the CN operator.

[0092] The MME 162 may be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an S1 interface and may act as a control node. For example, the MME 162 may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, activating / deactivating bearers, selecting a particular serving gateway during initial attachment of the WTRUs 102a, 102b, 102c, etc. The MME 162 may provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.

[0093] The SGW 164 may be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via an S1 interface. In general, the SGW 164 may route and forward user data packets to the WTRUs 102a, 102b, 102c. The SGW 164 may perform other functions such as, for example, anchoring the user plane during inter-eNode B handovers, triggering paging when DL data is available to the WTRUs 102a, 102b, 102c, managing and storing the context of the WTRUs 102a, 102b, 102c, etc.

[0094] The SGW 164 may be connected to a PGW 166, which may provide the WTRUs 102a, 102b, 102c with access to a packet-switched network, such as the Internet 110, to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices.

[0095] The CN 106 may facilitate communication with other networks. For example, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to a circuit-switched network, such as the PSTN 108, to facilitate communication between the WTRUs 102a, 102b, 102c and traditional landline communication devices. For example, the CN 106 may include or communicate with an IP gateway (e.g., an IMS (IP Multimedia Subsystem) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers.

[0096] Although the WTRU is described in Figures 13A-13D as a wireless terminal, it is expected that in certain representative aspects, such a terminal may use a wired communication interface with the communication network (e.g., temporarily or permanently).

[0097] In an exemplary aspect, the other network 112 may be a WLAN.

[0098] A WLAN in infrastructure Basic Service Set (BSS) mode may have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP may have access or an interface to a distribution system (DS) or another type of wired / wireless network that carries traffic into and out of the BSS. Traffic to a STA originating from outside the BSS may arrive through the AP and be delivered to the STA. Traffic originating from a STA to a destination outside the BSS may be sent to the AP and delivered to the respective destination. Traffic between STAs within a BSS may be sent through the AP, for example, where a source STA may send traffic to the AP, and the AP may deliver the traffic to the destination STA. Traffic between STAs within a BSS may be considered and / or referred to as peer-to-peer traffic. Peer-to-peer traffic may be sent between (e.g., directly between) a source and destination STA via a direct link setup (DLS). In certain exemplary aspects, the DLS may use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an Independent BSS (IBSS) mode may not have an AP, and STAs within or using the IBSS (e.g., all STAs) may communicate directly with each other. The IBSS mode of communication may sometimes be referred to herein as an "ad hoc" mode of communication.

[0099] When using the 802.11ac infrastructure mode of operation or a similar mode of operation, an AP may transmit beacons on a fixed channel, such as a primary channel. The primary channel may be a fixed width (e.g., a 20 MHz wide bandwidth) or a width that is dynamically set by signaling. The primary channel may be the operating channel of the BSS and may be used by STAs to establish a connection with the AP. In one exemplary aspect, CSMA / CA (Carrier Sense Multiple Access / Collision Avoidance) may be implemented in an 802.11 system, for example. With CSMA / CA, STAs (e.g., every STA), including the AP, may sense the primary channel. If the primary channel is sensed / detected and / or determined to be busy by a particular STA, the particular STA may back off. One STA (e.g., only one station) may transmit in a given BSS at any given time.

[0100] For example, a HT (high throughput) STA may use a 40 MHz wide channel for communication by combining a 20 MHz primary channel with adjacent or non-adjacent 20 MHz channels to form a 40 MHz wide channel.

[0101] A Very High Throughput (VHT) STA may support channels of 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz width. 40 MHz and / or 80 MHz channels may be constructed by combining contiguous 20 MHz channels. A 160 MHz channel may be constructed by combining eight contiguous 20 MHz channels or by combining two non-contiguous 80 MHz channels, which may result in an 80+80 configuration. For the 80+80 configuration, after channel encoding, the data may pass through a segment parser, which may split the data into two streams. IFFT (inverse fast Fourier transform) processing and time-domain processing may be performed on each stream separately. The streams may be mapped onto two 80 MHz channels, and the data may be transmitted by the transmitting STA. At the receiver of the receiving STA, the operations for the 80+80 configuration described above may be reversed and the combined data may be sent to the MAC (Media Access Control).

[0102] Sub-1 GHz modes of operation are supported by 802.11af and 802.11ah. The operating bandwidths of the channels and carriers are reduced in 802.11af and 802.11ah compared to those used in 802.11n and 802.11ac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidths in the TVWS (TV White Space) spectrum, while 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidths using non-TVWS spectrum. According to a representative aspect, 802.11ah may support Meter Type Control / Machine-Type Communication, such as for MTC devices in macro coverage areas. MTC devices may have limited capabilities, including support (e.g., only support) for some and / or limited bandwidths. An MTC device may include a battery with a battery life that exceeds a threshold (eg, to maintain a very long battery life).

[0103] A WLAN system that may support multiple channels and channel bandwidths, such as 802.11n, 802.11ac, 802.11af, and 802.11ah, includes a channel that may be designated as a primary channel. The primary channel may have a bandwidth equal to the largest common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel may be set and / or limited by the STA that supports the smallest bandwidth operating mode among all STAs operating in the BSS. In the 802.11ah example, the primary channel may be 1 MHz wide for a STA (e.g., an MTC-type device) that supports (e.g., only supports) the 1 MHz mode, even if the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes. Carrier sensing and / or NAV (Network Allocation Vector) setting may depend on the state of the primary channel. If the primary channel is busy, for example, due to a STA (that only supports a 1 MHz mode of operation) transmitting to the AP, the entire available frequency band may be considered busy even though most of the frequency band may remain idle and available.

[0104] In the United States, the available frequency bands that may be used by 802.11ah are 902MHz to 928MHz. In South Korea, the available frequency bands are 917.5MHz to 923.5MHz. In Japan, the available frequency bands are 916.5MHz to 927.5MHz. The total bandwidth available for 802.11ah is 6MHz to 26MHz, depending on the country code.

[0105] 13D is a system diagram illustrating the RAN 113 and the CN 115 according to an aspect. As mentioned above, the RAN 113 may employ NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. Additionally, the RAN 113 may be in communication with the CN 115.

[0106] The RAN 113 may include gNBs 180a, 180b, and 180c, although it will be understood that the RAN 113 may include any number of gNBs without being inconsistent with an aspect. The gNBs 180a, 180b, and 180c may each include one or more transceivers for communicating with the WTRUs 102a, 102b, and 102c over the air interface 116. In one aspect, the gNBs 180a, 180b, and 180c may implement MIMO techniques. For example, the gNBs 180a, 180b may utilize beamforming to transmit and / or receive signals to the gNBs 180a, 180b, and 180c. Thus, for example, the gNB 180a may use multiple antennas to transmit and / or receive wireless signals to the WTRU 102a. In an aspect, the gNBs 180a, 180b, and 180c may implement carrier aggregation techniques. For example, the gNB 180a may transmit multiple component carriers to the WTRU 102a (not shown). A subset of the aforementioned component carriers may be on a non-licensed spectrum, while the remaining component carriers may be on a licensed spectrum. In an aspect, the gNBs 180a, 180b, and 180c may implement Coordinated Multi-Point (CoMP) techniques. For example, the WTRU 102a may receive coordinated transmissions from the gNBs 180a and 180b (and / or 180c).

[0107] The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using transmissions associated with scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing may vary for separate transmissions, separate cells, and / or separate portions of the wireless transmission spectrum. The WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., including various numbers of OFDM symbols and / or various persistent absolute time lengths).

[0108] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing another RAN (e.g., eNode-Bs 160a, 160b, 160c, etc.). In a standalone configuration, the WTRUs 102a, 102b, 102c may utilize one or more of the gNBs 180a, 180b, 180c as mobility anchor points. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in unlicensed bands. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect with a gNB 180a, 180b, 180c while also communicating / connecting with another RAN, such as, for example, an eNode-B 160a, 160b, 160c. For example, the WTRUs 102a, 102b, 102c may implement DC principles to communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously. In a non-standalone configuration, the eNode-Bs 160a, 160b, 160c may act as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput in serving the WTRUs 102a, 102b, 102c.

[0109] Each of the gNBs 180a, 180b, 180c may be associated with a particular cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, support for network slicing, dual connectivity, interworking between NR and E-UTRA, routing of user plane data to User Plane Functions (UPFs) 184a, 184b, routing of control plane information to Access and Mobility Management Functions (AMFs) 182a, 182b, etc. As shown in FIG. 13D, the gNBs 180a, 180b, 180c may communicate with each other via an Xn interface.

[0110] The CN 115 shown in Figure 13D may include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one SMF (Session Management Function) 183a, 183b, and possibly a DN (Data Network) 185a, 185b. While each of the above elements is depicted as part of the CN 115, it will be understood that any of the just-mentioned elements may be owned and / or operated by an entity other than the CN operator.

[0111] The AMF 182a, 182b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N2 interface and may act as a control node. For example, the AMF 182a, 182b may be responsible for authenticating users of the WTRUs 102a, 102b, 102c, supporting network slicing (e.g., handling sessions of separate PDUs with separate requirements), selecting a particular SMF 183a, 183b, managing registration areas, terminating NAS signaling, mobility management, etc. Network slicing may be used by the AMF 182a, 182b to customize CN support for the WTRUs 102a, 102b, 102c based on the type of service being utilized for the WTRUs 102a, 102b, 102c. For example, separate network slices may be established for separate use cases, such as services relying on Ultra-Reliable Low-Latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, services related to machine-type communications (MTC) access, etc. The AMF 162 may provide control plane functionality for switching between the RAN 113 and other RANs (not shown) employing other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.

[0112] The SMFs 183a and 183b may be connected to the AMFs 182a and 182b in the CN 115 via an N11 interface. Furthermore, the SMFs 183a and 183b may be connected to the UPFs 184a and 184b in the CN 115 via an N4 interface. The SMFs 183a and 183b may select and control the UPFs 184a and 184b and configure the routing of traffic through the UPFs 184a and 184b. The SMFs 183a and 183b may perform other functions, such as managing and assigning UE IP addresses, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notification, etc. The PDU session type may be IP-based, non-IP-based, Ethernet-based, etc.

[0113] The UPFs 184a, 184b may be connected to one or more of the gNBs 180a, 180b, 180c in the RAN 113 via an N3 interface and may provide the WTRUs 102a, 102b, 102c with access to a packet-switched network such as the Internet 110 to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPFs 184, 184b may perform other functions such as, for example, routing and forwarding packets, enforcing user plane policies, supporting multi-homed PDU sessions, handling user plane QoS, buffering downlink packets, providing mobility anchoring, etc.

[0114] The CN 115 may facilitate communication with other networks. For example, the CN 115 may include or communicate with an IP gateway (e.g., an IMS (IP Multimedia Subsystem) server) that acts as an interface between the CN 115 and the PSTN 108. In addition, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one aspect, the WTRUs 102a, 102b, 102c may be connected to local DNs (data networks) 185a, 185b through an N3 interface to the UPFs 184a, 184b, and an N6 interface between the UPFs 184a, 184b and the DNs 185a, 185b.

[0115] 13A-13D and the corresponding descriptions thereof, one or more or all of the functions described herein in association with one or more of the WTRUs 102a-d, base stations 114a-b, eNode-Bs 160a-c, MME 162, SGW 164, PGW 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other device(s) described herein may be performed by one or more emulation devices (not shown). The emulation devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, the emulation devices may be used to test other devices and / or to simulate network and / or WTRU functionality.

[0116] The emulation device may be designed to implement one or more tests of other devices in a lab environment and / or in an operator's network environment. For example, one or more emulation devices may perform one or more, or all, functions while fully or partially implemented and / or deployed as part of a wired and / or wireless communications network to test other devices in the communications network. One or more emulation devices may perform one or more, or all, functions while temporarily implemented / deployed as part of a wired and / or wireless communications network. The emulation device may be directly coupled to another device for testing purposes and / or may perform testing using over-the-air (OTA) wireless communications.

[0117] The one or more emulation devices may perform one or more functions, inclusive, while not being implemented / deployed as part of a wired and / or wireless communications network. For example, the emulation devices may be utilized in testing laboratories and / or testing scenarios in undeployed (e.g., testing) wired and / or wireless communications networks to implement testing of one or more components. The one or more emulation devices may be test equipment. Direct RF coupling and / or wireless communication via RF circuitry (which may, for example, include one or more antennas) may be used by the emulation devices to transmit and / or receive data.

[0118] The processes and techniques described herein may be implemented in a computer program, software, and / or firmware embodied in a computer-readable medium for execution by a computer and / or processor. Examples of computer-readable media include, but are not limited to, electrical signals (transmitted over wired and / or wireless connections) and / or computer-readable recording media. Examples of computer-readable recording media include, but are not limited to, ROM (described), RAM (random access memory), registers, cache memory, semiconductor memory devices, magnetic media such as, but not limited to, internal hard disks and removable disks, magneto-optical media, and / or optical media such as, for example, CD-ROM disks and / or digital versatile disks (DVDs). A processor associated with software may be used to implement a radio frequency transceiver for use in a WTRU, a terminal, a base station, an RNC, and / or any host computer. [Explanation of symbols]

[0119] 302 Video Bitstream 308 Entropy Decoding Unit 310 Inverse Quantization Unit 312 Inverse Conversion Unit 360 Spatial Prediction Unit 362 Time Prediction Units 364 Reference Picture Store 366 In-Loop Filtering

Claims

1. obtaining a combined inter-merge-intra prediction (CIIP) indication that indicates whether CIIP is applied to the video block; determining, based on the CIIP indication, that a CIIP is applied to the video block; determining a merge mode is enabled for a plurality of video blocks including the video block, the merge mode enabling a triangulation of the video block to be predicted using unidirectional predictive merge candidates associated with the video block; determining to disable the merge mode for the video block based on determining that CIIP applies to the video block; Decoding the video block based on disabling the merge mode for the video block. Processor configured as 1. A device for video decoding, comprising:

2. The device of claim 1 , wherein the merge mode is a triangle merge mode.

3. 10. The device of claim 1 , wherein the unidirectional predictive merge candidate associated with the video block is a first unidirectional predictive merge candidate associated with the video block, the triangulation is a first partition of the video block, the video block including the first partition and a second partition to be predicted using a second unidirectional predictive merge candidate associated with the video block.

4. The device of claim 3 , wherein the first unidirectional predictive merge candidate and the second unidirectional predictive merge candidate are different.

5. The device of claim 3 , wherein the processor is further configured to obtain a unidirectional prediction candidate list including the first unidirectional prediction merge candidate and the second unidirectional prediction merge candidate.

6. 6. The device of claim 5, wherein the processor is further configured to generate a unidirectional prediction candidate list based on a plurality of video blocks spatially adjacent to the video block and a plurality of video blocks temporally co-located with the video block.

7. 10. The device of claim 1, wherein the processor is further configured to determine, based on the determination to disable the merge mode for the video block, that no indication for the merge mode is signaled in video data.

8. Obtaining a video block, determining that combined inter-merge-intra prediction (CIIP) is applied to the video block; determining a merge mode is enabled for a plurality of video blocks including the video block, the merge mode enabling a triangulation of the video block to be predicted using unidirectional predictive merge candidates associated with the video block; determining to disable the merge mode for the video block based on determining that CIIP applies to the video block; encoding the video block based on disabling the merge mode for the video block; Processor configured as 1. A device for video encoding, comprising:

9. The device of claim 8 , wherein the merge mode is a triangle merge mode.

10. 9. The device of claim 8, wherein the unidirectional predictive merge candidate associated with the video block is a first unidirectional predictive merge candidate associated with the video block, the triangulation is a first partition of the video block, the video block including the first partition and a second partition to be predicted using a second unidirectional predictive merge candidate associated with the video block.

11. obtaining a combined inter-merge-intra prediction (CIIP) indication that indicates whether CIIP is applied to the video block; determining, based on the CIIP indication, that CIIP is applied to the video block; determining a merge mode is valid for a plurality of video blocks including the video block, the merge mode enabling a triangulation of the video block to be predicted using uni-predictive merge candidates associated with the video block; determining, based on determining that CIIP applies to the video block, to disable the merge mode for the video block; decoding the video block based on disabling the merge mode for the video block; 1. A method for video decoding, comprising:

12. The method of claim 11 , wherein the merge mode is a triangle merge mode.

13. 12. The method of claim 11 , wherein the unidirectionally predictive merge candidate associated with the video block is a first unidirectionally predictive merge candidate associated with the video block, the triangulation is a first partition of the video block, and the video block includes the first partition and a second partition that will be predicted using a second unidirectionally predictive merge candidate associated with the video block.

14. A method for obtaining a video block, comprising: determining that combined inter-merge-intra prediction (CIIP) is applied to the video block; determining a merge mode is valid for a plurality of video blocks including the video block, the merge mode enabling a triangulation of the video block to be predicted using uni-predictive merge candidates associated with the video block; determining, based on determining that CIIP applies to the video block, to disable the merge mode for the video block; encoding the video block based on disabling the merge mode for the video block; 1. A method for video encoding, comprising:

15. 15. The method of claim 14, wherein the unidirectionally predictive merge candidate associated with the video block is a first unidirectionally predictive merge candidate associated with the video block, the triangulation is a first partition of the video block, and the video block includes the first partition and a second partition that will be predicted using a second unidirectionally predictive merge candidate associated with the video block.

Citation Information

Patent Citations

  • Method and apparatus for video coding

    US20200154101A1

  • Image signal encoding / decoding method and apparatus therefor

    WO2020096389A1

  • Method and apparatus for video coding

    WO2020117619A1