Combined intra- and inter-frame prediction
By combining inter-frame and intra-frame prediction tools, incorporating motion vector differences and combined inter-frame merging/intra-frame prediction indicators, the coding efficiency of the video coding system in mixed video content areas is optimized, solving the problem of poor combination of inter-frame prediction and intra-frame prediction in the existing technology, and achieving more efficient coding performance.
Patent Information
- Application Number
- CN201980086741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-31
- Filing Date
- 2019-12-20
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2039-12-20
AI Technical Summary
Existing video coding systems fail to achieve optimal coding efficiency when processing mixed video content through a combination of inter-frame prediction and intra-frame prediction, especially in mixed areas of old and newly appeared objects, resulting in poor coding performance.
A combined inter-frame and intra-frame prediction tool is used. By receiving the motion vector difference (MMVD) mode indication and the combined inter-frame merging/intra-frame prediction (CIIP) indication, the triangle merging mode is dynamically enabled or disabled, and the combination of affine motion model, alternative temporal motion vector prediction, cross-component linear model (CCLM) and other technologies are combined to optimize the combination of intra-frame and inter-frame prediction.
The coding efficiency of video coding is improved, the coding overhead is reduced, and the coding performance is improved, especially the coding quality in mixed video content areas.
Smart Images

Figure CN113228634B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 786,653, filed on December 31, 2018, the contents of which are incorporated herein by reference in their entirety. Background Art
[0003] Video coding systems are widely used to compress digital video signals to reduce the storage requirements and / or transmission bandwidth of such signals. Among various types of video coding systems (e.g., block-based systems, wavelet-based systems, and object-based systems), block-based hybrid video coding systems are widely used and deployed. Summary of the Invention
[0004] Disclosed are systems, methods, and means for combined inter- and intra-frame prediction. A video encoding device may receive a motion vector difference (MMVD) mode indication, the MMVD mode indicating whether to use MMVD mode to generate an inter-frame prediction for a coding unit. For example, when the MMVD mode indication indicates that MMVD mode is not used to generate the inter-frame prediction for the coding unit, the video encoding device may receive a combined inter-merging / intra-frame prediction (CIIP) indication. The video encoding device may determine whether to use triangle merge mode for the coding unit based on the MMVD mode indication and / or the CIIP indication. When the MMVD mode indicates that MMVD mode is used for the coding unit, the CIIP indication may not be received. The MMVD mode indication may be received for each coding unit.
[0005] If the CIIP indicates that CIIP is applied to the coding unit, the video encoding device may disable the triangle merge mode for the coding unit. If the MMVD mode indication indicates that MMVD mode is used to generate the inter-frame prediction, the video encoding device may disable the triangle merge mode for the coding unit. If the CIIP indication indicates that CIIP is not applied to the coding unit and the MMVD mode indicates that MMVD mode is not used to generate the inter-frame prediction, the video encoding device may enable the triangle merge mode for the coding unit. The video encoding device may enable the triangle merge mode for the coding unit without receiving a triangle merge flag. The video encoding device may infer whether the triangle merge mode is enabled for the coding unit based on one or more of the MMVD mode indication or the CIIP indication. When the MMVD mode indication indicates that MMVD mode is used for the coding unit, the CIIP indication may not be received. The MMVD mode indication may be received on a per-coding unit basis. BRIEF DESCRIPTION OF DRAWINGS
[0006] Figure 1 An exemplary diagram of a block-based video encoder is shown.
[0007] Figure 2 Examples associated with block partitioning and quad partitioning in a multi-type tree structure, examples associated with block partitioning and vertical binary partitioning in a multi-type tree structure, examples associated with block partitioning and horizontal binary partitioning in a multi-type tree structure, examples associated with block partitioning and vertical ternary partitioning in a multi-type tree structure, and examples associated with block partitioning and horizontal ternary partitioning in a multi-type tree structure are shown.
[0008] Figure 3 An exemplary diagram of a block-based video decoder is shown.
[0009] Figure 4 Examples associated with combined inter and intra prediction and DC plane mode, examples associated with combined inter and intra prediction and horizontal mode, and examples associated with combined inter and intra prediction and vertical mode are shown.
[0010] Figure 5 Examples associated with alternative temporal motion vector prediction are shown.
[0011] Figure 6 Examples associated with affine motion field modeling are shown.
[0012] Figure 7 Examples associated with motion compensation prediction based on diagonal triangle partitioning and inverse diagonal triangle partitioning are shown.
[0013] Figure 8 Examples associated with generating uni-prediction motion vector (MV) in triangle mode are shown.
[0014] Figure 9 Examples associated with intra mode derivation are shown.
[0015] Figure 10 Examples associated with gradient computation of template samples for intra mode derivation are shown.
[0016] Figure 11 Examples associated with chroma prediction for combined inter and intra prediction using cross-component linear model (CCLM) chroma prediction samples are shown.
[0017] Figure 12An example is shown that is associated with chroma prediction for combined inter and intra prediction, which combines chroma inter prediction samples and CCLM chroma prediction samples.
[0018] Figure 13A is a system diagram of an example communication system in which one or more disclosed embodiments can be implemented.
[0019] Figure 13B is a system diagram of an example wireless transmit / receive unit (WTRU) that can be used within the communications system Figure 13A illustrated in FIG. 1.
[0020] Figure 13C is a system diagram of an example radio access network (RAN) and an example core network (CN) that can be used within the communications system Figure 13A illustrated in FIG. 1.
[0021] Figure 13D is a system diagram of another example RAN and another example CN that can be used within the communications system Figure 13A illustrated in FIG. 1. DETAILED DESCRIPTION
[0022] Encoding tools that can provide higher coding efficiency and moderate implementation complexity can include one or more of the following: affine motion model, alternative temporal motion vector prediction or advanced temporal motion vector prediction (ATMVP), integer motion vector (IMV), general bi-prediction (GBi), bi-directional optical flow (BDOF), combined inter prediction / intra prediction (CIIP) with motion vector difference (MMVD), pair-wise average merge candidate, triangle inter prediction for inter coding; cross component linear model (CCLM), multi-line intra prediction, current picture reference (CPR) for intra prediction; enhanced multiple transform (EMT), correlated quantization for quantization and transform coding, and adaptive loop filtering (ALF) for in-loop filters.
[0023] Figure 1A block-based hybrid video coding system 200 is shown. Input video signals 202 can be processed on a block-by-block basis. Extended block sizes (e.g., referred to as coding units or CUs) can be used to compress high resolution (e.g., 1080p and / or above) video signals. CUs can contain sizes up to 128x128 pixels. Blocks can be partitioned based on quad-trees. Coding tree units (CTUs) can be partitioned into multiple CUs based on quad / binary / triple-trees to adapt to varying local characteristics. CUs can or can not be partitioned into prediction units or PUs that can apply separate predictions. CUs can be used as the basic unit for prediction and transform without further partitioning. In a multi-type tree structure, a CTU (e.g., one) can be partitioned (e.g., can be partitioned first) by a quad-tree structure. Quad-tree leaf nodes (e.g., each quad-tree leaf node) can be further partitioned by binary and ternary tree structures. As shown in Figure 2 There can be one or more (e.g., five) partition types as shown in the following. One or more of the following can be example partition types: quad-partition (e.g., (a)), horizontal binary partition (e.g., (c)), vertical binary partition (e.g., (b)), horizontal ternary partition (e.g., (e)), and vertical ternary partition (e.g., (d)).
[0024] Referring to Figure 1 For an input video block (e.g., a macroblock (MB) or a CU), spatial prediction 260 or motion prediction 262 can be performed. Spatial prediction (e.g., or intra prediction) can use pixels from coded neighboring blocks in the same video picture and / or slice to predict the current video block. Spatial prediction can reduce spatial redundancy inherent in the video signal. Motion prediction (e.g., referred to as inter prediction or temporal prediction) can use pixels from coded video pictures to predict the current video block. Motion prediction can reduce temporal redundancy inherent in the video signal. The motion prediction signal for a given video block can be signaled by a motion vector that indicates the amount and / or direction of motion between the current block and its reference block. If multiple reference pictures are supported, a reference picture index for a video block can be signaled to the decoder. The reference index can be used to identify which reference picture in the reference picture store 464 the temporal prediction signal can come from.
[0025] After spatial and / or motion prediction, mode decision 280 in the encoder can select a prediction mode, e.g., based on rate-distortion optimization. At 216, the prediction block can be subtracted from the current video block. The prediction residual can be de-correlated using transform module 204 and quantization module 206 to achieve a target bitrate. The quantized residual coefficients can be inverse quantized at 210 and inverse transformed at 212 to form a reconstructed residual. At 226, the reconstructed residual can be added back to the prediction block to form a reconstructed video block. At 266, in-loop filters such as a deblocking filter and / or an adaptive loop filter can be applied to the reconstructed video block before it is placed into the reference picture store 264. The reference pictures in the reference picture store 264 can be used to encode future video blocks. An output video bitstream 220 can be formed. The encoding mode (e.g., inter or intra), prediction mode information, motion information, and / or quantized residual coefficients can be sent to entropy encoding unit 208 to be compressed and packed to form the bitstream 220.
[0026] Figure 3 A general block diagram of an example block-based video decoder is shown. A video bitstream 302 can be received, unpacked, and / or entropy decoded at entropy decoding unit 308. The encoding mode and / or prediction information can be sent to a spatial prediction unit 360 (e.g., if intra coded) and / or to a temporal prediction unit 362 (e.g., if inter coded). A prediction block can be formed by spatial prediction unit 360 and / or temporal prediction unit 362. Residual transform coefficients can be sent to inverse quantization unit 310 and inverse transform unit 312 to reconstruct a residual block. The prediction block and residual block can be added at 326. The reconstructed block can undergo in-loop filtering 366 and can be stored in a reference picture store 364. The reconstructed video in the reference picture store 364 can be used to drive a display device and / or to predict future video blocks.
[0027] One or more encoding modules, e.g., associated with inter prediction, can be enhanced to improve inter coding efficiency. One or more inter coding tools can be described herein.
[0028] Combined inter and intra prediction can be performed.
[0029] As Figure 1 and 3As shown, inter prediction and intra prediction can be used to exploit the temporal and spatial redundancy present in a video signal. In an example, a PU can exploit the correlation of the original video in the temporal domain or spatial domain. Considering the characteristics of inter prediction and intra prediction, such a scheme can not be optimal for certain video content. For example, for a video region with a mix of old objects and newly appearing objects, better coding efficiency can be expected if there is a way to combine inter prediction and intra prediction together. Based on such a consideration, a combined inter and intra prediction tool can be performed. The combined inter and intra prediction tool can combine intra prediction with inter prediction generated from merge mode. In an example, for each CU coded in merge mode, a flag can be signaled to indicate whether the combined inter and intra mode is applied. When the flag is true, additional syntax can be signaled, e.g., to select an intra mode from a predefined list of intra mode candidates. For luma component, the intra mode candidates can include 4 frequently selected intra modes, e.g., planar, DC, horizontal, and vertical. For chroma component, DM mode (e.g., indicating that the chroma component reuses the luma intra mode to generate its prediction samples) can be applied without any signaling. Weights can be applied to combine the inter prediction samples and the intra prediction samples. One or more of the following can be applied.
[0030] For CUs predicted by DC or planar mode and CUs with width or height less than or equal to 4, equal weights, e.g., 0.5, can be applied.
[0031] If a CU is predicted by horizontal or vertical mode and the CU is larger than 4 samples in width and height, the CU can be split in the horizontal or vertical direction (e.g., depending on the applied intra mode); the split can be divided into four equal size regions. One weight combination, denoted as (w_intra i ,w_inter i ), where i = 0, …, 3. In an example, (w_intra0, w_inter0) = (0.75, 0.25), (w_intra1, w_inter1) = (0.625, 0.375), (w_intra2, w_inter2) = (0.375, 0.625), and (w_intra3, w_inter3) = (0.25, 0.75), which (w_intra0, w_inter0) can correspond to the region closest to the reconstructed neighboring samples (e.g., intra reference samples), and (w_intra3, w_inter3) can correspond to the region farthest from the reconstructed neighboring samples. Figure 4Some exemplary weights are shown that may be applied to combine inter-predicted samples and intra-predicted samples for the combined inter- and intra-prediction modes. For example, (a) shows exemplary weights that may be applied in DC and / or planar modes; (b) shows example weights that may be applied in horizontal mode; and (c) shows example weights that may be applied in vertical mode. Table 1 shows an example coding unit syntax table after incorporating additional syntax elements for combined inter- and intra-prediction.
[0032] Table 1 has coding unit syntax for combined inter and intra prediction.
[0033]
[0034] With reference to Table 1, one or more of the following may apply. mh_intra_flag[x0][y0] may indicate whether combined inter and intra prediction is applied for the current coding unit. When mh_intra_flag[x0][y0] is not present, it may be inferred to be equal to 0. mh_intra_luma_mpm_flag[x0][y0] and mh_intra_luma_mpm_idx[x0][y0] may indicate the intra prediction mode for luma samples, for example, by invoking the derivation process for luma intra prediction mode for MH intra mode with sample position (x0, y0), width of the current coding block, height of the current coding block, mh_intra_luma_mpm_flag[x0][y0], and mh_intra_luma_mpm_idx[x0][y0] as input.
[0035] For example, when the current CU is encoded by conventional merge mode (e.g., five spatial and two temporal neighbors of merge mode for HEVC), the combined inter and intra prediction mode may be enabled. For inter CUs that signal the MV in the bitstream or are encoded in other merge modes (e.g., affine merge, ATMVP, MMVD, and triangle prediction), the combined inter and intra prediction may be disabled.
[0036] A sub-block merging mode may be performed.
[0037] A CU (e.g., each CU) coded by the merge mode can have a set of motion parameters (e.g., one motion vector and one reference picture index) for each prediction direction. One or more merge candidates that enable the derivation of motion information at a sub-block level can be included in the merge mode. Categories of sub-block merge candidates can include an alternative temporal motion vector prediction (ATMVP). ATMVP can be built based on the same concept of temporal motion vector prediction (TMVP) tool and can allow a CU to extract its sub-blocks' motion information from multiple small blocks from its temporally neighboring pictures (e.g., collocated reference pictures). Categories of sub-block merge candidates can include an affine merge mode, which can model the motion of sub-blocks within a CU based on an affine model.
[0038] ATMVP can be performed.
[0039] In ATMVP, for example, temporal motion vector prediction can be improved by allowing a block to derive motion information (e.g., multiple motion information, including motion vectors and reference indices) for sub-blocks in a current block. Motion information for each sub-block can be derived from a corresponding small block of a temporally neighboring picture of the current picture.
[0040] ATMVP can derive motion information for sub-blocks of a block. One or more of the following can be applied. A corresponding block (e.g., which can be referred to as a collocated block) of the current block can be identified in a selected temporal reference picture. The current block can be partitioned into sub-blocks, and motion information for each sub-block can be derived from a corresponding small block in the collocated picture, e.g., as shown in Figure 5 .
[0041] Figure 5 An example sub-block motion information derivation 500 is depicted. The current block can be partitioned into sub-blocks, and motion information for each sub-block can be derived from a corresponding small block in the collocated picture, e.g., as shown in Figure 5 . The selected temporal reference picture can be referred to as a collocated picture. One or more of the following can be applied. The collocated block and the collocated picture can be identified by the motion information of the spatial neighboring blocks of the current block. Figure 5 An example associated with ATMVP is shown. Referring to Figure 5 , block A can be identified as the first available merge candidate in a merge candidate list of a current block. The corresponding motion vector (e.g., MV A ) and its reference index of block A can be used to identify the collocated picture and the collocated block. The position of the collocated block in the collocated picture can be determined by adding the motion vector (MV A ) of block A to the coordinates of the current block.
[0042] The current block can be partitioned into multiple sub-blocks, and the motion information of each sub-block can be derived from the corresponding sub-block in the collocated picture, e.g., as Figure 5 indicated by the red arrows in Figure 5 In an example, the motion information of each sub-block in the current block can be derived from its corresponding sub-block in the collocated block (e.g., as indicated by the red arrows in
[0043] An affine model can be used to indicate the motion information. There can be one or more types of motion in a video sequence, e.g., one or more of the following: translational motion, zoom-in / zoom-out, rotation, perspective motion, or other irregular motion. Motion compensated prediction based on affine motion field modeling can be applied. Figure 6 An example of affine motion field modeling 600 is shown. As Figure 6 indicated, the affine motion field of the block can be described by one or more (e.g., three) control point motion vectors. Based on the three control point motions, the motion field of one affine block can be described as:
[0044]
[0045] The motion vector (v 0x ,v 0y ) can be the motion vector of the top-left control point, and the motion vector (v 1x ,v 1y ) can be the motion vector of the top-right control point. When a video block is coded by the affine mode, its motion field can be derived based on the granularity of 4x4 blocks. To derive the motion vector of each 4x4 block, the motion vector of the center sample of each 4x4 sub-block can be calculated according to (1). It can be rounded to 1 / 16-pixel precision. The derived motion vector can be used in the motion compensation stage to generate the prediction signal for each sub-block inside the current block.
[0046] Triangular inter prediction can be performed. Figure 7 Example triangular prediction partitions 700, 702 are depicted.
[0047] In some content (e.g., natural video content), the boundary between two moving objects can not be horizontal or vertical (e.g., purely horizontal or vertical), which can be difficult to approximate accurately by rectangular blocks. Triangular prediction can be applied, e.g., in order to enable triangular partitioning for motion compensated prediction. As Figure 7As shown in FIG. 1, triangle prediction can partition a CU into one or more (e.g., two) triangle prediction units, e.g., in a diagonal or anti-diagonal direction. A triangle prediction unit (e.g., each triangle prediction unit in a CU) can be inter predicted using its own uni-prediction motion vector and reference frame index, which can be derived from a uni-prediction candidate list.
[0048] Figure 8 An example uni-prediction motion vector candidate derivation 800 is depicted. The uni-prediction candidate list can include one or more (e.g., five) uni-prediction motion vector candidates. The uni-prediction motion vector candidates can be derived from spatial / temporal neighboring blocks similar (e.g., identical) to those used for the merge process (e.g., the merge process of HEVC). In some examples, the uni-prediction MV candidates can be derived from Figure 8 the five spatial neighboring blocks and two temporally collocated blocks as shown. See Figure 8 The motion vectors of these seven neighboring blocks can be collected and placed in the uni-prediction MV candidate list in the following order: the L0 motion vector of the neighboring block, the Ll motion vector of the neighboring block, and the average motion vector of the L0 motion vector and the Ll motion vector of the neighboring block (e.g., if the neighboring block is bi-predicted). If the number of MV candidates is less than five, a zero (0) motion vector is added to the MV candidate list.
[0049] Cross-component prediction for chroma intra prediction can be performed. Certain video content (e.g., natural video content) can have a correlation between the luma component and the chroma component. A cross-component linear model (CCLM) prediction mode can be used for chroma intra prediction. In the CCLM mode prediction mode, chroma samples can be predicted from the reconstructed luma samples of a block (e.g., the same block) by using a linear model (e.g., (2)).
[0050] pred C (i,j) = a · rec L (i,j) + b (2)
[0051] Referring to (2), pred C (i,j) can indicate a prediction of the chroma samples in a block, and rec L (i,j) can indicate reconstructed luma samples of the same block at the same resolution as the chroma block, which can be down-sampled for 4:2:0 chroma format content. The parameters a and b can indicate a scaling parameter and an offset of the linear model, respectively.
[0052] As described herein, for luma components, one or more (e.g., up to four frequently used) intra modes may be supported (e.g., via combined inter and intra prediction modes), including, for example, planar, DC, horizontal, and vertical modes. An encoding device (e.g., an encoder) may examine multiple intra modes. The encoding device may select the intra mode that provides the best performance (e.g., in terms of rate-distortion tradeoff) among the multiple intra modes. The encoding device may signal (e.g., explicitly signal to a decoder) the selected intra mode. For combined inter and intra prediction modes, a non-negligible (e.g., significant) amount of bit rate may be expended on encoding the intra mode. In examples, due to increased computing power, modern devices (e.g., even battery-powered devices, such as wireless mobile devices equipped with a decoder) may perform some complex operations. An intra mode for combined inter and intra prediction may be derived for use on the decoder side. If the decoder-side derivation is accurate, signaling of the intra mode may be skipped, and coding efficiency may be improved.
[0053] When encoding a CU in merge mode, combined inter and intra prediction may be enabled. Merge mode may include using five spatial and two temporal neighbors for the merge mode (e.g., Figure 8 ). For inter CUs predicted by other merge modes (e.g., affine merging, ATMVP, MMVD, and triangle prediction), the combined inter and intra prediction may be disabled (e.g., may always be disabled). For example, since the MMVD mode is most often selected in true bi-prediction scenarios (e.g., there are forward and backward predictions from reference lists L0 and L1), the combined inter and intra prediction for MMVD may be disabled. Motion compensated prediction (e.g., inter prediction) can accurately predict the current CU. Additional intra prediction may not be necessary. For other merge modes, combined inter and intra prediction may not be disabled. For one or more merge modes (e.g., ATMVP and triangle prediction mode), sub-block motion derived from spatially neighboring blocks in the same reference picture or collocated blocks in a temporal reference picture may be inaccurate. In those cases, enabling a combination of combined inter and intra prediction with those merge modes may be beneficial (e.g., in terms of coding performance).
[0054] The coding performance of the combined inter- and intra-prediction modes may be improved. Intra-mode derivation may skip the overhead of signaling the intra-mode, for example by leveraging the computational power of an encoding device (e.g., a decoder). Chroma coding for the combined inter- and intra-prediction modes may be improved. The application of the combined inter- and intra-prediction modes may be extended with one or more coding tools, including, for example, triangular inter-prediction and / or sub-block merging modes.
[0055] For example, combined inter and intra prediction can be performed with decoder intra mode derivation. In combined inter and intra prediction, for example, a selected intra mode (e.g., planar, DC, horizontal, or vertical mode) for a luma component can be signaled to and / or from an encoding device (e.g., an encoder or a decoder). Signaling the selected intra mode can occupy a non-negligible portion of a bitstream (e.g., an output bitstream) and / or can reduce overall encoding performance. The intra mode for the combined inter and intra prediction can be derived at an encoding device (e.g., a decoder), which can reduce overhead. In other words, the decoder can derive the intra mode for the combined inter and intra prediction (e.g., based on one or more neighboring reconstructed samples). When the combined inter and intra prediction is applied to a CU (e.g., instead of directly signaling the intra mode in a bitstream), the intra mode can be derived, for example, from neighboring reconstructed samples of the CU.
[0056] Figure 9 An example intra mode derivation 900 (e.g., decoder-side derivation of the intra mode for the combined inter and intra mode) is shown. As shown in (1) of Figure 9 A current CU to which the combined inter and intra mode is applied can include a size of NxN (width x height), as shown in (2) of Figure 9 The shaded area in (3) can represent a template. The template size can be represented as the number of samples within an L-shaped area extending to the top and left side of the target block, e.g., L. For example, when there is a strong correlation between the samples of the current CU and its neighboring blocks (e.g., the template), a gradient analysis can be applied on the top of the template samples. The gradient analysis can be used to estimate the intra mode for the combined inter and intra prediction of the current CU. In an example, as shown in (3), a 3x3 Sobel filter can be applied to calculate the horizontal and vertical gradients of the samples in the template.
[0057]
[0058] Two matrices (e.g., M x and M y ) can be multiplied by a 3x3 window of samples (e.g., samples centered at the template samples). Two gradient values G x and G y and an activity value Act h may be calculated by summing the horizontal and vertical gradients at each template sample, which can be performed, for example, by using (4).
[0059]
[0060] Refer to (4), and The horizontal gradient and vertical gradient of the template sample at the coordinate (i, j) can be indicated respectively. Ω can indicate the coordinate set of the sample in the template. Figure 10 An example gradient calculation of a template sample for intra mode derivation is shown 1000. Given the gradient value and activity value calculated in (4), the final intra mode (eg, the final intra mode applied to the current CU) can be determined by (5).
[0061]
[0062] Reference (5), th g and th act Two predefined thresholds for gradient and activity values may be indicated, which may be fixed and used at an encoding device (e.g., an encoder and / or a decoder). In an example, the thresholds for gradient and activity values may be determined by an encoding device (e.g., an encoder) and signaled in a bitstream to another encoding device (e.g., a decoder).
[0063] exist Figure 9 In the example, reconstructed samples forming an L-shaped area (e.g., the reconstructed samples closest to the current CU) can be used as a template. In the example, templates with different shapes and / or sizes can be selected, which can provide different complexity / performance trade-offs. When a large template size is selected, the samples of the template may be far away from the target block. The correlation between the template and the target block may not be sufficient. A large template size may increase the encoding and decoding complexity (e.g., assuming that more samples can be considered for gradient and activity calculations). When noise is present, a large template size can produce reliable estimates. In the example, an adaptive template size can be used for intra-frame mode estimation. For example, the template size can be determined based on the CU size. The template size can be Figure 9 In an example, a first template size may be used for certain CU sizes (e.g., less than 64 samples), and a second template size may be used for other CU sizes (e.g., CUs with greater than or equal to 64 samples). The first template size may be 3 (e.g., Figure 9 The second template size may be 5 (e.g., Figure 9 L=5 shown).
[0064] The decoder can select a luma intra mode for the combined inter and intra mode. For example, the luma intra mode for the combined inter and intra mode can be selected by minimizing a difference between inter-predicted samples and intra-predicted samples. The intra mode selection for the combined inter and intra mode can include computing a cost measured between the inter-predicted signal and the intra-predicted signal (e.g., for each intra-prediction mode). One or more of the following cost measurements (e.g., template cost measurements) can be applied, such as sum of absolute difference (SAD), sum of squared difference (SSD), and / or sum of absolute transformed difference (SATD). Based on the comparison of the costs, the intra-prediction mode that generates the smallest template cost can be selected as the intra-prediction mode for the current CU (e.g., the best intra-prediction mode for the current CU). In an example, the intra-prediction mode associated with the lowest measured cost can be selected (e.g., by the decoder).
[0065] The combined inter and intra prediction can be performed with the chroma CCLM mode. In the combined inter and intra prediction, a direct mode (DM) can be applied to, for example, one or more chroma components without signaling. In the DM, the luma component intra mode can be reused for one or more chroma components. There can be an inter-channel correlation between the luma component and the chroma component. For example, the luma component can have an inter-channel correlation with the chroma component. The chroma component can have an inter-channel correlation with the luma component. The CCLM can be used for chroma intra prediction. For example, the CCLM can be used for chroma intra prediction, where the chroma samples are predicted from the corresponding luma samples, where sub-sampling is applied based on a linear model. One or more CCLM model parameters can be derived from one or more neighboring luma samples and / or chroma samples (e.g., the casual neighboring luma and chroma samples around the current CU). For example, when the combined inter and intra prediction is applied to a CU, the DM mode can be replaced with the CCLM mode for intra prediction of the chroma samples. In other words, the CCLM mode (e.g., instead of the DM mode) can be applied to one or more chroma components.
[0066] The CCLM mode can be applied to the chroma component in one or more of the following ways. In an example, one or more luma prediction samples of a CU can be generated by blending prediction samples generated from intra-prediction (e.g., based on an intra mode signaled in a bitstream) and inter-prediction (e.g., based on MVs of neighboring blocks indicated by a merge index). The one or more luma prediction samples and residual samples of the luma component can be added together, for example, to generate (e.g., form) one or more reconstructed luma samples. The one or more reconstructed luma samples can be sub-sampled. Based on the CCLM mode, the one or more reconstructed luma samples can be used to generate one or more chroma prediction samples.
[0067] In an example, the one or more chroma prediction samples (e.g., CCLM prediction samples) can be combined with the chroma inter prediction samples, e.g., to generate the final prediction samples for the chroma components. The one or more chroma prediction samples can be referred to as chroma intra prediction samples.
[0068] Figure 11 An example chroma prediction 1100 for combined inter and intra prediction using (e.g., directly using) CCLM chroma prediction samples is shown. In a first example method of applying CCLM mode to chroma components (e.g., as shown in Figure 11 The chroma LM prediction samples can be directly used.
[0069] Figure 12 An example chroma prediction 1200 for combined inter and intra prediction by combining chroma inter prediction samples and CCLM chroma prediction samples is shown. In a second example method of applying CCLM mode to chroma components (e.g., as shown in Figure 12 The chroma intra prediction samples can be combined with the chroma inter prediction samples.
[0070] The DM mode and the CCLM mode can be enabled for chroma intra prediction of the combined inter and intra prediction. For example, a flag can be signaled when the combined inter and intra prediction is enabled for a current CU. The flag can indicate whether to apply the DM mode or the CCLM mode. One or more of the following can be applied. When the flag is set to true, the DM mode can be applied to the chroma components, e.g., such that the same intra mode of the luma component is reused to generate the chroma intra prediction samples. When the flag is set to false, the CCLM mode can be applied to the intra prediction of the chroma samples, e.g., the chroma intra prediction samples will be generated from the sub-sampled luma reconstructed samples based on the linear mode.
[0071] The combined inter and intra prediction can interact with other coding tools. The combined inter and intra prediction can be enabled when a CU is coded by the merge mode (e.g., regular merge mode). The combined inter and intra prediction mode can interact with one or more other inter coding tools (e.g., including affine merge, ATMVP, and triangle prediction). One or more interactions (e.g., synergy of the interactions) between the combined inter and intra prediction mode and one or more other inter coding tools can be improved. One or more of the following can be applied.
[0072] The combined inter and intra prediction can be enabled for subblock merge mode. Motion information can be derived at subblock level. For example, subblock derivation of the motion information can be achieved by subblock merge mode (e.g., ATMVP and affine mode). In an example, the combined inter and intra prediction can be disabled (e.g., for a CU predicted by the subblock mode). The subblock motion information can be derived (e.g., purely derived) from one or more spatial neighbors (e.g., in affine merge mode). The subblock motion information can be derived from one or more temporal neighbors (e.g., in ATMVP mode). The subblock motion information can not be accurate enough to generate predicted samples for the current CU. The combined inter and intra prediction for subblock merge mode can be enabled. In an example, the combined inter and intra mode can be enabled (e.g., only enabled) for ATMVP mode and can be disabled (e.g., always disabled) for affine merge mode. For example, a combined inter and intra prediction flag can be signaled to enable the combined inter and intra prediction. The combined inter and intra prediction flag can be signaled after a subblock merge index. The combined inter and intra prediction flag can be skipped when the subblock merge index refers to an affine candidate. The combined inter and intra prediction flag can be inferred to be false (e.g., the combined inter / intra prediction is disabled) when the subblock merge index refers to an affine candidate. If the subblock merge index refers to an ATMVP candidate, the combined inter / intra flag can be signaled to indicate whether the combined inter and intra prediction is enabled for the current CU.
[0073] The combined inter and intra prediction can be enabled for triangle mode (e.g., triangle merge mode). In triangle prediction, a PU (e.g., each PU) in a current CU can be inter predicted using MVs (e.g., motion vector candidates) from a single prediction candidate list of the current CU. The MV candidates in the single prediction candidate list can be derived from spatial and / or temporal neighboring blocks of the current CU, which can not be accurate enough to describe the motion of the CU. The combined inter and intra prediction can be used for triangle prediction mode, for example, to improve the efficiency of triangle inter prediction. An intra prediction mode (e.g., a single intra prediction mode) can be signaled for one or more (e.g., two) PUs. The intra prediction mode can be signaled, for example, separately for each PU.
[0074] For the combined inter and intra prediction and / or MMVD mode, triangle mode can be disabled. Triangle mode can be referred to as triangle merge mode. Triangle merge mode can be enabled when CIIP is not applied and MMVD mode is not used. For example, a video encoding device (e.g., as described in FIG. 1) can encode a video sequence including a current CU. The video encoding device can determine whether to apply CIIP and MMVD mode for the current CU. The video encoding device can encode the current CU using triangle merge mode when CIIP is not applied and MMVD mode is not used. The video encoding device can encode the current CU using the combined inter and intra prediction when CIIP is applied and MMVD mode is used. Figure 1) may determine whether to enable the triangle merge mode. When the encoder determines to enable the triangle merge mode, the encoder may determine not to apply CIIP and not to use the MMVD mode. For example, the video encoding device may enable the triangle merge mode without receiving a triangle merge mode indication (e.g., a triangle merge mode flag). The video encoding device (e.g., Figure 3 The decoder of the configuration shown in FIG. 1 may receive an MMVD mode indication (e.g., an MMVD flag). For example, an encoder may send the MMVD mode indication to the video encoding device. The video encoding device may be a WTRU (e.g., Figures 13A-13D The video encoding device may include a WTRU (e.g., Figures 13A-13D WTRU102 shown in ). The MMVD mode indication may indicate whether MMVD mode is used to generate inter-frame prediction for the CU. The MMVD mode indication may be received on a per-coding-unit basis. The video encoding device may receive a combined inter-frame merging / intra-frame prediction (CIIP) indication (e.g., a CIIP flag). For example, the encoder may send the CIIP indication to the video encoding device. When the MMVD mode indication uses MMVD mode for a CU, the CIIP indication may not be received. The CIIP indication may indicate whether CIIP is applied to the CU.
[0075] In an example, the triangle merge flag can be signaled after the subblock merge flag, the combined inter / intra flag, and / or the MMVD flag, for example. The triangle flag can not be signaled if one or more of the above flags are set to true. The triangle merge flag can not be signaled when the CIIP indication indicates that CIIP is applied to the CU. For example, when the CIIP indication indicates that CIIP is applied for the CU, the encoder can determine not to include the triangle merge flag in the message to the video coding device (e.g., decoder). The triangle merge flag can not be signaled when the MMVD mode indication indicates that MMVD mode is used to generate the inter prediction. For example, when the MMVD mode indication indicates that MMVD mode is used to generate the inter prediction, the encoder can determine not to include the triangle merge flag in the message to the video coding device (e.g., decoder). When the triangle merge flag is not signaled, the triangle merge flag can be inferred to be false (e.g., the triangle mode is disabled). For example, the video coding device can infer whether the triangle merge mode is enabled for the CU based on the MMVD mode indication and / or the CIIP indication. For example, when the MMVD mode indication indicates that MMVD mode is used to generate the inter prediction, the video coding device can disable the triangle merge mode for the CU. For example, when the CIIP indication indicates that CIIP is applied to the CU, the video coding device can disable the triangle merge mode. If the three flags are set to false, the triangle flag can be signaled, for example, to indicate whether the triangle mode is applied to the current CU.
[0076] Figure 13A FIG. 1 shows a diagram of an example communications system 100 in which one or more disclosed embodiments can be implemented. The communications system 100 can be a multiple access system that provides content, such as voice, data, video, messaging, broadcast, etc., to multiple wireless users. The communications system 100 can enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communications systems 100 can employ one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), zero-tail
[0077] As Figure 13AAs shown, the communication system 100 can include wireless transmit / receive units (WTRUs) 102a, 102b, 102c, 102d, a RAN 104 / 113, a CN 106 / 115, a public switched telephone network (PSTN) 108, the Internet 110, and other networks 112, though it will be appreciated that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, 102d can be any type of device configured to
[0078] The communication system 100 can also include a base station 114a and / or a base station 114b. Each of the base stations 114a, 114b can be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, 102d to facilitate access to one or more communication networks, such as the CN 106 / 115, the Internet 110, and / or the other networks 112. By way of example, the base stations 114a, 114b can be a base transceiver station (BTS), a Node-B, an eNode B, a Home Node B, a Home eNode B, a gNB, a NR Node B, a site controller, an access point (AP), a wireless router, and the like. While the base stations 114a, 114b are each depicted as a single element, it will be appreciated that the base stations 114a, 114b can include any number of interconnected base stations and / or network elements.
[0079] The base stations 114a can be part of the RAN 104 / 113, which can also include other base stations and / or network elements (not shown), such as a base station controller (BSC), a radio network controller (RNC), relay nodes, etc. The base stations 114a and / or the base stations 114b can be configured to transmit and / or receive wireless signals on one or more carrier frequencies under the control of a base station controller (not shown), and / or to transmit and / or receive wireless signals on one or more carrier frequencies without the control of a base station controller. These frequencies can be in the licensed spectrum or the unlicensed spectrum, or a combination thereof. A cell can be a relatively fixed area or, possibly, a changing area depending on the location of the WTRUs 102a, 102b, 102c, 102d within the cell. The cell can further be split into multiple sectors, for example. Thus, in one embodiment, the base station 114a can include multiple transceivers, for example, each corresponding to a different sector of the cell. In an embodiment, the base station 114a can use multiple-input multiple-output (MIMO) techniques. Thus, the base station 114a can utilize, for example, beamforming to transmit and / or receive signals from and / or to the WTRUs 102a, 102b, 102c, 102d.
[0080] The base stations 114a, 114b can communicate with one or more of the WTRUs 102a, 102b, 102c, 102d over an air interface 116, which can be any suitable wireless communication link (e.g., radio frequency (RF), microwave, centimeter wave, micrometer wave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 116 can be established using any suitable radio access technology (RAT).
[0081] More specifically, as noted above, the communications system 100 can be a multiple access system and can employ one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, SC-FDMA, and the like. For example, the base station 114a in the RAN 104 / 113 and the WTRUs 102a, 102b, 102c can implement a radio technology such as UMTS Terrestrial Radio Access (UTRA), which can establish the air interface 115 / 116 / 117 using wideband CDMA (WCDMA). WCDMA can include communication protocols such as High-Speed Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA can include High-Speed Downlink (DL) Packet Access (HSDPA) and / or High-Speed UL Packet Access (HSUPA).
[0082] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which can establish the air interface 116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A) and / or LTE-Advanced Pro (LTE-A Pro).
[0083] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement a radio technology such as NR Radio Access, which can establish the air interface 116 using New Radio (NR).
[0084] In an embodiment, the base station 114a and the WTRUs 102a, 102b, 102c can implement multiple radio access technologies. For example, the base station 114a and WTRUs 102a, 102b, 102c can implement LTE radio access and NR radio access together, for instance using dual connectivity (DC) principles. Thus, the air interface utilized by WTRUs 102a, 102b, 102c can be characterized by multiple types of radio access technologies and / or transmissions sent to / from multiple types of base stations (e.g., a eNB and a gNB).
[0085] In other embodiments, the base station 114a and the WTRUs 102a, 102b, 102c can implement radio technologies such as IEEE 802.11 (i.e., Wireless Fidelity (WiFi), IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 IX, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile communications (GSM), Enhanced Data rates for GSM Evolution (EDGE), GSM EDGE (GERAN), and the like.
[0086] Figure 13AThe base station 114b in FIG. 1 A can be a wireless router, Home Node B, Home eNode B, or access point, for example, and can utilize any suitable RAT for facilitating wireless connectivity access to the Internet, such as IEEE 802.11. In one embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In an embodiment, the base station 114b and the WTRUs 102c, 102d can implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 114b and the WTRUs 102c, 102d can utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, LTE-A Pro, NR, etc.) to establish a picocell or femtocell. As shown in FIG. 1 A, the base station 114b can have a direct connection to the Internet 110. Thus, the base station 114b can not be required to access the Internet 110 via the CN 106 / 115. Figure 13A
[0087] The RAN 104 / 113 can be in communication with the CN 106 / 115, which can be any type of network configured to provide voice, data, applications, and / or voice over internet protocol (VoIP) services to one or more of the WTRUs 102a, 102b, 102c, 102d. The data can have varying quality of service (QoS) requirements, such as differing throughput requirements, latency requirements, error tolerance requirements, reliability requirements, data throughput requirements, mobility requirements, and the like. The CN 106 / 115 can provide call control, billing services, mobile location-based services, pre-paid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in FIG. 1 A, it will be appreciated that the RAN 104 / 113 and / or the CN 106 / 115 can be in direct or indirect communication with other RANs that employ the same RAT as the RAN 104 / 113 or a different RAT. For example, in addition to being connected to the RAN 104 / 113, which can be utilizing a NR radio technology, the CN 106 / 115 can also be in communication with another RAN (not shown) employing a GSM, UMTS, CDMA 2000, WiMAX, Wi-Fi radio technology, etc. Figure 13A
[0088] The CN 106 / 115 can also serve as a gateway for the WTRUs 102a, 102b, 102c, 102d to access the PSTN 108, the Internet 110, and / or the other networks 112. The PSTN 108 can include circuit-switched telephone networks that provide infrastructure for the provision of voice telephony. The Internet 110 can include a global system of interconnected computer networks and devices that use the Transmission Control Protocol (TCP) Internet Protocol (IP) suite of protocols to communicate with one another. The other networks 112 can include wired or wireless communications networks owned and / or operated by other service providers. For example, the other networks 112 can include another CN that can be similar to the CN 106 / 115 but that can be owned and / or operated by another service provider, such as a CN of another wireless service provider.
[0089] Some or all of the WTRUs 102a, 102b, 102c, 102d in the communications system 100 can include multi-mode capabilities, e.g., the WTRUs 102a, 102b, 102c, 102d can include multiple transceivers for communicating with different wireless networks over different wireless links. For example, the WTRU 102c shown in Figure 1 A can be configured to communicate with the base station 114a, which can employ a cellular-based radio technology, and with the base station 114b, which can employ an IEEE 802 radio technology. Figure 13A The WTRU 102c shown in Figure 1 A can be configured to communicate with the base station 114a using a cellular-based radio technology and with the base station 114b using an IEEE 802 radio technology.
[0090] Figure 13B Figure 1 B is a system diagram illustrating an example WTRU 102. As shown in Figure 13B The WTRU 102 can include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, non-removable memory 130, removable memory 132, a power source 134, a global positioning system (GPS) chipset 136, and / or other peripherals 138, among others. It will be appreciated that the WTRU 102 can include any sub-combination of the foregoing elements while remaining consistent with an embodiment.
[0091] The processor 118 can be a general purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASICs), Field Programmable Gate Array (FPGAs) circuits, any other type of integrated circuit (IC), a state machine, and the like. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled Figure 13B The processor 118 and the transceiver 120 are depicted as separate components, however, it will be appreciated that the processor 118 and the transceiver 120 can be integrated in an electronic package or chip.
[0092] The transmit / receive element 122 can be configured to transmit signals to, and receive signals from, a base station (e.g., the base station 114a) over the air interface 116. For example, in one embodiment, the transmit / receive element 122 can be an antenna configured to transmit and / or receive RF signals. In an embodiment, the transmit / receive element 122 can be an emitter / detector configured to transmit and / or receive IR, UV, or visible light signals, for example. In an embodiment, the transmit / receive element 122 can be configured to transmit and / or receive both RF and light signals. It will be appreciated that the transmit / receive element 122 can be configured to transmit and / or receive any combination of wireless signals.
[0093] Although the transmit / receive element 122 is depicted in the Figure 13B embodiments, the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can employ MIMO technology. Thus, in an embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 116.
[0094] The transceiver 120 can be configured to modulate the signals that are to be transmitted by the transmit / receive element 122 and to demodulate the signals that are received by the transmit / receive element 122. As noted above, the WTRU 102 can have multi-mode capabilities. Thus, the transceiver 120 can include multiple transceivers for enabling the WTRU 102 to communicate via multiple RATs, such as NR and IEEE 802.11, for example.
[0095] The processor 118 of the WTRU 102 can be coupled to, and can receive user input data from, the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or organic light-emitting diode (OLED) display unit). The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can access information from, and store data in, any type of suitable memory, such as the non-removable memory 130 and / or the removable memory 132. The non-removable memory 130 can include random-access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 can include a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, and the like. In other embodiments, the processor 118 can access information from, and store data in, memory that is not physically located on the WTRU 102, such as on a server or a home computer (not shown).
[0096] The processor 118 can receive power from the power source 134, and can be configured to distribute and / or control the power to the other components in the WTRU 102. The power source 134 can be any suitable device for powering the WTRU 102. For example, the power source 134 can include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, and the like.
[0097] The processor 118 can also be coupled to the GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or in lieu of, the information from the GPS chipset 136, the WTRU 102 can receive location information over the air interface 116 from a base station (e.g., base stations 114a, 114b) and / or determine its location based on
[0098] The processor 118 can further couple to other peripherals 138, which can include one or more software and / or hardware modules that provide additional features, functionality and / or wired or wireless connectivity. For example, the peripherals 138 can include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photographs or video), a universal serial bus (USB) port, a vibration device, a television transceiver, a hands- free headset, an A / V port, a subwoofer, a power meter, a memory stick, and the like. The peripheral device 138 can include one or more sensors, which can be one or more of a gyroscope, an accelerometer, a hall effect sensor, a magnetometer, a compass sensor, a proximity sensor, a temperature sensor, a time sensor, a geo-location sensor, an altimeter, a light sensor, a touch sensor, a magnetometer, a barometer, a gesture sensor, a biometric sensor, and / or a humidity sensor.
[0099] The WTRU 102 can include a full duplex radio for which transmission and reception of some or all signals (e.g., associated with particular subframes for both the UL and downlink) can be concurrent and / or simultaneous. The full duplex radio can include a transmission and reception interference management unit to reduce and / or substantially eliminate self-interference and / or other-interference. In an embodiment, the WTRU 102 can include a half duplex radio for which transmission and reception of some or all signals (e.g., associated with particular subframes for either the UL or the downlink) can be concurrent but not simultaneous.
[0100] Figure 13C is a system diagram illustrating the RAN 104 and the CN 106 according to an embodiment. As noted above, the RAN 104 can employ an E-UTRA radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 104 can also be in communication with the CN 106.
[0101] The RAN 104 can include eNode-Bs 160a, 160b, 160c, though it will be appreciated that the RAN 104 can include any number of eNode-Bs while remaining consistent with an embodiment. The eNode-Bs 160a, 160b, 160c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the eNode-Bs 160a, 160b, 160c can implement MIMO technology. Thus, the eNode-B 160a, for example, can use multiple antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a.
[0102] Each of the eNode-Bs 160a, 160b, 160c can be associated with a specific cell (not shown) and can be configured to handle radio resource management decisions, handover decisions, scheduling of users in the UL and / or DL, and the like. As shown, the eNode-Bs 160a, 160b, 160c can communicate with one another over an X2 interface. Figure 13C
[0103] Figure 13C The CN 106 can include a mobility management entity (MME) 162, a serving gateway (SGW) 164, and a packet data network (PDN) gateway (or PGW) 166. While each of the foregoing elements are depicted as part of the CN 106, it will be appreciated that any of these elements can be owned and / or operated by an entity other than the CN operator.
[0104] The MME 162 can be connected to each of the eNode-Bs 162a, 162b, 162c in the RAN 104 via an SI interface and can serve as a control node. For example, the MME 142 can be responsible for authenticating users of the WTRUs 102a, 102b, 102c, bearer activation / deactivation, selecting a particular serving gateway during an initial attach of the WTRUs 102a, 102b, 102c, and the like. The MME 162 can also provide a control plane function for switching between the RAN 104 and other RANs (not shown) that employ other radio technologies, such as GSM and / or WCDMA.
[0105] The SGW 164 can be connected to each of the eNode Bs 160a, 160b, 160c in the RAN 104 via the S I interface. The SGW 164 can generally route and forward user data packets to / from the WTRUs 102a, 102b, 102c. The SGW 164 can also perform other functions, such as anchoring user planes during inter-eNode B handovers, triggering paging when DL data is available for the WTRUs 102a, 102b, 102c, managing and storing contexts of the WTRUs 102a, 102b, 102c, and the like.
[0106] The SGW 164 can be connected to the PGW 166, which can provide the WTRUs 102a, 102b, 102c with access to packet-switched networks, such as the Internet 110, to facilitate communications between the WTRUs 102a, 102b, 102c and IP-enabled devices.
[0107] The CN 106 can facilitate communications with other networks. For example, the CN 106 can provide the WTRUs 102a, 102b, 102c with access to circuit-switched networks, such as the PSTN 108, to facilitate communications between the WTRUs 102a, 102b, 102c and traditional landline communications devices. For example, the CN 106 can include, or can communicate with, an IP gateway (e.g., an IP multimedia subsystem (IMS) server) that serves as an interface between the CN 106 and the PSTN 108. In addition, the CN 106 can provide the WTRUs 102a, 102b, 102c with access to the other networks 112, which can include other wired and / or wireless networks that are owned and / or operated by other service providers.
[0108] Although WTRUs are described herein as wireless terminals, it is contemplated that in certain representative embodiments such terminals can use (e.g., temporarily or permanently) wired communication interfaces with communication networks. Figures 13A-13D
[0109] In representative embodiments, the other network 112 can be a WLAN.
[0110] A WLAN using an infrastructure basic service set (BSS) mode can have an access point (AP) for the BSS and one or more stations (STAs) associated with the AP. The AP can have an interface to a distribution system (DS) or other type of wired / wireless network that serves to carry traffic in and / or out of the BSS. Traffic to STAs that originates from outside the BSS can arrive through the AP and can be delivered to the STAs. Traffic originating from STAs to destinations outside the BSS can be sent to the AP to be delivered to respective destinations. Traffic between STAs within the BSS can be sent through the AP, for example, a source STA can send traffic to the AP and the AP can deliver the traffic to a destination STA. The traffic between STAs within a BSS can be considered and / or referred to as peer-to-peer traffic. The peer-to-peer traffic can be sent between (e.g., directly between) the source and destination STAs using a direct link setup (DLS). In certain representative embodiments, the DLS can use 802.11e DLS or 802.11z tunneled DLS (TDLS). A WLAN using an IBSS mode can not have an AP, and STAs (e.g., all STAs) within the IBSS or using the IBSS can communicate directly with each other. Here, the IBSS mode of communication can sometimes be referred to as an“ad hoc” mode of communication.
[0111] In using an 802.1 lac infrastructure mode of operation or similar mode of operation, an AP can transmit beacons on a fixed channel (e.g., a primary channel). The primary channel can have a fixed width (e.g., 20 MHz bandwidth) or a dynamically set width by signaling. The primary channel can be the operating channel of the BSS and can be used by STAs to establish a connection with the AP. In certain representative embodiments, carrier sense multiple access with collision avoidance (CSMA / CA) (e.g., in 802.11 systems) can be implemented. For CSMA / CA, STAs, including the AP (e.g., each of the STAs) can sense the primary channel. If a particular STA senses / detects and / or determines that the primary channel is busy, the particular STA can back off. In a given BSS, only one STA (e.g., only one station) can transmit at any given time.
[0112] High Throughput (HT) STAs can use a 40 MHz wide channel for communication (e.g., by combining a 20 MHz primary channel with an adjacent or nonadjacent 20 MHz channel to form a 40 MHz wide channel).
[0113] Very High Throughput (VHT) STAs can support 20 MHz, 40 MHz, 80 MHz, and / or 160 MHz wide channels. 40 MHz and / or 80 MHz wide channels can be formed by combining contiguous 20 MHz channels. A 160 MHz wide channel can be formed by combining 8 contiguous 20 MHz channels, or by combining two noncontiguous 80 MHz channels, which can be referred to as an 80+80 configuration. For the 80+80 configuration, after channel coding, the data can be passed through a segment parser that can divide the data into two streams. Inverse Fast Fourier Transform (IFFT) processing and time domain processing can be done on each stream separately. The streams can be mapped on to the two 80 MHz channels and data can be transmitted by the transmitting STA. At the receiver of the receiving STA, the above described operations can be reversed for the 80+80 configuration and the combined data can be sent to the Medium Access Control (MAC).
[0114] 802.11af and 802.11ah support sub-1 GHz operating modes. In 802.11af and 802.11ah, there is a reduction in the channel operating bandwidth and carriers used as compared to 802.11η and 802.1 lac. 802.11af supports 5 MHz, 10 MHz, and 20 MHz bandwidth in TV White Space (TVWS) spectrum, and 802.11ah supports 1 MHz, 2 MHz, 4 MHz, 8 MHz, and 16 MHz bandwidth using non-TVWS spectrum. According to certain typical embodiments, 802.11ah can support meter type control / machine type communication, e.g., MTC devices in a macro coverage area. MTC can have certain capabilities, e.g., including limited capabilities that support (e.g., only support) certain and / or limited bandwidth. MTC devices can include a battery, and the battery life of the battery is above a threshold (e.g., to maintain a long battery life).
[0115] For WLAN systems that can support multiple channels and channel bandwidths (e.g., 802.11η, 802.1 lac, 802.11af, and 802.11ah), the WLAN system includes one channel that can be designated as a primary channel. The bandwidth of the primary channel can be equal to the maximum common operating bandwidth supported by all STAs in the BSS. The bandwidth of the primary channel can be set and / or limited by a STA that is derived from all STAs operating in the BSS that support the minimum bandwidth operating mode. In an example with respect to 802.11ah, even though the AP and other STAs in the BSS support 2 MHz, 4 MHz, 8 MHz, 16 MHz, and / or other channel bandwidth operating modes, the width of the primary channel can be 1 MHz for STAs (e.g., MTC type devices) that support (e.g., only support) 1 MHz mode. Carrier sensing and / or network allocation vector (NAV) settings can depend on the status of the primary channel. If the primary channel is busy (e.g., because a STA that only supports 1 MHz operating mode is transmitting to the AP), then the entire available frequency band can be considered busy even though most of the frequency band remains idle and available for use.
[0116] In the United States, the available frequency band for 802.11ah use is 902 MHz to 928 MHz. In Korea, the available frequency band is 917.5 MHz to 923.5 MHz. In Japan, the available frequency band is 916.5 MHz to 927.5 MHz. Depending on the country code, the total bandwidth available for 802.11ah is 6 MHz to 26 MHz.
[0117] Figure 13Dis a system diagram illustrating the RAN 113 and the CN 115 according to an embodiment. As described above, the RAN 113 can employ NR radio technology to communicate with the WTRUs 102a, 102b, 102c over the air interface 116. The RAN 113 can also be in communication with the CN 115.
[0118] The RAN 113 can include gNBs 180a, 180b, 180c, although it will be appreciated that the RAN 113 can include any number of gNBs while remaining consistent with an embodiment. The gNBs 180a, 180b, 180c can each include one or more transceivers for communicating with the WTRUs 102a, 102b, 102c over the air interface 116. In one embodiment, the gNBs 180a, 180b, 180c can implement MIMO technology. For example, gNBs 180a, 180b can utilize beamforming to transmit and / or receive signals to and / or from gNBs 180a, 180b, 180c. Thus, the gNB 180a, for example, can employ a number of antennas to transmit wireless signals to, and / or receive wireless signals from, the WTRU 102a. In an embodiment, the gNBs 180a, 180b, 180c can implement carrier aggregation technology. For example, the gNB 180a can transmit multiple component carriers (not shown) to the WTRU 102a. A subset of these component carriers can be on the licensed spectrum while the remaining component carriers can be on the unlicensed spectrum. In an embodiment, the gNBs 180a, 180b, 180c can implement Coordinated Multi-Point (CoMP) technology. For example, WTRU 102a can receive coordinated transmissions from gNBs 180a and 180b (and / or gNB 180c).
[0119] The WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using transmissions associated with a scalable numerology. For example, the OFDM symbol spacing and / or OFDM subcarrier spacing can vary for different transmissions, different cells, and / or different portions of the wireless transmission spectrum. WTRUs 102a, 102b, 102c can communicate with gNBs 180a, 180b, 180c using subframes or transmission time intervals (TTIs) of various or scalable lengths (e.g., containing varying number of OFDM symbols and / or lasting varying lengths of absolute time).
[0120] The gNBs 180a, 180b, 180c may be configured to communicate with the WTRUs 102a, 102b, 102c in a standalone configuration and / or a non-standalone configuration. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c without accessing other RANs (e.g., the eNode-Bs 160a, 160b, 160c). In a standalone configuration, the WTRUs 102a, 102b, 102c may use one or more of the gNBs 180a, 180b, 180c as mobility anchors. In a standalone configuration, the WTRUs 102a, 102b, 102c may communicate with the gNBs 180a, 180b, 180c using signals in an unlicensed band. In a non-standalone configuration, the WTRUs 102a, 102b, 102c may communicate / connect with the gNBs 180a, 180b, 180c while communicating / connecting with another RAN (e.g., the eNode-Bs 160a, 160b, 160c). For example, the WTRUs 102a, 102b, 102c may communicate with one or more gNBs 180a, 180b, 180c and one or more eNode-Bs 160a, 160b, 160c substantially simultaneously by implementing the DC principle. In a non-standalone configuration, the eNode-Bs 160a, 160b, 160c may serve as mobility anchors for the WTRUs 102a, 102b, 102c, and the gNBs 180a, 180b, 180c may provide additional coverage and / or throughput to serve the WTRUs 102a, 102b, 102c.
[0121] Each gNB 180a, 180b, 180c may be associated with a specific cell (not shown) and may be configured to handle radio resource management decisions, handover decisions, user scheduling in UL and / or DL, support network slicing, implement dual connectivity, implement interworking between NR and E-UTRA, route user plane data to user plane functions (UPFs) 184a, 184b, and route control plane information to access and mobility management functions (AMFs) 182a, 182b, etc. Figure 13D As shown, gNBs 180a, 180b, and 180c can communicate with each other via the Xn interface.
[0122] Figure 13DThe CN 115 shown can include at least one AMF 182a, 182b, at least one UPF 184a, 184b, at least one Session Management Function (SMF) 183a, 183b, and possibly a Data Network (DN) 185a, 185b. While each of the foregoing elements are depicted as part of the CN 115, it will be appreciated that any of these elements can be owned and / or operated by an entity other than the CN operator.
[0123] The AMF 182a, 182b can be connected to one or more of the RANs 113gNBs 180a, 180b, 180c in the CN 115 via an N2 interface and can serve as the control node. For example, the AMF 182a, 182b can be responsible for authenticating WTRUs 102a, 102b, 102c, support for network slicing (e.g., handling different PDU sessions with different requirements), selecting a particular SMF 183a, 183b, management of the WTRU 102a, 102b, 102c registration area, termination of NAS signaling, mobility management, and the like. The AMF 162 can utilize network slicing to tailor the CN support that it provides WTRUs 102a, 102b, 102c based on the type of service that the WTRUs 102a, 102b, 102c are using. For example, different network slices can be established for services relying on ultra-reliable low latency (URLLC) access, services relying on enhanced massive mobile broadband (eMBB) access, and / or services for machine type communication (MTC) access, among others. The AMF 162 can provide control plane functionality for switching between the RAN 113 and other RANs (not shown) that employ other radio technologies, such as LTE, LTE-A, LTE-A Pro, and / or non-3GPP access technologies, such as WiFi.
[0124] The SMF 183a, 183b can be connected to AMF 182a, 182b in the CN 115 via an N11 interface. The SMF 183a, 183b can also be connected to UPF 184a, 184b in the CN 115 via an N4 interface. The SMF 183a, 183b can select and control the UPF 184a, 184b and configure the routing of traffic through the UPF 184a, 184b. The SMF 183a, 183b can perform other functions, such as managing and allocating WTRU / UE IP address, managing PDU sessions, controlling policy enforcement and QoS, providing downlink data notifications, and the like. A PDU session type can be IP-based, non-IP based, Ethernet-based, and the like.
[0125] The UPF 184a, 184b can be connected to one or more gNBs 180a, 180b, 180c in the RAN 113 via the N3 interface, which can provide the WTRUs 102a, 102b, 102c with access to packet-switched networks (such as the Internet 110) to facilitate communication between the WTRUs 102a, 102b, 102c and IP-enabled devices. The UPF 184, 184b can perform other functions such as routing and forwarding packets, implementing user plane policies, supporting multi-host PDU sessions, handling user plane QoS, buffering downlink packets, and providing mobility anchor processing, etc.
[0126] The CN 115 may facilitate communications with other networks. For example, the CN 115 may include or may communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the CN 115 and the PSTN 108. Furthermore, the CN 115 may provide the WTRUs 102a, 102b, 102c with access to other networks 112, which may include other wired and / or wireless networks owned and / or operated by other service providers. In one embodiment, the WTRUs 102a, 102b, 102c may be connected to a local data network (DN) 185a, 185b through the UPFs 184a, 184b via an N3 interface to the UPFs 184a, 184b and an N6 interface between the UPFs 184a, 184b and the data network (DN) 185a, 185b.
[0127] In view of Figures 13A-13D and about Figures 13A-13D
[0015] As described herein, one or more or all of the functions described herein with respect to one or more of the following may be performed by one or more emulated devices (not shown): WTRUs 102a-d, base stations 114a-b, eNodeBs 160a-c, MMEs 162, SGWs 164, PGWs 166, gNBs 180a-c, AMFs 182a-b, UPFs 184a-b, SMFs 183a-b, DNs 185a-b, and / or any other device(s) described herein. These emulated devices may be one or more devices configured to emulate one or more or all of the functions described herein. For example, these emulated devices may be used to test other devices and / or emulate network and / or WTRU functions.
[0128] The emulation devices can be designed to implement one or more tests on other devices in a laboratory environment and / or a carrier network environment. For example, the one or more emulation devices can perform one or more or all functions while being implemented, at least in part, as part of a wired and / or wireless communication network to test other devices within the communication network. The one or more emulation devices can perform one or more or all functions while being temporarily implemented / deployed as part of a wired and / or wireless communication network. The emulation devices can directly couple to the other devices to perform the tests, and / or can perform the tests using over-the-air, wireless communications.
[0129] The one or more emulation devices can perform one or more functions, including all functions, while not being implemented / deployed as part of a wired and / or wireless communication network. For example, the emulation devices can be used in a test laboratory and / or a test field environment, where a wired and / or wireless communication network is not deployed (e.g., is under construction, is being tested, is being deployed, etc.) to implement tests on one or more components. The one or more emulation devices can be test equipment. The emulation devices can transmit and / or receive data using direct RF coupling and / or wireless communications via RF circuitry (which can include one or more antennas, for example).
[0130] The processes described herein can be implemented in a computer program, software, and / or firmware incorporated in a computer- readable medium for execution by a computer and / or processor. Examples of computer-readable media include but are not limited to electronic signals (optical, electrical or electromagnetic) transmitted over wired and / or wireless connections and / or computer- readable storage media. Examples of computer- readable storage media include but are not limited to removable / fixed memory devices, optical storage devices, and / or any other storage medium suitable for storing electronic instructions. The processor can embody any of a number of types of processors, including without limitation microprocessors, microcontrollers, digital signal processors, central processing units, arithmetic logic units, etc. The processor can be configured to implement a radio-frequency transceiver for a WTRU, terminal, base station, RNC, and / or any host computer.
Claims
1. A device for video decoding, comprising: A processor configured to: Obtaining a combined inter-merging / intra-prediction (CIIP) indication, which indicates whether CIIP is applied to a coding block; Based on the CIIP indication, determining that a CIIP is applied to the coding block; Applying CIIP to the coding block based on a weighted combination of inter prediction and intra planar prediction associated with the coding block; Determining, based on the CIIP indication, to disable triangle merging mode for the coding block; as well as Based on disabling the triangle merging mode for the coding block, the coding block is decoded. 2 . The apparatus of claim 1 , wherein the processor is further configured to determine that a triangle merging mode indication for the coding block is not signaled in the video data based on determining that CIIP is applied to the coding block.
3. The apparatus of any one of claims 1-2, wherein the apparatus comprises a memory.
4. A method for video decoding, comprising: Obtaining a combined inter-merging / intra-prediction (CIIP) indication, which indicates whether CIIP is applied to a coding block; as well as Based on the CIIP indication, determining that a CIIP is applied to the coding block; applying CIIP to the coding block based on a weighted combination of inter prediction and intra planar prediction associated with the coding block; Determining, based on the CIIP indication, to disable triangle merging mode for the coding block; as well as Based on disabling the triangle merging mode for the coding block, the coding block is decoded.
5. The method according to claim 4, further comprising: Based on determining that CIIP is applied to the coding block, determining that a triangle merging mode indication for the coding block is not signaled in the video data.
6. A device for video encoding, comprising: A processor configured to: Get the encoding block; determining that combined inter-merging / intra-prediction (CIIP) is applied to the coding block; applying CIIP to the coding block based on a weighted combination of inter prediction and intra planar prediction associated with the coding block; Based on determining that CIIP is applied to the coding block, determining to disable triangle merging mode for the coding block; as well as Based on disabling the triangle merging mode for the coding block, the coding block is encoded.
7. The apparatus of claim 6, wherein the processor is further configured to: A CIIP indication for the coding block is included in the video data to indicate that the CIIP is applied to the coding block.
8. The apparatus of any one of claims 6-7, wherein the apparatus comprises a memory.
9. A method for video encoding, comprising: Get the encoding block; determining that combined inter-merging / intra-prediction (CIIP) is applied to the coding block; applying CIIP to the coding block based on a weighted combination of inter prediction and intra planar prediction associated with the coding block; Based on determining that CIIP is applied to the coding block, determining to disable triangle merging mode for the coding block; as well as Based on disabling the triangle merging mode for the coding block, the coding block is encoded. 10 . The method of claim 9 , further comprising including a CIIP indication for the coding block in the video data to indicate that the CIIP is applied to the coding block.
11. A computer-readable medium comprising instructions for causing one or more processors to perform the method of any one of claims 4 to 5 and 9 to 10.