Method and apparatus for signaling syntax elements in video coding

By applying constraints on syntax elements in the picture header based on bi-predictive slices and considering spatial resolution and scaling ratio offsets, the method optimizes the use of temporal motion vector predictors, addressing inefficiencies in existing video encoding standards and enhancing encoding efficiency.

JP7709570B2Active Publication Date: 2025-07-16BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024076124
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-31
Filing Date
2024-05-08
Publication Date
2025-07-16
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

Existing video encoding standards like VVC face issues with redundant signaling of syntax elements and inefficient use of temporal motion vector predictors due to unresolved bitstream compliance constraints, leading to unnecessary decoding operations and reduced encoding efficiency.

Method used

Implementing constraints on syntax elements in the picture header based on the presence of bi-predictive slices and considering spatial resolution and scaling ratio offsets, ensuring that flags like ph_temporal_mvp_enabled_flag are set only when applicable, thereby optimizing the use of temporal motion vector predictors.

Benefits of technology

Enhances encoding efficiency by reducing redundant signaling and improving the utilization of temporal motion vector predictors, thereby improving video quality and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007709570000009
    Figure 0007709570000009
  • Figure 0007709570000010
    Figure 0007709570000010
  • Figure 0007709570000011
    Figure 0007709570000011
Patent Text Reader

Abstract

To provide methods and devices for video coding.SOLUTION: A method by a block-based hybrid video encoder includes that a decoder determines whether one or more reference picture lists are signaled in a picture header (PH) associated with a picture and whether the one or more reference picture lists indicate that one or more slices associated with the picture are bi-predictive, and that the decoder adds one or more constraints to one or more syntax elements in the PH in response to determining that the one or more reference picture lists are signaled in the PH and the one or more reference picture lists indicate that the one or more slices are not bi-predictive.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 003,226, entitled "Signaling of Syntax Elements in Video Coding," filed on March 31, 2020, the entire content of which is incorporated by reference in its entirety for all purposes.

[0002] This disclosure relates to video encoding and compression, and more particularly, but not limited to, methods and apparatus for signaling syntax elements in video encoding.

Background Art

[0003] Various video encoding techniques can be used to compress video data. Video encoding is performed according to one or more video encoding standards. For example, currently, some well-known video encoding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to its predecessor VP9. Audio Video Coding (AVS), which refers to digital audio and digital video compression standards, is another series of video compression standards developed by the Audio and Video Coding Standard Workgroup in China. Most of the existing video encoding standards are based on well-known hybrid video encoding frameworks, that is, they use block-based prediction methods (e.g., inter prediction, intra prediction) to reduce the redundancy present in video images or sequences, and use transform coding to compact the energy of the prediction error. An important goal of video encoding technology is to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. Summary of the Invention Problems to be Solved by the Invention

[0004] This disclosure provides examples of techniques related to the signaling of syntax elements in video encoding. Means for Solving the Problems

[0005] According to a first aspect of the present disclosure, a method for video encoding is provided. The method includes a decoder determining whether one or more reference picture lists are signaled by a picture header (PH) associated with a picture, and whether one or more reference picture lists indicate that one or more slices associated with the picture are bi-predictive. Further, the method includes the decoder adding one or more constraints to one or more syntax elements of the PH in response to determining that one or more reference picture lists are signaled by the PH and that one or more reference picture lists indicate that one or more slices are not bi-predictive.

[0006] According to a second aspect of the present disclosure, a method for video encoding is provided. The method includes a decoder using an enable flag to identify whether one or more temporal motion vector predictors are used for inter prediction of one or more slices associated with the PH of a picture. The method further includes the decoder constraining the value of the enable flag according to a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0007] According to a third aspect of the present disclosure, an apparatus for video encoding is provided. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. When executing the instructions, the one or more processors are configured to perform determining whether one or more reference picture lists are signaled in PH associated with a picture, and whether the one or more reference picture lists indicate that one or more slices associated with the picture are bi-predicted. Further, when determining that one or more reference picture lists are signaled in PH and the one or more reference picture lists indicate that one or more slices are not bi-predicted, the one or more processors are configured to perform adding one or more constraints to one or more syntax elements of the PH.

[0008] According to a fourth aspect of the present disclosure, an apparatus for video encoding is provided. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. When executing the instructions, the one or more processors are configured to perform using an enable flag to identify whether one or more temporal motion vector predictors are used for inter-prediction of one or more slices associated with the PH of a picture. The one or more processors are further configured to perform constraining a value of the enable flag according to a plurality of offsets applied to a size of the picture for scaling ratio calculation.

[0009] According to a fifth aspect of the present disclosure, a non-transitory computer-readable storage medium for video encoding storing computer-executable instructions is provided. When executed by one or more computer processors, the instructions cause the one or more computer processors to perform the method for video encoding according to the first aspect of the present disclosure.

[0010] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium for video encoding that stores computer-executable instructions. When the instructions are executed by one or more computer processors, the one or more computer processors are caused to execute a method for video encoding according to a second aspect of the present disclosure.

[0011] A more specific description of examples of the present disclosure will be made with reference to the specific examples shown in the accompanying drawings. Considering that these drawings show only some examples and are not considered to limit the scope, the examples will be described and explained in more specific and detailed manner using the accompanying drawings.

Brief Description of the Drawings

[0012]

Fig. 1

Fig. 2

Fig. 3

Fig. 4A

Fig. 4B

Fig. 4C

Fig. 4D

Fig. 5

Fig. 6

Fig. 7

DETAILED DESCRIPTION OF THE INVENTION

[0013] References to specific implementations are made here in detail, and examples thereof are shown in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to facilitate understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative examples may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented in many types of electronic devices having digital video capabilities.

[0014] References throughout this disclosure to "one embodiment", "an embodiment", "an example", "some embodiments", "some examples", or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are applicable to other embodiments as well, unless otherwise specified.

[0015] Throughout this disclosure, terms such as "first", "second", "third", etc. are used solely as a nomenclature for referring to related elements, such as devices, components, compositions, steps, etc., without implying any spatial or temporal order, unless otherwise specified. For example, "a first device" and "a second device" can refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be arbitrarily named.

[0016] The terms "module", "sub-module", "circuit", "sub-circuit", "circuitry", "sub-circuitry", "unit", or "sub-unit" may include memory (shared, dedicated, or group) that stores code or instructions that can be executed by one or more processors. A module may include one or more circuits with or without the stored code or instructions. A module or circuit may include one or more components that are directly or indirectly connected. These components may or may not be physically attached to or adjacent to each other.

[0017] As used herein, the terms "if" or "when" may be understood to mean "upon" or "in response to" depending on the context. These terms may not indicate that the associated limitations or features are conditional or optional when they appear in the claims. For example, a method may include: i) a step of performing function or action X' when or if condition X exists, and ii) a step of performing function or action Y' when or if condition Y exists. This method may be implemented with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be performed at each of multiple executions of this method.

[0018] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components that are directly or indirectly linked to each other to perform a particular function.

[0019] FIG. 1 shows a block diagram illustrating an exemplary block-based hybrid video encoder 100 that may be used in conjunction with many video coding standards that use block-based processing. In encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter prediction approach or an intra prediction approach. In inter prediction, one or more predictors are formed based on pixels of a previously reconstructed frame by motion estimation and motion compensation. In intra prediction, a predictor is formed based on reconstructed pixels within the current frame. By mode decision, the best predictor may be selected to predict the current block.

[0020] The prediction residual representing the difference between the current video block and its predictor is sent to the transform circuit 102. Next, the transform coefficients are sent from the transform circuit 102 to the quantization circuit 104 for entropy reduction. Next, the quantized coefficients are supplied to the entropy encoding circuit 106 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 110 from the inter prediction circuit and / or the intra prediction circuit 112, such as video block partitioning information, motion vectors, reference picture indices, and intra prediction modes, etc., is also supplied through the entropy encoding circuit 106 and stored in the compressed video bitstream 114.

[0021] In encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed through inverse quantization 116 and inverse transform circuit 118. This reconstructed prediction residual is combined with the block predictor 120 to generate the unfiltered reconstructed pixels of the current video block.

[0022] Intra prediction (also called "spatial prediction") uses pixels from samples of encoded neighboring blocks (referred to as reference samples) within the same video picture and / or slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal.

[0023] Inter prediction (also called "temporal prediction") uses reconstructed pixels from encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, one reference picture index is sent additionally, which is used to identify which reference picture in the reference picture store the temporal prediction signal is from.

[0024] After spatial prediction and / or temporal prediction is performed, the intra / inter mode determination circuit 121 within the encoder 100 selects the best prediction mode, for example, based on the rate distortion optimization method. Next, the block predictor 120 is subtracted from the current video block, and the resulting prediction residual is decorrelated using the transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are inverse quantized by the inverse quantization circuit 116 and inverse transformed by the inverse transform circuit 118 to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, in-loop filtering 115 such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF) can be applied to the reconstructed CU, after which it is placed in the reference picture store of the picture buffer 117 and used to encode future video blocks. To form the output video bitstream 114, the encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy encoding unit 106 to be further compressed and packed to form a bitstream.

[0025] For example, in the latest versions of AVC, HEVC, and VVC, a deblocking filter is available. In HEVC, an additional in-loop filter called SAO (sample adaptive offset) is defined to further improve the encoding efficiency. In the latest version of the VVC standard, yet another in-loop filter called ALF (adaptive loop filter) is being actively studied and has a good chance of being included in the final standard.

[0026] These in-loop filter operations are optional. Performing these operations helps improve the encoding efficiency and visual quality. They can also be turned off as a decision made by the encoder 100 to save computational complexity.

[0027] Intra prediction is typically based on the non-filtered reconstructed pixels, and it should be noted that inter prediction is based on the filtered reconstructed pixels if these filtering options are enabled by the encoder 100.

[0028] FIG. 2 is a block diagram showing an exemplary block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related part present in the encoder 100 of FIG. 1. In decoder 200, the incoming video bitstream 201 is first decoded by the entropy decoder 202 to derive the quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by the inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residual. The block prediction mechanism implemented in the intra / inter mode selector 212 is configured to perform either intra prediction 208 or motion compensation 210 based on the decoded prediction information. A set of non-filtered reconstructed pixels is obtained by summing the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block prediction mechanism using a summer 214.

[0029] The reconstructed block may further pass through the in-loop filter 209 and is then stored in the picture buffer 213 that functions as a reference picture store. The reconstructed video in the picture buffer 213 is sent to drive the display device and can be used to predict future video blocks. In the situation where the in-loop filter 209 is enabled, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0030] The aforementioned video encoding / decoding standards, for example, VVC, JEM, HEVC, MPEG-4, Part 10, are conceptually similar. For example, they all use block-based processing. The block partitioning methods in some standards will be described in detail below.

[0031] High Efficiency Video Coding (HEVC) HEVC is based on a hybrid block-based motion compensated transform coding architecture. The basic unit of compression is called a CTU. The maximum CTU size is defined as two blocks of up to 64×64 luma pixels and 32×32 chroma pixels for the 4:2:0 chroma format. Each CTU can contain one CU or be recursively divided into four smaller CUs until the predefined minimum CU size is reached. Each CU (also named leaf CU) contains a tree of one or more prediction units (PU: prediction unit) and transform units (TU: transform unit).

[0032] Generally, except for monochrome content, a CTU can contain one luma coding tree block (CTB: coding tree block) and two corresponding chroma CTBs, a CU can contain one luma coding block (CB: coding block) and two corresponding chroma CBs, a PU can contain one luma prediction block (PB: prediction block) and two corresponding chroma PBs, and a TU can contain one luma transform block (TB: transform block) and two corresponding chroma TBs. However, exceptions can occur because the minimum TB size is 4×4 for both luma and chroma (i.e., 2×2 chroma TBs are not supported for the 4:2:0 color format), and each intra-chroma CB always has only one intra-chroma PB regardless of the number of intra-luma PBs within the corresponding intra-luma CB.

[0033] For an intra CU, the luma CB can be predicted by one or four luma PBs, and each of the two chroma CBs is always predicted by one chroma PB, where each luma PB has one intra-luma prediction mode, and the two chroma PBs share one intra-chroma prediction mode. Further, for an intra CU, the TB size cannot be made larger than the PB size. In each PB, intra prediction is applied to predict the samples of each TB within the PB from the reconstructed samples in the vicinity of that TB. In each PB, in addition to 33 directed intra prediction modes, the DC mode and the planar mode are also supported, predicting flat regions and gradually changing regions respectively.

[0034] For each inter PU, one of three prediction modes including inter, skip, and merge can be selected. Generally speaking, a motion vector competition (MVC) scheme is introduced to select motion candidates from a given candidate set including spatial and temporal motion candidates. Multiple references for motion estimation enable finding the best reference in two possible reconstructed reference picture lists (i.e., list 0 and list 1). In the inter mode (referred to as the AMVP mode, where AMVP represents advanced motion vector prediction), an inter prediction indicator (list 0, list 1, or bi-directional prediction), a reference index, a motion candidate index, a motion vector difference (MVD), and a prediction residual are transmitted. For the skip mode and the merge mode, only the merge index is transmitted, and the current PU inherits the inter prediction indicator, the reference index, and the motion vector from a neighboring PU referred to by the encoded merge index. For a skip-encoded CU, the residual signal is also omitted.

[0035] Versatile Video Coding (VVC) At the 10th JVET meeting held in San Diego, the United States from April 10th to 20th, 2018, JVET defined the first draft of Versatile Video Coding (VVC) and the VVC Test Model 1 (VTM1) as its reference software implementation. As the first new coding feature of VVC, it was decided to include a quad tree with a nested multi-type tree. The multi-type tree is a coding block partitioning structure that includes both binary and ternary partitions. Since then, a reference software VTM with both encoding and decoding processes implemented has been developed and updated through subsequent JVET meetings.

[0036] In VVC, a picture of the input video is partitioned into blocks called Coding Tree Units (CTUs). A CTU is divided into Coding Units (CUs) using a quad tree with a nested multi-type tree structure, and a CU defines a region of pixels that share the same prediction mode (e.g., intra or inter). The term "unit" can define a region of an image that covers all components such as luma and chroma. The term "block" can be used to define a region that covers a specific component (e.g., luma), and considering chroma sampling formats such as 4:2:0, blocks of different components (e.g., luma vs. chroma) can have different spatial positions.

[0037] Partitioning of a Picture into CTUs FIG. 3 shows an example of a picture 300 divided into a plurality of CTUs 302 according to some implementations of the present disclosure.

[0038] In VCC, a picture is partitioned into a sequence of CTUs. The concept of CTU is the same as that of HEVC. For a picture with three sample arrays, a CTU is composed of a block of N×N luma samples and a corresponding block of two chroma samples.

[0039] The maximum allowable size of the luma block within a CTU is specified as 128×128 (however, the maximum size of the luma transform block is 64×64).

[0040] Partitioning of CTUs Using a Tree Structure In HEVC, to adapt to various local characteristics, CTUs are divided into CUs using a quaternary-tree structure called an encoding tree. The decision on whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode a picture area is made at the leaf CU level. Each leaf CU can be further divided into one, two, or four PUs depending on the PU partitioning type. Within one PU, the same prediction process is applied, and relevant information is sent to the decoder for each PU. After obtaining the residual block by applying the prediction process based on the PU partitioning type, the leaf CU can be divided into transform units (TUs) according to another quaternary-tree structure similar to the encoding tree of that CU. One of the important features of the HEVC structure is having the concept of multiple partitioning including CUs, PUs, and TUs.

[0041] In VVC, a quadtree with a nested multi-type tree using a binary and ternary segmentation structure replaces the concept of multiple partitioning unit types, that is, except for appropriately removing CUs with sizes too large for the maximum transform length, the distinction between the concepts of CUs, PUs, and TUs is removed, and higher flexibility of the CU partitioning shape is supported. In the encoding tree structure, a CU can have either a square or rectangular shape. The CTU is first partitioned by a quaternary-tree (also known as a quadtree) structure. Then, the leaf nodes of the quaternary-tree can be further partitioned by a multi-type tree structure.

[0042] Figures 4A to 4D are schematic diagrams showing the multi-type tree splitting mode according to some implementations of the present disclosure. As shown in Figures 4A to 4D, there are four splitting types in the multi-type tree structure, namely, vertical binary split 402 (SPLIT_BT_VER), horizontal binary split 404 (SPLIT_BT_HOR), vertical ternary split 406 (SPLIT_TT_VER), and horizontal ternary split 408 (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called CUs, and this segmentation is used for prediction and conversion processing without further segmentation, except when the CU is too large for the maximum transform length. This means that in most cases, CUs, PUs, and TUs have the same block size in a quad tree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color component of the CU.

[0043] VVC Syntax In VVC, the first layer of the syntax-signaling bitstream is the Network Abstraction Layer (NAL), where the bitstream is split into a set of NAL units. Some NAL units signal common control parameters such as the Sequence Parameter Set (SPS) and the Picture Parameter Set (PPS) to the decoder. Others contain video data. The Video Coding Layer (VCL) NAL units contain slices of the encoded video. The encoded picture is called an access unit and can be encoded as one or more slices.

[0044] The symbolized video sequence starts from an Instantaneous Decoder Refresh (IDR) picture. All subsequent video pictures are encoded as slices. A new IDR picture signals the end of the previous video segment and the start of a new one. Each NAL unit starts with a 1-byte header, followed by a Raw Byte Sequence Payload (RBSP). The RBSP contains the encoded slice. Since the slice is binary, it can be padded with zero bits to have a length that is an integer number of bytes. A slice consists of a slice header and slice data. The slice data is defined as a series of CUs.

[0045] The concept of a picture header that is transmitted once per picture as the first VCL NAL unit of the picture was adopted at the 16th JVET meeting. It was also proposed to group some syntax elements that were previously in the slice header into this picture header. Syntax elements that only need to be transmitted once per picture functionally can be moved to the picture header instead of being transmitted multiple times in the slices of a given picture.

[0046] In the VVC specification, the syntax table specifies the superset of the syntax of all allowed bitstreams. Additional constraints on the syntax can be specified directly or indirectly in other clauses. Table 1 below is the syntax table for the slice header and picture header in VVC. The semantics of some syntax are also shown after the syntax table.

Table 1

[0047] Semantics of the selected syntax element The ph_temporal_mvp_enabled_flag specifies whether a temporal motion vector predictor can be used for inter prediction of slices associated with PH. If the ph_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slice associated with PH are assumed to be constrained such that a temporal motion vector predictor is not used during slice decoding. Otherwise (if the ph_temporal_mvp_enabled_flag is equal to 1), a temporal motion vector predictor can be used during decoding of the slice associated with PH. If the value of the ph_temporal_mvp_enabled_flag does not exist, it is assumed to be equal to 0. If a reference picture having the same spatial resolution as the current picture is not in the Decoded Picture Buffer (DPB), the value of the ph_temporal_mvp_enabled_flag shall be equal to 0.

[0048] The maximum number of sub-block based merge MVP candidates, MaxNumSubblockMergeCand, is derived as follows. if(sps_affine_enabled_flag) MaxNumSubblockMergeCand = 5 - five_minus_max_num_subblock_merge_cand else MaxNumSubblockMergeCand = sps_sbtmvp_enabled_flag && ph_temporal_mvp_enabled_flag; Here, it is assumed that the value of MaxNumSubblockMergeCand is in the range from 0 to 5 (inclusive of both ends).

[0049] A slice_collocated_from_l0_flag equal to 1 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. A slice_collocated_from_l0_flag equal to 0 specifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1.

[0050] When slice_type is equal to B or P, ph_temporal_mvp_enabled_flag is equal to 1, and slice_collocated_from_l0_flag does not exist, the following applies. - When rpl_info_in_ph_flag is equal to 1, slice_collocated_from_l0_flag is presumed to be equal to ph_collocated_from_l0_flag. - Otherwise (when rpl_info_in_ph_flag is equal to 0 and slice_type is equal to P), the value of slice_collocated_from_l0_flag is presumed to be equal to 1.

[0051] slice_collocated_ref_idx specifies the reference index of the collocated picture used for temporal motion vector prediction.

[0052] When slice_type is equal to P, or when slice_type is equal to B and slice_collocated_from_l0_flag is equal to 1, slice_collocated_ref_idx refers to an entry in reference picture list 0, and the value of slice_collocated_ref_idx shall be in the range from 0 to NumRefIdxActive[0] - 1, inclusive.

[0053] When slice_type is equal to B and slice_collocated_from_l0_flag is equal to 0, slice_collocated_ref_idx shall refer to an entry in the reference picture list 1, and the value of slice_collocated_ref_idx shall be in the range from 0 to NumRefIdxActive[1] - 1, inclusive.

[0054] If slice_collocated_ref_idx does not exist, the following applies. - When rpl_info_in_ph_flag is equal to 1, the value of slice_collocated_ref_idx is presumed to be equal to ph_collocated_ref_idx. - Otherwise (when rpl_info_in_ph_flag is equal to 0), the value of slice_collocated_ref_idx is presumed to be equal to 0.

[0055] It is a bitstream compliance requirement that the picture referred to by slice_collocated_ref_idx must be the same for all slices of the encoded picture.

[0056] It is a bitstream compliance requirement that the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture referred to by slice_collocated_ref_idx shall be equal to the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively, and that RprConstraintsActive[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] shall be equal to 0.

[0057] The value of RprConstraintsActive[i][j] is derived in Section 8.3.2 of the VVC specification. The derivation of the value of RprConstraintsActive[i][j] is described below.

[0058] Decoding process for reference picture list construction The decoding process for reference picture list construction is called at the start of the decoding process for each slice of a non-IDR picture.

[0059] Reference pictures are addressed by reference indices. A reference index is an index into a reference picture list. When decoding an I slice, the reference picture list is not used during the decoding of slice data. When decoding a P slice, only reference picture list 0 (i.e., RefPicList[0]) is used during the decoding of slice data. When decoding a B slice, both reference picture list 0 and reference picture list 1 (i.e., RefPicList[1]) are used during the decoding of slice data.

[0060] At the start of the decoding process for each slice of a non-IDR picture, reference picture lists RefPicList[0] and RefPicList[1] are derived. The reference picture lists are used when marking reference pictures as specified in the video coding standard or when decoding slice data.

[0061] For an I slice of a non-IDR picture that is not the first slice of a picture, RefPicList[0] and RefPicList[1] may be derived for bitstream compliance checking purposes, but their derivation is not required for the decoding of the current picture or pictures later in the decoding order than the current picture. For a P slice that is not the first slice of a picture, RefPicList[1] may be derived for bitstream compliance checking purposes, but its derivation is not required for the decoding of the current picture or pictures later in the decoding order than the current picture.

[0062] With reference to RefPicList[0] and RefPicList[1], reference picture scaling ratios RefPicScale[i][j][0] and RefPicScale[i][j][1], and reference picture scaling flags RprConstraintsActive[0][j] and RprConstraintsActive[1][j], they are derived as follows.

Number

[0063] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the offsets applied to the picture size for scaling ratio calculation. If the values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset do not exist, they are assumed to be equal to pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset respectively.

[0064] It is assumed that the value of SubWidthC * (scaling_win_left_offset + scaling_win_right_offset) is less than pic_width_in_luma_samples, and the value of SubHeightC * (scaling_win_top_offset + scaling_win_bottom_offset) is less than pic_height_in_luma_samples.

[0065] The variables PicOutputWidthL and PicOutputHeightL are derived as follows. PicOutputWidthL = pic_width_in_luma_samples - SubWidthC * (scaling_win_right_offset + scaling_win_left_offset) PicOutputHeightL = pic_height_in_luma_samples - SubWidthC * (scaling_win_bottom_offset + scaling_win_top_offset)

[0066] Let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference picture of the current picture referring to this PPS, respectively. It is a requirement conforming to the bitstream that all of the following conditions are satisfied. - Assume that PicOutputWidthL * 2 is greater than or equal to refPicWidthInLumaSamples. - Assume that PicOutputHeightL * 2 is greater than or equal to refPicHeightInLumaSamples. - Assume that PicOutputWidthL is less than or equal to refPicWidthInLumaSamples * 8. - Assume that PicOutputHeightL is less than or equal to refPicHeightInLumaSamples * 8. - Assume that PicOutputWidthL * pic_width_max_in_luma_samples is greater than or equal to refPicOutputWidthL * (pic_width_in_luma_samples - Max(8, MinCbSizeY)). - Let PicOutputHeightL*pic_height_max_in_luma_samples be greater than or equal to refPicOutputHeightL*(pic_height_in_luma_samples - Max(8, MinCbSizeY)).

[0067] In the current VVC, mvd_l1_zero_flag is signaled in the PH without conditional constraints. However, the function controlled by the flag mvd_l1_zero_flag is applicable only when the slice is a bi-predictive slice (B slice). Therefore, when the slice associated with the picture header is not a B slice, the signaling of the flag is redundant.

[0068] In other examples, ph_disable_bdof_flag and ph_disable_dmvr_flag are signaled in the PH only when the corresponding enabling flags (sps_bdof_pic_present_flag, sps_dmvr_pic_present_flag) signaled in the sequence parameter set (SPS) are true. However, as shown in Table 2 below, the functions controlled by the flags ph_disable_bdof_flag and ph_disable_dmvr_flag are applicable only when the slice is a bi-predictive slice (B slice). Therefore, when the slice associated with the picture header is not a B slice, the signaling of these two flags is redundant or useless.

Table 2

[0069] The third problem is related to the syntax ph_temporal_mvp_enabled_flag. In the current VVC, since the resolution of the collocated pictures selected for temporal motion vector prediction (TMVP) derivation must be the same as that of the current picture, there are bitstream compliance constraints for checking the value of ph_temporal_mvp_enabled_flag as described below. If there is no reference picture with the same spatial resolution as the current picture in the DPB, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.

[0070] However, in the current VVC, not only does the resolution of the collocated pictures affect the enabling of TMVP, but also the offset applied to the picture size for scaling ratio calculation affects the enabling of TMVP. However, in the current VVC, the offset is not considered in terms of bitstream compliance of ph_temporal_mvp_enabled_flag.

[0071] Furthermore, there is a bitstream compliance requirement that the picture referred to by slice_collocated_ref_idx must be the same for all slices of the encoded picture. However, if the encoded picture has multiple slices and there is no common reference picture for all these slices, there is no chance to meet this bitstream compliance. Also, in such a case, ph_temporal_mvp_enabled_flag needs to be constrained to 0.

[0072] To address the above problems, several methods are proposed. Note that the proposed methods can be applied independently or in combination.

[0073] The functions controlled by the flags mvd_l1_zero_flag, ph_disable_bdof_flag, and ph_disable_dmvr_flag are only applicable when the slice is a bi-predicted slice (B slice). Therefore, according to the method of the present disclosure, it is proposed to signal these flags only when the associated slice is a B slice. It should be noted that when the reference picture list is signaled in PH (e.g., rpl_info_in_ph_flag = 1), this means that all slices of the encoded picture use the same reference picture that is signaled in PH. Therefore, when the reference picture list is signaled in PH and the current picture is signaled not to be bi-predicted as indicated by the signaled reference picture list, the flags mvd_l1_zero_flag, ph_disable_bdof_flag, and ph_disable_dmvr_flag do not need to be signaled.

[0074] In some examples, some conditions are added to the syntax set in PH to prevent redundant signaling or undefined decoding operations due to inappropriate values sent for a part of the syntax in the picture header. Some examples are shown below, where the variable num_ref_entries[i][RplsIdx[i]] represents the number of reference pictures in list i. If(rpl_info_in_ph_flag&&num_ref_entries[0][RplsIdx[0]]>1&&num_ref_entries[1][RplsIdx[1]]>1) mvd_l1_zero_flag; If(sps_bdof_pic_present_flag&&rpl_info_in_ph_flag&&num_ref_entries[0][RplsIdx[0]]>1&&num_ref_entries[1][RplsIdx[1]]>1) ph_disable_bdof_flag

[0075] In the current VVC, not only can the resolution of collocated pictures affect the enabling of TMVP, but also the offset applied to the picture size for scaling ratio calculation can affect the enabling of TMVP. However, in the current VVC, the offset is not considered in accordance with the bitstream for ph_temporal_mvp_enabled_flag. In some examples, as described below, it is proposed to add to the current VVC bitstream compliance constraints that require the value of ph_temporal_mvp_enabled_flag to depend on the offset applied to the picture size for scaling ratio calculation. When there is no reference picture in the DPB whose spatial resolution and the offset applied to the picture size for scaling ratio calculation are the same as those of the current picture, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.

[0076] The above bitstream compliance constraints can also be described in other ways as follows. When there is no reference picture in the DPB for which the associated variable value RprConstraintsActive[i][j] is equal to 0, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.

[0077] In the current VVC, there is a bitstream compliance requirement that the picture referred to by slice_collocated_ref_idx must be the same for all slices of the encoded picture. However, when the encoded picture has multiple slices and there is no common reference picture for all these slices, this bitstream compliance has no chance of being satisfied.

[0078] In some examples, the bitstream compliance requirements regarding ph_temporal_mvp_enabled_flag are changed to consider whether there is a common reference picture for all slices of the current picture.

[0079] The ph_temporal_mvp_enabled_flag specifies whether a temporal motion vector predictor can be used for inter prediction of slices associated with PH. If the ph_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slices associated with PH shall be constrained such that the temporal motion vector predictor is not used during slice decoding. Otherwise (if the ph_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used during decoding of the slices associated with PH. The value of the ph_temporal_mvp_enabled_flag, if not present, is assumed to be equal to 0. If there is no reference picture within the DPB that has the same spatial resolution as the current picture, the value of the ph_temporal_mvp_enabled_flag shall be equal to 0. If there is no common reference picture for all slices associated with PH, the value of the ph_temporal_mvp_enabled_flag shall be equal to 0.

[0080] In some examples, bitstream compliance regarding slice_collocated_ref_idx is simplified such that the bitstream compliance requirement is that RprConstraintsActive[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] shall be equal to 0.

[0081] FIG. 5 is a block diagram illustrating an exemplary apparatus for video encoding according to some implementations of the present disclosure. The apparatus 500 can be a terminal such as, for example, a cellular phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.

[0082] As shown in FIG. 5, device 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0083] The processing component 502 generally controls the overall operation of device 500, such as operations associated with display, telephone, data communication, camera operation, and recording operations. The processing component 502 may include one or more processors 520 for executing instructions to complete all or part of the steps of the above methods. Further, the processing component 502 may include one or more modules to facilitate interactions between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate interactions between the multimedia component 508 and the processing component 502.

[0084] The memory 504 is configured to store various types of data to support the operation of device 500. Examples of such data include instructions for any application or method operating on device 500, contact data, phone book data, messages, photos, videos, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, and the memory 504 may be a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or a compact disk.

[0085] The power supply component 506 supplies power to various components of the device 500. The power supply component 506 may include a power management system, one or more power supplies, and other components related to the generation, management, and distribution of power for the device 500.

[0086] The multimedia component 508 includes a screen that provides an output interface between the device 500 and the user. In some examples, the screen may include a liquid crystal display (LCD) and a touch panel (TP). When the screen includes a touch panel, the screen may be implemented as a touch screen that receives input signals from the user. The touch panel may include one or more touch sensors for sensing touches, swipes, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or swipe operation but also the duration and pressure associated with the touch or swipe operation. In some examples, the multimedia component 508 may include a front camera and / or a rear camera. When the device 500 is in an operation mode such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data.

[0087] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC). When the device 500 is in an operation mode such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received voice signals can be further stored in the memory 504 or transmitted via the communication component 516. In some examples, the audio component 510 further includes a speaker for outputting audio signals.

[0088] The I / O interface 512 provides an interface between the processing component 502 and the peripheral interface module. The peripheral interface module may be, for example, a keyboard, click wheel, buttons, etc. These buttons may include, but are not limited to, a home button, volume buttons, start button, and lock button.

[0089] The sensor component 514 includes one or more sensors for providing a state assessment of the device 500 in different modes. For example, the sensor component 514 may detect the on / off state of the device 500 and the relative positions of the components. For example, the components are the display and keypad of the device 500. The sensor component 514 may also detect a change in the position of the device 500 or components of the device 500, the presence or absence of user contact with the device 500, the orientation or acceleration / deceleration of the device 500, and a change in the temperature of the device 500. The sensor component 514 may include a proximity sensor configured to detect the presence of nearby objects without physical contact. The sensor component 514 may further include an optical sensor such as a CMOS or CCD image sensor used for imaging applications. In some examples, the sensor component 514 may further include an acceleration sensor, gyroscope sensor, magnetic sensor, pressure sensor, or temperature sensor.

[0090] The communication component 516 is configured to facilitate wired or wireless communication between the device 500 and other devices. The device 500 may access a wireless network based on communication standards such as, for example, WiFi, 4G, or a combination thereof. In one example, the communication component 516 may receive a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In one example, the communication component 516 may further include a Near Field Communication (NFC) module for facilitating short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, Ultra-Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0091] In one example, the device 500 may be implemented by one or more of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic elements for performing the above methods.

[0092] The non-transitory computer-readable storage medium may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a hybrid drive or a solid state hybrid drive (SSHD), a read only memory (ROM), a compact disc read only memory (CD-ROM), magnetic tape, a floppy disk, etc.

[0093] FIG. 6 is a flowchart showing an exemplary process of video encoding according to some implementations of the present disclosure.

[0094] In step 602, the processor 520 determines whether one or more reference picture lists are signaled in the PH associated with the picture, and whether one or more slices associated with the picture are indicated to be bi-predicted by one or more reference picture lists.

[0095] In step 604, in response to the processor 620 determining that one or more reference picture lists are signaled in the PH and that one or more slices are indicated not to be bi-predicted by one or more reference picture lists, the processor 620 adds one or more constraints to one or more syntax elements of the PH.

[0096] In some examples, the one or more constraints include skipping the parsing of one or more syntax elements.

[0097] In some examples, the one or more syntax elements include one or more flags applicable to one or more slices.

[0098] The processor 520 may further use an enabling flag such as the above-mentioned mvd_l1_zero_flag to determine whether the corresponding motion vector difference (MVD) encoding syntax structure is not parsed, and whether two variables are set to zero for one or more slices associated with the PH, where the two variables respectively identify the difference between the list vector component and the prediction corresponding to the list vector component.

[0099] In some examples, mvd_l1_zero_flag being equal to 1 indicates that the mvd_coding(x0,y0,1) syntax structure is not parsed, and for compIdx = 0 or 1, and cpIdx = 0, 1, or 2, MvdL1[x0][y0][compIdx] and MvdCpL1[x0][y0][cpIdx][compIdx] are set equal to 0. Further, mvd_l1_zero_flag being equal to 0 indicates that the mvd_coding(x0,y0,1) syntax structure is parsed. The mvd_coding(x0,y0,1) syntax structure is the corresponding MVD coding syntax structure. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture.

[0100] Further, the variable MvdLX[x0][y0][compIdx] where X is 0 or 1 identifies the difference between the list X vector component used and its prediction. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. compIdx = 0 is assigned to the horizontal motion vector component difference, and compIdx = 1 is assigned to the vertical motion vector component.

[0101] Further, the variable MvdCpLX[x0][y0][cpIdx][compIdx] where X is 0 or 1 identifies the difference between the list X vector component used and its prediction. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. The array index cpIdx identifies the index of the control point. compIdx = 0 is assigned to the horizontal motion vector component difference, and compIdx = 1 is assigned to the vertical motion vector component.

[0102] In response to determining that the activation flag is equal to 0, the processor 520 may further constrain one or more syntax elements such that the MVD coding syntax structure is parsed for one or more slices.

[0103] In response to determining that the activation flag is equal to 1, the processor 520 may further determine to skip parsing the MVD syntax structure when decoding one or more slices.

[0104] The processor 520 further uses an inactivation flag such as the above-mentioned ph_disable_bdof_flag to identify whether the bi-directional optical flow (BDOF) inter-prediction-based inter-biprediction for one or more slices associated with the PH is inactivated. In response to determining that the inactivation flag is equal to 0, the processor 520 may constrain one or more syntax elements such that the BDOF inter-prediction-based inter-biprediction is activated when decoding one or more slices. In response to determining that the inactivation flag is equal to 1, the processor 520 may inactivate the BDOF inter-prediction-based inter-biprediction when decoding one or more slices.

[0105] The processor 520 further uses an inactivation flag such as the above-mentioned ph_disable_dmvr_flag to identify whether the decoder motion vector refinement (DMVR) - based inter-biprediction for one or more slices associated with the PH is inactivated. In response to determining that the inactivation flag is equal to 0, the processor 520 may constrain one or more syntax elements such that the DMVR-based inter-biprediction is activated when decoding one or more slices. In response to determining that the inactivation flag is equal to 1, the processor 520 may inactivate the DMVR-based inter-biprediction when decoding one or more slices.

[0106] FIG. 7 is a flowchart showing an exemplary process of video encoding according to some implementations of the present disclosure.

[0107] At step 702, the processor 520 uses an enable flag to determine whether one or more temporal motion vector predictors are used for inter prediction of one or more slices associated with the PH of the picture.

[0108] At step 704, the processor 520 constrains the value of the enable flag according to a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0109] In response to determining that there is no reference picture in the DPB that has the same spatial resolution and the same offset as the picture, the processor 520 may set the enable flag to 0. Further, the offset may be applied to the size of the picture for scaling ratio calculation.

[0110] In response to determining that there is no common reference picture for one or more slices, the processor 520 may set the enable flag to 0.

[0111] In response to determining that there is no reference picture in the DPB for which the reference picture scaling flag is equal to 0, the processor 520 may set the enable flag to 0.

[0112] The processor 520 may derive a reference picture scaling flag based on a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0113] In some examples, an apparatus for video encoding is provided. The apparatus includes one or more processors 520 and a memory 504 configured to store instructions executable by the one or more processors. The processor is configured to execute the method shown in FIG. 6 when the instructions are executed.

[0114] In some examples, an apparatus for video encoding is provided. The apparatus includes one or more processors 520 and a memory 504 configured to store instructions executable by the one or more processors. The processors are configured to execute the method shown in FIG. 7 when the instructions are executed.

[0115] In some other examples, a non-transitory computer-readable storage medium 504 storing instructions is provided. When the instructions are executed by one or more processors 520, the instructions cause the processors to execute the method shown in FIG. 6.

[0116] In some other examples, a non-transitory computer-readable storage medium 504 storing instructions is provided. When the instructions are executed by one or more processors 520, the instructions cause the processors to execute the method shown in FIG. 7.

[0117] The description of the present disclosure is presented for illustrative purposes and is not intended to be exhaustive or to limit the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and associated drawings.

[0118] Examples have been chosen and described in order to explain the principles of the present disclosure and so that others skilled in the art may best utilize the present disclosure with respect to various implementations, with the underlying principles and various implementations with various modifications suitable for the intended particular uses. Accordingly, it is to be understood that the scope of the present disclosure should not be limited to the particular examples disclosed, and that modifications and other implementations are intended to be included within the scope of the present disclosure.

Claims

Claim 1 A method for video decoding, comprising: obtaining, by a decoder, an enabling flag that specifies whether one or more temporal motion vector predictors are enabled for inter prediction of one or more slices associated with a picture header (PH) of a picture; wherein, if a plurality of offsets applied to the size of the picture for scaling ratio calculation satisfy a first condition, the value of the enabling flag is equal to 0. Claim 2 The method according to claim 1, wherein in response to a determination that no reference picture exists in a decoded picture buffer (DPB) having the same spatial resolution and the same offset as the picture, the enabling flag is set to 0, and the offset is applied to the size of the picture for the scaling ratio calculation. Claim 3 The method according to claim 1, wherein in response to a determination that no common reference picture exists for the one or more slices, the enabling flag is set to 0. Claim 4 The method according to claim 1, wherein in response to a determination that no reference picture having a reference picture scaling flag equal to 0 exists in a decoded picture buffer (DPB), the enabling flag is set to 0. Claim 5 The method according to claim 4, wherein the reference picture scaling flag is derived based on the plurality of offsets applied to the size of the picture for the scaling ratio calculation. Claim 6 An apparatus for video decoding, comprising: one or more processors; and a memory configured to store instructions executable by the one or more processors, wherein the one or more processors are configured to execute the method according to any one of claims 1 to 5 when executing the instructions. Claim 7 A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to execute the method according to any one of claims 1 to 5. Claim 8 A computer program comprising a plurality of instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to any one of claims 1 to 5. Claim 9 A method of storing a bit stream decoded by the decoding method according to any one of claims 1 to 5.

10. A method of receiving a bit stream decoded by the decoding method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Decoding based on BI-directional picture condition

    WO2021201759A1