Method and apparatus for signaling syntax elements in video coding

By constraining syntax elements and managing temporal motion vector predictors based on bi-predictive slices and common reference pictures, the patent addresses inefficiencies in video coding standards, enhancing encoding efficiency and maintaining video quality at lower bitrates.

JP7717222B2Active Publication Date: 2025-08-01BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024076139
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-31
Filing Date
2024-05-08
Publication Date
2025-08-01
Estimated Expiration
2041-03-23

AI Technical Summary

Technical Problem

Existing video coding standards like VVC face issues with redundant signaling of syntax elements and inefficient use of temporal motion vector predictors due to unresolved bitstream compliance constraints, leading to unnecessary decoding operations and reduced encoding efficiency.

Method used

Implementing constraints on syntax elements in the picture header based on the presence of bi-predictive slices and common reference pictures, and using enable flags to manage temporal motion vector predictors, ensuring bitstream compliance and optimizing decoding processes.

Benefits of technology

Enhances encoding efficiency by reducing redundant signaling and improving decoding accuracy, thereby maintaining video quality at lower bitrates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007717222000009
    Figure 0007717222000009
  • Figure 0007717222000010
    Figure 0007717222000010
  • Figure 0007717222000011
    Figure 0007717222000011
Patent Text Reader

Abstract

To provide methods and apparatuses for signaling of syntax elements in video coding.SOLUTION: The methods include a decoder determining whether one or more reference picture lists are signaled in a picture header (PH) associated with a picture and whether the one or more reference picture lists indicate that one or more slices associated with the picture are bi-predictive. The methods further include adding one or more constraints to one or more syntax elements in the PH by the decoder in response to determining that the one or more reference picture lists are signaled in the PH and that the one or more reference picture lists indicate that the one or more slices are not bi-predictive.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 003,226, titled "Signaling of Syntax Elements in Video Coding," filed on Mar. 31, 2020, the entire content of which is hereby incorporated by reference for all purposes.

[0002] This disclosure relates to video encoding and compression, and more particularly, but not limited to, methods and apparatus for signaling syntax elements in video encoding.

Background Art

[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, currently, some well-known video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC, also known as H.265 or MPEG-H Part 2), Advanced Video Coding (AVC, also known as H.264 or MPEG-4 Part 10), which are jointly developed by ISO / IEC MPEG and ITU-T VECG. AOMedia Video 1 (AV1) was developed by the Alliance for Open Media (AOM) as a successor to its predecessor VP9. Audio Video Coding (AVS), which refers to digital audio and digital video compression standards, is another series of video compression standards developed by the Audio and Video Coding Standard Working Group in China. Most of the existing video coding standards are based on well-known hybrid video coding frameworks, that is, they use block-based prediction methods (e.g., inter prediction, intra prediction) to reduce the redundancy present in video images or sequences, and use transform coding to compact the energy of prediction errors. An important goal of video coding technology is to compress video data into a form that uses a lower bitrate while avoiding or minimizing the degradation of video quality. Summary of the Invention Problems to be Solved by the Invention

[0004] This disclosure provides examples of techniques related to the signaling of syntax elements in video coding. Means for Solving the Problems

[0005] According to a first aspect of the present disclosure, a method for video encoding is provided. The method includes a decoder determining whether one or more reference picture lists are signaled by a picture header (PH) associated with a picture, and whether one or more reference picture lists indicate that one or more slices associated with the picture are bi-predictive. Further, the method includes the decoder adding one or more constraints to one or more syntax elements of the PH in response to determining that one or more reference picture lists are signaled by the PH and that one or more reference picture lists indicate that one or more slices are not bi-predictive.

[0006] According to a second aspect of the present disclosure, a method for video encoding is provided. The method includes a decoder using an enable flag to identify whether one or more temporal motion vector predictors are used for inter prediction of one or more slices associated with the PH of a picture. The method further includes the decoder constraining the value of the enable flag according to a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0007] According to a third aspect of the present disclosure, an apparatus for video encoding is provided. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. When executing the instructions, the one or more processors are configured to execute determining whether one or more reference picture lists are signaled in PH associated with a picture and whether one or more reference picture lists indicate that one or more slices associated with the picture are bi-predicted. Further, in response to determining that one or more reference picture lists are signaled in PH and one or more reference picture lists indicate that one or more slices are not bi-predicted, the one or more processors are configured to execute adding one or more constraints to one or more syntax elements of PH.

[0008] According to a fourth aspect of the present disclosure, an apparatus for video encoding is provided. The apparatus includes one or more processors and a memory configured to store instructions executable by the one or more processors. When executing the instructions, the one or more processors are configured to execute using an enable flag to identify whether one or more temporal motion vector predictors are used for inter-prediction of one or more slices associated with the PH of a picture. The one or more processors are further configured to execute constraining the value of the enable flag according to a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0009] According to a fifth aspect of the present disclosure, a non-transitory computer-readable storage medium for video encoding storing computer-executable instructions is provided. When executed by one or more computer processors, the instructions cause the one or more computer processors to execute a method for video encoding according to the first aspect of the present disclosure.

[0010] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium for video encoding that stores computer-executable instructions is provided. The instructions, when executed by one or more computer processors, cause the one or more computer processors to execute a method for video encoding according to a second aspect of the present disclosure.

[0011] A more specific description of examples of the present disclosure will be made with reference to specific examples shown in the accompanying drawings. Considering that these drawings show only some examples and are not considered to limit the scope, the examples will be described and explained more specifically and in detail using the accompanying drawings.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4A

Figure 4B

Figure 4C

Figure 4D

Figure 5

Figure 6

Figure 7

[0013] References to specific implementations are made herein in detail, and examples thereof are shown in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in the understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternative examples may be used. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented in many types of electronic devices having digital video capabilities.

[0014] References throughout this disclosure to "one embodiment", "an embodiment", "an example", "some embodiments", "some examples", or similar language mean that a particular feature, structure, or characteristic described is included in at least one embodiment or example. Features, structures, elements, or characteristics described in connection with one or some embodiments are applicable to other embodiments as well, unless otherwise specified.

[0015] Throughout this disclosure, terms such as "first", "second", "third", etc. are used throughout solely as a nomenclature for referring to related elements, such as devices, components, compositions, steps, etc., without implying any spatial or temporal order, unless otherwise specified. For example, "a first device" and "a second device" may refer to two separately formed devices, or two parts, components, or operating states of the same device, and may be arbitrarily named.

[0016] The terms "module", "sub-module", "circuit", "sub-circuit", "circuitry", "sub-circuitry", "unit", or "sub-unit" may include memory (shared, dedicated, or group) that stores code or instructions executable by one or more processors. A module may include one or more circuits with or without the stored code or instructions. A module or circuit may include one or more components connected directly or indirectly. These components may or may not be physically attached to or adjacent to each other.

[0017] As used herein, the terms "if" or "when" may be understood to mean "upon" or "in response to" depending on the context. These terms may not indicate that the associated limitations or features are conditional or optional when they appear in a claim. For example, a method may include: i) a step of performing function or action X' when or if condition X exists, and ii) a step of performing function or action Y' when or if condition Y exists. This method may be implemented with both the ability to perform function or action X' and the ability to perform function or action Y'. Thus, functions X' and Y' may both be performed at each time during multiple executions of this method.

[0018] A unit or module may be implemented purely by software, purely by hardware, or by a combination of hardware and software. In a pure software implementation, for example, a unit or module may include functionally related code blocks or software components directly or indirectly linked to each other to perform a particular function.

[0019] FIG. 1 shows a block diagram illustrating an exemplary block-based hybrid video encoder 100 that can be used in conjunction with many video coding standards that use block-based processing. In encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter prediction approach or an intra prediction approach. In inter prediction, one or more predictors are formed based on pixels of a previously reconstructed frame by motion estimation and motion compensation. In intra prediction, a predictor is formed based on reconstructed pixels within the current frame. By mode decision, the best predictor can be selected to predict the current block.

[0020] The prediction residual, representing the difference between the current video block and its predictor, is sent to transformation circuit 102. The transformation coefficients are then sent from transformation circuit 102 to quantization circuit 104 for entropy reduction. The quantized coefficients are then supplied to entropy encoding circuit 106 to generate a compressed video bit stream. As shown in FIG. 1, prediction-related information 110 from inter prediction circuit and / or intra prediction circuit 112, such as video block partitioning information, motion vectors, reference picture indices, and intra prediction modes, etc., is also supplied through entropy encoding circuit 106 and stored in the compressed video bit stream 114.

[0021] In encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed through inverse quantization 116 and inverse transformation circuit 118. This reconstructed prediction residual is combined with block predictor 120 to generate the unfiltered reconstructed pixels of the current video block.

[0022] Intra prediction (also called "spatial prediction") predicts the current video block using pixels from samples of already encoded neighboring blocks (referred to as reference samples) within the same video picture and / or slice. Spatial prediction reduces the spatial redundancy inherent in the video signal.

[0023] Inter prediction (also called "temporal prediction") predicts the current video block using reconstructed pixels from already encoded video pictures. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given coding unit (CU) or coded block is typically signaled by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Further, if multiple reference pictures are supported, an additional reference picture index is sent, which is used to identify which reference picture in the reference picture store the temporal prediction signal is from.

[0024] After spatial prediction and / or temporal prediction is performed, the intra / inter-mode decision circuit 121 within the encoder 100 selects the best prediction mode, for example, based on the rate distortion optimization method. Next, the block predictor 120 is subtracted from the current video block, and the resulting prediction residual is decorrelated using the transform circuit 102 and quantization circuit 104. The resulting quantized residual coefficients are inverse quantized by the inverse quantization circuit 116 and inverse transformed by the inverse transform circuit 118 to form a reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, in-loop filtering 115 such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF) can be applied to the reconstructed CU, and then stored in the reference picture store of the picture buffer 117 and used to encode future video blocks. To form the output video bitstream 114, the encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy encoding unit 106 to be further compressed and packed to form a bitstream.

[0025] For example, in the latest versions of AVC, HEVC, and VVC, a deblocking filter is available. In HEVC, an additional in-loop filter called SAO (sample adaptive offset) is defined to further improve the encoding efficiency. In the latest version of the VVC standard, yet another in-loop filter called ALF (adaptive loop filter) is being actively studied and has a good chance of being included in the final standard.

[0026] These in-loop filter operations are optional. Performing these operations helps to improve the encoding efficiency and visual quality. They can also be turned off as a decision made by the encoder 100 to save computational complexity.

[0027] Intra prediction is typically based on the non-filtered reconstructed pixels, and it should be noted that inter prediction is based on the filtered reconstructed pixels if these filtering options are enabled by the encoder 100.

[0028] FIG. 2 is a block diagram showing an exemplary block-based video decoder 200 that can be used in conjunction with many video coding standards. This decoder 200 is similar to the reconstruction-related part present in the encoder 100 of FIG. 1. In the decoder 200, the incoming video bitstream 201 is first decoded by the entropy decoder 202 to derive the quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by the inverse quantization 204 and the inverse transform 206 to obtain the reconstructed prediction residual. The block prediction mechanism implemented in the intra / inter mode selector 212 is configured to perform either intra prediction 208 or motion compensation 210 based on the decoded prediction information. A set of non-filtered reconstructed pixels is obtained by summing the reconstructed prediction residual from the inverse transform 206 and the prediction output generated by the block prediction mechanism using a summer 214.

[0029] The reconstructed block may further pass through the in-loop filter 209 and is then stored in the picture buffer 213 that functions as a reference picture store. The reconstructed video in the picture buffer 213 is sent to drive the display device and can be used to predict future video blocks. In the situation where the in-loop filter 209 is enabled, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.

[0030] The aforementioned video encoding / decoding standards, such as VVC, JEM, HEVC, MPEG-4, Part 10, are conceptually similar. For example, they all use block-based processing. The block partitioning methods in some standards will be described in detail below.

[0031] High Efficiency Video Coding (HEVC) HEVC is based on a hybrid block-based motion compensated transform coding architecture. The basic unit of compression is called a CTU. The maximum CTU size is defined as two blocks of up to 64×64 luma pixels and 32×32 chroma pixels for the 4:2:0 chroma format. Each CTU may contain one CU or may be recursively divided into four smaller CUs until the pre-defined minimum CU size is reached. Each CU (also named leaf CU) contains a tree of one or more prediction units (PU: prediction unit) and transform units (TU: transform unit).

[0032] Generally, except for monochrome content, a CTU may contain one luma coding tree block (CTB: coding tree block) and two corresponding chroma CTBs, a CU may contain one luma coding block (CB: coding block) and two corresponding chroma CBs, a PU may contain one luma prediction block (PB: prediction block) and two corresponding chroma PBs, and a TU may contain one luma transform block (TB: transform block) and two corresponding chroma TBs. However, exceptions can occur because the minimum TB size is 4×4 for both luma and chroma (i.e., 2×2 chroma TBs are not supported for the 4:2:0 color format), and each intra-chroma CB always has exactly one intra-chroma PB regardless of the number of intra-luma PBs within the corresponding intra-luma CB.

[0033] In the case of an intra CU, the luma CB can be predicted by one or four luma PBs, and each of the two chroma CBs is always predicted by one chroma PB, where each luma PB has one intra-luma prediction mode, and the two chroma PBs share one intra-chroma prediction mode. Further, in the case of an intra CU, the TB size cannot be made larger than the PB size. In each PB, intra prediction is applied to predict the samples of each TB within the PB from the reconstructed samples in the vicinity of that TB. In each PB, in addition to 33 directed intra prediction modes, the DC mode and the planar mode are also supported, predicting flat regions and gradually changing regions respectively.

[0034] For each inter PU, one of three prediction modes including inter, skip, and merge can be selected. Generally speaking, a motion vector competition (MVC) scheme is introduced to select motion candidates from a given candidate set including spatial and temporal motion candidates. Multiple references for motion estimation make it possible to find the best reference in two possible reconstructed reference picture lists (i.e., list 0 and list 1). In the inter mode (referred to as the AMVP mode, where AMVP represents advanced motion vector prediction), an inter prediction indicator (list 0, list 1, or bidirectional prediction), a reference index, a motion candidate index, a motion vector difference (MVD), and a prediction residual are transmitted. For the skip mode and the merge mode, only the merge index is transmitted, and the current PU inherits the inter prediction indicator, the reference index, and the motion vector from a neighboring PU referred to by the encoded merge index. In the case of a skip-encoded CU, the residual signal is also omitted.

[0035] Versatile Video Coding (VVC) At the 10th JVET meeting held in San Diego, USA from April 10 to 20, 2018, JVET defined the first draft of Versatile Video Coding (VVC) and its reference software implementation, the VVC Test Model 1 (VTM1). As the first new coding feature of VVC, it was decided to include a quad tree with nested multi-type trees. The multi-type tree is an encoding block partitioning structure that includes both binary and ternary partitions. Since then, a reference software VTM that implements both encoding and decoding processes has been developed and updated through subsequent JVET meetings.

[0036] In VVC, a picture of the input video is partitioned into blocks called Coding Tree Units (CTUs). A CTU is divided into Coding Units (CUs) using a quad tree with a nested multi-type tree structure, and a CU defines a region of pixels that share the same prediction mode (e.g., intra or inter). The term "unit" can define a region of an image that covers all components such as luma and chroma. The term "block" can be used to define a region that covers a specific component (e.g., luma), and considering chroma sampling formats such as 4:2:0, blocks of different components (e.g., luma vs. chroma) can have different spatial positions.

[0037] Partitioning of a Picture into CTUs FIG. 3 shows an example of a picture 300 divided into a plurality of CTUs 302 according to some implementations of the present disclosure.

[0038] In VCC, a picture is partitioned into a sequence of CTUs. The concept of CTU is the same as that of HEVC. For a picture with three sample arrays, a CTU is composed of a block of N×N luma samples and two corresponding blocks of chroma samples.

[0039] The maximum allowable size of the luma block within a CTU is specified as 128×128 (however, the maximum size of the luma transform block is 64×64).

[0040] Partitioning of CTUs Using a Tree Structure In HEVC, to adapt to various local characteristics, CTUs are divided into CUs using a quaternary-tree structure called an encoding tree. The decision of whether to use inter-picture (temporal) or intra-picture (spatial) prediction to encode a picture area is made at the leaf CU level. Each leaf CU can be further divided into one, two, or four PUs depending on the PU partitioning type. Within one PU, the same prediction process is applied, and relevant information is sent to the decoder for each PU. After obtaining the residual block by applying the prediction process based on the PU partitioning type, the leaf CU can be divided into transform units (TUs) according to another quaternary-tree structure similar to the encoding tree of that CU. One of the important features of the HEVC structure is having the concept of multiple partitioning including CUs, PUs, and TUs.

[0041] In VVC, a quadtree with a nested multi-type tree using a binary and ternary segmentation structure replaces the concept of multiple partitioning unit types, that is, excluding CUs with sizes too large for the maximum transform length as appropriate, removing the distinction between the concepts of CUs, PUs, and TUs, and supporting higher flexibility in CU partitioning shapes. In the encoding tree structure, a CU can have either a square or rectangular shape. A CTU is first divided by a quaternary-tree (also known as a quadtree) structure. Then, the leaf nodes of the quaternary tree can be further divided by a multi-type tree structure.

[0042] Figures 4A to 4D are schematic diagrams showing the multi-type tree splitting mode according to some implementations of the present disclosure. As shown in Figures 4A to 4D, there are four splitting types in the multi-type tree structure, namely, vertical binary split 402 (SPLIT_BT_VER), horizontal binary split 404 (SPLIT_BT_HOR), vertical ternary split 406 (SPLIT_TT_VER), and horizontal ternary split 408 (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called CUs, and this segmentation is used for prediction and conversion processing without further division, except when the CU is too large for the maximum transform length. This means that, in most cases, CUs, PUs, and TUs have the same block size in a quad tree with a nested multi-type tree coding block structure. An exception occurs when the maximum supported transform length is smaller than the width or height of the color component of the CU.

[0043] Syntax of VVC In VVC, the first layer of the syntax-signaling bitstream is the Network Abstraction Layer (NAL), where the bitstream is split into a set of NAL units. Some NAL units signal common control parameters, such as the Sequence Parameter Set (SPS) and the Picture Parameter Set (PPS), to the decoder. Others contain video data. The Video Coding Layer (VCL) NAL units contain slices of the encoded video. The encoded picture is called an access unit and can be encoded as one or more slices.

[0044] The symbolized video sequence starts from an Instantaneous Decoder Refresh (IDR) picture. All subsequent video pictures are encoded as slices. A new IDR picture signals the end of the previous video segment and the start of a new one. Each NAL unit starts with a 1-byte header, followed by a Raw Byte Sequence Payload (RBSP). The RBSP contains the encoded slice. Since the slice is byte-aligned, it can be padded with zero bits to have an integer number of bytes. A slice consists of a slice header and slice data. The slice data is defined as a series of CUs.

[0045] The concept of a picture header that is sent once per picture as the first VCL NAL unit of the picture was adopted at the 16th JVET meeting. It was also proposed to group some syntax elements that were previously in the slice header into this picture header. Syntax elements that only need to be sent once per picture functionally can be moved to the picture header instead of being sent multiple times in the slices of a given picture.

[0046] In the VVC specification, the syntax table specifies the superset of the syntax of all allowed bitstreams. Additional constraints on the syntax can be specified directly or indirectly in other clauses. Table 1 below is the syntax table for the slice header and picture header in VVC. The semantics of some syntax are also shown after the syntax table.

Table 1

[0047] Semantics of the selected syntax element The ph_temporal_mvp_enabled_flag specifies whether a temporal motion vector predictor can be used for inter prediction of slices associated with PH. If the ph_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slice associated with PH shall be constrained such that the temporal motion vector predictor is not used during decoding of the slice. Otherwise (if the ph_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used during decoding of the slice associated with PH. If the value of the ph_temporal_mvp_enabled_flag does not exist, it is assumed to be equal to 0. If there is no reference picture in the Decoded Picture Buffer (DPB) that has the same spatial resolution as the current picture, the value of the ph_temporal_mvp_enabled_flag shall be equal to 0.

[0048] The maximum number of subblock-based merge MVP candidates MaxNumSubblockMergeCand is derived as follows. if(sps_affine_enabled_flag) MaxNumSubblockMergeCand = 5 - five_minus_max_num_subblock_merge_cand else MaxNumSubblockMergeCand = sps_sbtmvp_enabled_flag && ph_temporal_mvp_enabled_flag; Here, it is assumed that the value of MaxNumSubblockMergeCand is within the range from 0 to 5 (including both ends).

[0049] A slice_collocated_from_l0_flag equal to 1 identifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 0. A slice_collocated_from_l0_flag equal to 0 identifies that the collocated picture used for temporal motion vector prediction is derived from reference picture list 1.

[0050] When slice_type is equal to B or P, ph_temporal_mvp_enabled_flag is equal to 1, and slice_collocated_from_l0_flag does not exist, the following applies. - When rpl_info_in_ph_flag is equal to 1, slice_collocated_from_l0_flag is assumed to be equal to ph_collocated_from_l0_flag. - Otherwise (when rpl_info_in_ph_flag is equal to 0 and slice_type is equal to P), the value of slice_collocated_from_l0_flag is assumed to be equal to 1.

[0051] slice_collocated_ref_idx identifies the reference index of the collocated picture used for temporal motion vector prediction.

[0052] When slice_type is equal to P, or when slice_type is equal to B and slice_collocated_from_l0_flag is equal to 1, slice_collocated_ref_idx refers to an entry in reference picture list 0, and the value of slice_collocated_ref_idx shall be in the range from 0 to NumRefIdxActive[0] - 1, inclusive at both ends.

[0053] When slice_type is equal to B and slice_collocated_from_l0_flag is equal to 0, slice_collocated_ref_idx shall refer to an entry in reference picture list 1, and the value of slice_collocated_ref_idx shall be in the range from 0 to NumRefIdxActive[1] - 1, inclusive.

[0054] If slice_collocated_ref_idx does not exist, the following applies. - When rpl_info_in_ph_flag is equal to 1, the value of slice_collocated_ref_idx is assumed to be equal to ph_collocated_ref_idx. - Otherwise (when rpl_info_in_ph_flag is equal to 0), the value of slice_collocated_ref_idx is assumed to be equal to 0.

[0055] It is a bitstream compliance requirement that the picture referred to by slice_collocated_ref_idx must be the same for all slices of the coded picture.

[0056] It is a bitstream compliance requirement that the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the reference picture referred to by slice_collocated_ref_idx must be equal to the values of pic_width_in_luma_samples and pic_height_in_luma_samples of the current picture, respectively, and that RprConstraintsActive[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] must be equal to 0.

[0057] The value of RprConstraintsActive[i][j] is derived in Section 8.3.2 of the VVC specification. The derivation of the value of RprConstraintsActive[i][j] is described below.

[0058] Decoding process for reference picture list construction The decoding process for reference picture list construction is called at the start of the decoding process for each slice of a non-IDR picture.

[0059] Reference pictures are addressed by reference indices. A reference index is an index into a reference picture list. When decoding an I slice, the reference picture list is not used during the decoding of slice data. When decoding a P slice, only reference picture list 0 (i.e., RefPicList[0]) is used during the decoding of slice data. When decoding a B slice, both reference picture list 0 and reference picture list 1 (i.e., RefPicList[1]) are used during the decoding of slice data.

[0060] At the start of the decoding process for each slice of a non-IDR picture, reference picture lists RefPicList[0] and RefPicList[1] are derived. The reference picture lists are used when marking reference pictures as specified in the video coding standard or when decoding slice data.

[0061] For an I slice of a non-IDR picture that is not the first slice of the picture, RefPicList[0] and RefPicList[1] may be derived for bitstream compliance checking purposes, but their derivation is not required for the decoding of the current picture or pictures after the current picture in decoding order. For a P slice that is not the first slice of the picture, RefPicList[1] may be derived for bitstream compliance checking purposes, but its derivation is not required for the decoding of the current picture or pictures after the current picture in decoding order.

[0062] With reference to RefPicList[0] and RefPicList[1], reference picture scaling ratios RefPicScale[i][j][0] and RefPicScale[i][j][1], and reference picture scaling flags RprConstraintsActive[0][j] and RprConstraintsActive[1][j], they are derived as follows.

Number

[0063] scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset specify the offsets applied to the picture size for scaling ratio calculation. If the values of scaling_win_left_offset, scaling_win_right_offset, scaling_win_top_offset, and scaling_win_bottom_offset do not exist, they are presumed to be equal to pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset, respectively.

[0064] The value of SubWidthC * (scaling_win_left_offset + scaling_win_right_offset) is assumed to be less than pic_width_in_luma_samples, and the value of SubHeightC * (scaling_win_top_offset + scaling_win_bottom_offset) is assumed to be less than pic_height_in_luma_samples.

[0065] The variables PicOutputWidthL and PicOutputHeightL are derived as follows. PicOutputWidthL = pic_width_in_luma_samples - SubWidthC * (scaling_win_right_offset + scaling_win_left_offset) PicOutputHeightL = pic_height_in_luma_samples - SubWidthC * (scaling_win_bottom_offset + scaling_win_top_offset)

[0066] Let refPicOutputWidthL and refPicOutputHeightL be the PicOutputWidthL and PicOutputHeightL of the reference picture of the current picture referring to this PPS, respectively. It is a requirement compliant with the bitstream that all of the following conditions are satisfied. - Assume that PicOutputWidthL * 2 is greater than or equal to refPicWidthInLumaSamples. - Assume that PicOutputHeightL * 2 is greater than or equal to refPicHeightInLumaSamples. - Assume that PicOutputWidthL is less than or equal to refPicWidthInLumaSamples * 8. - Assume that PicOutputHeightL is less than or equal to refPicHeightInLumaSamples * 8. - Assume that PicOutputWidthL * pic_width_max_in_luma_samples is greater than or equal to refPicOutputWidthL * (pic_width_in_luma_samples - Max(8, MinCbSizeY)). - Let PicOutputHeightL*pic_height_max_in_luma_samples be greater than or equal to refPicOutputHeightL*(pic_height_in_luma_samples - Max(8, MinCbSizeY)).

[0067] In the current VVC, mvd_l1_zero_flag is signaled in the PH without conditional constraints. However, the function controlled by the flag mvd_l1_zero_flag is applicable only when the slice is a bi-predictive slice (B slice). Therefore, when the slice associated with the picture header is not a B slice, the signaling of the flag is redundant.

[0068] In other examples, ph_disable_bdof_flag and ph_disable_dmvr_flag are signaled in the PH only when the corresponding enabling flags (sps_bdof_pic_present_flag, sps_dmvr_pic_present_flag) signaled in the sequence parameter set (SPS) are true. However, as shown in Table 2 below, the functions controlled by the flags ph_disable_bdof_flag and ph_disable_dmvr_flag are applicable only when the slice is a bi-predictive slice (B slice). Therefore, when the slice associated with the picture header is not a B slice, the signaling of these two flags is redundant or useless.

Table 2

[0069] The third problem is related to the syntax ph_temporal_mvp_enabled_flag. In the current VVC, since the resolution of the collocated pictures selected for temporal motion vector prediction (TMVP) derivation must be the same as that of the current picture, there are bitstream compliance constraints for checking the value of ph_temporal_mvp_enabled_flag as described below. If there is no reference picture with the same spatial resolution as the current picture in the DPB, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.

[0070] However, in the current VVC, not only does the resolution of the collocated pictures affect the enabling of TMVP, but also the offset applied to the picture size for scaling ratio calculation affects the enabling of TMVP. However, in the current VVC, the offset is not considered in terms of bitstream compliance of ph_temporal_mvp_enabled_flag.

[0071] Furthermore, there is a bitstream compliance requirement that the picture referred to by slice_collocated_ref_idx must be the same for all slices of the encoded picture. However, if the encoded picture has multiple slices and there is no common reference picture for all these slices, there is no chance to meet this bitstream compliance. Also, in such a case, ph_temporal_mvp_enabled_flag needs to be restricted to 0.

[0072] To address the above problems, several methods are proposed. Note that the proposed methods can be applied independently or in combination.

[0073] The functions controlled by the flags mvd_l1_zero_flag, ph_disable_bdof_flag, and ph_disable_dmvr_flag are only applicable when the slice is a bi-predicted slice (B slice). Therefore, according to the method of the present disclosure, it is proposed to signal these flags only when the associated slice is a B slice. Note that when the reference picture list is signaled in PH (e.g., rpl_info_in_ph_flag = 1), this means using the same reference picture for which all slices of the encoded picture are signaled in PH. Thus, when the reference picture list is signaled in PH and the current picture is not bi-predicted as indicated by the signaled reference picture list, the flags mvd_l1_zero_flag, ph_disable_bdof_flag, and ph_disable_dmvr_flag need not be signaled.

[0074] In some examples, some conditions are added to the syntax set in PH to prevent redundant signaling or undefined decoding operations due to inappropriate values sent for a part of the syntax within the picture header. Some examples are shown below, where the variable num_ref_entries[i][RplsIdx[i]] represents the number of reference pictures in list i. If(rpl_info_in_ph_flag&&num_ref_entries[0][RplsIdx[0]]>1&&num_ref_entries[1][RplsIdx[1]]>1) mvd_l1_zero_flag; If(sps_bdof_pic_present_flag&&rpl_info_in_ph_flag&&num_ref_entries[0][RplsIdx[0]]>1&&num_ref_entries[1][RplsIdx[1]]>1) ph_disable_bdof_flag

[0075] In the current VVC, not only can the resolution of collocated pictures affect the activation of TMVP, but also the offset applied to the picture size for scaling ratio calculation can affect the activation of TMVP. However, in the current VVC, the offset is not considered in accordance with the bitstream for ph_temporal_mvp_enabled_flag. In some examples, as described below, it is proposed to add to the current VVC bitstream compliance constraints that require the value of ph_temporal_mvp_enabled_flag to depend on the offset applied to the picture size for scaling ratio calculation. When there is no reference picture in the DPB whose spatial resolution and offset applied to the picture size for scaling ratio calculation are the same as those of the current picture, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.

[0076] The above bitstream compliance constraints can also be described in other ways as follows. When there is no reference picture in the DPB for which the associated variable value RprConstraintsActive[i][j] is equal to 0, the value of ph_temporal_mvp_enabled_flag shall be equal to 0.

[0077] In the current VVC, there is a bitstream compliance requirement that the picture referred to by slice_collocated_ref_idx must be the same for all slices of the encoded picture. However, if the encoded picture has multiple slices and there is no common reference picture for all these slices, this bitstream compliance has no chance of being met.

[0078] In some examples, the bitstream compliance requirements regarding ph_temporal_mvp_enabled_flag are changed to consider whether there is a common reference picture for all slices of the current picture.

[0079] The ph_temporal_mvp_enabled_flag specifies whether a temporal motion vector predictor can be used for inter prediction of slices associated with PH. If the ph_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the slice associated with PH shall be constrained such that the temporal motion vector predictor is not used during slice decoding. Otherwise (if the ph_temporal_mvp_enabled_flag is equal to 1), the temporal motion vector predictor may be used during decoding of the slice associated with PH. If the value of the ph_temporal_mvp_enabled_flag does not exist, it is assumed to be equal to 0. If there is no reference picture in the DPB that has the same spatial resolution as the current picture, the value of the ph_temporal_mvp_enabled_flag shall be equal to 0. If there is no common reference picture for all slices associated with PH, the value of the ph_temporal_mvp_enabled_flag shall be equal to 0.

[0080] In some examples, bitstream compliance for slice_collocated_ref_idx is simplified such that the bitstream compliance requirement is that RprConstraintsActive[slice_collocated_from_l0_flag?0:1][slice_collocated_ref_idx] must be equal to 0.

[0081] FIG. 5 is a block diagram illustrating an exemplary apparatus for video encoding according to some implementations of the present disclosure. The apparatus 500 can be a terminal such as, for example, a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.

[0082] As shown in FIG. 5, the apparatus 500 may include one or more of the following components: a processing component 502, a memory 504, a power component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.

[0083] The processing component 502 generally controls the overall operation of the apparatus 500, such as operations associated with display, telephone, data communication, camera operation, and recording operations. The processing component 502 may include one or more processors 520 for executing instructions to complete all or part of the steps of the above methods. Further, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.

[0084] The memory 504 is configured to store various types of data to support the operation of the apparatus 500. Examples of such data include instructions for any application or method operating on the apparatus 500, contact data, phone book data, messages, photos, videos, etc. The memory 504 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, and the memory 504 may be a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or a compact disk.

[0085] The power component 506 supplies power to various components of the device 500. The power component 506 may include a power management system, one or more power sources, and other components related to the generation, management, and distribution of power for the device 500.

[0086] The multimedia component 508 includes a screen that provides an output interface between the device 500 and the user. In some examples, the screen may include a liquid crystal display (LCD) and a touch panel (TP). When the screen includes a touch panel, the screen may be implemented as a touch screen that receives input signals from the user. The touch panel may include one or more touch sensors for sensing touches, slides, and gestures on the touch panel. The touch sensors can detect not only the boundaries of a touch or slide operation but also the duration and pressure associated with the touch or slide operation. In some examples, the multimedia component 508 may include a front camera and / or a rear camera. When the device 500 is in an operation mode such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data.

[0087] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC). When the device 500 is in an operation mode such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received voice signals can be further stored in the memory 504 or transmitted via the communication component 516. In some examples, the audio component 510 further includes a speaker for outputting audio signals.

[0088] The I / O interface 512 provides an interface between the processing component 502 and the peripheral interface module. The above-mentioned peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0089] The sensor component 514 includes one or more sensors for providing state evaluation in different modes of the device 500. For example, the sensor component 514 can detect the on / off state of the device 500 and the relative positions of the components. For example, the components are the display and the keypad of the device 500. The sensor component 514 can also detect a change in the position of the device 500 or the components of the device 500, the presence or absence of user contact with the device 500, the orientation or acceleration / deceleration of the device 500, and a change in the temperature of the device 500. The sensor component 514 can include a proximity sensor configured to detect the presence of nearby objects without physical contact. The sensor component 514 can further include an optical sensor such as a CMOS or CCD image sensor used for imaging applications. In some examples, the sensor component 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0090] The communication component 516 is configured to facilitate wired or wireless communication between the device 500 and other devices. The device 500 can access a wireless network based on communication standards such as, for example, WiFi, 4G, or a combination thereof. In one example, the communication component 516 can receive a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In one example, the communication component 516 can further include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared data association (IrDA) technology, Ultra-Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0091] In one example, the device 500 can be implemented by one or more of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic elements for performing the above methods.

[0092] The non-transitory computer-readable storage medium can be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a hybrid drive or a solid state hybrid drive (SSHD), a read-only memory (ROM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, etc.

[0093] FIG. 6 is a flowchart showing an exemplary process of video encoding according to some implementations of the present disclosure.

[0094] In step 602, the processor 520 determines whether one or more reference picture lists are signaled in the PH associated with the picture, and whether one or more reference picture lists indicate that one or more slices associated with the picture are bi-predicted.

[0095] In step 604, in response to determining that one or more reference picture lists are signaled in the PH and one or more reference picture lists indicate that one or more slices are not bi-predicted, the processor 620 adds one or more constraints to one or more syntax elements of the PH.

[0096] In some examples, one or more constraints include skipping the parsing of one or more syntax elements.

[0097] In some examples, one or more syntax elements include one or more flags applicable to one or more slices.

[0098] The processor 520 may further use an enabling flag such as the above-mentioned mvd_l1_zero_flag to identify whether the corresponding motion vector difference (MVD) encoding syntax structure is not parsed, and whether two variables are set to zero for one or more slices associated with the PH, where the two variables respectively identify the difference between the list vector component and the prediction corresponding to the list vector component.

[0099] In some examples, mvd_l1_zero_flag being equal to 1 indicates that the mvd_coding(x0,y0,1) syntax structure is not parsed, and for compIdx = 0 or 1, and cpIdx = 0, 1, or 2, MvdL1[x0][y0][compIdx] and MvdCpL1[x0][y0][cpIdx][compIdx] are set equal to 0. Further, mvd_l1_zero_flag being equal to 0 indicates that the mvd_coding(x0,y0,1) syntax structure is parsed. The mvd_coding(x0,y0,1) syntax structure is the corresponding MVD coding syntax structure. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.

[0100] Further, the variable MvdLX[x0][y0][compIdx] where X is 0 or 1 identifies the difference between the list X vector component used and its prediction. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. compIdx = 0 is assigned to the horizontal motion vector component difference, and compIdx = 1 is assigned to the vertical motion vector component.

[0101] Further, the variable MvdCpLX[x0][y0][cpIdx][compIdx] where X is 0 or 1 identifies the difference between the list X vector component used and its prediction. The array indices x0, y0 identify the position (x0,y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. The array index cpIdx identifies the index of the control point. compIdx = 0 is assigned to the horizontal motion vector component difference, and compIdx = 1 is assigned to the vertical motion vector component.

[0102] In response to determining that the activation flag is equal to 0, the processor 520 may further constrain one or more syntax elements such that the MVD coding syntax structure is parsed for one or more slices.

[0103] In response to determining that the activation flag is equal to 1, the processor 520 may further determine to skip parsing the MVD syntax structure when decoding one or more slices.

[0104] The processor 520 further uses an inactivation flag such as the above-mentioned ph_disable_bdof_flag to identify whether bi-directional optical flow (BDOF) - based inter - bi - prediction for one or more slices associated with PH is inactivated. In response to determining that the inactivation flag is equal to 0, the processor 520 may constrain one or more syntax elements such that BDOF - based inter - bi - prediction is activated when decoding one or more slices. In response to determining that the inactivation flag is equal to 1, the processor 520 may inactivate BDOF - based inter - bi - prediction when decoding one or more slices.

[0105] The processor 520 further uses an inactivation flag such as the above-mentioned ph_disable_dmvr_flag to identify whether decoder motion vector refinement (DMVR) - based inter - bi - prediction for one or more slices associated with PH is inactivated. In response to determining that the inactivation flag is equal to 0, the processor 520 may constrain one or more syntax elements such that DMVR - based inter - bi - prediction is activated when decoding one or more slices. In response to determining that the inactivation flag is equal to 1, the processor 520 may inactivate DMVR - based inter - bi - prediction when decoding one or more slices.

[0106] FIG. 7 is a flowchart illustrating an exemplary process of video encoding according to some implementations of the present disclosure.

[0107] At step 702, the processor 520 uses an enable flag to determine whether one or more temporal motion vector predictors are used for inter prediction of one or more slices associated with the PH of the picture.

[0108] At step 704, the processor 520 constrains the value of the enable flag according to a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0109] In response to determining that there is no reference picture in the DPB that has the same spatial resolution and the same offset as the picture, the processor 520 may set the enable flag to 0. Further, the offset may be applied to the size of the picture for scaling ratio calculation.

[0110] In response to determining that there is no common reference picture for one or more slices, the processor 520 may set the enable flag to 0.

[0111] In response to determining that there is no reference picture in the DPB for which the reference picture scaling flag is equal to 0, the processor 520 may set the enable flag to 0.

[0112] The processor 520 may derive a reference picture scaling flag based on a plurality of offsets applied to the size of the picture for scaling ratio calculation.

[0113] In some examples, an apparatus for video encoding is provided. The apparatus includes one or more processors 520 and a memory 504 configured to store instructions executable by the one or more processors. The processor is configured to execute the method shown in FIG. 6 when the instructions are executed.

[0114] In some examples, an apparatus for video encoding is provided. The apparatus includes one or more processors 520 and a memory 504 configured to store instructions executable by the one or more processors. The processors are configured to execute the method shown in FIG. 7 when the instructions are executed.

[0115] In some other examples, a non-transitory computer-readable storage medium 504 storing instructions is provided. When the instructions are executed by one or more processors 520, the instructions cause the processors to execute the method shown in FIG. 6.

[0116] In some other examples, a non-transitory computer-readable storage medium 504 storing instructions is provided. When the instructions are executed by one or more processors 520, the instructions cause the processors to execute the method shown in FIG. 7.

[0117] The description of the present disclosure is presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings presented in the foregoing description and related drawings.

[0118] Examples have been chosen and described in order to explain the principles of the present disclosure and to enable others skilled in the art to best utilize the present disclosure in various implementations with various modifications made to suit the particular applications contemplated and the underlying principles. Therefore, it is to be understood that the scope of the present disclosure should not be limited to the specific examples disclosed, but that modifications and other implementations are intended to be included within the scope of the present disclosure.

Claims

1. A method for video encoding, comprising: dividing a plurality of pictures into a plurality of encoding units; signaling a first syntax element for determining whether information of one or more reference picture lists exists in a picture header (PH) associated with a picture among the plurality of pictures; when the first syntax element has a first value, signaling the information of the one or more reference picture lists in the PH; wherein the information of the one or more reference picture lists is signaled in the PH, and when one or more slices associated with the picture are determined not to be bi-predicted from the information of the one or more reference picture lists, one or more second syntax elements in the PH are not parsed, and the one or more second syntax elements include one or more flags applicable to the one or more slices; The method, wherein the one or more flags include a first flag for specifying whether a motion vector difference (MVD) encoding syntax structure is to be parsed.

2. Setting the value of the first flag to specify whether the MVD encoding syntax structure is not to be parsed and whether two variables are set to zero for one or more slices associated with the PH, wherein the two variables each further include specifying a difference between a list vector component and a prediction corresponding to the list vector component; setting the value of the first flag to be equal to 0 to specify that the MVD encoding syntax structure is to be parsed for the one or more slices, or setting the value of the first flag to be equal to 1 to specify that the MVD encoding syntax structure is not to be parsed for the one or more slices, according to the method of Claim 1.

3. further comprising setting the value of a second flag of the one or more second syntax elements to specify whether bi-directional optical flow (BDOF) inter-prediction-based inter bi-prediction is disabled for one or more slices associated with the PH. To specify that the BDOF inter prediction-based inter dual prediction is enabled for the one or more slices, the value of the second flag of the one or more second syntax elements is set equal to 0, or, To specify that the BDOF inter prediction-based inter dual prediction is disabled for the one or more slices, the value of the second flag of the one or more second syntax elements is set equal to 1, the method according to claim 1.

4. Further including setting the value of a third flag of one or more second syntax elements to identify whether decoder motion vector refinement (DMVR)-based inter dual prediction is disabled for one or more slices associated with the PH, To specify that the DMVR-based inter dual prediction is enabled for the one or more slices, the value of the third flag of the one or more second syntax elements is set equal to 0, or, To specify that the DMVR-based inter dual prediction is disabled for the one or more slices, the value of the third flag of the one or more second syntax elements is set equal to 1, the method according to claim 1.

5. The information of the one or more reference picture lists includes a variable identifying the number of entries in the one or more reference picture lists, the method according to claim 1.

6. An apparatus for video encoding, One or more processors; A memory configured to store instructions executable by the one or more processors, comprising: The one or more processors are configured to execute the method according to any one of claims 1 to 5 when executing the instructions, an apparatus.

7. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to execute the method according to any one of claims 1 to 5.

8. A computer program including a plurality of instructions that, when executed by one or more processors, cause the one or more processors to execute the method according to any one of claims 1 to 5.

9. A method of storing a bitstream generated by the encoding method according to any one of claims 1 to 5.

10. A method of transmitting a bitstream generated by the encoding method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and device for signaling of syntax element in video coding

    JP2024099843A

  • Decoding based on BI-directional picture condition

    WO2021201759A1