Method and apparatus for high level syntax in video coding
By rearranging and grouping the syntax elements in the video codec standard, ensuring that intra-frame prediction-related syntax elements are defined before inter-frame prediction-related syntax elements, the problems of standardization and efficiency in the prior art are solved, and more efficient codec performance is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
- Filing Date
- 2021-03-23
- Publication Date
- 2026-05-22
Smart Images

Figure CN118646871B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application filed on March 23, 2021, with application number 202180024900.8 and entitled "Method and apparatus for advanced syntax in video encoding and decoding".
[0002] Cross-reference to related applications
[0003] This application is based on and claims priority to Provisional Application No. 63 / 003,229, filed on March 31, 2020, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0004] This disclosure relates to video encoding / decoding and compression. More specifically, this application relates to methods and apparatus for using high-level syntax in video bitstreams applicable to one or more video encoding / decoding standards. Background Technology
[0005] Various video codec technologies can be used to compress video data. Video codecs are performed according to one or more video codec standards. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (H.265 / HEVC), High-Level Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) codecs. Video codecs typically use prediction methods that utilize redundancy present in video images or sequences (e.g., inter-frame prediction, intra-frame prediction, etc.). A key goal of video codec technologies is to compress video data to a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0006] Examples of this disclosure provide methods and apparatus for high-level syntax encoding and decoding in video encoding and decoding.
[0007] According to a first aspect of this disclosure, a method for decoding a video signal is provided. The method may include: receiving an arranged partition constraint syntax element at a Sequence Parameter Set (SPS) level. The arranged partition constraint syntax elements may be arranged such that syntax elements associated with intra-frame prediction may be defined before syntax elements associated with inter-frame prediction. The decoder may also obtain a first reference picture I associated with a video block in the bitstream. (0) Second reference image I (1) According to the display order, the first reference image I (0) It can be before the current image, and the second reference image I (1)It can be after the current image. The decoder can also be based on the first reference image I. (0) The reference block in the video block is used to obtain the first prediction sample I of the video block. (0) (i,j). i and j can represent the coordinates of a sample with the current image. The decoder can also be based on the second reference image I. (1) The reference block in the video block is used to obtain the second prediction sample I of the video block. (1) (i,j). The decoder can also be based on the arranged partition constraint syntax elements and the first predicted sample I. (0) (i,j) and the second predicted sample I (1) (i,j) is used to obtain bidirectional prediction samples.
[0008] According to a second aspect of this disclosure, a computing device is provided. The computing device may include: one or more processors; and a non-transitory computer-readable storage medium storing instructions executable by the one or more processors. The one or more processors may be configured to receive an ordered partition constraint syntax element at an SPS level. The ordered partition constraint syntax elements may be arranged such that syntax elements associated with intra-frame prediction may be defined before syntax elements associated with inter-frame prediction. The one or more processors may also be configured to obtain a first reference picture I associated with a video block in a bitstream. (0) Second reference image I (1) According to the display order, the first reference image I (0) It can be before the current image, and the second reference image I (1) This can occur after the current image. The one or more processors can also be configured to, based on the first reference image I... (0) The reference block in the video block is used to obtain the first prediction sample I of the video block. (0) (i,j). i and j can represent the coordinates of a sample with the current image. The one or more processors can also be configured to, based on the second reference image I (1) The reference block in the video block is used to obtain the second prediction sample I of the video block. (1) (i,j). The one or more processors may also be configured to, based on the arranged partition constraint syntax elements, the first prediction sample I (0) (i,j) and the second predicted sample I (1) (i,j) is used to obtain bidirectional prediction samples.
[0009] According to a third aspect of this disclosure, a non-transitory computer-readable storage medium having instructions stored therein is provided. When the instructions are executed by one or more processors of the device, the instructions can cause the device to receive an arranged partition constraint syntax elements at the SPS level. The arranged partition constraint syntax elements can be arranged such that syntax elements associated with intra-frame prediction are defined before syntax elements associated with inter-frame prediction. The instructions can also cause the device to obtain a first reference picture I associated with a video block in the bitstream. (0) Second reference image I (1) According to the display order, the first reference image I (0) It can be before the current image, and the second reference image I (1) It can be after the current image. The instruction can also cause the device to adjust according to the first reference image I. (0) The reference block in the video block is used to obtain the first prediction sample I of the video block. (0) (i,j). i and j can represent the coordinates of a sample with the current image. The instruction can also cause the device to adjust the coordinates according to the second reference image I. (1) The reference block in the video block is used to obtain the second prediction sample I of the video block. (1) (i,j). The instructions can also cause the device to base its operation on the arranged partition constraint syntax elements and the first prediction sample I. (0) (i,j) and the second predicted sample I (1) (i,j) is used to obtain bidirectional prediction samples.
[0010] It should be understood that the general description above and the detailed description below are exemplary and explanatory only, and are not intended to limit the scope of this disclosure. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with this disclosure and, together with the specification, serve to explain the principles of this disclosure.
[0012] Figure 1 This is an example encoder block diagram based on the content of this disclosure.
[0013] Figure 2 This is a block diagram of a decoder based on an example of this disclosure.
[0014] Figure 3A This is a diagram illustrating block partitioning in a multi-type tree structure, as an example of this disclosure.
[0015] Figure 3B This is a diagram illustrating block partitioning in a multi-type tree structure, as an example of this disclosure.
[0016] Figure 3C This is a diagram illustrating block partitioning in a multi-type tree structure, as an example of this disclosure.
[0017] Figure 3D This is a diagram illustrating block partitioning in a multi-type tree structure, as an example of this disclosure.
[0018] Figure 3E This is a diagram illustrating block partitioning in a multi-type tree structure, as an example of this disclosure.
[0019] Figure 4 This is an example of a method for decoding video signals based on the present disclosure.
[0020] Figure 5 This is an example of a method for decoding video signals based on the present disclosure.
[0021] Figure 6 This is a diagram illustrating a computing environment coupled with a user interface, as an example of this disclosure. Detailed Implementation
[0022] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the drawings, wherein, unless otherwise stated, the same numerals in different drawings denote the same or similar elements. The implementations set forth in the following description of exemplary embodiments do not represent all implementations consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with aspects related to this disclosure, as recited in the appended claims.
[0023] The terminology used in this disclosure is for the purpose of describing the objectives of particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein is intended to express and include any and all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various types of information, the information should not be limited by these terms. These terms are used only to distinguish one type of information from another. For example, without departing from the scope of this disclosure, first information may be referred to as second information; and similarly, second information may be referred to as first information. As used herein, the term “if” may be understood to mean “when,” “in,” or “in response to a judgment,” depending on the specific context.
[0025] The first version of the HEVC standard was finalized in October 2013, offering approximately 50% bitrate savings or equivalent perceived quality compared to its predecessor, H.264 / MPEG AVC. Despite the significant codec improvements offered by HEVC, evidence suggested that even greater codec efficiency could be achieved using additional codec tools. Based on this, VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-T VECG and ISO / IEC MPEG established a Joint Video Exploration Team (JVET) to initiate a major study of advanced technologies that would significantly enhance codec efficiency. JVET maintains a reference software, known as the Joint Exploration Model (JEM), by integrating several additional codec tools on top of the HEVC Test Model (HM).
[0026] In October 2017, the ITU-T and ISO / IEC issued a joint call for proposals (CfP) on video compression with capabilities exceeding HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, confirming a 40% compression efficiency gain compared to HEVC. Based on these evaluations, JVET launched a new project to develop a next-generation video codec standard called Universal Video Codec (VVC). In the same month, a reference software codebase (called the VVC Test Model (VTM)) was established to demonstrate a reference implementation of the VVC standard.
[0027] Like HEVC, VVC is built on top of a block-based hybrid video codec framework.
[0028] Figure 1 An overall diagram of a block-based video encoder for VVC is shown. Specifically, Figure 1A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-frame mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction related information 142, intra-frame prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy encoding / decoding 138, and bitstream 144.
[0029] In encoder 100, video frames are divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method.
[0030] The prediction residual, representing the difference between the current video block (a portion of video input 110) and its predicted value (a portion of block prediction value 140), is sent from adder 128 to transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantization coefficients are then fed to entropy coding 138 to generate a compressed video bitstream. Figure 1 As shown, prediction-related information 142 (such as video block segmentation information, motion vectors (MV), reference picture indexes, and intra-prediction modes) from intra-frame / inter-frame mode decision 116 is also fed through entropy coding 138 and saved to a compressed bitstream 144. The compressed bitstream 144 includes the video bitstream.
[0031] In encoder 100, circuitry associated with the decoder is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed via inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with block prediction values 140 to generate unfiltered reconstructed pixels for the current video block.
[0032] Spatial prediction (or "intra-frame prediction") uses pixels from samples of encoded neighboring blocks (called reference samples) in the same video frame as the current video block to predict the current video block.
[0033] Timing prediction (also known as "inter-frame prediction") uses reconstructed pixels from encoded video frames to predict the current video block. Timing prediction reduces the temporal redundancy inherent in the video signal. The timing prediction signal for a given codec unit (CU) or codec block is typically communicated via one or more MV signals, indicating the amount and direction of motion between the current CU and its timing reference. Additionally, if multiple reference frames are supported, a reference frame index is sent to identify which reference frame in the reference frame storage unit the timing prediction signal originates from.
[0034] Motion estimation 114 acquires the video input 110 and the signal from the image buffer 120, and outputs the motion estimation signal to motion compensation 112. Motion compensation 112 acquires the video input 110, the signal from the image buffer 120, and the motion estimation signal from motion estimation 114, and outputs the motion compensation signal to intra / inter-frame mode decision 116.
[0035] After performing spatial and / or temporal predictions, the intra / inter-frame mode decision 116 in encoder 100 selects the optimal prediction mode, for example, based on a rate-distortion optimization method. Then, the block prediction value 140 is subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantization residual coefficients are dequantized by inverse quantization 134 and inverse transformed by inverse transform 136 to form the reconstruction residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Furthermore, in-loop filtering 122 (e.g., deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF)) can be applied to the reconstructed CU before it is placed in the reference picture storage unit of picture buffer 120 and used for encoding and decoding future video blocks. To form the output video bitstream 144, the encoding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are all sent to entropy coding unit 138 for further compression and packing to form the bitstream.
[0036] Figure 1 A block diagram of a general block-based hybrid video coding system is presented. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128x128 pixels. However, unlike HEVC, which is based solely on quadtree-based block partitioning, in VVC, a coding tree unit (CTU) is split into CUs to accommodate different local characteristics based on quadtree / binary / tritree structures. Furthermore, the concept of multiple partitioned unit types in HEVC is removed; that is, in VVC, there is no longer a distinction between CUs, prediction units (PUs), and transform units (TUs); instead, each CU is always used as the basic unit for prediction and transform without further partitioning. In the multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary and ternary tree structures.
[0037] like Figure 3A , 3B As shown in 3C, 3D, and 3E, there are five splitting types: quadruple splitting, horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.
[0038] Figure 3AA diagram showing block quadrilateral partitioning in a multi-type tree structure according to this disclosure is provided.
[0039] Figure 3B A diagram showing block vertical binary partitioning in a multi-type tree structure according to this disclosure is provided.
[0040] Figure 3C A diagram showing block-level binary partitioning in a multi-type tree structure according to this disclosure is provided.
[0041] Figure 3D A diagram of block vertical ternary partitioning in a multi-type tree structure according to this disclosure is shown.
[0042] Figure 3E A diagram showing block-level ternary partitioning in a multi-type tree structure according to this disclosure is illustrated.
[0043] exist Figure 1 In the encoder, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples of encoded neighboring blocks (called reference samples) in the same video picture / strip to predict the current video block. Spatial prediction reduces the inherent spatial redundancy in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from encoded video pictures to predict the current video block. Temporal prediction reduces the inherent temporal redundancy in the video signal. The temporal prediction signal for a given CU is typically signaled via one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference pictures are supported, a reference picture index is sent to identify which reference picture in the reference picture storage unit the temporal prediction signal originates from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on rate-distortion optimization methods. The predicted block is then subtracted from the current video block; and the prediction residual is decorrelated using transform and quantization. The quantization residual coefficients are dequantized and inversely transformed to form the reconstruction residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Furthermore, in-loop filtering (such as deblocking filters, Sample Adaptive Offset (SAO), and Adaptive In-Loop Filter (ALF)) can be applied to the reconstructed CU before it is placed into the reference image storage unit and used to encode future video blocks. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantization residual coefficients are all sent to the entropy coding unit for further compression and packing to form the bitstream.
[0044] Figure 2 A general block diagram of a video decoder for VVC is shown. Specifically, Figure 2A typical block diagram of decoder 200 is shown. Decoder 200 has a bitstream 210, entropy decoding 212, dequantization 214, inverse transform 216, adder 218, intra / inter-frame mode selection 220, intra-frame prediction 222, memory 230, in-loop filter 228, motion compensation 224, image buffer 226, prediction-related information 234, and video output 232.
[0045] Decoder 200 is similar to residing in Figure 1 The reconstruction-related part is located in the encoder 100. In the decoder 200, the incoming video bitstream 210 is first decoded by entropy decoding 212 to derive the quantization coefficient level and prediction-related information. Then, the quantization coefficient level is processed by inverse quantization 214 and inverse transform 216 to obtain the reconstruction prediction residual. The block prediction mechanism implemented in the intra / inter-frame mode selector 220 is configured to perform intra-frame prediction 222 or motion compensation 224 based on the decoded prediction information. The unfiltered reconstructed pixel set is obtained by adding the reconstruction prediction residual from inverse transform 216 and the prediction output generated by the block prediction mechanism using adder 218.
[0046] The reconstructed blocks can be further passed through the in-loop filter 228 before being stored in the image buffer 226, which acts as a reference image storage unit. The reconstructed video in the image buffer 226 can be sent to drive a display device, as well as to predict future video blocks. With the in-loop filter 228 open, filtering operations are performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0047] Figure 2 A general block diagram of a block-based video decoder is presented. First, the video bitstream is entropy-decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (if intra-frame coded) or the temporal prediction unit (if inter-frame coded) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. The prediction blocks and residual blocks are then combined. Before storing the reconstructed blocks in the reference image storage unit, the reconstructed blocks can be further filtered through an in-loop process. Finally, the reconstructed video from the reference image storage unit is emitted to drive the display device, along with the video blocks used to predict future blocks.
[0048] Typically, the basic intra prediction scheme used in VVC is the same as that in HEVC, except for further extensions and / or improvements to several modules, such as Matrix Weighted Intra Prediction (MIP) coding mode, Intra-Segmentation Sub-Partition (ISP) coding mode, Extended Intra Prediction with Wide-Angle Intra-Direction, Position-Related Intra Prediction Combination (PDPC), and 4-Tap Intra-Interpolation. The main focus of this disclosure is to improve the existing high-level syntax design in the VVC standard. The relevant background information is set forth in the following sections.
[0049] Like HEVC, VVC uses a bitstream structure based on Network Abstraction Layer (NAL) units. The encoded bitstream is divided into NAL units, which should be smaller than the Maximum Transmission Unit (MTB) size when transmitted over a lossy packet network. Each NAL unit consists of a NAL unit header and a subsequent NAL unit payload. There are two conceptual categories of NAL units: Video Coding Layer (VCL) NAL units containing encoded sample data (e.g., encoded stripe NAL units), and non-VCL NAL units containing metadata, which typically belong to more than one encoded picture, or where association with a single encoded picture would be meaningless (e.g., parameter set NAL units), or where the decoding process does not require the information (e.g., SEI NAL units).
[0050] In VVC, a two-byte NAL unit header is introduced, with the design intended to support future expansions. The syntax and associated semantics of the NAL unit header in the current VVC draft specification are shown in Tables 1 and 2, respectively. How to read Table 1 is explained in the appendix of this invention, which can also be found in the VVC specification.
[0051] Table 1. NAL Unit Header Syntax
[0052]
[0053] Table 2. Semantics of NAL Unit Headers
[0054]
[0055]
[0056] Table 3. NAL Unit Type Codes and NAL Unit Type Categories
[0057]
[0058]
[0059] VVC inherits the parameter set concept from HEVC, with some modifications and additions. A parameter set can be part of the video bitstream or received by the decoder through other means (including out-of-band transmission using reliable channels, hard encoding / decoding in the encoder and decoder, etc.). The parameter set contains identifiers referenced directly or indirectly from the stripe header, as will be discussed in detail later. This referencing process is called "activation." Depending on the parameter set type, activation is performed by picture or by sequence. The concept of activation by reference was introduced because, in the case of out-of-band transmission, implicit activation by means of the information's position in the bitstream (common for other syntax elements of the video codec) is unavailable, and for other reasons.
[0060] Video Parameter Sets (VPSs) are introduced to transmit information applicable to multiple layers and sublayers. VPSs were introduced to address these shortcomings and to enable a clean and scalable high-level design for multi-layer codecs. Each layer of a given video sequence (regardless of whether they have the same or different Sequence Parameter Sets (SPs)) references the same VPS. The syntax and associated semantics of video parameter sets in the current VVC draft specification are illustrated in Tables 4 and 5, respectively. How to read Table 4 is explained in the appendix of this invention, which can also be found in the VVC specification.
[0061] Table 4. Video Parameter Set RBSP Syntax
[0062]
[0063]
[0064]
[0065] Table 5. Semantics of Video Parameter Set RBSP
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073] In VVC, the SPS contains information applicable to all stripes of the encoded video sequence. The encoded video sequence begins with the Instant Decode Refresh (IDR) picture, or BLA picture, or CRA picture, which is the first picture in the bitstream, and includes all subsequent pictures that are not IDR or BLA pictures. The bitstream consists of one or more encoded video sequences. The content of the SPS can be broadly categorized into six types: 1) self-references (its own ID); 2) decoder operation point information (profile, level, picture size, number of sublayers, etc.); 3) flags in the enable profile for certain tools, and, if tools are enabled, the associated codec tool parameters.
[0074] 4) Information on limiting structural flexibility and transform coefficient encoding; 5) Temporal scalability control; and 6) Visual usability information (VUI), which includes HRD information. Tables 6 and 7 illustrate the syntax and associated semantics of the sequence parameter set in the current VVC draft specification. How to read Table 6 is explained in the appendix of this invention, which can also be found in the VVC specification.
[0075] Table 6. Sequence Parameter Set (RBSP) Syntax
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082] Table 7. Sequence Parameter Set (RBSP) Semantics
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098] The Picture Parameter Set (PPS) of VVC contains information that may vary from picture to picture. The PPS includes information roughly equivalent to a portion of the PPS in HEVC, including: 1) self-reference; 2) initial picture control information, such as initial quantization parameters (QP), multiple flags indicating the use or presence of certain tools or control information in the strip header; and 3) tiling information. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are shown in Tables 8 and 9, respectively. How to read Table 8 is explained in the appendix of this invention, which can also be found in the VVC specification.
[0099] Table 8. Image Parameter Set RBSP Syntax
[0100]
[0101]
[0102]
[0103]
[0104] Table 9. Image Parameter Set RBSP Semantics
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114] The strip header contains information that can be changed strip-by-strip, as well as relatively small or only image-related information specific to a particular strip or image type. The size of the strip header can be significantly larger than the PPS, especially when there are tile or wavefront entry point offsets in the strip header and modifications to the RPS, prediction weights, or reference image list are explicitly signaled. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are shown in Tables 10 and 11, respectively. How to read Table 10 is explained in the appendix of this invention and can also be found in the VVC specification.
[0115] Table 10. Syntax of Image Header Structure
[0116]
[0117]
[0118]
[0119]
[0120]
[0121] Table 11. Semantic Structure of Image Headers
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131] Improvements to syntactic elements
[0132] In the current VVC, when similar syntax elements exist for intra-frame and inter-frame prediction in different places, the syntax elements related to inter-frame prediction are defined before those related to intra-frame prediction. Given that intra-frame prediction is allowed in all picture / strip types, but not inter-frame prediction, this order may not be preferred. From a standardization perspective, it would be beneficial to always define the syntax related to intra-frame prediction before defining the syntax for inter-frame prediction.
[0133] It was also observed that in the current VVC, some highly related syntactic elements are defined in different places in a diffuse manner. From a standardization perspective, grouping some syntactic elements together would also be beneficial.
[0134] The proposed method
[0135] In this disclosure, methods for simplifying and / or further improving existing designs of high-level grammars are provided to address the problems identified in the "Problem Statement" section. It should be noted that the innovative methods can be applied independently or in combination.
[0136] Grouping partition constraint syntax elements by prediction type
[0137] This disclosure proposes a rearrangement of syntax elements such that syntax elements related to intra-frame prediction are defined before those related to inter-frame prediction. According to this disclosure, partition constraint syntax elements are grouped by prediction type, with those related to intra-frame prediction preceding those related to inter-frame prediction. In one embodiment, the order of partition constraint syntax elements in the SPS is consistent with the order of partition constraint syntax elements in the picture header. An example of the decoding process for the VVC draft is shown in Table 12 below. Changes to the VVC draft are indicated using bold and italics. Strikethrough text has been removed from this process.
[0138] Table 12. Proposed Sequence Parameter Set (RBSP) Syntax
[0139]
[0140]
[0141] Figure 4 A method for decoding a video signal according to this disclosure is shown. For example, this method can be applied to a decoder.
[0142] In step 410, the decoder may receive the arranged partition constraint syntax elements at the SPS level. The arranged partition constraint syntax elements are arranged such that syntax elements related to intra-frame prediction are defined before syntax elements related to inter-frame prediction.
[0143] In step 412, the decoder can obtain the first reference image I associated with the video block in the bitstream. (0) Second reference image I (1) In the order of display, the first reference image I (0) Before the current image, and in the second reference image I (1) Following the current image.
[0144] In step 414, the decoder can determine the first reference image I. (0) The reference block in the video block is used to obtain the first predicted sample I of the video block. (0) (i,j). i and j represent the coordinates of a sample with the current image.
[0145] In step 416, the decoder can determine the second reference image I. (1) The reference block in the video block is used to obtain the second predicted sample I of the video block. (1) (i,j).
[0146] In step 418, the decoder can be based on the permuted partition constraint syntax elements and the first predicted sample I. (0) (i,j) and the second prediction sample I (1) (i,j) is used to obtain bidirectional prediction samples.
[0147] Figure 5 A method for decoding a video signal according to this disclosure is illustrated. For example, this method can be applied to a decoder. In step 510, the decoder may receive a bitstream including VPS, SPS, PPS, picture header, and stripe header for encoded video data. In step 512, the decoder may decode the VPS. In step 514, the decoder may decode the SPS and obtain the arranged partition constraint syntax elements at the SPS level. In step 516, the decoder may decode the PPS. In step 518, the decoder may decode the picture header. In step 520, the decoder may decode the stripe header. In step 522, the decoder may decode the video data based on the VPS, SPS, PPS, picture header, and stripe header.
[0148] Grouping of two-tree chroma syntax elements
[0149] In this disclosure, it is proposed to group syntax elements related to dual-tree chroma types. In one embodiment, partition constraint syntax elements in the SPS for dual-tree chroma should be signaled together in the dual-tree chroma case. An example of the decoding process for the VVC draft is shown in Table 13 below. Changes to the VVC draft are shown in bold and italics. Strikethrough text has been removed from this process.
[0150] Table 13. Proposed Sequence Parameter Set (RBSP) Syntax
[0151]
[0152]
[0153] If we also consider defining the syntax related to intra-frame prediction before the syntax related to inter-frame prediction, then another example of the decoding process on the VVC draft is shown in Table 14 below, according to the method of this disclosure. Changes to the VVC draft are shown in bold and italics. Strikethrough text has been removed from this process.
[0154] Table 14. Proposed Sequence Parameter Set (RBSP) Syntax
[0155]
[0156]
[0157] Conditionally use signals to notify syntax elements related to inter-frame prediction
[0158] As mentioned in the preceding description, current VVC allows intra-frame prediction in all picture / strip types, but not inter-frame prediction. According to this disclosure, it is proposed to add a flag to the VVC syntax at a specific coding level to indicate whether inter-frame prediction is used in sequences, pictures, and / or stripes. When inter-frame prediction is not used, the syntax related to inter-frame prediction is not signaled at the corresponding coding level (e.g., sequence, picture, and / or stripe level).
[0159] In one example, according to the method of this disclosure, a flag is added to the SPS to indicate whether inter-frame prediction is used when encoding the current video sequence. If it is not used, the syntax elements related to inter-frame prediction are not signaled in the SPS. The decoding process on the VVC draft is shown in Table 15 below. Changes to the VVC draft are indicated using bold and italics.
[0160] Table 15. Proposed Sequence Parameter Set (RBSP) Syntax
[0161]
[0162]
[0163] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics
[0164] A value of 0 for sps_inter_slice_used_flag indicates that all encoded slices in the video sequence have a slice_type of 2. A value of 1 for sps_inter_slice_used_flag indicates that one or more encoded slices with a slice_type of 0 or 1 may or may not exist in the video sequence.
[0165] The methods described above can be implemented using an apparatus comprising one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. This apparatus can combine these circuits with other hardware or software components to perform the methods described above. Each module, submodule, unit, or subunit disclosed above can be implemented at least partially using one or more circuits.
[0166] Figure 6 A computing environment 610 coupled to a user interface 660 is shown. The computing environment 610 may be part of a data processing server. The computing environment 610 includes a processor 620, memory 640, and I / O interface 650.
[0167] Processor 620 typically controls the overall operation of computing environment 610, such as operations associated with display, data acquisition, data communication, and image processing. Processor 620 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Furthermore, processor 620 may include one or more modules that facilitate interaction between processor 620 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, GPU, etc.
[0168] Memory 640 is configured to store various types of data to support the operation of computing environment 610. Memory 640 may include predefined software 642. Examples of such data include instructions for any application or method operating on computing environment 610, video datasets, image data, etc. Memory 640 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0169] I / O interface 650 provides an interface between processor 620 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). These buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 650 can be coupled to encoders and decoders.
[0170] In some embodiments, a non-transitory computer-readable storage medium is also provided, which includes a plurality of programs (such as those included in memory 640) that can be executed by processor 620 in computing environment 610 to perform the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0171] A non-transitory computer-readable storage medium has multiple programs stored therein for execution by a computing device having one or more processors, wherein the multiple programs, when executed by the one or more processors, cause the computing device to perform the aforementioned method for motion prediction.
[0172] In some embodiments, the computing environment 610 may be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above methods.
[0173] Other examples of this disclosure will be apparent to those skilled in the art based on consideration of the description and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include deviations from the scope of known practice or custom in the art. This specification and examples are intended to be illustrative only.
[0174] It will be understood that this disclosure is not limited to the exact examples shown above and in the accompanying drawings, and that various modifications and changes may be made without departing from its scope.
Claims
1. A method for encoding video signals, comprising: Obtain the partition constraint syntax elements at the Sequence Parameter Set (SPS) level, wherein the partition constraint syntax elements include syntax elements related to intra-frame prediction and syntax elements related to inter-frame prediction; The partition constraint syntax elements are arranged such that the syntax elements related to intra-frame prediction are defined before the syntax elements related to inter-frame prediction; and A bitstream is formed based on the arranged partition constraint syntax elements and then sent. The arranged partition constraint syntax elements include multiple partition constraint syntax elements for dual-tree chroma in the SPS level, and the multiple partition constraint syntax elements for dual-tree chroma are defined together in the dual-tree case, and the multiple partition constraint syntax elements for dual-tree chroma are defined before the syntax elements related to inter-frame prediction. The arranged partition constraint syntax elements are arranged in the following manner: Determine the qtbtt_dual_tree_intra_flag flag; When the qtbtt_dual_tree_intra_flag flag is true, the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma is set; and Set the value of sps_log2_diff_min_qt_min_cb_inter_slice.
2. The method according to claim 1, wherein, At least a portion of the arranged partition constraint syntax elements are arranged in an order consistent with the order of the partition constraint syntax elements in the image header.
3. The method according to claim 1, wherein, The arranged partition constraint syntax elements are arranged in the following way: Set the value of log2_min_luma_coding_block_size_minus2; Determine the value of sps_max_mtt_hierarchy_depth_intra_slice_luma; When the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0, set the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma. When the ChromaArrayType value is not equal to 0, the qtbtt_dual_tree_intra_flag flag is notified by a signal; and When the qtbtt_dual_tree_intra_flag flag is true, set the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma.
4. The method according to claim 1, wherein, The arranged partition constraint syntax elements are arranged in the following way: Set the value of log2_min_luma_coding_block_size_minus2; Determine the value of sps_max_mtt_hierarchy_depth_intra_slice_luma; When the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0, set the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma. When the ChromaArrayType value is not equal to 0, the qtbtt_dual_tree_intra_flag flag is notified by a signal; When the qtbtt_dual_tree_intra_flag flag is true, the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma is set; and Set the value of sps_log2_diff_min_qt_min_cb_inter_slice.
5. The method according to claim 1, wherein, The arranged partition constraint syntax elements include a flag indicating the use of the Universal Video Coding (VVC) syntax for inter-frame prediction.
6. The method according to claim 5, wherein, The VVC syntax flag is signaled at the SPS level and indicates the use of inter-frame prediction when encoding the current video sequence.
7. The method according to claim 5, wherein, The arranged partition constraint syntax elements are arranged in the following way: The `sps_inter_slice_used_flag` flag is signaled. Ensure that the ChromaArrayType value is not equal to 0; When the ChromaArrayType value is not equal to 0, the qtbtt_dual_tree_intra_flag flag is notified by a signal. Make sure the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0; When the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0, set the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma. Determine that the sps_inter_slice_used_flag flag is not true; and When the sps_inter_slice_used_flag flag is not true, set the sps_log2_diff_min_qt_min_cb_inter_slice value.
8. A method for decoding a video signal, comprising: The received bitstream contains partition constraint syntax elements at the Sequence Parameter Set (SPS) level, wherein the partition constraint syntax elements include syntax elements related to intra-frame prediction and syntax elements related to inter-frame prediction, and wherein the partition constraint syntax elements are arranged such that the syntax elements related to intra-frame prediction are defined before the syntax elements related to inter-frame prediction; and The partition constraint syntax element includes multiple partition constraint syntax elements for dual-tree chroma in the SPS level, and the multiple partition constraint syntax elements for dual-tree chroma are defined together in the dual-tree case, and the multiple partition constraint syntax elements for dual-tree chroma are defined before the syntax elements related to inter-frame prediction. The partition constraint syntax elements are arranged in the following manner: Determine the qtbtt_dual_tree_intra_flag flag; When the qtbtt_dual_tree_intra_flag flag is true, determine the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma; and Determine the value of sps_log2_diff_min_qt_min_cb_inter_slice.
9. The method according to claim 8, wherein, At least some of the partition constraint syntax elements are arranged in an order consistent with the order of the partition constraint syntax elements in the image header.
10. The method according to claim 8, wherein, The partition constraint syntax elements are arranged in the following manner: Determine the value of log2_min_luma_coding_block_size_minus2; Determine the value of sps_max_mtt_hierarchy_depth_intra_slice_luma; When the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0, determine the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma; When the ChromaArrayType value is not equal to 0, the qtbtt_dual_tree_intra_flag flag is signaled; and When the qtbtt_dual_tree_intra_flag flag is true, determine the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma.
11. The method according to claim 8, wherein, The partition constraint syntax elements are arranged in the following manner: Determine the value of log2_min_luma_coding_block_size_minus2; Determine the value of sps_max_mtt_hierarchy_depth_intra_slice_luma; When the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0, determine the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma; When the ChromaArrayType value is not equal to 0, the qtbtt_dual_tree_intra_flag flag is signaled. When the qtbtt_dual_tree_intra_flag flag is true, determine the value of sps_log2_diff_min_qt_min_cb_intra_slice_chroma; and Determine the value of sps_log2_diff_min_qt_min_cb_inter_slice.
12. The method according to claim 8, wherein, The partition constraint syntax element includes a flag indicating the use of the Universal Video Coding (VVC) syntax for inter-frame prediction.
13. The method according to claim 12, wherein, The VVC syntax flag is signaled at the SPS level and indicates the use of inter-frame prediction when decoding the current video sequence.
14. The method according to claim 12, wherein, The partition constraint syntax elements are arranged in the following manner: The sps_inter_slice_used_flag flag is signaled; Ensure that the ChromaArrayType value is not equal to 0; When the ChromaArrayType value is not equal to 0, the qtbtt_dual_tree_intra_flag flag is signaled. Make sure the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0; When the value of sps_max_mtt_hierarchy_depth_intra_slice_luma is not 0, determine the value of sps_log2_diff_max_bt_min_qt_intra_slice_luma; Determine that the sps_inter_slice_used_flag flag is not true; and When the sps_inter_slice_used_flag flag is not true, determine the value of sps_log2_diff_min_qt_min_cb_inter_slice.
15. A computing device, comprising: One or more processors; as well as A non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 1 to 7.
16. A computing device, comprising: One or more processors; as well as A non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 8 to 14.
17. A non-transitory computer-readable storage medium storing a plurality of programs for execution by a computing device having one or more processors, wherein, When the plurality of programs are executed by the one or more processors, they cause the computing device to perform the method according to any one of claims 1 to 14.
18. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the method according to any one of claims 1 to 14.
19. A computer-readable storage medium having stored thereon instructions and a bit stream, wherein the instructions, when executed by a processor, implement the encoding method according to any one of claims 1 to 7 to generate the bit stream.
20. A method for storing a bit stream, comprising: Perform the encoding method according to any one of claims 1 to 7 to generate a bit stream; as well as Store the bit stream.