Method and device for encoding and decoding video signal, and storage medium

By rearranging and grouping VVC syntax elements, the signal transmission order of intra-frame and inter-frame prediction is optimized, which solves the problem of unreasonable syntax design in VVC and improves encoding and decoding efficiency and video quality.

CN121967719APending Publication Date: 2026-05-01BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
Filing Date
2021-04-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video codec standards such as VVC suffer from unreasonable grouping of syntax elements in their advanced syntax design, leading to low encoding and decoding efficiency. Furthermore, the order in which intra-frame prediction and inter-frame prediction are allowed is not optimized across different picture/strip types.

Method used

The syntax elements are rearranged so that intra-frame prediction-related syntax elements are grouped in the general video codec (VVC) syntax at the coding level, and inter-frame prediction-related syntax elements are signaled according to predefined conditions at a specific coding level. Flags are added to indicate whether inter-frame prediction is allowed, thus optimizing the grouping of syntax elements and the order of signal transmission.

Benefits of technology

It improves the efficiency of video encoding and decoding, simplifies advanced syntax design, optimizes the encoding and decoding process, and enhances encoding efficiency and video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967719A_ABST
    Figure CN121967719A_ABST
Patent Text Reader

Abstract

The invention provides a method, a device and a storage medium for encoding and decoding a video signal. A decoder may receive ranked syntax elements in a sequence parameter set (SPS) level through a bitstream. The ranked syntax elements in the SPS level are ranked such that functions of related syntax elements are grouped in a generic video coding (VVC) syntax of the coding level. The decoder may receive, through the bitstream and in response to the plurality of syntax elements satisfying a predefined condition, a second syntax element immediately after the plurality of syntax elements. The decoder may perform a relevant syntax element function over the bitstream on video data from the bitstream according to the plurality of syntax elements and the second syntax element.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application filed on April 30, 2021, with application number 202180032251.6 and invention title "High-level syntax for video encoding and decoding". Technical Field

[0002] This disclosure relates to video encoding / decoding and compression. More specifically, this application relates to high-level syntax in video bitstreams applicable to one or more video encoding / decoding standards. Background Technology

[0003] Various video codec techniques can be used to compress video data. Video codecs are performed according to one or more video codec standards. Examples of video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (H.265 / HEVC), High-Level Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) codecs. Video codecs typically use prediction methods (e.g., inter-frame prediction, intra-frame prediction), which utilize redundancy present in video images or sequences. A key goal of video codec techniques is to compress video data to a lower bitrate while avoiding or minimizing video quality degradation. Summary of the Invention

[0004] The examples disclosed herein provide methods and apparatuses for high-level syntax in video encoding and decoding.

[0005] According to a first aspect of this disclosure, a method for decoding a video signal is provided. The method may include: a decoder receiving arranged syntax elements in a Sequence Parameter Set (SPS) level, wherein the arranged syntax elements in the SPS level are arranged such that the functions of related syntax elements are grouped in a Generic Video Codec (VVC) syntax at the coding level. The decoder may also receive a second syntax element immediately following the plurality of syntax elements in response to a plurality of syntax elements satisfying predefined conditions. The decoder may also perform related syntax element functions on video data from a bitstream based on the plurality of syntax elements and the second syntax element.

[0006] According to a second aspect of this disclosure, a method for decoding a video signal is provided. The method may include: a decoder receiving arranged syntax elements in a Sequence Parameter Set (SPS) level, wherein the arranged syntax elements in the SPS level are arranged such that inter-frame prediction related syntax elements are grouped in a General Video Coding (VVC) syntax at the coding level. The decoder may also obtain a first reference picture associated with a video block in the bitstream. Second reference image In order of display, the first reference image Before the current image, and the second reference image. Following the current image, the decoder can also access the first reference image. The reference block in the video block obtains the first prediction sample of the video block. i and j can represent the coordinates of a sample point in the current image. The decoder can also obtain coordinates from the second reference image. The reference block in the video block obtains the second prediction sample. The decoder can also base its work on the arranged syntax elements in the SPS level and the first prediction sample. and the second predicted sample Obtain bidirectional prediction samples.

[0007] According to a third aspect of this disclosure, a computing device is provided. The computing device may include: one or more processors; and a non-transitory computer-readable storage medium storing instructions executable by the one or more processors. The one or more processors may be configured to receive arranged syntax elements in a Sequence Parameter Set (SPS) level. The arranged syntax elements in the SPS level are arranged such that the functions of related syntax elements are grouped in a Generic Video Coding (VVC) syntax at the coding level. The one or more processors may also be configured to receive a second syntax element immediately following the plurality of syntax elements in response to the plurality of syntax elements satisfying predefined conditions. The one or more processors may also be configured to perform the functions of related syntax elements on video data from a bitstream based on the plurality of syntax elements and the second syntax element.

[0008] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided, on which instructions are stored. When executed by one or more processors of a device, the instructions cause the device to receive arranged syntax elements at the Sequence Parameter Set (SPS) level, wherein the arranged syntax elements at the SPS level are arranged such that inter-frame prediction related syntax elements are grouped in a generic video codec (VVC) syntax at the coding level. The instructions also cause the device to obtain a first reference picture associated with a video block in the bitstream. Second reference image In order of display, the first reference image Before the current image, and the second reference image. Following the current image, the instruction causes the device to move from the first reference image. The reference block in the video block obtains the first prediction sample of the video block. i and j represent the coordinates of a sample point in the current image. The instruction enables the device to access the second reference image. The reference block in the video block obtains the second prediction sample. The instructions enable the device to base its analysis on the arranged syntax elements in the SPS level and the first prediction sample. and the second predicted sample Obtain bidirectional prediction samples.

[0009] It should be understood that the general description above and the detailed description below are exemplary and illustrative only and are not intended to limit this disclosure. Attached Figure Description

[0010] The accompanying drawings are incorporated in and form a part of this specification. The drawings illustrate examples consistent with this disclosure and, together with the specification, serve to explain the principles of this disclosure.

[0011] Figure 1 This is a block diagram of an encoder based on an example of this disclosure.

[0012] Figure 2 This is a block diagram of a decoder based on an example of this disclosure.

[0013] Figure 3A This is a diagram illustrating block segmentation in a multi-type tree structure according to an example of this disclosure.

[0014] Figure 3B This is a diagram illustrating block segmentation in a multi-type tree structure according to an example of this disclosure.

[0015] Figure 3C This is a diagram illustrating block segmentation in a multi-type tree structure according to an example of this disclosure.

[0016] Figure 3D This is a diagram illustrating block segmentation in a multi-type tree structure according to an example of this disclosure.

[0017] Figure 3E A diagram illustrating block segmentation in a multi-type tree structure according to an example of this disclosure.

[0018] Figure 4 This is an example of a method for decoding video signals according to this disclosure.

[0019] Figure 5 This is an example of a method for decoding video signals according to this disclosure.

[0020] Figure 6 This is an example of a method for decoding video signals according to this disclosure.

[0021] Figure 7 This is a diagram illustrating a computing environment coupled to a user interface according to an example of this disclosure. Detailed Implementation

[0022] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, wherein like reference numerals in different drawings denote like or similar elements unless otherwise indicated. The implementations set forth in the following description of the embodiments do not represent all implementations consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with aspects related to this disclosure as set forth in the appended claims.

[0023] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. As used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein is intended to represent and include any or all possible combinations of one or more of the associated listed items.

[0024] It should be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various types of information, the information should not be limited by these terms. These terms are used only to distinguish one type of information from another. For example, without departing from the scope of this disclosure, first information may be referred to as second information; and similarly, second information may be referred to as first information. As used herein, depending on the context, the term “if” may be understood to mean “when,” “once,” or “in response to a judgment.”

[0025] The first version of the HEVC standard was completed in October 2013. Compared to its predecessor, H.264 / MPEG-AVC, the first version of HEVC offered approximately 50% bitrate savings or equivalent perceived quality. Despite the significant codec improvements offered by HEVC compared to its predecessor, evidence suggested that superior codec efficiency could be achieved using additional codec tools. Based on this, both VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-T VECG and ISO / IEC MPEG formed a Joint Video Exploration Group (JVET) to begin important research into advanced technologies that could significantly improve codec efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating several additional codec tools on top of the HEVC test model (HM).

[0026] In October 2017, the ITU-T and ISO / IEC issued a joint call for proposals (CfP) on video compression capabilities exceeding HEVC. In April 2018, at the 10th JVET meeting, 23 CfP responses were received and evaluated, demonstrating compression efficiency gains exceeding HEVC by approximately 40%. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard named Universal Video Codec (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0027] Similar to HEVC, VVC is built on a block-based hybrid video codec framework.

[0028] Figure 1 A general diagram of a block-based video encoder for VVC is shown. Specifically, Figure 1 A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-frame mode decision 116, block prediction value 140, adder 128, transform 130, quantization 132, prediction related information 142, intra-frame prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, loop filter 122, entropy coding 138, and bitstream 144.

[0029] In encoder 100, video frames are divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method.

[0030] The prediction residual, representing the difference between the current video block (a portion of video input 110) and its predicted value (a portion of block prediction value 140), is sent from adder 128 to transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantized coefficients are then fed into entropy coding 138 to generate a compressed video bitstream. Figure 1 As shown, prediction-related information 142 from the intra / inter-frame mode decision 116 (such as video block segmentation information, motion vectors (MV), reference picture index, and intra-frame prediction mode) is also fed into and stored in the compressed bitstream 144 via entropy coding 138. The compressed bitstream 144 includes the video bitstream.

[0031] In encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed via inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with block prediction values ​​140 to generate unfiltered reconstructed pixels for the current video block.

[0032] Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of encoded neighboring blocks in the same video frame as the current video block to predict the current video block.

[0033] Timing prediction (also known as "inter-frame prediction") uses reconstructed pixels from encoded video frames to predict the current video block. Timing prediction reduces the inherent temporal redundancy in the video signal. The timing prediction signal for a given coding unit (CU) or coding block is typically represented by one or more MVs, which indicate the amount and direction of motion between the current CU and its timing reference. Additionally, if multiple reference frames are supported, a reference frame index is sent to identify which reference frame in the reference frame storage device the timing prediction signal originates from.

[0034] Motion estimation 114 receives video input 110 and signals from image buffer 120, and outputs the motion estimation signal to motion compensation 112. Motion compensation 112 receives video input 110, signals from image buffer 120, and motion estimation signals from motion estimation 114, and outputs the motion compensation signal to intra / inter-frame mode decision 116.

[0035] After performing spatial and / or temporal predictions, the intra / inter-frame mode decision 116 in encoder 100 selects the optimal prediction mode, for example, based on a rate-distortion optimization method. The block prediction value 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantized residual coefficients are dequantized by inverse quantization 134 and inverse transformed by inverse transform 136 to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed CU signal. Before the reconstructed CU is placed in the reference picture storage device of picture buffer 120 and used for encoding and decoding future video blocks, loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive loop filter (ALF), can be further applied to the reconstructed CU. To form the output video bitstream 144, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are all sent to entropy coding unit 138 for further compression and packing to form the bitstream.

[0036] Figure 1A block diagram of a general block-based hybrid video coding system is presented. The input video signal is processed block by block (called CU). In VTM-1.0, a CU can be up to 128×128 pixels. However, unlike HEVC, which partitions blocks based solely on quadtrees, in VVC, a coding tree unit (CTU) is split into multiple CUs based on quadtrees / binaries / tritrees to accommodate varying local characteristics. Furthermore, the concept of multiple partitioning unit types in HEVC is removed; that is, in VVC, there is no longer a distinction between CUs, prediction units (PUs), and transform units (TUs); instead, each CU is always used as the basic unit for both prediction and transform, without further partitioning.

[0037] In a multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary and ternary tree structures.

[0038] like Figure 3A , Figure 3B , Figure 3C , Figure 3D and Figure 3E As shown, there are five types of splitting: quadruple splitting, horizontal binary splitting, vertical binary splitting, horizontal ternary splitting, and vertical ternary splitting.

[0039] Figure 3A A diagram illustrating block quadrilateral partitioning in a multi-type tree structure according to this disclosure is shown.

[0040] Figure 3B The diagram illustrates a block vertical binary partitioning in a multi-type tree structure according to the present disclosure.

[0041] Figure 3C The diagram illustrates a block-level binary partitioning in a multi-type tree structure according to the present disclosure.

[0042] Figure 3D The diagram illustrates a block vertical ternary segmentation in a multi-type tree structure according to the present disclosure.

[0043] Figure 3E The diagram illustrates a block-level ternary partitioning in a multi-type tree structure according to the present disclosure.

[0044] exist Figure 1In video signals, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples (called reference samples) of coded neighboring blocks in the same video picture / strip to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion-compensated prediction") uses reconstructed pixels from coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal.

[0045] The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Additionally, if multiple reference images are supported, a reference image index is sent to identify which reference image in the reference image repository the temporal prediction signal originates from. Following spatial and / or temporal prediction, a mode decision block in the encoder selects the optimal prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block, and the prediction residuals are decorrelated using a transform and quantized.

[0046] The quantized residual coefficients are dequantized and inversely transformed to form the reconstructed residuals, which are then added back to the prediction block to form the reconstructed CU signal. Furthermore, before the reconstructed CU is placed in the reference image repository and used for encoding and decoding future video blocks, loop filtering, such as deblocking filters, Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF), can be applied to it. To form the output video bitstream, the coding mode (inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit for further compression and packing to form the bitstream.

[0047] Figure 2 A general block diagram of a video decoder for VVC is shown. Specifically, Figure 2 A typical block diagram of decoder 200 is shown. Decoder 200 has a bitstream 210, entropy decoding 212, dequantization 214, inverse transform 216, adder 218, intra / inter-frame mode selection 220, intra-frame prediction 222, memory 230, loop filter 228, motion compensation 224, image buffer 226, prediction-related information 234, and video output 232.

[0048] Decoder 200 is similar to residing in Figure 1The reconstruction-related part is located in the encoder 100. In the decoder 200, the input video bitstream 210 is first decoded by entropy decoding 212 to derive the quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by inverse quantization 214 and inverse transform 216 to obtain the reconstructed prediction residuals. The block prediction mechanism implemented in the intra / inter-frame mode selector 220 is configured to perform intra-frame prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residuals from inverse transform 216 and the prediction output generated by the block prediction mechanism using adder 218.

[0049] Before the reconstructed blocks are stored in image buffer 226, which serves as a reference image repository, the reconstructed blocks can be further passed through loop filter 228. The reconstructed video in image buffer 226 can be sent to drive a display device, as well as to predict future video blocks. With loop filter 228 open, filtering operations are performed on these reconstructed pixels to produce the final reconstructed video output 232.

[0050] Figure 2 A general block diagram of a block-based video decoder is given. The video bitstream is first entropy-decoded at the entropy decoding unit. The coding mode and prediction information are sent to the spatial prediction unit (if intra-frame coded) or the temporal prediction unit (if inter-frame coded) to form prediction blocks. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual blocks. The prediction blocks and residual blocks are then added together. The reconstructed blocks can be further filtered through a loop before being stored in the reference picture repository. The reconstructed video in the reference picture repository is then sent to drive the display device, along with the video blocks used to predict future blocks.

[0051] Typically, the basic intra prediction scheme used in VVC is the same as that used in HEVC, except that the basic intra prediction scheme used in VVC is further extended and / or improved in several modules. Examples include Matrix Weighted Intra Prediction (MIP) coding mode, Intra-Segmentation (ISP) coding mode, Extended Intra Prediction with Wide-Angle Intra-Direction, Position-Related Intra Prediction Combination (PDPC), and 4-Tap Frame Interpolation. The main focus of this disclosure is to improve the existing high-level syntax design in the VVC standard. The relevant background information is described in detail in the following sections.

[0052] Like HEVC, VVC uses a bitstream structure based on NAL units. The encoded and decoded bitstream is divided into NAL units, which should be smaller than the maximum transmission unit size when transmitted over lossy packet networks. Each NAL unit consists of a NAL unit header and a subsequent NAL unit payload. There are two conceptual categories of NAL units: Video Coding Layer (VCL) NAL units containing encoded sample data, such as encoded stripe NAL units, and non-VCL NAL units containing metadata that typically belong to more than one encoded picture, or non-VCL NAL units whose association with a single encoded picture would be meaningless, such as parameter set NAL units, or non-VCL NAL units for which information is not needed during the decoding process, such as SEI NAL units.

[0053] In VVC, a two-byte NAL unit header is introduced, and this design is expected to be sufficient to support future expansions. The syntax and associated semantics of the NAL unit header in the current VVC draft specification are shown in Tables 1 and 2, respectively. Instructions for reading Table 1 can be found in the VVC specification.

[0054] Table 1. NAL Unit Header Syntax

[0055] Table 2. Semantics of NAL Unit Headers

[0056] Table 3. NAL Unit Type Codes and NAL Unit Type Categories

[0057]

[0058] VVC inherits the parameter set concept from HEVC and makes some modifications and additions. A parameter set can be part of the video bitstream or can be received by the decoder through other means, including out-of-band transmission using reliable channels, hard encoding / decoding in the encoder and decoder, etc. The parameter set contains identifiers that are referenced directly or indirectly from the stripe header, as discussed in more detail later. This referencing process is called "activation." Depending on the parameter set type, activation occurs by picture or by sequence. The concept of activation by reference was introduced, among other reasons, because implicit activation by means of positional information in the bitstream (common for other syntax elements of video codecs) is not available in the case of out-of-band transmission.

[0059] Video Parameter Sets (VPSs) are introduced to pass information applicable to multiple layers and sublayers. VPSs were introduced to address these shortcomings and to achieve a concise and scalable high-level design for multi-layer codecs. Each layer of a given video sequence (regardless of whether it has the same or different Sequence Parameter Sets (SPSs)) references the same VPS. Table 4 shows the syntax of video parameter sets in the current VVC draft specification. How to read Table 4 is shown in the appendix of this publication, which can also be found in the VVC specification.

[0060] Table 4. Video Parameter Set RBSP Syntax

[0061]

[0062]

[0063] In VVC, the SPS contains information applicable to all stripes of an encoded video sequence. An encoded video sequence begins with an Instantaneous Decoding Refresh (IDR) picture, BLA picture, or CRA picture, which is the first picture in the bitstream, and includes all subsequent pictures that are not IDR or BLA pictures. The bitstream consists of one or more encoded video sequences. The contents of the SPS can be roughly subdivided into six categories: 1) Self-reference (its own ID); 2) Decoder operation point information (profile, level, picture size, number of sublayers, etc.); 3) Enable flags for certain tools within the profile, and associated codec tool parameters when tools are enabled; 4) Information limiting the flexibility of encoding and decoding of structure and transform coefficients; 5) Temporal scalability controls; and 6) Visual Usability Information (VUI), which includes HRD information. The syntax and associated semantics of the sequence parameters set in the current VVC draft specification are shown in Tables 5 and 6, respectively. How to read Table 5 is shown in the appendix of this disclosure and can also be found in the VVC specification.

[0064] Table 5. Sequence Parameter Set (RBSP) Syntax

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071] Table 6. Sequence Parameter Set (RBSP) Semantics

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095] The Picture Parameter Set (PPS) of VVC contains this information that can be changed between pictures. The PPS includes information roughly equivalent to a portion of the PPS in HEVC, including: 1) self-reference; 2) initial picture control information, such as initial quantization parameters (QP), multiple flags indicating the use or presence of control information in certain tools or strip headers; and 3) tile information. The syntax and related semantics of the Picture Parameter Set in the current VVC draft specification are shown in Tables 7 and 8, respectively. How to read Table 7 is shown in the appendix of this disclosure, which can also be found in the VVC specification.

[0096] Table 7. Image Parameter Set RBSP Syntax

[0097]

[0098]

[0099]

[0100] Table 8. Image Parameter Set RBSP Semantics

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114] The strip header contains information that can be changed strip-by-strip, as well as relatively small or only relevant image-related information specific to a particular strip or image type. The size of the strip header can be significantly larger than the PPS, especially when there are tile or wavefront entry point offsets in the strip header and modifications to the RPS, prediction weights, or reference image list are explicitly transmitted via the signal. Table 10 shows the syntax of the image header in the current VVC draft specification. How to read Table 10 is shown in the appendix of this disclosure, which can also be found in the VVC specification.

[0115] Table 10. Image Header Structure Syntax

[0116]

[0117]

[0118]

[0119]

[0120] Improvements to syntactic elements In the current VVC, when similar syntax elements exist for intra-frame prediction and inter-frame prediction respectively, in some places, the syntax elements related to inter-frame prediction are defined before those related to intra-frame prediction. Given that intra-frame prediction is allowed in all picture / strip types but inter-frame prediction is not, this order may not be preferred. From a standardization perspective, it would be beneficial to always define the intra-frame prediction-related syntax before the syntax for inter-frame prediction.

[0121] It was also observed that in the current VVC, some highly related syntactic elements are defined in different locations in an extended manner. From a standardization perspective, grouping some syntactic elements together is also beneficial.

[0122] The proposed method In this disclosure, methods for simplifying and / or further improving existing designs of high-level grammars are provided to address the problems identified in the "Problem Statement" section. Note that the methods of this disclosure may be applied independently or in combination.

[0123] Grouping the segmentation constraint syntax elements according to prediction type In this disclosure, a rearrangement of syntax elements is proposed such that syntax elements related to intra-frame prediction are defined before those related to inter-frame prediction. According to this disclosure, segmentation constraint syntax elements are grouped by prediction type, with intra-frame prediction-related syntax elements preceding inter-frame prediction-related syntax elements. In one embodiment, the order of segmentation constraint syntax elements in the SPS is consistent with the order of segmentation constraint syntax elements in the image header. An example of the decoding process on the VVC draft is shown in Table 11 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0124] Table 11 presents the RBSP syntax for the sequence parameter set.

[0125] Grouping of two-tree chroma syntax elements In this disclosure, grouping of syntax elements associated with dual-tree chroma types is proposed. In one embodiment, the segmentation constraint syntax elements for dual-tree chroma in SPS should be sent together via signaling in the dual-tree chroma case. An example of the decoding process on the VVC draft is shown in Table 12 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0126] Table 12. Proposed Sequence Parameter Set (RBSP) Syntax

[0127] If we also consider defining intra-prediction related syntax before inter-prediction related syntax, then another example of the decoding process on the VVC draft is shown in Table 13 below, according to the method of this disclosure. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0128] Table 13. Proposed Sequence Parameter Set (RBSP) Syntax

[0129] Conditionally transmit inter-frame prediction related syntax elements via signaling As previously described, current VVC allows intra-frame prediction in all picture / strip types, but not inter-frame prediction in all picture / strip types. According to this disclosure, it is proposed to add a flag to the VVC syntax at a specific coding level to indicate whether inter-frame prediction is allowed in sequences, pictures, and / or stripes. If inter-frame prediction is not allowed, the inter-frame prediction-related syntax is not signaled at the corresponding coding level (e.g., sequence, picture, and / or strip level).

[0130] According to this disclosure, it is also proposed to add flags to the VVC syntax at a specific coding level to indicate whether inter-frame striping, such as P-striping and B-striping, is allowed in sequences, pictures, and / or stripes. If inter-frame striping is not allowed, the inter-frame striping-related syntax is not signaled at the corresponding coding level (e.g., sequence, picture, and / or stripe level).

[0131] The following sections provide some examples based on the proposed inter-frame stripe allow flag. Furthermore, the proposed inter-frame prediction allow flag can be used in a similar manner.

[0132] When proposed inter-frame striping allowance flags are added at different levels, these flags can be signaled in a hierarchical manner. When a flag signaled at a higher level indicates that inter-frame striping is not allowed, it is not necessary to signal a flag at a lower level, and it can be inferred that the flag is 0 (meaning that inter-frame striping is not allowed).

[0133] In one example, according to the method of this disclosure, a flag is added to the SPS to indicate whether inter-frame striping is allowed when encoding the current video sequence. If inter-frame striping is not allowed, inter-striping related syntax elements are not signaled in the SPS. An example of the decoding process on the VVC draft is shown in Table 14 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font. It should be noted that there are syntax elements other than those introduced in the example. For example, there are many inter-frame striping (or inter-frame prediction tool) related syntax elements such as `sps_weighted_pred_flag`, `sps_temporal_mvp_enabled_flag`, `sps_amvr_enabled_flag`, `sps_bdof_enabled_flag`, etc.; there are also syntax elements related to the list of reference pictures, such as `long_term_ref_pics_flag`, `inter_layer_ref_pics_present_flag`, `sps_idr_rpl_present_flag`, etc. All these syntax elements related to inter-frame prediction can be optionally controlled by the proposed flags.

[0134] Table 14. Proposed Sequence Parameter Set (RBSP) Syntax

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141] 7.4.3.3 Sequence Parameter Set (RBSP) Semantics A `sps_inter_slice_allowed_flag` value of 0 specifies that all encoded slices in the video sequence have a `slice_type` value of 2 (indicating that the encoded slices are I slices). A `sps_inter_slice_allowed_flag` value of 1 specifies that one or more encoded slices in the video sequence may or may not have a `slice_type` value of 0 (indicating that the encoded slices are P slices) or 1 (indicating that the encoded slices are B slices).

[0142] In another example, according to the method of this disclosure, a flag is added to the Picture Parameter Set (PPS) to indicate whether inter-frame striping is allowed when encoding pictures associated with that PPS. If inter-frame striping is not allowed, the selected inter-frame prediction related syntax elements are not signaled in the PPS.

[0143] In yet another example, according to the method of this disclosure, an inter-slice allow flag can be sent via signaling in a layered manner. A flag (e.g., sps_inter_slice_allowed_flag) is added to the SPS to indicate whether inter-sliceing is allowed when encoding the picture associated with that SPS. When sps_inter_slice_allowed_flag equals 0 (meaning inter-sliceing is not allowed), the inter-slice allow flag in the picture header sent via signaling can be omitted and inferred to be 0. An example of the decoding process on the VVC draft is shown in Table 15 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0144] Table 15. Proposed Sequence Parameter Set (RBSP) Syntax

[0145] 7.4.3.7 Image Header Structure Semantics A `ph_inter_slice_allowed_flag` value of 0 indicates that all encoded slices in the image have a `slice_type` of 2. A `ph_inter_slice_allowed_flag` value of 1 indicates that one or more encoded slices with a `slice_type` of 0 or 1 may or may not exist in the image. When `ph_inter_slice_allowed_flag` does not exist, it is inferred that the value of `ph_inter_slice_allowed_flag` is equal to 0.

[0146] Grouping inter-frame related syntax elements In this disclosure, a rearrangement of syntax elements is proposed such that inter-frame prediction-related syntax elements are grouped within the VVC syntax at a specific coding level (e.g., sequence, picture, and / or stripe level). According to this disclosure, a rearrangement of syntax elements related to inter-frame stripes in the Sequence Parameter Set (SPS) is proposed. An example of the decoding process on the VVC draft is shown in Table 16 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0147] Table 16. Proposed Sequence Parameter Set (RBSP) Syntax

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154] Table 17 below shows another example of the decoding process on the VVC draft. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0155] Table 17. Proposed Sequence Parameter Set (RBSP) Syntax

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162] In yet another example, the decoding process on the VVC draft is shown in Table 18 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0163] Table 18. Proposed Sequence Parameter Set (RBSP) Syntax

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170] In yet another example, the decoding process on the VVC draft is shown in Table 19 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0171] Table 19. Proposed Sequence Parameter Set (RBSP) Syntax

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178] Figure 4 A method for decoding video signals according to this disclosure is shown. For example, the method can be applied to a decoder.

[0179] In step 410, the decoder can receive the arranged syntax elements in the Sequence Parameter Set (SPS) level via a bitstream. The arranged syntax elements in the SPS level can be arranged such that the functionality of the related syntax elements is grouped in the general video coding (VVC) syntax at the coding level.

[0180] In step 412, the decoder can receive a second syntax element immediately following the multiple syntax elements via the bitstream and in response to multiple syntax elements satisfying a predefined condition. For example, the multiple syntax elements may include the `sps_mmvd_enabled_flag` and `sps_fpel_mmvd_enabled_flag` flags. For example, the predefined condition may include the `sps_mmvd_enabled_flag` flag being equal to 1.

[0181] In step 414, the decoder can perform relevant syntax element functions on the video data from the bitstream based on multiple syntax elements and a second syntax element.

[0182] According to this disclosure, it is also proposed to add a flag in the VVC syntax at a specific coding level to indicate whether inter-sliceing, such as P-slices and B-slices, is allowed in sequences, pictures, and / or slices. If inter-sliceing is not allowed, the inter-slice-related syntax is not signaled at the corresponding coding level (e.g., sequence, picture, and / or slice level). In one example, according to the method of this disclosure, the flag `sps_inter_slice_allowed_flag` is added to the SPS to indicate whether inter-sliceing is allowed when encoding the current video sequence. If not allowed, the inter-slice-related syntax elements are not signaled in the SPS. An example of the decoding process on the VVC draft is shown in Table 20 below. Added parts are shown in bold and italic font, while deleted parts are shown in strikethrough font.

[0183] Table 20. Proposed Sequence Parameter Set (RBSP) Syntax

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190] Table 21 below shows another example of the decoding process on the VVC draft. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0191] Table 21. Proposed Sequence Parameter Set (RBSP) Syntax

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198] Grouping similar functional syntax elements This disclosure proposes rearranging syntax elements such that similar functions (e.g., intra-frame tools, inter-frame tools, screen content tools, transform tools, quantization tools, loop filter tools, and / or segmentation tools) and related syntax elements are grouped within a specific coding level (e.g., sequence, picture, and / or stripe level) of the VVC syntax. According to this disclosure, a rearrangement of syntax elements in the Sequence Parameter Set (SPS) is proposed such that related syntax elements with similar functions are grouped. An example of the decoding process on the VVC draft is shown in Table 23 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0199] Table 23. Proposed Sequence Parameter Set (RBSP) Syntax

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206] Table 24 below shows another example of the decoding process on the VVC draft. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0207] Table 24. Proposed Sequence Parameter Set RBSP Syntax

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215] According to this disclosure, a rearrangement of the grammatical elements in the Picture Parameter Set (PPS) is proposed, such that related grammatical elements with similar functions are grouped together. An example of the decoding process on the VVC draft is shown in Table 25 below. Changes to the VVC draft are shown in bold and italic font, while deleted portions are shown in strikethrough font.

[0216] Table 25. Proposed Sequence Parameter Set (RBSP) Syntax

[0217]

[0218]

[0219]

[0220] Table 26 below shows another example of the decoding process on the VVC draft. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0221] Table 26. Proposed Sequence Parameter Set (RBSP) Syntax

[0222]

[0223]

[0224]

[0225] In yet another example, the decoding process on the VVC draft is shown in Table 27 below. Changes to the VVC draft are shown in bold and italics, while deleted portions are shown in strikethrough font.

[0226] Table 27. Proposed Sequence Parameter Set (RBSP) Syntax

[0227]

[0228]

[0229]

[0230] Figure 5 A method for decoding a video signal according to this disclosure is shown. This method can be applied, for example, to a decoder.

[0231] In step 510, the decoder may receive the arranged syntax elements in the SPS level, such that the inter-frame prediction related syntax elements are grouped in the VVC syntax at the coding level.

[0232] In step 512, the decoder obtains the first reference picture associated with the video block in the bitstream. Second reference image The first reference image is shown in the order it is displayed. A second reference image preceding the current image. Following the current image.

[0233] In step 514, the decoder can obtain the first reference image. The reference block in the video block obtains the first prediction sample of the video block. i and j represent the coordinates of a sample point in the current image.

[0234] In step 516, the decoder can obtain the second reference image. The reference block in the video block obtains the second prediction sample. .

[0235] In step 518, the decoder may base its work on the permuted syntax elements in the SPS level and the first prediction samples. Second prediction sample Obtain bidirectional prediction samples.

[0236] Figure 6 A method for decoding a video signal according to this disclosure is shown. This method can be applied, for example, to a decoder.

[0237] In step 610, the decoder may receive a bitstream including VPS, SPS, PPS, picture header, and stripe header for the encoded video data.

[0238] In step 612, the decoder can decode the VPS.

[0239] In step 614, the decoder can decode the SPS and obtain the arranged segmentation constraint syntax elements in the SPS level.

[0240] In step 616, the decoder can decode the PPS.

[0241] In step 618, the decoder can decode the image header.

[0242] In step 620, the decoder can decode the stripe header.

[0243] In step 622, the decoder can decode the video data based on VPS, SPS, PPS, image header, and strip header.

[0244] The methods described above can be implemented using a device comprising one or more circuits, including application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The device can be used in combination with other hardware or software components to perform the methods described above. Each module, submodule, unit, or subunit disclosed above can be implemented at least partially using one or more circuits.

[0245] Figure 7 A computing environment 710 coupled to a user interface 760 is shown. The computing environment 710 may be part of a data processing server. The computing environment 710 includes a processor 720, memory 740, and I / O interface 750.

[0246] Processor 720 typically controls the overall operation of computing environment 710, such as operations associated with display, data acquisition, data communication, and image processing. Processor 720 may include one or more processors to execute instructions to perform all or some of the steps in the methods described above. Furthermore, processor 720 may include one or more modules that facilitate interaction between processor 720 and other components. The processor may be a central processing unit (CPU), microprocessor, microcontroller, GPU, etc.

[0247] Memory 740 is configured to store various types of data to support the operation of computing environment 710. Memory 740 may include predefined software 742. Examples of such data include instructions for any application or method operating on computing environment 710, video datasets, image data, etc. Memory 740 can be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0248] I / O interface 750 provides an interface between processor 720 and peripheral interface modules (such as keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 750 can be coupled to encoders and decoders.

[0249] In some embodiments, a non-transitory computer-readable storage medium is also provided, which includes a plurality of programs, such as those included in memory 740 and executable by processor 720 in computing environment 710, for performing the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0250] A non-transitory computer-readable storage medium stores a plurality of programs, which are executed by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the methods for motion prediction described above.

[0251] In some embodiments, the computing environment 710 may be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0252] In light of the specification and practice of this disclosure herein, other examples of this disclosure will be apparent to those skilled in the art. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure, and includes such deviations from this disclosure within the scope of known or customary practice in the art. The specification and examples are intended to be considered illustrative only.

[0253] It should be understood that this disclosure is not limited to the exact examples shown above and in the accompanying drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for decoding a video signal, comprising: The sequence parameter set (SPS) level is received via a bitstream, wherein the arranged syntax elements in the SPS level are obtained by arranging and grouping syntax elements related to predefined functions into groups corresponding to the predefined functions.

2. The method according to claim 1, wherein the predefined functions include intra-frame tools, inter-frame tools, screen content tools, transformation tools, quantization tools, loop filter tools, or segmentation tools.

3. The method of claim 1, wherein the arranged syntax elements in the received sequence parameter set SPS level comprise: In response to the first syntax element in the arranged syntax elements satisfying a predefined condition, a second syntax element immediately following the first syntax element in the arranged syntax elements is received. The predefined function is executed on the video data from the bitstream according to the first syntax element and the second syntax element.

4. The method according to claim 3, further comprising: In response to the first syntax element not satisfying the predefined condition, the value of the second syntax element is set.

5. The method of claim 3, wherein the first syntax element is the sps_mmvd_enabled_flag flag, the second syntax element is a flag related to the use of integer sample precision in the motion vector difference-based merging mode, and the predefined condition includes sps_mmvd_enabled_flag equal to 1; or The first syntax element is the `sps_affine_enabled_flag` flag, the second syntax element is the value of `five_minus_max_num_subblock_merge_cand` or the `sps_affine_prof_enabled_flag` flag, and the predefined condition includes the `sps_affine_enabled_flag` flag being equal to 1; or The first syntax element is the sps_affine_prof_enabled_flag flag, the second syntax element is the sps_prof_control_present_in_ph_flag flag, and the predefined condition includes sps_affine_prof_enabled_flag equal to 1.

6. The method according to claim 1, further comprising: Receive the arranged syntax elements in the PPS level of the image parameter set, wherein the arranged syntax elements in the PPS related to the predefined function are grouped in the PPS level.

7. A method for encoding video signals, comprising: Obtain the arranged syntax elements in the sequence parameter set SPS level, wherein the arranged syntax elements in the SPS level are obtained by arranging and grouping the syntax elements related to the predefined functions into groups corresponding to the predefined functions; Generate a bitstream including the arranged syntax elements; and Send the bit stream.

8. The method of claim 7, wherein the predefined functions include intra-frame tools, inter-frame tools, screen content tools, transformation tools, quantization tools, loop filter tools, or segmentation tools.

9. The method of claim 7, wherein obtaining the arranged syntax elements in the sequence parameter set SPS level comprises: In response to the first syntax element in the arranged syntax elements satisfying a predefined condition, the second syntax element immediately following the first syntax element in the arranged syntax elements is obtained. The predefined function is executed on the video data according to the first syntax element and the second syntax element.

10. The method of claim 9, further comprising: In response to the first syntax element not satisfying the predefined condition, the value of the second syntax element is set.

11. The method of claim 9, wherein the first syntax element is the sps_mmvd_enabled_flag flag, the second syntax element is a flag related to the use of integer sample precision in the motion vector difference-based merging mode, and the predefined condition includes sps_mmvd_enabled_flag equal to 1; or The first syntax element is the `sps_affine_enabled_flag` flag, the second syntax element is the value of `five_minus_max_num_subblock_merge_cand` or the `sps_affine_prof_enabled_flag` flag, and the predefined condition includes the `sps_affine_enabled_flag` flag being equal to 1; or The first syntax element is the sps_affine_prof_enabled_flag flag, the second syntax element is the sps_prof_control_present_in_ph_flag flag, and the predefined condition includes sps_affine_prof_enabled_flag equal to 1.

12. The method of claim 7, further comprising: Obtain the arranged syntax elements in the image parameter set PPS level, wherein the arranged syntax elements in the PPS related to the predefined function are grouped in the PPS level.

13. A computing device, comprising: One or more processors; as well as A non-transitory computer-readable storage medium storing instructions executable by the one or more processors. When executed by the one or more processors, the instructions cause the computing device to perform a method for decoding a video signal according to any one of claims 1-6 or a method for encoding a video signal according to any one of claims 7-12.

14. A non-transitory computer-readable storage medium storing a bit stream formed by instructions, wherein the instructions, when executed by a computing device having one or more processors, cause the one or more processors to perform a method for encoding a video signal according to any one of claims 7-12 to generate the bit stream.

15. A computer program product having instructions for storing a bitstream, wherein, The bit stream includes: Video data decoded by the method for decoding video signals according to any one of claims 1-6; or Video data generated by the method for encoding video signals according to any one of claims 7-12.

16. A method for storing a bit stream, comprising: Perform the method for encoding a video signal according to any one of claims 7-12 to generate a bitstream; as well as Store the bit stream.

17. A method for transmitting a bit stream, comprising: Perform the method for encoding a video signal according to any one of claims 7-12 to generate a bitstream; as well as Send the bit stream.