Method and apparatus for advanced syntax in video coding

By optimizing the reception and processing of syntax elements in the video bitstream and adopting predefined functions, the problem of low encoding and decoding efficiency of high-resolution video data is solved, and more efficient video data processing is achieved.

CN120475178APending Publication Date: 2025-08-12BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510612042.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-04-10
Filing Date
2021-04-12
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are inefficient when processing high-resolution video data, making it difficult to achieve more efficient encoding and decoding while maintaining image quality.

Method used

In the video bitstream, syntax elements are received and set according to predefined conditions, and predefined functions such as intra prediction, inter prediction and merging mode are used to optimize the encoding and decoding process of video data.

Benefits of technology

It improves the encoding efficiency of video data, reduces redundant information, and improves the decoding performance of high-resolution video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475178A_ABST
    Figure CN120475178A_ABST
Patent Text Reader

Abstract

An electronic device performs a method of encoding video data. The method includes encoding a plurality of syntax elements of one or more of a picture parameter set (PPS) level and a slice level into a bitstream, where the plurality of syntax elements are associated with a predefined function and are sequentially arranged in the bitstream; in accordance with a determination that the plurality of syntax elements satisfy a predefined condition: encoding a second syntax element after the plurality of syntax elements into the bitstream; in accordance with a determination that the plurality of syntax elements do not meet the predefined condition: setting a default value as the second syntax element; and performing the predefined function on video data from the bitstream according to the plurality of syntax elements and the second syntax element, where the predefined function is an intra prediction function, an inter prediction function, or a merge mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 008,648, filed on April 10, 2020, entitled “HIGH-LEVEL SYNTAX FOR VIDEOCODING,” which is incorporated by reference in its entirety. Technical Field

[0003] The present application relates generally to video data coding and compression, and more particularly to methods and systems for video coding and decoding high-level syntax in a video bitstream applicable to one or more video codec standards. Background Art

[0004] Various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. support digital video. Electronic devices transmit, receive, encode, decode and / or store digital video data by implementing video compression / decompression standards defined by MPEG-4, ITU-T H.263, ITU-TH.264 / MPEG-4 Part 10, Advanced Video Codec (AVC), High Efficiency Video Codec (HEVC) and Versatile Video Codec (VVC) standards. Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or remove redundancy inherent in video data. For block-based video codecs, a video frame is partitioned into one or more slices, each slice having multiple video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain a coding unit (CU) or be recursively partitioned into smaller CUs until a predefined minimum CU size is reached. Each CU (also called a leaf-CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be coded in intra, inter, or IBC mode. Video blocks in an intra (I) slice of a video frame are coded using spatial prediction with respect to reference samples in neighboring blocks within the same video frame. Video blocks in an inter (P or B) slice of a video frame can use spatial prediction with respect to reference samples in neighboring blocks within the same video frame or temporal prediction with respect to reference samples in other previous and / or future reference video frames.

[0005] A prediction block for the current video block to be coded is generated based on spatial or temporal prediction of previously coded reference blocks (e.g., neighboring blocks). The process of finding the reference block can be accomplished using a block matching algorithm. The residual data representing the pixel differences between the current block to be coded and the prediction block is called a residual block or prediction error. Inter-coded blocks are encoded based on motion vectors pointing to reference blocks in the reference frames that form the prediction block, and the residual block. The process of determining the motion vector is typically called motion estimation. Intra-coded blocks are encoded based on the intra-frame prediction mode and the residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, such as the frequency domain, to produce residual transform coefficients, which can then be quantized. The quantized transform coefficients, initially arranged in a two-dimensional array, can be scanned to produce a one-dimensional vector of transform coefficients, which can then be entropy coded into the video bitstream to achieve even greater compression.

[0006] The coded video bitstream is then stored in a computer-readable storage medium (e.g., a flash memory) for access by another electronic device with digital video capabilities, or is directly transmitted to the electronic device in a wired or wireless manner. The electronic device then performs video decompression (which is the reverse process of the video compression described above) by, for example, parsing the coded video bitstream to obtain syntax elements from the bitstream and reconstructing digital video data from the coded video bitstream into its original format based at least in part on the syntax elements obtained from the bitstream, and renders the reconstructed digital video data on a display of the electronic device.

[0007] As digital video quality increases from HD to 4K×2K or even 8K×4K, the amount of video data to be encoded / decoded increases exponentially. There has always been a challenge in how to encode / decode video data more efficiently while maintaining the image quality of the decoded video data. Summary of the Invention

[0008] The embodiments described herein relate to video data encoding and decoding, and more particularly, to methods and systems for encoding and decoding high-level syntax in a video bitstream applicable to one or more video codec standards.

[0009] According to a first aspect of the present application, a method for decoding video data includes: receiving a plurality of syntax elements at a sequence parameter set (SPS) level from a bitstream, wherein the plurality of syntax elements are associated with predefined functions and are sequentially arranged in the bitstream; based on a determination that the plurality of syntax elements satisfy a predefined condition: receiving a second syntax element immediately following the plurality of syntax elements from the bitstream; based on a determination that the plurality of syntax elements do not satisfy the predefined condition: setting a default value to the second syntax element; and performing the predefined function on the video data from the bitstream based on the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of an intra-frame prediction function, an inter-frame prediction function, and a merge mode.

[0010] According to a second aspect of the present application, a method for decoding video data includes: receiving a plurality of syntax elements at one or more of a picture parameter set (PPS) level and a slice level from a bitstream, wherein the plurality of syntax elements are associated with predefined functions and are sequentially arranged in the bitstream; based on a determination that the plurality of syntax elements satisfy a predefined condition: receiving a second syntax element immediately following the plurality of syntax elements from the bitstream; based on a determination that the plurality of syntax elements do not satisfy the predefined condition: setting a default value to the second syntax element; and performing the predefined function on the video data from the bitstream based on the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of a quantization function, an intra-frame prediction function, and an inter-frame prediction function.

[0011] According to a third aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When the programs are executed by the one or more processing units, the electronic device performs the method for decoding video data as described above.

[0012] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units. When the programs are executed by the one or more processing units, the electronic device performs the method for decoding video data as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings, which are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification, illustrate the described embodiments and together with the description serve to explain the basic principles. Like reference numerals designate corresponding parts.

[0014] Figure 1is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure.

[0015] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0016] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0017] Figures 4A to 4E is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks of different sizes and shapes according to some embodiments of the present disclosure.

[0018] Figure 5 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure.

[0019] Figure 6 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure.

[0020] Figure 7 is a flow chart illustrating an exemplary process by which a video decoder implements techniques for decoding video data, according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth to facilitate understanding of the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives may be employed without departing from the scope of the claims, and that the subject matter may be practiced without these specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0022] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel according to some embodiments of the present disclosure. Figure 1As shown, system 10 includes a source device 12 that generates and encodes video data to be decoded at a later time by a destination device 14. Source device 12 and destination device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some implementations, source device 12 and destination device 14 are equipped with wireless communication capabilities.

[0023] In some embodiments, target device 14 may receive encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of moving encoded video data from source device 12 to target device 14. In one example, link 16 may include a communication medium for enabling source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated and transmitted to target device 14 according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that may be used to facilitate communication from source device 12 to target device 14.

[0024] In some other embodiments, the encoded video data can be transferred from the output interface 22 to a storage device 32. The encoded video data in the storage device 32 can then be accessed by the target device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile memory or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In a further example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The target device 14 can access the stored video data from the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing encoded video data and transmitting the encoded video data to the target device 14. Exemplary file servers include a web server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The target device 14 can access the encoded video data through any standard data connection, including a wireless channel suitable for accessing encoded video data stored on a file server (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both. The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination of both.

[0025] like Figure 1 As shown, source device 12 includes a video source 18, a video encoder 20, and an output interface 22. Video source 18 can include a source such as a video capture device, for example, a camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as source video, or a combination of these sources. As an example, if video source 18 is a camera of a security monitoring system, source device 12 and target device 14 can form a camera phone or a video phone. However, the embodiments described in this application can be generally applicable to video encoding and decoding and can be applied to wireless and / or wired applications.

[0026] Captured, pre-captured, or computer-generated video can be encoded by video encoder 20. The encoded video data can be transmitted directly to target device 14 via output interface 22 of source device 12. The encoded video data can also (or alternatively) be stored on storage device 32 for later access by target device 14 or other devices for decoding and / or playback. Output interface 22 may further include a modem and / or a transmitter.

[0027] Target device 14 includes input interface 28, video decoder 30, and display device 34. Input interface 28 may include a receiver and / or a modem and receives encoded video data via link 16. The encoded video data transmitted via link 16 or provided on storage device 32 may include various syntax elements generated by video encoder 20 for use by video decoder 30 in decoding the video data. Such syntax elements may be included in the encoded video data transmitted over a communication medium, stored on a storage medium, or stored in a file server.

[0028] In some implementations, the target device 14 may include a display device 34, which may be an integrated display device or an external display device configured to communicate with the target device 14. The display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0029] The video encoder 20 and the video decoder 30 may operate according to proprietary or industry standards such as VVC, HEVC, MPEG-4 Part 10, Advanced Video Codec (AVC), or extensions of such standards. It will be appreciated that the present application is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is generally contemplated that the video encoder 20 of the source device 12 may be configured to encode video data according to any of these current or future standards. Similarly, it is also generally contemplated that the video decoder 30 of the target device 14 may be configured to decode video data according to any of these current or future standards.

[0030] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video encoding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, any of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0031] Figure 2is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described herein. The video encoder 20 can perform intra-frame prediction and inter-frame prediction codecs on video blocks within a video frame. Intra-frame prediction codecs rely on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter-frame prediction codecs rely on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence.

[0032] like Figure 2 As shown, video encoder 20 includes video data memory 40, prediction processing unit 41, decoded picture buffer (DPB) 64, summer 50, transform processing unit 52, quantization unit 54, and entropy coding unit 56. Prediction processing unit 41 further includes motion estimation unit 42, motion compensation unit 44, partitioning unit 45, intra-prediction processing unit 46, and intra-block copy (BC) unit 48. In some embodiments, video encoder 20 also includes inverse quantization unit 58, inverse transform processing unit 60, and summer 62 for video block reconstruction. A deblocking filter (not shown) may be located between summer 62 and DPB 64 to filter block boundaries to remove blocking artifacts from the reconstructed video. In addition to the deblocking filter, a loop filter (not shown) may also be used to filter the output of summer 62. Video encoder 20 may take the form of fixed or programmable hardware units, or may be divided among one or more of the fixed or programmable hardware units shown.

[0033] The video data memory 40 can store video data to be encoded by the components of the video encoder 20. The video data in the video data memory 40 can be obtained, for example, from the video source 18. The DPB 64 is a buffer that stores reference video data for use by the video encoder 20 when encoding the video data (e.g., in intra-frame prediction codec mode or inter-frame prediction codec mode). The video data memory 40 and the DPB 64 can be formed by any of a variety of memory devices. In various examples, the video data memory 40 can be on-chip with the other components of the video encoder 20, or off-chip relative to those components.

[0034] like Figure 2As shown, after receiving the video data, the partition unit 45 within the prediction processing unit 41 partitions the video data into video blocks. The partitioning may also include partitioning the video frame into slices, tiles, or other larger coding units (CUs) according to a predefined partitioning structure (such as a quadtree structure associated with the video data). The video frame can be divided into multiple video blocks (or sets of video blocks called tiles). The prediction processing unit 41 can select one of multiple possible prediction codec modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of multiple intra-frame prediction codec modes or one of multiple inter-frame prediction codec modes. The prediction processing unit 41 can provide the resulting intra-frame prediction codec block or inter-frame prediction codec block to the summer 50 to generate a residual block, and to the summer 62 to reconstruct the code block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra-frame mode indicators, partition information, and other such syntax information to the entropy coding unit 56.

[0035] To select an appropriate intra-frame prediction codec mode for the current video block, intra-frame prediction processing unit 46 within prediction processing unit 41 may perform intra-frame prediction codec on the current video block relative to one or more neighboring blocks in the same frame as the current block to be coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-frame prediction codec on the current video block relative to one or more prediction blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may perform multiple encoding passes, for example, to select an appropriate coding mode for each block of video data.

[0036] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for the current video frame by generating motion vectors based on a predetermined pattern within a sequence of video frames. The motion vectors indicate the displacement of a prediction unit (PU) of a video block within the current video frame relative to a prediction block within a reference video frame. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimates the motion of a video block. A motion vector may, for example, indicate the displacement of a PU of a video block within the current video frame or picture relative to a prediction block (or other coded unit) within a reference frame, the prediction block being relative to the current block (or other coded unit) being encoded within the current frame. The predetermined pattern may designate video frames in the sequence as P-frames or B-frames. Intra BC unit 48 may determine vectors, such as block vectors, for use in intra BC encoding and decoding in a manner similar to the manner in which motion estimation unit 42 determines motion vectors for inter-prediction, or may utilize motion estimation unit 42 to determine block vectors.

[0037] A prediction block is a block of a reference frame that is considered to closely match the PU of the video block to be coded in terms of pixel difference, which can be determined by sum of absolute difference (SAD), sum of squared difference (SSD), or other difference metrics. In some embodiments, video encoder 20 can calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 can perform motion searches relative to full pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0038] Motion estimation unit 42 calculates a motion vector for a PU of a video block in an inter-prediction coded frame by comparing the position of the PU to the position of a prediction block of a reference frame selected from either the first reference frame list (List 0) or the second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy encoding unit 56.

[0039] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the PU of the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, retrieve the prediction block from DPB 64, and forward the prediction block to summer 50. Summer 50 then forms a residual video block having pixel difference values by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being coded. The pixel difference values forming the residual video block may include luma difference components, chroma difference components, or both. Motion compensation unit 44 may also generate syntax elements associated with the video block of the video frame for use by video decoder 30 when decoding the video block of the video frame. The syntax elements may include, for example, syntax elements defining a motion vector for identifying the prediction block, any flags indicating a prediction mode, or any other syntax information described herein. Note that motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are illustrated separately for conceptual purposes.

[0040] In some embodiments, the intra BC unit 48 may generate a vector and obtain a prediction block in a manner similar to that described above in conjunction with the motion estimation unit 42 and the motion compensation unit 44, but wherein the prediction block is in the same frame as the current block being encoded and wherein the vector is referred to as a block vector relative to the motion vector. Specifically, the intra BC unit 48 may determine an intra prediction mode to use for encoding the current block. In some examples, the intra BC unit 48 may encode the current block using various intra prediction modes, for example, during separate encoding passes, and test their performance using rate-distortion analysis. Next, the intra BC unit 48 may select an appropriate intra prediction mode to use from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values using the rate-distortion analysis for the various tested intra prediction modes and select the intra prediction mode with the best rate-distortion characteristics from the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis typically determines the amount of distortion (or error) between a coded block and the original, uncoded block (coded to produce the coded block) and the bit rate (i.e., the number of bits) used to produce the coded block. Intra BC unit 48 may calculate a ratio based on the distortion and rate for each coded block to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0041] In other examples, intra BC unit 48 may utilize, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, the prediction block may be a block that is considered to closely match the block to be coded in terms of pixel difference, which may be determined by sum of absolute difference (SAD), sum of squared difference (SSD), or other difference metrics, and identification of the prediction block may include calculating values for sub-integer pixel positions.

[0042] Regardless of whether the prediction block is from the same frame according to intra-frame prediction or from different frames according to inter-frame prediction, video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded and decoded to form pixel difference values. The pixel difference values forming the residual video block may include luma component difference values and chroma component difference values.

[0043] As described above, the intra-prediction processing unit 46 may perform intra-prediction on the current video block as an alternative to the inter-prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra-block copy prediction performed by the intra BC unit 48. In particular, the intra-prediction processing unit 46 may determine an intra-prediction mode to use for encoding the current block. To this end, the intra-prediction processing unit 46 may encode the current block using various intra-prediction modes, for example, during separate encoding passes, and the intra-prediction processing unit 46 (or, in some examples, a mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. The intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra-prediction mode in the bitstream.

[0044] After prediction processing unit 41 determines a prediction block for the current video block via inter-prediction or intra-prediction, summer 50 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to transform processing unit 52. Transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0045] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan on the matrix comprising the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.

[0046] After quantization, entropy coding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioned entropy (PIPE) coding, or other entropy coding methods or techniques. The encoded bitstream may then be transmitted to video decoder 30 or archived in storage device 32 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also entropy encode motion vectors and other syntax elements for the current video frame being encoded.

[0047] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual video block in the pixel domain to generate a reference block used to predict other video blocks. As described above, motion compensation unit 44 may generate a motion compensated prediction block from one or more reference blocks of a frame stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0048] Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in DPB 64. The reference block may then be used as a prediction block by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 to inter-predict another video block in a subsequent video frame.

[0049] Figure 3 is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra-frame prediction processing unit 84, and an intra-frame BC unit 85. The video decoder 30 may perform operations generally in conjunction with the above. Figure 2 The decoding process is the reverse of the encoding process described with respect to video encoder 20. For example, motion compensation unit 82 may generate prediction data based on motion vectors received from entropy decoding unit 80, and intra-prediction unit 84 may generate prediction data based on intra-prediction mode indicators received from entropy decoding unit 80.

[0050] In some examples, units of the video decoder 30 may be assigned to perform embodiments of the present application. Likewise, in some examples, embodiments of the present disclosure may be divided between one or more units of the video decoder 30. For example, the intra BC unit 85 may perform embodiments of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.

[0051] The video data memory 79 can store video data to be decoded by other components of the video decoder 30, such as an encoded video bitstream. For example, the video data stored in the video data memory 79 can be obtained from the storage device 32, a local video source (such as a camera) via a wired or wireless network transmission of the video data or by accessing a physical data storage medium (such as a flash drive or hard disk). The video data memory 79 may include a coded picture buffer (CPB) that stores encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use by the video decoder 30 when decoding the video data (for example, in an intra-frame prediction codec mode or an inter-frame prediction codec mode). The video data memory 79 and the DPB 92 can be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are shown in FIG. Figure 3 92 as two distinct components of video decoder 30. However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or separate memory devices. In some examples, video data memory 79 may be on-chip with other components of video decoder 30, or off-chip relative to those components.

[0052] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 may receive syntax elements at the video frame level and / or the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors and other syntax elements to the prediction processing unit 81.

[0053] When a video frame is encoded and decoded as an intra-frame prediction codec (I) frame or an intra-frame codec prediction block in other types of frames, the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video frame based on the intra-frame prediction mode transmitted by the signal and the reference data from the previously decoded block of the current frame.

[0054] When the video frame is encoded as an inter-frame prediction codec (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks for the video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each prediction block can be generated from a reference frame in one of the reference frame lists. The video decoder 30 can construct the reference frame lists: List 0 and List 1 using a default construction technique based on the reference frames stored in the DPB 92.

[0055] In some examples, when a video block is encoded or decoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block defined by video encoder 20.

[0056] The motion compensation unit 82 and / or the intra BC unit 85 determine prediction information for a video block of the current video frame by parsing the motion vector and other syntax elements, and then use the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine a prediction mode (e.g., intra prediction or inter prediction) for encoding or decoding the video block of the video frame, an inter prediction frame type (e.g., B or P), construction information for one or more reference frame lists in the reference frame list of the frame, a motion vector for each inter-frame prediction-encoded video block of the frame, an inter-frame prediction state for each inter-frame prediction-encoded video block of the frame, and other information for decoding the video block in the current video frame.

[0057] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., flags) to determine that the current video block is predicted using the intra BC mode, construction information that the video blocks of the frame are within the reconstructed region and should be stored in the DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction status for each intra BC predicted video block of the frame, and other information for decoding the video blocks in the current video frame.

[0058] Motion compensation unit 82 may also perform interpolation using interpolation filters to calculate interpolated values for sub-integer pixels of a reference block, as used by video encoder 20 during encoding of the video block. In this case, motion compensation unit 82 may determine the interpolation filters used by video encoder 20 from received syntax elements and use the interpolation filters to produce the prediction block.

[0059] Inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by entropy decoding unit 80, using the same quantization parameters that were calculated by video encoder 20 for each video block in the video frame to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process) to the transform coefficients to reconstruct the residual block in the pixel domain.

[0060] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vectors and other syntax elements, the summer 90 reconstructs the decoded video block of the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter (not shown) can be positioned between the summer 90 and the DPB 92 to further process the decoded video block. The decoded video block in a given frame is then stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video block. The DPB 92 or a memory device separate from the DPB 92 can also store the decoded video for later presentation on a video device such as a video frame. Figure 1 On display devices such as display device 34.

[0061] In a typical video encoding and decoding process, a video sequence typically comprises an ordered set of frames or pictures. Each frame may comprise three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In other instances, a frame may be monochrome and therefore comprise only a two-dimensional array of luma samples.

[0062] like Figure 4A As shown, the video encoder 20 (or more specifically, the partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs sequentially ordered from left to right and from top to bottom in a raster scan order. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, i.e., one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a particular size. Figure 4BAs shown, each CTU may include one coding tree block (CTB) of luma samples, two corresponding coding tree blocks of chroma samples, and syntax elements for encoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coding blocks of pixels and how the video sequence can be reconstructed at the video decoder 30, and the syntax elements include inter-frame prediction or intra-frame prediction, intra-frame prediction mode, motion vector and other parameters. In a monochrome picture or a picture with three separate color planes, a CTU may include a single coding tree block and syntax elements for encoding and decoding the samples of the coding tree block. The coding tree block may be an N×N sample block.

[0063] To achieve better performance, the video encoder 20 may recursively perform tree partitioning (such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination thereof) on the coding tree blocks of the CTU and partition the CTU into smaller coding units (CUs). Figure 4C As shown, a 64×64 CTU 400 is first divided into four smaller CUs, each with a block size of 32×32. Of the four smaller CUs, CU 410 and CU 420 are each divided into four 16×16 CUs by block size. The two 16×16 CUs 430 and 440 are each further divided into four 8×8 CUs by block size. Figure 4D Depicted diagram Figure 4C The quadtree data structure is the final result of the partitioning process of the CTU 400 depicted in FIG. 4 , where each leaf node of the quadtree corresponds to a CU with a corresponding size ranging from 32×32 to 8×8. Figure 4B Each CU may include a coding block (CB) of luma samples and two corresponding coding blocks of chroma samples of the same size frame, as well as syntax elements for encoding and decoding the samples of the coding blocks. In a monochrome picture or a picture with three separate color planes, a CU may include a single coding block and syntax structures for encoding and decoding the samples of the coding block. It should be noted that Figure 4C and Figure 4D The quadtree partitioning depicted in FIG is for illustration purposes only, and a CTU can be split into multiple CUs to accommodate different local characteristics based on quadtree / ternary tree / binary tree partitioning. In the multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned by a binary tree structure or a ternary tree structure. Figure 4E As shown, there are five types of partitions, namely, quadruple partition, horizontal binary partition, vertical binary partition, horizontal triple partition, and vertical triple partition.

[0064] In some embodiments, the video encoder 20 may further partition the coding block of the CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction (inter or intra) is applied. The prediction unit (PU) of a CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements for predicting the prediction blocks. In a monochrome picture or a picture with three separate color planes, a PU may include a single prediction block and a syntax structure for predicting the prediction block. The video encoder 20 may generate predicted luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.

[0065] Video encoder 20 may use intra prediction or inter prediction to generate the prediction block for a PU. If video encoder 20 uses intra prediction to generate the prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of the frame associated with the PU. If video encoder 20 uses inter prediction to generate the prediction block for a PU, video encoder 20 may generate the prediction block for the PU based on decoded samples of one or more frames other than the frame associated with the PU.

[0066] After the video encoder 20 generates the predicted luma, Cb, and Cr blocks for one or more PUs of a CU, the video encoder 20 may generate a luma residual block for the CU by subtracting the predicted luma block of the CU from its original luma coding block, such that each sample in the luma residual block of the CU indicates the difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, the video encoder 20 may generate a Cb residual block and a Cr residual block for the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.

[0067] In addition, if Figure 4CAs illustrated, the video encoder 20 may use quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. A transform unit (TU) of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some examples, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0068] Video encoder 20 may apply one or more transforms to the luma transform block of a TU to generate a luma coefficient block for the TU. A coefficient block may be a two-dimensional array of transform coefficients. A transform coefficient may be a scalar. Video encoder 20 may apply one or more transforms to the Cb transform block of a TU to generate a Cb coefficient block for the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of a TU to generate a Cr coefficient block for the TU.

[0069] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients, thereby providing further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 may entropy encode syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Ultimately, the video encoder 20 may output a bitstream comprising a sequence of bits forming a representation of a coded frame and associated data, which may be stored in the storage device 32 or transmitted to the target device 14.

[0070] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the opposite of the encoding process performed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient blocks associated with the TUs of the current CU to reconstruct the residual blocks associated with the TUs of the current CU. The video decoder 30 also reconstructs the coding blocks of the current CU by adding samples of the prediction blocks of the PUs of the current CU to corresponding samples of the transform blocks of the TUs of the current CU. After reconstructing the coding blocks of each CU of the frame, the video decoder 30 can reconstruct the frame.

[0071] In general, the basic intra prediction scheme used in VVC remains the same as that of HEVC, except that several modules have been further extended and / or improved, such as the matrix-weighted intra prediction (MIP) codec mode, the intra sub-partitioning (ISP) codec mode, extended intra prediction using wide-angle intra direction, position-dependent intra prediction combination (PDPC), and 4-tap intra interpolation. The main focus of this disclosure is to improve the existing high-level syntax design in the VVC standard. The relevant background knowledge is elaborated in detail in the following sections.

[0072] Like HEVC, VVC uses a bitstream structure based on Network Abstraction Layer (NAL) units. The codec bitstream is partitioned into multiple NAL units that should be smaller than the maximum transport unit size when transmitted over a lossy packet network. Each NAL unit consists of a NAL unit header followed by a NAL unit payload. There are two conceptual categories of NAL units. Video Codec Layer (VCL) NAL units that contain codec sample data, such as codec slice NAL units, whereas non-VCL NAL units contain metadata that typically belongs to more than one codec picture, or in cases where association with a single codec picture would be meaningless, such as parameter set NAL units, or in cases where the information is not needed for the decoding process, such as SEINAL units.

[0073] In VVC, a two-byte NAL unit header is introduced, and it is expected that this design will be sufficient to support future extensions. The syntax and associated semantics of the NAL unit header in the current VVC draft specification are illustrated in Table 1 and Table 2, respectively. How to read Table 1 is illustrated in the appendix of this invention, which can also be found in the VVC specification.

[0074] Table 1. NAL unit header syntax

[0075]

[0076] Table 2. NAL unit header semantics

[0077]

[0078] Table 3. NAL unit type codes and NAL unit type categories

[0079]

[0080] VVC inherits the parameter set concept of HEVC with some modifications and additions. Parameter sets can be part of the video bitstream or can be received by the decoder through other means (including out-of-band transmission using a reliable channel, hard encoding and decoding in the encoder and decoder, etc.). Parameter sets contain identifiers referenced directly or indirectly from the slice header, as discussed in more detail later. The reference process is called "activation". Depending on the parameter set type, activation occurs per picture or per sequence. Among other reasons, the concept of activation by reference is introduced because implicit activation cannot be performed by virtue of the location of information in the bitstream (such as other syntax elements common to the video codec) in the case of out-of-band transmission.

[0081] Video Parameter Sets (VPS) were introduced to convey information applicable to multiple layers and sub-layers. VPS was introduced to address these shortcomings and enable a clean and scalable high-level design of multi-layer codecs. Regardless of whether each layer of a given video sequence has the same or different Sequence Parameter Sets (SPSs), they all reference the same VPS. The syntax and associated semantics of the video parameter sets in the current VVC draft specification are illustrated in Tables 4 and 5, respectively. How to read Table 4 is illustrated in the Appendix of this disclosure, which can also be found in the VVC specification.

[0082] Table 4. Video parameter set RBSP syntax

[0083]

[0084]

[0085]

[0086]

[0087] Table 5. Video parameter set RBSP semantics

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097] In VVC, the SPS contains information applicable to all slices of a coded video sequence. A coded video sequence starts with an instantaneous decoding refresh (IDR) picture or a BLA picture, or a CRA picture that is the first picture in the bitstream, and includes all subsequent pictures that are not IDR or BLA pictures. The bitstream consists of one or more coded video sequences. The content of the SPS can be roughly divided into six categories: 1) self-reference (its own ID); 2) decoder operating point related information (profile, level, picture size, number of sub-layers, etc.); 3) enablement flags for certain tools within the profile, and associated codec tool parameters when the tools are enabled; 4) information that defines the flexibility of the structure and transform coefficient coding; 5) temporal scalability control; and 6) visual usability information (VUI) including HRD information. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are illustrated in Tables 6 and 7, respectively. How to read Table 6 is illustrated in the appendix of this invention, which can also be found in the VVC specification.

[0098] Table 6. Sequence parameter set RBSP syntax

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] Table 7. Sequence parameter set RBSP semantics

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126] VVC's picture parameter set (PPS) contains such information that can change between pictures. The PPS includes information roughly equivalent to a portion of the PPS in HEVC, including: 1) self-reference; 2) initial picture control information, such as the initial quantization parameter (QP), the number of flags indicating the use or presence of certain tools, or control information in the slice header; and 3) tile information. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are illustrated in Table 8 and Table 9, respectively. How to read Table 8 is illustrated in the appendix of this invention, which can also be found in the VVC specification.

[0127] Table 8. Picture parameter set RBSP syntax

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] Table 9. Picture parameter set RBSP semantics

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145] The slice header contains information that can vary between slices, as well as picture-related information that is relatively small or only relevant for a specific slice or picture type. The size of the slice header can be significantly larger than the PPS, especially when there are tile or wavefront entry point offsets in the slice header and RPS, prediction weights or reference picture list modifications are explicitly signaled. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are illustrated in Table 10 and Table 11, respectively. How to read Table 10 is illustrated in the appendix of this invention, which can also be found in the VVC specification.

[0146] Table 10. Picture header structure syntax

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153] Table 11. Image header structure semantics

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166] In the current VVC, when similar syntax elements exist for intra prediction and inter prediction, syntax elements related to inter prediction are defined before syntax elements related to intra prediction in some places. Considering the fact that intra prediction is allowed in all picture / slice types but inter prediction is not allowed, such an order may not be preferable. From a standardization perspective, it would be beneficial to always define syntax related to intra prediction before syntax for inter prediction.

[0167] It is also observed that in the current VVC, some syntax elements that are highly related to each other are defined in a sprawling manner in different places. From a standardization perspective, it would also be beneficial to group some syntax together.

[0168] In the present disclosure, in order to solve the problems pointed out in the "Problem Statement" section, methods are provided to simplify and / or further improve the existing designs of high-level grammars. It should be noted that the invented methods can be applied independently or in combination.

[0169] In the present disclosure, a rearrangement of syntax elements is designed so that intra-prediction related syntax elements are defined before inter-prediction related syntax elements. According to the present disclosure, partition constraint syntax elements are grouped by prediction type, with intra-prediction related syntax elements first and inter-prediction related syntax elements later. In one embodiment, the order of partition constraint syntax elements in the SPS is consistent with the order of partition constraint syntax elements in the picture header. An example of the decoding process for the VVC draft is illustrated in Table 12 below. Changes to the VVC draft are shown in italic font.

[0170] Table 12. Designed Sequence Parameter Set RBSP Syntax

[0171]

[0172] Figure 5 5 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure. For example, the method may be applied to a decoder.

[0173] The decoder may receive the permutation partition constraint syntax elements at the SPS level in step 510. These permutation partition constraint syntax elements are arranged such that intra prediction related syntax elements are defined before inter prediction related syntax elements.

[0174] In step 512, the decoder may obtain a first reference picture I associated with a video block in the bitstream. (0) and the second reference picture I (1) In the display order, the first reference picture I (0) Before the current picture and the second reference picture I (1) After the current image.

[0175] In step 514, the decoder can obtain the first reference picture I (0) The reference block in (i, j) obtains the first prediction sample I of the video block (0) . i and j represent the coordinates of a sample point in the current image.

[0176] In step 516, the decoder can obtain the second reference picture I (1) The reference block in (i, j) obtains the second prediction sample I of the video block (1) .

[0177] In step 518, the decoder may arrange the partition constraint syntax element, the first prediction sample I (0) (i, j) and the second predicted sample point I (1) (i, j) to obtain bidirectional prediction samples.

[0178] Figure 66 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure. For example, the method can be applied to a decoder. In step 610, the decoder can receive a bitstream including a VPS, SPS, PPS, picture header, and slice header for encoded and decoded video data. In step 612, the decoder can decode the VPS. In step 614, the decoder can decode the SPS and obtain an SPS-level permutation partition constraint syntax element. In step 616, the decoder can decode the PPS. In step 618, the decoder can decode the picture header. In step 620, the decoder can decode the slice header. In step 622, the decoder can decode the video data based on the VPS, SPS, PPS, picture header, and slice header.

[0179] In this disclosure, a design is proposed to group syntax elements related to dual-tree chroma types. In one embodiment, the partition constraint syntax elements for dual-tree chroma in the SPS should be signaled together in the dual-tree chroma case. An example of the decoding process for the VVC draft is illustrated in Table 13 below. Changes to the VVC draft are shown in italic font.

[0180] Table 13. Designed Sequence Parameter Set RBSP Syntax

[0181]

[0182] If it is also considered that the intra prediction related syntax is defined before the inter prediction related syntax, according to the method of the present disclosure, another example of the decoding process with respect to the VVC draft is illustrated in the following Table 14. Changes to the VVC draft are shown using italic font.

[0183] Table 14. Designed Sequence Parameter Set RBSP Syntax

[0184]

[0185] As mentioned earlier, according to the current VVC, intra prediction is allowed in all picture / slice types, but inter prediction is not allowed. According to the present disclosure, a flag is added to the VVC syntax at a specific codec level to indicate whether inter prediction is used in a sequence, picture, and / or slice. If inter prediction is not used, inter prediction-related syntax is not signaled at the corresponding codec level (e.g., sequence, picture, and / or slice level).

[0186] In one example, according to the method of the present disclosure, a flag is added to the SPS to indicate whether inter-frame prediction is used in encoding and decoding the current video sequence. If not used, inter-frame prediction related syntax elements are not signaled in the SPS. An example of the decoding process for the VVC draft is illustrated in Table 15 below. Changes to the VVC draft are shown in italic font.

[0187] Table 15. Designed Sequence Parameter Set RBSP Syntax

[0188]

[0189]

[0190] Sequence parameter set RBSP semantics

[0191] sps_inter_slice_used_flag equal to 0 specifies that all coded slices of the video sequence have slice_type equal to 2. sps_inter_slice_used_flag equal to 1 specifies that one or more coded slices with slice_type equal to 0 or 1 may or may not be present in the video sequence.

[0192] In some embodiments, syntax elements are arranged so that syntax elements with similar functions (e.g., intra tools, inter tools, screen content tools, transform tools, quantization tools, loop filter tools, and / or partition tools) are grouped at certain codec levels (e.g., sequence, picture, and / or slice levels) according to VVC syntax. According to some embodiments, syntax elements in a sequence parameter set (SPS) are arranged so that syntax elements with similar functions are grouped. An example of a decoding process for VVC is illustrated in Table 16 below. Table 16 shows a syntax for grouping syntax elements with similar functions at the SPS level.

[0193] Table 16. Syntax for grouping syntax elements with similar functions at the SPS level.

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202] Another example of a decoding process for VVC is illustrated in the following Table 17. Table 17 shows a syntax for grouping syntax elements having similar functions at the SPS level.

[0203] Table 17. Syntax for grouping syntax elements with similar functions at the SPS level

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212] According to some embodiments, the syntax elements in the picture parameter set (PPS) are arranged so that similar functional related syntax elements are grouped. An example of a decoding process for VVC is illustrated in Table 18 below. Table 18 shows the syntax for grouping syntax elements with similar functions at the PPS level.

[0213] Table 18. Syntax for grouping syntax elements with similar functions at the PPS level

[0214]

[0215]

[0216]

[0217]

[0218] Another example of a decoding process for VVC is illustrated in the following Table 19. Table 19 shows a syntax for grouping syntax elements having similar functions at a PPS level.

[0219] Table 19. Syntax for grouping syntax elements with similar functions at the PPS level

[0220]

[0221]

[0222]

[0223]

[0224]

[0225] Figure 7 is a flowchart 700 illustrating an exemplary process by which a video decoder (eg, video decoder 30 ) implements techniques for decoding video data, according to some embodiments of the present disclosure.

[0226] like Figure 7 As shown, in some embodiments, video decoder 30 receives a plurality of syntax elements at a sequence parameter set (SPS) level from a bitstream, and the plurality of syntax elements are associated with predefined functions and are sequentially arranged in the bitstream (710).

[0227] Based on a determination that the plurality of syntax elements satisfy the predefined condition, video decoder 30 receives a second syntax element from the bitstream that immediately follows the plurality of syntax elements ( 720 ).

[0228] Based on a determination that the plurality of syntax elements do not satisfy the predefined condition, video decoder 30 sets a default value to the second syntax element ( 730 ).

[0229] Video decoder 30 performs a predefined function on the video data from the bitstream according to the plurality of syntax elements and the second syntax element, and the predefined function is a predefined function selected from the group consisting of an intra prediction function, an inter prediction function, and a merge mode (740).

[0230] In some embodiments, the plurality of syntax elements associated with the intra prediction function and sequentially arranged in the bitstream include at least:

[0231] sps_isp_enabled_flag,

[0232] sps_mrl_enabled_flag,

[0233] sps_mip_enabled_flag,

[0234] sps_palette_enabled_flag and

[0235] sps_ibc_enabled_flag,

[0236] And in some embodiments, the sequentially arranged plurality of syntax elements associated with the intra prediction function do not include sps_bcw_enabled_flag.

[0237] In some embodiments, video decoder 30 receives a set of syntax elements associated with inter-prediction functionality from the bitstream after receiving a plurality of syntax elements associated with intra-prediction functionality.

[0238] In some embodiments, the plurality of syntax elements associated with the inter-frame prediction function and sequentially arranged in the bitstream include at least:

[0239] sps_weighted_pred_flag,

[0240] sps_weighted_brpred_flag,

[0241] log_term_ref_pics_flag,

[0242] inter_layer_ref_pics_present_flag、

[0243] sps_idr_rpl_present_flag,

[0244] rpl1_same_as_rpl0_flag,

[0245] sps_log2_diff_min_qt_min_cb_inter_slice,

[0246] sps_max_mtt_hierarchy_depth_inter_slice,

[0247] sps_ref_wraparound_enabled_flag,

[0248] sps_temporal_mvp_enabled_flag、

[0249] sps_amvr_enabled_flag,

[0250] sps_bdof_enabled_falg,

[0251] sps_ssmvd_enabled_flag,

[0252] sps_dmvr_enabled_flag,

[0253] sps_dmvr_enabled_flag,

[0254] six_minus_max_num+merge_cand、

[0255] sps_sbt_enabled_flag,

[0256] sps_affine_enabled_flag,

[0257] sps_bcw_enabled_flag,

[0258] sps_ciip_enabled_flag and

[0259] log2_parallel_merge_level_minus2 In some embodiments, a plurality of syntax elements associated with the merge mode and sequentially arranged in the bitstream include at least:

[0260] sps_mmvd_enabled_flag

[0261] six_minus_max_num_merge_cand、

[0262] sps_sbt_enabled_flag,

[0263] sps_affine_enabled_flag,

[0264] sps_bcw_enabled_flag,

[0265] sps_ciip_enabled_flag and

[0266] log2_parallel_merge_level_minus2.

[0267] Several syntax elements associated with merge mode are also associated with inter prediction functionality.

[0268] In some embodiments, video decoder 30 receives a set of syntax elements associated with a quantization function from the bitstream prior to receiving a plurality of syntax elements associated with an intra prediction function.

[0269] In some embodiments, the video decoder 30 receives a plurality of syntax elements at one or more of a picture parameter set (PPS) level and a slice level from a bitstream, wherein the plurality of syntax elements are associated with predefined functions and are sequentially arranged in the bitstream; based on a determination that the plurality of syntax elements satisfy a predefined condition, the video decoder 30: receives a second syntax element immediately following the plurality of syntax elements from the bitstream; based on a determination that the plurality of syntax elements do not satisfy the predefined condition, the video decoder 30: sets a default value to the second syntax element; and the video decoder 30 performs the predefined function on the video data from the bitstream based on the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of a quantization function, an intra-frame prediction function, and an inter-frame prediction function.

[0270] In some embodiments, before video decoder 30 receives a set of syntax elements associated with an inter-prediction function, video decoder 30 further receives a set of syntax elements associated with a quantization function.

[0271] In some embodiments, video decoder 30 further receives a set of syntax elements associated with a quantization function after receiving a set of syntax elements associated with an inter-prediction function.

[0272] In some embodiments, the plurality of syntax elements at the PPS level associated with the inter-frame prediction function and sequentially arranged in the bitstream include at least:

[0273] rpl1_idx_present_flag,

[0274] pps_weighted_pred_flag,

[0275] pps_weighted_bipred_flag,

[0276] rpl_info_in_ph_flag and

[0277] pps_ref_wraparound_enabled_flag.

[0278] In some embodiments, the plurality of syntax elements at the PPS level include: pps_weighted_ored_flag, pps_weighted_bipred_flag, and rpl_info_in_ph_flag; the predefined condition is: (pps_weighted_pred_flag is true or pps_weighted_bipred_flag is true) and rpl_info_in_ph_flag is true; and the second syntax is wp_info_in_ph_flag.

[0279] The above method can be implemented using an apparatus including one or more circuits, including an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. The apparatus can use these circuits in combination with other hardware or software components for performing the above method. Each module, submodule, unit, or subunit disclosed above can be implemented at least in part using one or more circuits.

[0280] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media corresponding to tangible media such as data storage media or communication media including any media that facilitates, for example, transferring a computer program from one place to another according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) non-transitory tangible computer-readable storage media or (2) communication media such as signals or carrier waves. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, codes, and / or data structures for implementing the embodiments described in this application. A computer program product may include computer-readable media.

[0281] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "and / or" as used herein refer to and encompass any and all possible combinations of one or more of the associated enumerated items. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of stated features, elements, and / or parts, but do not exclude the presence or addition of one or more other features, elements, parts, and / or groups thereof.

[0282] It should also be understood that although the terms first, second, etc. may be used to describe various elements in this article, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, without departing from the scope of the embodiment, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but the first electrode and the second electrode are not the same electrode.

[0283] The description of the present application has been presented for purposes of illustration and description, and the description is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings. The embodiments are chosen and described in order to best explain the principles of the invention, the practical application, and to enable others skilled in the art to understand the various embodiments of the invention and to best utilize the basic principles as well as the various embodiments with various modifications suitable for the particular use contemplated. Therefore, it should be understood that the scope of the claims should not be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.

Claims

1. A method for encoding video data, comprising: encoding a plurality of syntax elements at one or more of a picture parameter set (PPS) level and a slice level into a bitstream, wherein the plurality of syntax elements are associated with predefined functions and are sequentially arranged in the bitstream; According to determining that at least one syntax element among the plurality of syntax elements satisfies a predefined condition: encoding a second syntax element following the plurality of syntax elements into the bitstream; According to determining that the at least one syntax element among the plurality of syntax elements does not satisfy the predefined condition: Setting the value of the second syntax element to a default value; and performing the predefined function on video data from the bitstream according to at least one syntax element of the plurality of syntax elements and the second syntax element, wherein the predefined function is a quantization function, an intra-frame prediction function, or an inter-frame prediction function; The method further comprises: After encoding a set of syntax elements associated with the inter-frame prediction function at the PPS level into the bitstream, encoding a set of syntax elements associated with the quantization function at the PPS level into the bitstream, the set of syntax elements associated with the inter-frame prediction function at the PPS level in the bitstream comprising at least: rpl1_idx_present_flag, pps_weighted_pred_flag, pps_weighted_bipred_flag and pps_ref_wraparound_enabled_flag, wherein the rpl1_idx_present_flag specifies whether rpl_sps_flag[1] and rpl_idx[1] are present in the PH syntax structure or the slice header of the picture referencing the PPS, rpl_sps_flag[i] is equal to 1 specifies that the RPLi in ref_pic_lists() is derived based on one of the syntax structures ref_pic_list_struct(listIdx, rplsIdx), where listIdx is equal to i in the SPS, and rpl_sps_flag[i] is equal to 0 specifies that the RPLi in the picture is derived based on the syntax structure ref_pic_list_struct(listIdx, rplsIdx) included in the ref_pic_lists(), where listIdx is equal to i, and the ref_pic_lists( ) appears in the PH syntax structure or slice header, the rpl_idx[i] specifies the index of the syntax structure ref_pic_list_struct (listIdx, rplsIdx) with listIdx equal to i in the list of the syntax structure ref_pic_list_struct (listIdx, rplsIdx) with listIdx equal to i contained in the SPS, which is used to derive the RPL1 of the current picture or slice, the pps_weighted_pred_flag specifies whether weighted prediction is applied to the P slice referencing the PPS, the pps_weighted_bipred_flag specifies whether explicit weighted prediction is applied to the B slice referencing the PPS, and the pps_ref_wraparound_enabled_flag specifies whether horizontal wraparound motion compensation is applied in inter prediction.

2. The method according to claim 1, wherein The plurality of syntax elements at the PPS level associated with the inter-frame prediction function and sequentially arranged in the bitstream include at least: rpl1_idx_present_flag, pps_weighted_pred_flag, pps_weighted_bipred_flag, rpl_info_in_ph_flag and pps_ref_wraparound_enabled_flag, Among them, the rpl_info_in_ph_flag is equal to 1, which specifies that the reference picture list information exists in the PH syntax structure and does not exist in the slice header of the referenced PPS that does not contain the PH syntax structure; the rpl_info_in_ph_flag is equal to 0, which specifies that the reference picture list information does not exist in the PH syntax structure and can exist in the slice header of the referenced PPS that does not contain the PH syntax structure.

3. The method according to claim 1, wherein: The plurality of syntax elements at the PPS level include: pps_weighted_pred_flag, pps_weighted_bipred_flag, and rpl_info_in_ph_flag; The predefined conditions are: pps_weighted_pred_flag is true or pps_weighted_bipred_flag is true, and rpl_info_in_ph_flag is true; and The second syntax element is wp_info_in_ph_flag, Among them, the rpl_info_in_ph_flag is equal to 1, specifying that the reference picture list information exists in the PH syntax structure and does not exist in the slice header of the referenced PPS that does not contain the PH syntax structure; the rpl_info_in_ph_flag is equal to 0, specifying that the reference picture list information does not exist in the PH syntax structure and can exist in the slice header of the referenced PPS that does not contain the PH syntax structure; the wp_info_in_ph_flag is equal to 1, specifying that the weighted prediction information can exist in the PH syntax structure and does not exist in the slice header of the referenced PPS that does not contain the PH syntax structure; the wp_info_in_ph_flag is equal to 0, specifying that the weighted prediction information does not exist in the PH syntax structure and can exist in the slice header of the referenced PPS that does not contain the PH syntax structure.

4. An electronic device comprising: one or more processing units; a memory coupled to the one or more processing units; as well as A plurality of programs stored in the memory, which, when executed by the one or more processing units, cause the electronic device to perform the method according to any one of claims 1 to 3.

5. A non-transitory computer-readable storage medium storing a plurality of programs for execution by an electronic device having one or more processing units, wherein: The plurality of programs, when executed by the one or more processing units, causes the electronic device to perform the method according to any one of claims 1 to 3.

6. A computer program product comprising instructions for execution by a computing device having one or more processors, wherein when the instructions are executed by the one or more processors, the computing device performs the method according to any one of claims 1 to 3.

7. A method for storing a bitstream, comprising: Execute the method according to any one of claims 1 to 3 to generate a bit stream; The bitstream is stored.

8. A method for transmitting a bit stream, comprising: Execute the method according to any one of claims 1 to 3 to generate a bit stream; The bit stream is transmitted.