Methods and apparatus for advanced syntax in video coding and decoding

By processing syntax elements in the video encoding and decoding standard in the video bitstream and performing predefined video encoding/decoding functions, the challenge of efficient encoding and decoding is solved, and the encoding/decoding efficiency is achieved while maintaining image quality.

CN115398919BActive Publication Date: 2025-06-17BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180027271.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-10
Filing Date
2021-04-12
Publication Date
2025-06-17
Estimated Expiration
2041-04-12

AI Technical Summary

Technical Problem

With the improvement of digital video quality, the amount of video data to be encoded/decoded has increased exponentially. How to encode/decode more efficiently while maintaining image quality has become a challenge.

Method used

In a video bitstream suitable for one or more video codec standards, predefined conditions are determined and corresponding predefined functions are performed, such as groups composed of free intra prediction, inter prediction and merge mode, by receiving and processing multiple syntax elements at the Sequence Parameter Set (SPS) level and picture parameter Set (PPS) level.

Benefits of technology

It realizes video encoding/decoding more efficiently while maintaining the quality of video images, and is suitable for various video encoding and decoding standards, improving the efficiency of encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115398919B_ABST
    Figure CN115398919B_ABST
Patent Text Reader

Abstract

A method for an electronic device to perform decoding of video data. The method includes: receiving, from a bitstream, a plurality of syntax elements at one or more levels of a sequence parameter set (SPS), a picture parameter set (PPS) level, and a slice level, wherein the plurality of syntax elements are associated with a predefined function and are sequentially arranged in the bitstream; based on a determination that the plurality of syntax elements satisfy a predefined condition: receiving, from the bitstream, a second syntax element immediately following the plurality of syntax elements; based on a determination that the plurality of syntax elements do not satisfy the predefined condition: setting a default value as the second syntax element; and performing the predefined function on video data from the bitstream based on the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of an intra prediction function, an inter prediction function, and a merge mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority of U.S. Provisional Patent Application No. 63 / 008,648, entitled "HIGH - LEVEL SYNTAX FOR VIDEO CODING [Advanced Syntax for Video Coding and Decoding]", filed on April 10, 2020, which is incorporated herein by reference in its entirety. Technical field

[0003] This application generally relates to video data coding, decoding, and compression, and particularly to methods and systems for video coding and decoding of advanced syntax in a video bitstream applicable to one or more video coding standards. Background art

[0004] Various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smart phones, video teleconferencing devices, video streaming devices, etc. support digital video. Electronic devices send, receive, encode, decode, and / or store digital video data by implementing video compression / decompression standards defined by, for example, MPEG - 4, ITU - T H.263, ITU - T H.264 / MPEG - 4 Part 10, Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC) standards. Video compression typically includes performing spatial (intra - frame) prediction and / or temporal (inter - frame) prediction to reduce or remove the redundancy inherent in video data. For block - based video coding, a video frame is partitioned into one or more strips, each strip having multiple video blocks, which may also be referred to as coding tree units (CTUs). Each CTU may contain one coding unit (CU) or be recursively divided into smaller CUs until a predefined minimum CU size is reached. Each CU (also called a leaf CU) contains one or more transform units (TUs), and each CU also contains one or more prediction units (PUs). Each CU can be coded and decoded in intra - frame, inter - frame, or IBC mode. Intra - frame coding (I) strips of video blocks in a video frame are coded using spatial prediction relative to reference samples in adjacent blocks within the same video frame. Video blocks in inter - frame coding (P or B) strips of a video frame can be coded using spatial prediction relative to reference samples in adjacent blocks within the same video frame or using temporal prediction relative to reference samples in other previous and / or future reference video frames.

[0005] Spatial or temporal prediction based on previously encoded reference blocks (e.g., neighboring blocks) generates a prediction block for a current video block to be coded / decoded. The process of finding the reference blocks can be done by a block matching algorithm. Residual data representing the pixel difference between the current block to be coded / decoded and the prediction block is referred to as a residual block or prediction error. An inter-coded block is coded based on a motion vector pointing to a reference block in a reference frame forming the prediction block and the residual block. The process of determining the motion vector is typically referred to as motion estimation. An intra-coded block is coded based on an intra-prediction mode and the residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, such as the frequency domain, to generate residual transform coefficients, which can then be quantized. The quantized transform coefficients initially arranged as a two-dimensional array can be scanned to produce a one-dimensional vector of the transform coefficients, and then entropy coded into a video bitstream for more compression.

[0006] Then, the encoded video bitstream is saved in a computer-readable storage medium (e.g., flash memory) to be accessed by another electronic device having digital video capabilities or directly transmitted to the electronic device in a wired or wireless manner. Then, the electronic device performs video decompression (which is a process opposite to the video compression described above) by, for example, parsing the encoded video bitstream to obtain syntax elements from the bitstream and reconstructing the digital video data from the encoded video bitstream into its original format at least partially based on the syntax elements obtained from the bitstream, and rendering the reconstructed digital video data on a display of the electronic device.

[0007] As digital video quality goes from high definition to 4K×2K or even 8K×4K, the amount of video data to be encoded / decoded grows exponentially. There have been challenges in how to more efficiently encode / decode video data while maintaining the image quality of the decoded video data. Summary of the Invention

[0008] Embodiments described in this application relate to video data encoding and decoding, and more particularly to methods and systems for video encoding and decoding of advanced syntax in a video bitstream applicable to one or more video codec standards.

[0009] According to a first aspect of the present application, a method for decoding video data includes: receiving, from the bitstream, a plurality of syntax elements at the sequence parameter set (SPS) level, wherein the plurality of syntax elements are associated with a predefined function and are sequentially arranged in the bitstream; determining, based on the plurality of syntax elements satisfying a predefined condition, receiving, from the bitstream, a second syntax element immediately following the plurality of syntax elements; determining, based on the plurality of syntax elements not satisfying the predefined condition, setting a default value as the second syntax element; and performing the predefined function on the video data from the bitstream according to the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of an intra prediction function, an inter prediction function, and a merge mode.

[0010] According to a second aspect of the present application, a method for decoding video data includes: receiving, from the bitstream, a plurality of syntax elements at one or more levels of the picture parameter set (PPS) level and the slice level, wherein the plurality of syntax elements are associated with a predefined function and are sequentially arranged in the bitstream; determining, based on the plurality of syntax elements satisfying a predefined condition, receiving, from the bitstream, a second syntax element immediately following the plurality of syntax elements; determining, based on the plurality of syntax elements not satisfying the predefined condition, setting a default value as the second syntax element; and performing the predefined function on the video data from the bitstream according to the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of a quantization function, an intra prediction function, and an inter prediction function.

[0011] According to a third aspect of the present application, an electronic device includes one or more processing units, a memory, and a plurality of programs stored in the memory. When executed by the one or more processing units, the programs cause the electronic device to perform the method for decoding video data as described above.

[0012] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium stores a plurality of programs for execution by an electronic device having one or more processing units. When executed by the one or more processing units, the programs cause the electronic device to perform the method for decoding video data as described above. Description of the Drawings

[0013] The drawings, which are included to provide a further understanding of the embodiments and are incorporated herein and constitute a part of this specification, illustrate the described embodiments and, together with the specification, serve to explain the basic principles. Like reference numerals refer to corresponding parts.

[0014] Figure 1is a block diagram illustrating an exemplary video encoding and decoding system according to some embodiments of the present disclosure.

[0015] Figure 2 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure.

[0016] Figure 3 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure.

[0017] Figures 4A to 4E is a block diagram illustrating how a frame is recursively partitioned into multiple video blocks having different sizes and shapes according to some embodiments of the present disclosure.

[0018] Figure 5 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure.

[0019] Figure 6 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure.

[0020] Figure 7 is a flowchart illustrating an exemplary process by which a video decoder implements techniques for decoding video data according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to assist in understanding the subject matter presented herein. However, it will be apparent to one of ordinary skill in the art that various alternative solutions may be used and the subject matter may be practiced without these specific details. For example, it will be apparent to one of ordinary skill in the art that the subject matter presented herein may be implemented on many types of electronic devices having digital video capabilities.

[0022] Figure 1 is a block diagram illustrating an exemplary system 10 for encoding and decoding video blocks in parallel. As Figure 1As shown, system 10 includes a source device 12 that generates and encodes video data to be decoded by a destination device 14 at a later time. The source device 12 and the destination device 14 can include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smart phones, set-top boxes, digital televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some embodiments, the source device 12 and the destination device 14 are equipped with wireless communication capabilities.

[0023] In some embodiments, the destination device 14 can receive the encoded video data to be decoded via a link 16. The link 16 can include any type of communication medium or device capable of moving the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 can include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium can include routers, switches, base stations, or any other devices that can be used to facilitate communication from the source device 12 to the destination device 14.

[0024] In some other embodiments, the encoded video data can be transmitted from the output interface 22 to the storage device 32. Subsequently, the encoded video data in the storage device 32 can be accessed by the destination device 14 via the input interface 28. The storage device 32 can include any of a variety of distributed or locally accessible data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memories, volatile memories, or non-volatile memories, or any other suitable digital storage media for storing the encoded video data. In a further example, the storage device 32 can correspond to a file server or another intermediate storage device that can hold the encoded video data generated by the source device 12. The destination device 14 can access the stored video data from the storage device 32 via streaming or downloading. The file server can be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device 14 can access the encoded video data via any standard data connection, including a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination of both.

[0025] As Figure 1 shown, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 can include sources such as a video capture device, e.g., a camera, a video archive containing previously captured video, a video feed interface for receiving video from a video content provider, and / or a computer graphics system for generating computer graphics data as the source video, or a combination of these sources. As an example, if the video source 18 is a camera of a security monitoring system, the source device 12 and the destination device 14 can form a camera phone or a video phone. However, the embodiments described in this application can generally be applicable to video coding and decoding and can be applied to wireless and / or wired applications.

[0026] The captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data can be transmitted directly from the output interface 22 of the source device 12 to the destination device 14. The encoded video data can also (or alternatively) be stored on the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 can further include a modem and / or a transmitter.

[0027] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 can include a receiver and / or a modem and receives the encoded video data via a link 16. The encoded video data transmitted via the link 16 or provided on a storage device 32 can include various syntax elements generated by the video encoder 20 for use by the video decoder 30 when decoding the video data. Such syntax elements can be included within the encoded video data transmitted on a communication medium, stored on a storage medium, or stored in a file server.

[0028] In some embodiments, the destination device 14 can include a display device 34, which can be an integrated display device and an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data to a user and can include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0029] The video encoder 20 and the video decoder 30 can operate according to proprietary or industry standards, such as VVC, HEVC, MPEG-4 Part 10, Advanced Video Coding (AVC), or extensions of such standards. It should be understood that the present application is not limited to a particular video coding / decoding standard and can be applicable to other video coding / decoding standards. It is generally contemplated that the video encoder 20 of the source device 12 can be configured to encode video data according to any of these current or future standards. Similarly, it is generally also contemplated that the video decoder 30 of the destination device 14 can be configured to decode video data according to any of these current or future standards.

[0030] The video encoder 20 and the video decoder 30 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the video coding / decoding operations disclosed in the present disclosure. Each of the video encoder 20 and the video decoder 30 can be included in one or more encoders or decoders, and any of the one or more encoders or decoders can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0031] Figure 2FIG. is a block diagram illustrating an exemplary video encoder 20 according to some embodiments described in the present application. The video encoder 20 may perform intra prediction coding and decoding and inter prediction coding and decoding on video blocks within a video frame. Intra prediction coding and decoding relies on spatial prediction to reduce or remove spatial redundancy of video data within a given video frame or picture. Inter prediction coding and decoding relies on temporal prediction to reduce or remove temporal redundancy of video data within adjacent video frames or pictures of a video sequence.

[0032] As Figure 2 shown, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some embodiments, the video encoder 20 further includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. A deblocking filter (not shown) may be located between the adder 62 and the DPB 64 to filter block boundaries to remove blocking artifact from the reconstructed video. In addition to the deblocking filter, a loop filter (not shown) may be used to filter the output of the adder 62. The video encoder 20 may take the form of fixed or programmable hardware units, or may be partitioned among one or more of the illustrated fixed or programmable hardware units.

[0033] The video data memory 40 may store video data to be encoded by components of the video encoder 20. The video data in the video data memory 40 may be obtained, for example, from a video source 18. The DPB 64 is a buffer that stores reference video data for use by the video encoder 20 when encoding video data (e.g., in intra prediction coding and decoding mode or inter prediction coding and decoding mode). The video data memory 40 and the DPB 64 may be formed of any of a variety of memory devices. In various examples, the video data memory 40 may be on-chip with other components of the video encoder 20 or off-chip relative to those components.

[0034] As Figure 2As shown, after receiving video data, the partitioning unit 45 in the prediction processing unit 41 partitions the video data into video blocks. This partitioning may also include partitioning the video frame into stripes, tiles, or other larger coding units (CUs) according to a predefined segmentation structure, such as a quadtree structure associated with the video data. The video frame may be divided into multiple video blocks (or sets of video blocks referred to as tiles). The prediction processing unit 41 may select one of multiple possible prediction coding modes for the current video block based on error results (e.g., coding rate and distortion level), such as one of multiple intra-prediction coding modes or one of multiple inter-prediction coding modes. The prediction processing unit 41 may provide the resulting intra-prediction coded block or inter-prediction coded block to the adder 50 to generate a residual block, and to the adder 62 to reconstruct the coded block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra-mode indicators, partitioning information, and other such syntax information to the entropy coding unit 56.

[0035] To select an appropriate intra-prediction coding mode for the current video block, the intra-prediction processing unit 46 in the prediction processing unit 41 may perform intra-prediction coding of the current video block relative to one or more adjacent blocks in the same frame as the current block to be coded and decoded, to provide spatial prediction. The motion estimation unit 42 and the motion compensation unit 44 in the prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more prediction blocks in one or more reference frames, to provide temporal prediction. The video encoder 20 may execute multiple coding channels, for example, in order to select an appropriate coding mode for each block of the video data.

[0036] In some embodiments, the motion estimation unit 42 determines the inter-prediction mode of the current video frame by generating motion vectors according to a predefined pattern within the video frame sequence, where the motion vectors indicate the displacement of the prediction unit (PU) of the video block within the current video frame relative to the prediction block within the reference video frame. The motion estimation performed by the motion estimation unit 42 is a process of generating motion vectors, which estimates the motion of the video block. The motion vector may, for example, indicate the displacement of the PU of the video block within the current video frame or picture relative to the prediction block (or other coded unit) within the reference frame, where the prediction block is relative to the current block (or other coded unit) coded and decoded within the current frame. The predefined pattern may designate the video frames in the sequence as P frames or B frames. The intra-BC unit 48 may determine the vector for performing intra-BC coding in a manner similar to the way the motion estimation unit 42 determines the motion vector for inter-prediction, e.g., a block vector, or may utilize the motion estimation unit 42 to determine the block vector.

[0037] A prediction block is a block in a reference frame that is considered to closely match, in terms of pixel differences, a PU of a video block to be coded or decoded, where the pixel differences can be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some embodiments, video encoder 20 may compute values at sub-integer pixel positions of reference frames stored in DPB 64. For example, video encoder 20 may insert values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference frame. Accordingly, motion estimation unit 42 may perform a motion search relative to full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.

[0038] Motion estimation unit 42 computes a motion vector for a PU of a video block in an inter-predicted coded frame by comparing the position of the PU with the position of a prediction block of a reference frame selected from a first reference frame list (list 0) or a second reference frame list (list 1), each of the lists identifying one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the computed motion vector to motion compensation unit 44, and then to entropy coding unit 56.

[0039] Motion compensation performed by motion compensation unit 44 may involve obtaining or generating a prediction block based on the motion vector determined by motion estimation unit 42. After receiving the motion vector for a PU of the current video block, motion compensation unit 44 may locate the prediction block pointed to by the motion vector in one of the reference frame lists, obtain the prediction block from DPB 64 and forward the prediction block to adder 50. Then, adder 50 forms a residual video block with pixel differences by subtracting the pixel values of the prediction block provided by motion compensation unit 44 from the pixel values of the current video block being coded or decoded. The pixel differences forming the residual video block may include a luminance difference component or a chrominance difference component or both. Motion compensation unit 44 may also generate syntax elements associated with video blocks of a video frame for use by video decoder 30 when decoding the video blocks of the video frame. The syntax elements may include, for example, syntax elements defining the motion vector for identifying the prediction block, any flags indicating a prediction mode, or any other syntax information described herein. Note that motion estimation unit 42 and motion compensation unit 44 may be highly integrated, but are shown separately for conceptual purposes.

[0040] In some embodiments, the intra BC unit 48 can generate vectors and obtain a prediction block in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, but where the prediction block is in the same frame as the current block being coded and where, relative to a motion vector, the vector is referred to as a block vector. In particular, the intra BC unit 48 can determine an intra prediction mode for coding the current block. In some examples, the intra BC unit 48 can, for example, code the current block using various intra prediction modes during a separate coding pass and test its performance through rate-distortion analysis. Next, the intra BC unit 48 can select an appropriate intra prediction mode to use among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 can use rate-distortion analysis for the various tested intra prediction modes to calculate rate-distortion values and select, among the tested modes, the intra prediction mode having the best rate-distortion characteristics as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original uncoded block (which was coded to produce the encoded block) and the bitrate (i.e., the number of bits) used to produce the encoded block. The intra BC unit 48 can calculate a ratio based on the distortion and rate of each encoded block to determine which intra prediction mode exhibits the best rate-distortion value for the block.

[0041] In other examples, the intra BC unit 48 can use the motion estimation unit 42 and the motion compensation unit 44, in whole or in part, to perform such functions for intra BC prediction in accordance with the embodiments described herein. In either case, for intra block copy, the prediction block can be a block that is considered to closely match the block to be coded and decoded in terms of pixel differences, which can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics, and the identification of the prediction block can include calculating values at sub-integer pixel positions.

[0042] Regardless of whether the prediction block is from the same frame according to intra prediction or from a different frame according to inter prediction, the video encoder 20 can form a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being coded and decoded, thereby forming pixel differences. The pixel differences forming the residual video block can include luminance component differences and chrominance component differences.

[0043] As described above, the intra prediction processing unit 46 may perform intra prediction on a current video block as an alternative to inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or intra block copy prediction performed by the intra BC unit 48. In particular, the intra prediction processing unit 46 may determine an intra prediction mode for encoding the current block. To this end, the intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes during a separate encoding pass, and the intra prediction processing unit 46 (or in some examples, a mode selection unit) may select an appropriate intra prediction mode from the tested intra prediction modes to use. The intra prediction processing unit 46 may provide information indicating the selected intra prediction mode of the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode in the bitstream.

[0044] After the prediction processing unit 41 determines a predicted block of a current video block via inter prediction or intra prediction, the adder 50 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be included in one or more transform units (TUs) and provided to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform.

[0045] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 may then perform a scan on the matrix including the quantized transform coefficients. Alternatively, the entropy encoding unit 56 may perform the scan.

[0046] After quantization, the entropy encoding unit 56 entropy encodes the quantized transform coefficients into a video bitstream using, for example, context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy coding methods or techniques. The encoded bitstream may then be transmitted to the video decoder 30 or archived in the storage device 32 for later transmission to or retrieval by the video decoder 30. The entropy encoding unit 56 may also entropy encode the motion vectors and other syntax elements of the currently decoded video frame.

[0047] The dequantization unit 58 and the inverse transform processing unit 60 respectively apply dequantization and inverse transform to reconstruct the residual video block in the pixel domain to generate a reference block for predicting other video blocks. As described above, the motion compensation unit 44 can generate a motion-compensated prediction block from one or more reference blocks of the frames stored in the DPB 64. The motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.

[0048] The adder 62 adds the reconstructed residual block to the motion-compensated prediction block generated by the motion compensation unit 44 to generate a reference block for storage in the DPB 64. The reference block can then be used as a prediction block by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 to perform inter prediction on another video block in a subsequent video frame.

[0049] Figure 3 is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, a dequantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction processing unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process generally opposite to the encoding process described above in connection with Figure 2 the video encoder 20. For example, the motion compensation unit 82 can generate prediction data based on the motion vectors received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.

[0050] In some examples, the units of the video decoder 30 can be assigned to perform the embodiments of the present application. Similarly, in some examples, the embodiments of the present disclosure can be divided among one or more units of the video decoder 30. For example, the intra BC unit 85 can perform the embodiments of the present application alone or in combination with other units of the video decoder 30, such as the motion compensation unit 82, the intra prediction processing unit 84, and the entropy decoding unit 80. In some examples, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 can be performed by other components of the prediction processing unit 81, such as the motion compensation unit 82.

[0051] The video data memory 79 can store video data to be decoded by other components of the video decoder 30, such as an encoded video bitstream. For example, the video data stored in the video data memory 79 can be obtained from a storage device 32, a local video source (such as a camera), by wire or wireless network transmission of the video data or by accessing a physical data storage medium (e.g., a flash drive or a hard disk). The video data memory 79 can include a coded picture buffer (CPB) that stores the encoded video data from the encoded video bitstream. The decoded picture buffer (DPB) 92 of the video decoder 30 stores reference video data for use by the video decoder 30 when decoding video data (e.g., in an intra prediction codec mode or an inter prediction codec mode). The video data memory 79 and the DPB 92 can be formed of any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For illustrative purposes, the video data memory 79 and the DPB 92 are depicted in Figure 3 as two different components of the video decoder 30. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 can be provided by the same memory device or separate memory devices. In some examples, the video data memory 79 can be on-chip with other components of the video decoder 30 or off-chip relative to those components.

[0052] During the decoding process, the video decoder 30 receives an encoded video bitstream representing video blocks of an encoded video frame and associated syntax elements. The video decoder 30 can receive syntax elements at the video frame level and / or at the video block level. The entropy decoding unit 80 of the video decoder 30 entropy decodes the bitstream to generate quantized coefficients, motion vectors, or intra prediction mode indicators and other syntax elements. The entropy decoding unit 80 then forwards the motion vectors and other syntax elements to the prediction processing unit 81.

[0053] When a video frame is coded as an intra prediction codec (I) frame or for an intra codec prediction block in other types of frames, the intra prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for video blocks of the current video frame based on the signal-transmitted intra prediction mode and reference data from previously decoded blocks of the current frame.

[0054] When a video frame is encoded / decoded as an inter-predicted (i.e., B or P) frame, the motion compensation unit 82 of the prediction processing unit 81 generates one or more prediction blocks of a video block of the current video frame based on the motion vectors and other syntax elements received from the entropy decoding unit 80. Each prediction block may be generated from a reference frame within one of the reference frame lists. The video decoder 30 may construct the reference frame lists: list 0 and list 1, using a default construction technique based on the reference frames stored in the DPB 92.

[0055] In some examples, when a video block is encoded / decoded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block for the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within a reconstructed region of the same picture as the current video block defined by the video encoder 20.

[0056] The motion compensation unit 82 and / or the intra BC unit 85 determine the prediction information of the video block of the current video frame by parsing the motion vectors and other syntax elements, and then use the prediction information to generate a prediction block of the decoded current video block. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra prediction or inter prediction) for encoding / decoding the video block of the video frame, the inter-predicted frame type (e.g., B or P), the construction information of one or more of the reference frame lists in the reference frame list of the frame, the motion vector of each inter-predicted encoded video block of the frame, the inter-predicted state of each inter-predicted encoded / decoded video block of the frame, and other information for decoding the video block in the current video frame.

[0057] Similarly, the intra BC unit 85 may use some of the received syntax elements (e.g., flags) to determine that the current video block is predicted using: the intra BC mode, the construction information that the video block of the frame is within the reconstructed region and should be stored in the DPB 92, the block vector of each intra BC predicted video block of the frame, the intra BC prediction state of each intra BC predicted video block of the frame, and other information for decoding the video block in the current video frame.

[0058] The motion compensation unit 82 may also perform interpolation using an interpolation filter as used by the video encoder 20 during encoding of the video block to calculate the interpolation values of the sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the video encoder 20 from the received syntax elements and use the interpolation filter to generate the prediction block.

[0059] The dequantization unit 86 dequantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameter for determining the quantization degree calculated by the video encoder 20 for each video block in the video frame. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT, inverse integer transform, or conceptually similar inverse transform process) to the transform coefficients in order to reconstruct the residual block in the pixel domain.

[0060] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block for the current video block based on the vectors and other syntax elements, the adder 90 reconstructs the decoded video block of the current video block by adding the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. A loop filter (not shown) may be located between the adder 90 and the DPB 92 to further process the decoded video block. Then the decoded video blocks in a given frame are stored in the DPB 92, which stores reference frames for subsequent motion compensation of the next video blocks. The DPB 92 or a memory device separate from the DPB 92 may also store the decoded video for later presentation on a display device such as Figure 1 the display device 34 as shown.

[0061] In a typical video encoding and decoding process, a video sequence typically includes an ordered set of frames or pictures. Each frame may include three arrays of samples, denoted as SL, SCb, and SCr, respectively. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other instances, a frame may be monochrome and thus include only one two-dimensional array of luminance samples.

[0062] As Figure 4A shown, the video encoder 20 (or more specifically, the partitioning unit 45) generates an encoded representation of a frame by first partitioning the frame into a set of coding tree units (CTUs). A video frame may include an integer number of CTUs sorted consecutively in raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size, i.e., one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that the present application is not necessarily limited to a specific size. As Figure 4BAs shown, each CTU may include one coding tree block (CTB) of luminance samples, two corresponding coding tree blocks of chrominance samples, and syntax elements for coding the samples of the coding tree blocks. The syntax elements describe the attributes of different types of units of the coded blocks of pixels and how the video sequence can be reconstructed at the video decoder 30. The syntax elements include inter-prediction or intra-prediction, intra-prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture with three separate color planes, the CTU may include a single coding tree block and syntax elements for encoding and decoding the samples of the coding tree block. The coding tree block may be an N×N sample block.

[0063] To achieve better performance, the video encoder 20 may recursively perform tree partitioning (such as binary tree partitioning, ternary tree partitioning, quadtree partitioning, or a combination of both) on the coding tree block of the CTU, and partition the CTU into smaller coding units (CUs). As Figure 4C depicted, first, the 64×64 CTU 400 is divided into four smaller CUs, each with a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four 16×16 CUs according to the block size. The two 16×16 CUs 430 and 440 are each further divided into four 8×8 CUs according to the block size. Figure 4D depicts a quadtree data structure that illustrates the final result of the partitioning process of the CTU 400 as depicted in Figure 4C . Each leaf node of the quadtree corresponds to a CU with a corresponding size in the range of 32×32 to 8×8. Similar to the Figure 4B depicted CTU, each CU may include a coded block (CB) of luminance samples and two corresponding coded blocks of chrominance samples of the same-size frame, as well as syntax elements for encoding and decoding the samples of the coded blocks. In a monochrome picture or a picture with three separate color planes, the CU may include a single coded block and a syntax structure for encoding and decoding the samples of the coded block. It should be noted that Figure 4C and Figure 4D the quadtree partitioning depicted in is for illustrative purposes only, and a CTU can be split into multiple CUs to adapt to different local characteristics based on quadtree / ternary tree / binary tree partitioning. In a multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned by a binary tree structure or a ternary tree structure. As Figure 4E shown, there are five partitioning types, namely quaternary partitioning, horizontal binary partitioning, vertical binary partitioning, horizontal ternary partitioning, and vertical ternary partitioning.

[0064] In some embodiments, video encoder 20 may further partition the coded block of a CU into one or more M×N prediction blocks (PBs). A prediction block is a rectangular (square or non-square) block of samples to which the same prediction (inter-frame or intra-frame) is applied. The prediction unit (PU) of a CU may include a prediction block of luma samples, two corresponding prediction blocks of chroma samples, and syntax elements for predicting the prediction block. In a monochrome picture or a picture with three separate color planes, the PU may include a single prediction block and a syntax structure for predicting the prediction block. Video encoder 20 may generate prediction luma, Cb, and Cr blocks for the luma, Cb, and Cr prediction blocks of each PU of the CU.

[0065] Video encoder 20 may use intra-frame prediction or inter-frame prediction to generate the prediction block of the PU. If video encoder 20 uses intra-frame prediction to generate the prediction block of the PU, then video encoder 20 may generate the prediction block of the PU based on the decoded samples of the frame associated with the PU. If video encoder 20 uses inter-frame prediction to generate the prediction block of the PU, then video encoder 20 may generate the prediction block of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.

[0066] After video encoder 20 generates the prediction luma, Cb, and Cr blocks of one or more PUs of a CU, video encoder 20 may generate a luma residual block of the CU by subtracting the prediction luma block of the CU from its original luma coded block, such that each sample in the luma residual block of the CU indicates the difference between the luma sample in one of the prediction luma blocks of the CU and the corresponding sample in the original luma coded block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block of the CU, respectively, such that each sample in the Cb residual block of the CU indicates the difference between the Cb sample in one of the prediction Cb blocks of the CU and the corresponding sample in the original Cb coded block of the CU, and each sample in the Cr residual block of the CU may indicate the difference between the Cr sample in one of the prediction Cr blocks of the CU and the corresponding sample in the original Cr coded block of the CU.

[0067] In addition, as Figure 4CAs illustrated, video encoder 20 may use quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of a CU into one or more luminance, Cb, and Cr transform blocks. A transform block is a rectangular (square or non-square) block of samples to which the same transform is applied. The transform unit (TU) of a CU may include a transform block of luminance samples, two corresponding transform blocks of chrominance samples, and syntax elements for transforming the samples of the transform block. Thus, each TU of a CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some examples, the luminance transform block associated with a TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture with three separate color planes, a TU may include a single transform block and a syntax structure for transforming the samples of the transform block.

[0068] Video encoder 20 may apply one or more transforms to the luminance transform block of a TU to generate a luminance coefficient block of the TU. A coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalars. Video encoder 20 may apply one or more transforms to the Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 may apply one or more transforms to the Cr transform block of the TU to generate a Cr coefficient block of the TU.

[0069] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to the process of quantizing transform coefficients to potentially reduce the amount of data used to represent the transform coefficients to provide further compression. After video encoder 20 quantizes the coefficient block, video encoder 20 may entropy code the syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform context-adaptive binary arithmetic coding (CABAC) on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a bit sequence representing the coded frame and associated data, which is stored in storage device 32 or transmitted to destination device 14.

[0070] After receiving the bitstream generated by the video encoder 20, the video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. The video decoder 30 may reconstruct a frame of video data at least in part based on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by the video encoder 20. For example, the video decoder 30 may perform an inverse transform on a coefficient block associated with a TU of a current CU to reconstruct a residual block associated with the TU of the current CU. The video decoder 30 also reconstructs an encoded block of the current CU by adding samples of a prediction block of the PU of the current CU to corresponding samples of a transform block of the TU of the current CU. After reconstructing the encoded block of each CU of a frame, the video decoder 30 may reconstruct the frame.

[0071] Generally, the basic intra prediction scheme applied in VVC remains the same as that of HEVC, except that several modules are further extended and / or improved, for example, matrix weighted intra prediction (MIP) coding / decoding mode, intra sub-partition (ISP) coding / decoding mode, extended intra prediction using wide-angle intra directions, position-dependent intra prediction combination (PDPC), and 4-tap intra interpolation. The main focus of this disclosure is to improve the existing advanced syntax design in the VVC standard. The relevant background knowledge is elaborated in detail in the following sections.

[0072] Similar to HEVC, VVC uses a bitstream structure based on network abstraction layer (NAL) units. The encoded bitstream is partitioned into multiple NAL units, which should be smaller than the maximum transmission unit size when transmitted over a lossy packet network. Each NAL unit consists of a NAL unit header and the subsequent NAL unit payload. There are two conceptual categories of NAL units. Video coding layer (VCL) NAL units that contain encoded sample data, such as encoded slice NAL units, and conversely, non-VCL NAL units that contain metadata that generally belongs to more than one encoded picture, or parameter set NAL units in cases where the association with a single encoded picture would be meaningless, or SEI NAL units in cases where information is not required during the decoding process.

[0073] In VVC, a two-byte NAL unit header is introduced, and this design is expected to be sufficient to support future extensions. The syntax and associated semantics of the NAL unit header in the current VVC draft specification are illustrated in Tables 1 and 2, respectively. How to read Table 1 is illustrated in the appendix section of this invention, which can also be found in the VVC specification.

[0074] Table 1. NAL unit header syntax

[0075]

[0076] Table 2. NAL Unit Header Semantics

[0077]

[0078] Table 3. NAL Unit Type Codes and NAL Unit Type Categories

[0079]

[0080] VVC inherits the parameter set concept of HEVC and makes some modifications and additions. The parameter set can be part of the video bitstream or can be received by the decoder through other means (including out-of-band transmission using a reliable channel, hard encoding / decoding in the encoder and decoder, etc.). The parameter set contains identifiers that are directly or indirectly referenced from the slice header, as discussed in more detail later. The referencing process is called "activation". Depending on the parameter set type, activation occurs per picture or per sequence. Among other reasons, the concept of activation by reference is introduced because in the case of out-of-band transmission, implicit activation cannot be achieved by relying on the position of information in the bitstream (as is common for other syntax elements of the video codec).

[0081] The Video Parameter Set (VPS) is introduced to convey information applicable to multiple layers as well as sub-layers. The VPS is introduced to address these shortcomings and enable a clean and scalable high-level design for multi-layer codecs. Regardless of whether each layer of a given video sequence has the same or different Sequence Parameter Sets (SPS), they all reference the same VPS. The syntax and associated semantics of the video parameter set in the current VVC draft specification are illustrated in Tables 4 and 5 respectively. How to read Table 4 is illustrated in the appendix section of the present invention, which can also be found in the VVC specification.

[0082] Table 4. Video Parameter Set RBSP Syntax

[0083]

[0084]

[0085]

[0086]

[0087] Table 5. Video Parameter Set RBSP Semantics

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097] In VVC, the SPS contains information applicable to all slices of the decoded video sequence. The decoded video sequence starts from an Instantaneous Decoding Refresh (IDR) picture or a BLA picture or a CRA picture of the first picture in the bitstream, and includes all subsequent pictures that are not IDR or BLA pictures. The bitstream consists of one or more decoded video sequences. The content of the SPS can be roughly divided into six categories: 1) self-reference (its own ID); 2) decoder operating point related information (profile, level, picture size, number of sub-layers, etc.); 3) enabling flags for certain tools within the profile, and associated codec tool parameters in the case of enabled tools; 4) information defining the flexibility of the structure and transform coefficient coding; 5) temporal scalability control; and 6) Visual Usability Information (VUI) including HRD information. The syntax and associated semantics of the Sequence Parameter Set in the current VVC draft specification are illustrated in Tables 6 and 7 respectively. How to read Table 6 is illustrated in the appendix section of the present invention, which can also be found in the VVC specification.

[0098] Table 6. Sequence Parameter Set RBSP Syntax

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106] Table 7. Sequence Parameter Set RBSP Semantics

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114]

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126] The picture parameter set (PPS) of VVC contains such information that can be changed between pictures. The PPS includes information that is roughly equivalent to a part of the PPS in HEVC, including: 1) self-reference; 2) initial picture control information, such as the initial quantization parameter (QP), the number of flags indicating the use or presence of certain tools, or the control information in the slice header; and 3) tile information. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are illustrated in Tables 8 and 9 respectively. How to read Table 8 is illustrated in the appendix section of the present invention, which can also be found in the VVC specification.

[0127] Table 8. Picture Parameter Set RBSP Syntax

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] Table 9. Picture Parameter Set RBSP Semantics

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145] The slice header contains information that can change between slices and such picture-related information that is relatively small or only relevant for a particular slice or picture type. The size of the slice header may be significantly larger than the PPS, especially when there are tile or wavefront entry point offsets in the slice header and RPS, explicitly signaling prediction weight or reference picture list modification. The syntax and associated semantics of the sequence parameter set in the current VVC draft specification are illustrated in Tables 10 and 11, respectively. How to read Table 10 is illustrated in the appendix section of the present invention, which can also be found in the VVC specification.

[0146] Table 10. Picture Header Structure Syntax

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153] Table 11. Semantics of Picture Header Structure

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166] In the current VVC, when there are similar syntax elements for intra prediction and inter prediction respectively, in some places the syntax elements related to inter prediction are defined before the syntax elements related to intra prediction. Considering the fact that intra prediction is allowed in all picture / strip types while inter prediction is not, such an order may not be preferred. From the perspective of standardization, it would be beneficial to always define the syntax related to intra prediction before the syntax used for inter prediction.

[0167] It is also observed that in the current VVC, some syntax elements that are highly related to each other are defined in a sprawling manner in different places. From the perspective of standardization, it would also be beneficial to group some syntax together.

[0168] In the present disclosure, to solve the problems pointed out in the "Problem Description" section, methods for simplifying and / or further improving the existing design of the high-level syntax are provided. It should be noted that the invented methods can be applied independently or jointly.

[0169] In the present disclosure, the syntax elements are rearranged such that the syntax elements related to intra prediction are defined before the syntax elements related to inter prediction. According to the present disclosure, the partition constraint syntax elements are grouped by prediction type, where those related to intra prediction come first and those related to inter prediction come later. In one embodiment, the order of the partition constraint syntax elements in the SPS is the same as the order of the partition constraint syntax elements in the picture header. An example of the decoding process for the VVC draft is illustrated in Table 12 below. The changes to the VVC draft are shown in italic font.

[0170] Table 12. Designed Sequence Parameter Set RBSP Syntax

[0171]

[0172] Figure 5 FIG. 500 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure. For example, the method can be applied to a decoder.

[0173] In step 510, the decoder may receive the arranged partition constraint syntax elements at the SPS level. These arranged partition constraint syntax elements are arranged such that the syntax elements related to intra prediction are defined before the syntax elements related to inter prediction.

[0174] In step 512, the decoder may obtain a first reference picture I associated with a video block in the bitstream (0) and a second reference picture I (1) . In display order, the first reference picture I (0) is before the current picture and the second reference picture I (1) is after the current picture.

[0175] In step 514, the decoder may obtain a first predicted sample I of the video block from a reference block in the first reference picture I (0) (i, j). i and j represent the coordinates of a sample in the current picture. (0)

[0176] In step 516, the decoder may obtain a second predicted sample I of the video block from a reference block in the second reference picture I (1) (i, j). (1)

[0177] In step 518, the decoder may obtain the bi - directionally predicted sample based on the arranged partition constraint syntax elements, the first predicted sample I (0) (i, j) and the second predicted sample I (1) (i, j).

[0178] Figure 6FIG. 600 is a flowchart illustrating an exemplary method for decoding a video signal according to some embodiments of the present disclosure. For example, the method may be applied to a decoder. In step 610, the decoder may receive a bitstream including a VPS, an SPS, a PPS, a picture header, and a slice header for decoded video data. In step 612, the decoder may decode the VPS. In step 614, the decoder may decode the SPS and obtain the arrangement partition constraint syntax elements at the SPS level. In step 616, the decoder may decode the PPS. In step 618, the decoder may decode the picture header. In step 620, the decoder may decode the slice header. In step 622, the decoder may decode the video data based on the VPS, the SPS, the PPS, the picture header, and the slice header.

[0179] In the present disclosure, syntax elements related to the dual-tree chroma type are designed to be grouped. In one embodiment, the partition constraint syntax elements for the dual-tree chroma in the SPS should be signaled together in the case of dual-tree chroma. An example of the decoding process for the VVC draft is illustrated in Table 13 below. Changes to the VVC draft are shown in italic font.

[0180] Table 13. Designed sequence parameter set RBSP syntax

[0181]

[0182] If the intra prediction related syntax is defined before the syntax related to inter prediction is also considered, another example of the decoding process for the VVC draft is illustrated in Table 14 below according to the method of the present disclosure. Changes to the VVC draft are shown in italic font.

[0183] Table 14. Designed sequence parameter set RBSP syntax

[0184]

[0185] As mentioned in the earlier description, according to the current VVC, intra prediction is allowed in all picture / slice types and inter prediction is not allowed. According to the present disclosure, flags are designed to be added to the VVC syntax at a specific coding / decoding level to indicate whether inter prediction is used in a sequence, a picture, and / or a slice. In the case where inter prediction is not used, the syntax related to inter prediction is not signaled at the corresponding coding / decoding level (e.g., sequence, picture, and / or slice level).

[0186] In one example, according to the method of the present disclosure, a flag is added in the SPS to indicate whether inter prediction is used in the encoding and decoding of the current video sequence. When not used, the syntax elements related to inter prediction are not signaled in the SPS. An example of the decoding process for the VVC draft is illustrated in Table 15 below. The changes to the VVC draft are shown in italic font.

[0187] Table 15. Designed Sequence Parameter Set RBSP Syntax

[0188]

[0189]

[0190] Sequence Parameter Set RBSP Semantics

[0191] When sps_inter_slice_used_flag is equal to 0, all decoded slices of the video sequence have a slice_type equal to 2. When sps_inter_slice_used_flag is equal to 1, one or more decoded slices with a slice_type equal to 0 or 1 may or may not exist in the video sequence.

[0192] In some embodiments, the syntax elements are arranged such that syntax elements related to similar functions (e.g., intra tools, inter tools, screen content tools, transform tools, quantization tools, loop filter tools, and / or partitioning tools) are grouped according to the VVC syntax at certain encoding and decoding levels (e.g., sequence, picture, and / or slice level). According to some embodiments, the syntax elements in the Sequence Parameter Set (SPS) are arranged such that syntax elements related to similar functions are grouped. An example of the decoding process for VVC is illustrated in Table 16 below. Table 16 shows the syntax for grouping syntax elements with similar functions at the SPS level.

[0193] Table 16. Syntax for Grouping Syntax Elements with Similar Functions at the SPS Level.

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202] Another example of the decoding process for VVC is illustrated in Table 17 below. Table 17 shows the syntax that groups syntax elements with similar functions at the SPS level.

[0203] Table 17. Syntax for Grouping Syntax Elements with Similar Functions at the SPS Level

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212] According to some embodiments, the syntax elements in the Picture Parameter Set (PPS) are arranged such that syntax elements related to similar functions are grouped. An example of the decoding process for VVC is illustrated in Table 18 below. Table 18 shows the syntax that groups syntax elements with similar functions at the PPS level.

[0213] Table 18. Syntax for Grouping Syntax Elements with Similar Functions at the PPS Level

[0214]

[0215]

[0216]

[0217]

[0218] Another example of the decoding process for VVC is illustrated in Table 19 below. Table 19 shows the syntax that groups syntax elements with similar functions at the PPS level.

[0219] Table 19. Syntax for Grouping Syntax Elements with Similar Functions at the PPS Level

[0220]

[0221]

[0222]

[0223]

[0224]

[0225] Figure 7 FIG. 700 is a flow chart illustrating an exemplary process according to some embodiments of the present disclosure, by which a video decoder (e.g., video decoder 30) implements techniques for decoding video data.

[0226] As Figure 7 shown, in some embodiments, video decoder 30 receives a plurality of syntax elements at the sequence parameter set (SPS) level from a bitstream, and the plurality of syntax elements are associated with predefined functions and arranged sequentially in the bitstream (710).

[0227] Based on a determination that the plurality of syntax elements satisfy a predefined condition, video decoder 30 receives a second syntax element from the bitstream that immediately follows the plurality of syntax elements (720).

[0228] Based on a determination that the plurality of syntax elements do not satisfy the predefined condition, video decoder 30 sets a default value as the second syntax element (730).

[0229] Video decoder 30 performs a predefined function on video data from the bitstream based on the plurality of syntax elements and the second syntax element, and the predefined function is a predefined function selected from the group consisting of an intra prediction function, an inter prediction function, and a merge mode (740).

[0230] In some embodiments, the plurality of syntax elements associated with the intra prediction function and arranged sequentially in the bitstream at least include:

[0231] sps_isp_enabled_flag,

[0232] sps_mrl_enabled_flag,

[0233] sps_mip_enabled_flag,

[0234] sps_palette_enabled_flag and

[0235] sps_ibc_enabled_flag,

[0236] And in some embodiments, the plurality of sequentially arranged syntax elements associated with the intra prediction function do not include the sps_bcw_enabled_flag.

[0237] In some embodiments, the video decoder 30 receives a set of syntax elements associated with the inter prediction function from the bitstream after receiving the plurality of syntax elements associated with the intra prediction function.

[0238] In some embodiments, the plurality of syntax elements associated with the inter prediction function and sequentially arranged in the bitstream at least include:

[0239] sps_weighted_pred_flag,

[0240] sps_weighted_bipred_flag,

[0241] long_term_ref_pics_flag,

[0242] inter_layer_ref_pics_present_flag,

[0243] sps_idr_rpl_present_flag,

[0244] rpll_same_as_rpl0_flag,

[0245] sps_log2_diff_min_qt_min_cb_inter_slice,

[0246] sps_max_mtt_hierarchy_depth_inter_slice,

[0247] sps_ref_wraparound_enabled_flag,

[0248] sps_temporal_mvp_enabled_flag,

[0249] sps_amvr_enabled_flag,

[0250] sps_bdof_enabled_flag,

[0251] sps_smvd_enabled_flag,

[0252] sps_dmvr_enabled_flag,

[0253] sps_mmvd_enabled_flag,

[0254] six_minus_max_num_merge_cand,

[0255] sps_sbt_enabled_flag,

[0256] sps_affine_enabled_flag,

[0257] sps_bcw_enabled_flag,

[0258] sps_ciip_enabled_flag and

[0259] log2_parallel_merge_level_minus2

[0260] In some embodiments, the plurality of syntax elements associated with the merge mode and sequentially arranged in the bitstream at least include:

[0261] sps_mmvd_enabled_flag

[0262] six_minus_max_num_merge_cand,

[0263] sps_sbt_enabled_flag,

[0264] sps_affine_enabled_flag,

[0265] sps_bcw_enabled_flag,

[0266] sps_ciip_enabled_flag and

[0267] log2_parallel_merge_level_minus2.

[0268] The plurality of syntax elements associated with the merge mode are also associated with the inter prediction function.

[0269] In some embodiments, the video decoder 30 receives a set of syntax elements associated with the quantization function from the bitstream before receiving the plurality of syntax elements associated with the intra prediction function.

[0270] In some embodiments, the video decoder 30 receives a plurality of syntax elements at one or more levels of a picture parameter set (PPS) level and a slice level from a bitstream, wherein the plurality of syntax elements are associated with a predefined function and are sequentially arranged in the bitstream; based on a determination that the plurality of syntax elements satisfy a predefined condition, the video decoder 30: receives a second syntax element immediately following the plurality of syntax elements from the bitstream; based on a determination that the plurality of syntax elements do not satisfy the predefined condition, the video decoder 30: sets a default value as the second syntax element; and the video decoder 30 performs a predefined function on video data from the bitstream based on the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of a quantization function, an intra prediction function, and an inter prediction function.

[0271] In some embodiments, before the video decoder 30 receives a set of syntax elements associated with an inter prediction function, the video decoder 30 further receives a set of syntax elements associated with a quantization function.

[0272] In some embodiments, after the video decoder 30 receives a set of syntax elements associated with an inter prediction function, the video decoder 30 further receives a set of syntax elements associated with a quantization function.

[0273] In some embodiments, the plurality of syntax elements at the PPS level associated with an inter prediction function and sequentially arranged in the bitstream at least include:

[0274] rpl1_idx_present_flag,

[0275] pps_weighted_pred_flag,

[0276] pps_weighted_bipred_flag,

[0277] rpl_info_in_ph_flag and

[0278] pps_ref_wraparound_enabled_flag.

[0279] In some embodiments, the plurality of syntax elements at the PPS level include: pps_weighted_pred_flag, pps_weighted_bipred_flag, and rplinfo_in_ph_flag; the predefined condition is: (pps_weighted_pred_flag is true or pps_weighted_bipred_flag is true) and rpl_info_in_ph_flag is true; and the second syntax is wp_info_in_ph_flag.

[0280] The above method can be implemented using an apparatus including one or more circuits, where the one or more circuits include application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus can use these circuits in combination with other hardware or software components for performing the above method. Each module, sub-module, unit, or sub-unit disclosed above can be implemented at least in part using one or more circuits.

[0281] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to tangible media such as data storage media or a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the embodiments described in this application. A computer program product may include a computer-readable medium.

[0282] The terms used in the description of the embodiments herein are for the purpose of describing particular embodiments only and are not intended to limit the scope of the claims. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof.

[0283] It should also be understood that although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the embodiments, the first electrode may be referred to as the second electrode, and similarly, the second electrode may be referred to as the first electrode. The first electrode and the second electrode are both electrodes, but the first electrode and the second electrode are not the same electrode.

[0284] The description of the present application has been presented for purposes of illustration and description, and is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications, variations, and alternative embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description and the teachings presented in the associated drawings. The embodiments were chosen and described in order to best explain the principles of the invention, the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments and to best utilize the basic principles and various embodiments with various modifications suitable for the particular purposes contemplated. Accordingly, it should be understood that the scope of the claims should not be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the appended claims.

Claims

1. A method for decoding video data, comprising: Receiving a plurality of syntax elements at the sequence parameter set (SPS) level from a bitstream, wherein the plurality of syntax elements are associated with a predefined function and are arranged sequentially in the bitstream; Determining that at least one of the plurality of syntax elements satisfies a predefined condition; Receiving a second syntax element from the bitstream after the plurality of syntax elements; Determining that the at least one of the plurality of syntax elements does not satisfy the predefined condition; Setting the value of the second syntax element to a default value; and Performing the predefined function on video data from the bitstream according to at least one of the plurality of syntax elements and the second syntax element, wherein the predefined function is a predefined function selected from the group consisting of an intra prediction function, an inter prediction function, and a merge mode, and the plurality of syntax elements associated with the intra prediction function and sequentially arranged in the bitstream at least include: sps_isp_enabled_flag, sps_mrl_enabled_flag, sps_mip_enabled_flag, sps_palette_enabled_flag, and sps_ibc_enabled_flag, wherein the plurality of syntax elements associated with the intra prediction function and sequentially arranged do not include sps_bcw_enabled_flag, and there is no syntax element sps_mts_enabled_flag between the sps_mip_enabled_flag and the sps_palette_enabled_flag. The sps_isp_enabled_flag specifies whether intra prediction with sub-partitions is enabled. The sps_mrl_enabled_flag specifies whether intra prediction with multiple reference lines is enabled. The sps_mip_enabled_flag specifies whether matrix-based intra prediction is enabled. The sps_palette_enabled_flag specifies whether pred_mode_plt_flag exists in the codec unit syntax, and the pred_mode_plt_flag specifies using the Palette mode in the current coding unit. The sps_ibc_enabled_flag specifies whether the IBC prediction mode is used in decoding pictures in CLVS. The sps_bcw_enabled_flag specifies whether bi-directional prediction using CU weights can be used for inter prediction. The sps_mts_enabled_flag specifies whether sps_explicit_mts_intra_enabled_flag and sps_explicit_mts_inter_enabled_flag exist in the sequence parameter set RBSP syntax. The sps_explicit_mts_intra_enabled_flag specifies whether mts_idx exists in the intra codec unit syntax. The sps_explicit_mts_inter_enabled_flag specifies whether mts_idx exists in the inter codec unit syntax. The mts_idx specifies the transform kernel used in the current coding unit along the horizontal and vertical directions of the relevant luminance transform block.

2. The method according to claim 1, further comprising: After receiving the plurality of syntax elements associated with the intra prediction function, a set of syntax elements associated with the inter prediction function is received from the bitstream.

3. The method according to claim 1, wherein, The plurality of syntax elements associated with the inter prediction function and sequentially arranged in the bitstream at least includes: sps_weighted_pred_flag, sps_weighted_bipred_flag, long_term_ref_pics_flag, inter_layer_ref_pics_present_flag, sps_idr_rpl_present_flag, rpl1_same_as_rpl0_flag, sps_log2_diff_min_qt_min_cb_inter_slice, sps_max_mtt_hierarchy_depth_inter_slice, sps_ref_wraparound_enabled_flag, sps_temporal_mvp_enabled_flag, sps_amvr_enabled_flag, sps_bdof_enabled_flag, sps_smvd_enabled_flag, sps_dmvr_enabled_flag, sps_mmvd_enabled_flag, six_minus_max_num_merge_cand, sps_sbt_enabled_flag, sps_affine_enabled_flag, sps_bcw_enabled_flag, sps_ciip_enabled_flag and log2_parallel_merge_level_minus2, Among them, the sps_weighted_pred_flag specifies whether weighted prediction is applied to P slices that reference the SPS, the sps_weighted_bipred_flag specifies whether explicit weighted prediction is applied to B slices that reference the SPS, the long_term_ref_pics_flag specifies whether long-term reference pictures (LTRP) are used for inter-prediction of decoded pictures in CLVS, the inter_layer_ref_pics_present_flag specifies whether inter-layer reference pictures (ILRP) are used for inter-prediction of decoded pictures in CLVS, the sps_idr_rpl_present_flag specifies whether reference picture list syntax elements are present in the slice header of an IDR picture, the rpl1_same_as_rpl0_flag being equal to 1 specifies that the syntax elements num_ref_pic_lists_in_sps[1] and the syntax structure ref_pic_list_struct(1, rplsIdx) do not exist and the following applies: it is inferred that the value of num_ref_pic_lists_in_sps[1] is equal to the value of num_ref_pic_lists_in_sps[0]; it is inferred that the value of each syntax element in ref_pic_list_struct(1, rplsIdx) is equal to the corresponding syntax element in ref_pic_list_struct(0, rplsIdx), where rplsIdx ranges from 0 to num_ref_pic_lists_in_sps[0] - 1, num_ref_pic_lists_in_sps[i] specifies the number of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i included in the SPS, the sps_log2_diff_min_qt_min_cb_inter_slice specifies the default difference between the logarithm to the base 2 of the minimum size in luma samples of a luma leaf block resulting from quadtree partitioning of a CTU and the logarithm to the base 2 of the minimum luma decoded block size in luma samples of a luma CU in a slice (B) with slice_type equal to 0 or a slice (P) with slice_type equal to 1 that references the SPS, the sps_max_mtt_hierarchy_depth_inter_slice specifies the default maximum hierarchical depth of coding units resulting from multi-type tree partitioning of quadtree leaves in a slice (B) with slice_type equal to 0 or a slice (P) with slice_type equal to 1 that references the SPS, and the sps_ref_wraparound_enabled_flag specifies whether horizontal wraparound motion compensation is applied in inter-prediction.The sps_temporal_mvp_enabled_flag specifies whether to use the temporal motion vector prediction value in CLVS. The sps_amvr_enabled_flag specifies whether to use the adaptive motion vector difference resolution in motion vector encoding and decoding. The sps_bdof_enabled_flag specifies whether to enable bidirectional optical flow inter-frame prediction. The sps_smvd_enabled_flag specifies whether to use the symmetric motion vector difference in motion vector decoding. The sps_dmvr_enabled_flag specifies whether to enable inter-frame bidirectional prediction based on decoder motion vector refinement. The sps_mmvd_enabled_flag specifies whether to enable the merge mode with motion vector difference. The six_minus_max_num_merge_cand specifies the maximum number of merge motion vector prediction candidates supported in the SPS subtracted from 6. The sps_sbt_enabled_flag specifies whether to enable sub-block transformation of the CU for inter-frame prediction. The sps_affine_enabled_flag specifies whether motion compensation based on the affine model can be used for inter-frame prediction. The sps_bcw_enabled_flag specifies whether bidirectional prediction using CU weights can be used for inter-frame prediction. The sps_ciip_enabled_flag specifies whether the ciip_flag exists in the codec unit syntax of the inter-frame codec unit. The ciip_flag specifies whether to apply combined inter-frame merge and intra-frame prediction to the current coding unit. The log2_parallel_merge_level_minus2 plus 2 specifies the value of the variable Log2ParMrgLevel, which is used in the process of obtaining spatial merge candidates., 4. The method according to claim 1, wherein, The plurality of syntax elements associated with the merge mode and sequentially arranged in the bitstream at least includes: sps_mmvd_enabled_flag, six_minus_max_num_merge_cand, sps_sbt_enabled_flag, sps_affine_enabled_flag, sps_bcw_enabled_flag, sps_ciip_enabled_flag and log2_parallel_merge_level_minus2, Among them, the sps_mmvd_enabled_flag specifies whether to enable the merge mode with motion vector difference, the six_minus_max_num_merge_cand specifies the maximum number of merge motion vector prediction candidates supported in the SPS subtracted from 6, the sps_sbt_enabled_flag specifies whether to enable the sub-block transform of the CU for inter prediction, the sps_affine_enabled_flag specifies whether motion compensation based on the affine model can be used for inter prediction, the sps_bcw_enabled_flag specifies whether bi-directional prediction using CU weights can be used for inter prediction, the sps_ciip_enabled_flag specifies whether the ciip_flag exists in the codec unit syntax of the inter-coded unit, the ciip_flag specifies whether to apply combined inter-frame merge and intra prediction to the current coding unit, log2_parallel_merge_level_minus2 plus 2 specifies the value of the variable Log2ParMrgLevel, and the Log2ParMrgLevel is used in the process of obtaining spatial merge candidates.

5. The method according to claim 1, further comprising: Before receiving the plurality of syntax elements associated with the intra prediction function, a set of syntax elements associated with the quantization function is received from the bitstream.

6. A method for decoding video data, comprising: A plurality of syntax elements of one or more levels in the picture parameter set PPS level and the slice level are received from the bitstream, wherein the plurality of syntax elements are associated with a predefined function and are sequentially arranged in the bitstream; According to determining that at least one syntax element among the plurality of syntax elements satisfies a predefined condition: A second syntax element after the plurality of syntax elements is received from the bitstream; According to determining that the at least one syntax element among the plurality of syntax elements does not satisfy the predefined condition: Set the value of the second syntax element to a default value; and According to at least one syntax element among the plurality of syntax elements and the second syntax element, perform the predefined function on the video data from the bitstream, wherein the predefined function is a predefined function selected from the group consisting of a quantization function, an intra prediction function, and an inter prediction function; Among them, the method further includes: After receiving a set of syntax elements associated with the inter prediction function at the PPS level in the bitstream, a set of syntax elements associated with the quantization function at the PPS level in the bitstream is received. The set of syntax elements associated with the inter prediction function at the PPS level in the bitstream at least includes: rpl1_idx_present_flag, pps_weighted_pred_flag, pps_weighted_bipred_flag, and pps_ref_wraparound_enabled_flag Among them, the rpl1_idx_present_flag specifies whether rpl_sps_flag[1] and rpl_idx[1] exist in the PH syntax structure or the slice header of a picture that references the PPS. rpl_sps_flag[i] being equal to 1 specifies that RPLi in ref_pic_lists() is derived based on one of the syntax structures ref_pic_list_struct(listIdx, rplsIdx), where listIdx is equal to i in the SPS. The rpl_sps_flag[i] being equal to 0 specifies that RPL i in the picture is derived based on the syntax structure ref_pic_list_struct(listIdx, rplsIdx) directly included in the ref_pic_lists(), where listIdx is equal to i, and the ref_pic_lists() appears in the PH syntax structure or the slice header. The rpl_idx[i] specifies the index of the syntax structure ref_pic_list_struct(listIdx, rplsIdx) with listIdx equal to i in the list of syntax structures ref_pic_list_struct(listIdx, rplsIdx) with listIdx equal to i included in the SPS, which is used to derive RPL i of the current picture or slice. The pps_weighted_pred_flag specifies whether weighted prediction is applied to P slices that reference the PPS. The pps_weighted_bipred_flag specifies whether explicit weighted prediction is applied to B slices that reference the PPS. The pps_ref_wraparound_enabled_flag specifies whether horizontal wraparound motion compensation is applied in inter prediction.

7. The method according to claim 6, wherein, The multiple syntax elements at the PPS level that are associated with the inter prediction function and are sequentially arranged in the bitstream at least include: rpl1_idx_present_flag, pps_weighted_pred_flag, pps_weighted_bipred_flag, rpl_info_in_ph_flag and pps_ref_wraparound_enabled_flag, Among them, the rpl_info_in_ph_flag being equal to 1 specifies that the reference picture list information exists in the PH syntax structure and does not exist in the slice header of a picture that references the PPS and does not contain the PH syntax structure. The rpl_info_in_ph_flag being equal to 0 specifies that the reference picture list information does not exist in the PH syntax structure and can exist in the slice header of a picture that references the PPS and does not contain the PH syntax structure.

8. The method according to claim 6, wherein: The multiple syntax elements at the PPS level include: pps_weighted_pred_flag, pps_weighted_bipred_flag, and rpl_info_in_ph_flag; The predefined condition is: pps_weighted_pred_flag is true or pps_weighted_bipred_flag is true, and rpl_info_in_ph_flag is true; and The second syntax element is wp_info_in_ph_flag, wherein, rpl_info_in_ph_flag being equal to 1 specifies that reference picture list information exists in the PH syntax structure and does not exist in the slice header that references the PPS and does not contain the PH syntax structure, rpl_info_in_ph_flag being equal to 0 specifies that the reference picture list information does not exist in the PH syntax structure and can exist in the slice header that references the PPS and does not contain the PH syntax structure, wp_info_in_ph_flag being equal to 1 specifies that weighted prediction information can exist in the PH syntax structure and does not exist in the slice header that references the PPS and does not contain the PH syntax structure, wp_info_in_ph_flag being equal to 0 specifies that the weighted prediction information does not exist in the PH syntax structure and can exist in the slice header that references the PPS and does not contain the PH syntax structure.

9. An electronic device, comprising: One or more processing units; A memory, the memory being coupled to the one or more processing units; And Multiple programs stored in the memory, the multiple programs, when executed by the one or more processing units, cause the electronic device to perform the method according to any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium storing a plurality of programs for execution by an electronic device having one or more processing units, wherein, The multiple programs, when executed by the one or more processing units, cause the electronic device to perform the method according to any one of claims 1 to 8.

11. A computer program product comprising instructions for execution by a computing device having one or more processors, wherein when the instructions are executed by the one or more processors, the computing device is caused to perform the method according to any one of claims 1 to 8.