Signaling of syntax element of non-picture level in picture level

The method addresses the inefficiencies in video coding by signaling non-picture-level syntax elements at the picture level, reducing redundancy and bit waste, and enhancing compression efficiency.

JP2025093987AActive Publication Date: 2025-06-24HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025034451
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-09-24
Filing Date
2025-03-05
Publication Date
2025-06-24
Estimated Expiration
2040-09-23

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing and decompressing video data, leading to high bandwidth requirements and storage inefficiencies, particularly due to redundancy in signaling non-picture-level syntax elements across multiple slices.

Method used

The proposed solution involves signaling non-picture-level syntax elements at the picture level rather than in each slice header, using flags to indicate their presence in the picture header or slice headers. This approach ensures that syntax elements are only included once per picture if they remain constant across all slices, reducing redundancy and bit waste in the encoded bitstream.

Benefits of technology

By moving the signaling of non-picture-level syntax elements to the picture level, the method achieves reduced redundancy and fewer wasted bits in the encoded bitstream, thereby improving compression efficiency and reducing bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093987000001_ABST
    Figure 2025093987000001_ABST
Patent Text Reader

Abstract

To provide a method for more reducing a waste bit in a bit stream to be encoded by reducing redundancy by moving a signaling of a syntax element of a non-picture level in a picture level, and provide a video decorder, a video encoder, a program, and a computer-readable medium.SOLUTION: A method to be mounted by a video decoder, contains a step for receiving a video bit stream containing an adaptive loop filter (ALF) flag. An ALF flag that is equal to a first value designates that an ALF signaling is existed in a picture header (PH), is not existed in a slice header. The ALF flag that is equal to a second value designates that the ALF signaling is not existed in the PH, and may be existed in the slice header. The method also includes a step of obtaining a picture that the bit stream is decoded to obtain a picture to be decoded on the basis of the ALF flag.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application is a divisional application of Japanese Patent Application No. 2022 - 518802, filed on Mar. 23, 2022, which is a continuation of International Application No. PCT / US2020 / 052281, filed on Sep. 23, 2020, which claims priority to U.S. Provisional Patent Application No. 62 / 905,228, filed on Sep. 24, 2019, entitled "Signaling of Non - Picture - Level Syntax Elements in Picture Headers in Video Coding", by Futureway Technologies, Inc. The entire disclosure of the above - mentioned application is incorporated herein by reference.

[0002] Technical Field The disclosed embodiments generally relate to video coding, and more particularly, to signaling of non - picture - level syntax elements at the picture level.

Background Art

[0003] The amount of video data required to depict relatively short videos can be quite substantial, which can pose difficulties when the data is streamed or otherwise communicated over a communication network with limited bandwidth capacity. Thus, video data is generally compressed before being communicated over modern telecommunications networks. The size of the video can also be a problem when the video is stored in a storage device because memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent the digital video image. The compressed data is then received at the destination by a video decompression device that decodes the video data. Improved compression and decompression techniques that improve the compression ratio without sacrificing much or any of the image quality are desirable because network resources are limited and the demand for higher video quality is constantly increasing.

Summary of the Invention

Means for Solving the Problems

[0004] A first aspect is a method implemented by a video decoder, the method including steps of: receiving, by the video decoder, a video bitstream including an RPL flag, wherein the RPL flag equal to a first value specifies that RPL signaling is present within a PH, and the RPL flag equal to a second value specifies that RPL signaling is not present within the PH and may be present within a slice header.

[0005] In an embodiment, syntax elements are included in the picture header if they are the same, and are included in the slice header if there are changes to those syntax elements. However, in some embodiments, the syntax elements may not be included in both. First, non-picture level syntax elements may be present in the PH. Non-picture level syntax elements are syntax elements at a level of the video bitstream other than the picture level. Second, for each category of non-picture level syntax elements, a flag specifies when the syntax elements of that category are present in the PH or the slice header. The flag may be within the PH. Non-picture level syntax elements include RPL signaling, the co-located CbCr flag, SAO tool enable and parameters, ALF tool enable and parameters, LMCS tool enable and parameters, and scaling list tool enable and parameters. Third, if non-picture level syntax elements are present in the PH, the corresponding syntax elements are not present in any slice of the picture associated with the picture header that contains those syntax elements. The values of the non-picture level syntax elements present in the PH apply to all slices of the picture associated with the picture header that contains those syntax elements. Fourth, if non-picture level syntax elements are not present in the PH, the corresponding syntax elements may be present in the slice header of a slice of the picture associated with the picture header. By moving the signaling of non-picture level syntax elements to the picture level, redundancy is reduced and there are fewer wasted bits in the encoded bitstream.

[0006] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that the RPL signaling is present in the PH.

[0007] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that the RPL signaling is not within the slice.

[0008] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that RPL signaling does not exist within the PH.

[0009] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that RPL signaling may exist within the slice header.

[0010] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the bitstream further includes an RPL SPS flag, where the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i in the SPS, or that RPL i is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx directly included and equal to i.

[0011] Optionally, in any of the above aspects, another implementation of this aspect provides that the bitstream further includes an RPL index, where the RPL index specifies the index into the list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i, which is included in the sequence parameter set (SPS) and is used for the derivation of RPL i of the current picture.

[0012] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the method further includes displaying the decoded picture on a display of an electronic device.

[0013] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that when non-picture-level syntax elements are present in the PH, the corresponding syntax elements do not exist in any slice of the picture associated with the PH that contains those syntax elements.

[0014] A second aspect is a method implemented by a video encoder, the method including the step of generating an RPL flag, where an RPL flag equal to a first value specifies that RPL signaling is present in the PH, and an RPL flag equal to a second value specifies that RPL signaling is not present in the PH and may be present in the slice header.

[0015] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that RPL signaling is present in the PH.

[0016] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that RPL signaling is not within the slice.

[0017] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that RPL signaling is not present in the PH.

[0018] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that RPL signaling may be present in the slice header.

[0019] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the method further includes generating an RPL SPS flag, where the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i in the SPS, or that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx directly equal to i.

[0020] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the method further includes generating an RPL index, where the RPL index specifies an index into a list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i, which is included in the sequence parameter set (SPS) and is used for deriving RPL i of the current picture.

[0021] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that when non-picture level syntax elements are present in the PH, the corresponding syntax elements are not present in any slice of the picture associated with the PH that contains those syntax elements.

[0022] A third aspect is a method implemented by a video decoder, the method comprising: receiving, by the video decoder, a video bitstream including an SAO flag, wherein an SAO flag equal to a first value specifies that SAO signaling is present within a PH, and an SAO flag equal to a second value specifies that SAO signaling is not present within a PH and may be present within a slice header; and decoding, by the video decoder, a picture encoded using the SAO flag and obtaining a decoded picture.

[0023] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the method further includes displaying the decoded picture on a display of an electronic device.

[0024] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that when non-picture-level syntax elements are present in a PH, the corresponding syntax elements are not present in any slice of a picture associated with the PH that includes those syntax elements.

[0025] A fourth aspect is a method implemented by a video encoder, the method comprising: generating an SAO flag, wherein an SAO flag equal to a first value specifies that SAO signaling is present within a PH, and an SAO flag equal to a second value specifies that SAO signaling is not present within a PH and may be present within a slice header; encoding, by the video encoder, the SAO flag into a video bitstream; and storing, by the video encoder, the video bitstream for communication to a video decoder.

[0026] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that when non-picture level syntax elements are present in the PH, the corresponding syntax elements are not present in any slice of the picture associated with the PH containing those syntax elements.

[0027] A fifth aspect is a method implemented by a video decoder, the method including: receiving, by the video decoder, a video bitstream including an ALF flag, wherein the ALF flag equal to a first value specifies that ALF signaling is present in the PH, and the ALF flag equal to a second value specifies that ALF signaling is not present in the PH and may be present in a slice header; and decoding, by the video decoder, a picture encoded using the ALF flag to obtain a decoded picture.

[0028] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the method further includes displaying the decoded picture on a display of an electronic device.

[0029] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that when non-picture level syntax elements are present in the PH, the corresponding syntax elements are not present in any slice of the picture associated with the PH containing the syntax elements.

[0030] A sixth aspect is a method implemented by a video encoder, the method including: generating an ALF flag, where an ALF flag equal to a first value specifies that ALF signaling is present within a PH, and an ALF flag equal to a second value specifies that ALF signaling is not present within the PH and may be present within a slice header; encoding, by the video encoder, the ALF flag into a video bitstream; and storing, by the video encoder, the video bitstream for communication to a video decoder.

[0031] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that when a non-picture-level syntax element is present in a PH, the corresponding syntax element is not present in any slice of the picture associated with the PH that contains the syntax element.

[0032] A seventh aspect is a method implemented by a video decoder, the method including: receiving, by the video decoder, a video bitstream that includes a syntax element, where the syntax element specifies whether information may or may not be present within a PH, or whether the information may or may not be present within a slice header; and decoding, by the video decoder, an encoded picture using the syntax element to obtain a decoded picture.

[0033] Any of the above embodiments may be combined with any of the other above embodiments to create a new embodiment. These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and the claims.

Brief Description of the Drawings

[0034] To more fully understand the present disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, in which like reference numerals represent like parts.

[0035]

Figure 1

[0036]

Figure 2

[0037]

Figure 3

[0038]

Figure 4

[0039]

Figure 5

[0040]

Figure 6

[0041]

Figure 7

[0042]

Figure 8

[0043]

Figure 9

[0044]

Figure 10

[0045]

Figure 11

[0046]

Figure 12

[0047]

Figure 13

Embodiments for Carrying Out the Invention

[0048] Exemplary implementations of one or more embodiments are provided below, but it should first be understood that the disclosed system and / or method can be implemented using any number of techniques, whether currently known or existing. This disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and the full scope of equivalents.

[0049] The following abbreviations apply: ALF: adaptive loop filter APS: adaptation parameter set ASIC: application-specific integrated circuit AU: access unit AUD: access unit delimiter BT: binary tree CABAC: context - adaptive binary arithmetic coding (context - adaptive binary arithmetic coding) CAVLC: context - adaptive variable - length coding (context - adaptive variable - length coding) Cb: blue difference chroma (blue difference chroma) CLVS: coded layer - wise video sequence (coded layer - wise video sequence) CLVS: coded layer video sequence (coded layer video sequence) CPU: central processing unit (central processing unit) Cr: red difference chroma (red difference chroma) CRA: clean random access (clean random access) CTB: coding tree block (coding tree block) CTU: coding tree unit (coding tree unit) CU: coding unit (coding unit) CVS: coded video sequence (coded video sequence) DC: direct current (direct current) DCI: decoding capability information (decoding capability information) DCT: discrete cosine transform (discrete cosine transform) DMM: depth modeling mode (depth modeling mode) DPB: decoded picture buffer (decoded picture buffer) DPS: decoding parameter set (decoding parameter set) DSP: digital signal processor (Digital Signal Processor) DST: discrete sine transform (Discrete Sine Transform) EO: electrical-to-optical (Electro-Optical) FPGA: field-programmable gate array (Field-Programmable Gate Array) GDR: gradual decoding refresh (Gradual Decoding Refresh) HEVC: High Efficiency Video Coding (High Efficiency Video Coding) ID: identifier (Identifier) IDR: instantaneous decoding refresh (Instantaneous Decoding Refresh) IEC: International Electrotechnical Commission (International Electrotechnical Commission) I / O: input / output (Input / Output) IRAP: intra random access pictures (Intra Random Access Pictures) ISO: International Organization for Standardization (International Organization for Standardization) ITU: International Telecommunication Union (International Telecommunication Union) ITU-T: ITU Telecommunication Standardization Sector (ITU Telecommunication Standardization Sector) LMCS: luma mapping with chroma scaling (Luma Mapping with Chroma Scaling) LTRP: long-term reference picture (Long-Term Reference Picture) MVP: motion vector predictor (Motion Vector Predictor) NAL: network abstraction layer (network abstraction layer) OE: optical-to-electrical (optical to electrical) PH: picture header (picture header) PIPE: probability interval partitioning entropy (probability interval partitioning entropy) POC: picture order count (picture order count) PPS: picture parameter set (picture parameter set) PU: picture unit (picture unit) QT: quad tree (quad tree) RADL: random access decodable leading (random access decodable leading) RAM: random-access memory (random access memory) RASL: random access skipped leading (random access skipped leading) RBSP: raw byte sequence payload (raw byte sequence payload) RDO: rate-distortion optimization (rate-distortion optimization) ROM: read-only memory (read-only memory) RPL: reference picture list (reference picture list) Rx: receiver unit (receiver unit) SAD: sum of absolute differences (sum of absolute differences) SAO: sample adaptive offset (sample adaptive offset) SBAC: syntax-based arithmetic coding (syntax-based arithmetic coding) SOP: sequence of pictures (sequence of pictures) SPS: sequence parameter set (sequence parameter set) SRAM: static RAM (static RAM) SSD: sum of squared differences (sum of squared differences) TCAM: ternary content-addressable memory (ternary content-addressable memory) TT: triple tree (triple tree) TU: transform unit (transform unit) Tx: transmitter unit (transmitter unit) VCL: video coding layer (video coding layer) VPS: video parameter set (video parameter set) VVC: Versatile Video Coding (Versatile Video Coding)

[0050] Unless modified otherwise herein, the following definitions apply: A bitstream is a sequence of bits containing video data compressed for transmission between an encoder and a decoder. An encoder is a device that uses an encoding process to compress video data into a bitstream. A decoder is a device that uses a decoding process to reconstruct video data from a bitstream for display. A picture is an array of luma samples or chroma samples that make up a frame or field. A picture being encoded or decoded may be referred to as the current picture. A reference picture contains reference samples that can be used when coding other pictures by reference according to inter prediction or inter-layer prediction. A reference picture list is a list of reference pictures used for inter prediction or inter-layer prediction. A flag is a variable or 1-bit syntax element that can take on one of two possible values (0 or 1). Some video coding systems utilize two reference picture lists, which may be denoted as reference picture list 1 and reference picture list 0. A reference picture list structure is an addressable syntax structure that contains multiple reference picture lists. Inter prediction is a mechanism for coding samples of the current picture by reference to indicated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer. A reference picture list structure entry is an addressable position within a reference picture list structure that indicates a reference picture associated with a reference picture list. A slice header is part of a coded slice and contains data elements relevant to all the video data within the tile represented within the slice. A PPS contains data relevant to an entire picture. More specifically, a PPS is a syntax structure that contains syntax elements applicable to zero or more coded pictures, determined by syntax elements found in each picture header. An SPS contains data relevant to a sequence of pictures.An AU is a set of one or more coded pictures associated with the same presentation time (e.g., the same picture order count) for output from the DPB (e.g., for presentation to a user). An AUD indicates the start of an AU or the boundary between AUs. A decoded video sequence is a sequence of pictures reconstructed by a decoder in preparation for presentation to a user.

[0051] FIG. 1 is a flowchart of an exemplary operational method 100 for coding a video signal. Specifically, the video signal is encoded in an encoder. The encoding process compresses the video signal by employing various mechanisms to reduce the video file size. A smaller file size enables the transmission of the compressed video file to a user, while reducing the associated bandwidth overhead. The decoder then decodes the compressed video file to reconstruct the original video signal for display to an end user. The decoding process generally mirrors the encoding process to enable the decoder to consistently reconstruct the video signal.

[0052] In step 101, a video signal is input to the encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device such as a video camera and encoded to support video live streaming. The video file can include both an audio component and a video component. The video component includes a series of image frames that give the impression of visual motion when viewed in sequence. The frames include pixels represented using a luminance (or luminance sample) and a color (or color sample) referred to herein as a chroma component. In some examples, the frames may also include depth values to support three-dimensional viewing.

[0053] In step 103, the video is partitioned into blocks. Partitioning includes subdividing the pixels within each frame into square and / or rectangular blocks for compression. For example, in HEVC, a frame can first be divided into coding tree units (CTUs) which are blocks of a predefined size (e.g., 64 pixels × 64 pixels). A CTU contains both luma samples and chroma samples. Using a coding tree, the CTU can be divided into blocks and the blocks can be recursively subdivided until a configuration that supports further encoding is achieved. For example, the luma component of a frame may be subdivided until the individual blocks contain relatively uniform illumination values. Additionally, the chroma component of a frame may be subdivided until the individual blocks contain relatively uniform color values. Thus, the partitioning mechanism varies depending on the content of the video frame.

[0054] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, inter prediction and / or intra prediction may be used. Inter prediction is designed to utilize the fact that objects within a common scene tend to appear in consecutive frames. Thus, blocks depicting objects in the reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a fixed position over multiple frames. Thus, the table is described once and adjacent frames can refer to the previous reference frame. A pattern matching mechanism can be used to match objects over multiple frames. Further, a moving object may be represented over multiple frames, for example, due to the movement of the object or the movement of the camera. As a specific example, a video may show a car moving across the screen over multiple frames. To describe such movement, motion vectors can be used. A motion vector is a two-dimensional vector that provides an offset from the coordinates of an object in a frame to the coordinates of the object in the reference frame. Thus, inter prediction can encode an image block in the current frame as a set of motion vectors indicating the offsets from the corresponding blocks in the reference frame.

[0055] Intra prediction encodes blocks within a common frame. Intra prediction exploits the fact that luma and chroma components tend to cluster in a frame. For example, a green patch in a part of a tree tends to be located adjacent to similar green patches. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), a planar mode, and a DC mode. The directional mode indicates that the current block is similar / same as the samples of neighboring blocks in the corresponding direction. The planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the end of the row. The planar mode effectively shows a smooth transition of brightness / color across the row / column by using a relatively constant slope when changing values. The DC mode is used for boundary smoothing and indicates that the block is similar / same as the average value associated with the samples of all neighboring blocks associated with the angular direction of the directional prediction mode. Thus, an intra prediction block can represent an image block as various relational prediction mode values instead of actual values. Further, an inter prediction block can represent an image block as a motion vector value instead of actual values. In either case, the prediction block may not accurately represent the image block in some cases. If there is a difference, it is stored in a residual block. A transform may be applied to the residual block to further compress the file.

[0056] In step 107, various filtering techniques can be applied. In HEVC, the filters are applied according to the in-loop filtering method. The above-mentioned block-based prediction may lead to the generation of block-shaped images in the decoder. Further, the block-based prediction method can encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering method sequentially and iteratively applies a noise reduction filter, a deblocking filter, an adaptive loop filter, and an SAO filter to blocks / frames. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. Further, these filters mitigate artifacts in the reconstructed reference block, so that the artifacts are less likely to generate additional artifacts in subsequent blocks encoded based on the reconstructed reference block.

[0057] Once the video signal is partitioned, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the above-mentioned data and any signaling data desirable to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission towards the decoder upon request. The bitstream may also be broadcast and / or multicast towards multiple decoders. The generation of the bitstream is a sequential and iterative process. Thus, steps 101, 103, 105, 107, and 109 can occur continuously and / or simultaneously over many frames and blocks. The order shown in FIG. 1 is shown for clarity and ease of discussion and is not intended to limit the video coding process to a particular order.

[0058] The decoder receives the bitstream in step 111 and starts the decoding process. Specifically, the decoder employs an entropy decoding method that converts the bitstream into the corresponding syntax and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the partition for the frame. The partition splitting should match the result of the block partition splitting in step 103. The entropy encoding / decoding used in step 111 is described herein. The encoder makes many selections during the compression process, such as selecting a block partition splitting method from several possible choices based on the spatial positioning of the values within the input image(s). Accurate signaling of the selection may use a large number of bins. As used herein, a bin is a binary value treated as a variable (e.g., a bit value that can vary depending on the context). Entropy coding enables the encoder to discard any options that are clearly not realistic for a particular case and leave a set of acceptable options. A codeword is then assigned to each acceptable option. The length of the codeword is based on the number of acceptable options (e.g., 1 bin for 2 options, 2 bins for 3 to 4 options, etc.). The encoder then encodes the codeword for the selected option. This method reduces the size of the codeword because the codeword uniquely indicates a selection from a small subset of acceptable options rather than from a potentially large set of all possible options. The decoder then decodes the selection in a similar manner as the encoder by determining the set of acceptable options. By determining the set of acceptable options, the decoder can read the codeword and determine the selection made by the encoder.

[0059] In step 113, the decoder performs block decoding. Specifically, the decoder uses inverse transformation to generate residual blocks. Then, the decoder uses the residual blocks and the corresponding prediction blocks to reconstruct the image blocks according to the partition splitting. The prediction blocks may include both intra-prediction blocks and inter-prediction blocks generated by the encoder in step 105. Then, the reconstructed image blocks are positioned within the frame of the reconstructed video signal according to the partition splitting data determined in step 111. The syntax for step 113 may also be signaled within the bitstream by entropy coding as described above.

[0060] In step 115, filtering is performed on the frame of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frame to remove blocking artifacts. Once the frame is filtered, the video signal can be output to the display in step 117 for viewing by the end user.

[0061] Figure 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, the codec system 200 provides functionality to support the implementation of the operation method 100. The codec system 200 is generalized to show components used in both the encoder and the decoder. The codec system 200 receives and partitions a video signal as described with respect to steps 101 and 103 of the operation method 100, thereby obtaining a partitioned video signal 201. Next, when acting as an encoder, the codec system 200 compresses the partitioned video signal 201 into a coded bitstream as described with respect to steps 105, 107, and 109 of method 100. When acting as a decoder, the codec system 200 generates an output video signal from the bitstream as described with respect to steps 111, 113, 115, and 117 in the operation method 100. The codec system 200 includes a general-purpose coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and CABAC component 231. Such components are coupled as shown. In Figure 2, the solid lines indicate the movement of data to be encoded / decoded, and the dashed lines indicate the movement of control data that controls the operation of other components. All components of the codec system 200 may be present within the encoder. The decoder may include a subset of the components of the codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. Here, these components will be described.

[0062] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to subdivide the blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into even smaller blocks. The blocks may be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is called the depth of the node / coding tree. The partitioned blocks can optionally be included in a CU. For example, a CU may be a sub-part of a CTU that includes a luma block, a Cr block, and a Cb block, along with the corresponding syntax instructions for that CU. The partitioning modes may include BT, TT, and QT, which are used to partition a node into two, three, or four child nodes of different shapes depending on the partitioning mode used. The partitioned video signal 201 is transferred to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.

[0063] The general coder control component 211 is configured to make decisions related to coding the images of a video sequence into a bitstream according to the constraints of the application. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size versus reconstructed quality. Such decisions may be made based on memory space / bandwidth availability and image resolution requirements. The general coder control component 211 also manages the utilization of the buffer in light of the transmission speed to mitigate buffer underrun and overrun problems. To manage these problems, the general coder control component 211 manages the partitioning, prediction, and filtering by other components. For example, the general-purpose coder control component 211 may dynamically increase the complexity of compression to increase the resolution and increase the use of bandwidth, or may decrease the complexity of compression to decrease the resolution and the use of bandwidth. Thus, the general coder control component 211 controls other components of the codec system 200 to balance the video signal reconstructed quality with the concern of bitrate. The general coder control component 211 generates control data for controlling the operation of other components. The control data is also transferred to the header formatter and the CABAC component 231 for being encoded in the bitstream to signal the parameters for decoding at the decoder.

[0064] The partitioned video signal 201 is also sent to the motion estimation component 221 and the motion compensation component 219 for inter prediction. A frame or slice of the partitioned video signal 201 may be divided into a plurality of video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter prediction coding of the received video block with respect to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 can perform a plurality of coding passes, for example, to select an appropriate coding mode for each block of video data.

[0065] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation component 221 is a process that generates motion vectors, which estimate the motion for video blocks. The motion vectors can indicate, for example, the displacement of the object to be coded with respect to the prediction block. The prediction block is a block that is found to closely match the block to be coded in terms of pixel differences. The prediction block may also be referred to as a reference block. Such pixel differences can be determined by SAD, SSD, or other difference metrics. HEVC uses several coded objects including CTUs, CTBs, and CUs. For example, a CTU can be divided into CTBs, and a CTB can be divided into CUs for inclusion in CUs. A CU can be encoded as a prediction unit containing prediction data and / or a TU containing transformed residual data for the CU. The motion estimation component 221 generates motion vectors, prediction units, and TUs by using rate-distortion analysis as part of the rate-distortion optimization process. For example, the motion estimation component 221 can determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and select the reference block, motion vector, etc. having the best rate-distortion characteristics. The best rate-distortion characteristics balance both the quality of video reconstruction (e.g., the amount of data loss due to compression) and the coding efficiency (e.g., the size of the final encoding).

[0066] In some examples, codec system 200 can calculate the values of sub-integer pixel positions of reference pictures stored in decode picture buffer component 223. For example, video codec system 200 can interpolate the values of 1 / 4 pixel positions, 1 / 8 pixel positions, or other fractional pixel positions of a reference picture. Accordingly, motion estimation component 221 can perform motion search for both full pixel positions and fractional pixel positions and output motion vectors with fractional pixel accuracy. Motion estimation component 221 calculates the motion vector for a prediction unit of a video block in an inter-coded slice by comparing the position of the prediction unit with the position of the prediction block of the reference picture.

[0067] Motion estimation component 221 outputs the calculated motion vector as motion data to header format and CABAC component 231 for encoding and outputs the motion to motion compensation component 219. Motion compensation performed by motion compensation component 219 may include fetching or generating a prediction block based on the motion vector determined by motion estimation component 221. Here too, in some examples, motion estimation component 221 and motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the current video block's prediction unit, motion compensation component 219 can identify the prediction block pointed to by the motion vector. Then, a residual video block is formed by subtracting the pixel values of the prediction block from the pixel values of the currently coded video block to form pixel difference values. Generally, motion estimation component 221 performs motion estimation for the luma component, and motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The prediction block and the residual block are transferred to transform scaling and quantization component 213.

[0068] The partitioned video signal 201 is also sent to the intra picture estimation component 215 and the intra picture prediction component 217. Similar to the motion estimation component 221 and the motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated but are shown separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 perform intra prediction of the current block for a block within the current frame instead of inter prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames as described above. In particular, the intra picture estimation component 215 determines the intra prediction mode to be used to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra prediction mode to encode the current block from a plurality of tested intra picture prediction modes. The selected intra prediction mode is then transferred to the header format and the CABAC component 231 for encoding.

[0069] For example, the intra picture estimation component 215 calculates rate - distortion values using rate - distortion analysis for various tested intra picture prediction modes, and selects an intra prediction mode having the best rate - distortion characteristics among the tested modes. Rate - distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for various encoded blocks, and determines which intra prediction mode exhibits the best rate - distortion value for that block. Additionally, the intra picture estimation component 215 may be configured to code depth blocks of the depth map using a DMM based on RDO.

[0070] The intra picture prediction component 217 can generate a residual block from a prediction block based on the selected intra picture prediction mode determined by the intra picture estimation component 215 when implemented in an encoder, and can read a residual block from a bitstream when implemented in a decoder. The residual block includes the difference in values between the prediction block and the original block represented as a matrix. The residual block is then transferred to the transform scaling and quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 can operate on both luma and chroma components.

[0071] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a DCT, DST, or a conceptually similar transform, to the residual block to generate a video block containing residual transform coefficient values. Wavelet transforms, integer transforms, subband transforms, or other types of transforms can also be used. The transform can transform the residual information from the pixel value domain to a transform domain, such as a frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling includes applying a scale factor to the residual information so that different frequency information is quantized at different granularities, which can affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be adjusted by adjusting the quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients. The quantized transform coefficients are transferred to the header format and the CABAC component 231 for encoding in the bitstream.

[0072] The scaling and inverse transformation component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transformation component 229 applies inverse scaling, transformation, and / or quantization to reconstruct the residual block in the pixel domain. For example, it is for later use as a reference block that can be a predicted block for another current block. The motion estimation component 221 and / or the motion compensation component 219 can calculate the reference block by adding the residual block to the corresponding predicted block for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference block to mitigate artifacts generated during scaling, quantization, and transformation. Otherwise, such artifacts can cause inaccurate predictions (and create additional artifacts) when subsequent blocks are predicted.

[0073] The filter control analysis component 227 and the in-loop filter component 225 apply a filter to the residual block and / or the reconstructed image block. For example, the transformed residual block from the scaling and inverse transform component 229 can be combined with the corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct the original image block. Then, the filter can be applied to the reconstructed image block. In some examples, the filter may alternatively be applied to the residual block. Similar to the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filter applied to the reconstructed reference block is applied to a specific spatial region and includes a plurality of parameters to adjust how such a filter is applied. The filter control analysis component 227 analyzes the reconstructed reference block to determine where such a filter should be applied and sets the corresponding parameters. Such data is transferred as filter control data to the header format and the CABAC component 231 for encoding. The in-loop filter component 225 applies such a filter based on the filter control data. The filter may include a deblocking filter, a noise suppression filter, an SAO filter, and an adaptive loop filter. Such a filter may be applied, depending on the example, in the spatial / pixel domain (e.g., the reconstructed pixel block) or in the frequency domain.

[0074] When operating as an encoder, the filtered reconstructed image block, residual block, and / or prediction block are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores the reconstructed and filtered blocks as part of the output video signal and transfers them towards the display. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.

[0075] The header format and CABAC component 231 receive data from various components of the codec system 200 and encode such data into a coded bitstream for transmission to the decoder. Specifically, the header format and CABAC component 231 generate various headers for encoding control data such as general control data and filter control data. Further, prediction data including intra prediction and motion data, as well as residual data in the form of quantized transform coefficient data, are all encoded into the bitstream. The final bitstream contains all the information desired by the decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra prediction modes, indications of partition information, etc. Such data may be encoded by using entropy coding. For example, the information may be encoded by using CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. Following entropy coding, the coded bitstream may be transmitted to another device (e.g., a video decoder) or archived for later transmission or retrieval.

[0076] Figure 3 is a block diagram showing an exemplary video encoder 300. The video encoder 300 may be used to implement the encoding function of the codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of the operation method 100. The encoder 300 partitions the input video signal, and as a result, a partitioned video signal 301 is obtained, which is substantially similar to the partitioned video signal 201. Next, the partitioned video signal 301 is compressed by components of the encoder 300 and encoded into a bitstream.

[0077] Specifically, the partitioned video signal 301 is transferred to the intra-picture prediction component 317 for intra prediction. The intra-picture prediction component 317 may be substantially similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. The partitioned video signal 301 is also transferred to the motion compensation component 321 for inter prediction based on the reference blocks in the decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction blocks and residual blocks from the intra-picture prediction component 317 and the motion compensation component 321 are transferred to the transform and quantization component 313 for transformation and quantization of the residual blocks. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual blocks and the corresponding prediction blocks (along with the relevant control data) are transferred to the entropy coding component 331 for coding into the bitstream. The entropy coding component 331 may be substantially similar to the header format and CABAC component 231.

[0078] The transformed and quantized residual blocks and / or corresponding prediction blocks are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 to be reconstructed into reference blocks for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. The in-loop filter within the in-loop filter component 325 is also applied to the residual blocks and / or the reconstructed reference blocks, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters, as discussed with respect to the in-loop filter component 225. The filtered blocks are then stored in the decoded picture buffer component 323 for use as reference blocks by the motion compensation component 321. The decoded picture buffer component 323 may be substantially the same as the decoded picture buffer component 223.

[0079] FIG. 4 is a block diagram showing an exemplary video decoder 400. The video decoder 400 may be used to implement the decoding function of the codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of the method of operation 100. The decoder 400 receives, for example, a bitstream from the encoder 300 and generates a reconstructed output video signal based on the bitstream for display to an end user.

[0080] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding method such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can use the header information to provide a context for interpreting the additional data encoded as codewords within the bitstream. The decoded information includes any desired information for decoding a video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from residual blocks. The quantized transform coefficients are transferred to the inverse transform and quantization component 429 for reconstruction into the residual blocks. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.

[0081] The reconstructed residual block(s) and / or prediction block(s) are transferred to the intra-picture prediction component 417 for reconstruction of the picture block based on the intra prediction operation. The intra-picture prediction component 417 may be similar to the intra-picture estimation component 215 and the intra-picture prediction component 217. Specifically, the intra-picture prediction component 417 uses a prediction mode to locate a reference block within the frame, and applies the residual block to the result to reconstruct the intra-predicted picture block. The reconstructed intra-predicted picture block(s) and / or residual block(s) and the corresponding inter-prediction data are transferred to the decoded picture buffer component 423 via the in-loop filter component 425. These may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225 respectively. The in-loop filter component 425 filters the reconstructed picture block(s), residual block(s) and / or prediction block(s), and such information is stored in the decoded picture buffer component 423. The reconstructed picture block from the decoded picture buffer component 423 is transferred to the motion compensation component 421 for inter prediction. The motion compensation component 421 may be substantially the same as the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using the motion vector from the reference block, and applies the residual block to the result to reconstruct the picture block. The resulting reconstructed block may be transferred to the decoded picture buffer component 423 via the in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed picture blocks that can be reconstructed into the frame via the partition information. Such a frame may be placed in the sequence. This sequence is output to the display as the reconstructed output video signal.

[0082] FIG. 5 shows an embodiment of a video bitstream 500. The video bitstream 500 may also be referred to as a coded video bitstream, a bitstream, or a variation thereof. The bitstream 500 includes at least one PU 501. Although three PUs 501 are shown in FIG. 5, in practical applications, different numbers of PUs 501 may be present in the bitstream 500. Each PU 501 is associated with each other according to a specified classification rule, is consecutive in decoding order, and is a set of NAL units including exactly one coded picture (for example, picture 514). In one embodiment, each PU 501 has a temporal ID 519 or is associated with the temporal ID 519.

[0083] In one embodiment, each PU 501 includes one or more of DCI 502, VPS 504, SPS 506, PPS 508, PH 512, and picture 514. Each of DCI 502, VPS 504, SPS 506, and PPS 508 may be collectively referred to as a parameter set. Other parameter sets not shown in FIG. 5, such as APS, may be included in the bitstream 500. APS is a syntax structure including syntax elements applied to zero or more slices determined by zero or more syntax elements found in the slice header 520.

[0084] DCI 502, sometimes also called DPS, is a syntax structure that contains syntax elements applied to the entire bitstream. DCI 502 contains parameters that remain constant over the lifetime of a video bitstream (e.g., bitstream 500) that can be converted over the lifetime of a session. DCI 502 can contain profile, level, and sub-profile information to determine the maximum interoperability point of complexity that is guaranteed never to be exceeded, even if splicing of the video sequence occurs within the session. It can further optionally contain constraint flags. The flags indicate that the video bitstream will be restricted in the use of certain functions indicated by their values. Thereby, the bitstream can be labeled as not using certain tools that allow resource allocation in decoder implementation, etc. Similar to all parameter sets, DCI 502 exists when first referenced and is referenced by the first picture of the video sequence. This implies that DCI 502 must be sent within the first few NAL units of the bitstream. There may be multiple DCI 502s within the bitstream, but the values of the syntax elements among them must not conflict when referenced.

[0085] VPS 504 contains decoding dependencies or information for the construction of reference picture sets in the enhancement layer. It provides an overall perspective or view of the scalable sequence, and it contains some other high-level characteristics of the bitstream that can be used as a basis for, such as what types of operation points are provided, the profile, layer, and level of those operation points, as well as session negotiation and content selection.

[0086] The SPS 506 contains data common to all pictures within the SOP. The SPS 506 is a syntax structure that includes syntax elements applicable to zero or more entire CLVS as determined by the content of the syntax elements found in the PPS that are referenced by the syntax elements found in each picture header. In contrast, the PPS 508 contains data common to the entire picture. The PPS 508 is a syntax structure that includes syntax elements applicable to zero or more entire coded pictures as determined by the syntax elements found in each picture header (e.g., PH 512).

[0087] The DCI 502, VPS 504, SPS 506, and PPS 508 are contained in different types of NAL units. The NAL unit is a syntax structure that includes an indication of the type of data to follow (e.g., coded video data). The NAL units are classified into VCL NAL units and non-VCL NAL units. The VCL NAL units contain data representing the values of samples within a video picture, and the non-VCL NAL units contain any relevant additional information such as parameter sets (important data applicable to several VCL NAL units) and supplementary enhancement information (timing information and other supplementary data that may enhance the usefulness of the decoded video signal but is not necessary for decoding the values of samples within the video picture).

[0088] The PH 512 is a syntax structure that includes syntax elements applicable to all slices (e.g., slice 518) of a coded picture (e.g., picture 514). In one embodiment, the PH 512 is within a new type of non-VCL NAL unit designated as a PH NAL unit. Thus, the PH NAL unit has a PH NUT (e.g., PH_NUT). In one embodiment, each PU 501 contains only one PH 512. That is, the PU 501 contains a single or isolated PH 512. In one embodiment, exactly one PH NAL unit exists for each picture 501 within the bitstream 500.

[0089] Picture 514 is an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats. In certain embodiments, each PU 501 contains only one picture 514. Thus, within each PU 501, there is only one PH 512 and only one picture 514 corresponding to that PH 512. That is, PU 501 contains a single or isolated picture 514.

[0090] Each picture 514 includes one or more slices 518. A slice 518 is an integer number of complete tiles, or an integer number of consecutive complete CTU rows within a tile of a picture (e.g., picture 514). Each slice 518 is exclusively contained within a single NAL unit (e.g., a VCL NAL unit). In some embodiments, a single NAL unit is associated with or has a layer ID 515. A tile (not shown) is a rectangular region of CTUs within a particular tile column and a particular tile row within a picture (e.g., picture 514). Tiles are partitioned portions of a picture created by horizontal and vertical boundaries. Tiles may be rectangular and / or square. Specifically, a tile includes four sides connected at right angles. The four sides include two pairs of parallel sides. Further, the sides of each pair of parallel sides are of equal length. Thus, a tile may be of any rectangular shape, and a square is a special case of a rectangle where all four sides are of equal length. An image / picture can include one or more tiles. A CTU (not shown) is a CTB of luma samples of a picture having three sample arrays, two corresponding CTBs of chroma samples, or a CTB of samples of a monochrome image or picture coded using three separate color planes and syntax structures for coding the samples. A CTB (not shown) is an N×N block of samples for some value of N such that the partitioning of a component into CTBs is a partition. A block (not shown) is an M×N (M columns × N rows) array of samples (e.g., pixels) or an M×N array of transform coefficients.

[0091] Pictures 514 and their slices 518 contain data related to the image or video to be encoded or decoded. Thus, pictures 514 and their slices 518 can simply be referred to as the payload or data carried within bitstream 500. PH 512 and slice headers 520 may include a flag 522. The flag 522 may be an RPL flag, an SAO flag, or an ALF flag, as will be described later.

[0092] The VVC specification defines only a few syntax elements at the picture level. However, when commonly used, there are additional syntax elements that may have different values between slices of the same picture, but usually have the same value for all slices of the same picture. Examples of such syntax elements are those related to RPL, the unified chroma coding flag, the SAO enable flag, the ALF enable flag and parameters, the LMCS enable flag and parameters, and the scaling list enable flag and parameters. Such non-picture-level syntax elements are not signaled in the PH and must be repeated in all slice headers of those slices, even if they have the same value for all slices of the same picture. In other words, in some approaches, these syntax elements are signaled in the slice header. This is because the data it carries may vary from slice to slice. However, in most cases, it is the same for the entire picture containing the slice. As a result, these elements are signaled several times per picture, but the values are generally the same, which is redundant and wastes bits in the encoded bitstream.

[0093] Disclosed herein are embodiments for signaling non-picture-level syntax at the picture level. In an embodiment, a syntax element is included in a picture header if the syntax elements are the same, and is included in a slice header if there is a change in the syntax elements. However, in some embodiments, the syntax elements may not be included in both. First, non-picture-level syntax elements may be present in the PH. Non-picture-level syntax elements are syntax elements at a level of a video bitstream other than the picture level. Second, for each category of non-picture-level syntax elements, a flag designates when the syntax elements of that category are present in the PH or the slice header. The flag may be within the PH. Non-picture-level syntax elements include those related to the signaling of RPL, the common CbCr sign flag, SAO tool enable and parameters, ALF tool enable and parameters, LMCS tool enable and parameters, and scaling list tool enable and parameters. Third, if non-picture-level syntax elements are present in the PH, the corresponding syntax elements do not exist in any slice of the picture related to the picture header containing those syntax elements. The values of the non-picture-level syntax elements present in the PH are applied to all slices of the picture associated with the picture header containing those syntax elements. Fourth, if non-picture-level syntax elements are not present in the PH, the corresponding syntax elements may be present in the slice header of a slice of the picture related to the picture header. By moving the signaling of non-picture-level syntax elements to the picture level, redundancy is reduced and there are fewer wasted bits in the encoded bitstream.

Table 1-1

Table 1-2

Table 1-3

Table 2-1

Table 2-2

Table 2-3

Table 2-4

Table 2-5

Table 2-6

[0094] Semantics of PH RBSP PH contains common information for all slices of the coded picture such that the next VCL NAL unit in decode order is the first slice to be coded.

[0095] For a given value of pic_type, pic_type indicates the characterization of the coded picture listed in Table 1. The value of pic_type is equal to 0 to 5 (inclusive) in a bitstream compliant with this version of this specification. Other values of pic_type are reserved for future use by ITU-T ISO / IEC. A decoder compliant with this version of this specification shall ignore reserved values of pic_type.

Table 3

[0096] pic_parameter_set_id specifies the value of pps_pic_parameter_set_id of the PPS in use. The value of pic_parameter_set_id is in the range of 0 to 63 (inclusive).

[0097] That non_reference_picture_flag is equal to 1 specifies that the picture associated with the PH is never used as a reference picture. That non_reference_picture_flag is equal to 0 specifies that the picture may or may not be used as a reference picture.

[0098] colour_plane_id specifies the color plane associated with the picture associated with the PH when separate_colour_plane_flag is equal to 1. The value of colour_plane_id is within the range of 0 to 2 (both ends included). The colour_plane_id values 0, 1, and 2 correspond to the Y, Cb, and Cr planes respectively. There is no dependency during the decoding process of pictures with different colour_plane_id values.

[0099] pic_order_cnt_lsb specifies the picture order count for the picture associated with the PH, with MaxPicOrderCntLsb as the modulus. The length of the pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits. The value of pic_order_cnt_lsb is within the range of 0 to MaxPicOrderCntLsb - 1 (both ends included).

[0100] The recovery_poc_cnt specifies the recovery points of the decoded pictures in the output order. In the CVS, if there is a picture picA after the current GDR picture in the decoding order and its PicOrderCntVal is equal to the value obtained by adding the value of recovery_poc_cnt to the PicOrderCntVal of the current GDR picture, then the picture picA is called a recovery point picture. Otherwise, the first picture in the output order with a PicOrderCntVal greater than the value obtained by adding the value of recovery_poc_cnt to the PicOrderCntVal of the current picture is called the recovery point picture. The recovery point picture does not precede the current GDR picture in the decoding order. The value of recovery_poc_cnt is within the range of 0 to MaxPicOrderCntLsb - 1 (both ends included).

[0101] The variable RpPicOrderCntVal is derived as RpPicOrderCntVal = PicOrderCntVal + recovery_poc_cnt.

[0102] The no_output_of_prior_pics_flag affects the output of the previously decoded pictures in the decoded picture buffer after decoding a non - first CLVSS picture in the bitstream as specified in Annex C.

[0103] The pic_output_flag affects the decoded picture output and removal process as defined in Annex C. If the pic_output_flag does not exist, it is presumed to be equal to 1.

[0104] That pic_rpl_present_flag is equal to 1 specifies that the RPL signaling exists within the PH. That pic_rpl_present_flag is equal to 0 specifies that the RPL signaling does not exist within the PH and may exist in the slice header of the slice of the picture. If it does not exist, the value of pic_rpl_present_flag is presumed to be equal to 0. The RPL signaling is the RPL information included in the video bitstream 500.

[0105] That pic_rpl_sps_flag[i] is equal to 1 specifies that the RPL i of the picture is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures. Here, listIdx is equal to i within the SPS. That ref_pic_list_sps_flag[i] is equal to 0 specifies that the reference picture list i of the picture is derived based on the ref_pic_list_struct(listIdx,rplsIdx) syntax structure. Here, listIdx is equal to i directly included in the picture header.

[0106] If pic_rpl_sps_flag[i] does not exist, the following applies: If num_ref_pic_lists_in_sps[i] is equal to 0, the value of pic_rpl_sps_flag[i] is presumed to be equal to 0. Otherwise (where num_ref_pic_lists_in_sps[i] is greater than 0), if rpl1_idx_present_flag is equal to 0, the value of pic_rpl_sps_flag[1] is presumed to be equal to pic_rpl_sps_flag[0]. Otherwise, the value of pic_rpl_sps_flag[i] is presumed to be equal to pps_ref_pic_list_sps_idc[i] - 1.

[0107] pic_rpl_idx[i] specifies the index into the list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i, which is used in the derivation of RPL i of the current picture and is included in the SPS. The syntax element pic_rpl_idx[i] is represented by Ceil(Log2(num_ref_pic_lists_in_sps[i])) bits. If it does not exist, the value of pic_rpl_idx[i] is assumed to be equal to 0. The value of pic_rpl_idx[i] is in the range of 0 to num_ref_pic_lists_in_sps[i] - 1 (inclusive). If pic_rpl_sps_flag[i] is equal to 1 and num_ref_pic_lists_in_sps[i] is equal to 1, the value of pic_rpl_idx[i] is assumed to be equal to 0. If pic_rpl_sps_flag[i] is equal to 1 and rpl1_idx_present_flag is equal to 0, the value of pic_rpl_idx[1] is assumed to be equal to pic_rpl_idx[0].

[0108] The variable PicRplsIdx[i] is derived as follows: PicRplsIdx[i] = pic_rpl_sps_flag[i]? pic_rpl_idx[i] : num_ref_pic_lists_in_sps[i].

[0109] pic_poc_lsb_lt[i][j] specifies the value of the picture order count modulo MaxPicOrderCntLsb for the j-th LTRP entry in the i-th reference picture list for the picture associated with PH. The length of the pic_poc_lsb_lt[i][j] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.

[0110] The variable PicPocLsbLt[i][j] is derived as follows: PicPocLsbLt[i][j] = ltrp_in_slice_header_flag[i][PicRplsIdx[i]]? pic_poc_lsb_lt[i][j] : rpls_poc_lsb_lt[listIdx][PicRplsIdx[i]][j].

[0111] That pic_delta_poc_msb_present_flag[i][j] is equal to 1 specifies that pic_delta_poc_msb_cycle_lt[i][j] exists. That pic_delta_poc_msb_present_flag[i][j] is equal to 0 specifies that pic_delta_poc_msb_cycle_lt[i][j] does not exist.

[0112] Let prevTid0Pic be the previous picture in decoding order that has the same nuh_layer_id as PH, has a TemporalId equal to 0, and is not a RASL or RADL picture. Let setOfPrevPocVals be the set consisting of: the PicOrderCntVal of prevTid0Pic; the PicOrderCntVal of each picture that is referenced by an entry in RefPicList[0] or RefPicList[1] of prevTid0Pic and has the same nuh_layer_id as the current picture; and the PicOrderCntVal of each picture that is after prevTid0Pic in decoding order, has the same nuh_layer_id as the current picture, and precedes the current picture in decoding order.

[0113] If there are two or more values in setOfPrevPocVals such that the value modulo MaxPicOrderCntLsb is equal to PicPocLsLt[i][j], the value of pic_delta_poc_msb_present_flag[i][j] is equal to 1.

[0114] pic_delta_poc_msb_cycle_lt[i][j] specifies the value of the variable PicFullPocLt[i][j] as follows:

Table 4

[0115] The value of pic_delta_poc_msb_cycle_lt[i][j] is in the range of 0 to 2(32 - log2_max_pic_order_cnt_lsb_minus4 - 4) (inclusive). If it does not exist, the value of pic_delta_poc_msb_cycle_lt[i][j] is assumed to be 0.

[0116] pic_temporal_mvp_enabled_flag specifies whether the temporal MVP can be used for inter prediction. If pic_temporal_mvp_enabled_flag is equal to 0, the syntax elements of the picture related to the picture header are constrained so that the temporal MVP is not used in the decoding of the picture. Otherwise (if pic_temporal_mvp_enabled_flag is equal to 1), the temporal MVP may be used in the decoding of the picture.

[0117] If pic_temporal_mvp_enabled_flag does not exist, the following applies: If sps_temporal_mvp_enabled_flag is equal to 0, the value of pic_temporal_mvp_enabled_flag is assumed to be 0. Otherwise (if sps_temporal_mvp_enabled_flag is equal to 1), the value of pps_temporal_mvp_enabled_flag is assumed to be equal to pps_temporal_mvp_enabled_idc - 1.

[0118] The fact that pic_level_joint_cbcr_sign_flag is equal to 1 specifies that slice_joint_cbcr_sign_flag does not exist in the slice header. The fact that pic_level_joint_cbcr_sign_flag is equal to 0 specifies that slice_joint_cbcr_sign_flag exists in the slice header. If it does not exist, the value of pic_level_joint_cbcr_sign_flag is presumed to be equal to 0.

[0119] The fact that pic_level_alf_enabled_flag is equal to 1 specifies that ALF is enabled for all slices belonging to the picture associated with PH and can be applied to the Y, Cb, or Cr color components within the slice. The fact that pic_level_alf_enabled_flag is equal to 0 specifies that ALF can be disabled for one or more, or all slices belonging to the picture associated with PH. If it does not exist, pic_level_alf_enabled_flag is presumed to be equal to 0.

[0120] pic_num_alf_aps_ids_luma specifies the number of ALF APSs that the slices belonging to the picture associated with PH refer to. The value of slice_num_alf_aps_ids_luma is within the range of 0 to 7 (inclusive).

[0121] pic_alf_aps_id_luma[i] specifies the adaptation_parameter_set_id of the i-th ALF APS that the luma component of the slice of the picture associated with PH refers to.

[0122] That pic_alf_chroma_idc is equal to 0 specifies that ALF is not applied to the Cb and Cr color components. That pic_alf_chroma_idc is equal to 1 indicates that ALF is applied to the Cb color component. That pic_alf_chroma_idc is equal to 2 indicates that ALF is applied to the Cr color component. That pic_alf_chroma_idc is equal to 3 indicates that ALF is applied to the Cb and Cr color components. If pic_alf_chroma_idc does not exist, it is presumed to be equal to 0.

[0123] pic_alf_aps_id_chroma specifies the adaptation_parameter_set_id of the ALF APS that the chroma components of the slices of the picture associated with the picture header refer to.

[0124] That pic_level_lmcs_enabled_flag is equal to 1 specifies that for all slices belonging to the picture associated with the picture header, the luma mapping with chroma scaling is enabled. That pic_level_lmcs_enabled_flag is equal to 0 specifies that for one, or more, or all slices belonging to the picture associated with the picture header, the luma mapping with chroma scaling may be disabled. If it does not exist, the value of pic_level_lmcs_enabled_flag is presumed to be equal to 0.

[0125] pic_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS that the slices of the picture associated with the picture header refer to.

[0126] The fact that pic_chroma_residual_scale_flag is equal to 1 specifies that chroma residual scaling is enabled for all slices belonging to the picture associated with the picture header. The fact that pic_chroma_residual_scale_flag is equal to 0 specifies that chroma residual scaling may be disabled for one, or more, or all slices belonging to the picture associated with the picture header. If pic_chroma_residual_scale_flag does not exist, it is assumed to be equal to 0.

[0127] The fact that pic_level_scaling_list_present_flag is equal to 1 specifies that the scaling list data used for the slices of the picture associated with the picture header is derived based on the scaling list data included in the referenced scaling list APS. The fact that pic_level_scaling_list_present_flag is equal to 0 specifies that the scaling list data used for one, or more, or all slices of the picture associated with the picture header is the default scaling list data derived and specified in Section 7.4.3.16. If it does not exist, the value of pic_level_scaling_list_present_flag is assumed to be equal to 0.

[0128] pic_scaling_list_aps_id specifies the adaptation_parameter_set_id of the scaling list APS.

[0129] Semantics of the slice header RBSP That slice_rpl_sps_flag[i] is equal to 1 specifies that the RPL i of the current slice is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i within the SPS. That slice_rpl_sps_flag[i] is equal to 0 specifies that the RPL i of the current slice is derived based on the ref_pic_list_struct(listIdx, rplsIdx) syntax structure with listIdx equal to i directly included in the slice header of the current picture.

[0130] If slice_rpl_sps_flag[i] does not exist, the following applies: If pic_rpl_present_flag is equal to 1, the value of slice_rpl_sps_flag[i] is assumed to be equal to pic_rpl_sps_flag[i]. Otherwise, if num_ref_pic_lists_in_sps[i] is equal to 0, the value of slice_rpl_sps_flag[i] is assumed to be equal to 0. Otherwise (num_ref_pic_lists_in_sps[i] is greater than 0), if rpl1_idx_present_flag is equal to 0, the value of slice_rpl_sps_flag[1] is assumed to be equal to slice_rpl_sps_flag[0]. Otherwise, the value of slice_rpl_sps_flag[i] is assumed to be equal to pps_ref_pic_list_sps_idc[i] - 1.

[0131] slice_rpl_idx[i] specifies the index into the list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i that are used in the derivation of the current slice's reference picture list i in the SPS. The syntax element slice_rpl_idx[i] is represented by Ceil(Log2(num_ref_pic_lists_in_sps[i])) bits. The value of slice_rpl_idx[i] is in the range of 0 to num_ref_pic_lists_in_sps[i]-1 (inclusive). If slice_rpl_sps_flag[i] is equal to 1 and num_ref_pic_lists_in_sps[i] is equal to 1, the value of slice_rpl_idx[i] is assumed to be 0. If slice_rpl_sps_flag[i] is equal to 1 and rpl1_idx_present_flag is equal to 0, the value of slice_rpl_idx[1] is assumed to be equal to slice_rpl_idx[0].

[0132] The variable RplsIdx[i] is derived as follows:

Table 5

[0133] slice_poc_lsb_lt[i][j] specifies the value of the picture order count modulo MaxPicOrderCntLsb for the j-th LTRP entry in the i-th reference picture list for the current slice. The length of the slice_poc_lsb_lt[i][j] syntax element is log2_max_pic_order_cnt_lsb_minus4 + 4 bits.

[0134] The variable PocLsbLt[i][j] is derived as follows:

Table 6

[0135] The fact that slice_delta_poc_msb_present_flag[i][j] is equal to 1 specifies that slice_delta_poc_msb_cycle_lt[i][j] exists. The fact that slice_delta_poc_msb_present_flag[i][j] is equal to 0 specifies that slice_delta_poc_msb_cycle_lt[i][j] does not exist.

[0136] Let prevTid0Pic be the previous picture in decoding order that has the same nuh_layer_id as the current picture, has a TemporalId equal to 0, and is not a RASL or RADL picture. Let setOfPrevPocVals be the set consisting of: the PicOrderCntVal of prevTid0Pic, the PicOrderCntVal of each picture that is referenced by an entry in RefPicList[0] or RefPicList[1] of prevTid0Pic and has the same nuh_layer_id as the current picture, and the PicOrderCntVal of each picture that is after prevTid0Pic in decoding order, has the same nuh_layer_id as the current picture, and precedes the current picture in decoding order.

[0137] When pic_rpl_present_flag is equal to 0 and there are two or more values in setOfPrevPocVals such that the value with MaxPicOrderCntLsb as the modulus is equal to PocLsbLt[i][j], the value of slice_delta_poc_msb_present_flag[i][j] is equal to 1.

[0138] slice_delta_poc_msb_cycle_lt[i][j] specifies the value of the variable FullPocLt[i][j] as follows: [Table 7]

[0139] The value of slice_delta_poc_msb_cycle_lt[i][j] is within the range of 0 to 2(32 - log2_max_pic_order_cnt_lsb_minus4 - 4) (inclusive). If it does not exist, the value of slice_delta_poc_msb_cycle_lt[i][j] is assumed to be equal to 0.

[0140] slice_joint_cbcr_sign_flag specifies whether the co - located residual samples of both chroma components have inverted signs in the transform unit where tu_joint_cbcr_residual_flag[x0][y0] is equal to 1. If tu_joint_cbcr_residual_flag[x0][y0] is equal to 1 for a certain transform unit, slice_joint_cbcr_sign_flag being equal to 0 specifies that the sign of each residual sample of the Cr (or Cb) component is the same as the sign of the co - located Cb (or Cr) residual sample, and slice_joint_cbcr_sign_flag being equal to 1 specifies that the sign of each residual sample of the Cr (or Cb) component is given by the inverted sign of the co - located Cb (or Cr) residual sample. If it does not exist, the value of slice_joint_cbcr_sign_flag is assumed to be equal to pic_level_joint_cbcr_sign_flag.

[0141] slice_sao_luma_flag being equal to 1 specifies that SAO is enabled for the luma component in the current slice. slice_sao_luma_flag being equal to 0 specifies that SAO is disabled for the luma component in the current slice. If slice_sao_luma_flag does not exist, it is assumed to be equal to pic_level_sao_luma_enabled_flag.

[0142] When slice_sao_chroma_flag is equal to 1, it specifies that SAO is enabled for the chroma components in the current slice. When slice_sao_chroma_flag is equal to 0, it specifies that SAO is disabled for the chroma components in the current slice. If slice_sao_chroma_flag does not exist, it is assumed to be equal to pic_level_sao_chroma_enabled_flag.

[0143] When slice_alf_enabled_flag is equal to 1, it specifies that ALF is enabled and may be applied to the Y, Cb, or Cr color components within the slice. When slice_alf_enabled_flag is equal to 0, it specifies that ALF is disabled for all color components within the slice. If it does not exist, the value of slice_alf_enabled_flag is assumed to be equal to pic_level_alf_enabled_flag.

[0144] slice_num_alf_aps_ids_luma specifies the number of ALF APSs referenced by the slice. When slice_alf_enabled_flag is equal to 1 and slice_num_alf_aps_ids_luma does not exist, the value of slice_num_alf_aps_ids_luma is assumed to be equal to the value of pic_num_alf_aps_ids_luma. The value of slice_num_alf_aps_ids_luma is in the range of 0 to 7 (inclusive).

[0145] slice_alf_aps_id_luma[i] specifies the adaptation_parameter_set_id of the i-th ALF APS that the luma component of the slice refers to. The TemporalId of the APS NAL unit where aps_params_type is equal to ALF_APS and adapdation_parameter_set_id is equal to slice_alf_aps_id_luma[i] is less than or equal to the TemporalId of the slice NAL unit being coded. When slice_alf_enabled_flag is equal to 1 and slice_alf_aps_id_luma[i] does not exist, the value of slice_alf_aps_id_luma[i] is presumed to be equal to the value of pic_alf_aps_id_luma[i].

[0146] For slices within an intra-slice and IRAP picture, slice_alf_aps_id_luma[i] shall not refer to an ALF APS associated with a picture other than the picture containing that intra-slice or IRAP picture.

[0147] slice_alf_chroma_idc being equal to 0 specifies that ALF is not applied to the Cb and Cr color components. slice_alf_chroma_idc being equal to 1 indicates that ALF is applied to the Cb color component. slice_alf_chroma_idc being equal to 2 indicates that ALF is applied to the Cr color component. slice_alf_chroma_idc being equal to 3 indicates that ALF is applied to the Cb and Cr color components. When slice_alf_chroma_idc does not exist, it is presumed to be equal to pic_alf_chroma_idc.

[0148] slice_alf_aps_id_chroma specifies the adaptation_parameter_set_id of the ALF APS that the chroma component of the slice refers to. The TemporalId of the APS NAL unit where aps_params_type is equal to ALF_APS and adapdation_parameter_set_id is equal to slice_alf_aps_id_chroma is less than or equal to the TemporalId of the slice NAL unit to be coded. When slice_alf_enabled_flag is equal to 1 and slice_alf_aps_id_chroma does not exist, the value of slice_alf_aps_id_chroma is assumed to be equal to the value of pic_alf_aps_id_chroma.

[0149] For slices in intra slices and IRAP pictures, slice_alf_aps_id_chroma shall not refer to the ALF APS associated with another picture, but rather to the picture that contains that intra slice or IRAP picture.

[0150] When slice_lmcs_enabled_flag is equal to 1, it specifies that the luma mapping with chroma scaling is enabled for the current slice. When slice_lmcs_enabled_flag is equal to 0, it specifies that the luma mapping with chroma scaling is not enabled for the current slice. When slice_lmcs_enabled_flag does not exist, it is assumed to be equal to pic_lmcs_enabled_flag.

[0151] The slice_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS that the slice refers to. The TemporalId of an APS NAL unit where aps_params_type is equal to LMCS_APS and adapdation_parameter_set_id is equal to slice_lmcs_aps_id is less than or equal to the TemporalId of the coded slice NAL unit. If slice_lmcs_enabled_flag is equal to 1 and slice_lmcs_aps_id does not exist, the value of slice_lmcs_aps_id is assumed to be equal to the value of pic_lmcs_aps_id.

[0152] If it exists, the value of slice_lmcs_aps_id is the same for all slices of the picture.

[0153] That slice_chroma_residual_scale_flag is equal to 1 specifies that chroma residual scaling is enabled for the current slice. That slice_chroma_residual_scale_flag is equal to 0 specifies that chroma residual scaling is not enabled for the current slice. If slice_chroma_residual_scale_flag does not exist, it is assumed to be equal to pic_chroma_residual_scale_flag.

[0154] The fact that slice_scaling_list_present_flag is equal to 1 specifies that the scaling list data used for the current slice is derived based on the scaling list data included in the referenced scaling list APS. The fact that slice_scaling_list_present_flag is equal to 0 specifies that the scaling list data used for the current picture is the default scaling list data specified in Section 7.4.3.16. If not present, the value of slice_scaling_list_present_flag is assumed to be equal to pic_level_scaling_list_present_flag.

[0155] slice_scaling_list_aps_id specifies the adaptation_parameter_set_id of the scaling list APS. The TemporalId of the APS NAL unit where aps_params_type is equal to SCALING_APS and adaptation_parameter_set_id is equal to slice_scaling_list_aps_id is less than or equal to the TemporalId of the coded slice NAL unit. If slice_scaling_list_enabled_flag is equal to 1 and slice_scaling_list_aps_id does not exist, the value of slice_scaling_list_aps_id is assumed to be equal to the value of pic_scaling_list_aps_id.

[0156] FIG. 6 is a flowchart showing a method 600 for decoding a bitstream according to the first embodiment. The decoder 400 may implement the method 600. In step 610, a video bitstream including an RPL flag is received. The RPL flag equal to the first value specifies that RPL signaling exists in the PH. The RPL flag equal to the second value specifies that RPL signaling does not exist in the PH and may exist in the slice header. Finally, in step 620, the coded picture is decoded using the RPL flag to obtain a decoded picture.

[0157] The method 600 may implement additional embodiments. For example, the first value is 1. The second value is 0. The bitstream further includes an RPL SPS flag, and the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i in the SPS, or is directly included and is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i. The bitstream further includes an RPL index, and the RPL index specifies the index into the list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i in the SPS, which is used for deriving the RPL i of the current picture. The decoded picture is displayed on the display of the electronic device.

[0158] FIG. 7 is a flowchart showing a method 700 for encoding a bitstream according to the first embodiment. The encoder 300 may implement the method 700. In step 710, an RPL flag is generated. The RPL flag equal to the first value specifies that RPL signaling is present in the PH. The RPL flag equal to the second value specifies that RPL signaling is not present in the PH and may be present in the slice header. Finally, in step 730, the video bitstream is stored for communication to the video decoder.

[0159] The method 700 may implement additional embodiments. For example, the first value is 1. The second value is 0. An RPL SPS flag is generated. Here, the RPL SPS flag specifies that the RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i in the SPS, or is directly included and derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i. An RPL index is generated. Here, the RPL index specifies the index into the list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures having a listIdx equal to i, which is included in the sequence parameter set (SPS) and is used for the derivation of the RPL i of the current picture.

[0160] FIG. 8 is a flowchart showing a method 800 for decoding a bitstream according to a second embodiment. Decoder 400 may implement method 800. At step 810, a video bitstream including an SAO flag is received. An SAO flag equal to a first value specifies that SAO signaling is present in the PH, and an SAO flag equal to a second value specifies that SAO signaling is not present in the PH and may be present in the slice header. Finally, at step 820, the coded picture is decoded using the SAO flag to obtain a decoded picture. Method 800 may implement additional embodiments. For example, the decoded picture may be displayed on a display of an electronic device.

[0161] FIG. 9 is a flowchart showing a method 900 for encoding a bitstream according to a second embodiment. Encoder 300 may implement method 900. At step 910, an SAO flag is generated. An SAO flag equal to a first value specifies that SAO signaling is present in the PH, and an SAO flag equal to a second value specifies that SAO signaling is not present in the PH and may be present in the slice header. At step 920, the RPL flag is encoded into the video bitstream. Finally, at step 930, the video bitstream is stored for communication to a video decoder.

[0162] FIG. 10 is a flowchart showing a method 1000 for decoding a bitstream according to a third embodiment. The decoder 400 may implement the method 1000. In step 1010, a video bitstream including an ALF flag is received. An ALF flag equal to a first value specifies that ALF signaling is present in the PH, and an ALF flag equal to a second value specifies that ALF signaling is not present in the PH and may be present in the slice header. Finally, in step 1020, the coded picture is decoded using the ALF flag, and a decoded picture is obtained. The method 1000 may implement additional embodiments. For example, the decoded picture may be displayed on a display of an electronic device.

[0163] FIG. 11 is a flowchart showing a method 1100 for encoding a bitstream according to a third embodiment. The encoder 300 may implement the method 1100. In step 1110, an ALF flag is generated. An ALF flag equal to a first value specifies that ALF signaling is present in the PH, and an ALF flag equal to a second value specifies that ALF signaling is not present in the PH and may be present in the slice header. In step 1120, the ALF flag is encoded into the video bitstream. Finally, in step 1130, the video bitstream is stored for communication to a video decoder.

[0164] FIG. 12 is a schematic diagram of a video coding apparatus 1200 (e.g., video encoder 300 or video decoder 400) according to an embodiment of the present disclosure. The video coding apparatus 1200 is suitable for implementing the disclosed embodiments. The video coding apparatus 1200 includes an input port 1210 and Rx 1220 for receiving data; a processor, logic unit, or CPU 1230 for processing the data; a Tx 1240 and an output port 1250 for transmitting the data; and a memory 1260 for storing the data. The video coding apparatus 1200 may also include OE components and EO components coupled to the input port 1210, the receiver unit 1220, the transmitter unit 1240, and the output port 1250 for the input and output of optical or electrical signals.

[0165] The processor 1230 is implemented by hardware and software. The processor 1230 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 1230 communicates with the input port 1210, Rx 120, Tx 1240, output port 1250, and memory 1260. The processor 1230 includes a coding module 1270. The coding module 1270 implements the disclosed embodiments. For example, the coding module 1270 implements, processes, prepares, or provides various codec functions. Thus, including the coding module 1270 provides a substantial improvement to the functionality of the video coding apparatus 1200 and performs a conversion of the video coding apparatus 1200 to different states. Alternatively, the coding module 1270 may be implemented as instructions stored in the memory 1260 and executed by the processor 1230.

[0166] In addition, the video coding device 1200 may include an input / output device 1280 for communicating data with a user. The input / output device 1280 may include output devices such as a display for displaying video data, a speaker for outputting audio data, etc. The input / output device 1280 may also include an input device such as a keyboard, a mouse, or a trackball, or a corresponding interface for interacting with such output devices.

[0167] The memory 1260 includes one or more disks, tape drives, and solid state drives, and may be used as an overflow data storage device, store a program when such a program is selected for execution, and store instructions and data read during program execution. The memory 1260 may be volatile and / or non-volatile, and may be ROM, RAM, TCAM, or SRAM.

[0168] FIG. 13 is a schematic diagram of an embodiment of the coding means 1300. In one embodiment, the coding means 1300 is implemented in a video coding device 1302 (e.g., a video encoder 300 or a video decoder 400). The video coding device 1302 includes receiving means 1301. The receiving means 1301 is configured to receive a picture to be encoded or a bitstream to be decoded.

[0169] The video coding device 1302 includes a transmission means 1307 coupled to a reception means 1301. The transmission means 1307 is configured to transmit a bitstream to a decoder or to transmit a decoded image to a display means (for example, one of the I / O devices 1280). The video coding device 1302 includes a storage means 1303. The storage means 1303 is coupled to at least one of the reception means 1301 or the transmission means 1307. The storage means 1303 is configured to store instructions. Further, the video coding device 1302 also includes a processing means 1305. The processing means 1305 is coupled to the storage means 1303. The processing means 1305 is configured to execute the instructions stored in the storage means 1303 in order to execute the methods disclosed herein.

[0170] In one embodiment, the reception means receives a video bitstream including an RPL flag. The RPL flag specifies whether RPL signaling is present or not present during PH, or specifies that RPL signaling may be present in the slice header. The processing means decodes the picture coded using the RPL flag and obtains the decoded picture.

[0171] The term "about" means a range including ±10% of the subsequent number unless otherwise specified. Although several embodiments are provided in the present disclosure, it can be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The examples of the present application are considered to be illustrative rather than restrictive, and the intention is not limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0172] Furthermore, techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined with or integrated into other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other items shown or described as being combined may be directly coupled or may communicate indirectly through any interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and modifications can be envisioned by those skilled in the art and can be made without departing from the spirit and scope disclosed herein.

Claims

1. 13. A method implemented by a video decoder, comprising: receiving, by the video decoder, a video bitstream including a Reference Picture List (RPL) flag, the RPL flag equal to a first value specifying that RPL signaling is present in a picture header (PH) and the RPL flag equal to a second value specifying that RPL signaling is not present in the PH and may be present in a slice header; and decoding, by the video decoder, the coded picture using the RPL flag to obtain a decoded picture. method.

2. The method of claim 1 , wherein the first value is one.

3. The method of claim 1 or 2, wherein the second value is 0.

4. The bitstream further includes an RPL Sequence Parameter Set (SPS) flag, the RPL SPS flag specifying that RPL i is derived based on one of the ref_pic_list_struct (listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS or based on one of the directly included ref_pic_list_struct (listIdx, rplsIdx) syntax structures with listIdx equal to i.

4. The method according to any one of claims 1 to 3.

5. 5. The method of claim 1, wherein the bitstream further comprises an RPL index, the RPL index specifying an index into a list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i contained in a Sequence Parameter Set (SPS) of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i that are used to derive the RPL i of the current picture.

6. 6. The method according to claim 1, further comprising the step of displaying the decoded picture on a display of the electronic device.

7. 7. The method of claim 1, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.

8. 13. A method implemented by a video encoder, comprising: generating a Reference Picture List (RPL) flag, the RPL flag equal to a first value specifying that RPL signaling is present in a Picture Header (PH) and the RPL flag equal to a second value specifying that RPL signaling is not present in the PH and may be present in a slice header; encoding, by the video encoder, the RPL flag into a video bitstream; storing, by the video encoder, the video bitstream for communication to a video decoder. method.

9. The method of claim 8 , wherein the first value is one.

10. 10. The method of claim 8 or 9, wherein the second value is 0.

11. generating an RPL sequence parameter set (SPS) flag, the RPL SPS flag specifying that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS or based on one of the directly contained ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i; 11. The method according to any one of claims 8 to 10.

12. 12. A method according to claim 8, further comprising the step of generating an RPL index, the RPL index specifying an index into a list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i contained in a Sequence Parameter Set (SPS) of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i used to derive the RPL i of the current picture.

13. 13. A method according to claim 8, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains that syntax element.

14. 13. A method implemented by a video decoder, comprising: receiving, by the video decoder, a video bitstream including a sample adaptive offset (SAO), wherein the SAO flag equal to a first value specifies that SAO signaling is present in a picture header (PH) and the SAO flag equal to a second value specifies that SAO signaling is not present in the PH and may be present in a slice header; and decoding, by the video decoder, the coded picture using the SAO flag to obtain a decoded picture. method.

15. 15. The method of claim 14, further comprising displaying the decoded picture on a display of the electronic device.

16. 16. The method of claim 14 or 15, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.

17. 13. A method implemented by a video encoder, comprising: generating a sample adaptive offset (SAO) flag, the SAO flag equal to a first value specifying that SAO signaling is present in a picture header (PH) and the SAO flag equal to a second value specifying that SAO signaling is not present in the PH and may be present in a slice header; encoding, by the video encoder, the SAO flag into a video bitstream; storing, by the video encoder, the video bitstream for communication to a video decoder. method.

18. 20. The method of claim 17, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.

19. 13. A method implemented by a video decoder, comprising: The video decoder receiving a video bitstream including an adaptive loop filter (ALF) flag, the ALF flag equal to a first value specifying that ALF signaling is present in a picture header (PH), and the ALF flag equal to a second value specifying that ALF signaling is not present in the PH and may be present in a slice header; and decoding the coded picture using the ALF flag to obtain a decoded picture. method.

20. 20. The method of claim 19, further comprising displaying the decoded picture on a display of the electronic device.

21. 21. The method of claim 19 or 20, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.

22. A memory configured to store instructions; a processor coupled to said memory and configured to execute said instructions to perform any one of claims 1 to 7, 14 to 16, and 19 to 21. Video decoder.

23. A computer program product having computer executable instructions for storage on a non-transitory medium which, when executed by a processor, causes a video decoder to perform any one of claims 1 to 7, 14 to 16, 19 to 21.

24. 13. A method implemented by a video encoder, comprising: generating an adaptive loop filter (ALF) flag, the ALF flag equal to a first value specifying that ALF signaling is present in a picture header (PH) and the ALF flag equal to a second value specifying that ALF signaling is not present in the PH and may be present in a slice header; encoding, by the video encoder, the ALF flag into a video bitstream; storing, by the video encoder, the video bitstream for communication to a video decoder. method.

25. 25. The method of claim 24, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.

26. A memory configured to store instructions; a processor coupled to said memory and configured to execute said instructions to perform any one of claims 8 to 13, 17 to 18 and 24 to 25. Video encoder.

27. A computer program product having computer-executable instructions for storage on a non-transitory medium which, when executed by a processor, causes a video encoder to perform any one of items 8-13, 17-18 and 24-25.

28. 13. A method implemented by a video decoder, comprising: receiving, by the video decoder, a video bitstream including a syntax element that specifies that information may or may not be present in a picture header (PH), or that the information may or may not be present in a slice header; and decoding, by the video decoder, the coded picture using the syntax element to obtain a decoded picture. method.

29. 30. A computer readable medium having a bitstream encoded or decoded by the method of any one of claims 1 to 21, 24 to 25 and 28.

Citation Information

Patent Citations

  • Image encoding method, image decoding method, image encoding device, image decoding device, and image encoding-decoding device

    WO2013042329A1

  • Signaling of reference picture lists in video coding

    WO2020112488A1

  • Systems and methods for signaling temporal sub-layer information in video coding

    WO2021045128A1