Signaling of non-picture level syntax elements at picture level
By signaling non-picture-level syntax elements at the picture level with flags, the method optimizes video compression by reducing redundancy, enhancing compression ratios and maintaining image quality.
Patent Information
- Application Number
- JP2025227175
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-04
AI Technical Summary
The challenge of efficiently compressing video data for transmission over limited bandwidth networks while maintaining high image quality, as existing video compression techniques result in redundant bit usage due to signaling of non-picture-level syntax elements in both picture and slice headers.
Implementing a method where non-picture-level syntax elements are signaled at the picture level, using flags to determine their presence in the picture header or slice header, reducing redundancy by ensuring they are only present in one header, thereby optimizing bitstream efficiency.
This approach reduces wasted bits in the encoded bitstream, enhancing compression ratios without sacrificing image quality, thus improving bandwidth utilization and storage efficiency.
Smart Images

Figure 2026035808000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a divisional application of Japanese Patent Application No. 2022-518802, filed March 23, 2022, which is a continuation of International Application No. PCT / US2020 / 052281, filed September 23, 2020, and which claims priority to U.S. Provisional Patent Application No. 62 / 905,228, entitled "Signaling Non-Picture-Level Syntax Elements in Picture Headers in Video Coding," filed September 24, 2019 by Futureway Technologies, Inc., which is incorporated by reference.
[0002] Technical Field The disclosed embodiments relate generally to video coding, and more particularly to signaling non-picture-level syntax elements at the picture level. [Background technology]
[0003] The amount of video data required to represent even a relatively short video can be substantial, which can create difficulties when the data is streamed or otherwise communicated over communications networks with limited bandwidth capacity. Thus, video data is typically compressed before being communicated over modern telecommunications networks. Video size can also be an issue when the video is stored on a storage device, as memory resources may be limited. Video compression devices often use software and / or hardware at the source to encode the video data before transmission or storage, thereby reducing the amount of data required to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. With limited network resources and ever-increasing demands for higher video quality, improved compression and decompression techniques that improve compression ratios with little or no sacrifice in image quality are desirable. Summary of the Invention [Means for solving the problem]
[0004] A first aspect relates to a method implemented by a video decoder, the method comprising: receiving, by the video decoder, a video bitstream including an RPL flag, wherein the RPL flag equal to a first value specifies that RPL signaling is present in a PH, and wherein the RPL flag equal to a second value specifies that RPL signaling is not present in a PH, but may be present in a slice header.
[0005] In embodiments, syntax elements are included in picture headers if they are the same and in slice headers if they vary. However, in some embodiments, syntax elements may not be included in both. First, non-picture level syntax elements may be present in the PH. Non-picture level syntax elements are syntax elements at levels of the video bitstream other than the picture level. Second, for each category of non-picture level syntax elements, a flag specifies when the syntax elements of that category are present in the PH or the slice header. The flag may be present in the PH. Non-picture level syntax elements include RPL signaling, joint Cb Cr code flag, SAO tool enable and parameters, ALF tool enable and parameters, LMCS tool enable and parameters, and scaling list tool enable and parameters. Third, if non-picture level syntax elements are present in the PH, the corresponding syntax elements are not present in any slices of the picture associated with the picture header containing those syntax elements. The values of non-picture level syntax elements present in the PH apply to all slices of the picture associated with the picture header containing those syntax elements. Fourth, if a non-picture level syntax element is not present in the PH, the corresponding syntax element may be present in the slice header of the slice of the picture associated with the picture header. By moving the signaling of non-picture level syntax elements to the picture level, redundancy is reduced, resulting in fewer wasted bits in the encoded bitstream.
[0006] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that RPL signaling is present in the PH.
[0007] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that no RPL signaling is present in the slice.
[0008] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that no RPL signaling is present in the PH.
[0009] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that RPL signaling may be present in the slice header.
[0010] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the bitstream further includes an RPL SPS flag, where the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS, or specifies that RPL i is derived based on one of the directly included ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i.
[0011] Optionally, in any of the above aspects, another implementation of this aspect provides that the bitstream further includes an RPL index, where the RPL index specifies an index into a list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i included in the sequence parameter set (SPS) of a ref_pic_list_struct(listIdx,rplsIdx) syntax structure with listIdx equal to i that is used to derive RPL i of the current picture.
[0012] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes displaying the decoded picture on a display of the electronic device.
[0013] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that if non-picture level syntax elements are present in a PH, then corresponding syntax elements are not present in any slice of the picture associated with the PH that contains those syntax elements.
[0014] A second aspect relates to a method implemented by a video encoder, comprising generating an RPL flag, wherein an RPL flag equal to a first value specifies that RPL signaling is present in a PH, and an RPL flag equal to a second value specifies that RPL signaling is not present in the PH but may be present in a slice header.
[0015] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that RPL signaling is present in the PH.
[0016] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 1 specifies that no RPL signaling is present in the slice.
[0017] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that no RPL signaling is present in the PH.
[0018] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that an RPL flag equal to 0 specifies that RPL signaling may be present in the slice header.
[0019] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes generating an RPL SPS flag, where the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS, or specifies that RPL i is derived based on one of the directly contained ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i.
[0020] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes generating an RPL index, where the RPL index specifies an index into a list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i included in the sequence parameter set (SPS) of a ref_pic_list_struct(listIdx,rplsIdx) syntax structure with listIdx equal to i that is used to derive RPL i of the current picture.
[0021] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that if non-picture level syntax elements are present in a PH, then corresponding syntax elements are not present in any slice of the picture associated with the PH that contains those syntax elements.
[0022] A third aspect relates to a method implemented by a video decoder, the method including: receiving, by the video decoder, a video bitstream including an SAO flag, wherein the SAO flag equal to a first value specifies that SAO signaling is present in a PH, and the SAO flag equal to a second value specifies that SAO signaling is not present in a PH but may be present in a slice header; and decoding, by the video decoder, a coded picture using the SAO flag to obtain a decoded picture.
[0023] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that the method further includes displaying the decoded picture on a display of the electronic device.
[0024] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that if non-picture level syntax elements are present in a PH, then corresponding syntax elements are not present in any slice of the picture associated with the PH that contains those syntax elements.
[0025] A fourth aspect is a method implemented by a video encoder, comprising: generating an SAO flag, where an SAO flag equal to a first value specifies that SAO signaling is present in a PH and an SAO flag equal to a second value specifies that SAO signaling is not present in the PH but may be present in a slice header; encoding, by the video encoder, the SAO flag into a video bitstream; and storing, by the video encoder, the video bitstream for communication to a video decoder.
[0026] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that if non-picture level syntax elements are present in a PH, then corresponding syntax elements are not present in any slice of the picture associated with the PH that contains those syntax elements.
[0027] A fifth aspect relates to a method implemented by a video decoder, the method comprising: receiving, by the video decoder, a video bitstream including an ALF flag, wherein the ALF flag equal to a first value specifies that ALF signaling is present in a PH, and the ALF flag equal to a second value specifies that ALF signaling is not present in a PH but may be present in a slice header; and decoding, by the video decoder, a picture coded using the ALF flag to obtain a decoded picture.
[0028] Optionally, in any of the foregoing aspects, another implementation aspect of this aspect provides that the method further includes displaying the decoded picture on a display of the electronic device.
[0029] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that if a non-picture level syntax element is present in a PH, the corresponding syntax element is not present in any slice of the picture associated with the PH that contains the syntax element.
[0030] A sixth aspect provides a method implemented by a video encoder, the method comprising: generating an ALF flag, where an ALF flag equal to a first value specifies that ALF signaling is present in a PH, and an ALF flag equal to a second value specifies that ALF signaling is not present in the PH but may be present in a slice header; encoding, by the video encoder, the ALF flag into a video bitstream; and storing, by the video encoder, the video bitstream for communication to a video decoder.
[0031] Optionally, in any of the aforementioned aspects, another implementation of this aspect provides that if a non-picture level syntax element is present in a PH, the corresponding syntax element is not present in any slice of the picture associated with the PH that contains the syntax element.
[0032] A seventh aspect relates to a method implemented by a video decoder, the method comprising: receiving, by the video decoder, a video bitstream including a syntax element, the syntax element specifying that information may or may not be present in a PH, or specifying that the information may or may not be present in a slice header; and decoding, by the video decoder, a coded picture using the syntax element to obtain a decoded picture.
[0033] Any of the above embodiments may be combined with any of the other above embodiments to create new embodiments. These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. [Brief explanation of the drawings]
[0034] For a more complete understanding of this disclosure, reference is now made to the following brief description taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.
[0035] [Figure 1] 1 is a flowchart of an exemplary method for encoding a video signal.
[0036] [Figure 2] 1 is a schematic diagram of an example of a coding and decoding (codec) system for video coding.
[0037] [Figure 3] 1 is a schematic diagram illustrating an exemplary video encoder.
[0038] [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary video decoder.
[0039] [Figure 5] 1 illustrates an embodiment of a video bitstream.
[0040] [Figure 6] 3 is a flowchart illustrating a method for decoding a bitstream according to a first embodiment.
[0041] [Figure 7] 3 is a flowchart illustrating a method for encoding a bitstream according to a first embodiment.
[0042] [Figure 8] 10 is a flowchart illustrating a method for decoding a bitstream according to a second embodiment.
[0043] [Figure 9] 10 is a flowchart illustrating a method for encoding a bitstream according to a second embodiment.
[0044] [Figure 10] 10 is a flowchart illustrating a method for decoding a bitstream according to a third embodiment.
[0045] [Figure 11] 10 is a flowchart illustrating a method for encoding a bitstream according to a third embodiment.
[0046] [Figure 12] 1 is a schematic diagram of a video coding device.
[0047] [Figure 13] 1 is a schematic diagram of an embodiment of a coding means; DETAILED DESCRIPTION OF THE INVENTION
[0048] While exemplary implementations of one or more embodiments are provided below, it should be understood at the outset that the disclosed systems and / or methods may be implemented using any number of technologies, whether currently known or existing. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and the full scope of equivalents.
[0049] The following abbreviations apply: ALF: adaptive loop filter APS: adaptation parameter set ASIC: application-specific integrated circuit AU: access unit AUD:access unit delimiter BT:binary tree CABAC: context-adaptive binary arithmetic coding CAVLC: context-adaptive variable-length coding Cb:blue difference chroma CLVS: coded layer-wise video sequence CLVS: coded layer video sequence CPU: central processing unit Cr: red difference chroma CRA: clean random access CTB: coding tree block CTU: coding tree unit CU: coding unit CVS: coded video sequence DC:direct current DCI: decoding capability information DCT: Discrete cosine transform DMM: depth modeling mode DPB: decoded picture buffer DPS: decoding parameter set DSP: digital signal processor DST: discrete sine transform EO: electrical-to-optical FPGA: field-programmable gate array GDR: gradual decoding refresh HEVC: High Efficiency Video Coding ID: identifier IDR: Instantaneous decoding refresh IEC: International Electrotechnical Commission I / O: input / output IRAP: intra random access pictures ISO: International Organization for Standardization ITU: International Telecommunication Union ITU-T: ITU Telecommunication Standardization Sector LMCS: luma mapping with chroma scaling LTRP: long-term reference picture MVP: motion vector predictor NAL: network abstraction layer OE: optical-to-electrical PH: picture header PIPE: probability interval partitioning entropy POC: picture order count PPS: picture parameter set PU: picture unit QT:quad tree RADL: random access decodable leading RAM: random-access memory RASL: random access skipped leading RBSP: raw byte sequence payload RDO: rate-distortion optimization ROM: read-only memory RPL: reference picture list Rx: receiver unit SAD: sum of absolute differences SAO: sample adaptive offset SBAC: syntax-based arithmetic coding SOP: sequence of pictures SPS:sequence parameter set SRAM:static RAM SSD:sum of squared differences TCAM: ternary content-addressable memory TT: triple tree TU: transform unit Tx: transmitter unit VCL: video coding layer VPS:video parameter set VVC: Versatile Video Coding
[0050] Unless otherwise modified, the following definitions apply: A bitstream is a sequence of bits containing compressed video data for transmission between an encoder and a decoder. An encoder is a device that uses an encoding process to compress video data into a bitstream. A decoder is a device that uses a decoding process to reconstruct video data from a bitstream for display. A picture is an array of luma or chroma samples that make up a frame or field. The picture being encoded or decoded is sometimes referred to as the current picture. A reference picture contains reference samples that can be used when coding other pictures by reference according to inter- or inter-layer prediction. A reference picture list is a list of reference pictures used for inter- or inter-layer prediction. A flag is a variable or one-bit syntax element that can take one of two possible values (0 or 1). Some video coding systems utilize two reference picture lists, which can be represented as reference picture list 1 and reference picture list 0. A reference picture list structure is an addressable syntax structure that contains multiple reference picture lists. Inter-prediction is the coding of samples of a current picture by reference to designated samples in a reference picture different from the current picture, where the reference picture and the current picture are in the same layer. A reference picture list structure entry is an addressable location within a reference picture list structure that indicates the reference picture associated with the reference picture list. A slice header is part of a coded slice and contains data elements related to all video data within the tile represented within the slice. A PPS contains data related to the entire picture. More specifically, a PPS is a syntax structure containing syntax elements that apply to zero or more entire coded pictures, determined by the syntax elements found in each picture header. An SPS contains data related to a sequence of pictures.An AU is a collection of one or more coded pictures associated with the same display time (e.g., the same picture order count) for output from the DPB (e.g., for display to a user). An AUD indicates the beginning of an AU or the boundary between AUs. A decoded video sequence is a sequence of pictures reconstructed by a decoder in preparation for display to a user.
[0051] FIG. 1 is a flowchart of an exemplary operational method 100 for coding a video signal. Specifically, a video signal is encoded in an encoder. The encoding process compresses the video signal by employing various mechanisms to reduce the video file size. The smaller file size allows the compressed video file to be transmitted to a user while reducing the associated bandwidth overhead. A decoder then decodes the compressed video file to reconstruct the original video signal for display to the end user. The decoding process generally mirrors the encoding process, allowing the decoder to consistently reconstruct the video signal.
[0052] In step 101, a video signal is input to an encoder. For example, the video signal may be an uncompressed video file stored in memory. As another example, the video file may be captured by a video capture device, such as a video camera, and encoded to support live streaming of the video. The video file may include both an audio component and a video component. The video component includes a series of image frames that, when viewed in sequence, create the impression of visual movement. The frames include pixels represented using brightness, referred to herein as luma components (or luma samples), and color, referred to herein as chroma components (or color samples). In some examples, the frames may also include depth values to support three-dimensional viewing.
[0053] In step 103, the video is partitioned into blocks. Partitioning involves subdividing pixels within each frame into square and / or rectangular blocks for compression. For example, in HEVC, a frame can first be divided into CTUs, which are blocks of a predefined size (e.g., 64 pixels by 64 pixels). A CTU contains both luma samples and chroma samples. A coding tree can be used to divide the CTUs into blocks, and then the blocks can be recursively subdivided until a configuration that supports further encoding is achieved. For example, the luma component of a frame can be subdivided until each block contains relatively uniform illumination values. Furthermore, the chroma component of a frame can be subdivided until each block contains relatively uniform color values. Thus, the partitioning scheme varies depending on the content of the video frame.
[0054] In step 105, various compression mechanisms are used to compress the image blocks partitioned in step 103. For example, inter-prediction and / or intra-prediction may be used. Inter-prediction is designed to take advantage of the fact that objects in a common scene tend to appear in consecutive frames. Thus, a block depicting an object in a reference frame need not be repeatedly described in adjacent frames. Specifically, an object such as a table may remain in a constant position across multiple frames. Thus, the table may be described once, and adjacent frames may reference the previous reference frame. Pattern matching mechanisms can be used to match objects across multiple frames. Furthermore, moving objects may be represented across multiple frames, for example, due to object motion or camera motion. As a specific example, a video may show a car moving across the screen over multiple frames. Motion vectors can be used to describe such motion. A motion vector is a two-dimensional vector that provides an offset from the object's coordinates in a frame to the object's coordinates in a reference frame. Thus, inter-prediction can encode an image block in a current frame as a set of motion vectors that indicate its offset from a corresponding block in a reference frame.
[0055] Intra prediction encodes blocks within a common frame. Intra prediction takes advantage of the fact that luma and chroma components tend to cluster in a frame. For example, a green patch in a part of a tree tends to be located adjacent to similar green patches. Intra prediction uses multiple directional prediction modes (e.g., 33 in HEVC), planar mode, and DC mode. Directional mode indicates that the current block is similar / the same as samples of neighboring blocks in the corresponding direction. Planar mode indicates that a series of blocks along a row / column (e.g., a plane) can be interpolated based on neighboring blocks at the ends of the row. Planar mode effectively indicates a smooth transition in brightness / color across the row / column by using a relatively constant slope in changing values. DC mode is used for boundary smoothing and indicates that the block is similar / the same as the average value associated with samples of all neighboring blocks associated with the angular direction of the directional prediction mode. Thus, intra prediction blocks can represent image blocks as various related prediction mode values instead of their actual values. Additionally, inter-predicted blocks can represent image blocks as motion vector values instead of actual values. In either case, the predicted blocks may not exactly represent the image blocks in some cases. Any differences are stored in residual blocks. Transforms may be applied to the residual blocks to further compress the file.
[0056] Various filtering techniques may be applied in step 107. In HEVC, filters are applied according to an in-loop filtering scheme. The block-based prediction described above can lead to the generation of blocky images at the decoder. Furthermore, the block-based prediction scheme can encode a block and then reconstruct the encoded block for later use as a reference block. The in-loop filtering scheme iteratively applies a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter to the block / frame. These filters mitigate such blocking artifacts so that the encoded file can be accurately reconstructed. Furthermore, these filters mitigate artifacts in the reconstructed reference block, so that the artifacts are less likely to generate additional artifacts in subsequent blocks that are encoded based on the reconstructed reference block.
[0057] Once the video signal has been partitioned, compressed, and filtered, the resulting data is encoded into a bitstream in step 109. The bitstream includes the data described above and any signaling data desired to support proper video signal reconstruction at the decoder. For example, such data may include partition data, prediction data, residual blocks, and various flags that provide coding instructions to the decoder. The bitstream may be stored in memory for transmission to the decoder upon request. The bitstream may also be broadcast and / or multicast to multiple decoders. Generating the bitstream is an iterative process. Thus, steps 101, 103, 105, 107, and 109 can occur sequentially and / or simultaneously across many frames and blocks. The order depicted in FIG. 1 is for clarity and ease of discussion and is not intended to limit the video coding process to any particular order.
[0058] The decoder receives the bitstream in step 111 and begins the decoding process. Specifically, the decoder employs an entropy decoding scheme to convert the bitstream into corresponding syntax and video data. In step 111, the decoder uses the syntax data from the bitstream to determine the partitions for the frame. The partitioning should match the results of the block partitioning in step 103. The entropy encoding / decoding used in step 111 is now described. The encoder makes many choices during the compression process, such as selecting a block partitioning scheme from several possible choices based on the spatial positioning of values within the input image(s). Signaling the exact selection may use multiple bins. As used herein, a bin is a binary value (e.g., a bit value that can change depending on the context) that is treated as a variable. Entropy coding allows the encoder to discard any options that are clearly not realistic for a particular case, leaving a set of acceptable options. Each acceptable option is then assigned a codeword. The length of the codeword is based on the number of allowable options (e.g., one bin for two options, two bins for three to four options, etc.). The encoder then encodes the codeword for the selected option. This scheme reduces the size of the codeword because the codeword is only as large as desired to uniquely indicate a selection from a small subset of allowable options, rather than uniquely indicating a selection from a potentially large set of all possible options. The decoder then decodes the selection by determining the set of allowable options in a similar manner to the encoder. By determining the set of allowable options, the decoder can read the codeword and determine the selection made by the encoder.
[0059] In step 113, the decoder performs block decoding. Specifically, the decoder uses an inverse transform to generate a residual block. The decoder then uses the residual block and a corresponding prediction block to reconstruct an image block according to the partitioning. The prediction block may include both intra-predicted and inter-predicted blocks generated by the encoder in step 105. The reconstructed image block is then positioned within a frame of the reconstructed video signal according to the partitioning data determined in step 111. The syntax for step 113 may also be signaled in the bitstream by entropy coding, as described above.
[0060] At step 115, filtering is performed on the frames of the reconstructed video signal in a manner similar to step 107 in the encoder. For example, a noise suppression filter, a deblocking filter, an adaptive loop filter, and an SAO filter may be applied to the frames to remove blocking artifacts. Once the frames have been filtered, the video signal can be output to a display at step 117 for viewing by an end user.
[0061] FIG. 2 is a schematic diagram of an exemplary coding and decoding (codec) system 200 for video coding. Specifically, codec system 200 provides functionality to support the implementation of operational method 100. Codec system 200 is generalized to illustrate components used in both encoders and decoders. Codec system 200 receives and partitions a video signal, as described with reference to steps 101 and 103 of operational method 100, resulting in partitioned video signal 201. When acting as an encoder, codec system 200 then compresses partitioned video signal 201 into a coded bitstream, as described with reference to steps 105, 107, and 109 of method 100. When acting as a decoder, codec system 200 generates an output video signal from the bitstream, as described with reference to steps 111, 113, 115, and 117 of operational method 100. Codec system 200 includes a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, an intra-picture prediction component 217, a motion compensation component 219, a motion estimation component 221, a scaling and inverse transform component 229, a filter control analysis component 227, an in-loop filter component 225, a decoded picture buffer component 223, and a header format and CABAC component 231. Such components are coupled as shown. In FIG. 2, black lines indicate the movement of data to be encoded / decoded, and dashed lines indicate the movement of control data that controls the operation of other components. All of the components of codec system 200 may reside within an encoder. A decoder may include a subset of the components of codec system 200.For example, the decoder may include an intra-picture prediction component 217, a motion compensation component 219, a scaling and inverse transform component 229, an in-loop filter component 225, and a decoded picture buffer component 223. These components are now described.
[0062] The partitioned video signal 201 is a captured video sequence that has been partitioned into blocks of pixels by a coding tree. The coding tree uses various partitioning modes to subdivide the blocks of pixels into smaller blocks of pixels. These blocks can then be further subdivided into smaller blocks. Blocks may be referred to as nodes on the coding tree. Larger parent nodes are divided into smaller child nodes. The number of times a node is subdivided is referred to as the depth of the node / coding tree. The partitioned blocks may be included in CUs, in some cases. For example, a CU may be a subpart of a CTU that includes a luma block, a Cr block, and a Cb block along with the corresponding syntax instructions for that CU. Partitioning modes may include BT, TT, and QT, which are used to partition a node into two, three, or four child nodes, each of different shapes depending on the partitioning mode used. The partitioned video signal 201 is forwarded to a general coder control component 211, a transform scaling and quantization component 213, an intra-picture estimation component 215, a filter control analysis component 227, and a motion estimation component 221 for compression.
[0063] The general coder control component 211 is configured to make decisions related to coding images of a video sequence into a bitstream according to application constraints. For example, the general coder control component 211 manages the optimization of bitrate / bitstream size versus reconstruction quality. Such decisions may be made based on storage space / bandwidth availability and image resolution requirements. The general coder control component 211 also manages buffer utilization in relation to transmission rate to mitigate buffer underrun and overrun issues. To manage these issues, the general coder control component 211 manages partitioning, prediction, and filtering by other components. For example, the general coder control component 211 may dynamically increase compression complexity to increase resolution and bandwidth usage, or decrease compression complexity to decrease resolution and bandwidth usage. Thus, the general coder control component 211 controls other components of the codec system 200 to balance video signal reconstruction quality with bitrate concerns. The general coder control component 211 generates control data that controls the operation of other components. Control data is also forwarded to the header format and CABAC component 231 to be encoded in the bitstream to signal parameters for decoding at the decoder.
[0064] The partitioned video signal 201 is also sent to a motion estimation component 221 and a motion compensation component 219 for inter-prediction. A frame or slice of the partitioned video signal 201 may be divided into multiple video blocks. The motion estimation component 221 and the motion compensation component 219 perform inter-predictive coding of the received video blocks relative to one or more blocks in one or more reference frames to provide temporal prediction. The codec system 200 may perform multiple coding passes, for example, to select an appropriate coding mode for each block of video data.
[0065] The motion estimation component 221 and the motion compensation component 219 may be highly integrated but are illustrated separately for conceptual purposes. Motion estimation, performed by the motion estimation component 221, is the process of generating motion vectors, which estimate motion for video blocks. A motion vector may indicate, for example, the displacement of an object being coded relative to a predictive block. A predictive block is a block that is found to closely match a block to be coded in terms of pixel difference. A predictive block is sometimes referred to as a reference block. Such pixel difference may be determined by SAD, SSD, or other difference metrics. HEVC uses several coded objects, including CTUs, CTBs, and CUs. For example, a CTU may be divided into CTBs, and a CTB may be divided into CBs for inclusion in a CU. A CU may be encoded as a prediction unit (PU), which contains prediction data, and / or a TU, which contains transformed residual data for the CU. The motion estimation component 221 generates motion vectors, prediction units, and TUs by using rate-distortion analysis as part of a rate-distortion optimization process. For example, the motion estimation component 221 may determine multiple reference blocks, multiple motion vectors, etc. for the current block / frame and select the reference block, motion vector, etc. with the best rate-distortion characteristics, which balance both the quality of the video reconstruction (e.g., the amount of data lost due to compression) and the coding efficiency (e.g., the size of the final encoding).
[0066] In some examples, the codec system 200 can calculate values for sub-integer pixel positions of a reference picture stored in the decoded picture buffer component 223. For example, the video codec system 200 can interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of a reference picture. Thus, the motion estimation component 221 can perform motion searches for whole-pixel and fractional pixel positions and output motion vectors with fractional-pixel precision. The motion estimation component 221 calculates motion vectors for prediction units of video blocks in inter-coded slices by comparing the positions of the prediction units with the positions of the predictive blocks of the reference pictures.
[0067] The motion estimation component 221 outputs the calculated motion vector as motion data to the header format and CABAC component 231 for encoding and outputs the motion to the motion compensation component 219. The motion compensation performed by the motion compensation component 219 may include fetching or generating a prediction block based on the motion vector determined by the motion estimation component 221. Again, in some examples, the motion estimation component 221 and the motion compensation component 219 may be functionally integrated. Upon receiving the motion vector for the prediction unit of the current video block, the motion compensation component 219 can locate the prediction block to which the motion vector points. A residual video block is then formed by subtracting pixel values of the prediction block from pixel values of the current video block being coded to form pixel difference values. Generally, the motion estimation component 221 performs motion estimation on the luma component, and the motion compensation component 219 uses the motion vector calculated based on the luma component for both the chroma and luma components. The prediction block and the residual block are forwarded to the transform scaling and quantization component 213.
[0068] The partitioned video signal 201 is also sent to an intra picture estimation component 215 and an intra picture prediction component 217. Like the motion estimation component 221 and motion compensation component 219, the intra picture estimation component 215 and the intra picture prediction component 217 may be highly integrated but are illustrated separately for conceptual purposes. The intra picture estimation component 215 and the intra picture prediction component 217 intra-predict the current block relative to blocks within the current frame, instead of the inter-prediction performed by the motion estimation component 221 and the motion compensation component 219 between frames, as described above. In particular, the intra picture estimation component 215 determines the intra-prediction mode to use to encode the current block. In some examples, the intra picture estimation component 215 selects an appropriate intra-prediction mode to encode the current block from multiple tested intra picture prediction modes. The selected intra-prediction mode is then forwarded to the header format and CABAC component 231 for encoding.
[0069] For example, the intra picture estimation component 215 calculates rate-distortion values for various tested intra picture prediction modes using rate-distortion analysis and selects the intra prediction mode with the best rate-distortion characteristics among the tested modes. The rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block encoded to generate the encoded block, and the bit rate (e.g., number of bits) used to generate the encoded block. The intra picture estimation component 215 calculates a ratio from the distortion and rate for the various encoded blocks and determines which intra prediction mode exhibits the best rate-distortion value for the block. In addition, the intra picture estimation component 215 may be configured to code the depth block of the depth map using a DMM based on RDO.
[0070] The intra picture prediction component 217, when implemented in an encoder, generates a residual block from the prediction block based on the selected intra picture prediction mode determined by the intra picture estimation component 215, and, when implemented in a decoder, can read the residual block from the bitstream. The residual block includes a matrix representation of the difference in values between the prediction block and the original block. The residual block is then forwarded to the transform scaling and quantization component 213. The intra picture estimation component 215 and the intra picture prediction component 217 can operate on both the luma and chroma components.
[0071] The transform scaling and quantization component 213 is configured to further compress the residual block. The transform scaling and quantization component 213 applies a transform, such as a DCT, DST, or a conceptually similar transform, to the residual block, generating a video block containing residual transform coefficient values. A wavelet transform, an integer transform, a subband transform, or other types of transforms may also be used. The transform may convert the residual information from the pixel value domain to a transform domain, such as the frequency domain. The transform scaling and quantization component 213 is also configured to scale the transformed residual information, for example, based on frequency. Such scaling may involve applying a scale factor to the residual information so that different frequency information is quantized with different granularities, which may affect the final visual quality of the reconstructed video. The transform scaling and quantization component 213 is also configured to quantize the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the transform scaling and quantization component 213 may then perform a scan of the matrix containing the quantized transform coefficients, which are forwarded to the header format and CABAC component 231 for encoding in the bitstream.
[0072] The scaling and inverse transform component 229 applies the inverse operations of the transform scaling and quantization component 213 to support motion estimation. The scaling and inverse transform component 229 applies inverse scaling, transform, and / or quantization to reconstruct residual blocks in the pixel domain, for example, for later use as reference blocks that may become predictive blocks for another current block. The motion estimation component 221 and / or motion compensation component 219 can calculate reference blocks by adding the residual blocks to the corresponding predictive blocks for use in motion estimation of subsequent blocks / frames. A filter is applied to the reconstructed reference blocks to mitigate artifacts generated during scaling, quantization, and transform. Such artifacts may otherwise cause inaccurate predictions (and may create additional artifacts) when subsequent blocks are predicted.
[0073] The filter control analysis component 227 and the in-loop filter component 225 apply filters to residual blocks and / or reconstructed image blocks. For example, a transformed residual block from the scaling and inverse transform component 229 can be combined with a corresponding prediction block from the intra-picture prediction component 217 and / or the motion compensation component 219 to reconstruct an original image block. The filter may then be applied to the reconstructed image block. In some examples, the filter may instead be applied to the residual block. Like the other components in FIG. 2, the filter control analysis component 227 and the in-loop filter component 225 are highly integrated and may be implemented together, but are shown separately for conceptual purposes. The filters applied to reconstructed reference blocks are applied to specific spatial regions and include multiple parameters to adjust how such filters are applied. The filter control analysis component 227 analyzes the reconstructed reference blocks to determine where such filters should be applied and set the corresponding parameters. Such data is forwarded as filter control data to the header format and CABAC component 231 for encoding. The in-loop filter component 225 applies such filters based on the filter control data. The filters may include deblocking filters, noise suppression filters, SAO filters, and adaptive loop filters. Such filters may be applied in the spatial / pixel domain (e.g., reconstructed pixel blocks) or in the frequency domain, depending on the example.
[0074] When operating as an encoder, the filtered reconstructed image blocks, residual blocks, and / or prediction blocks are stored in the decoded picture buffer component 223 for later use in motion estimation as described above. When operating as a decoder, the decoded picture buffer component 223 stores and forwards the reconstructed and filtered blocks toward the display as part of the output video signal. The decoded picture buffer component 223 may be any memory device capable of storing prediction blocks, residual blocks, and / or reconstructed image blocks.
[0075] The header format and CABAC component 231 receives data from various components of the codec system 200 and encodes such data into a coded bitstream for transmission to the decoder. Specifically, the header format and CABAC component 231 generates various headers for encoding control data, such as general control data and filter control data. Additionally, prediction data, including intra-prediction and motion data, and residual data in the form of quantized transform coefficient data are all encoded into the bitstream. The final bitstream contains all information desired by the decoder to reconstruct the original partitioned video signal 201. Such information may also include an intra-prediction mode index table (also called a codeword mapping table), definitions of encoding contexts for various blocks, indications of the most likely intra-prediction mode, indications of partition information, etc. Such data may be encoded using entropy coding. For example, the information may be encoded using CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. Following entropy coding, the coded bitstream may be transmitted to another device (eg, a video decoder) or archived for later transmission or retrieval.
[0076] 3 is a block diagram illustrating an exemplary video encoder 300. Video encoder 300 may be used to implement the encoding functionality of codec system 200 and / or to implement steps 101, 103, 105, 107, and / or 109 of operating method 100. Encoder 300 partitions an input video signal, resulting in partitioned video signal 301, which is substantially similar to partitioned video signal 201. Partitioned video signal 301 is then compressed and encoded into a bitstream by components of encoder 300.
[0077] Specifically, the partitioned video signal 301 is forwarded to an intra picture prediction component 317 for intra prediction. The intra picture prediction component 317 may be substantially similar to the intra picture estimation component 215 and the intra picture prediction component 217. The partitioned video signal 301 is also forwarded to a motion compensation component 321 for inter prediction based on reference blocks in a decoded picture buffer component 323. The motion compensation component 321 may be substantially similar to the motion estimation component 221 and the motion compensation component 219. The prediction block and residual block from the intra picture prediction component 317 and the motion compensation component 321 are forwarded to a transform and quantization component 313 for transforming and quantizing the residual block. The transform and quantization component 313 may be substantially similar to the transform scaling and quantization component 213. The transformed and quantized residual block and the corresponding prediction block (together with associated control data) are forwarded to an entropy coding component 331 for coding into a bitstream. The entropy coding component 331 may have a header format and be substantially similar to the CABAC component 231.
[0078] The transformed and quantized residual block and / or the corresponding prediction block are also transferred from the transform and quantization component 313 to the inverse transform and quantization component 329 for reconstructing into a reference block for use by the motion compensation component 321. The inverse transform and quantization component 329 may be substantially similar to the scaling and inverse transform component 229. An in-loop filter within the in-loop filter component 325 is also applied to the residual block and / or the reconstructed reference block, depending on the example. The in-loop filter component 325 may be substantially similar to the filter control analysis component 227 and the in-loop filter component 225. The in-loop filter component 325 may include multiple filters, as discussed with respect to the in-loop filter component 225. The filtered block is then stored in the decoded picture buffer component 323 for use as a reference block by the motion compensation component 321. The decoded picture buffer component 323 may be substantially similar to the decoded picture buffer component 223.
[0079] 4 is a block diagram illustrating an exemplary video decoder 400. Video decoder 400 may be used to implement the decoding functionality of codec system 200 and / or to implement steps 111, 113, 115, and / or 117 of method of operation 100. Decoder 400 receives a bitstream, for example, from encoder 300, and generates a reconstructed output video signal based on the bitstream for display to an end user.
[0080] The bitstream is received by the entropy decoding component 433. The entropy decoding component 433 is configured to implement an entropy decoding scheme, such as CAVLC, CABAC, SBAC, PIPE coding, or other entropy coding techniques. For example, the entropy decoding component 433 can use header information to provide context for interpreting additional data encoded as codewords in the bitstream. The decoded information includes any desired information for decoding the video signal, such as general control data, filter control data, partition information, motion data, prediction data, and quantized transform coefficients from the residual block. The quantized transform coefficients are forwarded to the inverse transform and quantization component 429 for reconstruction into the residual block. The inverse transform and quantization component 429 may be similar to the inverse transform and quantization component 329.
[0081] The reconstructed residual block and / or predictive block are forwarded to the intra picture prediction component 417 for reconstruction into an image block based on an intra prediction operation. The intra picture prediction component 417 may be similar to the intra picture estimation component 215 and the intra picture prediction component 217. Specifically, the intra picture prediction component 417 uses a prediction mode to locate a reference block within a frame and applies the residual block to the result to reconstruct an intra predicted image block. The reconstructed intra predicted image block and / or residual block and corresponding inter prediction data are forwarded to the decoded picture buffer component 423 via an in-loop filter component 425, which may be substantially similar to the decoded picture buffer component 223 and the in-loop filter component 225, respectively. The in-loop filter component 425 filters the reconstructed image block, residual block, and / or predictive block, and such information is stored in the decoded picture buffer component 423. The reconstructed image blocks from the decoded picture buffer component 423 are forwarded to the motion compensation component 421 for inter-prediction. The motion compensation component 421 may be substantially similar to the motion estimation component 221 and / or the motion compensation component 219. Specifically, the motion compensation component 421 generates a prediction block using a motion vector from a reference block and applies a residual block to the result to reconstruct an image block. The resulting reconstructed block may be forwarded to the decoded picture buffer component 423 via an in-loop filter component 425. The decoded picture buffer component 423 continues to store additional reconstructed image blocks that can be reconstructed into frames via the partition information. Such frames may be arranged in a sequence. This sequence is output to a display as a reconstructed output video signal.
[0082] FIG. 5 illustrates an embodiment of a video bitstream 500. The video bitstream 500 may also be referred to as a coded video bitstream, a bitstream, or variations thereof. The bitstream 500 includes at least one PU 501. Although three PUs 501 are shown in FIG. 5, in practical applications, a different number of PUs 501 may be present in the bitstream 500. Each PU 501 is a collection of NAL units that are associated with each other according to a specified classification rule, are consecutive in decoding order, and contain exactly one coded picture (e.g., picture 514). In one embodiment, each PU 501 has or is associated with a temporal ID 519.
[0083] In one embodiment, each PU 501 includes one or more of a DCI 502, a VPS 504, an SPS 506, a PPS 508, a PH 512, and a picture 514. Each of the DCI 502, the VPS 504, the SPS 506, and the PPS 508 may be collectively referred to as a parameter set. Other parameter sets not shown in FIG. 5, such as an APS, may be included in the bitstream 500. An APS is a syntax structure that includes syntax elements that apply to zero or more slices, as determined by zero or more syntax elements found in a slice header 520.
[0084] DCI 502, sometimes referred to as DPS, is a syntax structure containing syntax elements that apply to the entire bitstream. DCI 502 contains parameters that remain constant over the lifetime of a video bitstream (e.g., bitstream 500), which can be translated to the lifetime of a session. DCI 502 can include profile, level, and subprofile information to determine a maximum complexity interop point that is guaranteed never to be exceeded, even if splicing of video sequences occurs within a session. It can also optionally include constraint flags, which indicate that the video bitstream will be constrained to use certain features indicated by the values of those flags. This allows the bitstream to be labeled as not using certain tools, allowing resource allocation in decoder implementations, etc. Like all parameter sets, DCI 502 is present at the first reference and is referenced by the first picture of a video sequence. This implies that DCI 502 must be transmitted within the first NAL units of the bitstream. There may be multiple DCIs 502 in a bitstream, but the values of syntax elements within them must not be contradictory when referenced.
[0085] The VPS 504 contains decoding dependencies or information for construction of reference picture sets for enhancement layers. It provides an overall perspective or view of the scalable sequence, which includes what types of operation points are provided, the profile, tier, and level of those operation points, and several other high-level characteristics of the bitstream that can be used as the basis for session negotiation, content selection, etc.
[0086] The SPS 506 contains data common to all pictures in the SOP. The SPS 506 is a syntactic structure containing syntax elements that apply across zero or more CLVs, as determined by the content of syntax elements found in the PPS referenced by syntax elements found in each picture header. In contrast, the PPS 508 contains data common to an entire picture. The PPS 508 is a syntactic structure containing syntax elements that apply across zero or more coded pictures, as determined by the content of syntax elements found in each picture header (e.g., PH 512).
[0087] DCI 502, VPS 504, SPS 506, and PPS 508 are contained in different types of NAL units. A NAL unit is a syntactic structure that contains an indication of the type of data (e.g., coded video data) that follows. NAL units are classified as VCL NAL units and non-VCL NAL units. VCL NAL units contain data that represent the values of samples in a video picture, while non-VCL NAL units contain any relevant additional information, such as parameter sets (important data applicable to some VCL NAL units) and supplemental enhancement information (timing information and other supplemental data that may enhance the usefulness of the decoded video signal but are not necessary for decoding the values of samples in a video picture).
[0088] PH 512 is a syntax structure that includes syntax elements that apply to all slices (e.g., slice 518) of a coded picture (e.g., picture 514). In an embodiment, PH 512 is within a new type of non-VCL NAL unit designated as a PH NAL unit. Thus, a PH NAL unit has a PH NUT (e.g., PH_NUT). In an embodiment, each PU 501 includes only one PH 512. That is, a PU 501 includes a single or isolated PH 512. In an embodiment, exactly one PH NAL unit exists for each picture 501 in bitstream 500.
[0089] A picture 514 is an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats. In one embodiment, each PU 501 contains only one picture 514. Thus, within each PU 501, there is only one PH 512 and only one picture 514 corresponding to that PH 512. That is, a PU 501 contains a single or isolated picture 514.
[0090] Each picture 514 includes one or more slices 518. A slice 518 is an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture (e.g., picture 514). Each slice 518 is contained exclusively in a single NAL unit (e.g., a VCL NAL unit). In one embodiment, a single NAL unit is associated with or has a layer ID 515. A tile (not shown) is a rectangular region of CTUs within a particular tile column and a particular tile row within a picture (e.g., picture 514). A tile is a partitioned portion of a picture created by horizontal and vertical boundaries. A tile may be rectangular and / or square. Specifically, a tile includes four sides connected at right angles. The four sides include two pairs of parallel sides. Furthermore, the sides of each pair of parallel sides are of equal length. Thus, a tile may be any rectangular shape, with a square being a special case of a rectangle with all four sides of equal length. An image / picture may include one or more tiles. A CTU (not shown) is a CTB for luma samples of a picture with three sample arrays, two corresponding CTBs for chroma samples, or a CTB for samples of a monochrome image or picture coded using three separate color planes and a syntax structure used to code the samples. A CTB (not shown) is an N x N block of samples, for some value of N, such that the division of a component into CTBs is a partitioning. A block (not shown) is an M x N (M columns x N rows) array of samples (e.g., pixels) or an M x N array of transform coefficients.
[0091] The pictures 514 and their slices 518 contain data related to the image or video being encoded or decoded. Thus, the pictures 514 and their slices 518 may simply be referred to as the payload or data carried within the bitstream 500. The PH 512 and slice headers 520 may include flags 522. The flags 522 may be RPL flags, SAO flags, or ALF flags, as described below.
[0092] Although the VVC specification defines only a few picture-level syntax elements, there are additional syntax elements whose values, in common use, can vary between slices of the same picture but are typically the same for all slices of the same picture. Examples of such syntax elements are syntax elements related to the RPL, congruent chroma code flag, SAO enable flag, ALF enable flag and parameter, LMCS enable flag and parameter, and scaling list enable flag and parameter. Non-picture-level syntax elements such as these are not signaled in the PH and must be repeated in the slice headers of all slices of the same picture, even if they have the same value for all slices of those slices. In other words, in some approaches, these syntax elements were signaled in the slice header because the data they carried could vary from slice to slice, but in most cases, they are the same for the entire picture that contains the slice. As a result, these elements are signaled several times per picture, but their values are typically the same, which is redundant and wastes bits in the encoded bitstream.
[0093] Disclosed herein are embodiments for signaling non-picture level syntax at the picture level. In embodiments, syntax elements are included in picture headers if the syntax elements are the same and in slice headers if the syntax elements vary. However, in some embodiments, syntax elements may not be included in both. First, non-picture level syntax elements may be present in the PH. Non-picture level syntax elements are syntax elements at a level of the video bitstream other than the picture level. Second, for each category of non-picture level syntax elements, a flag specifies when the syntax elements of that category are present in the PH or slice header. The flag may be present in the PH. Non-picture level syntax elements include those related to signaling the RPL, joint Cb Cr code flag, SAO tool enable and parameters, ALF tool enable and parameters, LMCS tool enable and parameters, and scaling list tool enable and parameters. Third, if non-picture level syntax elements are present in the PH, the corresponding syntax elements are not present in any slices of the picture associated with the picture header containing those syntax elements. The values of non-picture level syntax elements present in the PH apply to all slices of the picture associated with the picture header containing those syntax elements. Fourth, if a non-picture level syntax element is not present in the PH, the corresponding syntax element may be present in the slice headers of the slices of the picture associated with the picture header. Moving the signaling of non-picture level syntax elements to the picture level reduces redundancy and results in fewer wasted bits in the encoded bitstream. [Table 1-1] [Table 1-2] [Table 1-3] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6]
[0094] PH RBSP Semantics The PH contains information common to all slices of a coded picture, such that the next VCL NAL unit in decoding order is the first coded slice.
[0095] pic_type indicates the characterization of the coded picture as listed in Table 1 for a given value of pic_type. Values of pic_type shall be equal to 0 to 5 (inclusive) in bitstreams conforming to this version of this specification. Other values of pic_type are reserved for future use by ITU-T ISO / IEC. Decoders conforming to this version of this specification shall ignore reserved values of pic_type. [Table 3]
[0096] The pic_parameter_set_id specifies the value of the pps_pic_parameter_set_id of the PPS in use. The value of pic_parameter_set_id is in the range of 0 to 63 (inclusive).
[0097] non_reference_picture_flag equal to 1 specifies that the picture associated with PH is never used as a reference picture. non_reference_picture_flag equal to 0 specifies that the picture may or may not be used as a reference picture.
[0098] colour_plane_id specifies the colour plane associated with the picture associated with PH when separate_colour_plane_flag is equal to 1. The value of colour_plane_id is in the range 0 to 2 (inclusive). colour_plane_id values 0, 1 and 2 correspond to the Y, Cb and Cr planes respectively. There is no dependency between the decoding process of pictures with different colour_plane_id values.
[0099] pic_order_cnt_lsb specifies the picture order count modulo MaxPicOrderCntLsb for the picture associated with the PH. The length of the pic_order_cnt_lsb syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits. The value of pic_order_cnt_lsb is in the range of 0 to MaxPicOrderCntLsb-1 (inclusive).
[0100] recovery_poc_cnt Specifies the recovery point of decoded pictures in output order. If there is a picture picA after the current GDR picture in decoding order in the CVS with a PicOrderCntVal equal to the PicOrderCntVal of the current GDR picture plus the value of recovery_poc_cnt, then picture picA is called the recovery point picture. Otherwise, the first picture in output order with a PicOrderCntVal greater than the PicOrderCntVal of the current picture plus the value of recovery_poc_cnt is called the recovery point picture. A recovery point picture does not precede the current GDR picture in decoding order. The value of recovery_poc_cnt is in the range of 0 to MaxPicOrderCntLsb-1 (inclusive).
[0101] The variable RpPicOrderCntVal is derived as follows: RpPicOrderCntVal=PicOrderCntVal+recovery_poc_cnt.
[0102] The no_output_of_prior_pics_flag affects the output of previously decoded pictures in the decode picture buffer after decoding a CLVSS picture that is not the first picture in the bitstream, as specified in Annex C.
[0103] pic_output_flag affects the decoded picture output and removal process as specified in Annex C. If pic_output_flag is not present, it is inferred to be equal to 1.
[0104] pic_rpl_present_flag equal to 1 specifies that RPL signaling is present in the PH. pic_rpl_present_flag equal to 0 specifies that RPL signaling is not present in the PH but may be present in the slice header of the picture's slice. If not present, the value of pic_rpl_present_flag is inferred to be equal to 0. RPL signaling is the RPL information included in the video bitstream 500.
[0105] pic_rpl_sps_flag[i] equal to 1 specifies that the RPL i of picture is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures, where listIdx is equal to i in the SPS. ref_pic_list_sps_flag[i] equal to 0 specifies that the Reference Picture List i of picture is derived based on the ref_pic_list_struct(listIdx,rplsIdx) syntax structure, where listIdx is equal to i directly included in the picture header.
[0106] If pic_rpl_sps_flag[i] is not present, the following applies: if num_ref_pic_lists_in_sps[i] is equal to 0, the value of pic_rpl_sps_flag[i] is inferred to be equal to 0. Otherwise (num_ref_pic_lists_in_sps[i] is greater than 0) and rpl1_idx_present_flag is equal to 0, the value of pic_rpl_sps_flag[1] is inferred to be equal to pic_rpl_sps_flag[0]. Otherwise, the value of pic_rpl_sps_flag[i] is inferred to be equal to pps_ref_pic_list_sps_idc[i]-1.
[0107] pic_rpl_idx[i] specifies the index into the list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i contained in the SPS of the ref_pic_list_struct(listIdx,rplsIdx) syntax structure with listIdx equal to i that is used to derive RPL i for the current picture. The syntax element pic_rpl_idx[i] is represented in Ceil(Log2(num_ref_pic_lists_in_sps[i])) bits. If not present, the value of pic_rpl_idx[i] is inferred to be equal to 0. The value of pic_rpl_idx[i] is in the range 0 to num_ref_pic_lists_in_sps[i]-1 (inclusive). If pic_rpl_sps_flag[i] is equal to 1 and num_ref_pic_lists_in_sps[i] is equal to 1, the value of pic_rpl_idx[i] is inferred to be equal to 0. If pic_rpl_sps_flag[i] is equal to 1 and rpl1_idx_present_flag is equal to 0, the value of pic_rpl_idx[1] is inferred to be equal to pic_rpl_idx[0].
[0108] The variable PicRplsIdx[i] is derived as follows: PicRplsIdx[i]=pic_rpl_sps_flag[i] ? pic_rpl_idx[i]:num_ref_pic_lists_in_sps[i].
[0109] pic_poc_lsb_lt[i][j] specifies the value of the picture order count modulo MaxPicOrderCntLsb of the j-th LTRP entry in the ith reference picture list for the picture associated with the PH. The length of the pic_poc_lsb_lt[i][j] syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits.
[0110] The variable PicPocLsbLt[i][j] is derived as follows: PicPocLsbLt[i][j]=ltrp_in_slice_header_flag[i][PicRplsIdx[i]] ? pic_poc_lsb_lt[i][j]:rpls_poc_lsb_lt[listIdx][PicRplsIdx[i]][j].
[0111] pic_delta_poc_msb_present_flag[i][j] equal to 1 specifies that pic_delta_poc_msb_cycle_lt[i][j] is present. pic_delta_poc_msb_present_flag[i][j] equal to 0 specifies that pic_delta_poc_msb_cycle_lt[i][j] is not present.
[0112] Let prevTid0Pic be the previous picture in decoding order that has the same nuh_layer_id as PH, has TemporalId equal to 0, and is not a RASL or RADL picture. Let setOfPrevPocVals be the set consisting of: the PicOrderCntVal of prevTid0Pic; the PicOrderCntVal of each picture referenced by an entry in RefPicList[0] or RefPicList[1] for prevTid0Pic and that has the same nuh_layer_id as the current picture; and the PicOrderCntVal of each picture that is after prevTid0Pic in decoding order, has the same nuh_layer_id as the current picture, and precedes the current picture in decoding order.
[0113] If there are two or more values in setOfPrevPocVals whose value modulo MaxPicOrderCntLsb is equal to PicPocLsLt[i][j], then the value of pic_delta_poc_msb_present_flag[i][j] is equal to 1.
[0114] pic_delta_poc_msb_cycle_lt[i][j] specifies the value of the variable PicFullPocLt[i][j] as follows: [Table 4]
[0115] The value of pic_delta_poc_msb_cycle_lt[i][j] is in the range of 0 to 2(32-log2_max_pic_order_cnt_lsb_minus4-4), inclusive. If not present, the value of pic_delta_poc_msb_cycle_lt[i][j] is inferred to be equal to 0.
[0116] pic_temporal_mvp_enabled_flag specifies whether temporal MVP can be used for inter prediction. If pic_temporal_mvp_enabled_flag is equal to 0, the picture syntax elements associated with the picture header constrain the temporal MVP from being used in decoding the picture. Otherwise (pic_temporal_mvp_enabled_flag is equal to 1), the temporal MVP may be used in decoding the picture.
[0117] If pic_temporal_mvp_enabled_flag is not present, the following applies: if sps_temporal_mvp_enabled_flag is equal to 0, the value of pic_temporal_mvp_enabled_flag is inferred to be equal to 0. Otherwise (sps_temporal_mvp_enabled_flag is equal to 1), the value of pps_temporal_mvp_enabled_flag is inferred to be equal to pps_temporal_mvp_enabled_idc-1.
[0118] pic_level_joint_cbcr_sign_flag equal to 1 specifies that slice_joint_cbcr_sign_flag is not present in the slice header. pic_level_joint_cbcr_sign_flag equal to 0 specifies that slice_joint_cbcr_sign_flag is present in the slice header. If not present, the value of pic_level_joint_cbcr_sign_flag is inferred to be equal to 0.
[0119] pic_level_alf_enabled_flag equal to 1 specifies that ALF is enabled for all slices belonging to the picture associated with PH and may be applied to Y, Cb, or Cr color components within the slice. pic_level_alf_enabled_flag equal to 0 specifies that ALF may be disabled for one or more, or all slices belonging to the picture associated with PH. If not present, pic_level_alf_enabled_flag is inferred to be equal to 0.
[0120] pic_num_alf_aps_ids_luma specifies the number of ALF APSs referenced by slices belonging to the picture associated with the PH. The value of slice_num_alf_aps_ids_luma is in the range of 0 to 7 (inclusive).
[0121] pic_alf_aps_id_luma[i] specifies the adaptation_parameter_set_id of the i-th ALF APS referenced by the luma component of the slice of the picture associated with PH.
[0122] pic_alf_chroma_idc equal to 0 specifies that ALF is not applied to the Cb and Cr color components. pic_alf_chroma_idc equal to 1 indicates that ALF is applied to the Cb color component. pic_alf_chroma_idc equal to 2 indicates that ALF is applied to the Cr color component. pic_alf_chroma_idc equal to 3 indicates that ALF is applied to the Cb and Cr color components. If pic_alf_chroma_idc is not present, it is inferred to be equal to 0.
[0123] pic_alf_aps_id_chroma specifies the adaptation_parameter_set_id of the ALF APS referenced by the chroma components of the slice of the picture associated with the picture header.
[0124] pic_level_lmcs_enabled_flag equal to 1 specifies that luma mapping with chroma scaling is enabled for all slices belonging to the picture associated with the picture header. pic_level_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling may be disabled for one, multiple, or all slices belonging to the picture associated with the picture header. If not present, the value of pic_level_lmcs_enabled_flag is inferred to be equal to 0.
[0125] pic_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS referenced by the slice of the picture associated with the picture header.
[0126] pic_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for all slices belonging to the picture associated with the picture header. pic_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling may be disabled for one, multiple, or all slices belonging to the picture associated with the picture header. If pic_chroma_residual_scale_flag is not present, it is inferred to be equal to 0.
[0127] pic_level_scaling_list_present_flag equal to 1 specifies that the scaling list data used for the slices of the picture associated with the picture header is derived based on the scaling list data contained in the referenced scaling list APS. pic_level_scaling_list_present_flag equal to 0 specifies that the scaling list data used for one, more, or all slices of the picture associated with the picture header is the default scaling list data derived and specified in Section 7.4.3.16. If not present, the value of pic_level_scaling_list_present_flag is inferred to be equal to 0.
[0128] pic_scaling_list_aps_id specifies the adaptation_parameter_set_id of the scaling list APS.
[0129] Slice Header RBSP Semantics slice_rpl_sps_flag[i] equal to 1 specifies that RPL i of the current slice is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i in the SPS. slice_rpl_sps_flag[i] equal to 0 specifies that RPL i of the current slice is derived based on the ref_pic_list_struct(listIdx,rplsIdx) syntax structure with listIdx equal to i that is directly included in the slice header of the current picture.
[0130] If slice_rpl_sps_flag[i] is not present, the following applies: if pic_rpl_present_flag is equal to 1, the value of slice_rpl_sps_flag[i] is inferred to be equal to pic_rpl_sps_flag[i]. Otherwise, if num_ref_pic_lists_in_sps[i] is equal to 0, the value of slice_rpl_sps_flag[i] is inferred to be equal to 0. Otherwise (num_ref_pic_lists_in_sps[i] is greater than 0), if rpl1_idx_present_flag is equal to 0, the value of slice_rpl_sps_flag[1] is inferred to be equal to slice_rpl_sps_flag[0]. Otherwise, the value of slice_rpl_sps_flag[i] is inferred to be equal to pps_ref_pic_list_sps_idc[i]-1.
[0131] slice_rpl_idx[i] specifies the index into the list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i contained in the SPS of the ref_pic_list_struct(listIdx,rplsIdx) syntax structure with listIdx equal to i, which is used to derive reference picture list i for the current slice. The syntax element slice_rpl_idx[i] is represented by Ceil(Log2(num_ref_pic_lists_in_sps[i])) bits. The value of slice_rpl_idx[i] is in the range of 0 to num_ref_pic_lists_in_sps[i]-1, inclusive. If slice_rpl_sps_flag[i] is equal to 1 and num_ref_pic_lists_in_sps[i] is equal to 1, the value of slice_rpl_idx[i] is inferred to be equal to 0. If slice_rpl_sps_flag[i] is equal to 1 and rpl1_idx_present_flag is equal to 0, the value of slice_rpl_idx[1] is inferred to be equal to slice_rpl_idx[0].
[0132] The variable RplsIdx[i] is derived as follows: [Table 5]
[0133] slice_poc_lsb_lt[i][j] specifies the value of the picture order count modulo MaxPicOrderCntLsb of the j-th LTRP entry in the ith reference picture list for the current slice. The length of the slice_poc_lsb_lt[i][j] syntax element is log2_max_pic_order_cnt_lsb_minus4+4 bits.
[0134] The variable PocLsbLt[i][j] is derived as follows: [Table 6]
[0135] slice_delta_poc_msb_present_flag[i][j] equal to 1 specifies that slice_delta_poc_msb_cycle_lt[i][j] is present. slice_delta_poc_msb_present_flag[i][j] equal to 0 specifies that slice_delta_poc_msb_cycle_lt[i][j] is not present.
[0136] Let prevTid0Pic be the previous picture in decoding order that has the same nuh_layer_id as the current picture, has TemporalId equal to 0, and is not a RASL or RADL picture. Let setOfPrevPocVals be the set consisting of: the PicOrderCntVal of prevTid0Pic, the PicOrderCntVal of each picture referenced by an entry in RefPicList[0] or RefPicList[1] for prevTid0Pic and that has the same nuh_layer_id as the current picture, and the PicOrderCntVal of each picture that is after prevTid0Pic in decoding order, has the same nuh_layer_id as the current picture, and precedes the current picture in decoding order.
[0137] When pic_rpl_present_flag is equal to 0 and there are two or more values in setOfPrevPocVals such that the value modulo MaxPicOrderCntLsb is equal to PocLsbLt[i][j], the value of slice_delta_poc_msb_present_flag[i][j] is equal to 1.
[0138] slice_delta_poc_msb_cycle_lt[i][j] specifies the value of the variable FullPocLt[i][j] as follows: [Table 7]
[0139] The value of slice_delta_poc_msb_cycle_lt[i][j] is in the range of 0 to 2(32-log2_max_pic_order_cnt_lsb_minus4-4), inclusive. If not present, the value of slice_delta_poc_msb_cycle_lt[i][j] is inferred to be equal to 0.
[0140] slice_joint_cbcr_sign_flag specifies whether the co-located residual samples of both chroma components have inverted signs in transform units where tu_joint_cbcr_residual_flag[x0][y0] is equal to 1. If tu_joint_cbcr_residual_flag[x0][y0] is equal to 1 for a transform unit, slice_joint_cbcr_sign_flag equal to 0 specifies that the sign of each residual sample of the Cr (or Cb) component is the same as the sign of its co-located Cb (or Cr) residual sample, and slice_joint_cbcr_sign_flag equal to 1 specifies that the sign of each residual sample of the Cr (or Cb) component is given by the inverted sign of its co-located Cb (or Cr) residual sample. If not present, the value of slice_joint_cbcr_sign_flag is inferred to be equal to pic_level_joint_cbcr_sign_flag.
[0141] slice_sao_luma_flag equal to 1 specifies that SAO is enabled for the luma component in the current slice. slice_sao_luma_flag equal to 0 specifies that SAO is disabled for the luma component in the current slice. If slice_sao_luma_flag is not present, it is inferred to be equal to pic_level_sao_luma_enabled_flag.
[0142] slice_sao_chroma_flag equal to 1 specifies that SAO is enabled for the chroma components in the current slice. slice_sao_chroma_flag equal to 0 specifies that SAO is disabled for the chroma components in the current slice. If slice_sao_chroma_flag is not present, it is inferred to be equal to pic_level_sao_chroma_enabled_flag.
[0143] slice_alf_enabled_flag equal to 1 specifies that ALF is enabled and may be applied to the Y, Cb, or Cr color components in the slice. slice_alf_enabled_flag equal to 0 specifies that ALF is disabled for all color components in the slice. If not present, the value of slice_alf_enabled_flag is inferred to be equal to pic_level_alf_enabled_flag.
[0144] slice_num_alf_aps_ids_luma specifies the number of ALF APSs referenced by the slice. If slice_alf_enabled_flag is equal to 1 and slice_num_alf_aps_ids_luma is not present, the value of slice_num_alf_aps_ids_luma is inferred to be equal to the value of pic_num_alf_aps_ids_luma. The value of slice_num_alf_aps_ids_luma is in the range of 0 to 7 (inclusive).
[0145] slice_alf_aps_id_luma[i] specifies the adaptation_parameter_set_id of the i-th ALF APS to which the luma component of the slice refers. The TemporalId of the APS NAL unit whose aps_params_type is equal to ALF_APS and whose adaptation_parameter_set_id is equal to slice_alf_aps_id_luma[i] is less than or equal to the TemporalId of the slice NAL unit being coded. If slice_alf_enabled_flag is equal to 1 and slice_alf_aps_id_luma[i] is not present, the value of slice_alf_aps_id_luma[i] is inferred to be equal to the value of pic_alf_aps_id_luma[i].
[0146] For slices within intra slices and IRAP pictures, slice_alf_aps_id_luma[i] must not reference an ALF APS associated with a picture other than the picture containing the intra slice or IRAP picture.
[0147] slice_alf_chroma_idc equal to 0 specifies that ALF is not applied to the Cb and Cr color components. slice_alf_chroma_idc equal to 1 indicates that ALF is applied to the Cb color component. slice_alf_chroma_idc equal to 2 indicates that ALF is applied to the Cr color component. slice_alf_chroma_idc equal to 3 indicates that ALF is applied to the Cb and Cr color components. If slice_alf_chroma_idc is not present, it is inferred to be equal to pic_alf_chroma_idc.
[0148] slice_alf_aps_id_chroma specifies the adaptation_parameter_set_id of the ALF APS to which the chroma components of the slice refer. The TemporalId of the APS NAL unit whose aps_params_type is equal to ALF_APS and whose adaptation_parameter_set_id is equal to slice_alf_aps_id_chroma is less than or equal to the TemporalId of the slice NAL unit being coded. If slice_alf_enabled_flag is equal to 1 and slice_alf_aps_id_chroma is not present, the value of slice_alf_aps_id_chroma is inferred to be equal to the value of pic_alf_aps_id_chroma.
[0149] For slices in intra slices and IRAP pictures, slice_alf_aps_id_chroma must not reference an ALF APS associated with another picture other than the picture containing the intra slice or IRAP picture.
[0150] slice_lmcs_enabled_flag equal to 1 specifies that luma mapping with chroma scaling is enabled for the current slice. slice_lmcs_enabled_flag equal to 0 specifies that luma mapping with chroma scaling is not enabled for the current slice. If slice_lmcs_enabled_flag is not present, it is inferred to be equal to pic_lmcs_enabled_flag.
[0151] slice_lmcs_aps_id specifies the adaptation_parameter_set_id of the LMCS APS to which the slice refers. The TemporalId of the APS NAL unit whose aps_params_type is equal to LMCS_APS and whose adaptation_parameter_set_id is equal to slice_lmcs_aps_id is less than or equal to the TemporalId of the slice NAL unit being coded. If slice_lmcs_enabled_flag is equal to 1 and slice_lmcs_aps_id is not present, the value of slice_lmcs_aps_id is inferred to be equal to the value of pic_lmcs_aps_id.
[0152] If present, the value of slice_lmcs_aps_id is the same for all slices of a picture.
[0153] slice_chroma_residual_scale_flag equal to 1 specifies that chroma residual scaling is enabled for the current slice. slice_chroma_residual_scale_flag equal to 0 specifies that chroma residual scaling is not enabled for the current slice. If slice_chroma_residual_scale_flag is not present, it is inferred to be equal to pic_chroma_residual_scale_flag.
[0154] slice_scaling_list_present_flag equal to 1 specifies that the scaling list data used for the current slice is derived based on the scaling list data contained in the referenced scaling list APS. slice_scaling_list_present_flag equal to 0 specifies that the scaling list data used for the current picture is the default scaling list data specified in Section 7.4.3.16. If not present, the value of slice_scaling_list_present_flag is inferred to be equal to pic_level_scaling_list_present_flag.
[0155] slice_scaling_list_aps_id specifies the adaptation_parameter_set_id of the scaling list APS. The TemporalId of the APS NAL unit whose aps_params_type is equal to SCALING_APS and whose adaptation_parameter_set_id is equal to slice_scaling_list_aps_id is less than or equal to the TemporalId of the slice NAL unit being coded. If slice_scaling_list_enabled_flag is equal to 1 and slice_scaling_list_aps_id is not present, the value of slice_scaling_list_aps_id is inferred to be equal to the value of pic_scaling_list_aps_id.
[0156] 6 is a flowchart illustrating a method 600 for decoding a bitstream according to the first embodiment. The decoder 400 may implement the method 600. In step 610, a video bitstream including an RPL flag is received. An RPL flag equal to a first value specifies that RPL signaling is present in the PH. An RPL flag equal to a second value specifies that RPL signaling is not present in the PH but may be present in the slice header. Finally, in step 620, the coded picture is decoded using the RPL flag to obtain a decoded picture.
[0157] Method 600 may implement additional embodiments. For example, the first value is 1. The second value is 0. The bitstream further includes an RPL SPS flag, which specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i in the SPS, or based on one of the directly included ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i. The bitstream further includes an RPL index, which specifies an index into a list of ref_pic_list_struct(listIdx,rplsIdx) syntax structures with listIdx equal to i included in the SPS of the ref_pic_list_struct(listIdx,rplsIdx) syntax structure with listIdx equal to i that is used to derive RPL i of the current picture. The decoded picture is displayed on the display of the electronic device.
[0158] 7 is a flowchart illustrating a method 700 for encoding a bitstream according to a first embodiment. The encoder 300 may implement the method 700. In step 710, an RPL flag is generated. An RPL flag equal to a first value specifies that RPL signaling is present in the PH. An RPL flag equal to a second value specifies that RPL signaling is not present in the PH, but may be present in the slice header. Finally, in step 730, the video bitstream is stored for communication to a video decoder.
[0159] Method 700 may implement additional embodiments. For example, the first value is 1. The second value is 0. An RPL SPS flag is generated, where the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS, or based on one of the directly included ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i. An RPL index is generated, where the RPL index specifies an index into a list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i included in the sequence parameter set (SPS) of the ref_pic_list_struct(listIdx, rplsIdx) syntax structure with listIdx equal to i that is used to derive RPL i of the current picture.
[0160] 8 is a flowchart illustrating a method 800 for decoding a bitstream according to a second embodiment. The decoder 400 may implement the method 800. In step 810, a video bitstream including an SAO flag is received. An SAO flag equal to a first value specifies that SAO signaling is present in the PH, and an SAO flag equal to a second value specifies that SAO signaling is not present in the PH but may be present in a slice header. Finally, in step 820, the coded picture is decoded using the SAO flag to obtain a decoded picture. The method 800 may implement additional embodiments. For example, the decoded picture may be displayed on a display of an electronic device.
[0161] 9 is a flowchart illustrating a method 900 for encoding a bitstream according to the second embodiment. The encoder 300 may implement the method 900. At step 910, an SAO flag is generated. An SAO flag equal to a first value specifies that SAO signaling is present in the PH, and an SAO flag equal to a second value specifies that SAO signaling is not present in the PH but may be present in the slice header. At step 920, an RPL flag is encoded into the video bitstream. Finally, at step 930, the video bitstream is stored for communication to a video decoder.
[0162] FIG. 10 is a flowchart illustrating a method 1000 for decoding a bitstream according to a third embodiment. The decoder 400 may implement the method 1000. In step 1010, a video bitstream including an ALF flag is received. An ALF flag equal to a first value specifies that ALF signaling is present in the PH, and an ALF flag equal to a second value specifies that ALF signaling is not present in the PH but may be present in the slice header. Finally, in step 1020, the coded picture is decoded using the ALF flag to obtain a decoded picture. The method 1000 may implement additional embodiments. For example, the decoded picture may be displayed on a display of an electronic device.
[0163] 11 is a flowchart illustrating a method 1100 for encoding a bitstream according to the third embodiment. The encoder 300 may implement the method 1100. In step 1110, an ALF flag is generated. An ALF flag equal to a first value specifies that ALF signaling is present in the PH, and an ALF flag equal to a second value specifies that ALF signaling is not present in the PH but may be present in the slice header. In step 1120, the ALF flag is encoded into the video bitstream. Finally, in step 1130, the video bitstream is stored for communication to a video decoder.
[0164] 12 is a schematic diagram of a video coding device 1200 (e.g., video encoder 300 or video decoder 400) according to an embodiment of the present disclosure. Video coding device 1200 is suitable for implementing the disclosed embodiments. Video coding device 1200 includes an ingress port 1210 and an Rx 1220 for receiving data; a processor, logic unit, or CPU 1230 for processing the data; a Tx 1240 and an egress port 1250 for transmitting the data; and a memory 1260 for storing the data. Video coding device 1200 may also include optical and electrical components coupled to ingress port 1210, receiver unit 1220, transmitter unit 1240, and egress port 1250 for inputting and outputting optical or electrical signals.
[0165] The processor 1230 is implemented in hardware and software. The processor 1230 may be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 1230 communicates with the ingress port 1210, the Rx 120, the Tx 1240, the egress port 1250, and the memory 1260. The processor 1230 includes a coding module 1270. The coding module 1270 implements the disclosed embodiments. For example, the coding module 1270 implements, processes, prepares, or provides various codec functions. Thus, the inclusion of the coding module 1270 provides a substantial improvement to the functionality of the video coding device 1200 and performs transformations of the video coding device 1200 into different states. Alternatively, the coding module 1270 is implemented as instructions stored in the memory 1260 and executed by the processor 1230.
[0166] Video coding device 1200 may also include input / output devices 1280 for communicating data to and from a user. Input / output devices 1280 may include output devices such as a display for displaying video data, speakers for outputting audio data, etc. Input / output devices 1280 may also include input devices such as a keyboard, mouse, or trackball, or corresponding interfaces for interacting with such output devices.
[0167] Memory 1260 may include one or more disks, tape drives, and solid state drives, may be used as overflow data storage, store programs when such programs are selected for execution, and store instructions and data retrieved during program execution. Memory 1260 may be volatile and / or non-volatile and may be ROM, RAM, TCAM, or SRAM.
[0168] 13 is a schematic diagram of one embodiment of a coding means 1300. In one embodiment, the coding means 1300 is implemented in a video coding device 1302 (e.g., video encoder 300 or video decoder 400). The video coding device 1302 includes a receiving means 1301. The receiving means 1301 is configured to receive a picture to encode or a bitstream to decode.
[0169] The video coding device 1302 includes a transmitting means 1307 coupled to the receiving means 1301. The transmitting means 1307 is configured to transmit the bitstream to a decoder or transmit the decoded image to a display means (e.g., one of the I / O devices 1280). The video coding device 1302 includes a storage means 1303. The storage means 1303 is coupled to at least one of the receiving means 1301 or the transmitting means 1307. The storage means 1303 is configured to store instructions. The video coding device 1302 also includes a processing means 1305. The processing means 1305 is coupled to the storage means 1303. The processing means 1305 is configured to execute the instructions stored in the storage means 1303 to perform the methods disclosed herein.
[0170] In one embodiment, the receiving means receives a video bitstream including an RPL flag, the RPL flag specifying whether RPL signaling is present or absent in a PH or specifying that RPL signaling may be present in a slice header, and the processing means decodes the coded picture using the RPL flag to obtain a decoded picture.
[0171] The term "about," unless otherwise specified, refers to a range including ±10% of the subsequent number. While several embodiments are provided in this disclosure, it will be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples herein are intended to be illustrative rather than restrictive, and the intention is not to be limited to the details provided herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0172] Furthermore, techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, components, techniques, or methods without departing from the scope of the present disclosure. Other items shown or described as coupled may be directly coupled, or may be indirectly coupled or in communication through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and alterations may be discernible by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.
Claims
1. 1. A method implemented by a video decoder: receiving, by the video decoder, a video bitstream including a reference picture list (RPL) flag, wherein the RPL flag equal to a first value specifies that RPL signaling is present in a picture header (PH), and the RPL flag equal to a second value specifies that RPL signaling is not present in the PH and may be present in a slice header; and decoding, by the video decoder, the coded picture using the RPL flag to obtain a decoded picture. method.
2. The method of claim 1 , wherein the first value is one.
3. The method of claim 1 or 2, wherein the second value is 0.
4. The bitstream further includes an RPL Sequence Parameter Set (SPS) flag, wherein the RPL SPS flag specifies that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS, or based on one of the directly included ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i.
4. The method according to any one of claims 1 to 3.
5. 5. The method of claim 1, wherein the bitstream further comprises an RPL index, the RPL index specifying an index into a list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i contained in a sequence parameter set (SPS) of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i that are used to derive RPL i for the current picture.
6. 6. The method of claim 1, further comprising the step of displaying the decoded picture on a display of the electronic device.
7. 7. The method of claim 1, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.
8. 1. A method implemented by a video encoder, comprising: generating a Reference Picture List (RPL) flag, wherein the RPL flag equal to a first value specifies that RPL signaling is present in a Picture Header (PH), and the RPL flag equal to a second value specifies that RPL signaling is not present in the PH and may be present in a Slice Header; encoding, by the video encoder, the RPL flag into a video bitstream; storing, by the video encoder, the video bitstream for communication to a video decoder. method.
9. The method of claim 8 , wherein the first value is one.
10. 10. The method of claim 8 or 9, wherein the second value is 0.
11. generating an RPL Sequence Parameter Set (SPS) flag, the RPL SPS flag specifying that RPL i is derived based on one of the ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i in the SPS, or based on one of the directly contained ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i; 11. The method according to any one of claims 8 to 10.
12. 12. The method of claim 8, further comprising generating an RPL index, wherein the RPL index specifies an index into a list of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i contained in a sequence parameter set (SPS) of ref_pic_list_struct(listIdx, rplsIdx) syntax structures with listIdx equal to i that are used to derive the RPL i of the current picture.
13. 13. The method of claim 8, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.
14. 1. A method implemented by a video decoder: receiving, by the video decoder, a video bitstream including a sample adaptive offset (SAO), wherein the SAO flag equal to a first value specifies that SAO signaling is present in a picture header (PH), and the SAO flag equal to a second value specifies that SAO signaling is not present in the PH and may be present in a slice header; and decoding, by the video decoder, the coded picture using the SAO flag to obtain a decoded picture. method.
15. 15. The method of claim 14, further comprising displaying the decoded picture on a display of the electronic device.
16. 16. The method of claim 14 or 15, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of the picture associated with the PH that contains the syntax element.
17. 1. A method implemented by a video encoder, comprising: generating a sample adaptive offset (SAO) flag, wherein the SAO flag equal to a first value specifies that SAO signaling is present in a picture header (PH), and the SAO flag equal to a second value specifies that SAO signaling is not present in the PH and may be present in a slice header; encoding, by the video encoder, the SAO flag into a video bitstream; storing, by the video encoder, the video bitstream for communication to a video decoder. method.
18. 18. The method of claim 17, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.
19. 1. A method implemented by a video decoder: by the video decoder receiving a video bitstream including an adaptive loop filter (ALF) flag, wherein the ALF flag equal to a first value specifies that ALF signaling is present in a picture header (PH), and the ALF flag equal to a second value specifies that ALF signaling is not present in the PH but may be present in a slice header; and decoding the coded picture using the ALF flag to obtain a decoded picture. method.
20. 20. The method of claim 19, further comprising displaying the decoded picture on a display of the electronic device.
21. 21. The method of claim 19 or 20, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of the picture associated with the PH that contains the syntax element.
22. a memory configured to store instructions; a processor coupled to the memory and configured to execute the instructions to perform any one of claims 1 to 7, 14 to 16, and 19 to 21. Video decoder.
23. A computer program product having computer-executable instructions for storage on a non-transitory medium which, when executed by a processor, causes a video decoder to perform any one of claims 1 to 7, 14 to 16, and 19 to 21.
24. 1. A method implemented by a video encoder, comprising: generating an adaptive loop filter (ALF) flag, wherein the ALF flag equal to a first value specifies that ALF signaling is present in a picture header (PH), and the ALF flag equal to a second value specifies that ALF signaling is not present in the PH but may be present in a slice header; encoding, by the video encoder, the ALF flag into a video bitstream; storing, by the video encoder, the video bitstream for communication to a video decoder. method.
25. 25. The method of claim 24, wherein when a non-picture level syntax element is present in the PH, no corresponding syntax element is present in any slice of a picture associated with the PH that contains the syntax element.
26. a memory configured to store instructions; a processor coupled to the memory and configured to execute the instructions to perform any one of claims 8-13, 17-18 and 24-25. Video encoder.
27. A computer program product having computer-executable instructions for storage on a non-transitory medium which, when executed by a processor, causes a video encoder to perform any one of items 8-13, 17-18, and 24-25.
28. 1. A method implemented by a video decoder: receiving, by the video decoder, a video bitstream including a syntax element specifying that information may or may not be present in a picture header (PH), or specifying that the information may or may not be present in a slice header; and decoding, by the video decoder, the coded picture using the syntax elements to obtain a decoded picture. method.
29. 29. A computer readable medium having a bitstream encoded or decoded by the method of any one of claims 1 to 21, 24 to 25 and 28.