Signaling of subpicture level information in video coding
By using sub-image indexes instead of IDs in the VVC design and standardizing SEI message nesting rules, problems in sub-image level information signaling notifications are solved, improving video encoding and decoding efficiency and flexibility, and supporting the flexibility of multi-layer video encoding and decoding.
Patent Information
- Application Number
- CN202180041817.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-09
- Filing Date
- 2021-06-08
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2041-06-08
AI Technical Summary
In existing VVC designs, sub-picture level information signaling notifications suffer from issues such as improper handling of padding payloads due to sub-picture ID changes, improper association of SLI SEI messages, and insufficient nesting of SEI messages, which affect the efficiency and flexibility of video encoding and decoding.
By using sub-image indexes instead of sub-image IDs, the removal of padding payloads during sub-bitstream extraction is prohibited. This standardizes SEI message nesting rules, ensures the applicability of SLI SEI messages, and explicitly defines NAL unit types in scalable nested SEI messages, supporting the flexibility of multi-layer video encoding and decoding.
It improves the accuracy and efficiency of sub-picture level information signaling notification in the video encoding and decoding process, supports the flexibility of multi-layer video encoding and decoding, and optimizes the sub-bitstream extraction process.
Smart Images

Figure CN115699740B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application is based on International Patent Application No. PCT / US2021 / 036349 filed June 8, 2021, which claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 036,743 filed June 9, 2020. All of the above-identified applications are hereby incorporated by reference in their entirety into this document. TECHNICAL FIELD
[0003] This patent document relates to image and video coding and decoding. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue growing. SUMMARY
[0005] This document discloses techniques that can be used by video encoders and decoders to process coded representations of video or images.
[0006] In one example aspect, a method of video processing is disclosed. The method includes performing a conversion between a video comprising one or more subpictures and a bitstream of the video, wherein during the conversion, one or more supplemental enhancement information messages with filler payload are processed according to a format rule, and wherein the format rule disallows the one or more supplemental enhancement information messages with filler payload in scalable-nested supplemental enhancement information messages.
[0007] In another example aspect, a method of video processing is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein during the conversion, one or more syntax elements are processed according to a format rule, and wherein the format rule specifies that the one or more syntax elements are used to indicate subpicture information for a layer of the video that has a picture with multiple subpictures.
[0008] In another example aspect, a method of video processing is disclosed. The method includes performing a conversion between a video comprising multiple subpictures and a bitstream of the video, wherein during the conversion, scalable-nested supplemental enhancement information messages are processed according to a format rule, and wherein the format rule specifies that one or more subpicture indexes are used to associate one or more subpictures to the scalable-nested supplemental enhancement information messages.
[0009] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more subpictures and a bitstream of the video according to a format rule, wherein the format rule specifies that, responsive to a scalable-nested supplemental enhancement information message including one or more subpicture level information supplemental enhancement information messages, a first syntax element in the scalable-nested supplemental enhancement information message in the bitstream is set to a particular value, and wherein the particular value of the first syntax element indicates that the scalable-nested supplemental enhancement information message includes one or more scalable-nested supplemental enhancement information messages that apply to a particular output video layer set.
[0010] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising a plurality of subpictures and a bitstream of the video, wherein the conversion is according to a format rule that specifies that a scalable-nested supplemental enhancement information message is not allowed to include a first supplemental enhancement information message of a first payload type and a second supplemental enhancement information message of a second payload type.
[0011] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the conversion is performed according to a format rule that specifies, responsive to a supplemental enhancement information network abstraction layer unit including a scalable-nested supplemental enhancement information message, the supplemental enhancement information network abstraction layer unit including a network abstraction layer unit type equal to a prefix supplemental enhancement information network abstraction layer unit type, wherein the scalable-nested supplemental enhancement information message includes a supplemental enhancement information message that is not associated with a particular payload type.
[0012] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video and a bitstream of the video, wherein the conversion is performed according to a format rule that specifies, responsive to a supplemental enhancement information network abstraction layer unit including a scalable-nested supplemental enhancement information message, the supplemental enhancement information network abstraction layer unit including a network abstraction layer unit type equal to a suffix supplemental enhancement information network abstraction layer unit type, wherein the scalable-nested supplemental enhancement information message includes a supplemental enhancement information message that is associated with a particular payload type.
[0013] In another example aspect, a video processing method is disclosed. The method includes performing a conversion between a video comprising one or more subpictures or one or more subpicture sequences and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies whether or how scalable-nested supplemental enhancement information (SEI) is included in the coded representation.
[0014] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement a method recited above.
[0015] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder comprises a processor configured to implement a method described above.
[0016] In yet another example aspect, a computer readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0017] These and other features are described in this document. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 An example of raster-scan tile partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster-scan slices.
[0019] Figure 2 An example of rectangular slice partitioning of a picture is shown, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0020] Figure 3 An example of a picture partitioned into tiles and rectangular slices is shown, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0021] Figure 4 A picture partitioned into 15 tiles, 24 slices, and 24 sub-pictures is shown.
[0022] Figure 5 A block diagram of an example video processing system.
[0023] Figure 6 A block diagram of a video processing apparatus.
[0024] Figure 7 A flowchart of an example method of video processing.
[0025] Figure 8 is a block diagram illustrating a video coding system in accordance with some embodiments of the disclosure.
[0026] Figure 9 is a block diagram illustrating an encoder in accordance with some embodiments of the disclosure.
[0027] Figure 10 is a block diagram illustrating a decoder in accordance with some embodiments of the disclosure.
[0028] Figure 11 An example of a typical sub-picture based viewport dependent 360° video coding scheme is shown.
[0029] Figure 12 A viewport dependent 360° video coding scheme based on sub-pictures and spatial scalability is shown.
[0030] Figures 13 to 19 is a flowchart of an example method of processing video data. DETAILED DESCRIPTION
[0031] The use of section headings in this document is for ease of understanding and does not limit the applicability of techniques and embodiments disclosed in a section to only the section. Also, the use of H.266 terminology in some descriptions is merely for ease of understanding and is not intended to limit the scope of the disclosed techniques. Thus, the techniques described herein are applicable to other video codec protocols and designs as well. In this document, editorial changes to the text relative to the current draft of the VVC specification are shown by struck-out text for cancelled text and highlighted text (including boldface italics) for added text.
[0032] 1. INTRODUCTION
[0033] This document is related to video coding technology. Specifically, it is about specifying and signaling level information for sub-picture sequences. It can be applied to any video coding standard or non-standard video codec that supports single-layer video coding and multi-layer video coding, such as the Versatile Video Coding (VVC) that is under development.
[0034] 2. ABBREVIATIONS
[0035] APS (Adaptation Parameter Set) adaptation parameter set
[0036] AU (Access Unit) access unit
[0037] AUD (Access Unit Delimiter) access unit delimiter
[0038] AVC (Advanced Video Coding) advanced video coding
[0039] BP (Buffering Period) buffering period
[0040] CLVS (Coded Layer Video Sequence) coded layer video sequence
[0041] CPB (Coded Picture Buffer) coded picture buffer
[0042] CRA (Clean Random Access) clean random access
[0043] CTU (Coding Tree Unit) coding tree unit
[0044] CVS (Coded Video Sequence) coded video sequence
[0045] DPB (Decoded Picture Buffer) decoded picture buffer
[0046] DPS (Decoding Parameter Set) decoding parameter set
[0047] DUI (Decoding Unit Information) decoding unit information
[0048] EOB (End Of Bitstream) end of bitstream
[0049] EOS (End Of Sequence) end of sequence
[0050] GCI (General Constraints Information) general constraints information
[0051] GDR (Gradual Decoding Refresh) gradual decoding refresh
[0052] HEVC (High Efficiency Video Coding) high efficiency video coding
[0053] HRD (Hypothetical Reference Decoder) hypothetical reference decoder
[0054] IDR (Instantaneous Decoding Refresh) instantaneous decoding refresh
[0055] JEM (Joint Exploration Model) joint exploration model
[0056] MCTS (Motion-Constrained Tile Sets) motion-constrained tile sets
[0057] NAL (Network Abstraction Layer) network abstraction layer
[0058] OLS (Output Layer Set) output layer set
[0059] PH (Picture Header) picture header
[0060] PPS (Picture Parameter Set) picture parameter set
[0061] PT (Picture Timing) picture timing
[0062] PTL (Profile, Tier and Level) profile, tier and level
[0063] PU (Picture Unit) picture unit
[0064] RRP (Reference Picture Resampling) reference picture resampling
[0065] RBSP (Raw Byte Sequence Payload) raw byte sequence payload
[0066] SEI (Supplemental Enhancement Information) supplemental enhancement information
[0067] SH (Slice Header) slice header
[0068] SLI (Subpicture Level Information) subpicture level information
[0069] SPS (Sequence Parameter Set) sequence parameter set
[0070] SVC (Scalable Video Coding) scalable video coding
[0071] VCL (Video Coding Layer) video coding layer
[0072] VPS (Video Parameter Set) video parameter set
[0073] VTM (VVC Test Model) VVC test model
[0074] VUI (Video Usability Information) video usability information
[0075] VVC (Versatile Video Coding) versatile video coding
[0076] 3. Preliminary discussion
[0077] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263 standards, ISO / IEC produced the MPEG-1 and MPEG-4 Visual standards, and the two organizations jointly produced the H.262 / MPEG-2 Video standard and the H.264 / MPEG-4 Advanced Video Coding (AVC) standard and the H.265 / HEVC standard. From H.262 onwards, the video coding standards are based on the hybrid video coding structure, where temporal prediction plus transform coding are utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, the JVET has adopted a number of new methods and incorporated them into the reference software named Joint Exploration Model (JEM). The JVET meeting is held every quarter, and the new coding standard targets at 50% bitrate reduction compared to HEVC. The new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was also released at that time. As the VVC standardization is ongoing, new coding techniques are being adopted into the VVC standard in every JVET meeting. The VVC working draft and test model VTM are updated after every meeting. The VVC project now targets at the Final Draft International
[0078] 3.1. Picture partitioning schemes in HEVC
[0079] HEVC includes four different picture partitioning schemes, namely regular slices, non- independent slices, tiles, and Wavefront Parallel Processing (WPP), which can be applied for Maximum Transfer Unit (MTU) size matching, parallel processing, and reducing end-to-end delay.
[0080] A regular slice is similar to that in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependency across slice boundaries are disabled. Therefore, a regular slice can be reconstructed independently from other regular slices within the same picture (although there can still be inter-dependencies due to in-loop filtering operations).
[0081] Conventional slices are the only tool available for parallelization, which is also available in H.264 / AVC in almost the same form. Conventional slice-based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding a predictive coded picture, which is typically much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, using conventional slices can incur a lot of coding overhead due to the bit cost of the slice header and the prediction discontinuity across slice boundaries. In addition, due to the intra-picture independence of conventional slices and the fact that each conventional slice is encapsulated in its own NAL, conventional slices can also serve as the key mechanism for bitstream segmentation to match the MTU size requirement (compared to other tools mentioned below). In many cases, the goals of parallelization and MTU size matching are contradictory in terms of the requirements on the slice layout in a picture. The implementation of this situation leads to the development of the parallelization tools mentioned below.
[0082] Non-independent slices have short slice headers and allow bitstream segmentation at tree block boundaries without breaking any intra-picture prediction. Essentially, non-independent slices split a conventional slice into multiple NAL units, reducing the end-to-end delay by allowing a portion of a conventional slice to be sent before the encoding of the entire conventional slice is complete.
[0083] In WPP, a picture is segmented into single rows of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other segments. Parallel processing can be done by parallel decoding of CTB rows, with a two-CTB start delay to ensure that data related to CTBs above and to the right of the subject CTB can be obtained before the subject CTB is decoded. Using this staggered start (which looks like a wavefront when represented graphically), as many processors / cores as there are pictures containing CTB rows can be parallelized. Because intra-picture prediction between adjacent tree block rows is allowed, the inter-processor / core communication required to implement intra-picture prediction can be substantial. WPP segmentation does not result in additional NAL units compared to not applying WPP segmentation, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, conventional slices can be used with WPP, but at a certain coding overhead.
[0084] Tile definitions segment a picture into horizontal and vertical boundaries of tile columns and tile rows. Tile columns extend from the top of the picture to the bottom of the picture. Likewise, tile rows extend from the left side of the picture to the right side of the picture. The number of tiles in a picture can simply be derived by multiplying the number of tile columns by the number of tile rows.
[0085] The scan order of the CTBs is changed to the local scan order within the slice (in the order of CTB raster scan of the slice) before decoding the left-top CTB of the next slice in the order of slice raster scan of the picture. Similar to regular slices, slices break the inter-picture prediction dependencies as well as the entropy decoding dependencies. However, they do not need to be contained in independent NAL units (same as WPP in this regard); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case of a slice that spans multiple tiles, the inter-processor / core communication required for inter-picture prediction between processing units of neighboring slices is limited to communicating the shared slice header and loop filtering related to sharing of reconstructed samples and metadata. When more than one slice or WPP segment is contained in a slice, the entry point byte offset of each slice or WPP segment other than the first one in the slice is signaled in the slice header.
[0086] For simplicity, restrictions on the application of the four different picture partitioning schemes are specified in HEVC. For most profiles specified in HEVC, a given coded video sequence cannot contain both slices and wavefronts. For each slice and tile, one or both of the following conditions must be met: 1) all the coding tree blocks in the slice belong to the same tile; 2) all the coding tree blocks in the tile belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is used, if a slice starts within a CTB row, the slice must end in the same CTB row.
[0087] A recent revision of HEVC is specified in JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)". Within this revision, HEVC specifies three MCT-related SEI messages, namely the temporal MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nesting SEI message.
[0088] The temporal MCTS SEI message indicates the presence of MCTSs in the bitstream and signals the MCTSs. For each MCTS, the motion vectors are restricted to point to full-sample positions within the MCTS and to fractional-sample positions that only require full-sample positions within the MCTS for interpolation, and the use of motion vector candidates for temporal motion vector prediction that are derived from blocks outside the MCTS is not allowed. In this way, each MCTS can be decoded independently without the presence of slices that are not included in the MCTS.
[0089] The MCTS extraction information set SEI message provides supplemental information that can be used for MCTS sub-bitstream extraction (part of the semantics of the SEI message) to generate a bitstream conforming to an MCTS group. The information consists of multiple extraction information sets, each defining multiple MCTS groups and containing RBSP bytes of replacement VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and slice headers need to be slightly updated as one or all of the syntax elements related to slice address (including first_slice_segment_in_pic_flag and slice_segment_address) generally need to have different values.
[0090] 3.2. Partitioning of pictures in VVC
[0091] In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs covering a rectangular region of the picture. The CTUs in a tile are scanned in the tile in a raster scan order.
[0092] A slice consists of an integer number of complete tiles within a tile of a picture or an integer number of consecutive complete CTU rows.
[0093] Two slice modes are supported, namely, raster-scan slice mode and rectangular slice mode. In the raster-scan slice mode, a slice contains a complete sequence of tiles in the raster scan of a picture. In the rectangular slice mode, a slice contains multiple complete tiles that collectively form a rectangular region of a picture or multiple consecutive complete CTU rows of one tile that collectively form a rectangular region of a picture. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular region corresponding to the slice.
[0094] A subpicture contains one or more slices that collectively cover a rectangular region of a picture.
[0095] Figure 1 An example of a raster-scan slice partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster-scan slices.
[0096] Figure 2 An example of a rectangular slice partitioning of a picture is shown, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular tiles.
[0097] Figure 3 An example of a picture partitioned into tiles and rectangular slices is shown, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0098] Figure 4An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 tiles, 12 tiles on the left (each covering one slice of 4x4 CTUs) and 6 tiles on the right (each covering 2 slices of 2x2 CTUs vertically stacked), resulting in 24 slices and 24 sub-pictures of different sizes (each slice is a sub-picture).
[0099] 3.3. Picture resolution change within a sequence
[0100] In AVC and HEVC, the spatial resolution of a picture cannot change unless a new sequence with a new SPS is used starting with an IRAP picture. VVC allows changing the picture resolution within a sequence at a location without encoding an IRAP picture, which is always intra coded. This feature is sometimes referred to as reference picture resampling (RPR) because it requires resampling of the reference picture for inter prediction when the reference picture has a different resolution than the current picture being decoded.
[0101] The scaling ratio is limited to be greater than or equal to 1 / 2 (2x downsampling from the reference picture to the current picture) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied to the scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as the case of the motion compensation interpolation filter. In fact, the normal MC interpolation process is a special case of the resampling process, where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.
[0102] Other aspects of the VVC design that are different from HEVC to support this feature include: i) the picture resolution and the corresponding conformance window are signaled in the PPS instead of the SPS, while the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture store (a slot in the DPB that stores a decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.
[0103] 3.4. Scalable video coding (SVC) in general and in VVC
[0104] Scalable Video Coding (SVC, sometimes also referred to as scalability in video coding) refers to video coding using a base layer (BL) (sometimes referred to as a reference layer (RL)) and one or more scalable enhancement layers (ELs). In SVC, the base layer can carry video data at a base level of quality. The one or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise (SNR) levels. An enhancement layer can be defined relative to a previously coded layer. For example, a bottom layer can serve as a BL, while a top layer can serve as an EL. An intermediate layer can serve as an EL or an RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest layer nor the highest layer) can be an EL of the layers below the intermediate layer (e.g., the base layer or any intervening enhancement layers), and at the same time serve as an RL for one or more enhancement layers above the intermediate layer. Similarly, in the Multiview or 3D extensions of the HEVC standard, there can be multiple views, and information of one view can be used to code (e.g., encode or decode) information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).
[0105] In SVC, the parameters used by an encoder or decoder are grouped into parameter sets based on the coding level at which they can be used (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be used by one or more coded video sequences of different layers in a bitstream can be included in a video parameter set (VPS), and parameters that can be used by one or more pictures in a coded video sequence can be included in a sequence parameter set (SPS). Similarly, parameters that are used by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters specific to a single slice can be included in a slice header. Similarly, an indication of which parameter set(s) a particular layer uses at a given time can be provided at various coding levels.
[0106] Thanks to the support of reference picture resampling (RPR) in VVC, it is possible to design the support of bitstreams containing multiple layers without any additional signaling of the level of the coding tools, e.g., two layers with SD and HD resolutions in VVC, since the upsampling needed for the spatial scalability support can only use the RPR upsampling filter. However, to support scalability, high level syntax changes (compared to the non-scalable case) are needed. The scalability support is specified in VVC version 1. The design of VVC scalability is friendly to the single-layer decoder design as much as possible, unlike the scalability support in any earlier video coding standards, including the extensions of AVC and HEVC. The decoding capability of a multi-layer bitstream is specified in a way as if there is only a single layer in the bitstream. For example, the decoding capability such as the DPB size is specified in a way independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream can decode a multi-layer bitstream without much change. Compared to the design of the multi-layer extensions of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, an IRAP AU needs to contain pictures of every layer present in the CVS.
[0107] 3.5. Subpicture-based viewport-dependent 360° video streaming
[0108] In the streaming of 360° video (also known as omnidirectional video), at any particular time, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user, while the user can turn his / her head at any time to change the viewing direction, thus changing the current viewport. While it is desirable to have at least some lower quality representation of the areas not covered by the current viewport at the client ready to be presented to the user in case the user suddenly changes his / her viewing direction to any location on the sphere, only the high quality representation of the omnidirectional video is needed for the current viewport that is currently being presented to the user. Partitioning the high quality representation of the entire omnidirectional video into subpictures at an appropriate granularity can enable this optimization. Using VVC, the two representations can be coded as two layers independent of each other.
[0109] A typical subpicture-based viewport-dependent 360° video streaming scheme is shown in Figure 11 , where the higher resolution representation of the complete video is composed of subpictures, while the lower resolution representation of the complete video does not use subpictures and can be coded using random access points at a lower frequency than the higher resolution representation. The client receives the lower resolution complete video, while for the higher resolution video, it only receives and decodes the subpicture covering the current viewport.
[0110] The latest VVC draft specification also supports an improved 360° video coding scheme as shown in Figure 12 . Compared to Figure 11The only difference compared to the illustrated method is that inter-layer prediction (ILP) is applied Figure 12 The illustrated method.
[0111] 3.6. Parameter sets
[0112] AVC, HEVC and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS and VPS. SPS and PPS are supported by all AVC, HEVC and VVC. VPS is introduced from HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0113] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that does not change frequently. With SPS and PPS, there is no need to repeat the information that does not change frequently for each sequence or picture, so redundant signaling of this information can be avoided. In addition, the use of SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission, but also improving error recovery capability.
[0114] VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0115] APS is introduced to carry picture-level or slice-level information that requires a considerable number of bits for coding, can be shared by multiple pictures, and can have many different changes in a sequence.
[0116] 3.7. Specification and signaling of nested SEI messages for subpicture sequences in VVC
[0117] In the latest VVC draft text, the specification and signaling of nested SEI messages for subpicture sequences in VVC is done through scalable nesting SEI messages. Subpicture sequences are defined in the semantics of the subpicture level information (SLI) SEI message. Subpicture sequences can be extracted from the bitstream by applying the subpicture sub-bitstream extraction process specified in clause C.7 of VVC.
[0118] The syntax and semantics of scalable nesting SEI messages in the latest VVC draft text are as follows.
[0119] D.6.1 Scalable nesting SEI message syntax
[0120]
[0121]
[0122] D.6.2 Scalable nesting SEI message semantics
[0123] Scalable nesting SEI messages provide a mechanism to associate SEI messages with a particular OLS or a particular layer, and to associate SEI messages with a particular subpicture group.
[0124] A scalable nesting SEI message contains one or more SEI messages. SEI messages contained in a scalable nesting SEI message are also referred to as being scalable-nested.
[0125] The following restrictions apply to SEI messages contained in a scalable nesting SEI message as a requirement of bitstream conformance:
[0126] - SEI messages with payloadType equal to 132 (Decoded Picture Hash) can only be contained in a scalable nesting SEI message where sn_subpic_flag is equal to 1.
[0127] - SEI messages with payloadType equal to 133 (Scalable nesting) shall not be contained in a scalable nesting SEI message.
[0128] - When a scalable nesting SEI message contains a BP, PT, or DUI SEI message, the scalable nesting SEI message shall not contain any other SEI messages where payloadType is not equal to 0 (BP), 1 (PT), or 130 (DUI).
[0129] The following restrictions apply to the value of nal_unit_type of the SEI NAL unit containing a scalable nesting SEI message as a requirement of bitstream conformance:
[0130] - When a scalable nesting SEI message contains SEI messages where payloadType is equal to 0 (BP), 1 (PT), 130 (DUI), 145 (DRAP indication), or 168 (Frame field information), the SEI NAL unit containing the scalable nesting SEI message shall have nal_unit_type equal to PREFIX_SEI_NUT.
[0131] - When a scalable nesting SEI message contains SEI messages where payloadType is equal to 0 (BP), 1 (PT), 130 (DUI), 145 (DRAP indication), or 168 (Frame field information), the SEI NAL unit containing the scalable nesting SEI message shall have nal_unit_type equal to PREFIX_SEI_NUT.
[0132] equal to 1 specifies that the scalable-nested SEI message applies to a particular OLS. sn_ols_flag equal to 0 specifies that the scalable-nested SEI message applies to a particular layer.
[0133] The following restrictions apply to the value of sn_ols_flag as a requirement of bitstream conformance:
[0134] – When the scalable-nested SEI message contains an SEI message with payloadType equal to 0 (BP), 1 (PT), or 130 (DUI), the value of sn_ols_flag shall be equal to 1.
[0135] – When the scalable-nested SEI message contains an SEI message with payloadType equal to a value in VclAssociatedSeiList, the value of sn_ols_flag shall be equal to 0.
[0136] equal to 1 specifies that the scalable-nested SEI message that applies to a specified OLS or layer applies only to a particular subpicture of the specified OLS or layer. sn_subpic_flag equal to 0 specifies that the scalable-nested SEI message that applies to a particular OLS or layer applies to all subpictures of the specified OLS or layer.
[0137] plus 1 specifies the number of OLSs to which the scalable-nested SEI message applies. The value of sn_num_olss_minusl shall be in the range of 0 to TotalNumOlss - 1, inclusive.
[0138] [i] The variable NestingOlsIdx[ i ] that specifies the OLS index of the i-th OLS to which the scalable-nested SEI message applies when sn_ols_flag is equal to 1 is derived as follows:
[0139] The value of sn_ols_idx_delta_minusl[ i ] shall be in the range of 0 to TotalNumOlss - 2, inclusive.
[0140] The variable NestingOlsIdx[ i ] is derived as follows:
[0141]
[0142] equal to 1 specifies that the scalable-nested SEI message applies to all layers with nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. sn_all_layers_flag equal to 0 specifies that the scalable-nested SEI message can or can not apply to all layers with nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit.
[0143] plusl specifies the number of layers to which the scalable-nested SEI message applies
[0144] The value of sn num layers minusl shall be in the range of 0 to vps max layers minusl - GeneralLayerld[ nuh layer id ], inclusive, where nuh layer id is the nuh layer id of the current SEI NAL unit.
[0145] When sn all layers flag is equal to 0, specifies the nuh layer id values that apply to the i-th layer of the scalable-nested SEI message. The value of sn layer id[ i ] shall be greater than nuh layer id, where nuh layer id is the nuh layer id of the current SEI NAL unit.
[0146] When sn ols flag is equal to 0, specifies the variable nestingNumLayers that specifies the number of layers to which the scalable-nested SEI message applies, and a list nestingLayerld[ i ] that specifies a list of nuh layer id values that apply to the layers of the scalable-nested SEI message for i in the range of 0 to nestingNumLayers - 1, inclusive, is derived as follows, where nuh layer id is the nuh layer id of the current SEI NAL unit:
[0147]
[0148] plusl specifies the number of sub-pictures to which the scalable-nested SEI message applies. The value of sn num subpics minusl shall be less than or equal to the value of sps num subpics minusl in the SPS in which the picture referred to in the CLVS.
[0149] plusl specifies the number of sub-pictures to which the scalable-nested SEI message applies. The value of sn num subpics minusl shall be less than or equal to the value of sps num subpics minusl in the SPS in which the picture referred to in the CLVS.
[0150] The requirement of bitstream conformance is that the value of sn subpic id len minusl shall be the same for all scalable-nested SEI messages present in the CLVS.
[0151] [i] indicates the i-th subpicture ID associated with the scalable-nested SEI message. The length of the sn_subpic_id[ i ] syntax element is sn_subpic_id_len_minus1 + 1 bits.
[0152] Add 1 specifies the number of scalable-nested SEI messages. The value of sn_num_seis_minus1 shall be in the range of 0 to 63, inclusive.
[0153] shall be equal to 0.
[0154] 4. Technical problems solved by the disclosed technical solutions
[0155] The existing VVC design of specifying and signaling nested SEI messages for subpictures and subpicture sequences through scalable-nested SEI messages has the following problems:
[0156] 1) In order to associate the scalable-nested SEI messages to one or more subpictures, the scalable-nested SEI messages use the subpicture ID. However, the duration range of the scalable-nested SEI messages can be multiple consecutive AUs, while the subpicture ID of the subpicture with a specific subpicture index in a layer can change within a CLVS. Therefore, the subpicture index should be used in the scalable-nested SEI messages instead of using the subpicture ID.
[0157] 2) When the related subpicture is removed, the filler payload SEI message (if present) needs to be removed from the output bitstream in the subpicture sub-bitstream extraction process. However, when the filler payload SEI message can be included in the scalable-nested SEI message, removing the filler payload SEI message in the subpicture sub-bitstream extraction process sometimes needs to extract some scalable-nested SEI messages from the scalable-nested SEI message.
[0158] 3) Since the SLI SEI message is applicable to OLS, as the other three HRD-related SEI messages (i.e., BP / PT / DUI SEI messages), the value of sn_ols_flag needs to be equal to 1 when the SLI SEI message is scalable-nested. In addition, since the SLI SEI message specifies the information of all subpictures of the pictures in the OLS to which the SLI SEI message is applicable, it is meaningless for the value of sn_subpic_flag to be equal to 1 for the scalable-nested SEI message containing the SLI SEI message.
[0159] 4) The lack of a constraint that when a scalable-nested SEI message contains a BP, PT, DUI, or SLI SEI message, the scalable-nested SEI message shall not contain any other SEI message where payloadType is not equal to 0 (BP), 1 (PT), 130 (DUI), or 203 (SLI).
[0160] 5) The provision that when a scalable-nested SEI message contains an SEI message where payloadType is equal to 0 (BP), 1 (PT), 130 (DUI), 145 (DRAP indication), or 168 (frame-field information), the SEI NAL unit containing the scalable-nested SEI message shall have nal_unit_type equal to PREFIX_SEI_NUT.
[0161] However, when nesting many other SEI messages, the value of the scalable-nested SEI message shall also have nal_unit_type equal to PREFIX_SEI_NUT.
[0162] 6) The lack of a constraint that when a scalable-nested SEI message contains an SEI message where payloadType is equal to 132 (Decoded Picture Hash), the SEI NAL unit containing the scalable-nested SEI message shall have nal_unit_type equal to SUFFIX_SEI_NUT.
[0163] 7) The semantics of sn_num_subpics_minus1 and sn_subpic_idx[i] need to be specified in such a way that the syntax elements are about the subpictures of the layers that have multiple subpictures per picture, in order to be able to support the case where OLS has some layers that have multiple subpictures per picture and some other layers that have a single subpicture per picture.
[0164] 5. List of solutions and embodiments
[0165] To solve the above and other problems, methods as summarized below are disclosed. The solution items should be considered as examples to explain the general concepts and should not be interpreted in a narrow way. Furthermore, these items can be used individually or combined in any way.
[0166] 1) To solve the first problem, use the subpicture index (instead of using the subpicture ID) to associate the subpicture to the SEI messages that are scalable-nested in the scalable-nested SEI message.
[0167] a. In one example, the syntax element sn_subpic_id[i] is changed to sn_subpic_idx[i] and the sn_subpic_id_len_minus1 syntax element is removed as a consequence.
[0168] 2) To solve the second problem, the filler payload SEI message is prohibited to be scalable-nested, i.e. contained in a scalable-nested SEI message.
[0169] 3) To solve the third problem, a constraint is added such that when a scalable-nested SEI message contains one or more SLI SEI messages, the value of sn_ols_flag shall be equal to 1.
[0170] a. In one example, in addition, or alternatively, a constraint is added such that when a scalable-nested SEI message contains one or more SLI SEI messages, the value of sn_subpic_flag shall be equal to 0.
[0171] 4) To solve the 4th problem, it is required that when a scalable-nested SEI message contains a BP, PT, DUI or SLI SEI message, the scalable-nested SEI message shall not contain any other SEI message with payloadType not equal to 0 (BP), 1 (PT), 130 (DUI) or 203 (SLI).
[0172] 5) To solve the 5th problem, it is specified that when a scalable-nested SEI message contains a SEI message with payloadType not equal to 3 (filler payload) or 132 (decoded picture hash), the SEI NAL unit containing the scalable-nested SEI message shall have nal_unit_type equal to PREFIX_SEI_NUT.
[0173] 6) To solve the 6th problem, a constraint is added such that when a scalable-nested SEI message contains a SEI message with payloadType equal to 132 (decoded picture hash), the SEI NAL unit containing the scalable-nested SEI message shall have nal_unit_type equal to SUFFIX_SEI_NUT.
[0174] 7) To solve the 7th problem, the semantics of sn_num_subpics_minus1 and sn_subpic_idx[i] are specified in such a way that the syntax elements specify information about the sub-pictures of the layer having multiple sub-pictures per picture.
[0175] 6. Embodiments
[0176]
[0177] 6.1. Example 1
[0178] This example is directed to items 1 to 5 and some of their sub-items.
[0179] D.6.1 Scalable nesting SEI message syntax
[0180]
[0181]
[0182] D.6.2 Scalable nesting SEI message semantics
[0183] Scalable nesting SEI messages provide a mechanism to associate SEI messages with a particular OLS or a particular layer, and to associate SEI messages with a particular sub-picture group.
[0184] A scalable nesting SEI message contains one or more SEI messages. SEI messages contained in a scalable nesting SEI message are also referred to as scalable-nested SEI messages.
[0185] The requirement of bitstream conformance is that the following restrictions apply to SEI messages contained in a scalable nesting SEI message:
[0186] [[– SEI messages with payloadType equal to 132 (Decoded Picture Hash) can only be contained in a scalable nesting SEI message in which sn_subpic_flag is equal to 1.]]
[0187] – SEI messages with payloadType equal to 133 (scalable nesting) shall not be contained in a scalable nesting SEI message.
[0188] – When a scalable nesting SEI message contains BP, PT, [[or]] DUI SEI messages, the scalable nesting SEI message shall not contain any other SEI messages in which payloadType is not equal to 0 (BP), 1 (PT), [[or]] 130 (DUI)
[0189] .
[0190] The requirement of bitstream conformance is that the following restrictions apply to the value of nal_unit_type of the SEI NAL unit containing a scalable nesting SEI message:
[0191] – When a scalable nesting SEI message contains SEI messages in which payloadType
[0192] When an SEI message is [[equal to 0 (BP), 1 (PT), 130 (DUI), 145 (DRAP indication) or 168 (frame field information)]], the SEI NAL unit containing scalable nested SEI messages should have a nal_unit_type equal to PREFIX_SEI_NUT.
[0193]
[0194] A value of 1 indicates that the scalable nested SEI message applies to a specific OLS. A value of 0 indicates that the scalable nested SEI message applies to a specific layer.
[0195] The requirements for bitstream consistency are as follows, and the following constraints apply to the value of sn_ols_flag:
[0196] –When a scalable nested SEI message contains a payloadType equal to 0 (BP), 1 (PT),
[0197] [[or]]130(DUI) When receiving an SEI message, the value of sn_ols_flag should be equal to 1.
[0198] – When a scalable nested SEI message contains a payloadType equal to a value in VclAssociatedSeiList When receiving an SEI message, the value of sn_ols_flag should be equal to 0.
[0199] A value of 1 indicates that the scalable nested SEI message applicable to a specified OLS or layer applies only to a specific subpicture of the specified OLS or layer. A value of 0 indicates that the scalable nested SEI message applicable to a specific OLS or layer applies to all subpictures of the specified OLS or layer.
[0200]
[0201] Add 1 to specify the number of OLS applicable to the scalable nested SEI messages. The value of sn_num_olss_minus1 should be in the range of 0 to TotalNumOlss-1 (inclusive).
[0202] [i] The variable NestingOlsIdx[ i ] is derived as follows for determining the OLS index that applies to the i-th OLS of the scalably-nested SEI message when sn_ols_flag is equal to 1.
[0203] The value of sn_ols_idx_delta_minus1[ i ] shall be in the range of 0 to TotalNumOlss - 2, inclusive.
[0204] The variable NestingOlsIdx[ i ] is derived as follows:
[0205]
[0206] Equal to 1 specifies that the scalably-nested SEI message applies to all layers with nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit. sn_all_layers_flag equal to 0 specifies that the scalably-nested SEI message can or can not apply to all layers with nuh_layer_id greater than or equal to the nuh_layer_id of the current SEI NAL unit.
[0207] Plus 1 specifies the number of layers to which the scalably-nested SEI message applies
[0208] The value of sn_num_layers_minus1 shall be in the range of 0 to vps_max_layers_minus1 - GeneralLayerIdx[ nuh_layer_id ], inclusive, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.
[0209] Specifies the nuh_layer_id value of the i-th layer to which the scalably-nested SEI message applies when sn_all_layers_flag is equal to 0. The value of sn_layer_id[ i ] shall be greater than nuh_layer_id, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit.
[0210] When sn_ols_flag equals 0, the variable nestingNumLayers specifies the number of layers applicable to the scalable nested SEI messages, and nestingLayerId[i] is a list of nuh_layer_id values applicable to the layers of the scalable nested SEI messages for i in the range of 0 to nestingNumLayers-1 (inclusive), as deduced below, where nuh_layer_id is the nuh_layer_id of the current SEI NAL unit:
[0211]
[0212]
[0213]
[0214] Adding 1 specifies that [[applies to scalable nested SEI messages]] The number of subpics. The value of sn_num_subpics_minus1 should be less than or equal to the value of sn_num_subpics_minus1 in the SPS referenced in multiSubpicLayers[[CLVS]].
[0215] Adding 1 specifies the number of bits used to represent the syntax element sn_subpic_id[i]. The value of sn_subpic_id_len_minus1 should be in the range of 0 to 15 (inclusive).
[0216] The requirement for bitstream consistency is that the value of sn_subpic_id_len_minus1 should be the same for all scalable nested SEI messages present in CLVS.
[0217] [i][[instruction]] [[Associated with the scalable nested SEI message]]'s i-th sub-image[[ID]]
[0218] The length of the syntax element [[sn_subpic_id[i] is sn_subpic_id_len_minus1+1 bits]].
[0219] The value of sn num seis minusl shall be in the range of zero to sixty-three, inclusive.
[0220] SHOULDEQUALTO0.
[0221] Figure 5 is a block diagram of an example video processing system 1900 that can implement the various techniques disclosed herein. Various implementations can include some or all of the components of system 1900. System 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be received in a compressed or encoded format. Input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0222] System 1900 can include a codec component 1904 that can implement various coding or encoding methods described in this document. Codec component 1904 can reduce the average bitrate of video from input 1902 to the output of codec component 1904 to produce a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via a connected communication, as represented by component 1906. The bitstream (or coded) representation of the video that is stored or communicated at input 1902 can be used by component 1908 to generate pixel values or displayable video that is sent to a display interface 1910. The process of generating user- visible video from a bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0223] Examples of peripheral bus interfaces or display interfaces can include universal serial bus (USB) or high definition multimedia interface (HDMI) or Displayport, etc. Examples of storage interfaces include SATA (serial advanced technology attachment), PCI, IDE interfaces, etc. The techniques described in this document can be implemented in various electronic devices such as mobile phones, laptops, smart phones, or other digital data processing and / or video display capable devices.
[0224] Figure 6 is a block diagram of a video processing device 3600. The device 3600 can be used to implement one or more of the methods described herein. The device 3600 can be implemented in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in the present document. The memory(ies) 3604 can be used for storing data and code used during operation of the present device 3600. The video processing hardware 3606 can be used to implement some techniques described in the present document in hardware circuitry.
[0225] Figure 8 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0226] As shown in FIG. 1, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 generates encoded video data, which can be referred to as a video encoding device. Destination device 120 can decode the encoded video data generated by source device 110, and the destination device 120 can be referred to as a video decoding device. Figure 8
[0227] Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0228] Video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system generating video data, or a combination of such sources. Video data can comprise one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. Associated data can include sequence parameter sets, picture parameter sets, and other syntax elements. I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 by source device 110 via I / O interface 116 and network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by destination device 120.
[0229] Destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0230] The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the destination device 120, or can be external to the destination device 120 configured to interface with the external display device.
[0231] The video encoder 114 and the video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other current and / or further standards.
[0232] Figure 9 is a block diagram illustrating an example of a video encoder 200 that can be Figure 8 the video encoder 114 in the system 100 illustrated in FIG. 1.
[0233] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 9 examples, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0234] The functional components of the video encoder 200 can include a partition unit 201, a prediction unit 202 (which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0235] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can make predictions in IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0236] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately in Figure 5 examples for explanatory purposes.
[0237] The partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[0238] The mode selection unit 203 can select one of the intra or inter coding modes, e.g., based on the error results, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra and inter prediction (CIIP) mode, in which the prediction is based on both an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select a resolution of a motion vector for a block, e.g., sub-pixel or integer-pixel precision, for the case of inter prediction.
[0239] To perform inter prediction for a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 to the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.
[0240] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0241] In some examples, the motion estimation unit 204 can perform single prediction for a current video block, and the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of List 0 or List 1. The motion estimation unit 204 can then generate a reference index indicating the reference video block is in the reference picture of List 0 or List 1 and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0242] In other examples, motion estimation unit 204 can perform bi-prediction of the current video block, motion estimation unit 204 can search for a reference video block for the current video block in the reference pictures of list 0 and also can search for another reference video block for the current video block in the reference pictures of list 1. Motion estimation unit 204 can then generate a reference index indicating the reference video block in the reference pictures of list 0 or list 1 and a motion vector indicating a spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and the motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0243] In some examples, motion estimation unit 204 can output the entire set of motion information for the decoding process of the decoder.
[0244] In some examples, motion estimation unit 204 can not output the entire set of motion information for the current video. Instead, motion estimation unit 204 can signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0245] In one example, motion estimation unit 204 can indicate in a syntax structure associated with the current video block a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0246] In another example, motion estimation unit 204 can identify in a syntax structure associated with the current video block another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between the motion vector of the current video block and a motion vector that indicates the video block. Video decoder 300 can use the motion vector that indicates the video block and the motion vector difference to determine the motion vector of the current video block.
[0247] As discussed above, video encoder 200 can predictively signal motion vectors. Two examples of predictively signaling techniques that can be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0248] Intra prediction unit 206 can perform intra prediction of the current video block. When intra prediction unit 206 performs intra prediction of the current video block, intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0249] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., represented by a minus sign) the prediction video block(s) from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0250] In other examples, such as in a skip mode, there can be no residual data for the current video block for the current video block, and residual generation unit 207 can not perform a subtraction operation.
[0251] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0252] After transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0253] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transforms, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0254] After reconstruction unit 212 reconstructs the video block, loop filtering operations can be performed to reduce blocking artifacts in the video block.
[0255] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0256] Figure 10 is a block diagram illustrating an example of a video decoder 300 that can be Figure 8 the video decoder 114 in the system 100 illustrated in FIG.
[0257] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 10 In examples where video decoder 300 includes multiple functional components, the techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0258] In Figure 10 In the example of FIG. 3, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, video decoder 300 can perform a decoding process generally inversely to the encoding process described with respect to video encoder 200. Figure 9 ) described with respect to video encoder 200.
[0259] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video and, from the entropy decoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and merge mode.
[0260] Motion compensation unit 302 can generate a motion compensated block, possibly interpolated based on an interpolation filter. An identifier of an interpolation filter to be used at sub-pixel precision can be included in a syntax element.
[0261] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolated values of sub-integer pixel positions of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from received syntax information and use the interpolation filter to generate the prediction block.
[0262] Motion compensation unit 302 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information to decode the encoded video sequence.
[0263] Intra-prediction unit 303 can form a prediction block from spatial neighboring blocks using, for example, an intra-prediction mode received in the bitstream. Inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transformation unit 303 applies an inverse transform.
[0264] The reconstruction unit 306 can sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. As desired, a deblocking filter can also be applied to the filtered decoded block in order to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and also produces decoded video for presentation on a display device.
[0265] Next, a list of some embodiments preferred solutions is provided.
[0266] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., all items).
[0267] 1. A method of video processing, comprising performing a conversion between a video comprising one or more subpictures or one or more sequences of subpictures and a coded representation of the video, wherein the coded representation conforms to a format rule that specifies whether or how supplemental enhancement information (SEI) that is scalably-nested is included in the coded representation.
[0268] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0269] 2. The method of solution 1, wherein the format rule specifies that the coded representation uses a subpicture index to associate a subpicture to corresponding scalably-nested SEI information.
[0270] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 2).
[0271] 3. The method of any of solutions 1-2, wherein the format rule does not allow a filter payload SEI message to be used in a scalably-nested manner.
[0272] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0273] 4. The method of any of solutions 1-3, wherein the format rule specifies that, for a scalably-nested SEI message that includes one or more subpicture level information SEI messages, a flag is included in the coded representation to indicate a presence thereof.
[0274] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 4).
[0275] 5. The method of any of solutions 1-4, wherein the format rule prohibits inclusion of a nested SEI message of a particular type of payload in a scalable SEI message containing a particular type of message.
[0276] The following solution shows an example embodiment of the techniques discussed in the previous section (e.g., item 5).
[0277] 6. The method of any of solutions 1-5, wherein the format rule specifies that SEI messages that are not filler payload type or decoded picture hash type need to have a particular network abstraction layer unit type.
[0278] The following solution shows an example embodiment of the techniques discussed in the previous section (e.g., item 6).
[0279] 7. The method of any of solutions 1-6, wherein the format rule specifies that SEI messages of decoded picture hash type need to have a particular network abstraction layer unit type.
[0280] The following solution shows an example embodiment of the techniques discussed in the previous section (e.g., item 7).
[0281] 8. The method of any of solutions 1-7, wherein the format rule specifies that a syntax element specifies information about sub-pictures of a layer that has multiple sub-pictures per picture.
[0282] 9. The method of any of solutions 1 to 8, wherein the converting comprises encoding the video into a coded representation.
[0283] 10. The method of any of solutions 1 to 8, wherein the converting comprises decoding a coded representation to generate pixel values of the video.
[0284] 11. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 10.
[0285] 12. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 10.
[0286] 13. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement a method recited in any of solutions 1 to 10.
[0287] 14. The method, apparatus or system described in the present document.
[0288] In the solutions described herein, an encoder can conform to a format rule by generating a coded representation according to the format rule. In the solutions described herein, a decoder can use a format rule to parse syntax elements in a coded representation, understand the presence and absence of syntax elements according to the format rule to generate a decoded video.
[0289] Figure 13 is a flowchart of an example method 1300 of processing video data. Operation 1302 includes performing a conversion between a video comprising one or more subpictures and a bitstream of the video, wherein during the conversion, one or more supplemental enhancement information messages having a filler payload are processed according to a format rule, and wherein the format rule does not allow the one or more supplemental enhancement information messages having the filler payload in a scalably-nested supplemental enhancement information message.
[0290] In some embodiments of the method 1300, the one or more supplemental enhancement information messages having the filler payload include a payload type equal to 3. In some embodiments of the method 1300, the format rule does not allow one or more second supplemental enhancement information messages having a scalable nesting in the scalably-nested supplemental enhancement information message.
[0291] Figure 14 is a flowchart of an example method 1400 of processing video data. Operation 1402 includes performing a conversion between a video and a bitstream of the video, wherein during the conversion, one or more syntax elements are processed according to a format rule, and wherein the format rule specifies the one or more syntax elements for indicating subpicture information of a layer of the video having a picture with a plurality of subpictures.
[0292] In some embodiments of the method 1400, the one or more syntax elements include a first syntax element, and wherein a value of the first syntax element plus one specifies a number of subpictures in the picture with the plurality of subpictures. In some embodiments of the method 1400, a value of the first syntax element is less than or equal to a value of a second syntax element in a sequence parameter set referred to by the picture in the plurality of subpicture layers. In some embodiments of the method 1400, the one or more syntax elements include a third syntax element, and wherein the third syntax element indicates a subpicture index of an i-th subpicture in each of the picture with the plurality of subpictures. In some embodiments of the method 1400, the first syntax element is labeled as sn num subpics minusl. In some embodiments of the method 1400, the third syntax element is labeled as sn subpic idx [i].
[0293] Figure 15is a flowchart of an example method 1500 of processing video data. Operation 1502 includes performing a conversion between a video comprising a plurality of subpictures and a bitstream of the video, wherein, during the conversion, a scalably-nested supplemental enhancement information message is processed according to a format rule, and wherein the format rule specifies that one or more subpicture indexes are used to associate one or more subpictures to the scalably-nested supplemental enhancement information message.
[0294] In some embodiments of the method 1500, the format rule does not allow one or more subpicture identifiers to be used to associate one or more subpictures to the scalably-nested supplemental enhancement information message. In some embodiments of the method 1500, the format rule replaces a first syntax element in the scalably-nested supplemental enhancement information message with a second syntax element, the format rule removes a third syntax element from the scalably-nested supplemental information message, the first syntax element indicates a subpicture identifier of an i-th subpicture in each picture in one or more video layers, the second syntax element indicates a subpicture index of the i-th subpicture in each picture in the one or more video layers, and the third syntax element specifies a number of bits used to represent the first syntax element plus 1.
[0295] Figure 16 is a flowchart of an example method 1600 of processing video data. Operation 1602 includes performing a conversion between a video comprising one or more subpictures and a bitstream of the video according to a format rule, wherein the format rule specifies that, in response to a scalably-nested supplemental enhancement information message including one or more subpicture level information supplemental enhancement information messages, a first syntax element in the scalably-nested supplemental enhancement information message in the bitstream is set to a particular value, and wherein the particular value of the first syntax element indicates that the scalably-nested supplemental enhancement information message includes one or more scalably-nested supplemental enhancement information messages that are applicable to a particular output video layer set.
[0296] In some embodiments of the method 1600, the format rule specifies that, responsive to the scalable-nested supplemental enhancement information message including one or more subpicture level information supplemental enhancement information messages, a value of a second syntax element in the bitstream is equal to 0, and the value of the second syntax element equal to 0 specifies that the scalable-nested supplemental enhancement information message includes one or more scalable-nested supplemental enhancement information messages that are applicable to one or more output video layer sets or one or more video layers, the one or more output video layer sets or the one or more video layers being applicable to all subpictures of the one or more output video layer sets or the one or more video layers. In some embodiments of the method 1600, the format rule specifies that a payload type of the one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message is 203, the payload type 203 of the supplemental enhancement information message indicating that the supplemental enhancement information message is a subpicture level information supplemental enhancement information message. In some embodiments of the method 1600, the particular value of the first syntax element is equal to 1.
[0297] Figure 17 is a flowchart of an example method 1700 of processing video data. Operation 1702 includes performing a conversion between a video comprising a plurality of subpictures and a bitstream of the video, wherein the conversion is according to a format rule that specifies that a scalable-nested supplemental enhancement information message is not allowed to include a first supplemental enhancement information message of a first payload type and a second supplemental enhancement information message of a second payload type.
[0298] In some embodiments of the method 1700, the first payload type includes a payload type of a buffering period supplemental enhancement information message. In some embodiments of the method 1700, the first payload type includes a payload type of a picture timing supplemental enhancement information message. In some embodiments of the method 1700, the first payload type includes a payload type of a decoding unit information supplemental enhancement information message. In some embodiments of the method 1700, the first payload type includes a payload type of a subpicture level information supplemental enhancement information message. In some embodiments of the method 1700, the second payload type includes a payload type that is not one of: (i) the payload type of the buffering period supplemental enhancement information message, (ii) the payload type of the picture timing supplemental enhancement information message, (iii) the payload type of the decoding unit information supplemental enhancement information message, and (iv) the payload type of the subpicture level information supplemental enhancement information message.
[0299] Figure 18is a flowchart of an example method 1800 of processing video data. Operation 1802 includes performing a conversion between video content and a bitstream of the video content, wherein the conversion is performed according to a format rule that specifies, responsive to a supplemental enhancement information network abstraction layer unit including a scalable-nested supplemental enhancement information message, the supplemental enhancement information network abstraction layer unit including a network abstraction layer unit type that is equal to a prefix supplemental enhancement information network abstraction layer unit type, wherein the scalable-nested supplemental enhancement information message includes a supplemental enhancement information message that is not associated with a particular payload type.
[0300] In some embodiments of the method 1800, the network abstraction layer unit type is equal to PREFIX SEI NUT.
[0301] Figure 19 is a flowchart of an example method 1900 of processing video data. Operation 1902 includes performing a conversion between video content and a bitstream of the video content, wherein the conversion is performed according to a format rule that specifies, responsive to a supplemental enhancement information network abstraction layer unit including a scalable-nested supplemental enhancement information message, the supplemental enhancement information network abstraction layer unit including a network abstraction layer unit type that is equal to a suffix supplemental enhancement information network abstraction layer unit type, wherein the scalable-nested supplemental enhancement information message includes a supplemental enhancement information message that is associated with a particular payload type.
[0302] In some embodiments of the method 1900, the network abstraction layer unit type is equal to SUFFIX SEI NUT.
[0303] In some embodiments of the method(s) 1800-1900, the particular payload type is a payload type of a decoded picture hash supplemental enhancement information message. In some embodiments of the method(s) 1800-1900, the particular payload type is associated with a value equal to 132.
[0304] In some embodiments of the method(s) 1300-1900, performing the conversion includes encoding the video into a bitstream. In some embodiments of the method(s) 1300-1900, performing the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of the method(s) 1300-1900, performing the conversion includes decoding the video from the bitstream. In some embodiments, a video decoding apparatus includes a processor configured to implement the method(s) 1300-1900 or embodiments thereof. In some embodiments, a video encoding apparatus includes a processor configured to implement the method(s) 1300-1900 or embodiments thereof. In some embodiments, a computer program product having computer instructions stored therein, which, when executed by a processor, cause the processor to implement the method(s) 1300-1900 or embodiments thereof. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to the method(s) 1300-1900 or embodiments thereof. In some embodiments, a non-transitory computer-readable storage medium stores instructions that cause a processor to implement the method(s) 1300-1900 or embodiments thereof. In some embodiments, a method of bitstream generation includes generating a bitstream of a video according to the method(s) 1300-1900 or embodiments thereof, and storing the bitstream on a computer-readable program medium. In some embodiments, a method, apparatus, bitstream generated according to the disclosed methods or systems described in this document.
[0305] Some embodiments of the disclosed technology include a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder will use or implement the tool or mode in the processing of blocks of the video, but does not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on the decision or determination, the conversion from blocks of the video to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, a decoder will process a bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from a bitstream representation of the video to blocks of the video will be performed using the video processing tool or mode enabled based on the decision or determination.
[0306] Some embodiments of the disclosed technology include a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode in the conversion of blocks of the video to a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder will process a bitstream using the knowledge that the bitstream has not been modified based on the decision or determination to disable the video processing tool or mode.
[0307] In this document, the term“video processing” can refer to video encoding, video decoding, video compression, or video decompression. For example, during a conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm can be applied, and vice versa. As defined by the syntax, a bitstream representation of a current video block can correspond, for example, to bits that are co-located or scattered at different locations within the bitstream. For example, a macroblock can be encoded according to transformed and coded error residual values and also using bits in the header and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination, knowing that some fields can or can not be present, as described in the above solutions. Similarly, an encoder can determine to include or not include certain syntax fields and generate the coded representation accordingly by including or excluding the syntax fields from the coded representation.
[0308] The disclosed and other solutions, examples, embodiments, modules and functional operations set forth in this document can be realized in digital electronic circuitry, or in a computer software, firmware, or hardware, including the structural equivalents of such disclosure as set forth in this document and as such are to be interpreted within the ordinary meaning of the terms employed. The disclosed and other embodiments can be realized as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term“data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0309] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or code portions). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0310] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, and / or by a combination of computer and special purpose logic circuitry. Devices suitable for the execution of a computer program include, by way of example, general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0311] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0312] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of a particular art. In this patent document, certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in various suitable sub-combinations. Furthermore, although features may be described above as operating in certain combinations and even initially claimed in the same manner, in certain circumstances one or more features from the claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0313] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or to perform all the operations shown, in order to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0314] Only a few implementations and examples are described, and other implementations, enhancements and variations can be made based on what is described and shown in this patent document.
Claims
1. A method of processing video data, comprising: performing a conversion between a video comprising one or more subpictures and a bitstream of the video according to a format rule, wherein the format rule specifies, in response to a scalable-nested supplemental enhancement information message comprising one or more subpicture level information supplemental enhancement information messages, setting a first syntax element in the scalable-nested supplemental enhancement information message in the bitstream to a particular value, wherein the particular value of the first syntax element indicates that one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message are applicable to a particular output video layer set, and wherein a payload type 203 of a supplemental enhancement information message indicates that the supplemental enhancement information message is the subpicture level information supplemental enhancement information message.
2. The method of claim 1, wherein wherein the format rule specifies, in response to the scalable-nested supplemental enhancement information message comprising the one or more subpicture level information supplemental enhancement information messages, a value of a second syntax element in the bitstream is equal to 0, and wherein the value of the second syntax element being equal to 0 specifies that the one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message that are applicable to a specified one or more output video layer sets or a specified one or more video layers are applicable to all subpictures of the specified one or more output video layer sets or all subpictures of the specified one or more video layers.
3. The method of claim 1, wherein, the particular value of the first syntax element is equal to 1.
4. The method of claim 1, wherein, the format rule specifies that the scalable-nested supplemental enhancement information message is not allowed to comprise a first supplemental enhancement information message of a first payload type and a second supplemental enhancement information message of a second payload type, wherein the first payload type comprises a payload type of a buffering period supplemental enhancement information message, a payload type of a picture timing supplemental enhancement information message, a payload type of a decoding unit information supplemental enhancement information message, or a payload type of a subpicture level information supplemental enhancement information message, and wherein the second payload type comprises a payload type that is not one of: (i) the payload type of the buffering period supplemental enhancement information message, (ii) the payload type of the picture timing supplemental enhancement information message, (iii) the payload type of the decoding unit information supplemental enhancement information message, and (iv) the payload type of the subpicture level information supplemental enhancement information message.
5. The method of claim 1, wherein, performing the conversion comprises encoding the video into the bitstream.
6. The method of claim 1, wherein, performing the conversion comprises decoding the video from the bitstream.
7. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, the instructions, when executed by the processor, cause the processor to: perform a conversion between a video comprising one or more subpictures and a bitstream of the video according to a format rule, wherein the format rule specifies, in response to a scalable-nested supplemental enhancement information message comprising one or more subpicture level information supplemental enhancement information messages, setting a first syntax element in the scalable-nested supplemental enhancement information message in the bitstream to a particular value, wherein the particular value of the first syntax element indicates that one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message are applicable to a particular output video layer set, and wherein a particular value of the first syntax element indicates that one or more of the scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message are applicable to a particular output video layer set, and wherein a payload type of a supplemental enhancement information message 203 indicates that the supplemental enhancement information message is the subpicture level information supplemental enhancement information message.
8. The apparatus of claim 7, wherein, the format rule specifies that, in response to the scalable-nested supplemental enhancement information message including the one or more subpicture level information supplemental enhancement information messages, a value of a second syntax element in the bitstream is equal to 0, and wherein the value of the second syntax element equal to 0 specifies that the one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message that are applicable to a specified one or more output video layer sets or a specified one or more video layers are applicable to all subpictures of the specified one or more output video layer sets or all subpictures of the specified one or more video layers.
9. The apparatus of claim 7, wherein, a particular value of the first syntax element is equal to 1.
10. The apparatus of claim 7, wherein, the format rule specifies that the scalable-nested supplemental enhancement information message is not allowed to include a first supplemental enhancement information message of a first payload type and a second supplemental enhancement information message of a second payload type, wherein the first payload type includes a payload type of a buffering period supplemental enhancement information message, a payload type of a picture timing supplemental enhancement information message, a payload type of a decoding unit information supplemental enhancement information message, or a payload type of a subpicture level information supplemental enhancement information message, and wherein the second payload type includes a payload type that is not one of: (i) the payload type of the buffering period supplemental enhancement information message, (ii) the payload type of the picture timing supplemental enhancement information message, (iii) the payload type of the decoding unit information supplemental enhancement information message, and (iv) the payload type of the subpicture level information supplemental enhancement information message.
11. A non-transitory computer-readable storage medium storing instructions that cause a processor to: perform a conversion between a video including one or more subpictures and a bitstream of the video according to a format rule, wherein, the format rule specifies that, in response to a scalable-nested supplemental enhancement information message including one or more subpicture level information supplemental enhancement information messages, a first syntax element in the scalable-nested supplemental enhancement information message in the bitstream is set to a particular value, wherein a particular value of the first syntax element indicates that one or more of the scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message are applicable to a particular output video layer set, and wherein a payload type of a supplemental enhancement information message 203 indicates that the supplemental enhancement information message is the subpicture level information supplemental enhancement information message.
12. The non-transitory computer-readable storage medium of claim 11, wherein the format rule specifies that, in response to the scalable-nested supplemental enhancement information message including the one or more subpicture-level information supplemental enhancement information messages, a value of a second syntax element in the bitstream is equal to 0, and wherein the value of the second syntax element equal to 0 specifies that the one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message that are applicable to a specified one or more output video layer sets or a specified one or more video layers are applicable to all subpictures of the specified one or more output video layer sets or all subpictures of the specified one or more video layers.
13. The non-transitory computer-readable storage medium of claim 11, wherein, a particular value of the first syntax element is equal to 1.
14. The non-transitory computer-readable storage medium of claim 11, wherein, the format rule specifies that the scalable-nested supplemental enhancement information message is not allowed to include a first supplemental enhancement information message of a first payload type and a second supplemental enhancement information message of a second payload type, wherein the first payload type includes a payload type of a buffering period supplemental enhancement information message, a payload type of a picture timing supplemental enhancement information message, a payload type of a decoding unit information supplemental enhancement information message, or a payload type of a subpicture-level information supplemental enhancement information message, and wherein the second payload type includes a payload type that is not one of: (i) the payload type of the buffering period supplemental enhancement information message, (ii) the payload type of the picture timing supplemental enhancement information message, (iii) the payload type of the decoding unit information supplemental enhancement information message, and (iv) the payload type of the subpicture-level information supplemental enhancement information message.
15. A method for storing a bitstream of a video, comprising: generating the bitstream of the video including one or more subpictures according to a format rule; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the format rule specifies that, in response to a scalable-nested supplemental enhancement information message including one or more subpicture-level information supplemental enhancement information messages, a first syntax element in the scalable-nested supplemental enhancement information message in the bitstream is set to a particular value, wherein the particular value of the first syntax element indicates that one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message are applicable to a particular output video layer set, and wherein a payload type of a supplemental enhancement information message 203 indicates that the supplemental enhancement information message is the subpicture-level information supplemental enhancement information message.
16. The method of claim 15, wherein the format rule specifies that, in response to the scalable-nested supplemental enhancement information message including the one or more subpicture-level information supplemental enhancement information messages, a value of a second syntax element in the bitstream is equal to 0, and wherein the value of the second syntax element equal to 0 specifies that the one or more scalable-nested supplemental enhancement information messages in the scalable-nested supplemental enhancement information message that are applicable to a specified one or more output video layer sets or a specified one or more video layers are applicable to all subpictures of the specified one or more output video layer sets or all subpictures of the specified one or more video layers. wherein a value of the second syntax element equal to 0 specifies that the one or more scalably nested supplemental enhancement information messages applicable to the specified one or more output video layer sets or the specified one or more video layers, of the one or more supplemental enhancement information messages that are scalably nested, are applicable to all sub-pictures of the specified one or more output video layer sets or all sub-pictures of the specified one or more video layers.
17. The method of claim 15, wherein, a particular value of the first syntax element is equal to 1.
18. The method of claim 15, wherein, the format rule specifies that the scalably nested supplemental enhancement information messages are not allowed to include a first supplemental enhancement information message of a first payload type and a second supplemental enhancement information message of a second payload type, wherein the first payload type comprises a payload type of a buffering period supplemental enhancement information message, a payload type of a picture timing supplemental enhancement information message, a payload type of a decoding unit information supplemental enhancement information message, or a payload type of a sub-picture level information supplemental enhancement information message, and wherein the second payload type comprises a payload type that is not one of: (i) the payload type of the buffering period supplemental enhancement information message, (ii) the payload type of the picture timing supplemental enhancement information message, (iii) the payload type of the decoding unit information supplemental enhancement information message, and (iv) the payload type of the sub-picture level information supplemental enhancement information message.
19. A video decoding apparatus comprising a processor configured to implement a method recited in any of claims 1 to 6.
20. A video encoding apparatus comprising a processor configured to implement a method recited in any of claims 1 to 6.