Sublayer signaling in video coding
By introducing sub-picture level information (SEI) messages into video encoding and decoding, the problem of managing sub-picture sequence level information in multi-layer video encoding and decoding is solved, which improves encoding and decoding efficiency and resource utilization, and achieves more efficient bitstream management and decoder compatibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN CO LTD
- Filing Date
- 2021-06-07
- Publication Date
- 2026-08-04
AI Technical Summary
Existing video encoding and decoding technologies struggle to effectively manage the level information of signaling notification sub-image sequences when processing multi-layer video encoding and decoding, resulting in low encoding and decoding efficiency and insufficient resource utilization.
By introducing Subpicture Level Information (SEI) messages, the level information of subpicture sequences is specified, including the maximum number, level information existence, cycle score, and level indicator, ensuring effective management and signaling notification of the level information of subpicture sequences in the bitstream.
It improves efficiency and resource utilization in the video encoding and decoding process, especially in multi-layer video encoding and decoding, achieving more efficient bitstream management and decoder compatibility.
Smart Images

Figure CN115699728B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application promptly claims priority and interest in U.S. Provisional Patent Application No. 63 / 036,365, filed June 8, 2020. The entire disclosure of the aforementioned application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] The patent document relates to image and video encoding and decoding. Background Technology
[0004] Digital video consumes the largest share of bandwidth in the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This patent document discloses a technology that can be used by video encoders and decoders to process encoded or decoded representations of video or images.
[0006] In one example aspect, a method for processing video data is disclosed. The method includes performing a conversion between the video and a bitstream of the video comprising one or more Output Layer Sets (OLS) according to a rule. The rule specifies that a Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of a subpicture sequence in a set of encoded and decoded video sequences to which the SLI SEI message is applied. The syntax structure of the SLI SEI message includes (1) a first syntax element specifying the maximum number of sublayers of the subpicture sequence, (2) a second syntax element specifying whether the level information of the subpicture sequence exists in one or more sublayer representations, and (3) a loop of multiple sublayers, each sublayer associated with a fraction of the bitstream level limit and a level indicator indicating the level to which each subpicture sequence conforms.
[0007] In another example, a method for processing video data is disclosed. The method includes performing a conversion between the current access unit of a video comprising one or more Output Layer Sets (OLSs) and the video bitstream according to a rule. The rule specifies that Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) messages include information about the level of subpicture sequences in the set of encoded and decoded video sequences to which the SLI SEI message is applied. The SLI SEI message persists from the current access unit in decoding order until the end of the bitstream, or until the next access unit contains a subsequent SLI SEI message containing content different from the SLI SEI message.
[0008] In another example, a method for processing video data is disclosed. The method includes performing a representation conversion between the currently accessed unit of a video comprising one or more Output Layer Sets (OLSs) and the bitstream of the video, according to a rule. A Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of a subpicture sequence in the set of codec video sequences of one or more OLSs to which the SLI SEI message is applied. A layer in one or more OLSs whose number of subpictures, indicated by variables in its reference sequence parameter set, is greater than one is called a multi-subpicture layer. The codec video sequences in the OLS set are called target codec video sequences (CVSs). The rule specifies that a subpicture sequence includes (1) all subpictures in the target CVS that have the same subpicture index and belong to a layer in a multi-subpicture layer, and (2) all subpictures in the target CVS that have a subpicture index of 0 and belong to a layer in the OLS but not to a multi-subpicture layer.
[0009] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video and a video codec representation comprising one or more video sublayers, wherein the codec representation conforms to a format rule; wherein the format rule specifies a syntax structure including loops over multiple sublayers in the codec representation and one or more syntax fields indicating each sublayer included in the syntax structure, wherein the syntax structure includes information about a score for signaling notification and a reference level indicator.
[0010] In another example, a different video processing method is disclosed. This method includes performing a conversion between a video comprising one or more sub-pictures and a codec representation of the video, wherein the conversion uses or generates supplementary enhancement information at the one or more sub-picture level.
[0011] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.
[0012] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.
[0013] In yet another example, a computer-readable medium on which code is stored is disclosed. This code implements one of the methods described herein in the form of processor-executable code.
[0014] These and other features will be described in this document. Attached Figure Description
[0015] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 tiles and 3 raster scan strips.
[0016] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.
[0017] Figure 3 An example of an image divided into slices and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0018] Figure 4 The image is shown as being divided into 15 slices, 24 strips, and 24 sub-images.
[0019] Figure 5 This is a block diagram of an example video processing system.
[0020] Figure 6 This is a block diagram of a video processing device.
[0021] Figure 7 This is a flowchart of an example method for video processing.
[0022] Figure 8 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0023] Figure 9 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0024] Figure 10 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0025] Figure 11 An example of a typical sub-picture-based viewport-dependent 360° video encoding / decoding scheme is shown.
[0026] Figure 12 A viewport-dependent 360° video encoding and decoding scheme based on sub-pictures and spatial scalability is presented.
[0027] Figure 13 This is a flowchart of a method for processing video data according to one or more embodiments of the present technology.
[0028] Figure 14 This is a flowchart illustrating another method for processing video data according to one or more embodiments of the present technology.
[0029] Figure 15 This is a flowchart illustrating another method for processing video data according to one or more embodiments of the present technology. Detailed Implementation
[0030] Chapter headings are used in this document for ease of understanding, not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter only. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding, not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. In this document, edits are displayed in the text relative to the current draft of the VVC specification, with strikethrough indicating deleted text and highlighting (including bold and italic) indicating added text.
[0031] 1. Overview
[0032] This document relates to video codec technology. Specifically, it concerns the level information for specifying and signaling sub-picture sequences. It can be applied to any video codec standard or non-standard video codec that supports single-layer and multi-layer video codecs, such as the under-development Multi-Functional Video Codec (VVC).
[0033] 2. Abbreviation
[0034] APS Adaptive Parameter Set
[0035] AU Access Unit
[0036] AUD Access Unit Delimiter
[0037] AVC Advanced Video Codec
[0038] BP buffer period
[0039] CLVS codec layer video sequence
[0040] CPB encoded image buffer
[0041] CRA (Clean Random Access)
[0042] CTU (Codec Tree Unit)
[0043] CVS codec video sequence
[0044] DPB Decoding Image Buffer
[0045] DPS Decoding Parameter Set
[0046] DUI Decoding Unit Information
[0047] EOB End of Bitstream
[0048] End of EOS sequence
[0049] GCI General Constraint Information
[0050] GDR gradually decoded and refreshed
[0051] HEVC High-Efficiency Video Encoding and Decoding
[0052] HRD Hypothetical Reference Decoder
[0053] IDR Instant Decoding and Refresh
[0054] JEM Collaborative Exploration Mode
[0055] MCTS Motion Constraint Pieces
[0056] NAL Network Abstraction Layer
[0057] OLS Output Layer Set
[0058] PH image header
[0059] PPS Image Parameter Set
[0060] PT Image Time Sequence
[0061] PTL refers to profiles, tiers, and levels.
[0062] PU Image Unit
[0063] RRP reference image resampling
[0064] RBSP raw byte sequence payload
[0065] SEI Supplemental Enhancement Information
[0066] SH strip header
[0067] SLI sub-image level information
[0068] SPS Sequence Parameter Set
[0069] SVC Scalable Video Codec
[0070] VCL (Video Codec Layer)
[0071] VPS Video Parameter Set
[0072] VTM VVC Test Model
[0073] VUI Video Availability Information
[0074] VVC Multi-Functional Video Encoding and Decoding
[0075] 3. Preliminary Discussion
[0076] Video coding standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture, utilizing time prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new coding standard is to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Universal Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted into the VVC standard at every JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.
[0077] 3.1. Image Segmentation Schemes in HEVC
[0078] HEVC includes four different image segmentation schemes: regular striping, subordinate striping, slice, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end latency.
[0079] The regular stripes are similar to those in H.264 / AVC. Each regular stripe is encapsulated in its own NAL unit, and intra-image predictions (intra-sample prediction, motion information prediction, coding pattern prediction) and entropy coding dependencies across stripe boundaries are disabled. Therefore, regular stripes can be reconstructed independently of other regular stripes within the same image (although interdependencies may still exist due to loop filtering operations).
[0080] Regular striping is the only tool available for parallelization, and it is also available in almost the same form in H.264 / AVC. Parallelization based on regular striping requires minimal inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictive encode-decode images, which is typically much heavier than inter-processor or inter-core data sharing due to intra-image prediction). However, for the same reason, using regular striping results in significant encoding / decoding overhead due to the bit cost of the stripe header and the lack of prediction across stripe boundaries. Furthermore, due to the intra-image independence of regular striping and the fact that each regular stripe is encapsulated in its own NAL unit, regular striping (compared to other tools mentioned below) also serves as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching present conflicting requirements for stripe layout in images. This recognition led to the development of the parallelization tools mentioned below.
[0081] Slave stripes have short stripe headers and allow the bitstream to be split at tree block boundaries without disrupting any in-picture predictions. Essentially, slave stripes provide the ability to fragment regular stripes into multiple NAL units, thereby reducing end-to-end latency by allowing a portion of the regular stripe to be sent before the entire regular stripe's encoding is complete.
[0082] In WPP, images are segmented into single-row codec tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding a CTB row is delayed by two CTBs, ensuring that data related to the top of the CTB and the right side of the subject CTB is available before the subject CTB is decoded. Using this staggered start (which looks like a wavefront when graphically represented), parallelization can use as many processors / cores as the number of CTB rows contained in the image. Because intra-image prediction between adjacent tree block rows within an image is allowed, the inter-processor / inter-core communication required to implement intra-image prediction can be substantial. WPP partitioning does not generate additional NAL units compared to when it is not applied, therefore WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, but with some encoding / decoding overhead.
[0083] A slice defines the horizontal and vertical boundaries that divide an image into slice columns and rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right of the image. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0084] Before decoding the top-left CTB of the next slice in the order of slice raster scans of the image, the scan order of the CTBs is changed to be local within the slice (in the order of slice CTB raster scans). Similar to regular stripes, slices break the intra-image prediction dependencies and entropy decoding dependencies. However, they do not need to be included in a single NAL unit (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-image prediction between processing units decoding adjacent slices is limited to transmitting a shared stripe header when the stripe spans more than one slice, and loop filtering associated with the sharing of reconstructed samples and metadata. When a stripe includes more than one slice or WPP segment, the entry point byte offset of each slice or WPP segment in the stripe, except for the first slice or WPP segment, is signaled in the stripe header.
[0085] For simplicity, HEVC specifies restrictions on the application of four different image segmentation schemes. A given codec video sequence cannot simultaneously include both slices and wavefronts of most of the levels specified in the HEVC standard. For each strip and slice, one or both of the following conditions must be met: 1) All coded tree blocks in a strip belong to the same slice; 2) All coded tree blocks in a slice belong to the same strip. Finally, a wavefront segment contains exactly one CTB line, and when using WPP, if a strip begins at a CTB line, it must end at the same CTB line.
[0086] The latest modifications to HEVC are specified in the JCT-VC output file JCTVC-AC1005 “HEVC Additional Supplemental Enhancement Information (Draft 4)” published by J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, and Y.-K. Wang (eds.) on October 24, 2017, and are publicly available here: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this modification, HEVC specifies three MCTS-related SEI messages: the Time MCTS SEI message, the MCTS Extracted Information Set SEI message, and the MCTS Extracted Information Nested SEI message.
[0087] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream, and signaling notifies the MCTS. For each MCTS, motion vectors are restricted to pointing to full-sampled positions within the MCTS and fractional-sampled positions that require interpolation only from full-sampled positions within the MCTS, and motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction are not allowed. In this way, each MCTS can be decoded independently, and there are no slices not included in the MCTS.
[0088] The MCTS Extraction Information Set (SEI) message provides supplementary information (specified as part of the SEI message semantics) that can be used in MCTS sub-bitstream extraction to generate a bitstream conforming to the MCTS set. This information consists of multiple extraction information sets, each defining multiple MCTS sets and containing RBSP bytes that will replace the VPS, SPS, and PPS during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) typically need to have different values.
[0089] 3.2. Image Segmentation in VVC
[0090] In VVC, an image is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the image. The CTUs in a slice are scanned in raster scan order within that slice.
[0091] A strip consists of an integer number of consecutive complete CTU lines from an integer number of complete slices or images.
[0092] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a complete sequence of stripes in a raster scan of the image. In rectangular stripe mode, a stripe contains multiple complete slices that together form a rectangular area of the image, or multiple consecutive complete CTU rows of a single slice that together form a rectangular area of the image. Slices within a rectangular stripe are scanned in slice raster scan order within the rectangular area corresponding to that stripe.
[0093] A sub-image contains one or more stripes that collectively cover a rectangular area of the image.
[0094] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 slices and 3 raster scan strips.
[0095] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 strips (6 strip columns and 4 strip rows) and 9 rectangular strips.
[0096] Figure 3 An example of an image divided into slices and rectangular strips is shown, where the image is divided into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.
[0097] Figure 4 An example of sub-image segmentation of an image is shown, where the image is segmented into 18 pieces: 12 pieces on the left-hand side, each covering a 4×4 CTU strip, and 6 pieces on the right-hand side, each covering a 2×2 CTU strip, forming two vertically stacked strips, resulting in a total of 24 strips and 24 sub-images of different dimensions (each strip being a sub-image).
[0098] 3.3. Changes in image resolution within a sequence
[0099] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence with a new SPS begins with an IRAP picture. VVC allows changing the picture resolution within a sequence at locations where IRAP pictures are not encoded; IRAP pictures are always intra-frame encoded and decoded. This feature is sometimes called Reference Picture Resampling (RPR) because it requires resampling the reference picture used for inter-frame prediction when the reference picture has a different resolution than the current picture being decoded.
[0100] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, the same as in motion-compensated interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process where the scaling ratio ranges from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height, as well as the left, right, top, and bottom scaling offsets specified for the reference and current images.
[0101] Other aspects of the VVC design that support this feature differ from HEVC include: i) Picture resolution and the corresponding consistency window are signaled in the PPS instead of the SPS, where the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture storage (a slot in the DPB used to store a decoded picture) occupies the buffer size required to store the decoded picture with the maximum picture resolution.
[0102] 3.4. Scalable Video Codec (SVC) in General and VVC
[0103] Scalable Video Coding (SVC, sometimes also called scalability in video coding) refers to video coding using a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a basic quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously encoded layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of the layer below the intermediate layer (such as a base layer or any intermediate enhancement layer) and simultaneously used as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, there can be multiple views, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).
[0104] In SVC, parameters used by the encoder or decoder are grouped into parameter sets based on the codec level (e.g., video level, sequence level, picture level, stripe level, etc.), and these sets may be utilized at that codec level. For example, parameters that can be utilized by one or more codec video sequences at different layers in a bitstream can be included in the Video Parameter Set (VPS), and parameters that can be utilized by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters used by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single strip can be included in the stripe header. Likewise, indications of which parameter set(s) a particular layer uses at a given time can be provided at various codec levels.
[0105] Because VVC supports Reference Picture Resampling (RPR), it's possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions in VVC) without requiring any additional signal processing-level codec tools, as the upsampling required for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires a higher level of syntax changes (compared to no scalability support). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standards (including extensions to AVC and HEVC), VVC's scalability is designed to be as friendly as possible to single-layer decoder designs. The decoding capability of a multi-layer bitstream is specified as if there were only one layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layer bitstreams don't require many changes to decode multi-layer bitstreams. Compared to the multi-layer extension designs of AVC and HEVC, HLS is significantly simplified at the expense of some flexibility. For example, IRAPU requires a picture of every layer present in CVS.
[0106] 3.5. Viewport-dependent 360° video streaming based on sub-images
[0107] In 360° video streaming (also known as omnidirectional video), at any given moment, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user, who can rotate their head at any time to change their viewing orientation, thus changing the current viewport. While it is desirable to have at least some lower-quality representations of areas not covered by the current viewport available at the client end, ready to be presented to the user in case they suddenly change their viewing orientation anywhere on the sphere, the high-quality representation of the omnidirectional video is only needed for the current viewport being presented to the user. This optimization is achieved by segmenting the high-quality representation of the entire omnidirectional video into sub-pictures with appropriate granularity. Using VVC, these two representations can be encoded as two independent layers.
[0108] A typical sub-image-based viewport-dependent 360° video transmission scheme is as follows: Figure 11 As shown, the higher resolution representation of the full video consists of sub-pictures, while the lower resolution representation does not use sub-pictures and can be encoded and decoded using less frequent random access points than the higher resolution representation. The client receives the lower resolution full video, while for the higher resolution video, it only receives and decodes the sub-pictures covering the current viewport.
[0109] The latest VVC draft specification also supports such Figure 12 The improved 360° video encoding / decoding scheme is shown. (Compared to...) Figure 11The only difference between the methods shown is that inter-layer prediction (ILP) is applied to... Figure 12 The method shown.
[0110] 3.6. Parameter Set
[0111] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. All versions of AVC, HEVC, and VVC support SPS and PPS. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0112] The Sequence-Level Prefix (SPS) is designed to carry sequence-level header information, while the Picture-Level Prefix (PPS) is designed to carry infrequently changing picture-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or picture, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thereby not only avoiding the need for redundant transmission but also improving error resilience.
[0113] The VPS was introduced to carry sequence-level header information shared by all layers in a multi-layer bitstream.
[0114] The purpose of APS is to carry such image-level or stripe-level information, which requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.
[0115] 3.7. Grades, Levels, and Classes
[0116] Video codec standards typically specify levels and grades. Some video codec standards also define layers, such as HEVC and the developing VVC.
[0117] Grades, levels, and tiers specify limitations on the bitstream, thus limiting the capabilities required to decode it. Grades, levels, and tiers can also be used to indicate points of interoperability between different decoder implementations.
[0118] Each grade specifies a subset of algorithmic features and limitations that all decoders conforming to that grade should support. Note that the encoder does not need to use all the codec tools or features supported in the grade, while the decoder conforming to the grade needs to support all codec tools or features.
[0119] Each level of a hierarchy specifies a set of restrictions on the values that bitstream syntax elements can take. All hierarchies typically use the same set of hierarchies and levels, but different implementations may support different hierarchies, and within a single hierarchy, each supported hierarchy may support different levels. For any given hierarchy, the level of the hierarchy typically corresponds to a specific decoder processing load and memory capacity.
[0120] The capabilities of a video decoder that conforms to a video codec specification are defined by its ability to decode video streams that conform to the constraints of the grade, level, and tier specified in the video codec specification. When expressing the capabilities of a decoder for a specific grade, the grades and tiers supported by that grade should also be expressed.
[0121] 3.8. Level information for specifying and signaling notification sub-picture sequences in VVC
[0122] In the latest VVC draft text, the level information of the subpicture sequence is specified and signaled in VVC through the Subpicture Level Information (SLI) SEI message. The subpicture sequence can be extracted from the bitstream by applying the subpicture subbitstream extraction procedure specified in Clause C.7 of VVC.
[0123] The syntax and semantics of the Subpicture Level Information (SEI) message in the latest VVC draft text are as follows.
[0124] D.7.1 Sub-picture level information SEI message syntax
[0125]
[0126] D.7.2 Sub-picture level information SEI message semantics
[0127] When testing the consistency of an extracted bitstream containing sub-picture sequences according to Appendix A, the Subpicture Level Information (SEI) message contains information about the level to which the subpicture sequences in the bitstream conform.
[0128] When a Subpicture Level Information (SEI) message exists in any picture within a CLVS, the SEI message should also exist in the first picture of the CLVS. SEI messages are stored in the current layer from the current picture in decoding order until the end of the CLVS. All SEI messages applied to the same CLVS should have the same content. A subpicture sequence consists of all subpictures within the CLVS that have the same subpicture index value.
[0129] The requirement for bitstream consistency is that when the subpicture level information (SEI) message exists in CLVS, the value of sps_subpic_treated_as_pic_flag[i] should be equal to 1 for each i value in the range of 0 to sps_num_subpics_minus1 (inclusive).
[0130] The increment 1 in sli_num_ref_levels_minus1 specifies the number of reference levels for signaling notifications for each of the 1+1 subpics in the sps_num_subpics_minus1 subpics.
[0131] A value of 0 for `sli_cbr_constraint_flag` indicates that the hypothetical stream scheduler (HSS) operates in intermittent bit rate mode in order to decode any sub-bitstream resulting from a sub-picture extracted according to clause C.7 by using an HRD of any CPB specification in the extracted sub-bitstream. A value of 1 for `sli_cbr_constraint_flag` indicates that the HSS operates in constant bit rate (CBR) mode.
[0132] A value of 1 for sli_explicit_fraction_present_flag indicates that the syntax element sli_ref_level_fraction_minus1[i] exists. A value of 0 for sli_explicit_fraction_present_flag indicates that the syntax element sli_ref_level_fraction_minus1[i] does not exist.
[0133] Incrementing 1 to sli_num_subpics_minus1 specifies the number of subpicks in the CLVS image. When present, the value of sli_num_subpics_minus1 should be equal to the value of sps_num_subpics_minus1 in the SPS referenced by the image in the CLVS.
[0134] sli_alignment_zero_bit should be equal to 0.
[0135] sli_non_subpic_layers_fraction[i] specifies the fraction of the bitstream level limit associated with the layer in the bitstream where sps_num_subpics_minus1 equals 0. sli_non_subpic_layers_fraction[i] should be 0 when vps_max_layers_minus1 equals 0 or when sps_num_subpics_minus1 equals 0 for no layers in the bitstream.
[0136] sli_ref_level_idc[i] indicates the level that each sub-picture specified in Appendix A conforms to. The bitstream should not contain values of sli_ref_level_idc other than those specified in Appendix A. Other values of sli_ref_level_idc[i] are reserved for future use by ITU-T|ISO / IEC. The requirement for bitstream conformance is that the value of sli_ref_level_idc[0] should be equal to the value of general_level_idc for the bitstream, and for any value where i is greater than 0 and k is greater than i, the value of sli_ref_level_idc[i] should be less than or equal to sli_ref_level_idc[k].
[0137] The increment 1 in sli_ref_level_fraction_minus1[i][j] specifies the score of the level limit associated with sli_ref_level_idc[i] that the j-th sub-picture conforms to, as specified in Clause A.4.1.
[0138] The variable SubpicSizeY[j] is set to equal to (sps_subpic_width_minus1[j]+1)*CtbSizeY*(sps_subpic_height_minus1[j]+1)*CtbSizeY.
[0139] When it does not exist, the value of sli_ref_level_fraction_minus1[i][j] is inferred to be equal to Ceil(256*SubpicSizeY[j]÷PicSizeInSamplesY*MaxLumaPs(general_level_idc)÷MaxLumaPs(sli_ref_level_idc[i])-1.
[0140] The variable LayerRefLevelFraction[i][j] is set to equal sli_ref_level_fraction_minus1[i][j]+1.
[0141] The variable OlsRefLevelFraction[i][j] is set to equal sli_non_subpic_layers_fraction[i]+(256-sli_non_subpic_layers_fraction[i])÷256*(sli_ref_level_fraction_minus1[i][j]+1).
[0142] The variables SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] are derived as follows:
[0143] SubpicCpbSizeVcl[i][j]=Floor(CpbVclFactor*MaxCPB*OlsRefLevelFraction[i][j]÷256) (D.6)
[0144] SubpicCpbSizeNal[i][j]=Floor(CpbNalFactor*MaxCPB*OlsRefLevelFraction[i][j]÷256) (D.7)
[0145] MaxCPB is derived from sli_ref_level_idc[i] as specified in Clause A.4.2.
[0146] The variables SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j] are derived as follows:
[0147] SubpicBitRateVcl[i][j]=Floor(CpbVclFactor*ValBR*OlsRefLevelFraction[0][j]÷256) (D.8)
[0148] SubpicBitRateNal[i][j]=Floor(CpbNalFactor*ValBR*OlsRefLevelFraction[0][j]÷256) (D.9)
[0149] The value of ValBR is derived as follows:
[0150] – When bit_rate_value_minus1[Htid][ScIdx] is available in the corresponding HRD parameter in the VPS or SPS, ValBR is set to equal to (bit_rate_value_minus1[Htid][ScIdx]+1)*2 (6+bit_rate_scale) , where Htid is the sub-level index under consideration, and ScIdx is the scheduling index under consideration.
[0151] Otherwise, ValBR is set to MaxBR derived from sli_ref_level_idc[0] as specified in Clause A.4.2.
[0152] Note 1: When extracting subpicks, the resulting bitstream has a CpbSize (indicated or inferred in VPS, SPS) greater than or equal to SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] and a bitrate (indicated or inferred in VPS, SPS) greater than or equal to SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j].
[0153] The requirement for bitstream consistency is that each layer in a bitstream generated from layers in the input bitstream of the extraction process that extract the j-th subpick for the range j (inclusive) from 0 to sps_num_subpics_minus1 (0 to num_ref_level_minus1) and whose level is equal to sli_ref_level_idc[i], and which has a general_tier_flag equal to 0 and a level equal to sli_ref_level_idc[i] for the range i (inclusive) from 0 to num_ref_level_minus1 (0 to num_ref_level_minus1), shall comply with the following constraints as specified in Appendix C for each bitstream consistency test:
[0154] Ceil(256*SubpicSizeY[j]÷LayerRefLevelFraction[i][j]) should be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1 for level sli_ref_level_idc[i].
[0155] The value of Ceil(256*(sps_subpic_width_minus1[j]+1)*CtbSizeY÷LayerRefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0156] The value of Ceil(256*(sps_subpic_height_minus1[j]+1)*CtbSizeY÷LayerRefLevelFraction[i][j]) should be less than or equal to Sqrt(MaxLumaPs*8).
[0157] The value of SubpicWidthInTiles[j] should be less than or equal to MaxTileCOLS, and the value of SubpicHeightInTiles[j] should be less than or equal to MaxTileRows, where MaxTileCOLS and MaxTileRows are specified in Table A.1 for level sli_ref_level_idc[i].
[0158] The value of SubpicWidthInTiles[j]*SubpicHeightInTiles[j] should be less than or equal to MaxTileCOLS*MaxTileRows*LayerRefLevelFraction[i][j], where MaxTileCOLS and MaxTileRows are specified in Table A.1 for level sli_ref_level_idc[i].
[0159] The requirement for bitstream consistency is that a bitstream generated from the extraction of the j-th subpick within the range of 0 to sps_num_subpics_minus1 (inclusive) and conforming to a level equal to ref_level_idc[i] with general_tier_flag equal to 0 and (for i within the range of 0 to sli_num_ref_level_minus1 (inclusive)) level equal to ref_level_idc[i], shall comply with the following constraints as specified in Appendix C for each bitstream consistency test:
[0160] The sum of the NumBytesInNalUnit variables corresponding to AU 0 of the j-th subpic should be less than or equal to the FormatCapabilityFactor*(Max(SubpicSizeY[j],fR*MaxLumaSr*OlsRefLevelFraction[i][j]÷256)+MaxLumaSr*(AuCpbRemovalTime[0]-AuNominalRemovalTime[0])*OlsRefLevelFraction[i][j])÷(256*MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3, respectively, applied to AU 0 of level sli_ref_level_idc[i], and MinCr is derived as indicated in A.4.2.
[0161] The sum of the NumBytesInNalUnit variables corresponding to the AU n (n > 0) of the j-th sub-image should be less than or equal to FormatCapabilityFactor * MaxLumaSr * (AuCpbRemovalTime[n] - AuCpbRemovalTime[n-1]) * OlsRefLevelFraction[i][j] ÷ (256 * MinCr), where MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3, respectively, applied to the AU n of level sli_ref_level_idc[i], and MinCr is derived as indicated in A.4.2.
[0162] The value of the subpic sequence level indicator SubpicLevelIdc is derived as follows:
[0163]
[0164] Subpic sequence bitstreams that meet the criteria of general_tier_flag equal to 0 and level equal to SubpicLevelIdc should comply with the following constraints as specified in Appendix C for each bitstream consistency test:
[0165] For the VCL HRD parameter, SubpicCpbSizeVcl[i] should be less than or equal to CpbVclFactor*MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in bits of CpbVclFactor.
[0166] For the NAL HRD parameter, SubpicCpbSizeNal[i] should be less than or equal to CpbNalFactor*MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in bits of CpbNalFactor.
[0167] For the VCL HRD parameter, SubpicBitRateVcl[i] should be less than or equal to CpbVclFactor*MaxBR, where CpbVclFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in bits of CpbVclFactor.
[0168] For the NAL HRD parameter, SubpicBitRateNal[i] should be less than or equal to CpbNalFactor*MaxBR, where CpbNalFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in bits of CpbNalFactor.
[0169] Note 2: When extracting subpicture sequences, the resulting bitstream has a CpbSize (indicated or inferred in VPS, SPS) greater than or equal to SubpicCpbSizeVcl[i][j] and SubpicCpbSizeNal[i][j] and a bitrate (indicated or inferred in VPS, SPS) greater than or equal to SubpicBitRateVcl[i][j] and SubpicBitRateNal[i][j].
[0170] 4. The technical problem solved by the disclosed technical solution
[0171] The existing VVC design used for specifying and signaling information at the sub-image sequence level has the following problems:
[0172] (1) SLI SEI messages only signal a single-level set of information for the sub-picture sequence, regardless of the value of the highest TemporalId. However, just as each picture has a bitstream of individual sub-pictures, different sub-layer representations can conform to different levels.
[0173] (2) SLI SEI messages are defined as available only by being located in the bitstream. However, similar to parameter sets and other HRD-related SEI information, SLI SEI information should also be made available externally.
[0174] (3) The persistence of an SLI SEI message is defined within one CVS. However, in most cases, an SLI SEI message will be applied to several consecutive CVSs, and usually the entire bitstream.
[0175] (4) The definition of a sub-image sequence does not cover the case where there are one or more layers where each image has a single sub-image.
[0176] (5) Lack of requirement that when the SLI SEI message exists in CVS, the value of sps_num_subpics_minus1 should be the same for all SPS referenced by the images in layers with multiple subpicks per image. Otherwise, it is meaningless to require that the value of sli_num_subpics_minus1 be equal to the value of sps_num_subpics_minus1.
[0177] (6) The semantics of sli_num_subpics_minus1 do not apply to cases where there are one or more layers where each image has multiple subpicks.
[0178] (7) The variables SubpicLevelIdc and SubpicLevelIdx need to be specified as subpic sequence-specific because different subpic sequences extracted from the same original bitstream can conform to different levels.
[0179] 5. Examples of solutions and implementation methods
[0180] To address the aforementioned and other issues, methods outlined below are disclosed. These items should be considered as examples for explaining general concepts, and not interpreted in a narrow sense. Furthermore, these items can be used individually or in combination in any way.
[0181] 1) To address the first issue, add sli_max_sublayers_minus1, sli_sub layers_info_present_flag, and loops for sublayers with scores and reference level indicators for signaling notifications to align with the signaling of level information in the PTL syntax structure.
[0182] a. In addition, in one example, sli_cbr_constraint_flag is also sub-layer specific, that is, it is changed to sli_cbr_constraint_flag[k] and moved into the loop of the sub-layer.
[0183] Furthermore, in one example, when the lower sublayer's sli_cbr_constraint_flag[k] does not exist, it is inferred to be equal to sli_cbr_constraint_flag[k+1].
[0184] b. Furthermore, in one example, when the score or reference level indicator of a lower sub-layer is not present, it is inferred to be the same as the next higher sub-layer.
[0185] 2) To address the second issue, SLI SEI messages are allowed to be available either within the bitstream or provided by external means, in order to be consistent with the parameter set and the other three consistency / HRD related SEI messages (i.e., PT, BP, and DUISEI messages).
[0186] 3) To address the third issue, the scope of existence is changed from one CVS to one or more CVSs to align with VPS and SPS, where level information is signaled or may be signaled.
[0187] 4) To address the fourth issue, the definition of the sub-image sequence was changed to cover the case where there are one or more layers where each image has a single sub-image.
[0188] 5) To address the fifth issue, it is required that when the SLI SEI message exists in CVS, the value of sps_num_subpics_minus1 should be the same for all SPS referenced by images in layers where each image has multiple subpicks.
[0189] 6) To address the sixth problem, the semantics of sli_num_subpics_minus1 are defined in such a way that the syntax elements are subpicks of a layer with multiple subpicks for each picture.
[0190] 7) To address the seventh issue, in the final constraint set of the SLI SEI message semantics, add array indices and subpic sequence indices to the variables SubpicLevelIdc and SubpicLevelIdx, as well as to the arrays SubpicCpbSizeVcl, SubpicCpbSizeNal, SubpicBitRateVcl, and SubpicBitRateNal.
[0191] 6. Example of an implementation plan
[0192] The following are some example embodiments of the inventions outlined above, which can be applied to the VVC specification. Most of the relevant additions or modifications are marked in bold, italics, and underlined, and some deleted parts are marked with [[]].
[0193] 6.1. First Embodiment
[0194] This example is used for projects 1 to 7 and some of their sub-projects.
[0195] D.7.1 Sub-picture level information SEI message syntax
[0196]
[0197] D.7.2 Sub-picture level information SEI message semantics
[0198] When testing the consistency of the extracted bitstream containing sub-image sequences according to Appendix A, sub-image level information... The message contains information about The information of the level that the sub-image sequences in the [[bitstream]] conform to.
[0199] [When a Subpicture Level Information (SEI) message exists in any picture within a CLVS, the SEI message should also exist in the first picture of the CLVS. SEI messages are stored in the current layer from the current picture in decoding order until the end of the CLVS. All SEI messages applied to the same CLVS should have the same content.]
[0200]
[0201] Sub-image sequence by Images with the same sub-image index value All sub-images composition.
[0202] The requirement for bitstream consistency is that, [[when the Sub-Picture Level Information (SEI) message exists in CLVS,]] For each value of i in the range of 0 to sps_num_subpics_minus1 (inclusive), the value of sps_subpic_treated_as_pic_flag[i] should be equal to 1.
[0203] Increasing 1 by 1 in sli_num_ref_levels_minus1 specifies the number of subpicks to be referenced (sps_num_subpics_minus1+1). The number of reference levels for each of the signaling notifications.
[0204] The `sli_cbr_constraint_flag` being equal to 0 specifies that any sub-picture extracted according to Clause C.7 should be decoded using the HRD of any CPB specification in the extracted sub-bitstream. The resulting sub-bitstreams are handled by a hypothetical stream scheduler (HSS) operating in intermittent bitrate mode. `sli_cbr_constraint_flag` equal to 1 specifies... HSS operates in constant bit rate (CBR) mode.
[0205] A value of 1 for sli_explicit_fraction_present_flag indicates that the syntax element sli_ref_level_fraction_minus1[i] exists. A value of 0 for sli_explicit_fraction_present_flag indicates that the syntax element sli_ref_level_fraction_minus1[i] does not exist.
[0206] Incrementing 1 by 1 specifies the subpics. The number of subpicks in the [[CLVS]] image. When present, the value of sli_num_subpics_minus1 should be equal to... The images in the text are referenced from The value of sps_num_subpics_minus1 in SPS.
[0207]
[0208] sli_alignment_zero_bit should be equal to 0.
[0209] Specify and The bitstream level limit associated with the layer where `sps_num_subpics_minus1` is equal to 0 in the `[[Bitstream]]`. Score. When vps_max_layers_minus1 equals 0 or when there are no layers in the bitstream, sps_num_subpics_minus1 equals 0. It should be equal to 0.
[0210] instruct As specified in Appendix A, each sub-image The conforming Level. In addition to the values specified in Appendix A, the bitstream should not contain... The value of . Other values are reserved for future use by ITU-T|ISO / IEC. Bitstream consistency requirements are... The value of should be equal to the value of general_level_idc of the bitstream, and for any value where i is greater than 0 and km is greater than i, The value should be less than or equal to
[0211] Add 1 to specify For a subpick in a layer where `sps_num_subpics_minus1` is greater than 0, the subpick index equal to `j`, [[as specified in Clause A.4.1 for the `j`th subpick]], is subject to the level constraint associated with `sli_ref_level_idc[i]`. Fraction.
[0212] The variable SubpicSizeY[j] is set to equal to (sps_subpic_width_minus1[j]+1)*CtbSizeY*(sps_subpic_height_minus1[j]+1)*CtbSizeY.
[0213] When it does not exist The value is inferred to be equal to
[0214] variable Set to equal to
[0215] variable Set to equal to
[0216] variable and The derivation is as follows:
[0217]
[0218]
[0219] MaxCPB is specified as in Clause A.4.2. It is deduced.
[0220] variable and The derivation is as follows:
[0221]
[0222]
[0223] The value of ValBR is derived as follows:
[0224] – In the corresponding HRD parameters in VPS or SPS When available, ValBR is set to equal to Where [[Htid is the sub-level index under consideration, and]]ScIdx is the scheduling index under consideration.
[0225] Otherwise, ValBR is set as specified in Clause A.4.2. The derived MaxBR.
[0226] Note 1: When extracting sub-images, the resulting bitstream has a bit rate greater than or equal to 1. and CpbSize (indicated or inferred in VPS, SPS) and greater than or equal to and The bit rate (indicated or inferred in VPS, SPS).
[0227] The requirement for bitstream consistency is that, Extract the j-th sub-image from the layers in the input bitstream of the extraction process where sps_num_subpics_minus1 is greater than 0, for each j within the range of 0 to sps_num_subpics_minus1 (inclusive). And the generated one that meets the condition that general_tier_flag is equal to 0 and (for 0 to 0) Within the range (including endpoints), level i) equals Each layer in the bitstream of a certain grade should comply with the following constraints as specified in Appendix C for each bitstream conformance test:
[0228] It should be less than or equal to MaxLumaPs, where MaxLumaPs is specified in Table A.1 for the level. specified.
[0229] The value should be less than or equal to Sqrt(MaxLumaPs*8).
[0230] The value should be less than or equal to Sqrt(MaxLumaPs*8).
[0231] The value of SubpicWidthInTiles[j] should be less than or equal to MaxTileCOLS, and the value of SubpicHeightInTiles[j] should be less than or equal to MaxTileRows, where MaxTileCOLS and MaxTileRows are specified in Table A.1 for each level. specified.
[0232] The value of SubpicWidthInTiles[j] * SubpicHeightInTiles[j] should be less than or equal to 1. MaxTileCOLS and MaxTileRows are listed in Table A.1 for different levels. specified.
[0233] The requirement for bitstream consistency is that, Extraction from 0 to The j-th sub-image within the range of j (including endpoints) And the generated one that meets the condition that general_tier_flag is equal to 0 and (for 0 to 0) Within the range (including endpoints), level i) equals The bitstreams of each grade should comply with the following constraints as specified in Appendix C for each bitstream conformance test:
[0234] Corresponding to the j-th sub-image The sum of the NumBytesInNalUnit variables for AU 0 should be less than or equal to the value of SubpicSizeInSamples for AU 0. Among them, MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3, respectively, and they are applied to the level. AU 0, and MinCr is derived as indicated in A.4.2.
[0235] Corresponding to the j-th sub-image The sum of the NumBytesInNalUnit variables of AU n (n greater than 0) should be less than or equal to MaxLumaSr and FormatCapabilityFactor are the values specified in Tables A.2 and A.3, respectively, and they are applied to the level. AU n, and MinCr are derived as indicated in A.4.2.
[0236] Sub-image sequence level indicator The value is derived as follows:
[0237]
[0238] Conforms to the condition where general_tier_flag equals 0 and level equals The level Sub-image [[bitstream]] The following constraints should be observed for each bitstream compliance test as specified in Appendix C:
[0239] For VCL HRD parameters, It should be less than or equal to CpbVclFactor*MaxCPB, where CpbVclFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in bits of CpbVclFactor.
[0240] For NAL HRD parameters It should be less than or equal to CpbNalFactor*MaxCPB, where CpbNalFactor is specified in Table A.3 and MaxCPB is specified in Table A.1 in bits of CpbNalFactor.
[0241] For VCL HRD parameters, It should be less than or equal to CpbVclFactor*MaxBR, where CpbVclFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in bits of CpbVclFactor.
[0242] For NAL HRD parameters It should be less than or equal to CpbNalFactor*MaxBR, where CpbNalFactor is specified in Table A.3 and MaxBR is specified in Table A.1 in bits of CpbNalFactor.
[0243] Note 2 When extracting Sub-image sequence The generated bitstream has a value greater than or equal to and CpbSize (indicated or inferred in VPS, SPS) and greater than or equal to and The bit rate (indicated or inferred in VPS, SPS).
[0244] Figure 5 This is a block diagram illustrating an example video processing system 1900, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0245] System 1900 may include codec component 1904, which may implement various codec or encoding methods described in this document. Codec component 1904 may reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. As indicated by component 1906, the output of codec component 1904 may be stored or transmitted via connected communication. Component 1908 may use the stored or communicated bitstream (or codec) representation of the video received at input 1902 to generate pixel values or displayable video sent to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and corresponding decoding tools or operations, the opposite of the encoding result, will be performed by the decoder.
[0246] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0247] Figure 6 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more methods described in this document. Memory (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0248] Figure 8 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.
[0249] like Figure 8 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0250] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0251] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems used to generate video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0252] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0253] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or it may be external to destination device 120, which is configured to interface with an external display device.
[0254] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Multi-Function Video Coding (VVM) standard, and other current and / or further standards.
[0255] Figure 9 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 8 The video encoder 114 in the system 100 shown.
[0256] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 9In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0257] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.
[0258] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0259] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 9 The examples are shown separately.
[0260] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0261] The mode selection unit 203 may, for example, select one of the coding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra-frame and inter-frame prediction (CIIP) modes, where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0262] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information from images other than those associated with the current video block and decoded samples from buffer 213.
[0263] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0264] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating a reference image in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0265] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in the reference images in list 1. Motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates a reference image in list 0 or list 1 containing the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0266] In some examples, the motion estimation unit 204 can output the complete set of motion information for the decoder's decoding processing.
[0267] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 can signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0268] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0269] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0270] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Combined Mode Signaling.
[0271] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0272] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0273] In other examples, there may be no residual data for the current video block, for example, in skip mode, and the residual generation unit 207 may not perform the subtraction operation.
[0274] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0275] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0276] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample from one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is then stored in buffer 213.
[0277] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.
[0278] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.
[0279] Figure 10 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 8 The video decoder 114 in the system 100 shown.
[0280] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 10 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0281] exist Figure 10 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform functions typically associated with the video encoder 200. Figure 9 The decoding process is the inverse of the encoding process described.
[0282] Entropy decoding unit 301 can retrieve encoded bitstreams. The encoded bitstreams may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-coded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. Motion compensation unit 302 can determine this information, for example, by performing AMVP and merging modes.
[0283] Motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The syntax elements can include identifiers of the interpolation filters to be used with sub-pixel precision.
[0284] The motion compensation unit 302 can use interpolation filters, such as those used by the video encoder 200 during the encoding of video blocks, to calculate interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate a prediction block.
[0285] The motion compensation unit 302 can use some syntax information to determine the size of the blocks of (multiple) frames and / or (multiple) stripes used to encode the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.
[0286] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0287] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.
[0288] The following is a list of preferred embodiments.
[0289] The first set of clauses illustrates example embodiments of the techniques discussed in the preceding section (e.g., item 1).
[0290] 1. A video processing method (e.g., Figure 7 The method 700 described in the text includes: performing (702) a conversion between a video and a video codec representation comprising one or more video sublayers, wherein the codec representation conforms to a format rule; wherein the format rule specifies a syntax structure that loops through multiple sublayers in the codec representation and one or more syntax fields indicating each sublayer included in the syntax structure, wherein the syntax structure includes information about the score and reference level indicator of signaling notification.
[0291] 2. According to the approach of Solution 1, the format rule stipulates that a particular fraction not explicitly included in the syntax structure is interpreted as having the same value as the next higher sub-level.
[0292] The following clauses illustrate example embodiments of the techniques discussed in the preceding sections (e.g., items 2, 5, 6).
[0293] 3. A video processing method, comprising: performing a conversion between a video comprising one or more sub-pictures and a video codec representation, wherein the conversion uses or generates supplementary enhancement information at the level of one or more sub-pictures.
[0294] 4. According to the method of Solution 3, the supplementary enhancement information is included in the codec representation.
[0295] 5. According to the method of Solution 3, the supplementary enhancement information is excluded from the codec representation, and the supplementary enhancement information is communicated between the encoding and decoding ends using a mechanism different from the codec representation.
[0296] 6. According to the method of Solution 4, wherein the encoding and decoding representation conforms to a format rule that specifies that the same value is signaled in each sequence parameter set, the sequence parameter set indicating the number of sub-pictures in a layer where each picture has multiple sub-pictures.
[0297] 7. The method according to any one of solutions 1 to 6, wherein the conversion includes encoding the video into a codec representation.
[0298] 8. The method according to any one of solutions 1 to 6, wherein the conversion includes decoding the encoding / decoding representation to generate pixel values of the video.
[0299] 9. A video decoding apparatus, comprising a processor configured to implement the method described in one or more of solutions 1 to 8.
[0300] 10. A video encoding apparatus comprising a processor configured to implement the method described in one or more of solutions 1 to 8.
[0301] 11. A computer program product having computer code stored thereon, which, when executed, causes a processor to implement the method described in any one of solutions 1 to 9.
[0302] 12. The methods, apparatus or systems described in this document.
[0303] Figure 13This is a flowchart of a method 1300 for processing video data according to one or more embodiments of the present technology. Method 1300 includes, at operation 1310, performing a conversion between video and a bitstream of video comprising one or more Output Layer Sets (OLS) according to a rule. The rule specifies that a Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of a subpicture sequence in a set of codec video sequences to which the SLI SEI message is applied. The syntax structure of the SLI SEI message includes (1) a first syntax element specifying the maximum number of sublayers of the subpicture sequence, (2) a second syntax element specifying whether the subpicture sequence level information exists in one or more sublayer representations, and (3) a loop of multiple sublayers, each sublayer associated with a score of bitstream level limits and a level indicator indicating the level to which each subpicture sequence conforms.
[0304] In some embodiments, the value of the first syntax element is in the range of 0 to 1 less than the maximum number of sublayers indicated in the video parameter set. In some embodiments, the second syntax element is inferred to be 0 in response to the absence of a second syntax element in the bitstream. In some embodiments, the score for the bitstream level limit associated with sublayer k is inferred to be equal to the score associated with sublayer k+1 in response to the absence of a level indicator associated with sublayer k. In some embodiments, the level indicator is inferred to be equal to the level indicator associated with sublayer k+1 in response to the absence of a level indicator associated with sublayer k.
[0305] In some embodiments, for each sublayer, the syntax structure further includes a third syntax element specifying a score for the bitstream level limit associated with the level indicator. In response to the absence of a third syntax element associated with sublayer k, the third syntax element is inferred to be equal to the syntax element associated with sublayer k+1. In some embodiments, the syntax structure further includes a fourth syntax element specifying the number of reference levels signaled for each sub-picture sequence. In some embodiments, the syntax structure further includes a fifth syntax element specifying whether the hypothetical stream scheduler (HSS) operates in intermittent bit rate mode or constant bit rate (CBR) mode for the sub-picture sequence.
[0306] Figure 14This is a flowchart of a method 1400 for processing video data according to one or more embodiments of the present technology. Method 1400 includes, at operation 1410, performing a conversion between a current access unit of video including one or more Output Layer Sets (OLS) and the video bitstream according to a rule. The rule specifies that Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) messages include level information about the subpicture sequence in the set of codec video sequences to which the SLI SEI message is applied. The SLI SEI message persists from the current access unit in decoding order until the end of the bitstream, or until the next access unit contains a subsequent SLI SEI message including content different from the SLI SEI message.
[0307] In some embodiments, the rule applies to all SLI SEI messages of the same CVS having the same content. In some embodiments, the SLI SEI message exists in the current access unit either within the bitstream or provided externally. In some embodiments, a first variable indicating the sub-picture level indicator is specified to include the value for each sub-picture sequence. In some embodiments, a second variable indicating the sub-picture level index is specified to include the value for each sub-picture sequence.
[0308] Figure 15 This is a flowchart of a method 1500 for processing video data according to one or more embodiments of the present technology. Method 1500 includes, at operation 1510, performing a conversion between a currently accessed unit of video comprising one or more Output Layer Sets (OLS) and a bitstream of video, according to a rule. A Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of a subpicture sequence in a set of codec video sequences of one or more OLS to which the SLI SEI message is applied. A layer in one or more OLS whose number of subpictures is greater than 1, as indicated by variables in its reference sequence parameter set, is called a multi-subpicture layer. The codec video sequences in the OLS set are called target codec video sequences (CVS). The rule specifies that a subpicture sequence includes (1) all subpictures within the target CVS that have the same subpicture index and belong to a layer in a multi-subpicture layer, and (2) all subpictures in the target CVS that have a subpicture index of 0 and belong to a layer in the OLS but are not in a multi-subpicture layer.
[0309] In some embodiments, the bitstream conforms to a format rule that specifies that, in response to the presence of an SLI SEI message in a codec video sequence, all sequence parameter sets referenced by the pictures in a multi-subpicture layer have the same number of subpictures. In some embodiments, in response to the presence of an SLI SEI message in any access unit of a codec video sequence (CVS) of one or more OLSs, the rule specifies that the SLI SEI message is present in the first access unit of the CVS. In some embodiments, syntax elements in the syntax structure of the SLI SEI message specify the number of subpictures in the pictures of a multi-subpicture layer in the target CVS.
[0310] In some embodiments, the conversion includes encoding the video into a bitstream. In some embodiments, the conversion includes decoding the video from the bitstream.
[0311] In the terms described herein, an encoder can conform to a format rule by generating a codec representation according to the format rule. In the terms described herein, a decoder can use the format rule to parse the syntax elements into a codec representation based on the presence or absence of the syntax elements known according to the format rule, in order to produce a decoded video.
[0312] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. As defined in the syntax, the bitstream representation of the current video block can, for example, correspond to bits that are co-located or scattered at different locations within the bitstream. For example, a macroblock can be encoded based on the error residual values of the transformation and encoding, and also using bits in the header and other fields of the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether certain syntax fields are included or excluded, and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0313] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a combination of a machine-readable storage device, a machine-readable storage substrate, a memory device, a substance that implements a machine-readable propagating signal, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagating signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0314] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suited to a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as portions of multiple collaborative files (e.g., a file storing portions of one or more modules, subroutines, or code). Computer programs can be deployed to execute on a single computer or on multiple computers located in one place or distributed across multiple locations and interconnected via a communication network.
[0315] The processes and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and the devices can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0316] For example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, to receive data from or transfer data to, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented or incorporated therein by dedicated logic circuitry.
[0317] While this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or claimed content, but rather as descriptions of features characteristic of specific embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations, and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.
[0318] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0319] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
Claims
1. A method for processing video data, comprising: Perform conversions between video and video bitstreams that include one or more output layer sets (OLS) according to rules. The rule specifies that the Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of subpicture sequences in the set of encoded and decoded video sequences of one or more OLSs to which the SLI SEI message is applied. The syntax structure of the SLI SEI message includes (1) a first syntax element specifying the maximum number of temporal sublayers in the subpicture sequence, (2) a second syntax element specifying whether the level information of the subpicture sequence exists in the representation of the one or more sublayers, and (3) a loop of multiple sublayers, each sublayer associated with a score for bitstream level limits and a level indicator indicating the level to which each subpicture sequence conforms. Wherein, the second syntax element equal to 1 indicates that for one or more sub-layer representations in the range of 0 to the value of the first syntax element including the endpoints, there exists level information for the sub-image sequence, and the second syntax element equal to 0 indicates that for the p-th sub-layer representation, there exists level information for the sub-image sequence, where p is the value of the first syntax element.
2. The method of claim 1, wherein, In response to the fact that the second syntax element does not exist in the bitstream, the second syntax element is inferred to be 0.
3. The method according to claim 1, wherein, In response to the absence of a score for the bitstream level limit associated with sublayer k and k being less than the value of the first syntax element, the score is inferred to be equal to the score associated with sublayer k+1.
4. The method according to claim 1, wherein, In response to the absence of a level indicator associated with sublevel k and k being less than the value of the first syntax element, the level indicator is inferred to be equal to the level indicator associated with sublevel k+1.
5. The method according to claim 1, wherein, For each sub-layer, the syntax structure further includes a third syntax element that specifies a score for the bitstream level limit associated with the level indicator, and wherein, in response to the absence of a third syntax element associated with sub-layer k and k being less than the value of the first syntax element, the third syntax element is inferred to be equal to the syntax element associated with sub-layer k+1.
6. The method according to claim 1, wherein, The syntax structure also includes a fourth syntax element, which specifies the number of reference levels for signaling notification for each sub-picture sequence.
7. The method according to claim 1, wherein, The syntax structure also includes a fifth syntax element, which specifies whether the hypothetical stream scheduler HSS operates in intermittent bit rate mode or constant bit rate (CBR) mode for a sub-picture sequence.
8. The method according to any one of claims 1 to 7, wherein, The conversion includes encoding the video into the bitstream.
9. The method according to any one of claims 1 to 7, wherein, The conversion includes decoding the video from the bitstream.
10. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: Perform conversions between video and video bitstreams that include one or more output layer sets (OLS) according to rules. The rule specifies that the Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of subpicture sequences in the set of encoded and decoded video sequences of one or more OLSs to which the SLI SEI message is applied. The syntax structure of the SLI SEI message includes (1) a first syntax element specifying the maximum number of temporal sublayers in the subpicture sequence, (2) a second syntax element specifying whether the level information of the subpicture sequence exists in the representation of the one or more sublayers, and (3) a loop of multiple sublayers, each sublayer associated with a score for bitstream level limits and a level indicator indicating the level to which each subpicture sequence conforms. Wherein, the second syntax element equal to 1 indicates that for one or more sub-layer representations in the range of 0 to the value of the first syntax element including the endpoints, there exists level information for the sub-image sequence, and the second syntax element equal to 0 indicates that for the p-th sub-layer representation, there exists level information for the sub-image sequence, where p is the value of the first syntax element.
11. The apparatus according to claim 10, wherein, In response to the fact that the second syntax element does not exist in the bitstream, the second syntax element is inferred to be 0.
12. The apparatus according to claim 10, wherein, In response to the absence of a score for the bitstream level limit associated with sublayer k and k being less than the value of the first syntax element, the score is inferred to be equal to the score associated with sublayer k+1. In response to the absence of a level indicator associated with sub-level k and k being less than the value of the first syntax element, the level indicator is inferred to be equal to the level indicator associated with sub-level k+1.
13. The apparatus according to claim 10, wherein, For each sub-layer, the syntax structure further includes a third syntax element that specifies a score for the bitstream level limit associated with the level indicator, and wherein, in response to the absence of a third syntax element associated with sub-layer k and k being less than the value of the first syntax element, the third syntax element is inferred to be equal to the syntax element associated with sub-layer k+1.
14. The apparatus according to claim 10, wherein, The syntax structure also includes a fourth syntax element, which specifies the number of reference levels for signaling notification for each sub-picture sequence.
15. The apparatus according to claim 10, wherein, The syntax structure also includes a fifth syntax element, which specifies whether the hypothetical stream scheduler HSS operates in intermittent bit rate mode or constant bit rate (CBR) mode for a sub-picture sequence.
16. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: Perform conversions between video and video bitstreams that include one or more output layer sets (OLS) according to rules. in, The rule specifies that the Subpicture Level Information (SLI) Supplemental Enhancement Information (SEI) message includes information about the level of subpicture sequences in a set of encoded and decoded video sequences for one or more OLSs to which the SLI SEI message is applied, and wherein the syntax structure of the SLI SEI message includes (1) a first syntax element specifying the maximum number of temporal sublayers in the subpicture sequence, (2) a second syntax element specifying whether the level information of the subpicture sequence exists in the representation of the one or more sublayers, and (3) a loop of multiple sublayers, each sublayer being associated with a score of bitstream level limits and a level indicator indicating the level to which each subpicture sequence conforms. Wherein, the second syntax element equal to 1 indicates that for one or more sub-layer representations in the range of 0 to the value of the first syntax element including the endpoints, there exists level information for the sub-image sequence, and the second syntax element equal to 0 indicates that for the p-th sub-layer representation, there exists level information for the sub-image sequence, where p is the value of the first syntax element.
17. The non-transitory computer-readable storage medium according to claim 16, in, In response to the second syntax element not being present in the bitstream, the second syntax element is inferred to be 0. Wherein, in response to the absence of a score for the bitstream level constraint associated with sublayer k and k being less than the value of the first syntax element, the score is inferred to be equal to the score associated with sublayer k+1. Wherein, in response to the absence of a level indicator associated with sublevel k and k being less than the value of the first syntax element, the level indicator is inferred to be equal to the level indicator associated with sublevel k+1. For each sub-layer, the syntax structure further includes a third syntax element that specifies a score for the bitstream level limit associated with the level indicator, and wherein, in response to the absence of a third syntax element associated with sub-layer k and k being less than the value of the first syntax element, the third syntax element is inferred to be equal to the syntax element associated with sub-layer k+1. The syntax structure further includes a fourth syntax element, which specifies the number of reference levels for signaling notification for each sub-image sequence. The syntax structure also includes a fifth syntax element, which specifies whether the hypothetical stream scheduler HSS operates in intermittent bit rate mode or constant bit rate (CBR) mode for a sub-picture sequence.
18. A non-transitory computer-readable recording medium having a computer program and a bit stream stored thereon, wherein, When the computer program is executed by the video processing device, it generates the bitstream by implementing the method described in any one of claims 1-8.
19. A method for storing a video bitstream, comprising: The bit stream is generated by performing the method according to any one of claims 1-8; as well as The bitstream is stored in a non-transitory computer-readable recording medium.
20. A video decoding apparatus comprising a processor configured to implement the method of any one of claims 1 to 7, 9.
21. A video encoding apparatus comprising a processor configured to implement the method of any one of claims 1 to 8.
22. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 9.