Decoded picture buffer management and sub-pictures in video codecs

By improving DPB parameter signaling and sub-picture sequence-level information processing, the problems of insufficient DPB parameters and incorrect POC derivation in multi-layer video encoding and decoding are solved, and the efficiency, accuracy, adaptability and error elasticity of video encoding and decoding are improved.

CN115769571BActive Publication Date: 2025-08-15DOUYIN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180043484.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-17
Filing Date
2021-06-15
Publication Date
2025-08-15
Estimated Expiration
2041-06-15

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies process multi-layer video and sub-pictures, there are problems such as insufficient signaling notification of DPB parameters, incomplete sub-picture sequence-level information processing, inaccurate information removal during sub-bitstream extraction, and incorrect POC derivation, resulting in a decrease in decoding efficiency and accuracy.

Method used

By improving the signaling notification mechanism of DPB parameters, the processing of sub-picture sequence-level information is optimized, the non-scalable nested SEI messages are accurately removed, and the image sequence count is correctly derived during the sub-bitstream extraction process, ensuring the effectiveness and efficiency of multi-layer video encoding and decoding.

Benefits of technology

It improves the decoding efficiency and accuracy of multi-layer video encoding and decoding, reduces redundant information transmission, and enhances the error elasticity and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769571B_ABST
    Figure CN115769571B_ABST
Patent Text Reader

Abstract

Several techniques for video encoding and video decoding are described. An exemplary method includes performing conversion between a video and a bitstream of the video, wherein the bitstream includes one or more pictures, the one or more pictures including one or more sub-pictures, according to a rule, and wherein the rule specifies performing an overwriting operation on one or more sequence parameter sets referenced during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream in response to a condition.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is a Chinese national phase application of International Application No. PCT / US2021 / 037371, filed on June 15, 2021, and promptly claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 040,427, filed on June 17, 2020. The entire disclosure of the aforementioned application is incorporated herein by reference as a part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth used by digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of videos or images.

[0006] In one exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, wherein the bitstream includes one or more pictures, the one or more pictures including one or more sub-pictures, according to a rule, and wherein the rule specifies that, in response to a condition, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced thereto.

[0007] In another exemplary aspect, a video processing method is disclosed, comprising performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream comprises one or more output layer sets comprising a plurality of layers according to a format rule, wherein the format rule provides for indicating in the bitstream a single syntax structure applicable to each of the one or more output layer sets, and wherein the single syntax structure comprises decoded picture cache parameters.

[0008] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video, wherein the bitstream includes one or more pictures, the one or more pictures including one or more sub-pictures, according to a rule, and wherein the rule specifies that during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, in response to removing a sub-picture from the bitstream, all supplemental enhancement information network abstraction layer units including non-scalable nested supplemental enhancement information messages of a specific payload type are also removed.

[0009] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video and a bitstream of the video according to a rule, wherein the bitstream includes a plurality of access units, the plurality of access units including one or more pictures, and wherein the rule provides for removing a current picture of a current access unit from a decoded picture cache in response to the following conditions: (1) the current picture is not the first picture of the current access unit, and (2) the current access unit is the first access unit in decoding order in the bitstream.

[0010] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream includes a first syntax element having a value indicating the presence of a second syntax element in a picture header, the second syntax element specifying a value of a most significant bit of a picture sequence count of a current picture, wherein the bitstream conforms to a format rule responsive to the following conditions, the format rule specifying that the value of the first syntax element is equal to 1: (1) the current picture associated with the picture header is an intra random access point picture or a progressive decoding refresh picture, the current picture has an associated third syntax element equal to 0, which indicates that the current picture is a recovery point picture, (2) the picture order count difference between the current picture and a previous intra random access point picture or a previous progressive decoding refresh picture in the same layer in decoding order whose associated third syntax element is equal to 0 is equal to or greater than a variable indicating the maximum picture order count least significant bit divided by 2, and (3) the value of NumSubLayersInLayerInOLS[i][GeneralLayerIdx[nuh_layer_id]] is equal to 0 for any value of i in the range of 0 to TotalNumOlss-1, inclusive , and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, where NumSubLayersInLayerInOLS[i][GeneralLayerIdx[nuh_layer_id]] indicates the number of sublayers in the GeneralLayerIdx[nuh_layer_id]th output layer in the i-th output layer set, where TotalNumOlss specifies the total number of output layer sets specified by the video parameter set, where vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1 specifies that the layer with layer index equal to GeneralLayerIdx[nuh_layer_id] does not use inter-layer prediction, where GeneralLayerIdx[nuh_layer_id] specifies the layer index with layer identifier equal to nuh_layer_id, and where nuh_layer_id specifies the identifier of the layer to which the video codec layer network abstraction layer unit belongs or the identifier of the layer to which the non-video codec layer network abstraction layer unit applies.

[0011] In another exemplary aspect, a video processing method is disclosed. The method includes performing conversion between a video comprising one or more video layers and a codec representation of the video, the one or more video layers comprising one or more color components, wherein the codec representation conforms to a format rule that specifies that the codec representation includes a single syntax structure indicating decoded picture cache parameters applicable to each layer of a multi-layer output layer set of the codec representation.

[0012] In another exemplary aspect, another video processing method is disclosed, comprising: performing conversion between a video comprising one or more video layers and a codec representation of the video, the one or more video layers comprising one or more color components; wherein the codec representation complies with a rule related to processing the codec representation based on removing pictures or sub-pictures from a buffer during a decoding process.

[0013] In yet another exemplary aspect, a video encoder apparatus is disclosed, wherein the video encoder includes a processor configured to implement the above method.

[0014] In yet another exemplary aspect, a video decoder apparatus is disclosed, wherein the video decoder includes a processor configured to implement the above method.

[0015] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed, wherein the code is in the form of processor-executable code to implement one of the methods described herein.

[0016] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is partitioned into 12 slices and 3 raster scan strips.

[0018] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is partitioned into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0019] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is partitioned into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0020] Figure 4 A picture is shown partitioned into 15 slices, 24 slices, and 24 sub-pictures.

[0021] Figure 5 is a block diagram of an exemplary video processing system.

[0022] Figure 6 It is a block diagram of a video processing device.

[0023] Figure 7 is a flow chart of an exemplary method of video processing.

[0024] Figure 8 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0025] Figure 9 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0026] Figure 10 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0027] Figure 11 An example of a typical sub-picture-based viewport-dependent 360° video encoding and decoding scheme is shown.

[0028] Figure 12 A viewport-dependent 360° video encoding and decoding scheme based on sub-picture and spatial scalability is shown.

[0029] Figures 13 to 17 is a flow chart of an exemplary method of video processing. DETAILED DESCRIPTION

[0030] The section headings used in this document are for ease of understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is only for ease of understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein also apply to other video codec protocols and designs. In this document, editorial changes to text relative to the current draft of the VVC specification are indicated by deleting text to indicate cancellations and by highlighting text (including bold italics) to indicate additions.

[0031] 1. Introduction

[0032] This document relates to video codec technology. Specifically, it is about specifying and signaling level information for sub-picture sequences. It is applicable to any video codec standard or non-standard video codec that supports single-layer video codec and multi-layer video codec (e.g., the Versatile Video Codec (VVC) under development)

[0033] 2. Abbreviation

[0034] APS Adaptation Parameter Set

[0035] AU Access Unit

[0036] AUD Access Unit Delimiter

[0037] AVC Advanced Video Codec

[0038] BP buffer period

[0039] Layer video sequence for CLVS codec

[0040] CLVSS codec layer video sequence starts

[0041] CPB codec image cache

[0042] CRA Clean Random Access

[0043] CTU Codec Tree Unit

[0044] CVS encoded and decoded video sequence

[0045] DCI decoding capability information

[0046] DPB decoded image cache

[0047] DUI decoding unit information

[0048] EOB End of bitstream

[0049] EOS sequence end

[0050] GDR Progressive Decode Refresh

[0051] HEVC High-Efficiency Video Codec

[0052] HRD Hypothesized Reference Decoder

[0053] IDR Instant Decode Refresh

[0054] ILP inter-layer prediction

[0055] ILRP interlayer reference image

[0056] JEM Joint Exploration Model

[0057] LTRP Long Term Reference Picture

[0058] MCTS motion constraint set

[0059] NAL Network Abstraction Layer

[0060] OLS output layer set

[0061] PH Image Header

[0062] POC picture sequence counting

[0063] PPS picture parameter set

[0064] PT Picture Timing

[0065] PTL Profiles, Hierarchies, and Levels

[0066] PU picture unit

[0067] RAP Random Access Point

[0068] RBSP Raw Byte Sequence Payload

[0069] SEI Supplemental Enhancement Information

[0070] SLI sub-picture level information

[0071] SPS sequence parameter set

[0072] STRP Short-Term Reference Picture

[0073] AVC Scalable Video Codec

[0074] VCL video codec layer

[0075] VPS Video Parameter Set

[0076] VTM VVC test model

[0077] VUI video usage information

[0078] VVC Universal Video Codec

[0079] 3. Initial Discussion

[0080] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Video, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards are based on a hybrid video codec structure that utilizes temporal prediction plus transform coding. In order to explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by the Video Coding Experts Group (VCEG) and MPEG. Since then, JVET has adopted many new methods and put them into reference software called the Joint Exploration Model (JEM). JVET meets at the same time every quarter, and the goal for the new codec standard is a 50% reduction in bit rate compared to HEVC. At the JVET meeting in April 2018, the new video codec standard was officially named Versatile Video Codec (VVC), and the first version of the VVC Test Model (VTM) was released at this time. Due to the ongoing efforts to contribute to VVC standardization, new codec technologies are being used in the VVC standard at each JVET meeting. The VVC working draft and test model VTM are then updated after each meeting. The VVC project is now aiming for technical completion (FDIS) at the July 2020 meeting.

[0081] 3.1. Random Access and Support in HEVC and VVC

[0082] Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture in the bitstream in decoding order. In order to support tuning and channel switching in broadcast / multicast and multi-party video conferencing, seeking in local playback and streaming, and stream adaptation in streaming, the bitstream needs to include frequent random access points, which are usually intra-coded pictures, but can also be inter-coded pictures (for example, in the case of progressive decoding refresh).

[0083] HEVC includes signaling of internal random access point (IRAP) pictures via NAL unit types in the NAL unit header. Three types of IRAP pictures are supported, namely instantaneous decoder refresh (IDR), clean random access (CRA) and broken link access (BLA) pictures. IDR pictures restrict the inter-picture prediction structure so as not to reference any pictures before the current group of pictures (GOP) and are commonly referred to as closed GOP random access points. CRA pictures are less restricted by allowing certain pictures to reference pictures before the current GOP, all of which are discarded in the case of random access. CRA pictures are commonly referred to as open GOP random access points. BLA pictures typically result from the concatenation of two bitstreams or parts thereof at a CRA picture, for example during stream switching. To enable better system usage of IRAP pictures, a total of six different NAL units are defined to signal properties of IRAP pictures that can be used to better match stream access point types as defined in the ISO Base Media File Format (ISOBMFF) for random access support in Dynamic Adaptive Streaming over HTTP (DASH).

[0084] VVC supports three types of IRAP pictures, two types of IDR pictures (one type with an associated RADL picture or the other type without an associated RADL picture), and one type of CRA picture. These are essentially the same as in HEVC. The BLA picture type in HEVC is not included in VVC, primarily for two reasons: i) the basic functionality of BLA pictures can be implemented with CRA pictures plus the end of a sequence NAL unit, the presence of which indicates that the subsequent picture starts a new CVS in a single-layer bitstream; ii) during the development of VVC, it was desirable to specify fewer NAL unit types than in HEVC, as indicated by using five bits instead of six for the NAL unit type field in the NAL unit header.

[0085] Another key difference in random access support between VVC and HEVC is that GDR is supported in VVC in a more standardized way. In GDR, decoding of the bitstream can start from an inter-frame coded picture, and although not the entire picture area can be decoded correctly at the beginning, after multiple pictures, the entire picture area will be correct. AVC and HEVC also support GDR, using the recovery point SEI message for signaling of GDR random access points and recovery points. In VVC, a new NAL unit type is specified for the indication of GDR pictures, and the recovery point is signaled in the picture header syntax structure. CVS and bitstreams are allowed to start with GDR pictures. This means that the entire bitstream is allowed to include only inter-frame coded pictures and not a single intra-frame coded picture. The main benefit of specifying GDR support in this way is to provide consistent behavior for GDR. GDR enables the encoder to smooth the bitrate of the bitstream by distributing intra-coded slices or blocks across multiple pictures instead of intra-coding the entire picture, allowing for significant end-to-end latency reduction, which is considered more important today than ever before as ultra-low latency applications such as those based on wireless displays, online gaming, and drones become more popular.

[0086] Another GDR-related feature in VVC is virtual boundary signaling. The boundary between the refreshed area (i.e., correctly decoded area) and the non-refreshed area at the picture between the GDR picture and its recovery point can be signaled as a virtual boundary, and when signaled, in-loop filtering across the boundary will not be applied, so decoding mismatches of some samples at or near the boundary will not occur. This may be useful when the application determines during the GDR process that the correctly decoded area should be displayed.

[0087] IRAP pictures and GDR pictures may be collectively referred to as random access point (RAP) pictures.

[0088] 3.2. Image Segmentation Scheme in HEVC

[0089] HEVC includes four different picture partitioning schemes, namely regular strips, correlated strips, slices, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end delay.

[0090] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although they may still have dependencies due to loop filtering operations).

[0091] Regular strips are the only tool available for parallelization that is also available in H.264 / AVC in essentially the same form. Parallelization based on regular strips does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is generally much heavier than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, the use of regular strips can incur a large amount of codec overhead due to the bit cost of the slice header and due to the lack of prediction across slice boundaries. In addition, due to the intra-picture independence of regular strips and the fact that each regular strip is encapsulated in its own NAL unit, regular strips (compared to the other tools mentioned below) are also used as a key mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching place conflicting demands on the stripe layout in the picture. The realization of this situation led to the development of the parallelization tools mentioned below.

[0092] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without disrupting any intra-picture prediction. Essentially, dependent slices provide segmentation of a regular slice into multiple NAL units to provide reduced end-to-end latency by allowing a portion of a regular slice to be emitted before coding of the entire regular slice is complete.

[0093] In WPP, a picture is partitioned into a single row of codec treeblocks (CTBs). This allows entropy decoding and prediction to use data from CTBs in other partitions. Parallel processing is possible by decoding CTB rows in parallel, where the start of decoding a CTB row is delayed by two CTBs to ensure that data associated with the CTBs above and to the right of the target CTB is available before the target CTB is decoded. This staggered start (which, when represented graphically, looks like a wavefront) allows parallelization using as many processors / cores as the picture contains CTB rows. Because intra-picture prediction is permitted between adjacent treeblock rows within a picture, the required inter-processor / inter-core communication to implement intra-picture prediction can be substantial. WPP partitioning does not result in the generation of additional NAL units compared to when WPP is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, with some codec overhead.

[0094] Slices define the horizontal and vertical boundaries that divide a picture into slice columns and rows. Slice columns extend from the top to the bottom of the picture. Similarly, slice rows extend from the left to the right. The number of slices in a picture can be simply derived as the number of slice columns multiplied by the number of slice rows.

[0095] Before decoding the top left CTB of the next slice in the order of raster scan of the slices of the picture, the scan order of the CTBs is changed to be local within the slice (in the order of raster scan of the CTBs of the slice). Similar to regular slices, slices break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included in separate NAL units (the same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and in the case where a slice spans more than one slice, the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to transmitting a shared slice header, and loop filtering involves sharing of reconstructed samples and metadata. When multiple slices or WPP segments are included in a slice, the entry point byte offset of each slice or WPP segment except the first slice or WPP segment in the slice is signaled in the slice header.

[0096] For simplicity, HEVC specifies restrictions on the application of four different picture partitioning schemes. A video sequence for a given codec cannot include both slices and wavefronts for most profiles specified in HEVC. For each slice and slice, either or both of the following conditions must be met: 1) all codec treeblocks in a slice belong to the same slice; 2) all codec treeblocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is in use, if a slice starts within a CTB row, it must end in the same CTB row.

[0097] Recent modifications to HEVC are specified in the following: JCTVC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), "HEVC Additional Supplementary Enhancement Information (Draft 4)," October 24, 2017, publicly available at: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this modification, HEVC specifies three MCTS-related SEI messages: a temporal MCTS SEI message, an MCTS extraction information set SEI message, and an MCTS extraction information nesting SEI message.

[0098] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals the MCTS. For each MCTS, motion vectors are restricted to pointing to full sample positions within the MCTS and fractional sample positions that only require interpolation from full sample positions within the MCTS, and motion vector candidates derived from blocks outside the MCTS are not allowed to be used for temporal motion vector prediction. In this way, each MCTS can be decoded independently without the presence of slices not included in the MCTS.

[0099] The MCTS Extraction Information Set SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to produce a conforming bitstream for the MCTS set. This information consists of multiple extraction information sets, each of which defines multiple MCTS sets and contains RBS bytes that replace the VPS, SPS, and PPS to be used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the slice address-related syntax elements (including first_slice_segment_in_pic_flag and slice_segment_address) will typically need to have different values.

[0100] 3.3. Image Segmentation in VVC

[0101] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the picture. The CTUs in a slice are scanned in raster scan order within the slice.

[0102] A slice consists of an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of a picture.

[0103] Two modes of slices are supported: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice consists of a sequence of complete slices in a slice raster scan of a picture. In rectangular slice mode, a slice consists of several complete slices that together form a rectangular region of the picture, or several consecutive complete CTU rows that together form a slice of a rectangular region of the picture. Within the rectangular region corresponding to the rectangular slice, the slices within the rectangular strip are scanned in slice raster scan order.

[0104] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.

[0105] Figure 1 An example of raster scan striping partitioning of a picture is shown, where the picture is partitioned into 12 slices and 3 raster scan strips.

[0106] Figure 2 An example of rectangular strip partitioning of a picture is shown, where the picture is partitioned into 24 slices (6 slice columns and 4 slice rows) and 9 rectangular strips.

[0107] Figure 3 An example of a picture partitioned into slices and rectangular strips is shown, where the picture is partitioned into 4 slices (2 slice columns and 2 slice rows) and 4 rectangular strips.

[0108] Figure 4 An example of sub-picture partitioning of a picture is shown, where the picture is partitioned into 18 slices, with the 12 slices on the left-hand side each covering one strip of 4 by 4 CTUs, and the 6 slices on the right-hand side each covering 2 vertically stacked strips of 2 by 2 CTUs, resulting in a total of 24 slices and 24 sub-pictures of varying sizes (each strip being a sub-picture).

[0109] 3.4. Image resolution changes within a sequence

[0110] In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC enables changing the resolution of pictures within a sequence without encoding an IRAP picture, which is always intra-coded. This feature is sometimes called reference picture resampling (RPR) because it requires resampling the reference picture used for inter prediction when it has a different resolution than the current picture being decoded.

[0111] The scaling ratio is restricted to be greater than or equal to 1 / 2 (downsampling by a factor of 2 from the reference picture to the current picture) and less than or equal to 8 (upsampling by a factor of 8). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied to scaling ratios ranging from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as the case of motion compensated interpolation filters. In fact, the standard MC interpolation process is a special case of the resampling process for scaling ratios in the range of 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the left, right, top, and bottom scaling offsets specified for the reference picture and the current picture.

[0112] Other aspects of the VVC design that differ from HEVC in supporting this feature include: i) the picture resolution and corresponding consistency window are signaled in the PPS rather than in the SPS, where the maximum picture resolution is signaled; ii) for a single-layer bitstream, each picture store (a slot in the DPB used to store one decoded picture) occupies a buffer size required to store a decoded picture with the maximum picture resolution.

[0113] 3.5. General Scalable Video Codec (SVC) and SVC in VVC

[0114] Scalable Video Codec (SVC, sometimes also referred to simply as scalability in video codecs) refers to video codecs that use a base layer (BL) (sometimes referred to as a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data at a base quality level. The one or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previously coded layers. For example, the bottom layer can serve as the BL, while the top layer can serve as the EL. Intermediate layers can serve as the EL, the RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can serve as the EL for layers below the intermediate layer (such as the base layer or any intervening enhancement layers) while simultaneously serving as the RL for one or more enhancement layers above the intermediate layer. Similarly, in multi-view or 3D scaling in the HEVC standard, multiple views can be present, and information from one view can be used to encode (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).

[0115] In SVC, parameters used by an encoder or decoder are grouped into parameter sets based on the codec level at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be utilized by one or more codecs of different layers in a bitstream can be included in a video parameter set (VPS), and parameters that can be utilized by one or more pictures in a codec of a video sequence can be included in a sequence parameter set (SPS). Similarly, parameters that can be utilized by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters that are specific to individual slices can be included in a slice header. Similarly, indications of which parameter sets a particular layer is using at a given time can be provided at various codec levels.

[0116] Due to the support of reference picture resampling (RPR) in VVC, support for bitstreams containing multiple layers (for example, two layers with SD and HD resolutions in VVC) can be designed without the need for any additional signal processing level codec tools, because the upsampling required for spatial scalability support can use only RPR upsampling filters. However, high-level syntax changes (compared to not supporting scalability) are required for scalability support. Scalability support is specified in VVC version 1. Unlike scalability support in any earlier video codec standards including AVC and HEVC scaling, the design of VVC scalability has been made as friendly to single-layer decoder design as possible. The decoding capabilities of multi-layer bitstreams are specified as if there is only a single layer in the bitstream. For example, decoding capabilities such as DPB size are specified in a manner independent of the number of layers in the bitstream to be decoded. Basically, a decoder designed for a single-layer bitstream does not require too many changes to be able to decode multi-layer bitstreams. Compared to the design of multi-layer scaling of AVC and HEVC, the HLS aspect is significantly simplified at the expense of some flexibility. For example, IRAPAU requires images for every layer present in the CVS.

[0117] Parameter Set

[0118] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced from HEVC and is included in HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.

[0119] The SPS is designed to carry sequence-level header information, while the PPS is designed to carry infrequently changing picture-level header information. Using SPS and PPS eliminates the need to repeat infrequently changing information for every sequence or picture, thus avoiding redundant signaling of this information. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, eliminating the need for redundant transmission and improving error resilience.

[0120] VPS is introduced to carry sequence level header information that is common to all layers in a multi-layer bitstream.

[0121] APS is introduced to carry such picture-level or slice-level information, which requires quite a lot of bits to encode or decode, can be shared by multiple pictures, and can have quite a lot of different variations in a sequence.

[0122] 4. Technical problems solved by the disclosed technical solutions

[0123] The latest design of signaling of DPB parameters, removal of decoded pictures from DPB, signaling of level information of sub-picture sequence in SLI SEI message, processing of SEI messages in sub-picture extraction, and signaling of POC derivation in VVC has the following problems (in the description involving modifications to some existing text, additions are indicated by bold italic text, and deletions are marked by opening and closing double brackets (e.g., [[ ]]), where the deleted text is between the double brackets):

[0124] 1) The DPB parameter syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_minus8[i] are signaled in the VPS of each multi-layer OLS. However, in many cases, these parameters have the same values for multiple OLSs, while the DPB parameter syntax elements max_dec_pic_buffering_minus1[i], max_num_reorder_pics[i], and max_latency_increase_plus1[i] are signaled in the dpb_parameters() syntax structure. Each syntax element can be shared by two or more OLSs.

[0125] 2) JVET-S0099 proposes the following changes to the sub-picture sub-bitstream extraction processing: when outBitstream contains a SEI NAL unit (the SEI NAL unit contains a scalable nesting SEI message with sn_nesting_ols_flag equal to 1 and sn_nesting_subpic_flag equal to 1, and the SEI NAL unit is applicable to outBitstream), remove all SEI NAL units containing non-scalable nesting SEI messages with a payload type equal to 1 (picture timing), 130 (decoding unit information), or 132 (decoded picture hash).

[0126] However, this approach has the following problems:

[0127] a. The condition for applying such removal should be when there are some sub-pictures removed by an earlier step, rather than when outBitstream contains a SEINAL unit containing a scalable nesting SEI message with sn_nesting_ols_flag equal to 1 and sn_nesting_subpic_flag equal to 1 applicable to outBitstream.

[0128] b. SEI NAL units containing some other non-scalable nesting SEI messages that are no longer applicable to the output sub-bitstream should also be removed. These other non-scalable nesting SEI messages include BP SEI messages and SLI SEI messages.

[0129] 3) In the sub-picture sub-bitstream extraction process, when no explicit sub-picture ID mapping is signaled in the input bitstream, although a sequence of sub-pictures identified by sub-picture indices greater than 0 is extracted as the output bitstream, the sub-picture ID mapping (although there is only one sub-picture in each picture in the output bitstream) needs to be included in the SPS.

[0130] 4) JVET-S0233 proposes the following changes to the process of outputting and removing pictures from the DPB as specified in clause C.5.2.2 of the latest VVC draft specification:

[0131] - Otherwise (the current AU is not a CVSS AU, the current AU is CVSS AU 0, or the current AU is a CVSS AU, the CVSS AU is not AU 0 but the current picture is not the first picture of the current AU), all picture storage buffers containing pictures marked as "not required for output" and "not used for reference" are flushed (not output). For each flushed picture storage buffer, the DPB fullness is decremented by one. When one or more of the following conditions are true, the "buffering" process specified in clause C.5.2.4 is repeatedly invoked, while the DPB fullness is further decremented by one for each additional picture storage buffer flushed, until none of the following conditions are true:

[0132] – The number of pictures marked as “required for output” in the DPB is greater than max_num_reorder_pics[Htid].

[0133] – max_latency_increase_plus1[Htid] is not equal to 0, and there is at least one picture marked as “need to be output” in the DPB, for which the associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].

[0134] – The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1.

[0135] However, the change should be as follows: Otherwise (current AU is not a CVSS AU, or current AU is a CVSS AU [[not AU 0]], but current picture is not the first picture of the current AU)...

[0136] 5) JVET-S0241 proposes to add the following constraints to address the proposed defect, which may cause the POC derivation process to work incorrectly for IRAP pictures in independent non-output layers, where pictures other than IRAP pictures are not used for inter-layer prediction and will therefore be removed during sub-bitstream extraction:

[0137] a. Bitstream conformance requirement is that the value of sps_poc_msb_cycle_flag shall be equal to 1 when both of the following conditions are true:

[0138] –The value of sps_video_parameter_set_id is not equal to 0.

[0139] – The SPS is referenced by at least a VCL NAL unit with layer id equal to nuh_layer_id and NumSubLayersInLayerInOLS[i][GeneralLayerIdx[nuh_layer_id]] equal to 0 and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] equal to 1 for any value of i in the range of 0 to TotalNumOlss–1 (inclusive).

[0140] b. Bitstream conformance requirement is that the value of ph_poc_msb_cycle_present_flag shall be equal to 1 when all of the following conditions are true:

[0141] – The picture associated with the picture header is an IRAP picture or a GDR picture, its associated ph_recovery_poc_cnt is equal to 0, and the picture is not a CLVSS picture.

[0142] – The POC difference between the current picture and the previous IRAP picture or GDR picture with ph_recovery_poc_cnt equal to 0 in the same layer in decoding order is equal to or greater than MaxPicOrderCntLsb / 2.

[0143] – for i in the range of 0 to TotalNumOlss–1 (inclusive), the value of sps_video_parameter_set_id is greater than 0, and the value of NumSubLayersInLayerInOLS[i][GeneralLayerIdx[nuh_layer_id]] is equal to 0, and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1.

[0144] However, the SPS constraints (in item a) appear to be unnecessary, since the PH constraints would impose necessary constraints on the value of sps_poc_msb_cycle_flag, and the proposed constraints are stricter than necessary. For the PH constraints (in item b): For the first item, "not a CLVSS picture" should not exist, since a CRA or GDR picture is not typically a CLVSS picture, but can become a CLVSS picture when random access occurs starting from it. For the third item, "for i in the range 0 to TotalNumOlss-1 (inclusive)" should be "for any value of i in the range 0 to TotalNumOlss-1 (inclusive)", and finally "the value of sps_video_parameter_set_id is greater than 0" is not necessary, since the value of NumSubLayersInLayerInOLS[][][] is inferred when sps_video_parameter_set_id is equal to 0.

[0145] 5. List of Examples and Technical Solutions

[0146] In order to solve the above-mentioned problems and other problems, the method of following generalization is disclosed.These items should be considered as examples of explaining general concepts, and should not be interpreted in a narrow sense.In addition, these items can be applied individually or combined in any way.In the description relating to the change to some existing texts, addition is indicated with bold italic text, and deletion is marked by opening and closing double brackets (for example [[ ]]), and the deleted text is between the double brackets.

[0147] 1) To address issue 1, instead of signaling the set of width, height, chroma format, and bit depth of the DPB picture storage buffer for each multi-layer OLS, these DPB parameters are signaled in the dpb_parameters() syntax structure, each of which can be shared by two or more OLSs.

[0148] 2) To address issue 2, when there are some sub-pictures removed by steps in the sub-picture sub-bitstream extraction process, remove all SEINAL units containing non-scalable nested SEI messages with payloadType equal to 0 (BP), 1 (PT), 130 (DUI), 132 (decoded picture hash) or 203 (SLI).

[0149] a. Alternatively, when there are some sub-pictures removed by steps in the sub-picture sub-bitstream extraction process, remove all SEI NAL units containing non-scalable nested SEI messages with payloadType equal to 0 (BP), 130 (DUI), 132 (Decoded Picture Hash), or 203 (SLI).

[0150] i. Alternatively, furthermore, when general_same_pic_timing_in_all_ols_flag is equal to 0, all SEI NAL units containing non-scalable nesting SEI messages with payloadType equal to 1 (PT) are removed.

[0151] 3) In order to solve problem 3, in the sub-picture sub-bitstream extraction process, when the sub-picture ID mapping is not explicitly signaled in the input bitstream, when a sub-picture sequence identified by a sub-picture index greater than 0 is extracted as the output bitstream, a step of rewriting the reference SPS is added to include the sub-picture ID mapping into the SPS.

[0152] 4) To address issue 4, the process for outputting and removing pictures from the DPB specified in clause C.5.2.2 of the latest VVC draft specification is changed as follows:

[0153] - Otherwise (the current AU is not a CVSS AU, or the current AU is a CVSS AU [[not AU 0]], but the current picture is not the first picture of the current AU), all picture storage buffers containing pictures marked as "not needed for output" and "not used for reference" are flushed (not output). For each flushed picture storage buffer, the DPB fullness is decremented by one. When one or more of the following conditions are true, the "buffering" process specified in clause C.5.2.4 is repeatedly invoked, while the DPB fullness is further decremented by one for each additional picture storage buffer flushed, until none of the following conditions are true:

[0154] – The number of pictures marked as “required for output” in the DPB is greater than max_num_reorder_pics[Htid].

[0155] – max_latency_increase_plus1[Htid] is not equal to 0, and there is at least one picture marked as “need to be output” in the DPB, for which the associated variable PicLatencyCount is greater than or equal to MaxLatencyPictures[Htid].

[0156] – The number of pictures in the DPB is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1.

[0157] 5) To solve problem 5, add the following constraints to VVC:

[0158] Bitstream conformance requirement is that the value of ph_poc_msb_cycle_present_flag shall be equal to 1 when all of the following conditions are true:

[0159] – The picture associated with the picture header is an IRAP picture or a GDR picture with its associated ph_recovery_poc_cnt equal to 0 [[not a CLVSS picture]].

[0160] – The POC difference between the current picture and the previous IRAP picture or GDR picture with ph_recovery_poc_cnt equal to 0 in the same layer in decoding order is equal to or greater than MaxPicOrderCntLsb / 2.

[0161] – For values in the range 0 to TotalNumOlss-1 (inclusive) [[the value of sps_video_parameter_set_id is greater than 0 and]]NumSubLayersInLayerInOLS[i][GeneralLayerIdx[the value of nuh_layer_id]] is equal to 0, and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1.

[0162] 6. Examples

[0163] The following are some example embodiments of some aspects of the present disclosure outlined in Section 5 above that can be applied to the VVC specification. The changed text is based on the latest VVC text in JVET-S0152-v5. The most relevant parts that have been added or modified are indicated in bold italic text, and some deleted parts are marked by opening and closing double brackets (e.g., [[ ]]), with the deleted text between the double brackets. There may be some other changes that are editable in nature and therefore not highlighted.

[0164] 6.1. First embodiment

[0165] This embodiment is used for Project 1.

[0166] 7.3.2.2 Video Parameter Set RBSP Syntax

[0167]

[0168]

[0169] 7.3.2.3 Sequence Parameter Set RBSP Syntax

[0170]

[0171]

[0172] 7.3.4DPB Parameter Syntax

[0173]

[0174] 7.4.3.2 Video Parameter Set RBSP Semantics ...

[0176] [[vps_ols_dpb_pic_width[i] specifies the width in luma samples of each picture storage buffer used for the i-th multi-layer OLS. ]]

[0177] [[vps_ols_dpb_pic_height[i]] specifies the height in luma samples of each picture storage buffer used for the i-th multi-layer OLS. ]]

[0178] [[vps_ols_dpb_chroma_format[i] specifies the maximum allowed value of sps_chroma_format_idcc for all SPSs referenced by the CLVS in the CVS of the i-th multi-layer OLS. ]]

[0179] [[vps_ols_dpb_bitdepth_minus8[i] specifies the maximum allowed value of sps_bitdepth_minus8 for all SPSs referenced by the CLV in the CVS of the i-th multi-layer OLS. ]]

[0180] [[NOTE 2 - To decode the i-th multi-layer OLS, the decoder can safely allocate memory for the DPB according to the values of the syntax elements vps_ols_dpb_pic_width[i], vps_ols_dpb_pic_height[i], vps_ols_dpb_chroma_format[i], and vps_ols_dpb_bitdepth_minus8[i]. ]] ...

[0182] 7.4.3.3 Sequence Parameter Set RBSP Semantics ...

[0184] sps_chroma_format_idc specifies the chroma samples relative to the luma samples, as specified in clause 6.2.

[0185] When sps_video_parameter_set_id is greater than 0 and the SPS is referenced by a layer included in the i-th multi-layer OLS specified by VPS for any i in the range of 0 to NumMultiLayerOlss-1 (inclusive), the value of sps_chroma_format_idc required for the bitstream to conform shall be less than or equal to VPS[[vps_ols_dpb_chroma_format[i]]] ...

[0187] sps_pic_width_max_in_luma_samples specifies the maximum width of each decoded picture that references the SPS in units of luma samples. sps_pic_width_max_in_luma_samples shall not be equal to 0 but shall be an integer multiple of Max(8,MinCbSizeY).

[0188] When sps_video_parameter_set_id is greater than 0 and the SPS is referenced by a layer included in the i-th multi-layer OLS specified by VPS for any i in the range of 0 to NumMultiLayerOlss-1 (inclusive), the bitstream is required to conform to the value of sps_pic_width_max_in_luma_samples shall be less than or equal to VPS[[vps_ols_dpb_pic_width[i]]]

[0189] sps_pic_height_max_in_luma_samples specifies the maximum height of each decoded picture that references the SPS in units of luma samples. sps_pic_height_max_in_luma_samples shall not be equal to 0 but shall be an integer multiple of Max(8,MinCbSizeY).

[0190] When sps_video_parameter_set_id is greater than 0 and the SPS is referenced by a layer included in the i-th multi-layer OLS specified by VPS with any i in the range of 0 to NumMultiLayerOlss-1 (inclusive), the bitstream is required to conform to the value of sps_pic_height_max_in_luma_samples shall be less than or equal to ...

[0192] sps_bitdepth_minus8 specifies the bit depth BitDepth of the samples of the luminance and chrominance arrays, and the value of the luminance and chrominance quantization parameter range offset QpBdOffset, as follows:

[0193] BitDepth=8+sps_bitdepth_minus8 (47)

[0194] QpBdOffset=6*sps_bitdepth_minus8 (48)

[0195] sps_bitdepth_minus8 should be in the range of 0 to 8 (inclusive).

[0196] When sps_video_parameter_set_id is greater than 0 and the SPS is referenced by a layer included in the i-th multi-layer OLS specified by VPS with any i in the range of 0 to NumMultiLayerOlss-1 (inclusive), the bitstream is required to conform to the value of sps_bitdepth_minus8 which shall be less than or equal to ...

[0198] 7.4.5DPB Parameter Semantics ...

[0200] When the dpb_parameters() syntax structure is included in the VPS, the OLS to which the dpb_parameters() syntax structure applies is specified by the VPS. When the dpb_parameters() syntax structure is included in the SPS, it applies to the OLS of the layer that is the lowest layer among the layers that refer to the SPS, and the lowest layer is an independent layer.

[0201] ...

[0203] A.4.1 General Tiers and Level Restrictions ...

[0205] – For OLS with OLS index TargetOlsIdx, the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, and PicSizeMaxInSamplesY and the applicable dpb_parameters() syntax structure are derived as follows:

[0206] o If NumLayersInOls[TargetOlsIdx] is equal to 1, then set PicWidthMaxInSamplesY equal to sps_pic_width_max_in_luma_samples, set PicHeightMaxInSamplesY equal to sps_pic_height_max_in_luma_samples, and set PicSizeMaxInSamplesY equal to PicWidthMaxInSamplesY*PicHeightMaxInSamplesY, where sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are found in the SPS referenced by this layer in the OLS, where the applicable dpb_parameters() syntax structure is also found.

[0207] o Otherwise (NumLayersInOls[TargetOlsIdx] is greater than 1), set PicWidthMaxInSamplesY equal to dpb_ols_pic_width[[vps_ols_dpb_pic_width[MultiLayerOlsIdx[TargetOlsIdx]]]], set PicHeightMaxInSamplesY equal to dpb_ols_pic_height[[vps_ols_dpb_pic_height[MultiLayerOlsIdx[TargetOlsIdx]]]], set PicSizeMaxInSamplesY equal to PicWidthMaxInSamplesY * PicHeightMaxInSamplesY, and identify the applicable dpb_parameters() syntax structure by vps_ols_dpb_params_idx[MultiLayerOlsIdx[TargetOlsIdx]] found in the VPS. ...

[0209] C.1(HRD) General ...

[0211] For each bitstream conformance test, the CPB size (number of bits) is CpbSize[Htid][ScIdx] as specified in clause 7.4.6.3, where ScIdx and HRD parameters are specified above in this clause, and the DPB parameters max_dec_pic_buffering_minus1[Htid], max_num_reorder_pics[Htid], and MaxLatencyPictures[Htid] are found in or derived from the dpb_parameters() syntax structure applicable to the target OLS as follows:

[0212] - If NumLayersInOls[TargetOlsIdx] is equal to 1, the dpb_parameters() syntax structure is found in the SPS referred to as the layer in the target OLS, and the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat and MaxBitDepthMinus8 are set equal to sps_pic_width_max_in_luma_samples, sps_pic_height_max_in_luma_samples, sps_chroma_format_idc and ps_bitdepth_minus8, respectively, found in the SPS referenced by the layer in the target OLS.

[0213] Otherwise (NumLayersInOls[TargetOlsIdx] is greater than 1), the applicable dpb_parameters() syntax structure is identified by vps_ols_dpb_params_idx[MultiLayerOlsIdx[TargetOlsIdx]] found in the VPS, and the variables PicWidthMaxInSamplesY, PicHeightMaxInSamplesY, MaxChromaFormat, and MaxBitDepthMinus8 are set equal to dpb_ols_pic_width, dpb_ols_pic_height, and MaxBitDepthMinus8, respectively, found in the applicable dpb_parameters() syntax structure [[VPS]]. eight, dpb_ols_chroma_format and dpb_ols_bitdepth_minus8, [[vps_ols_dpb_pic_width[MultiLayerOlsIdx[TargetOlsIdx]], vps_ols_dpb_pic_height[MultiLaye rOlsIdx[TargetOlsIdx]], vps_ols_dpb_chroma_format[MultiLayerOlsIdx[TargetOlsIdx]] and vps_ols_dpb_bitdepth_minus8[MultiLayerOlsIdx[TargetOlsIdx]] ...

[0215] Figure 5 is a block diagram illustrating an exemplary video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and the like, as well as wireless interfaces such as Wi-Fi or a cellular interface.

[0216] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from the input 1902 to the output of codec component 1904 to produce a coded representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 can be stored or transmitted via connected communications, as represented by component 1906. The stored or transmitted bitstream (or coded) representation of the video received at input 1902 can be used by component 1908 to generate pixel values or displayable video that is sent to display interface 1910. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0217] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.

[0218] Figure 6 36 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more of the methods described herein. Device 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor 3602 can be configured to implement one or more of the methods described herein. Memory 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0219] Figure 8 is a block diagram illustrating an exemplary video coding system 100 that may utilize the techniques of this disclosure.

[0220] like Figure 8 As shown in , a video codec system 100 may include a source device 110 and a destination device. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.

[0221] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0222] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded and decoded pictures and associated data. The coded and decoded pictures are codec representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to a destination device via the I / O interface 116 via the network 130a. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device.

[0223] The destination device may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0224] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage media / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with the destination device or may be external to the destination device, configured to interface with an external display device.

[0225] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or other standards.

[0226] Figure 9 is a block diagram illustrating an example of a video encoder 200, which may be Figure 8 The video encoder 114 in the system 100 is shown in FIG.

[0227] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 9 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0228] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0229] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.

[0230] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated but are not shown for purposes of explanation. Figure 9 In the example, they are represented separately.

[0231] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0232] The mode selection unit 203 may, for example, select one of the intra or inter coding modes based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) modes, where prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select a resolution (e.g., sub-pixel or integer pixel precision) for the motion vector of the block in the case of inter prediction.

[0233] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information of pictures other than the picture associated with the current video block from the buffer 213 and decoded samples.

[0234] Motion estimation unit 204 and motion compensation unit 205 may perform different operations for the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0235] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures of list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that includes the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0236] In other examples, motion estimation unit 204 may perform bidirectional prediction on the current video block. Motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. Motion estimation unit 204 may then generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video block and a motion vector indicating the spatial displacement between the reference video block and the current video block. Motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0237] In some examples, motion estimation unit 204 may output the entire set of motion information for use in the decoding process of the decoder.

[0238] In some examples, motion estimation unit 204 may not output the entire set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information for the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.

[0239] In one example, motion estimation unit 204 may indicate, in a syntax structure associated with the current video block, a value that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0240] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0241] As discussed above, video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge mode signaling.

[0242] Intra-prediction unit 206 may perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, intra-prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0243] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0244] In other examples, such as in skip mode, there may be no residual data for the current video block, and residual generation unit 207 may not perform a subtraction operation.

[0245] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0246] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0247] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block for storage in buffer 213.

[0248] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0249] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0250] Figure 10 is a block diagram illustrating an example of a video decoder 300, which may be Figure 8 The video decoder 114 in the system 100 is shown in FIG.

[0251] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 10 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0252] exist Figure 10 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform operations generally similar to those described with respect to the video encoder 200 ( Figure 9 ) is a decoding round that is the inverse of the encoding described.

[0253] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 may decode the entropy-encoded video data, and the motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information from the entropy-decoded video data. For example, the motion compensation unit 302 may determine such information by performing AMVP and merge mode.

[0254] Motion compensation unit 302 may generate motion compensated blocks, possibly performing interpolation based on interpolation filters.Identifiers of interpolation filters to be used with sub-pixel precision may be included in the syntax elements.

[0255] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters as used during encoding and decoding of the video block by video encoder 200. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 from received syntax information and use the interpolation filters to produce a predictive block.

[0256] The motion compensation unit 302 may use some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence.

[0257] The intra-frame prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra-frame prediction mode received in the bitstream. The inverse quantization unit 303 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0258] The reconstruction unit 306 may sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter may also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0259] A list of preferred solutions for some embodiments is provided below.

[0260] The following solutions illustrate exemplary embodiments of the techniques discussed in the previous sections (eg, item 1).

[0261] 1. A video processing method (e.g., Figure 7 ), the method comprising: performing a conversion between a video comprising one or more video layers and a codec representation of the video, the one or more video layers comprising one or more color components; wherein the codec representation conforms to a format rule that specifies that the codec representation comprises a single grammatical structure indicating decoded picture cache parameters applicable to each layer of a multi-layer output layer set of the codec representation.

[0262] 2. The method according to solution 1, wherein the decoded picture buffer parameters include width, height, chroma format or bit depth of the decoded picture buffer.

[0263] The following solutions illustrate exemplary embodiments of the techniques discussed in the previous section (eg, items 2, 3, 4).

[0264] 3. A video processing method, the method comprising: performing a conversion between a video comprising one or more video layers and a codec representation of the video, the one or more video layers comprising one or more color components; wherein the codec representation complies with a rule associated with processing the codec representation based on removing pictures or sub-pictures from a buffer during a decoding process.

[0265] 4. The method according to solution 3, wherein the rule specifies that if a sub-picture is removed, then all Supplemental Enhancement Information Network Abstraction Layer (SEINAL) units containing non-scalable nested SEI messages with a payloadType equal to a predefined type are removed.

[0266] 5. The method according to solution 4, wherein the predetermined type is 0 (BP), 1 (PT), 130 (DUI), 132 (Decoded Picture Hash), or 203 (SLI).

[0267] 6. The method according to solution 4, wherein the predetermined type is 0 (BP), 130 (DUI), 132 (Decoded Picture Hash), or 203 (SLI).

[0268] 7. A method according to solution 3, wherein the rule specifies that in the case where a sub-picture sequence identified by a sub-picture index greater than zero is extracted as the output bitstream, since the codec representation does not have a sub-picture ID mapping, the referenced sequence parameter set (SPS) is rewritten to include the sub-picture ID mapped to the SPS.

[0269] 8. The method according to solution 3, wherein the rule provides for removing pictures from the decoded picture cache in a similar manner for all access units.

[0270] 9. The method according to any of solutions 1 to 8, wherein the conversion comprises encoding the video into the codec representation.

[0271] 10. The method according to any one of solutions 1 to 8, wherein the converting comprises decoding the codec representation to generate pixel values of the video.

[0272] 11. A video decoding device, comprising a processor configured to implement the method of one or more of solutions 1 to 10.

[0273] 12. A video encoding apparatus comprising a processor configured to implement the method of one or more of solutions 1 to 10.

[0274] 13. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of solutions 1 to 10.

[0275] 14. A method, apparatus or system as described in this document.

[0276] Figure 13 13 is a flow chart of an exemplary method 1300 for video processing. Operation 1302 comprises performing conversion between a video and a bitstream of the video, wherein the bitstream comprises one or more pictures, the one or more pictures comprising one or more sub-pictures, according to a rule, and wherein the rule specifies that, in response to a condition, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced.

[0277] In some embodiments of the method 1300, the condition is that no sub-picture identifier mapping is indicated in the bitstream, and the sub-picture sub-bitstream extraction process extracts a sub-picture sequence identified by a sub-picture index greater than zero. In some embodiments of the method 1300, the rewriting operation includes including the sub-picture identifier mapping in the referenced one or more sequence parameter sets.

[0278] Figure 14 14 is a flow chart of an exemplary method 1400 for video processing. Operation 1402 comprises performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream comprises one or more output layer sets comprising a plurality of layers according to a format rule, wherein the format rule provides for indicating in the bitstream a single syntax structure applicable to each of the one or more output layer sets, and wherein the single syntax structure comprises decoded picture cache parameters.

[0279] In some embodiments of the method 1400 , the decoded picture buffer parameters include a width, a height, a chroma format, or a bit depth of the decoded picture buffer.

[0280] Figure 15 is a flow chart of an exemplary method 1500 for video processing. Operation 1502 comprises performing conversion between a video and a bitstream of the video, wherein the bitstream comprises one or more pictures, the one or more pictures comprising one or more sub-pictures, according to a rule, and wherein the rule specifies that during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, in response to removing a sub-picture from the bitstream, all supplemental enhancement information network abstraction layer units including non-scalable nested supplemental enhancement information messages of a specific payload type are also removed.

[0281] In some embodiments of the method 1500, a first value of the specific payload type is 0, indicating that the specific payload type is a buffering period, wherein a second value of the specific payload type is 1, indicating that the specific payload type is picture timing, wherein a third value of the specific payload type is 130, indicating that the specific payload type is decoding unit information, wherein a fourth value of the specific payload type is 132, indicating that the specific payload type is a decoded picture hash, or wherein a fifth value of the specific payload type is 230, indicating that the specific payload type is sub-picture level information. In some embodiments of the method 1500, a first value of the specific payload type is 0, indicating that the specific payload type is a buffering period, wherein a second value of the specific payload type is 130, indicating that the specific payload type is decoding unit information, wherein a third value of the specific payload type is 132, indicating that the specific payload type is a decoded picture hash, or wherein a fourth value of the specific payload type is 230, indicating that the specific payload type is sub-picture level information. In some embodiments of the method 1500, the rule provides that in response to general_same_pic_timing_in_all_ols_flag being equal to 0, all supplemental enhancement information network abstraction layer units including non-scalable nested supplemental enhancement information messages having a specific payload type value of 1 are removed from the bitstream, and wherein the specific payload type value of 1 indicates that the specific payload type is picture timing.

[0282] Figure 16 is a flow chart of an exemplary method 1600 for video processing. Operation 1602 comprises performing conversion between a video and a bitstream of the video according to a rule, wherein the bitstream comprises a plurality of access units, the plurality of access units comprising one or more pictures, and wherein the rule provides for removing a current picture of a current access unit from a decoded picture cache in response to the following conditions: (1) the current picture is not the first picture of the current access unit, and (2) the current access unit is the first access unit in decoding order in the bitstream.

[0283] In some embodiments of method 1600, the number of pictures in the decoded picture cache marked as needing to be output is greater than max_num_reorder_pics[Htid], wherein max_num_reorder_pics[] indicates a maximum allowed number of pictures of the output layer set that precede any picture in the output layer set in the decoding order and follow the picture in the output order, and wherein Htid indicates a highest temporal sub-layer to be decoded. In some embodiments of method 1600, the decoded picture cache includes at least one picture marked as needing to be output, for which at least one picture has an associated variable picture latency count greater than or equal to MaxLatencyPictures[Htid], wherein MaxLatencyPictures[Htid] indicates a maximum number of pictures allowed to precede a particular picture in the output order and follow the particular picture in the decoding order, wherein Htid indicates the highest temporal sub-layer to be decoded, and wherein max_latency_increase_plus1[Htid] is not equal to 0. In some embodiments of method 1600, the number of pictures in the decoded picture cache is greater than or equal to max_dec_pic_buffering_minus1[Htid]+1, and wherein max_dec_pic_buffering_minus1[Htid]+1 specifies the maximum required size of the decoded picture cache in units of picture storage buffers, and wherein Htid indicates the highest temporal sub-layer to be decoded.

[0284] Figure 17is a flow chart of an exemplary method 1700 for video processing. Operation 1702 comprises: performing conversion between a video comprising one or more pictures and a bitstream of the video, wherein the bitstream comprises a first syntax element, the first syntax element having a value indicating the presence of a second syntax element in a picture header, the second syntax element specifying a value of a most significant bit of a picture sequence count of a current picture, wherein the bitstream conforms to a format rule responsive to the following conditions, the format rule specifying that the value of the first syntax element is equal to 1: (1) the current picture associated with the picture header is an intra random access point picture or a progressive decoding refresh picture, the current picture having an associated third syntax element equal to 0 an element indicating that the current picture is a recovery point picture, (2) the picture order count difference between the current picture and a previous intra random access point picture or a previous progressive decoding refresh picture in the same layer in decoding order whose associated third syntax element is equal to 0 is equal to or greater than a variable indicating the maximum picture order count least significant bit divided by 2, and (3) for any value of i in the range of 0 to TotalNumOlss-1, inclusive, the value of NumSubLayersInLayerInOLS[i][GeneralLayerIdx[nuh_layer_id]] is equal to 0, and vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1, where NumSubLayersInLayerInOLS[i][GeneralLayerIdx[nuh_layer_id]] indicates the number of sublayers in the GeneralLayerIdx[nuh_layer_id]th output layer in the i-th output layer set, where TotalNumOlss specifies the total number of output layer sets specified by the video parameter set, where vps_independent_layer_flag[GeneralLayerIdx[nuh_layer_id]] is equal to 1 specifies that the layer with layer index equal to GeneralLayerIdx[nuh_layer_id] does not use inter-layer prediction, where GeneralLayerIdx[nuh_layer_id] specifies the layer index with layer identifier equal to nuh_layer_id, and where nuh_layer_id specifies the identifier of the layer to which the video codec layer network abstraction layer unit belongs or the identifier of the layer to which the non-video codec layer network abstraction layer unit applies.

[0285] In some embodiments of methods 1300-1700, performing the conversion includes encoding the video into a bitstream. In some embodiments of methods 1300-1700, performing the conversion includes generating a bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium. In some embodiments of methods 1300-1700, performing the conversion includes decoding the video from the bitstream.

[0286] In some embodiments, a video decoding device includes a processor configured to implement operations associated with embodiments of methods 1300-1700. In some embodiments, a video encoding device includes a processor configured to implement operations associated with embodiments of methods 1300-1700. In some embodiments, a computer program product having computer instructions stored thereon, the instructions causing the processor to implement operations associated with embodiments of methods 1300-1700 when executed by the processor. In some embodiments, a non-transitory computer-readable storage medium stores a bitstream generated according to operations associated with embodiments of methods 1300-1700. In some embodiments, a non-transitory computer-readable storage medium stores instructions causing the processor to implement operations associated with embodiments of methods 1300-1700. In some embodiments, a method for generating a bitstream includes generating a bitstream of a video according to operations associated with embodiments of methods 1300-1700, and storing the bitstream on a computer-readable program medium. In some embodiments, methods, devices, and bitstreams generated according to the disclosed methods or systems described in this document are disclosed.

[0287] Some embodiments of the disclosed technology include making a decision or determining to enable a video processing tool or mode. In one example, when the video processing tool or mode is enabled, the encoder will use or implement the tool or mode when processing the video block, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when the video processing tool or mode is enabled based on the decision or determination, the conversion from the video block to the bitstream representation of the video will use the video processing tool or mode. In another example, when the video processing tool or mode is enabled, the decoder will process the bitstream in the case where the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to the video block will be performed using the video processing tool or mode enabled based on the decision or determination.

[0288] Some embodiments of the disclosed technology include making a decision or determination to deactivate a video processing tool or mode. In one example, when the video processing tool or mode is deactivated, the encoder will not use the tool or mode in converting video blocks into a bitstream representation of the video. In another example, when the video processing tool or mode is deactivated, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was deactivated based on the decision or determination.

[0289] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, during the conversion from a pixel representation of a video to a corresponding bitstream representation, a video compression algorithm may be applied, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits that are co-located or interspersed at different locations within the bitstream as defined by the syntax. For example, a macroblock may be encoded based on error residual values that have been transformed and encoded, and may also be encoded using bits in headers and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on the knowledge that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether certain syntax fields will be included and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0290] In the solution described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the solution described herein, a decoder can use the format rules to parse syntax elements in the codec representation, where the presence and absence of syntax elements are known according to the format rules to produce decoded video.

[0291] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of materials that implement a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0292] A computer program (also referred to as a program, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that stores other programs or data (e.g., one or more scripts stored in a markup language file), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer, or to execute on multiple computers that are located in one location or distributed in multiple locations and interconnected by a communication network.

[0293] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0294] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to one or more mass storage devices for storing data to receive data from or transfer data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CDROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0295] Although this patent document contains many details, these details should not be interpreted as limitations on the scope of any subject matter or the scope that may be claimed, but rather should be interpreted as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed as such, one or more features from a claimed combination may in some cases be removed from the combination, and a claimed combination may be directed to a subcombination or variations of a subcombination.

[0296] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0297] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing video data, comprising: Perform conversion between video and bitstream of said video, According to the rule, the bitstream includes one or more pictures, and the one or more pictures include one or more sub-pictures. wherein the rule specifies that, in response to a condition that no sub-picture identifier mapping is indicated in the bitstream and the sub-picture sub-bitstream extraction process extracts a sub-picture sequence identified by a sub-picture index greater than zero, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced thereto; and The rewriting operation includes including a sub-picture identifier mapping into one or more sequence parameter sets of the reference.

2. The method according to claim 1, wherein According to the format rules, the bitstream includes one or more output layer sets, the one or more output layer sets including a plurality of layers, wherein the format rules specify that a single syntax structure applicable to each of one or more sets of output layers is indicated in the bitstream, and The single syntax structure includes decoded picture cache parameters.

3. The method according to claim 2, wherein: The decoded picture buffer parameters include the width, height, chroma format or bit depth of the decoded picture buffer.

4. The method according to claim 1, wherein The rules further provide that during a sub-picture sub-bitstream extraction process of extracting the target output sub-bitstream from the bitstream, in response to removing a sub-picture from the bitstream, all supplemental enhancement information network abstraction layer units including non-scalable nested supplemental enhancement information messages of a specific payload type are also removed.

5. The method according to claim 4, wherein The first value of the specific payload type is 0, indicating that the specific payload type is a buffer period. The second value of the specific payload type is 1, indicating that the specific payload type is picture timing. The third value of the specific payload type is 130, indicating that the specific payload type is decoding unit information. The fourth value of the specific payload type is 132, indicating that the specific payload type is a decoded picture hash, or The fifth value of the specific payload type is 230, indicating that the specific payload type is sub-picture level information.

6. The method according to claim 4, wherein: The first value of the specific payload type is 0, indicating that the specific payload type is a buffer period. The second value of the specific payload type is 130, indicating that the specific payload type is decoding unit information. The third value of the specific payload type is 132, indicating that the specific payload type is a decoded picture hash, or The fourth value of the specific payload type is 230, indicating that the specific payload type is sub-picture level information.

7. The method according to claim 4, wherein: The rule states that, in response to general_same_pic_timing_in_all_ols_flag being equal to 0, all supplemental enhancement information network abstraction layer units including the non-scalable nested supplemental enhancement information message with the value of 1 for the specific payload type are removed from the bitstream, and The value of the specific payload type being 1 indicates that the specific payload type is picture timing.

8. The method according to claim 1, wherein Performing the conversion includes encoding the video into the bitstream.

9. The method according to claim 1, wherein Performing the conversion includes decoding the video from the bitstream.

10. A device for processing video data, comprising: processor; as well as a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Perform conversion between video and bitstream of said video, According to the rule, the bitstream includes one or more pictures, and the one or more pictures include one or more sub-pictures. wherein the rule specifies that, in response to a condition that no sub-picture identifier mapping is indicated in the bitstream and the sub-picture sub-bitstream extraction process extracts a sub-picture sequence identified by a sub-picture index greater than zero, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced thereto; and The rewriting operation includes including a sub-picture identifier mapping into one or more sequence parameter sets of the reference.

11. A non-transitory computer-readable storage medium having stored thereon instructions that cause a processor to: Perform conversion between video and bitstream of said video, in, According to a rule, the bitstream includes one or more pictures, and the one or more pictures include one or more sub-pictures. wherein the rule specifies that, in response to a condition that no sub-picture identifier mapping is indicated in the bitstream and the sub-picture sub-bitstream extraction process extracts a sub-picture sequence identified by a sub-picture index greater than zero, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced thereto; and The rewriting operation includes including a sub-picture identifier mapping into one or more sequence parameter sets of the reference.

12. A non-transitory computer-readable storage medium storing a bitstream, the bitstream being generated by a method performed by a video processing device, wherein: The method comprises: generating a bitstream of the video, According to the rule, the bitstream includes one or more pictures, and the one or more pictures include one or more sub-pictures. wherein the rule specifies that, in response to a condition that no sub-picture identifier mapping is indicated in the bitstream and the sub-picture sub-bitstream extraction process extracts a sub-picture sequence identified by a sub-picture index greater than zero, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced thereto; and The rewriting operation includes including a sub-picture identifier mapping into one or more sequence parameter sets of the reference.

13. A method for storing a bitstream, comprising: Generate video bitstream; as well as storing the bitstream in a non-transitory computer-readable medium, wherein, according to a rule, the bitstream includes one or more pictures, the one or more pictures including one or more sub-pictures, wherein the rule specifies that, in response to a condition that no sub-picture identifier mapping is indicated in the bitstream and the sub-picture sub-bitstream extraction process extracts a sub-picture sequence identified by a sub-picture index greater than zero, during a sub-picture sub-bitstream extraction process for extracting a target output sub-bitstream from the bitstream, a rewrite operation is performed on one or more sequence parameter sets referenced thereto; and The rewriting operation includes including a sub-picture identifier mapping into one or more sequence parameter sets of the reference.

14. A video decoding apparatus, comprising a processor configured to implement the method according to any one of claims 1 to 9.

15. A video encoding apparatus comprising a processor configured to implement the method according to any one of claims 1 to 9.

16. A non-transitory computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, cause the processor to implement the method according to any one of claims 1 to 9.

17. A non-transitory computer-readable storage medium storing a bitstream generated by a method performed by a video processing device, wherein: The method comprises the method according to any one of claims 1 to 9.

18. A method for generating a bitstream, comprising: Generating a video bitstream according to any one of claims 1 to 9, and The bitstream is stored on a computer readable program medium.

Citation Information

Patent Citations

  • Scaling list signaling and parameter sets activation

    CN105453569A

  • Multi-layer video coding

    CN106464924A