Expansion of Sub-Picture Sub-Bitstream Extraction Process

The sub-picture sub-bitstream extraction process is enhanced by implementing specific rules for SEI message handling, sub-picture indexing, and syntax element rewriting, addressing existing compliance issues and improving the process's reliability and efficiency.

JP7688053B2Active Publication Date: 2025-06-03BYTEDANCE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022575699
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-09
Filing Date
2021-06-08
Publication Date
2025-06-03
Estimated Expiration
2041-06-08

AI Technical Summary

Technical Problem

The existing sub-picture sub-bitstream extraction process in video coding technologies faces issues such as incorrect handling of SEI messages, inconsistent sub-picture indexing, and unclear conditions for rewriting syntax elements, leading to non-compliant bitstreams.

Method used

The proposed method involves specific rules for the sub-picture sub-bitstream extraction process, including the removal of non-applicable SEI NAL units, correct sub-picture indexing based on layers with multiple sub-pictures, and conditional rewriting of syntax elements like general_level_idc, sublayer_level_idc, CPB size, and bitrate values.

Benefits of technology

This approach ensures compliance with bitstream requirements by accurately handling SEI messages, maintaining consistent sub-picture indexing, and appropriately modifying syntax elements, thereby improving the reliability and efficiency of the sub-picture sub-bitstream extraction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007688053000001
    Figure 0007688053000001
  • Figure 0007688053000002
    Figure 0007688053000002
  • Figure 0007688053000003
    Figure 0007688053000003
Patent Text Reader

Abstract

A method for processing video data includes performing a conversion between a video and a video bitstream, the bitstream including a plurality of layers including one or more subpictures according to a rule, the rule specifying that during a subpicture sub-bitstream extraction process in which an output bitstream is extracted from the bitstream, supplemental enhancement information network abstraction layer units (SEI NAL units) including scalable nested SEI messages that are not applicable to the output bitstream are omitted in the output bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application claims the priority and benefit of 、2 U.S. Provisional Patent Application No. 63 / 036,908, filed on Jun. 9, 2020 This application claims priority based on International Patent Application No. PCT / US2021 / 036359, filed on Jun. 8, 2021. The entire disclosure of all the above patent applications is incorporated by reference.

[0002] [Technical Field] This patent document relates to image and video data processing.

Background Art

[0003] Digital video occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, it is assumed that the bandwidth demand for digital video utilization continues to increase.

Summary of the Invention

[0004] This document discloses techniques that can be used by video encoders and decoders to process coded representations of video or images.

[0005] 1 In one exemplary aspect, a method for processing video data is disclosed. The method includes performing a conversion between video and a bitstream of the video, the bitstream including a plurality of layers including one or more sub - pictures according to rules, the rules being scalable nesting that is not applicable to the output bitstream during the sub - picture sub - bitstream extraction process from which the output bitstream is extracted issued specifying that a Supplemental Enhancement Information Network Abstraction Layer Unit (SEI NAL Unit) including SEI messages is omitted in the output bitstream.

[0006] In one exemplary aspect, a method for processing video data is disclosed. The method includes performing a conversion between a video and a bitstream of the video, the bitstream including a plurality of layers each including one or more pictures including one or more sub-pictures according to a rule, the rule specifying that a first sub-picture index identifying a sub-picture sequence extracted by a sub-picture sub-bitstream extraction process for the bitstream is based on a second sub-picture index of a layer of the bitstream having a plurality of sub-pictures per picture.

[0007] In one exemplary aspect, a method for processing video data is disclosed. The method includes performing a conversion between a video and a bitstream of the video, the bitstream including one or more layers each including one or more sub-layers each including one or more pictures including one or more sub-pictures according to a rule, the rule specifying a rewrite operation selectively performed on one or more syntax structures during a sub-picture sub-bitstream extraction process from which an output target sub-bitstream is extracted from the bitstream, the one or more syntax elements including information of the output target sub-bitstream.

[0008] In one exemplary aspect, a method for processing video data is disclosed. The method includes performing a conversion between a video and a bitstream of the video, the bitstream including a plurality of layers each including one or more pictures including one or more sub-pictures according to a rule, the rule specifying a selection process of a first supplemental enhancement information network abstraction layer (SEI NAL) unit of a target output sub-picture sub-bitstream extracted during a sub-picture sub-bitstream extraction process according to a condition.

[0009] In yet another exemplary aspect, a video encoder device is disclosed. The video encoder includes a processor configured to implement the method described above.

[0010] In yet another exemplary aspect, a video decoder device is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0011] In yet another exemplary aspect, a computer-readable medium storing code is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.

[0012] These and other features are described throughout this document.

Brief Description of the Drawings

[0013]

FIG. 1

FIG. 2

FIG. 3

FIG. 4

FIG. 5

FIG. 6

FIG. 7

FIG. 8

FIG. 9

FIG. 10

FIG. 11

FIG. 12

FIG. 13A

FIG. 13B

FIG. 13C

FIG. 13D

DETAILED DESCRIPTION OF THE INVENTION

[0014] The section headings are used in this document to facilitate understanding and do not limit the applicability of the technologies and embodiments disclosed in each section to that section only. Further, the terms of H.266 are used only in some descriptions to facilitate understanding and are not used to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs. In this document, editorial changes are indicated in the text by strike-through ([[]] including) for cancelled text and highlighting (bold italic, including underline) for added text with respect to the current draft of the VVC specification.

[0015] 1. Introduction This document relates to video coding technology. Specifically, it is about specifying and signaling level information for sub-picture sequences. The concept may be applied to any video coding standard or non-standard video codec that supports single-layer video coding and multi-layer video coding, for example, the Versatile Video Coding (VVC) under development.

[0016] 2. Abbreviations APS Adaptation Parameter Set AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding CLVS Coded Layer Video Sequence CPB Coded Picture Buffer CRA Clean Random Access CTU Coding Tree Unit CVS Coded Video Sequence DCI Decoding Capability Information DPB Decoded Picture Buffer EOB End Of Bitstream EOS End Of Sequence GDR Gradual Decoding Refresh HEVC High Efficiency Video Coding HRD Hypothetical Reference Decoder IDR Instantaneous Decoding Refresh ILP Inter-Layer Prediction ILRP Inter-Layer Reference Picture JEM Joint Exploration Model LTRP Long-Term Reference Picture MCTS Motion-Constrained Tile Sets NAL Network Abstraction Layer OLS Output Layer Set PH Picture Header PPS Picture Parameter Set PTL Profile, Tier and Level PU Picture Unit RAP Random Access Point RBSP Raw Byte Sequence Payload SEI Supplemental Enhancement Information SLI Subpicture Level Information SPS Sequence Parameter Set STRP Short-Term Reference Picture SVC Scalable Video Coding VCL Video Coding Layer VPS Video Parameter Set VTM VVC Test Model VUI Video Usability Information VVC Versatile Video Coding

[0017] 3. Initial Discussion Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and the two organizations jointly created H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC), as well as the H.265 / HEVC [1] standard. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes transform coding in addition to temporal prediction. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into the reference software named the Joint Exploration Model (JEM) [2]. Currently, JVET meetings are held quarterly, and the new coding standard aims to reduce the bitrate by 50% compared to HEVC. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Since continuous efforts are being made to contribute to VVC standardization, new coding technologies are being adopted into the VVC standard for each JVET meeting. Subsequently, the VVC working draft and the test model VTM have been updated after each meeting. The VVC project is currently aiming for a Final Draft International Standard (FDIS) at the meeting in July 2020.

[0018] 3.1. Picture Partitioning Methods in HEVC HEVC can be applied to maximum transfer unit (MTU) size matching, parallel processing, and reduced end-to-end latency, and includes four different picture partitioning methods, namely, regular slice, dependent slice, tile, and wavefront parallel processing (WPP).

[0019] Regular slices are similar to H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and the intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (even though there may still be mutual dependencies due to loop filtering operations).

[0020] Regular slices are the only tool available for parallelization that can be used in substantially the same form in H.264 / AVC. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures. This is typically much heavier than inter-processor or inter-core data sharing by intra-picture prediction). However, for the same reason, the use of regular slices can cause significant coding overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. Furthermore, regular slices (in contrast to the other tools mentioned below) also function as an important mechanism for bitstream partitioning that matches the MTU size requirement due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit. In many cases, the goals of parallelization and MTU size matching impose conflicting requirements on the slice layout within a picture. The realization of this situation has led to the development of the parallelization tools mentioned below.

[0021] Dependent slices have short slice headers and enable partitioning of the bitstream at tree block boundaries without sacrificing any intra-picture prediction. Basically, dependent slices provide reduced end-to-end latency by enabling a part of a regular slice to be sent before the encoding of the entire regular slice is complete, by providing fragmentation of the regular slice into multiple NAL units.

[0022] In WPP, pictures are partitioned into a single row of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs within other partitions. Parallel processing is possible through parallel decoding of CTB rows, and the start of decoding of a CTB row is delayed by only two CTBs to ensure that data regarding the CTB to the upper right of the target CTB is available before the target CTB is decoded. Using this staggered start (which looks like a wavefront when graphically represented), parallelization is possible up to the same number of processors / cores as the picture has CTB rows. Since in-picture prediction between adjacent tree block rows within a picture is allowed, the necessary inter-processor / inter-core communication to enable in-picture prediction can be quite substantial. WPP partitioning does not result in the generation of additional NAL units compared to the case where it is not applied, so WPP is not a tool for MTU size matching. However, when MTU size matching is required, regular slices can be used with WPP with a specific coding overhead.

[0023] Tiles define horizontal and vertical boundaries that partition a picture into tile columns and rows. Tile columns extend from the top to the bottom of the picture. Similarly, tile rows extend from the left to the right of the picture. The number of tiles within a picture can be simply derived as the multiplication of the number of tile columns and the number of tile rows.

[0024] The scan order of CTBs is changed locally within a tile (in the order of CTB raster scan of the tile) before decoding the top-left CTB of the next tile in the order of tile raster scan of the picture. Similar to regular slices, tiles break the dependencies of intra-picture prediction and entropy decoding. However, since these do not need to be included in individual NAL units (the same as in WPP in this regard), tiles cannot be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent tiles is limited to the transmission of the shared slice header when a slice spans more than one tile and the sharing related to loop filtering of the reconstructed samples and metadata. When more than one tile or WPP segment is included in a slice, the entropy point byte offset for each tile or WPP segment other than the first one in the slice is signaled in the slice header.

[0025] Briefly, restrictions on the application of four different picture partitioning methods are specified in HEVC. A given coded video sequence cannot contain both tiles and waves for most of the profiles specified in HEVC. For each slice and tile, one or both of the following conditions must be met. 1) All coded tree blocks within a slice belong to the same tile. 2) All coded tree blocks within a tile belong to the same slice. Finally, a wave segment contains exactly one CTB row, and if WPP is used, if a slice starts within a CTB row, it must end within the same CTB row.

[0026] The most recent amendment to HEVC is specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)", dated October 24, 2017, which is publicly available from http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. With this amendment, HEVC specifies three MCTS-related SEI messages, namely, the Temporal MCTS SEI message, the MCTS Extraction Information Set SEI message, and the MCTS Extraction Information Nested SEI message.

[0027] The Temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vectors are restricted to indicate the full sample position within the MCTS and the fractional sample position that requires only the full sample position within the MCTS for interpolation, and the use of motion vector candidates for temporal motion vector prediction derived from blocks outside the MCTS is not permitted. Thus, each MCTS may be decoded independently without the presence of tiles not included in the MCTS.

[0028] The MCTS extraction information set SEI message provides supplementary information (specified as part of the semantics of the SEI message) that can be used in the MCTS sub-bitstream to generate a conforming bitstream for the MCTS set. The information is composed of multiple extraction information sets, and each extraction information set defines multiple MCTS sets and includes the RBSP bytes of the replacement VPS, SPS, and PPS used during the MCTS sub-bitstream extraction process. When extracting the sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because typically one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) need to have different values.

[0029] 3.2. Picture Partitioning in VVC In VVC, a picture is divided into one or more tile rows and one or more tile columns. A tile is a contiguous sequence of CTUs that cover a rectangular region of the picture. The CTUs within a tile are scanned in raster scan order within that tile.

[0030] A slice is composed of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of the picture.

[0031] Two modes of slices, namely, the raster scan slice mode and the rectangular slice mode, are supported. In the raster scan slice mode, a slice contains a sequence of complete tiles in the tile raster scan of the picture. In the rectangular slice mode, a slice contains a plurality of complete tiles that together form a rectangular region of the picture, or a plurality of consecutive complete CTU rows of one tile that together form a rectangular region of the picture. Tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to that slice.

[0032] A sub-picture contains one or more slices that together cover a rectangular region of the picture.

[0033] Figure 1 shows an example of raster scan slice partitioning of a picture, where the picture is divided into 12 tiles and 3 raster scan slices.

[0034] Figure 2 shows an example of rectangular slice partitioning of a picture, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.

[0035] Figure 3 shows an example of a picture partitioned into tiles and rectangular slices, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.

[0036] Figure 4 shows an example of sub-picture partitioning of a picture, where the picture is partitioned into 18 tiles. The 12 tiles on the left each cover one slice of 4×4 CTUs, and the 6 tiles on the right each cover two vertically stacked slices of 2×2 CTUs, resulting in 24 slices and 24 sub-pictures of different dimensions (each slice is a sub-picture).

[0037] 3.3. Change of Picture Resolution within a Sequence In AVC and HEVC, the spatial resolution of a picture cannot be changed unless a new sequence using a new SPS starts with an IRAP picture. VVC can change the picture resolution within a sequence at certain positions without encoding an IRAP picture that is always intra-coded. When a reference picture has a different resolution from the current picture being decoded, this function requires resampling of the reference picture used for inter prediction, so this function is sometimes called reference picture resampling (RPR).

[0038] The scaling ratio is limited to 1 / 2 or more (two downsamplings from the reference picture to the current picture) and 8 or less (eight upsamplings). To handle various scaling ratios between the reference picture and the current picture, three sets of resampling filters with different frequency cutoffs are specified. The three sets of resampling filters are applied for scaling ratios in the ranges of 1 / 2 to 1 / 1.75, 1 / 1.75 to 1 / 1.25, and 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, similar to the case of the motion compensation interpolation filter. In practice, the normal MC interpolation process is a special case of the resampling process with a scaling ratio in the range of 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height, and the left, right, up, and down scaling offsets specified for the reference picture and the current picture.

[0039] Other aspects of the VVC design to support this feature, different from HEVC, include the following. i) The picture resolution and the corresponding compliant window are signaled in the PPS rather than the SPS, where the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture store (slot in the DPB for storing one decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.

[0040] 3.4. General Scalable Video Coding (SVC) and SVC in VVC Scalable Video Coding (SVC, also sometimes referred to as scalability in video coding) refers to video coding in which a base layer (BL), sometimes called a reference layer (RL), and one or more scalable enhancement layers (EL) are used. In SVC, the base layer can carry video data of base-level quality. One or more enhancement layers can carry additional video data, for example, to support higher spatial, temporal, and / or signal-to-noise (SNR) levels. The enhancement layer may be defined with respect to a previously encoded layer. For example, a lower layer may function as the BL, and an upper layer may function as the EL. An intermediate layer may function as either the EL or the RL, or both. For example, an intermediate layer (e.g., a layer that is neither the lowest layer nor the highest layer) may be the EL for the layer below the intermediate layer, e.g., the base layer or any intervening enhancement layer, and at the same time may function as the RL for one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, multiple views may exist, and the information of one view may be utilized (e.g., encoded or decoded) to code the information of other views (e.g., motion estimation, motion vector prediction, and / or other redundancies).

[0041] In SVC, the parameters used by an encoder or a decoder are grouped into parameter sets based on the coding levels at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, the parameters that can be utilized by one or more coded video sequences of different layers within a bitstream may be included in a video parameter set (VPS), and the parameters that are utilized by one or more pictures within a coded video sequence may be included in a sequence parameter set (SPS). Similarly, the parameters that are utilized by one or more slices within a picture may be included in a picture parameter set (PPS), and other parameters specific to a single slice may be included in the slice header. Similarly, an indication of which parameter set a particular layer is using at a given time may be provided at various coding levels.

[0042] Thanks to the support of reference picture resampling (RPR) in VVC, the support of bitstreams containing multiple layers, for example, two layers with SD and HD resolutions in VVC, can be designed without the need for additional signal processing level coding tools because the upsampling required for spatial scalability only needs to use the RPR upsampling filter. Nevertheless, to support scalability, changes to high-level syntax are required (compared to the case where scalability is not supported). The support for scalability is specified in VVC version 1. Different from the support for scalability in any previous video coding standard, including extensions of AVC and HEVC, the design of VVC scalability is made to be suitable for the design of single-layer decoders as much as possible. The decoding capabilities for multi-layer bitstreams are specified as if only a single layer exists in the bitstream. For example, decoding capabilities such as the DPB size are specified in a way that does not depend on the number of layers of the bitstream being decoded. Basically, a decoder designed for a single-layer bitstream does not require major changes to be able to decode a multi-layer bitstream. Compared to the design of multi-layer extensions of AVC and HEVC, the HLS aspect is considerably simplified at the expense of some flexibility. For example, IRAP AUs are required to include each picture of the layers present in the CVS.

[0043] 3.5. Sub-picture based viewport dependent 360° video streaming In the streaming of 360° video, also known as omnidirectional video, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) at any given instant is rendered to the user, but the user can always turn their head to change the viewing direction and thus change the current viewport. In case the user suddenly changes the viewing direction somewhere on the sphere, it is desirable to make available to the client at least some low-quality representations of areas not covered by the current viewport and prepare the user for rendering. However, the high-quality representation of the omnidirectional video is only required for the current viewport being rendered for current use. Dividing the high-quality representation of the entire omnidirectional video into sub-pictures at an appropriate granularity enables such optimization. Using VVC, the two representations can be encoded as two layers independent of each other.

[0044] A typical sub-picture-based viewport-dependent 360° video delivery scheme is shown in FIG. 11. The high-resolution representation of the full video is composed of sub-pictures, while the low-resolution representation of the full video does not use sub-pictures and can be coded at less frequent random access points than the high-resolution representation. The client receives the low-resolution full video and for the high-resolution video, only receives and decodes the sub-pictures covering the current viewport.

[0045] The latest VVC draft specification also supports an improved 360° video coding scheme as shown in FIG. 12. The only difference when compared with the method shown in FIG. 11 is that inter-layer prediction (ILP) is applied to the method shown in FIG. 12.

[0046] 3.6. Parameter Set AVC, HEVC, and VVC specify parameter sets. The types of parameter sets include SPS, PPS, APS, and VPS. SPS and PPS are supported in all of AVC, HEVC, and VVC. VPS was introduced in HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC but is included in the latest VVC draft text.

[0047] SPS is designed to carry sequence-level header information, and PPS is designed to carry picture-level header information that rarely changes. In SPS and PPS, information that rarely changes for each sequence or picture does not need to be repeated, so redundant signaling of this information can be avoided. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, not only avoids the need for redundant transmission, but also improves error resilience.

[0048] VPS was introduced to carry sequence-level header information common to all layers of a multi-layer bitstream.

[0049] APS requires a significant number of bits for coding, can be shared by multiple pictures, and was introduced to carry picture-level or slice-level information where there can be a very large number of different variations in a sequence.

[0050] 3.7. Sub-picture Sub-bitstream Extraction Process In addition to Section C.7 of the latest VVC text, the sub-picture sub-bitstream extraction process of the proposed changes in Bytedance IDF P2005612001H_v0 is as follows. C.7 Sub-picture Sub-bitstream Extraction Process The input to this process is a list of the bitstream inBitstream, the target OLS index targetOlsIdx, the target maximum TemporalId value tIdTarget, and the target subpicture index values subpicIdxTarget[i] for i from 0 to NumLayersInOls[targetOLsIdx] - 1. The output of this process is the sub-bitstream outBitstream. The requirement for bitstream compliance for the input bitstream is that any output sub-bitstream that satisfies all of the following conditions is a compliant bitstream. - The output sub-bitstream is the output of the process specified in this section, with the bitstream, targetOlsIdx equal to the index into the list of OLSs specified by the VPS, tIdTarget equal to any value in the range from 0 to vps_max_sublayers_minus1, and the list subpicIdxTarget[i] for i from 0 to NumLayersInOls[targetOLsIdx] - 1 satisfying the following conditions as input. - All layers within the OLS of targetOLsIdx have the same spatial resolution, the same value for sps_num_subpics_minus1, and the same subpicture layout, and all subpictures have sps_subpic_treated_as_pic_flag[] equal to 1. - The values of subpicIdxTarget[i] for all values of i are the same and equal to a specific value within the range from 0 to sps_num_subpics_minus1. - If NumLayersInOls[targetOlsIdx] is greater than 1 and sps_num_subpics_minus1 is greater than 0, the sub-picture level information SEI message is assumed to be present in the scalable nesting SEI message with sn_ols_flag equal to 1, and NestingOlsIdx[i] is equal to targetOlsIdx for one value of i in the range from 0 to sn_num_olss_minus1. For use in multi-layer OLS, the SLI SEI message is assumed to be included in the scalable nesting SEI message and is indicated in the scalable nesting SEI message to apply to a specific OLS or all layers within a specific OLS. - The output sub-bitstream includes at least one VCL NAL unit having a nuh_layer_id equal to each of the nuh_layer_id values in the list LayerIdInOls[targetOlsIdx]. - The output sub-bitstream includes at least one VCL NAL unit having a TemporalId equal to tIdTarget. Note - The compliant bitstream includes one or more coded slice NAL units having a TemporalId equal to 0, but does not necessarily include a coded slice NAL unit having a nuh_layer_id equal to 0. - The output sub-bitstream includes at least one VCL NAL unit having a nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] for each i in the range from 0 to NumLayersInOls[targetOlsIdx] - 1 and having a sh_subpic_id equal to SubpicIdVal[subpicIdxTarget[i]]. The output sub-bitstream outBitstream is derived as follows. - The sub-bitstream extraction process specified in Annex C.6 is called with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the output of the process is assigned to outBitstream. - For each value of i in the range from 0 to NumLayersInOls[targetOLsIdx] - 1, all VCL NAL units in outBitstream that have a nuh_layer_id equal to LayerIdInOls[targetOLsIdx][i] and a sh_subpic_id not equal to SubpicIdVal[subpicIdxTarget[i]], and SEI NAL units including filler data NAL units and filler payload SEI messages related to them are deleted. - When sli_cbr_constraint_flag is equal to 0, all NAL units with a nal_unit_type equal to FD_NUT and SEI NAL units including filler payload SEI messages are deleted. - If some external means not specified in this specification are available to provide a replacement parameter set for the sub-bitstream outBitstream, all parameter sets are replaced with the replacement parameter set. - Otherwise, if a sub-picture level information SEI message exists in inBitstream, the following applies. - Rewrite the value of general_level_idc in the entry of the vps_ols_ptl_idx[targetOlsIdx] of the list of profile_tier_level() syntax structures in all referenced VPS NAL units to be equal to SubpicSetLevelIdc derived by Equation D.11 for the set of sub-pictures composed of sub-pictures with a sub-picture index equal to subpicIdx. - If the VCL HRD parameters or NAL HRD parameters exist, in the ols_hrd_parameters() syntax structure of the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] in all referenced VPS NAL units and in the ols_hrd_parameters() syntax structure in all SPS NAL units referenced by the i-th layer, rewrite the respective values of cpb_size_value_minus1[tIdTarget][j] and bit_rate_value_minus1[tIdTarget][j] of the j-th CPB so that they respectively correspond to SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx] and SubpicCpbSizeNal[SubpicSetLevelIdx][subpicIdx] derived by Equation D.6, and SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx] and SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx] derived by Equations D.8 and D.9 respectively. SubpicSetLevelIdx is derived by Equation D.11 for the subpicture having the subpicture index equal to subpicIdx, j is in the range from 0 to hrd_cpb_cnt_minus1, and i is in the range from 0 to NumLayersInOls[targetOlsIdx] - 1. - For each value of i in the range from 0 to NumLayersInOls[targetOlsIdx] - 1, the following applies. - The variable spIdx is set equal to subpicIdxTarget[i]. - For all referenced SPS NAL units having sps_ptl_dpb_hrd_params_present_flag equal to 1, rewrite the value of general_level_idc in the profile_tier_level() syntax structure to be equal to SubpicSetLevelIdc derived by Equation D.11 for the set of sub-pictures composed of sub-pictures having sub-picture index equal to spIdx. - The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are derived as follows. subpicWidthInLumaSamples = min((sps_subpic_ctu_top_left_x[spIdx]+sps_subpic_width_minus1[spIdx]+1)*CtbSizeY,pps_pic_width_in_luma_samples)-sps_subpic_ctu_top_left_x[spIdx]*CtbSizeY (C.24) subpicHeightInLumaSamples = min((sps_subpic_ctu_top_left_y[spIdx]+sps_subpic_height_minus1[spIdx]+1)*CtbSizeY,pps_pic_height_in_luma_samples)-sps_subpic_ctu_top_left_y[spIdx]*CtbSizeY (C.25) - Rewrite the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS NAL units, and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS NAL units, to be equal to subpicWidthInLumaSamples and subpicHeightInLumaSamples, respectively. - Rewrite the values of sps_num_subpics_minus1 in all referenced SPS NAL units and pps_num_subpics_minus1 in all referenced PPS NAL units to 0. - If present, rewrite the syntax elements sps_subpic_ctu_top_left_x[spIdx] and sps_subpic_ctu_top_left_y[spIdx] in all referenced SPS NAL units to 0. - In all referenced SPS NAL units, for each j not equal to spIdx, delete the syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j] and sps_subpic_id[j]. - Rewrite the syntax elements in all referenced PPSs for tile and slice signaling, and delete all tile rows, tile columns and slices not related to the subpicture having the subpicture index equal to spIdx. - The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset and subpicConfWinBottomOffset are derived as follows. subpicConfWinLeftOffset=sps_subpic_ctu_top_left_x[spIdx]==0?sps_conf_win_left_offset:0 (C.26) subpicConfWinRightOffset = (sps_subpic_ctu_top_left_x[spIdx] + sps_subpic_width_minus1[spIdx] + 1) * CtbSizeY >= sps_pic_width_max_in_luma_samples? sps_conf_win_right_offset : 0 (C.27) subpicConfWinTopOffset = sps_subpic_ctu_top_left_y[spIdx] == 0? sps_conf_win_top_offset : 0 (C.28) subpicConfWinBottomOffset = (sps_subpic_ctu_top_left_y[spIdx] + sps_subpic_height_minus1[spIdx] + 1) * CtbSizeY >= sps_pic_height_max_in_luma_samples? sps_conf_win_bottom_offset : 0 (C.29) Here, sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, sps_pic_height_max_in_luma_samples, sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in the above formulas are those of the original SPS before rewriting. - Rewrite the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in all referenced SPS NAL units, and the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset in all referenced PPS NAL units to be equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively. - The variables subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset are derived as follows. subpicScalWinLeftOffset = pps_scaling_win_left_offset - sps_subpic_ctu_top_left_x[spIdx] * CtbSizeY / SubWidthC (C.30) rightSubpicBd = (sps_subpic_ctu_top_left_x[spIdx] + sps_subpic_width_minus1[spIdx] + 1) * CtbSizeY subpicScalWinRightOffset = (rightSubpicBd >= sps_pic_width_max_in_luma_samples)? pps_scaling_win_right_offset : pps_scaling_win_right_offset - (sps_pic_width_max_in_luma_samples - rightSubpicBd) / SubWidthC (C.31) subpicScalWinTopOffset = pps_scaling_win_top_offset - sps_subpic_ctu_top_left_y[spIdx] * CtbSizeY / SubHeightC (C.32) botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] + sps_subpic_height_minus1[spIdx] + 1) * CtbSizeY subpicScalWinBotOffset = (botSubpicBd >= sps_pic_height_max_in_luma_samples)? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - (sps_pic_height_max_in_luma_samples - botSubpicBd) / SubHeightC (C.33) Here, sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, and sps_pic_height_max_in_luma_samples in the above formulas are those of the original SPS before rewriting, and the above pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset are those of the original PPS before rewriting. - Rewrite the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in all referenced PPS NAL units to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset, respectively. - When sli_cbr_constraint_flag is equal to 1, for j in the range of 0 to hrd_cpb_cnt_minus1, set cbr_flag[tIdTarget][j] in the ols_hrd_parameters() syntax structure of the j-th CPB in all referenced VPS NAL units and SPS NAL units to be equal to 1. Otherwise (when sli_cbr_constraint_flag is equal to 0), set cbr_flag[tIdTarget][j] to be equal to 0. - If outBitstream contains an SEI NAL unit that contains a scalable nesting SEI message with sn_ols_flag equal to 1 and sn_subpic_flag equal to 1 applicable to outBitstream, extract an appropriate not scalable and not nested SEI message with a payloadType equal to 1 (PT), 130 (DUI), or 132 (decoded picture hash) from the scalable nesting SEI message and place the extracted SEI message in outBitstream.

[0051] 4. Technical problems solved by the disclosed technical solution The latest design of the subpicture sub-bitstream extraction process has the following problems. 1) In a scalable nested SEI message having an sn_subpic_flag equal to 1 that is not applied to the output bitstream, the scalable nesting issued The SEI NAL unit containing the SEI message shall be removed from the output bitstream. 2) The subpicture index for identifying the subpicture sequence shall be the subpicture index of the subpicture to be extracted in a layer having multiple subpictures per picture, rather than in a layer having only one subpicture per picture. 3) For k in the range from 0 to tIdTarget - 1, the rewriting of sublayer_level_idc[k] is missing, and it is not clearly specified under what conditions the rewriting of the level information should be performed for the referenced VPS and / or the referenced SPS. 4) For k in the range from 0 to tIdTarget - 1, the rewriting of cpb_size_value_minus1[k][j] and bit_rate_value_minus1[k][j] is missing, and it is not clearly specified under what conditions the rewriting of the CPB size and bitrate information should be performed for the referenced VPS and / or the referenced SPS. 5) It is not clearly specified under what conditions the rewriting of cbr_flag[tIdTarget][j] should be performed for the referenced VPS and / or the referenced SPS. 6) Scalable nesting issued The SEI message not scalable and not nested The last step of making the SEI message has multiple problems. a. If the decoded picture hash SEI message is included in the scalable nested SEI message, the value of sn_ols_flag should be 0, but the current text of the last step assumes that sn_ols_flag is 1. b. SLI and BP SEI messages in the case of having an sn_ols_flag equal to 1 and an sn_subpic_flag equal to 1 are not covered. c. The SEI messages with sn_ols_flag equal to 0 and sn_subpic_flag equal to 1 are not covered. d. The not scalable and not nested location in the output bitstream where the resulting SEI message should be placed (within the SEI NAL unit, which SEI NAL unit it should be) is not specified. e. The original container SEI NAL unit should be removed from the output bitstream.

[0052] 5. List of solutions and embodiments To solve the above problems and the like, the methods summarized below are disclosed. The items of the solutions should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these items can be applied individually or combined in any way. 1) To solve Problem 1, in the sub-picture sub-bitstream extraction process, scalable nesting that is not applied to the output bitstream issued The SEI NAL unit containing the SEI message may be specified to be removed from the output bitstream. 2) To solve Problem 2, in the sub-picture sub-bitstream extraction process, the following may be specified. The sub-picture index for identifying the sub-picture sequence is specified as the sub-picture index of the sub-picture to be extracted in a layer having a plurality of sub-pictures per picture, rather than in a layer having only one sub-picture per picture. 3) To solve Problem 3, in the sub-picture sub-bitstream extraction process, the following may be specified. In the VPS that is referenced if it exists, and in the SPS that is referenced if NumLayersInOls[targetOLsIdx] is equal to 0, rewrite both general_level_idc and sublayer_level_idc[k] for k in the range from 0 to tIdTarget - 1 to appropriate values (e.g., the values described in this document). 4) To solve Problem 4, in the sub-picture sub-bitstream extraction process, the following may be specified. In the VPS that is referenced if it exists, and in the SPS that is referenced if NumLayersInOls[targetOLsIdx] is equal to 0, rewrite cpb_size_value_minus1[k][j] and bit_rate_value_minus1[k][j] for all values of k in the range from 0 to tIdTarget to appropriate values (e.g., the values described in this document). 5) To solve Problem 5, in the sub-picture sub-bitstream extraction process, the following may be specified. In the VPS that is referenced if it exists, and in the SPS that is referenced if NumLayersInOls[targetOLsIdx] is equal to 0, rewrite cbr_flag[tIdTarget][j] to an appropriate value (e.g., the value described in this document). 6) To solve Problem 6, under specific conditions, one or more of the following operations may be executed. a. Generate a new SEI NAL unit seiNalUnitB. b. Include seiNalUnitB in the PU that includes seiNalUnitA. c. Include seiNalUnitB immediately after seiNalUnitA in the PU that includes seiNalUnitA. d. Extract a plurality of scalable nesting issued SEI messages from the scalable nesting SEI message and include these ([[]] not scalable and not nested(As an SEI message) directly include it in seiNalUnitB. e. Delete seiNalUnitA from outBitstream. 7) In one example, the specific conditions in item 6) are as follows. outBitstream has the same set of layers as outBitstream and has a sn_subpic_flag equal to 1 applicable to a subpicture having the same set of subpictures as outBitstream, or when applicable to an OLS (when sn_ols_flag is equal to 1), the case where the SEI NAL unit seiNalUnitA includes a scalable nested SEI message. 8) In one example, in the subpicture sub-bitstream extraction process, when LayerIdInOls[targetOlsIdx] does not include all the values of nuh_layer_id in all NAL units in the bitstream and outBitstream includes the SEI NAL unit seiNalUnitA including a scalable nested SEI message, retain seiNalUnitA in the output bitstream.

[0053] 6. Embodiments The following are some exemplary embodiments of the aspects of the invention summarized above in Section 5 that are applicable to the VVC specification. The changed text is based on the latest VVC text of JVET - S0152 - v5. Most of the relevant parts added or changed are highlighted in underlined bold (or underlined), and some of the deleted parts are highlighted in [[italic double brackets]] (or [[]]). There may be some other changes that are not highlighted as they are essentially editorial changes. 6.1 First Embodiment This embodiment pertains to items 1 - 7 and their sub - items. C.7 Subpicture Sub - Bitstream Extraction Process The input to this process is the bitstream inBitstream, the target OLS index targetOlsIdx, the target maximum TemporalId value tIdTarget, and for i in the range from 0 to NumLayersInOls[targetOLsIdx] - 1 [[for each layer]] the target sub-picture index value subpicIdxTarget i in an [[array]] list . The output of this process is the sub-bitstream outBitstream. The OLS with OLS index targetOlsIdx is referred to as the target OLS. Among the layers in the target OLS, those having a referenced SPS with sps_num_subpics_minus1 greater than 0 are called multiSubpicLayers. The requirement for bitstream compliance for the input bitstream is that any output sub-bitstream that satisfies all of the following conditions is a compliant bitstream. - The output sub-bitstream is equal to the bitstream, the index into the list of OLSs specified by the VPS, targetOlsIdx, tIdTarget equal to any value in the range from 0 to vps_max_sublayers_minus1 and for i in the range from 0 to NumLayersInOls[targetOLsIdx] - 1 that satisfy the following conditions [[equal to the sub-picture index present in the input]] list subpicIdxTarget i that is as input the output of the process specified in this section. - The value of subpicIdxTarget[i] is equal to a value in the range from 0 to sps_num_subpics_minus1 such that sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] is equal to 1, and sps_num_subpics_minus1 and sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] are detected or estimated based on the SPS referred to by the layer having nuh_layer_id equal to LayerIdInOls[targetOLsIdx][i]. Note 1 - If sps_num_subpics_minus1 for the layer having nuh_layer_id equal to LayerIdInOls[targetOLsIdx][i] is equal to 0, the value of subpicIdxTarget[i] is always equal to 0. - For any two different integer values of m and n, for both layers having a nuh_layer_id equal to LayerIdInOls[targetOLsIdx][m] and LayerIdInOls[targetOLsIdx][n] respectively, when sps_num_subpics_minus1 is greater than 0, subpicIdxTarget[m] shall be equal to subpicIdxTarget[n]. - The output sub-bitstream contains at least one VCL NAL unit having an nuh_layer_id equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx]. List - The output sub-bitstream contains at least one VCL NAL unit having a TemporalId equal to tIdTarget. - The output sub-bitstream contains at least one VCL NAL unit having a TemporalId equal to tIdTarget. Note 2 - A compliant bitstream contains one or more coded slice NAL units having a TemporalId equal to 0, but does not necessarily contain a coded slice NAL unit having an nuh_layer_id equal to 0. - The output sub-bitstream contains at least one VCL NAL unit having a nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] for each i in the range from 0 to NumLayersInOls[targetOlsIdx] - 1 and having a sh_subpic_id equal to the value in SubpicIdVal[subpicIdxTarget[i]][[inside value]]. The output sub-bitstream outBitstream is [[derived as follows]] In the following order of steps derived. 1. The sub-bitstream extraction process specified in Annex C.6 is called with inBitstream, targetOlsIdx, and tIdTarget as inputs, and the output of the process is assigned to outBitstream. 2. For each value of i in the range from 0 to NumLayersInOls[targetOLsIdx] - 1, delete from outBitstream all VCL NAL units having a nuh_layer_id equal to LayerIdInOls[targetOLsIdx][i] and a sh_subpic_id not equal to SubpicIdVal[subpicIdxTarget[i]], and the related SEI NAL units including filler data NAL units and filler payload SEI messages related to these. 3. When the sli_cbr_constraint_flag of the SLI SEI message applied to the target OLS is equal to 0, delete all NAL units having a nal_unit_type equal to FD_NUT and the SEI NAL units including filler payload SEI messages. 4. Delete from outBitstream all SEI NAL units including scalable nesting SEI messages having a sn_subpic_flag equal to 1, and none of the sn_subpic_idx[j] values for j from 0 to sn_num_subpics_minus1 shall be equal to any of the subpicIdxTarget[i] values for the layers within multiSubpicLayers. 5. If any external means not specified in this specification is available to provide a replacement parameter set for the sub-bitstream outBitstream, all parameter sets are replaced with the replacement parameter set. Otherwise, if an SLI SEI message exists in inBitstream, the following The steps in the order of [[applies]]. a. The variable [[subpicIdx]] spIdx is For any layer within multiSubpicLayers set equal to the value of [[subpicIdxTarget[[NumLayersInOls[targetOlsIdx]-1]]]] subpicIdxTarget[i] . b. If present, within the list of profile_tier_level() syntax structures in all referenced VPS NAL units, and also within the profile_tier_level() syntax structure in all referenced SPS NAL units if NumLayersInOls[targetOLsIdx] is equal to 0, for k within the range of 0 to tIdTarget - 1 in the entry of vps_ols_ptl_idx[targetOlsIdx] The value of general_level_idc and sublayer_level_idc[k] is , for each subpicture sequence of spIdx rewritten as derived by Equation D.10 for the [[set of sub-pictures composed of sub-pictures having a sub-picture index equal to subpicIdx]] be made equal to SubpicLevelIdc[spIdx][tIdTarget] and SubpicLevelIdc[spIdx][k], respectively . c. For k within the range of 0 to tIdTarget, set spLvIdx equal to SubpicLevelIdx[spIdx][k], where SubpicLevelIdx[spIdx][k] is derived by Equation D.10 for the subpicture sequence of spIdx.If there are VCL HRD parameters or NAL HRD parameters, For k within the range of 0 to tIdTarget, if present the ols_hrd_parameters() syntax structure of the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] in all referenced VPS NAL units, and when NumLayersInOls[targetOLsIdx] is equal to 0 all referenced the ols_hrd_parameters() syntax structure in the SPS NAL unit k [[tIdTarget]]][j] and bit_rate_value_minus1 k [[tIdTarget]]][j], respectively, are rewritten to correspond to SubpicCpbSizeVcl spLvIdx [[SubpicLevelIdx]]] spIdx [[subpicIdx]]] [k] and SubpicCpbSizeNal spLvIdx [[SubpicLevelIdx]]] spIdx [[subpicIdx]]] [k] derived by equations D.6 and D.8 and D.9, respectively, and SubpicBitrateVcl spLvIdx [[SubpicLevelIdx]]] spIdx [[subpicIdx]]] [k] and SubpicBitrateNal spLvIdx [[SubpicLevelIdx]]] spIdx [[subpicIdx]]] [k] where SubpicSetLevelIdx is derived by equation D.10 for the subpicture with a subpicture index equal to subpicIdx, j is in the range from 0 to hrd_cpb_cnt_minus1, and i is in the range from 0 to NumLayersInOls[targetOlsIdx]-1. d. For the i-th layer having an i within the range of 0 to NumLayersInOls[targetOlsIdx] - 1 each within multiSubpicLayers the following For the rewriting of the SPS and PPS referenced by the pictures within that layer applies. the steps in the order of [[is applied.]] i. The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are derived as follows. subpicWidthInLumaSamples = min((sps_subpic_ctu_top_left_x spIdx [[subpicIdx]]] + sps_subpic_width_minus1 spIdx [[subpicIdx]]] + 1) * CtbSizeY, pps_pic_width_in_luma_samples) - sps_subpic_ctu_top_left_x s spIdx [[subpicIdx]]] * CtbSizeY (C.24) subpicHeightInLumaSamples = min((sps_subpic_ctu_top_left_y spIdx [[subpicIdx]]] + sps_subpic_height_minus1 spIdx [[subpicIdx]]] + 1) * CtbSizeY, pps_pic_height_in_luma_samples) - sps_subpic_ctu_top_left_y spIdx [[subpicIdx]]] * CtbSizeY (C.25) ii. Rewrite the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS NAL units, and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS NAL units, to be equal to subpicWidthInLumaSamples and subpicHeightInLumaSamples, respectively. iii. Rewrite the values of sps_num_subpics_minus1 in all referenced SPS NAL units and pps_num_subpics_minus1 in all referenced PPS NAL units to 0. iv. If present, rewrite the syntax elements sps_subpic_ctu_top_left_x spIdx [[subpicIdx]]] and sps_subpic_ctu_top_left_y spIdx [[subpicIdx]]] in all referenced SPS NAL units to 0. v. In all referenced SPS NAL units, spIdx for each j not equal to [[subpicIdx]], delete the syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j] and sps_subpic_id[j]. vi. Rewrite the syntax elements in all referenced PPSs for tile and slice signaling, spIdx and delete all tile rows, tile columns and slices not related to the subpicture having the subpicture index equal to [[subpicIdx]]. vii. The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset and subpicConfWinBottomOffset are derived as follows. subpicConfWinLeftOffset = sps_subpic_ctu_top_left_x spIdx [[subpicIdx]]]==0?sps_conf_win_left_offset:0 (C.26) subpicConfWinRightOffset=(sps_subpic_ctu_top_left_x spIdx [[subpicIdx]]]+sps_subpic_width_minus1 spIdx [[subpicIdx]]]+1)*CtbSizeY>=sps_pic_width_max_in_luma_samples?sps_conf_win_right_offset:0 (C.27) subpicConfWinTopOffset=sps_subpic_ctu_top_left_y spIdx [[subpicIdx]]==0?sps_conf_win_top_offset:0 (C.28) subpicConfWinBottomOffset=(sps_subpic_ctu_top_left_y spIdx [[subpicIdx]]]+sps_subpic_height_minus1 spIdx [[subpicIdx]]]+1)*CtbSizeY>=sps_pic_height_max_in_luma_samples?sps_conf_win_bottom_offset:0 (C.29) Here, sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, sps_pic_height_max_in_luma_samples, sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in the above formula are those of the original SPS before rewriting. Note 3 - For pictures within a layer in multiSubpicLayers in both the input bitstream and the output bitstream, the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are equal to pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples, respectively. Therefore, in the above formula, sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples can be replaced by pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples, respectively. viii. Rewrite the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in all referenced SPS NAL units and the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset in all referenced PPS NAL units to be equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively. ix. The variables subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset are derived as follows. subpicScalWinLeftOffset = pps_scaling_win_left_offset - sps_subpic_ctu_top_left_x[spIdx] * CtbSizeY / SubWidthC (C.30) rightSubpicBd = (sps_subpic_ctu_top_left_x[spIdx] + sps_subpic_width_minus1[spIdx] + 1) * CtbSizeY subpicScalWinRightOffset = (rightSubpicBd >= sps_pic_width_max_in_luma_samples)? pps_scaling_win_right_offset : pps_scaling_win_right_offset - (sps_pic_width_max_in_luma_samples - rightSubpicBd) / SubWidthC (C.31) subpicScalWinTopOffset = pps_scaling_win_top_offset - sps_subpic_ctu_top_left_y[spIdx] * CtbSizeY / SubHeightC (C.32) botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] + sps_subpic_height_minus1[spIdx] + 1) * CtbSizeY subpicScalWinBotOffset = (botSubpicBd >= sps_pic_height_max_in_luma_samples)? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - (sps_pic_height_max_in_luma_samples - botSubpicBd) / SubHeightC (C.33) Here, sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, and sps_pic_height_max_in_luma_samples in the above formula are those of the original SPS before rewriting, and pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset are those of the original PPS before rewriting. x. Rewrite the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in all referenced PPS NAL units to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset, respectively. [[Delete all VCL NAL units from the -outBitstream that have an nuh_layer_id equal to the nuh_layer_id of the i-th layer and a sh_subpic_id not equal to SubpicIdVal[subpicIdx].]] - When sli_cbr_constraint_flag is equal to 1 Case [[delete all NAL units with a nal_unit_type equal to FD_NUT and all filler payload SEI messages not related to the VCL NAL units of the subpictures in subpicIdTarget[]]] If it exists In all referenced VPS NAL units, [[for j in the range from 0 to hrd_cpb_cnt_minus1]]If NumLayersInOls[targetOLsIdx] is equal to 0, all referenced Within the SPS NAL unit, set cbr_flag[tIdTarget][j] of the j-th CPB in the ols_hrd_parameters() syntax structure of the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] to be equal to 1. Otherwise (when sli_cbr_constraint_flag is equal to 0), [[delete all NAL units having a nal_unit_type equal to FD_NUT and filler payload SEI messages]], and set cbr_flag[tIdTarget][j] to be equal to 0. In both cases, j is in the range of 0 or more and hrd_cpb_cnt_minus1 or less. 6. If the SEI NAL unit seiNalUnitA contains a scalable nested SEI message with a sn_subpic_flag equal to 1 applicable to a subpicture having the same set of layers as outBitstream and the same set of subpictures as outBitstream (when sn_ols_flag is equal to 0) or OLS (when sn_ols_flag is equal to 1), generate a new SEI NAL unit seiNalUnitB, include it immediately after seiNalUnitA in the PU containing seiNalUnitA, extract a plurality of scalable nested SEI messages from the scalable nested SEI message, directly include them in seiNalUnitB (as non-scalable nested SEI messages), and delete seiNalUnitA from outBitstream. [[If the outBitstream contains a SEI NAL unit that contains a scalable nested SEI message having a sn_ols_flag equal to 1 and a sn_subpic_flag equal to 1 applicable to the outBitstream, extract an appropriate SEI message having a payloadType equal to 1 (PT), 130 (DUI), or 132 (decoded picture hash) from the scalable nested SEI message, and place the extracted SEI message in the outBitstream.]] non-scalable nested

[0054] ​FIG. 5 is a block diagram showing an exemplary video processing system 1900 in which various techniques disclosed in this specification may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received, for example, in a raw or uncompressed format of 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0055] System 1900 may include a coding component 1904 that may implement various coding or encoding methods described in this document. Coding component 1904 may reduce the average bitrate of the video from input 1902 to the output of coding component 1904 and generate a coded representation of the video. Thus, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of coding component 1904 may be stored or transmitted via a connection represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at input 1902 may be used by component 1908 to generate pixel values or displayable video to be transmitted to display interface 1910. The process of generating user-viewable video from the bitstream is sometimes referred to as video decompression. Further, certain video processing operations are referred to as "coding" operations or tools, but it is recognized that coding tools or operations are used in an encoder and corresponding decoding tools or operations that are the reverse of the coding result are executed in a decoder.

[0056] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI (registered trademark)), a DisplayPort, etc. Examples of a storage interface include a Serial Advanced Technology Attachment (SATA), a PCI, an IDE interface, etc. The technology described in this document may be embodied in various electronic devices such as a mobile phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.

[0057] FIG. 6 is a block diagram of a video processing apparatus 3600. The apparatus 3600 may be used to implement one or more of the methods described herein. The apparatus 3600 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver. The apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor 3602 may be configured to execute one or more of the methods described in this document. The memory 3604 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to implement some of the techniques described in this document in hardware circuits.

[0058] FIG. 8 is a block diagram showing an exemplary video coding system 100 that may utilize the technology of this disclosure.

[0059] As shown in FIG. 8, video coding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0060] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0061] The video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that forms a coding representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coding representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 120 through the network 130a via the I / O interface 116. The encoded video data may also be stored in the storage medium / server 130b for access by the destination device 120.

[0062] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0063] The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120 configured to interface with an external display device.

[0064] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0065] FIG. 9 is a block diagram showing an example of a video encoder 200 that may be the video encoder 114 within the system 100 shown in FIG. 8.

[0066] The video encoder 200 may be configured to execute any or all of the techniques of this disclosure. In the example of FIG. 9, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 200. In some examples, the processor may be configured to execute any or all of the techniques described in this disclosure.

[0067] The functional components of the video encoder 200 may include a partitioning unit 201, a prediction unit 202 that may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a conversion unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse conversion unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.

[0068] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is the picture in which the current video block is located.

[0069] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 are shown separately for illustrative purposes in the example of FIG. 9, but may be highly integrated.

[0070] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0071] The mode selection unit 203 may select, for example, one of coding modes such as intra or inter based on the error result, and provide the resulting intra or inter coding block to a residual generation unit 207 for generating residual block data and a reconstruction unit 212 for reconstructing the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter prediction (CIIP) mode where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel accuracy) in the case of inter prediction.

[0072] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information of the pictures from the buffer 213 and the decoded samples other than the picture related to the current video block.

[0073] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations for the current video block depending on, for example, whether the current video block is in an I slice, a P slice, or a B slice.

[0074] In some examples, the motion estimation unit 204 may perform one - direction prediction for the current video block. The motion estimation unit 204 may also obtain a reference video block for the current video block and search for a reference picture in list 0 or list 1. Then, the motion estimation unit 204 may generate a reference index indicating the reference picture in list 0 or list 1 that includes the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0075] In other examples, the motion estimation unit 204 may perform bi - direction prediction for the current video block. The motion estimation unit 204 may obtain a reference video block for the current video block and search for a reference picture in list 0, and may also obtain another reference video block for the current video block and search for a reference picture in list 1. Then, the motion estimation unit 204 may generate a reference index indicating the reference pictures in lists 0 and 1 that include the reference video blocks, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0076] In some examples, the motion estimation unit 204 may output a complete set of motion information for decoder decoding processing.

[0077] In some examples, the motion estimation unit 204 may not output a complete set of motion information for the current video. Rather, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of other video blocks. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of an adjacent video block.

[0078] In one example, the motion estimation unit 204 may indicate, in the syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0079] In other examples, the motion estimation unit 204 may identify, in the syntax structure associated with the current video block, an other video block and a motion vector difference (MVD). The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0080] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0081] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks within the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0082] The residual generation unit 207 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block (e.g., indicated by a negative sign). The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples within the current video block.

[0083] In other examples, for example, in skip mode, there may be no residual data for the current video block for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0084] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0085] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0086] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block related to the current block stored in the buffer 213.

[0087] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts within the video block.

[0088] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0089] FIG. 10 is a block diagram showing an example of a video decoder 300 that may be the video decoder 114 within the system 100 shown in FIG. 8.

[0090] The video decoder 300 may be configured to execute any or all of the techniques of this disclosure. In the example of FIG. 10, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure may be shared among various components of the video decoder 300. In some examples, the processor may be configured to execute any or all of the techniques described in this disclosure.

[0091] In the example of FIG. 10, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 may perform a decoding path that is generally inverse to the encoding path described with respect to video encoder 200 (FIG. 9).

[0092] Entropy decoding unit 305 may obtain an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., an encoded block of video data). Entropy decoding unit 301 may decode the entropy-coded video data, and from the entropy-decoded video data, motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 may determine such information, for example, by performing AMVP and merge mode.

[0093] Motion compensation unit 302 may perform interpolation based on an interpolation filter in some cases to generate a motion-compensated block. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.

[0094] Motion compensation unit 302 may calculate interpolation values for sub-integer pixels of a reference block using the interpolation filter used by video encoder 200 during encoding of a video block. Motion compensation unit 302 may determine the interpolation filter used by video encoder 200 according to the received syntax information and use the interpolation filter to generate a prediction block.

[0095] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, the partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-encoded block, and other information for decoding the encoded video sequence.

[0096] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 303 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0097] The reconstruction unit 306 may add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and generates the decoded video for presentation on a display device.

[0098] Next, a list of solutions selected by some embodiments is provided.

[0099] The following solutions illustrate exemplary embodiments of the techniques discussed in the previous section (e.g., items 1-8).

[0100] 1. A video processing method (e.g., method 700 shown in FIG. 7), comprising A method including a step (702) of performing conversion between a video including one or more sub-pictures and a coded representation of the video, the coded representation being composed of one or more Network Abstraction Layer (NAL) units, and the conversion following rules that specify a sub-picture sub-bitstream extraction process in which a sub-bitstream of the sub-picture is constructed or extracted.

[0101] The following solutions show exemplary embodiments of the techniques described in the above section (e.g., item 1).

[0102] 2. The rules are scalable nesting not applicable to the output bitstream performed The method described in Solution 1, which specifies that a Supplemental Enhancement Information (SEI) NAL unit including an SEI message is removed from the output bitstream.

[0103] The following solutions show exemplary embodiments of the techniques described in the above section (e.g., item 2).

[0104] 3. The rules are as described in any of Solutions 1 to 2, which specify that the sub-picture index for identifying a sub-picture sequence corresponds to the sub-picture index of the sub-picture to be extracted in a video layer including a plurality of sub-pictures per picture.

[0105] The following solutions show exemplary embodiments of the techniques described in the above section (e.g., item 3).

[0106] 4. The rules are as described in any of Solutions 1 to 3, which specify that when the number of layers in the output layer set is 1, the first syntax element indicating the general level and the second syntax element indicating the layer level are rewritten to other values.

[0107] The following solutions show exemplary embodiments of the techniques described in the above section (e.g., item 4).

[0108] 5. The method according to any one of Solutions 1 to 4, wherein the rule specifies rewriting a first syntax element indicating the size of the coding picture buffer and a second syntax element indicating the bit rate to other values (for example, appropriate values described in this specification) when the number of layers in the output layer set is 1.

[0109] The following solutions show exemplary embodiments of the techniques described in the above sections (for example, items 5 to 8).

[0110] 6. The method according to any one of Solutions 1 to 5, wherein the rule specifies rewriting the value of a syntax field indicating the coding bit rate to another value (for example, an appropriate value described in this specification) in a reference video parameter set or a sequence parameter set.

[0111] 7. The method according to any one of Solutions 1 to 6, wherein the conversion includes encoding the video into a coded representation.

[0112] 8. The method according to any one of Solutions 1 to 6, wherein the conversion includes decoding the coded representation to generate pixel values of the video.

[0113] 9. A video decoding apparatus including a processor configured to implement the method according to one or more of Solutions 1 to 8.

[0114] 10. A video encoding apparatus including a processor configured to implement the method according to one or more of Solutions 1 to 8.

[0115] 11. A computer program product storing computer code, which, when executed by a processor, causes the processor to implement the method according to any one of Solutions 1 to 8. 12. The method, apparatus, or system described in this document.

[0116] 12. The method, apparatus, or system described in this document.

[0117] In the solution described in this specification, the encoder may generate a coding representation according to formatting rules, or may also be in accordance with formatting rules. In the solution described in this specification, the decoder may use formatting rules to recognize the presence or absence of syntax elements according to formatting rules, analyze the syntax elements in the coding representation, and generate a decoded video.

[0118] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation or vice versa. The current bitstream representation of a video block may correspond to bits that are at the same position or are scattered at different positions within the bitstream, as defined by the syntax, for example. For example, a macroblock may be encoded with respect to transformed and coded error residual values and using bits in headers and other fields within the bitstream. Further, during the conversion, the decoder may recognize the presence or absence of some fields and analyze the bitstream based on a decision, as described in the above solution. Similarly, the encoder may determine whether a particular syntax field should be included and, accordingly, generate a coding representation by including or excluding the syntax field from the coding representation.

[0119] In some preferred embodiments, the first set of solutions may be implemented as further described in items 1 and 2 of Section 5.

[0120] 1. A method of processing video data (e.g., the method shown in FIG. 13A), comprising: Performing a step (1302) of converting between a video and a bitstream of the video, the bitstream including a plurality of layers including one or more sub-pictures according to rules, the rules being scalable nesting not applicable to the output bitstream during a sub-picture sub-bitstream extraction process from which the output bitstream is extracted from the bitstream performed A method of specifying that a supplementary enhancement information network abstraction layer unit (SEI NAL unit) including an SEI message is omitted in the output bitstream.

[0121] 2. The output bitstream includes one or more output layers including sub-pictures identified by one or more target sub-picture indexes, and the SEI NAL unit is scalable nesting performed The method according to Solution 1, wherein the SEI includes a flag having a predetermined value, and is considered not applicable to the output bitstream in response to the first one or more sub-picture indexes of the SEI NAL unit not matching the second one or more sub-picture indexes in the output bitstream.

[0122] 3. The method according to Solution 2, wherein the flag having a predetermined value indicates that the SEI NAL unit is applicable to a specific sub-picture of a specific layer.

[0123] 4. The method according to Solutions 2 to 3, wherein the flag is sn_subpic_flag.

[0124] 5. The method according to any one of Solutions 2 to 4, wherein the first one or more sub-picture indexes are sn_subpic_idx[j], j is an integer having a value from 0 to sn_num_subpics_minus1, and the second one or more sub-picture indexes are subpicIdxTarget[i] for a layer in the output bitstream having a plurality of sub-pictures per picture, and i is an integer.

[0125] 6. A method for processing video data (e.g., method 1310 shown in FIG. 13), comprising: performing a conversion between a video and a bitstream of the video, the bitstream including a plurality of layers each including one or more pictures including one or more sub-pictures according to rules, the rules specifying that a first sub-picture index identifying a sub-picture sequence extracted by a sub-picture sub-bitstream extraction process for the bitstream is based on a second sub-picture index of a layer of the bitstream having a plurality of sub-pictures per picture.

[0126] 7. The method according to solution 6, wherein the rules specify that the first sub-picture index is in accordance with the sub-picture level information supplementary enhancement information (SLI SEI) message being included in the bitstream.

[0127] In some preferred embodiments, the second set of solutions may be implemented as further described in items 3, 4, and 5 of Section 5.

[0128] 1. A method for processing video data (e.g., method 1320 shown in FIG. 13C), comprising: performing a conversion (1322) between a video and a bitstream of the video, the bitstream including one or more layers each including one or more sub-layers each including one or more pictures including one or more sub-pictures according to rules, the rules specifying a rewrite operation selectively performed on one or more syntax structures during a sub-picture sub-bitstream extraction process from which an output target sub-bitstream is extracted from the bitstream, and one or more syntax elements including information of the output target sub-bitstream.

[0129] 2. The one or more syntax structures include: (a) a first syntax structure indicating the coding level to which the output target sub-bitstream conforms; and (b) a second syntax structure indicating the coding level to which the sublayer sequence within the output target sub-bitstream having index values from 0 to tIdTarget-1 conforms, where tIdTarget is an integer representing the highest temporal layer identifier of the sublayers within the output target sub-bitstream, the method according to Solution 1.

[0130] 3. The one or more syntax elements include: (a) a first syntax structure indicating the coding picture buffer size of each sublayer sequence within the output target sub-bitstream; and (b) a second syntax structure indicating the bitrate value of each sublayer sequence within the output target sub-bitstream, the method according to Solution 1.

[0131] 4. The one or more syntax elements include a first syntax structure indicating whether each sublayer sequence within the output target sub-bitstream is to be processed as having a constant bitrate, the method according to Solution 1.

[0132] 5. The first syntax structure and the second syntax structure are included in the video parameter set referred to by the output target sub-bitstream, the method according to Solutions 1 to 2.

[0133] 6. When the output target sub-bitstream includes a single layer, the first syntax structure and the second syntax structure are included in the sequence parameter set referred to by the output target sub-bitstream, the method according to Solutions 1 to 2.

[0134] In some preferred embodiments, the third set of solutions may be implemented as further described in items 6, 7, and 8 of Section 5.

[0135] 1. A method for processing video data (e.g., method 1330 shown in FIG. 13D), comprising: performing a conversion between a video and a bitstream of the video (1332), the bitstream including one or more layers each including one or more pictures each including one or more sub - pictures according to rules, the rules specifying a selection process of a first supplementary enhancement information network abstraction layer (SEI NAL) unit of a target output sub - picture sub - bitstream extracted during a sub - picture sub - bitstream extraction process according to conditions.

[0136] 2. The method according to solution 1, wherein the processing includes generating a first SEI NAL unit.

[0137] 3. The method according to any one of solutions 1 - 2, wherein the processing includes adding the first SEI NAL unit to a picture unit including a second SEI NAL unit.

[0138] 4. The method according to any one of solutions 1 - 3, wherein the processing includes adding the first SEI NAL unit immediately after the second SEI NAL unit to a picture unit including the second SEI NAL unit.

[0139] 5. The method according to any one of solutions 1 - 4, wherein the processing includes extracting a scalable nesting SEI message from a scalable nesting SEI message in the second SEI NAL unit and including the extracted scalable nesting SEI message as the first SEI NAL unit as an SEI message. performed extracting a scalable nesting SEI message from a scalable nesting SEI message in the second SEI NAL unit and including the extracted scalable nesting SEI message as the first SEI NAL unit as an SEI message. performed SEI message into non-scalable nested SEI message into the first SEI NAL unit as an SEI message.

[0140] 6. The method according to any one of solutions 1 - 5, wherein the processing includes deleting the second SEI NAL unit from the target output sub - picture sub - bitstream.

[0141] 7. The condition includes that (a) the target output sub-picture sub-bitstream includes a second SEI NAL unit including a scalable nesting SEI message, (b) the syntax field in the scalable nesting SEI message is set to a value indicating that the scalable nesting SEI message is applicable to the same set of layers in the target output sub-picture sub-bitstream, and (c) the scalable nesting SEI message is applicable to the same set of sub-pictures in the target output sub-picture sub-bitstream, and the method according to any one of Solutions 1 to 6.

[0142] 8. The condition includes that (a) the list of layers in the target output sub-picture sub-bitstream does not include all the layers in the bitstream, and (b) the target output sub-picture sub-bitstream includes a scalable nesting SEI message, and the process includes holding the first SEI NAL unit in the target output sub-picture sub-bitstream without modification, and the method according to Solution 1.

[0143] 9. The first SEI NAL unit is seiNalUnitB, and the method according to any one of Solutions 1 to 8.

[0144] 10. The second SEI NAL unit is seiNalUnitA, and the method according to any one of Solutions 1 to 9.

[0145] Referring to the above first, second, and third sets of solutions, in some embodiments, the video includes 360-degree video. In some embodiments, the conversion includes encoding the video into a bitstream. In some examples, the conversion includes decoding the bitstream to generate pixel values of the video.

[0146] Some embodiments may include a video decoding device including a processor configured to implement the method described in the solutions of the first, second, or third list.

[0147] Some embodiments may include a video encoding device including a processor configured to implement the method described in one or more of Solutions 1-8.

[0148] In some embodiments, a method of storing a bitstream representing video on a computer-readable recording medium may be implemented. The method includes generating a bitstream from the video according to the method described in any one or more of the above solutions, and storing the bitstream on a computer-readable recording medium.

[0149] Some embodiments may include a computer-readable medium storing a bitstream generated according to any one or more of the above solutions.

[0150] Some embodiments may include a computer program product storing computer code, which, when executed by a processor, causes the processor to implement the method described in any one of the above solutions.

[0151] The disclosures and other solutions, examples, modules and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware or hardware, including the structures disclosed in this document and their structural equivalents or combinations of one or more of them. The disclosures and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that generates a machine-readable propagated signal, or a combination of one or more of these. The term "data processing apparatus" includes, by way of example, all apparatus, devices and machines for processing data, including programmable processors, computers or multiple processors or computers. The apparatus can include, in addition to hardware, code for creating an execution environment for the computer program in question, e.g., processor firmware, protocol stack, database management system, operating system or code that constitutes a combination of one or more of these. A propagated signal is an artificially generated signal, e.g., an electrical, optical or electromagnetic signal generated by a machine, generated to encode information for transmission to an appropriate receiver device.

[0152] A computer program (also known as a program, software, software application, script, or code) includes compiled or interpreted languages, can be written in any form of programming language, and can be deployed in any form as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), or in a single file dedicated to the program in question, or in multiple related files (e.g., files that hold one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer, or can be deployed to be executed on multiple computers arranged in one location or distributed across multiple locations and interconnected by a communication network.

[0153] The processes and logical flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data to generate output. The processes and logical flows can also be executed by dedicated logic circuits, such as an FPGA (field programmable gate array) or ASIC (application specific integrated circuit), and the apparatus can also be implemented as dedicated logic circuits.

[0154] Processors suitable for the execution of a computer program include, by way of example, any one or more processors of both general and special purpose microprocessors, and any one of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as, magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from or transfer data to a mass storage device. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, such as, EPROM, EEPROM, and flash memory devices, including all forms of non-volatile memory, media, and memory devices. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0155] Although this patent document contains many details, these should not be construed as limitations on the scope of any object or of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. The particular features described in this patent document with respect to separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described with respect to a single embodiment may be implemented separately in multiple embodiments or in any suitable subcombination. Further, features that have been described and even initially claimed above as acting in a particular combination may in some cases have one or more features excluded from the claimed combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0156] Similarly, although the operations are shown in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order or sequence shown, or that all illustrated operations be performed, to achieve the desired result. Further, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0157] Only several implementations and examples are described, and other implementations, extensions, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing video data, comprising: executing a conversion between a video and a bitstream of the video, wherein the bitstream includes a plurality of layers including one or more sub-pictures according to rules, the rules specify that a scalable nested SEI message included in supplementary enhancement information network abstraction layer unit (SEI NAL unit) that is not applicable to the output bitstream is deleted from the output bitstream during a sub-picture sub-bitstream extraction process in which the output bitstream is extracted from the bitstream, the output bitstream includes one or more output layers including sub-pictures identified by one or more target sub-picture indexes, the SEI NAL unit has a scalable nested SEI message having a flag with a predetermined value, and is considered not applicable to the output bitstream in response to that a first one or more sub-picture identifiers of the SEI NAL unit do not match second one or more sub-picture identifiers in the output bitstream, the rules further specify a rewriting operation selectively executed for one or more syntax elements during the sub-picture sub-bitstream extraction process, the one or more syntax elements include information of the output bitstream, the one or more syntax elements include: (a) a first syntax element indicating a coding picture buffer (CPB) size of each sub-layer sequence in the output bitstream, and (b) a second syntax element indicating a bit rate value of each sub-layer sequence in the output bitstream, each sub-layer sequence has an index value k, k ranges from 0 or more to tIdTarget or less, and tIdTarget is an integer representing a target maximum temporal layer identifier of sub-layers in the output bitstream.

2. The method according to claim 1, wherein the flag with the predetermined value indicates that a scalable nested SEI message in the scalable nested SEI message is applicable to a specific sub-picture of a specific layer.

3. The method according to claim 1 or 2, wherein the flag having the predetermined value is an sn_subpic_flag equal to 1.

4. The method according to any one of claims 1 to 3, wherein the rule further specifies that one or more sub-picture indexes for identifying a sub-picture sequence extracted by the sub-picture sub-bitstream extraction process are based on the one or more target sub-picture indexes of a layer having a plurality of sub-pictures per picture.

5. The method according to claim 4, wherein the layer having a plurality of sub-pictures per picture is a layer having a syntax element sps_num_subpics_minus1 whose referenced sequence parameter set is greater than 0.

6. The method according to any one of claims 1 to 5, wherein the video includes a 360-degree video.

7. The method according to any one of claims 1 to 6, wherein the conversion includes encoding the video into the bitstream.

8. The method according to any one of claims 1 to 6, wherein the conversion includes decoding the bitstream to generate pixel values of the video.

9. An apparatus for processing video data, including a processor and a non-transitory memory having instructions, wherein when the instructions are executed by the processor, the processor is caused to perform a conversion between a video and a bitstream of the video, the bitstream includes a plurality of layers including one or more sub-pictures according to a rule, the rule specifies that a scalable nesting SEI message that is not applicable to the output bitstream is removed from the output bitstream during a sub-picture sub-bitstream extraction process in which the output bitstream is extracted from the bitstream, and the supplementary enhancement information network abstraction layer unit (SEI NAL unit) includes the scalable nesting SEI message, the output bitstream includes one or more output layers including sub-pictures identified by one or more target sub-picture indexes. The SEI NAL unit has a scalable nested SEI message having a flag with a predetermined value, and is considered inapplicable to the output bitstream in response to the first one or more sub-picture identifiers of the SEI NAL unit not matching the second one or more sub-picture identifiers in the output bitstream. The rule further specifies a rewriting operation that is selectively executed on one or more syntax elements during the sub-picture sub-bitstream extraction process. The one or more syntax elements include information of the output bitstream. The one or more syntax elements include (a) a first syntax element indicating the coding picture buffer (CPB) size of each sub-layer sequence in the output bitstream, and (b) a second syntax element indicating the bit rate value of each sub-layer sequence in the output bitstream. Each sub-layer sequence has an index value k, where k ranges from 0 to tIdTarget, and tIdTarget is an integer representing the target highest temporal layer identifier of the sub-layers in the output bitstream. Claim 10 A non-transitory computer-readable storage medium storing instructions, wherein the instructions cause a processor to perform a conversion between video and a bitstream of the video, the bitstream includes a plurality of layers including one or more sub-pictures according to a rule, the rule specifies that a supplementary enhancement information network abstraction layer unit (SEI NAL unit) including a scalable nested SEI message inapplicable to the output bitstream is removed from the output bitstream during a sub-picture sub-bitstream extraction process in which the output bitstream is extracted from the bitstream, the output bitstream includes one or more output layers including sub-pictures identified by one or more target sub-picture indexes. The SEI NAL unit has a scalable nested SEI message having a flag with a predetermined value, and is considered inapplicable to the output bitstream in response to the first one or more sub-picture identifiers of the SEI NAL unit not matching the second one or more sub-picture identifiers in the output bitstream. The rule further specifies a rewriting operation that is selectively executed on one or more syntax elements during the sub-picture sub-bitstream extraction process. The one or more syntax elements include information of the output bitstream. The one or more syntax elements include (a) a first syntax element indicating the coding picture buffer (CPB) size of each sub-layer sequence in the output bitstream, and (b) a second syntax element indicating the bit rate value of each sub-layer sequence in the output bitstream. Each sub-layer sequence has an index value k, where k ranges from 0 to tIdTarget, and tIdTarget is an integer representing the target highest time layer identifier of the sub-layer in the output bitstream, a non-transitory computer-readable storage medium. Claim 11 A method for storing a bitstream of a video, the method comprising: generating the bitstream of the video; storing the bitstream in a non-transitory computer-readable recording medium. The bitstream includes a plurality of layers including one or more sub-pictures according to a rule, wherein the rule specifies that a supplementary enhancement information network abstraction layer unit (SEI NAL unit) including a scalable nested SEI message inapplicable to the output bitstream is removed from the output bitstream during a sub-picture sub-bitstream extraction process in which the output bitstream is extracted from the bitstream, and the output bitstream includes one or more output layers including sub-pictures identified by one or more target sub-picture indexes. ​ ​ The SEI NAL unit is regarded as not applicable to the output bitstream in response to that the scalable nested SEI message having a flag with a predetermined value and that the first one or more sub-picture identifiers of the SEI NAL unit do not match the second one or more sub-picture identifiers in the output bitstream. The rule further specifies a rewriting operation that is selectively executed on one or more syntax elements during the sub-picture sub-bitstream extraction process. The one or more syntax elements include information of the output bitstream. The one or more syntax elements include: (a) a first syntax element indicating the coding picture buffer (CPB) size of each sublayer sequence in the output bitstream, and (b) a second syntax element indicating the bitrate value of each sublayer sequence in the output bitstream, where each sublayer sequence has an index value k, k ranges from 0 to tIdTarget, and tIdTarget is an integer representing the target highest temporal layer identifier of the sublayers in the output bitstream.

Citation Information

Patent Citations

  • JPP7511681B