Scaling window in sub-picture sub-bitstream extraction processing
By specifying scaling window parameters in the video bitstream, the problem of inconsistent parameters in sub-image sub-bitstream extraction is solved, improving encoding and decoding efficiency and quality. It is suitable for multi-layer video encoding and decoding standards and 360° video transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-21
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies have a problem in effectively handling scaling window parameters when extracting sub-image sub-bitstreams from video bitstreams, especially when extracting rectangular regions that cover one or more sub-images. This leads to inconsistent scaling window parameters, affecting encoding and decoding efficiency and quality.
By specifying scaling window parameters according to rules in the video processing method, including extracting network abstraction layer units from the bitstream and determining the parameter combination applicable to sub-pictures, the output is ensured to conform to a predefined format. At the same time, scalable nested supplementary enhancement information and dependencies of encoding and decoding tools are handled, and inconsistent use of reference pictures is disabled.
It achieves consistency of scaling window parameters during sub-bitstream extraction, improving encoding and decoding efficiency and quality, and is suitable for multi-layer video encoding and decoding standards such as VVC, especially for efficient transmission of 360° video.
Smart Images

Figure CN115699729B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is the China National Phase of International Patent Application No. PCT / CN2021 / 095179, filed May 21, 2021, which claims priority to and the benefit of International Patent Application No. PCT / CN2020 / 091696, filed May 22, 2020. The entire disclosure of the foregoing applications is incorporated by reference as part of the disclosure of this application. TECHNICAL FIELD
[0003] This patent document relates to image and video coding and decoding. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, it is expected that the bandwidth demand for digital video usage will continue to grow. SUMMARY
[0005] This document discloses techniques that can be used by video encoders and decoders for processing coded representations of video using control information useful for decoding the coded representations.
[0006] In one example aspect, a method of video processing is disclosed. The method includes converting, according to a rule, between a video comprising one or more video pictures comprising one or more subpictures and a bitstream of the video, wherein the rule specifies determining, from one or more syntax elements, one or more parameters of a scaling window applicable to a subpicture during subpicture subbitstream extraction processing.
[0007] In another example aspect, another method of video processing is disclosed. The method includes converting, according to a rule, between a video comprising one or more layers comprising one or more video pictures comprising one or more subpictures and a bitstream of the video, wherein the rule defines that network abstraction layer (NAL) units are to be extracted from the bitstream during subbitstream extraction processing to output a subbitstream, and wherein the rule further specifies one or more inputs to the subbitstream extraction processing and / or which combination of subpictures of different layers of the bitstream to use such that an output of the subbitstream extraction processing conforms to a predefined format.
[0008] In another example aspect, another video processing method is disclosed. The method includes converting, according to a rule, between a video and a bitstream of the video, wherein the bitstream comprises a first layer comprising pictures having a plurality of subpictures and a second layer comprising pictures each having a single subpicture, and wherein the rule specifies a combination of subpictures of the first layer and the second layer that, when extracted, results in an output bitstream that conforms to a predefined format.
[0009] In another example aspect, another video processing method is disclosed. The method includes converting, according to a rule, between a video and a bitstream of the video, wherein the bitstream comprises a first layer comprising pictures having a plurality of subpictures and a second layer comprising pictures each having a single subpicture, and wherein the rule specifies a combination of subpictures of the first layer and the second layer that, when extracted, results in an output bitstream that conforms to a predefined format.
[0010] In another example aspect, another video processing method is disclosed. The method includes converting, according to a rule, between a video and a bitstream of the video, wherein the bitstream comprises a first layer comprising pictures having a plurality of subpictures and a second layer comprising pictures each having a single subpicture, and wherein the rule specifies a combination of subpictures of the first layer and the second layer that, when extracted, results in an output bitstream that conforms to a predefined format.
[0011] In another example aspect, another video processing method is disclosed. The method includes converting, according to a rule, between a video and a bitstream of the video, wherein the bitstream comprises a first layer comprising pictures having a plurality of subpictures and a second layer comprising pictures each having a single subpicture, and wherein the rule specifies a combination of subpictures of the first layer and the second layer that, when extracted, results in an output bitstream that conforms to a predefined format.
[0012] In another example aspect, another video processing method is disclosed. The method includes converting, according to a rule, between a video and a bitstream of the video, wherein the bitstream comprises a first layer comprising pictures having a plurality of subpictures and a second layer comprising pictures each having a single subpicture, and wherein the rule specifies a combination of subpictures of the first layer and the second layer that, when extracted, results in an output bitstream that conforms to a predefined format.
[0013] In another example aspect, another video processing method is disclosed. The method includes converting, according to a rule, between a video comprising one or more video pictures in a video layer and a bitstream of the video, wherein the rule specifies that a supplemental enhancement information (SEI) network abstraction layer (NAL) unit containing a SEI message with a particular payload type does not contain another SEI message with a payload type different from the particular payload type.
[0014] In yet another example aspect, a video encoder apparatus is disclosed. The video encoder comprises a processor configured to implement the above-described method.
[0015] In yet another example aspect, a video decoder apparatus is disclosed. The video decoder comprises a processor configured to implement the above-described method.
[0016] In yet another example aspect, a computer readable medium having code stored thereon is disclosed. The code embodies one of the methods described herein in the form of processor-executable code.
[0017] These and other features are described throughout this document. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 An example of raster-scan tile partitioning of a picture is shown, where the picture is divided into 12 tiles and 3 raster-scan slices.
[0019] Figure 2 An example of rectangular slice partitioning of a picture is shown, where the picture is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0020] Figure 3 An example of a picture partitioned into tiles and rectangular slices is shown, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0021] Figure 4 A picture partitioned into 15 tiles, 24 slices, and 24 sub-pictures is shown.
[0022] Figure 5 A typical sub-picture based viewport dependent 360° video coding scheme is shown.
[0023] Figure 6 An improved viewport dependent 360° video coding scheme based on sub-pictures and spatial scalability is shown.
[0024] Figure 7 is a block diagram of an example video processing system.
[0025] Figure 8is a block diagram of a video processing device.
[0026] Figure 9 is a flowchart of an example method for video processing.
[0027] Figure 10 is a block diagram illustrating a video coding system in accordance with some embodiments of the disclosure.
[0028] Figure 11 is a block diagram illustrating an encoder in accordance with some embodiments of the disclosure.
[0029] Figure 12 is a block diagram illustrating a decoder in accordance with some embodiments of the disclosure.
[0030] Figure 13 is shown a flowchart of an example method for video processing based on some implementations of the disclosed technology.
[0031] Figures 14A to 14C is shown a flowchart of an example method for video processing based on some implementations of the disclosed technology.
[0032] Figures 15A to 15C is shown a flowchart of an example method for video processing based on some implementations of the disclosed technology.
[0033] Figure 16 is shown a flowchart of an example method for video processing based on some implementations of the disclosed technology. DETAILED DESCRIPTION
[0034] The use of section headings in this document is for convenience only and not to be construed as limiting the applicability of the technology and embodiments disclosed in each section to the section alone. Also, the use of H.266 terminology in some descriptions is for convenience only and is not intended to limit the scope of the disclosed technology. Thus, the technology described herein is applicable to other video coding protocols and designs as well.
[0035] 1. INTRODUCTION
[0036] This document relates to video coding technology. In particular, it is about subpicture sub-bitstream extraction processing, scalable nesting SEI messages, and subpicture level information SEI messages. These ideas can be applied individually or in various combinations to any video coding standard or non-standard video codec that supports multi-layer video coding, e.g., the Versatile Video Coding (VVC) that is under development.
[0037] 2. ABBREVIATIONS
[0038] APS adaptation parameter set
[0039] AU access unit
[0040] AUD access unit delimiter
[0041] AVC advanced video coding
[0042] CLVS coded layer video sequence
[0043] CPB coded picture buffer
[0044] CRA clean random access
[0045] CTU coding tree unit
[0046] CVSC coded video sequence
[0047] DCI decoding capability information
[0048] DPB decoded picture buffer
[0049] EOB end of bitstream
[0050] EOS end of sequence
[0051] GDR gradual decoding refresh
[0052] HEVC high efficiency video coding
[0053] HRD hypothetical reference decoder
[0054] IDR instantaneous decoding refresh
[0055] ILP inter-layer prediction
[0056] ILRP inter-layer reference picture
[0057] JEM joint exploration model
[0058] LTRP long-term reference picture
[0059] MCTS motion-constrained tile set
[0060] NAL network abstraction layer
[0061] OLS output layer set
[0062] PH picture header
[0063] PPS picture parameter set
[0064] PTLP profile, tier, and level
[0065] PU picture unit
[0066] RAP random access point
[0067] RBSP raw byte sequence payload
[0068] SEI supplemental enhancement information
[0069] SPS sequence parameter set
[0070] STRP short-term reference picture
[0071] SVC scalable video coding
[0072] VCL video coding layer
[0073] VPS video parameter set
[0074] VTM VVC test model
[0075] VUI video usability information
[0076] VVC versatile video coding
[0077] 3. Initial discussion
[0078] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure with motion prediction in the temporal domain and transform coding in the spatial domain. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by the JVET and introduced into the reference software called Joint Exploration Model (JEM). The JVET meeting is held simultaneously every quarter, and the new coding standard targets 50% bit rate reduction compared to HEVC. The new video coding standard was officially named as Versatile Video Coding (VVC) in the April 2018 JVET meeting, and the first version of VVC test model (VTM) was released at that time. As there is continuous effort on VVC standardization, new coding techniques are adopted in the VVC standard in every JVET meeting. The VVC working draft and test model VTM are then updated after every meeting. The VVC project is currently targeting completion of the technical specification (FDIS) in the July 2020 meeting.
[0079] 3.1. Picture partitioning schemes in HEVC
[0080] HEVC includes four different picture partitioning schemes, namely regular slices, dependent slices, tiles, and Wavefront Parallel Processing (WPP), which can be applied for Maximum Transmission Unit (MTU) size matching, parallel processing, and reduced end-to-end delay.
[0081] Regular slices are similar to in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Thus, a regular slice can be reconstructed independently from other regular slices within the same picture (although there can still be interdependencies due to in-loop filtering operations).
[0082] Regular slices are the only tool available for parallelization, which is also available in H.264 / AVC in practically the same form. Regular slice based parallelization does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predicted pictures, which is typically much more heavy than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reason, using regular slices can incur a large coding overhead due to the bit cost of the slice header and due to the lack of prediction over slice boundaries. Furthermore, due to the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit, regular slices also serve as a key mechanism for bitstream partitioning to match MTU size requirements (in contrast to the other tools mentioned below). In many cases, the goals of parallelization and MTU size matching contradict the requirements on the slice layout in a picture. The implementation of this situation led to the development of the parallelization tools mentioned below.
[0083] Dependent slices have a short slice header and allow for bitstream partitioning at treeblock boundaries without breaking any intra-picture prediction. Basically, dependent slices segment regular slices into multiple NAL units to provide reduced end-to-end delay by allowing to send out a part of a regular slice before the encoding of the entire regular slice is completed.
[0084] In WPP, a picture is partitioned into individual rows of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible through parallel decoding of CTB rows, with the start of decoding of a CTB row delayed by two CTBs in order to ensure that data related to CTBs above and to the right of the subject CTB are available before the subject CTB is decoded. Using this staggered start (which looks like a wavefront when represented graphically), parallelization is possible for up to as many processors / cores as there are rows of pictures. Because intra-picture prediction between adjacent tree block rows within a picture is allowed, inter-processor / core communication needed to enable intra-picture prediction can be significant. WPP partitioning does not result in additional NAL units being generated compared to when no WPP partitioning is applied, so WPP is not a tool for MTU size matching. However, regular slices with some coding overhead can be used with WPP if MTU size matching is needed.
[0085] Tile definitions partition a picture into horizontal and vertical boundaries of tile columns and rows. Tile columns extend from the top of the picture to the bottom of the picture. Likewise, tile rows extend from the left of the picture to the right of the picture. The number of tiles in a picture can simply be derived as the number of tile columns multiplied by the number of tile rows.
[0086] The scan order of CTBs is changed to local within a tile (in the order of CTB raster scan of the tile) before decoding the top-left CTB of the next tile in the order of tile raster scan of the picture. Similar to regular slices, tiles break intra-picture prediction dependencies as well as entropy decoding dependencies. However, they do not need to be included in separate NAL units (same as WPP in this regard); thus, tiles cannot be used for MTU size matching. Each tile can be processed by one processor / core, and inter-processor / core communication needed between processing units decoding neighboring tiles for intra-picture prediction is limited to passing a shared slice header in the case of a slice that spans more than one tile, and loop filtering involves sharing of reconstructed samples and metadata. When more than one tile or WPP segment is included in a slice, an entry point byte offset for each tile or WPP segment other than the first tile or WPP segment in the slice is signaled in the slice header.
[0087] For simplicity, restrictions on the application of the four different picture partitioning schemes have been specified in HEVC. A given coded video sequence cannot include tiles and wavefronts for most of the profiles specified in HEVC. For each slice and tile, one or both of the following conditions must be met: 1) all coded tree blocks in a slice belong to the same tile; 2) all coded tree blocks in a tile belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is used, if a slice starts within a CTB row, it must end in the same CTB row.
[0088] Recent modifications to HEVC are specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)" Oct. 24, 2017, available at: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wgl l / JCTVC-AC1005-v2.zip. With the inclusion of this modification, HEVC specifies three MCTS-related SEI messages, namely the temporal MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nesting SEI message.
[0089] The temporal MCTS SEI message indicates the presence of MCTSs in the bitstream and signals the MCTSs. For each MCTS, the motion vectors are restricted to full-sample positions pointing inside the MCTS and fractional-sample positions that only require full-sample positions inside the MCTS for interpolation, and the use of motion vector candidates derived from blocks outside the MCTS for temporal motion vector prediction is not allowed. In this way, each MCTS can be decoded independently without the presence of tiles that are not included in the MCTS.
[0090] The MCTS extraction information set SEI message provides supplemental information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for an MCTS set. The information consists of multiple extraction information sets, each defining multiple MCTS sets, and containing the RBSP bytes of the replacement VPS, SPS, and PPS used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated as one or all of the syntax elements related to slice address (including first_slice_segment_in_pic_flag and slice_segment_address) will typically need to have different values.
[0091] 3.2. Picture partitioning in VVC
[0092] In VVC, a picture is partitioned into one or multiple tile rows and one or multiple tile columns. A tile is a sequence of CTUs that cover a rectangular region of a picture. The CTUs in a tile are scanned in the tile in a raster scan order.
[0093] A slice consists of an integer number of complete tiles within a tile of a picture or an integer number of consecutive complete CTU rows.
[0094] Two slice modes are supported, namely raster-scan slice mode and rectangular slice mode. In the raster-scan slice mode, a slice contains a sequence of complete tiles in the raster scan of the tiles of a picture. In the rectangular slice mode, a slice contains multiple complete tiles that collectively form a rectangular region of a picture or multiple consecutive complete CTU rows of one tile that collectively form a rectangular region of a picture. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular region corresponding to the slice.
[0095] A subpicture contains one or multiple slices that collectively cover a rectangular region of a picture.
[0096] Figure 1 An example of a raster-scan slice partitioning of a picture is shown, where the picture is partitioned into 12 tiles and 3 raster-scan slices.
[0097] Figure 2 An example of a rectangular slice partitioning of a picture is shown, where the picture is partitioned into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular slices.
[0098] Figure 3An example of a picture that is partitioned into tiles and rectangular slices is shown, where the picture is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular slices.
[0099] Figure 4 An example of subpicture partitioning of a picture is shown, where the picture is partitioned into 18 tiles, each tile on the left covering 12 tiles of one slice of 4 by 4 CTUs, and each tile on the right covering 6 tiles of 2 vertically stacked slices of 2 by 2 CTUs, resulting in 24 slices and 24 subpictures of different sizes (each slice is a subpicture).
[0100] 3.3. In-Sequence Picture Resolution Change
[0101] In AVC and HEVC, the spatial resolution of a picture cannot change unless a new sequence using a new SPS starts with an IRAP picture. VVC allows picture resolution change at in-sequence locations without encoding IRAP pictures that are always intra-coded. This feature is sometimes referred to as reference picture resampling (RPR) because it requires resampling of reference pictures used for inter prediction when the reference picture has a different resolution from the current picture being decoded.
[0102] The scaling ratio is limited to be greater than or equal to 1 / 2 (2 times down-sampling from the reference picture to the current picture) and less than or equal to 8 (8 times up-sampling). Three sets of resampling filter are specified with different frequency cutoffs to handle various scaling ratios between the reference picture and the current picture. The three sets of resampling filters are applied for scaling ratios from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, which is the same as the case of motion compensation interpolation filters. In fact, the normal MC interpolation process is a special case of the resampling process with scaling ratios ranging from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the picture width and height and the scaling offsets of the left, right, top, and bottom specified for the reference picture and the current picture.
[0103] Other aspects of the VVC design to support this feature, different from HEVC, include: i) the picture resolution and the corresponding conformance window are signaled in the PPS instead of the SPS, while the maximum picture resolution is signaled in the SPS. ii) For a single-layer bitstream, each picture store (a slot in the DPB for storing one decoded picture) occupies the buffer size required to store a decoded picture with the maximum picture resolution.
[0104] 3.4. General Scalable Video Coding (SVC) and SVC in VVC
[0105] Scalable Video Coding (SVC, sometimes also referred to as scalability in video coding) refers to video coding in which a base layer (BL), sometimes referred to as a reference layer (RL), and one or more scalable enhancement layers (ELs) are used. In SVC, the base layer can carry video data with a base level of quality. The one or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. An enhancement layer can be defined with respect to a previous coded layer. For example, a bottom layer can serve as a BL, while a top layer can serve as an EL. An intermediate layer can serve as an EL or an RL or both. For example, an intermediate layer (e.g., a layer that is neither the lowest layer nor the highest layer) can be an EL of the layers below the intermediate layer (e.g., the base layer or any intervening enhancement layers) and at the same time serve as an RL of the one or more enhancement layers above the intermediate layer. Similarly, in the Multiview or 3D extension of the HEVC standard, there can be multiple views, and information of one view can be used to code (e.g., encode or decode) information of another view (e.g., motion estimation, motion vector prediction, and / or other redundancies).
[0106] In SVC, parameters used by an encoder or a decoder are grouped into parameter sets based on the coding level at which they can be utilized (e.g., video level, sequence level, picture level, slice level, etc.). For example, parameters that can be used by one or more coded video sequences of different layers in a bitstream can be included in a video parameter set (VPS), while parameters used by one or more pictures in a coded video sequence can be included in a sequence parameter set (SPS). Similarly, parameters utilized by one or more slices in a picture can be included in a picture parameter set (PPS), and other parameters specific to a single slice can be included in a slice header. Similarly, an indication of the parameter set(s) used by a particular layer at a given time can be provided at various coding levels.
[0107] Because of the support for Reference Picture Resampling (RPR) in VVC, it is possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions) in VVC without requiring any additional signal processing level codec tools, as the upsampling required for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires high-level syntax changes (compared to no scalability support). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standards, including extensions of AVC and HEVC, VVC scalability was designed to be as friendly as possible to single-layer decoder designs. Decoding capabilities for multi-layer bitstreams are specified as if there were only a single layer in the bitstream. For example, decoding capabilities, such as DPB size, are specified in a way that is independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layer bitstreams do not require many changes to decode multi-layer bitstreams. Compared to the multi-layer extensions of AVC and HEVC, the HLS aspect has been significantly simplified at the expense of some flexibility. For example, IRAPU is required to contain the picture for each layer present in CVS.
[0108] 3.5. Viewport-dependent 360° video streaming based on sub-images
[0109] In 360° video (also known as omnidirectional video) streaming, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user at any given time, while the user can turn his / her head at any time to change their viewing orientation and thus change the current viewport. While it is desirable to have at least some lower-quality representations of areas available at the client that are not covered by the current viewport and are ready to be presented to the user only if the user suddenly changes their viewing orientation to anywhere on the sphere, a high-quality representation of the omnidirectional video is required for the currently used viewport. This optimization is achieved by dividing the high-quality representation of the entire omnidirectional video into sub-pictures with appropriate granularity. Using VVC, these two representations can be encoded as two independent layers.
[0110] Figure 5 The diagram illustrates a typical subpicture-based viewport-dependent 360° video delivery scheme, where a higher-resolution representation of the full video consists of subpictures, while a lower-resolution representation of the full video does not use subpictures and can be encoded and decoded using less frequent random access points represented by the higher-resolution representation. The client receives the lower-resolution full video, while for the higher-resolution video, it only receives and decodes the subpictures covering the current viewport.
[0111] The latest VVC draft specification also supports, for example Figure 6 The improved 360° video encoding / decoding scheme is shown. (Compared to...)Figure 5 The only difference between the methods shown is for... Figure 6 The method shown applies inter-layer prediction (ILP).
[0112] 3.6. Parameter Set
[0113] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in AVC, HEVC, and VVC. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0114] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or image, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving error resilience.
[0115] A VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0116] APS was introduced to carry image-level or stripe-level information that requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.
[0117] 3.7. Sub-image Sub-bitstream Extraction and Processing
[0118] Clause C.7 of the latest VVC documentation specifies the sub-image sub-bitstream extraction process as follows:
[0119] C.7 Sub-image Sub-bitstream Extraction Processing
[0120] The inputs to this process are the bitstream inBitstream, the target OLS index targetOlsIdx, the highest target TemporalId value tIdTarget, and the array subpicIdxTarget[] for the target subpick index values for each layer.
[0121] The output of this process is the sub-bit stream outBitstream.
[0122] For a given input bitstream, any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0123] - The output sub-bitstream is the output of the process specified in this clause with the bitstream for which targetOlsIdx is equal to an index of the OLS list specified by the VPS and subpicIdxTarget[ ] is equal to the subpicture indexes present in the OLS as input.
[0124] - The output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to each of the nuh layer id values in LayerIdInOls[ targetOlsIdx ].
[0125] - The output sub-bitstream contains at least one VCL NAL unit with Temporalld equal to tldTarget.
[0126] NOTE - The conforming bitstream contains one or more coded slice NAL units with Temporalld equal to 0, but is not necessarily required to contain coded slice NAL units with nuh layer id equal to 0.
[0127] - The output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to LayerIdInOls[ targetOlsIdx ][ i ] for each i in the range of 0 to NumLayersInOls[ targetOlsIdx ] - 1, inclusive, and where sh subpic id is equal to the value in SubpicldVal[ subpicIdxTarget[ i ] ].
[0128] The output sub-bitstream outBitstream is derived as follows:
[0129] - The sub-bitstream extraction process specified in Annex C.6 is invoked with inBitstream, targetOlsIdx, and tldTarget as inputs, and the output of this process is assigned to outBitstream.
[0130] - If some external means not specified in this Specification can be used to provide alternative parameter sets for the sub-bitstream outBitstream, all parameter sets are replaced with the alternative parameter sets.
[0131] - Otherwise, when subpicture level information SEI messages are present in inBitstream, the following applies:
[0132] - The variable subpicIdx is set equal to the value of subpicIdxTarget[ [ NumLayersInOls[ targetOlsIdx ] - 1 ] ].
[0133] - The value of general_level_idc in the vps_ols_ptl_idx[targetOlsIdx]th entry in the list of profile_tier_level( ) syntax structures in all referenced VPS NAL units is overwritten with a value equal to SubpicSetLevelIdc derived in Equation D.11 for the subpicture set consisting of subpictures with subpicture index equal to subpicIdx.
[0134] - When VCL HRD parameters or NAL HRD parameters are present, the respective values of cpb_size_value_minus1[ tldTarget ][ j ] and bit_rate_value_minus1[ tldTarget ][ j ] for the jth CPB in the vps_ols_hrd_idx[ MultiLayerOlsIdx[ targetOlsIdx ] ]th ols_hrd_parameters( ) syntax structure in all referenced VPS NAL units and in the ols_hrd_parameters( ) syntax structure in all SPS NAL units referred by the ith layer are overwritten so that they correspond to SubpicCpbSizeVcl[ SubpicSetLevelIdx ][ subpicIdx ] and SubpicCpbSizeNal[ SubpicSetLevelIdx ][ subpicIdx ] derived in Equations D.6 and D.7, respectively, SubpicBitrateVcl[ SubpicSetLevelIdx ][ subpicIdx ] and SubpicBitrateNal[ SubpicSetLevelIdx ][ subpicIdx ] derived in Equations D.8 and D.9, respectively, where SubpicSetLevelIdx is derived in Equation D.11 for subpictures with subpicture index equal to subpicIdx, j is in the range of 0 to hrd_cpb_cnt_minus1, inclusive, and i is in the range of 0 to NumLayersInOls[ targetOlsIdx ] - 1, inclusive.
[0135] The following applies for the ith layer, i in the range of 0 to NumLayersInOls[ targetOlsIdx ] - 1.
[0136] - The value of general_level_idc in the profile_tier_level( ) syntax structure in all reference SPS NAL units with sps_ptl_dpb_hrd_params_present_flag equal to 1 is overwritten to be equal to SubpicSetLevelIdc derived by Equation D.11 for the subpicture set consisting of the subpicture with subpicture index equal to subpicldx.
[0137] - The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are derived as follows:
[0138] subpicWidthInLumaSamples = min( (sps_subpic_ctu_top_left_x[ subpicldx ] + (C.24)
[0139] sps_subpic_width_minusl[ subpicldx ] + 1 ) * CtbSizeY,
[0140] pps_pic_width_in_luma_samples ) - 1
[0141] sps_subpic_ctu_top_left_x[ subpicldx ] * CtbSizeY
[0142] subpicHeightInLumaSamples = min( (sps_subpic_ctu_top_left_y[ subpicldx ] + (C.25)
[0143] sps_subpic_height_minusl[ subpicldx ] + 1 ) * CtbSizeY,
[0144] pps_pic_height_in_luma_samples ) - 1
[0145] sps_subpic_ctu_top_left_y[ subpicldx ] * CtbSizeY
[0146] - The values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS NAL units and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS NAL units are rewritten to be equal to subpicWidthlnLumaSamples and subpicHeightlnLumaSamples, respectively.
[0147] - The values of sps_num_subpics_minus1 in all referenced SPS NAL units and pps_num_subpics_minus1 in all referenced PPS NAL units are rewritten to be equal to 0.
[0148] - When present, the syntax elements sps_subpic_ctu_top_left_x[ subpicldx ] and sps_subpic_ctu_top_left_y[ subpicldx ] in all referenced SPS NAL units are rewritten to be equal to 0.
[0149] - The syntax elements sps_subpic_ctu_top_left_x[ j ], sps_subpic_ctu_top_left_y[ j ], sps_subpic_width_minus1[ j ], sps_subpic_height_minus1[ j ], sps_subpic_treated_as_pic_flag[ j ], sps_loop_filter_across_subpic_enabled_flag[ j ], and sps_subpic_id[ j ] for each j not equal to subpicldx in all referenced SPS NAL units are removed.
[0150] - The syntax elements in all referenced PPS are rewritten for signaling tiles and slices to remove all tile rows, tile columns, and slices that are not associated with the subpicture with subpicture index equal to subpicldx.
[0151] - The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are derived as follows:
[0152] subpicConfWinLeftOffset = sps_subpic_ctu_top_left_x[ subpicldx ] == 0? (C.26)
[0153] sps_conf_win_left_offset : 0
[0154] subpicConfWinRightOffset = (sps_subpic_ctu_top_left_x[ subpicldx ] + (C.27)
[0155] sps_subpic_width_minusl [ subpicldx ] + 1 ) * CtbSizeY > = sps_pic_width_max_in_luma_samples? sps_conf_win_right_offset : 0
[0156] subpicConfWinTopOffset = sps_subpic_ctu_top_left_y[ subpicldx ] == 0? (C.28)
[0157] sps_conf_win_top_offset : 0
[0158] subpicConfWinBottomOffset = (sps_subpic_ctu_top_left_y[ subpicldx ] + (C.29)
[0159] sps_subpic_height_minusl [ subpicldx ] + 1 ) * CtbSizeY > = sps_pic_height_max_in_luma_samples? sps_conf_win_bottom_offset : 0
[0160] - The values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset and sps_conf_win_bottom_offset in all referenced SPS NAL units and the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset and pps_conf_win_bottom_offset in all referenced PPS NAL units are respectively overwritten with equalities to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset and subpicConfWinBottomOffset.
[0161] - Remove from outBitstream all VCL NAL units having nuh layer id equal to nuh layer id of the i-th layer and having sh subpic id not equal to SubpicIdVal[subpicldx].
[0162] - When sli cbr constraint flag is equal to 1, remove all NAL units having nal unit type equal to FD NUT and filler payload SEI messages not associated with VCL NAL units of subpictures in subpicIdTarget[ ], and set cbr flag[tldTarget][j] equal to 1 in the j-th CPB of all referenced VPS NAL units and SPS NAL units in vps ols hrd idx[MultiLayerOlsIdx[targetOlsIdx]] equal to 1, and j is in the range of 0 to hrd cpb cnt minusl. Otherwise, (sli cbr constraint flag is equal to 0), clear all NAL units having nal unit type equal to FD NUT and filler payload SEI messages, and set cbr flag[tldTarget][j] equal to 0.
[0163] - When outBitstream contains SEI NAL units containing scalable-nested SEI messages with sn_ols_flag equal to 1 and sn_subpic_flag equal to 1 applicable to outBitstream, extract from the scalable-nested SEI messages the appropriate non-scalable-nested SEI messages with payloadType equal to 1 (PT), 130 (DUI) or 132 (Decoded Picture Hash), and put the extracted SEI messages into outBitstream.
[0164] 4. Technical problem solved by the technical solution
[0165] The existing design of the subpicture sub-bitstream extraction process in the latest VVC text (in JVET-R2001-vA / v10) has the following problems:
[0166] 1) When extracting a rectangular region from a picture sequence where the region covers one or more subpictures, in order to keep the scaling window identical to the one used during encoding of the original bitstream, the scaling window offset parameters in PPS will need to be overwritten as the picture width and / or height of the extracted picture has changed. This is particularly needed in the scenario of the improved 360o video coding scheme as shown in Figure 6 However, the subpicture sub-bitstream extraction process in the latest VVC text does not necessarily overwrite the scaling window offset parameters.
[0167] 2) In the subpicture sub-bitstream extraction process in the latest VVC text, any output sub-bitstream that satisfies all the following conditions is a conforming bitstream:
[0168] - The output sub-bitstream is the output of the process specified in this clause with, as input, for the bitstream, targetOlsIdx equal to an index of the OLS list specified by the VPS and subpicIdxTarget[] equal to the subpicture indexes present in the OLS.
[0169] - The output sub-bitstream contains at least one VCL NAL unit with nuh_layer_id equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0170] - The output sub-bitstream contains at least one VCL NAL unit with Temporalld equal to tIdTarget.
[0171] NOTE - A conforming bitstream contains one or more coded slice NAL units with Temporalld equal to 0, but does not necessarily contain coded slice NAL units with nuh layer id equal to 0.
[0172] - The output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to LayerldInOls[targetOlsldx][i] and sh subpic id equal to the value in SubpicldVal[subpicldxTarget[i]] for each i in the range of 0 to NumLayersInOls[targetOlsldx] - 1, inclusive.
[0173] However, the above constraints have the following problems:
[0174] a. In the first item, the input tldTarget is missing.
[0175] b. Another problem associated with the first item is as follows. Only certain combinations of sub-pictures of pictures from different layers can form a conforming bitstream, not all combinations.
[0176] 3) The removal of VCL NAL units and their associated filler data NAL units and associated filler payload SEI messages, etc. is only done when there is no external means for replacing parameter sets. However, such removal is also needed when there is an external means for replacing parameter sets.
[0177] 4) The current removal of filler payload SEI messages can involve the rewriting of SEI NAL units.
[0178] 5) The sub-picture level information (SLI) SEI message is specified as layer-specific. However, the SLI SEI message is used in the sub-picture sub-bitstream extraction process as if the information applies to sub-pictures of all layers.
[0179] 6) The scalable nesting SEI message can be used to nest SEI messages for certain extracted sub-picture sequences of certain OLSs. However, the semantics of sn num subpics minusl and sn subpic id len minusl are specified as layer-specific due to the use of the phrase "in the CLVS" or "in a CLVS".
[0180] 7) Changes are needed to support the extraction of bitstreams, such as Figure 6the lower part of the picture, wherein in the extracted bitstream sent to the decoder, the pictures containing multiple sub-pictures in the original bitstream now contain fewer sub-pictures, while the pictures containing only one sub-picture remain unchanged.
[0181] 5. Examples of solutions and embodiments
[0182] To solve the above problems and other problems, the methods as outlined below are disclosed. These items should be considered as examples to explain the general concepts and should not be interpreted in a narrow way. Furthermore, these items can be applied individually or in any combination. In the following description, the parts that are added or modified from the related specification are highlighted in bold and italic, and some parts that are deleted are marked with double brackets (e.g. [[a]] means the deletion of the character “a”). There can be some other changes that are substantially editorial and thus not highlighted.
[0183] 1) To solve problem 1, the calculation and rewriting of the scaling window offset parameters (such as in PPS) are specified as part of the sub-picture sub-bitstream extraction process.
[0184] a. In one example, the scaling window offset parameters are calculated and rewritten as follows:
[0185] - The variables subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset and subpicScalWinBotOffset are derived as follows:
[0186] subpicScalWinLeftOffset = pps_scaling_win_left_offset - (C.30)
[0187] sps_subpic_ctu_top_left_x[spIdx] * CtbSizeY / SubWidthC
[0188] rightSubpicBd = (sps_subpic_ctu_top_left_x[spIdx] + (C.31)
[0189] sps_subpic_width_minus1[spIdx] + 1) * CtbSizeY
[0190] subpicScalWinRightOffset = (C.32)
[0191] ( rightSubpicBd >= sps_pic_width_max_in_luma_samples )? ( C.31 )
[0192] pps_scaling_win_right_offset: pps_scaling_win_right_offset
[0193] ( sps_pic_width_max_in_luma_samples - rightSubpicBd ) / SubWidthC
[0194] SubWidthC
[0195] subpicScalWinRightOffset = pps_scaling_win_right_offset - ( C.32 )
[0196] sps_subpic_ctu_top_left_x[ spIdx ] * CtbSizeY / SubHeightC
[0197] botSubpicBd = ( sps_subpic_ctu_top_left_y[ spIdx ] + CtbSizeY ) * SubWidthC
[0198] sps_subpic_height_minus1[ spIdx ] + 1
[0199] subpicScalWinBotOffset = pps_scaling_win_bottom_offset - ( C.33 )
[0200] ( botSubpicBd >= sps_pic_height_max_in_luma_samples )? ( C.33 )
[0201] pps_scaling_win_bottom_offset: pps_scaling_win_bottom_offset
[0202] ( sps_pic_height_max_in_luma_samples - botSubpicBd ) / SubHeightC
[0203] SubHeightC
[0204] where sps_subpic_ctu_top_left_x[spldx], sps_subpic_width_minusl[spldx], sps_subpic_height_minusl[spldx], sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in the above equations are from the original SPS before they are overwritten, and pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset and pps_scaling_win_bottom_offset in the above are from the original PPS before they are overwritten.
[0205] - the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset and pps_scaling_win_bottom_offset in all the referred PPS NAL units are overwritten to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset and subpicScalWinBotOffset, respectively.
[0206] b. In one example, sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in the above equations are replaced by pps_pic_width_in_luma_samples and pps_pic_width_in_luma_samples, respectively, and pps_pic_width_in_luma_samples is the value in the original PPS before they are overwritten.
[0207] c. In addition, in one example, the reference sample enabling flag (e.g., sps_ref_pic_resampling_enabled_flag) in the SPS is overwritten when changed, and the resolution change allowed flag (e.g., sps_res_change_in_clvs_allowed_flag) in the SPS is overwritten when changed.
[0208] 2) To address problem 2a and 2b, propose the following method.
[0209] a. To address problem 2a, add the missing input tIdTarget, for example, by changing
[0210] - The output sub-bitstream is the output of the process specified in this clause, where for the bitstream targetOlsIdx is equal to an index of the OLS list specified by the VPS and subpicIdxTarget[] is equal to the subpicture indices present in the OLS as input.
[0211] to the following:
[0212] - The output sub-bitstream is the output of the process specified in this clause, where for the bitstream targetOlsIdx is equal to an index of the OLS list specified by the VPS and subpicIdxTarget[] is equal to the subpicture indices present in the OLS as input.
[0213] b. To address problem 2b, clearly specify which combination of subpictures of different layers needs to be a conforming bitstream when extracted, for example as follows:
[0214] The bitstream conformance requirement of the input bitstream is that any output sub-bitstream that satisfies all the following conditions shall be a conforming bitstream:
[0215] - The output sub-bitstream is the output of the process specified in this clause, where for the bitstream targetOlsIdx is equal to an index of the OLS list specified by the VPS, tIdTarget is equal to any value in the range of 0 to vps_max_sublayers_minus1, inclusive, and as input:
[0216] - The output sub-bitstream contains at least one VCL NAL unit with nuh_layer_id equal to each of the nuh_layer_id values in the list LayerIdInOls[targetOlsIdx].
[0217] - The output sub-bitstream contains at least one VCL NAL unit with Temporalld equal to tIdTarget.
[0218] NOTE - A conforming bitstream contains one or more coded slice NAL units with Temporalld equal to 0, but does not necessarily contain coded slice NAL units with nuh layer id equal to 0.
[0219] - The output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to LayerldInOls[targetOlsldx][i] and sh subpic id equal to SubpicldVal[subpicldxTarget[i]] for each i in the range of 0 to NumLayersInOls[targetOlsldx] - 1, inclusive.
[0220] 3) To address issue 3, removal of VCL NAL units and their associated filler data NAL units and associated filler payload SEI messages, etc. is performed regardless of whether there is an external means for replacing parameter sets.
[0221] 4) To address issue 4, a constraint is added to require that an SEI NAL unit containing a filler payload SEI message does not contain other types of SEI messages, and removal of a filler payload SEI message is specified as removal of an SEI NAL unit containing the filler payload SEI message.
[0222] 5) To address issue 5, it is specified such that one or more of the values of SubpicSetLevelldc, SubpicCpbSizeVcl[SubpicSetLevelldx][subpicldx], SubpicCpbSizeNal[SubpicSetLevelldx][subpicldx], SubpicBitrateVcl[SubpicSetLevelldx][subpicldx], SubpicBitrateNal[SubpicSetLevelldx][subpicldx], and sli cbr constraint flag derived based on the SLI SEI message or found in the SLI SEI message are the same for all layers in the target OLS of the extraction process, respectively.
[0223] a. In one example, a constraint is that, for use with a multi-layer OLS, the SLI SEI message shall be contained in a scalable nesting SEI message and shall be indicated in the scalable nesting SEI message to apply to a particular OLS that includes at least the multi-layer OLS (i.e., when sn ols flag is equal to 1).
[0224] i. Furthermore, the value 203 (payloadType value of SLI SEI message) is removed from the list VclAssociatedSeiList to enable the SLI message to be included in a scalable nesting SEI message with sn_ols_flag equal to 1.
[0225] b. In one example, it is constrained that, for use with multi-layer OLS, the SLI SEI message shall be included in a scalable nesting SEI message and shall be indicated in the scalable nesting SEI message to apply to a particular list of layers consisting of all layers of the multi-layer OLS (i.e., sn_ols_flag equal to 1).
[0226] c. In one example, the SLI SEI message is specified as OLS-specific, e.g., by adding an OLS index to the SLI SEI message syntax to indicate the OLS to which the information carried in the SEI message applies.
[0227] i. Optionally, a list of OLS indexes is added to the SLI SEI message syntax to indicate a list of OLSes to which the information carried in the SEI message applies.
[0228] d. In one example, it is required that the values of one or more of SubpicSetLevelldx, SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx], SubpicCpbSizeNal[SubpicSetLevelIdx][subpicIdx], SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx], SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx], and sli_cbr_constraint_flag in all SLI SEI messages for all layers of an OLS shall be respectively the same.
[0229] 6) To address issue 6, the semantics from:
[0230] sn_num_subpics_minus1 plus 1 specifies the number of sub-pictures to which the scalable nesting SEI message applies. The value of sn_num_subpics_minus1 shall be less than or equal to the value of sps_num_subpics_minus1 in the SPS referred to by the pictures in the CLVS.
[0231] sn_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sn_subpic_id[ i ]. The value of sn_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.
[0232] It is a requirement of bitstream conformance that the value of sn_subpic_id_len_minus1 shall be the same for all scalable-nested SEI messages present in a CLVS.
[0233] Change to the following:
[0234] sn_num_subpics_minus1 plus 1 specifies the number of subpictures in that the scalable-nested SEI message applies to. The value of sn_num_subpics_minus1 shall be less than or equal to the value of sps_num_subpics_minus1 in the SPS referred to by the [[CLVS in which]]
[0235] sn_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sn_subpic_id[ i ]. The value of sn_subpic_id_len_minus1 shall be in the range of 0 to 15, inclusive.
[0236] It is a requirement of bitstream conformance that the value of sn_subpic_id_len_minus1 shall be the same for all scalable-nested SEI messages present in a [[CLVS]] .
[0237] 7) To address issue 7, the following approach is proposed:
[0238] a. In one example, it is clearly specified which combination of subpictures of different layers when extracted is required to be a conforming bitstream, e.g., as follows, where there are at least two layers, where the pictures of one layer comprise multiple subpictures and the pictures of the other layer comprise only one subpicture:
[0239] For bitstream conformance of an input bitstream, any output sub-bitstream shall be a conforming bitstream that meets all the following conditions:
[0240] - the output sub-bitstream is the output of the processing specified in this clause, where for the bitstream, targetOlsIdx is equal to the index of the OLS list specified by the VPS,
[0241] - The output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to each of the values in nuh layer id in Ols [ targetOlsIdx ].
[0242] - The output sub-bitstream contains at least one VCL NAL unit with Temporalld equal to tldTarget.
[0243] NOTE 2 - A conforming bitstream contains one or more coded slice NAL units with Temporalld equal to 0, but is not required to contain coded slice NAL units with nuh layer id equal to 0.
[0244] - The output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to LayerldlnOls [ targetOlsIdx ] [ i ] for each i in the range of 0 to NumLayerslnOls [ targetOlsIdx ] - 1, inclusive, and with sh subpic id equal to the value in SubpicldVal [ subpicIdxTarget [ i ] ].
[0245] b. The semantics of the scalable nesting SEI message are specified such that when sn_ols_flag and sn_subpic_flag are both equal to 1, the list of sn_subpic_id [ ] values specifies subpicture IDs of subpictures in pictures in the applicable OLS, which each contain multiple subpictures. In this way, it is allowed that the OLS to which the scalable nesting SEI message applies has layers in which each picture contains only one subpicture for which the subpicture ID is not indicated in the scalable nesting SEI message.
[0246] 8) If the first reference picture is partitioned into subpictures in a subpicture layout that is different from the subpicture layout of the current picture, it is required that the first reference picture is not set as a collocated picture of the current picture.
[0247] a. In one example, the first reference picture and the current picture are in different layers.
[0248] 9) If coding tool X depends on the first reference picture, and the first reference picture is partitioned into subpictures in a subpicture layout that is different from the subpicture layout of the current picture, it is required that the coding tool X is disabled for the current picture.
[0249] a. In one example, the first reference picture and the current picture are in different layers.
[0250] b. In one example, the coding tool X is bi-directional optical flow (BDOF).
[0251] c. In one example, the coding tool X is decoder-side motion vector refinement (DMVR).
[0252] d. In one example, the coding tool X is prediction refinement with optical flow (PROF).
[0253] 6. Embodiments
[0254] Below are some example embodiments for some aspects of the present invention that can be applied to the VVC specification outlined in Section 5 above. The modified text is based on the latest VVC text in JVET-R2001-vA / vl0. Most of the added or modified relevant parts are highlighted in bold slant, and some of the deleted parts are marked with double brackets (e.g., [[a]] means the deleted character “a”). There can be some other changes that are editorial in nature and thus not highlighted.
[0255] 6.1. First Embodiment
[0256] This embodiment is directed to items 1, 1.a, 2a, 2b, 3, 4, 5, 5a, 5.a.i, and 6.
[0257] C.7 Subpicture sub-bitstream extraction process
[0258] The inputs of this process are the bitstream inBitstream, the target OLS index targetOlsIdx, the target highest Temporalld value tldTarget, and, for i in the range of [[an array of]] target subpicture index values for each layer subpicldTarget[i] for i in the range of
[0259] The output of this process is the sub-bitstream outBitstream.
[0260] For bitstream conformance of the input bitstream, any output sub-bitstream that satisfies all the following conditions shall be a conforming bitstream:
[0261] - the output sub-bitstream is the output of the process specified in this clause for the bitstream where targetOlsIdx is equal to the index of the OLS list specified by the VPS, and [[Equal to the subpicture index present in the OLS as input]] as input:
[0262]
[0263] For use with multi-layer OLSs, the SLI SEI message shall be contained in a scalable nesting SEI message and shall be indicated in the scalable nesting SEI message to apply to a particular OLS or to apply to all layers in a particular OLS.
[0264] - the output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to each of the values in the list LayerldlnOls [ targetOlsldx ].
[0265] - the output sub-bitstream contains at least one VCL NAL unit with Temporalld equal to tldTarget.
[0266] NOTE - A conforming bitstream contains one or more coded slice NAL units with Temporalld equal to 0, but is not necessarily required to contain coded slice NAL units with nuh layer id equal to 0.
[0267] - the output sub-bitstream contains at least one VCL NAL unit with nuh layer id equal to LayerldlnOls [ targetOlsldx ][ i ] and sh subpic id equal to the value in SubpicldVal [ subpicldxTarget [ i ] ] for each i in the range of 0 to NumLayerslnOls [ targetOlsldx ] - 1, inclusive.
[0268] The output sub-bitstream outBitstream is derived as follows:
[0269] - the sub-bitstream extraction process specified in Annex C.6 is invoked with inBitstream, targetOlsldx, and tldTarget as inputs, and the output of this process is assigned to outBitstream.
[0270]
[0271] - if some external means not specified in this Specification can be used to provide alternative parameter sets for the sub-bitstream outBitstream, all parameter sets are replaced with the alternative parameter sets.
[0272] - Otherwise, when subpicture level information SEI messages are present in inBitstream, the following applies:
[0273] - [[The variable subpicIdx is set equal to the value of subpicIdxTarget[[NumLayersInOls[targetOlsIdx] - 1]]]].
[0274] - The value of general_level_idc in the vps_ols_ptl_idx[targetOlsIdx]th entry in the list of profile_tier_level( ) syntax structures in all referenced VPS NAL units is overwritten with the value equal to SubpicSetLevelIdc derived in Equation D.11 for the subpicture set consisting of subpictures with subpicture index equal to subpicIdx.
[0275] - When VCL HRD parameters or NAL HRD parameters are present, in all referenced VPS NAL units, in the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]]th ols_hrd_parameters( ) syntax structure and in the ols_hrd_parameters( ) syntax structure in all SPS NAL units referred by the i-th layer, the respective values of cpb_size_value_minus1[tIdTarget][j] and bit_rate_value_minus1[tIdTarget][j] for the j-th CPB are overwritten so that they correspond to SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx] and SubpicCpbSizeNal[SubpicSetLevelIdx][subpicIdx] derived in Equations D.6 and D.7, respectively, SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx] and SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx] derived in Equations D.8 and D.9, respectively, where SubpicSetLevelIdx is derived in Equation D.11 for subpictures with subpicture index equal to subpicIdx, j is in the range of 0 to hrd_cpb_cnt_minus1, inclusive, and i is in the range of 0 to NumLayersInOls[targetOlsIdx] - 1, inclusive.
[0276] - For each value of [[i-th layer]] in the range of 0 to NumLayersInOls[targetOlsIdx]-1, the following applies.
[0277]
[0278] - Rewrite the value of general_level_idc in the profile_tier_level() syntax structure of all referenced SPS NAL units where sps_ptl_dpb_hrd_params_present_flag is equal to 1, to be equal to SubpicSetLevelIdc derived from Equation D.11 for the set of subpicks whose subpick index is equal to subpicIdx.
[0279] - The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are exported as follows:
[0280] subpicWidthInLumaSamples=min((sps_subpic_ctu_top_left_x[spIdx]+(C.24)
[0281] sps_subpic_width_minus1[spIdx]+1)*CtbSizeY,pps_pic_width_in_luma_samples)-
[0282] sps_subpic_ctu_top_left_x[spIdx]*CtbSizeY
[0283] subpicHeightInLumaSamples=min((sps_subpic_ctu_top_left_y[spIdx]+(C.25)
[0284] sps_subpic_height_minus1[spIdx]+1)*CtbSizeY,pps_pic_height_in_luma_samples)-
[0285] sps_subpic_ctu_top_left_y[spIdx]*CtbSizeY
[0286] - The values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS NAL units and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS NAL units are rewritten to be equal to subpicWidthlnLumaSamples and subpicHeightlnLumaSamples, respectively.
[0287] - The values of sps_num_subpics_minus1 in all referenced SPS NAL units and pps_num_subpics_minus1 in all referenced PPS NAL units are rewritten to be equal to 0.
[0288] - When present, the syntax elements sps_subpic_ctu_top_left_x[ spIdx ] and sps_subpic_ctu_top_left_y[ spIdx ] in all referenced SPS NAL units are rewritten to be equal to 0.
[0289] - The syntax elements sps_subpic_ctu_top_left_x[ j ], sps_subpic_ctu_top_left_y[ j ], sps_subpic_width_minus1[ j ], sps_subpic_height_minus1[ j ], sps_subpic_treated_as_pic_flag[ j ], sps_loop_filter_across_subpic_enabled_flag[ j ], and sps_subpic_id[ j ] for each j not equal to subpicIdx in all referenced SPS NAL units are removed.
[0290] - The syntax elements in all referenced PPS are rewritten for signaling tiles and slices to remove all tile rows, tile columns, and slices that are not associated with the subpicture with subpicture index equal to subpicIdx.
[0291] - The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are derived as follows:
[0292] subpicConfWinLeftOffset = sps_subpic_ctu_top_left_x[ spIdx ] == 0? (C.26)
[0293] sps_conf_win_left_offset : 0
[0294] subpicConfWinRightOffset = (sps_subpic_ctu_top_left_x[ spIdx ] + (C.27)
[0295] sps_subpic_width_minus1[ spIdx ] + 1 ) * CtbSizeY > =
[0296] sps_pic_width_max_in_luma_samples? sps_conf_win_right_offset : 0
[0297] subpicConfWinTopOffset = sps_subpic_ctu_top_left_y[ spIdx ] == 0? (C.28)
[0298] sps_conf_win_top_offset : 0
[0299] subpicConfWinBottomOffset = (sps_subpic_ctu_top_left_y[ spIdx ] + (C.29)
[0300] sps_subpic_height_minus1[ spIdx ] + 1 ) * CtbSizeY > =
[0301]
[0302] - The values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset and sps_conf_win_bottom_offset in all referenced SPS NAL units and the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset and pps_conf_win_bottom_offset in all referenced PPS NAL units are respectively overwritten with equalities to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset and subpicConfWinBottomOffset.
[0303]
[0304] - [[Remove from outBitstream all VCL NAL units with nuh layer id equal to the nuh layer id of the i-th layer and with sh subpic id not equal to SubpicIdVal[subpicIdx].]]
[0305] - [[When]] if sli cbr constraint flag is equal to 1 [[then]] [[remove all NAL units with nal unit type equal to FD NUT and associated with no VCL NAL units of subpictures in subpicIdTarget[ ], and]] set cbr flag[tIdTarget][j] equal to 1 in the j-th CPB of all ols hrd parameters() syntax structures in all referenced VPS NAL units and SPS NAL units with vps ols hrd idx[MultiLayerOlsIdx[targetOlsIdx]] equal to i for j in the range of 0 to hrd cpb cnt minusl. Otherwise, (sli cbr constraint flag is equal to 0), [[clear all NAL units with nal unit type equal to FD NUT and associated with no VCL NAL units of subpictures in subpicIdTarget[ ], and]] set cbr flag[tIdTarget][j] equal to 0.
[0306] - When outBitstream contains an SEI NAL unit containing a scalable-nested SEI message with sn_ols_flag equal to 1 and sn_subpic_flag equal to 1 applicable to outBitstream, extract from the scalable-nested SEI message the appropriate non-scalable-nested SEI message with payloadType equal to 1 (PT), 130 (DUI) or 132 (Decoded Picture Hash), and put the extracted SEI message into outBitstream.
[0307] D2.2 General SEI payload semantics ...
[0309] The list VclAssociatedSeiList is set to consist of the payloadType values 3, 19, 45, 129, 132, 137, 144, 145, 147 to 150, inclusive, 153 to 156, inclusive, 168, [[203,]] and 204.
[0310] The list PicUnitRepConSeiList is set to include the payloadType values 0, 1, 19, 45, 129, 132, 133, 137, 147 to 150, inclusive, 153 to 156, inclusive, 168, 203 and 204.
[0311] NOTE 4 - VclAssociatedSeiList consists of the payloadType values of SEI messages that, when non-scalable-nested and contained in an SEI NAL unit, infer constraints on the NAL unit header of the SEI NAL unit based on the NAL unit header of the associated VCL NAL unit. PicUnitRepConSeiList consists of the payloadType values of SEI messages that are subject to 4 times repetition per PU.
[0312] The requirement of bitstream conformance is that the following restrictions apply to the inclusion of SEI messages in an SEI NAL unit:
[0313] - When the SEI NAL unit contains a non-scalable-nested BP SEI message, a non-scalable-nested PT SEI message or a non-scalable-nested DUI SEI message, the SEI NAL unit shall not contain any other SEI message with a payloadType not equal to 0 (BP), 1 (PT) or 130 (DUI).
[0314] - When an SEI NAL unit contains a scalable-nested BP SEI message, a scalable-nested PT SEI message, or a scalable-nested DUI SEI message, the SEI NAL unit shall not contain any other SEI messages with payloadType not equal to 0 (BP), 1 (PT), 130 (DUI), or 133 (scalable-nested).
[0315] ...
[0317] D.6.2 Scalable-nested SEI message semantics ...
[0319] sn num subpics minusl plusl specifies the number of subpictures in each picture in the OLS (when sn ols flag is equal to 1) or layer (when sn ols flag is equal to 0) to which the scalable-nested SEI message applies. The value of sn num subpics minusl shall be less than or equal to the value of sps num subpics minusl in each of the SPSs referred to by the pictures in the OLS (when sn ols flag is equal to 1) or layer (when sn ols flag is equal to 0) in the [[CLVS]].
[0320] sn subpic id len minusl plusl specifies the number of bits used to represent the syntax element sn subpic id [ i ]. The value of sn subpic id len minusl shall be in the range of 0 to 15, inclusive.
[0321] A requirement for bitstream conformance is that the value of sn subpic id len minusl shall be the same for all scalable-nested SEI messages that apply to pictures within the [[CLVS]]. ...
[0323] Figure 7is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 1900. The system 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be received in a compressed or encoded format. The input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0324] The system 1900 can include a codec component 1904, which can implement various coding or encoding methods described in this document. The codec component 1904 can reduce the average bitrate of the video from the input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 can be stored or transmitted via a connected communication as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 can be used by component 1908 for generating pixel values or displayable video that is sent to a display interface 1910. The process of generating user- visible video from the bitstream representation is sometimes referred to as video decompression. Also, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used at an encoder, and a decoder will perform corresponding decoding tools or operations that reverse the results of that coding.
[0325] Examples of peripheral bus interfaces or display interfaces can include a universal serial bus (USB) or a high-definition multimedia interface (HDMI) or display port, etc. Examples of storage interfaces include SATA (serial advanced technology attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices, such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0326] Figure 8is a block diagram of a video processing apparatus 3600. The apparatus 3600 can be used to implement one or more of the methods described herein. The apparatus 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in the present document. The memory(ies) 3604 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement, in hardware circuitry, some of the techniques described in the present document.
[0327] Figure 10 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0328] As shown in Figure 10 , the video coding system 100 can include a source device 110 and a destination device 120. The source device 110, which can be referred to as a video encoding device, generates encoded video data. The destination device 120, which can be referred to as a video decoding device, can decode the encoded video data generated by the source device 110.
[0329] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0330] The video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 by the network 130a via the I / O interface 116. The encoded video data can also be stored onto a storage medium / server 130b for access by the destination device 120.
[0331] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0332] The I / O interface 126 can include a receiver and / or a modem. The I / O interface 126 can obtain encoded video data from the source device 110 or the storage medium / server 130b. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to a user. The display device 122 can be integrated with the destination device 120, or can be external to the destination device 120 configured to interface with an external display device.
[0333] The video encoder 114 and the video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[0334] Figure 11 is a block diagram illustrating an example of a video encoder 200 that can be Figure 10 the video encoder 114 in the system 100 illustrated in FIG.
[0335] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 11 example, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0336] The functional components of the video encoder 200 can include a partitioning unit 201, a prediction unit 202 (which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0337] In other examples, the video encoder 200 can include more, less, or different functional components. In one example, the prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0338] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately for explanatory purposes. Figure 11 In examples, the motion estimation unit 204 and the motion compensation unit 205 are separate components. In other examples, the motion estimation unit 204 and the motion compensation unit 205 are highly integrated and can be viewed as a single component.
[0339] The partitioning unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0340] The mode selection unit 203 can select one of intra or inter coding modes (e.g., based on the error results), and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 also selects a resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).
[0341] To perform inter prediction for a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 to the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples from a picture of the buffer 213 that is different from the picture associated with the current video block.
[0342] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0343] In some examples, the motion estimation unit 204 can perform single prediction for a current video block, and the motion estimation unit 204 can search a reference picture of list 0 or list 1 for a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index indicating the reference picture of list 0 or list 1 that contains the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0344] In other examples, the motion estimation unit 204 can bi-predictively encode the current video block, the motion estimation unit 204 can search a reference picture in List 0 for a reference video block for the current video block, and also search a reference picture in List 1 for another reference video block for the current video block. The motion estimation unit 204 can then generate a reference index that indicates the reference pictures in List 0 and List 1 that contain the reference video blocks and a motion vector that indicates a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0345] In some examples, the motion estimation unit 204 can output the full set of motion information for use in the decoding process at the decoder.
[0346] In some examples, the motion estimation unit 204 can not output the full set of motion information for the current video. Instead, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0347] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0348] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0349] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of techniques of predictive signaling that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.
[0350] The intra prediction unit 206 can intra-predict the current video block. When the intra prediction unit 206 intra-predicts the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0351] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the predicted video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0352] In other examples, such as in skip mode, there can be no residual data for the current video block, and residual generation unit 207 can not perform a subtraction operation.
[0353] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0354] After transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0355] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the predicted video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0356] After reconstruction unit 212 reconstructs the video block, in-loop filtering operations can be performed to reduce video block artifacts in the video block.
[0357] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0358] Figure 12 is a block diagram illustrating an example of a video decoder 300 that can be Figure 10 the video decoder 114 in the system 100 illustrated in FIG.
[0359] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 12In examples of the video decoder 300, the video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0360] In Figure 12 In examples of the video decoder 300, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding pass generally reciprocal to the encoding pass described with respect to the video encoder 200. Figure 11 ) described with respect to the video encoder 200.
[0361] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data and the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference picture list indices, and other motion information, from the entropy decoded video data. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge mode.
[0362] The motion compensation unit 302 can generate a motion compensated block, possibly with interpolation. An identifier of an interpolation filter to be used at sub-pixel precision can be included in the syntax elements.
[0363] The motion compensation unit 302 can use an interpolation filter as used by the video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 from the received syntax information and use the interpolation filter to generate the predictive block.
[0364] The motion compensation unit 302 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how to partition each macroblock of the pictures of the encoded video sequence, modes indicating how to encode each partition, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence.
[0365] The intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0366] The reconstruction unit 306 can sum the residual block with a corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and also produces decoded video for presentation on a display device.
[0367] The list of solutions describes some embodiments of the disclosed technology.
[0368] A first set of solutions is next provided. The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0369] 1. A method of video processing (e.g., item 1 of section A above), comprising: Figure 9 The method 900 depicted in item 2 of section A above, comprises:
[0370] 2. The method of solution 1, wherein the one or more parameters comprise a left offset, a right offset, a top offset, or a bottom offset of the scaling window applicable to the subpicture.
[0371] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 3).
[0372] 3. A method of video processing, comprising: converting between a video comprising one or more video pictures in a video coding layer and a coded representation of the video, wherein the converting conforms to a rule that specifies removal of a network abstraction layer unit and filter data and filler supplemental enhancement information data regardless of whether parameter sets are replaceable externally.
[0373] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 4).
[0374] 4. A method of video processing, comprising converting between a video comprising one or more video pictures in a video coding layer and a coded representation of the video, wherein the converting conforms to a format rule that specifies excluding other supplemental enhancement information (SEI) from a network abstraction layer that includes a filler SEI message payload.
[0375] 5. The method of solution 1, wherein the converting conforms to a rule that specifies removing the filler payload SEI message as an SEI NAL unit.
[0376] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 7).
[0377] 6. A method of video processing, comprising converting between a video comprising one or more layers comprising one or more video pictures comprising one or more subpictures and a coded representation of the video, wherein the converting is according to a format rule that specifies which combination of subpictures of different layers are required to conform to the format rule, wherein there are at least two layers such that pictures of one layer are each composed of multiple subpictures and pictures of another layer are each composed of only one subpicture.
[0378] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 8).
[0379] 7. A method of video processing, comprising converting between a video comprising one or more layers comprising one or more video pictures comprising one or more subpictures and a coded representation of the video, wherein the converting conforms to a rule that specifies that, for a first picture that is partitioned into subpictures using a subpicture layout, a second picture cannot be used as a collocated reference picture in a case where the second picture is partitioned according to a different subpicture layout than the first picture.
[0380] 8. The method of solution 7, wherein the first picture and the second picture are in different layers.
[0381] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 9).
[0382] 9. A video processing method, comprising: converting between a video comprising one or more layers comprising one or more video pictures comprising one or more subpictures, and a coded representation of the video, wherein the conversion conforms to a rule that specifies that, for a first picture that is partitioned into subpictures using a subpicture layout, a coding tool is prohibited from being used during encoding, or a corresponding decoding tool is prohibited from being used during decoding, in case the coding tool depends on a different subpicture layout than a second picture that is used as a reference picture for the first picture.
[0383] 10. The method of solution 9, wherein the first picture and the second picture are in different layers.
[0384] 11. The method of any of solutions 1 to 10, wherein the conversion comprises encoding the video into the coded representation.
[0385] 12. The method of any of solutions 1 to 10, wherein the conversion comprises decoding the coded representation to generate pixel values of the video.
[0386] 13. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 12.
[0387] 14. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 12.
[0388] 15. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement a method recited in any of solutions 1 to 12.
[0389] 16. A method, an apparatus or a system described in the present document.
[0390] The second set of solutions shows example embodiments of the techniques discussed in the preceding section (e.g., item 1).
[0391] 1. A video processing method (e.g., method 1300 as shown in Figure 13 FIG. 13), comprising: converting 1302, between a video comprising one or more video pictures comprising one or more subpictures, and a bitstream of the video, wherein the conversion conforms to a rule that specifies to determine, from one or more syntax elements, one or more parameters of a scaling window applicable to a subpicture during subpicture sub-bitstream extraction processing.
[0392] 2. The method of solution 1, wherein the one or more syntax elements are included in a picture parameter set and / or a sequence parameter set.
[0393] 3. The method of solution 1 or 2, wherein the original picture parameter set and / or the original sequence parameter set change with the computing and overwriting of the one or more parameters.
[0394] 4. The method of any of solutions 1 to 3, wherein the one or more parameters comprise at least one of a left offset, a right offset, a top offset, or a bottom offset of the scaling window applicable to the sub-picture.
[0395] 5. The method of any of solutions 1 to 4, wherein a left offset of the scaling window (subpicScalWinLeftOffset) is derived as follows:
[0396] subpicScalWinLeftOffset = pps_scaling_win_left_offset - sps_subpic_ctu_top_left_x[spIdx] * CtbSizeY / SubWidthC, and
[0397] where pps_scaling_win_left_offset indicates a left offset applied to picture size for scaling ratio calculation, sps_subpic_ctu_top_left_x[spIdx] indicates an x-coordinate of a coding tree unit located at the top-left corner of the sub-picture, CtbSizeY is a width or height of a luma coding tree block or coding tree unit, and SubWidthC indicates a width of a video block and is obtained from a table according to a chroma format of a picture including the video block.
[0398] 6. The method of any of solutions 1 to 4, wherein a right offset of the scaling window (subpicScalWinRightOffset) is derived as follows:
[0399] subpicScalWinRightOffset = (rightSubpicBd >= sps_pic_width_max_in_luma_samples)? pps_scaling_win_right_offset : pps_scaling_win_right_offset - (sps_pic_width_max_in_luma_samples - rightSubpicBd) / SubWidthC, and
[0400] where pps_scaling_win_right_offset indicates a right offset applied to picture size for scaling ratio calculation, rightSubpicBd indicates a width of the sub-picture, sps_pic_width_max_in_luma_samples indicates a maximum picture width in luma samples, and SubWidthC indicates a width of a video block and is obtained from a table according to a chroma format of a picture including the video block.
[0401] where sps_pic_width_max_in_luma_samples specifies the maximum width of each decoded picture of the reference sequence parameter set in luma samples, pps_scaling_win_right_offset indicates a right offset applied to picture size for scaling ratio calculation, SubWidthC indicates the width of a video block and is obtained from a table depending on the chroma format of the picture including the video block, and
[0402] rightSubpicBd = (sps_subpic_ctu_top_left_x[spIdx] + sps_subpic_width_minus1[spIdx] + 1) * CtbSizeY, and
[0403] where sps_subpic_ctu_top_left_x[spIdx] indicates the x-coordinate of a coding tree unit located at the top-left corner of the subpicture, sps_subpic_width_minus1[spIdx] indicates the width of the subpicture, and CtbSizeY is the width or height of a luma coding tree block or coding tree unit.
[0404] 7. The method according to any of solutions 1 to 4, wherein the top offset of the scaling window, subpicScalWinTopOffset, is derived as follows:
[0405] subpicScalWinTopOffset = pps_scaling_win_top_offset - sps_subpic_ctu_top_left_y[spIdx] * CtbSizeY / SubHeightC,
[0406] where pps_scaling_win_top_offset indicates a top offset applied to picture size for scaling ratio calculation, sps_subpic_ctu_top_left_y[spIdx] indicates the y-coordinate of a coding tree unit located at the top-left corner of the subpicture, CtbSizeY is the width or height of a luma coding tree block or coding tree unit, and SubHightC indicates the height of a video block and is obtained from a table depending on the chroma format of the picture including the video block.
[0407] 8. The method according to any of solutions 1 to 4, wherein the bottom offset of the scaling window, ubpicScalWinBotOffset, is derived as follows:
[0408] subpicScalWinBotOffset = (botSubpicBd >= sps_pic_height_max_in_luma_samples)? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - (sps_pic_height_max_in_luma_samples - botSubpicBd) / SubHeightC, and
[0409] where sps_pic_height_max_in_luma_samples indicates the maximum height of each decoded picture of the reference sequence parameter set in luma samples, pps_scaling_win_bottom_offset indicates the bottom offset applied to the picture size for scaling ratio calculation, SubHeightC indicates the height of a video block and is obtained from a table according to the chroma format of the picture including said video block, and
[0410] where botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] + SubHeightC) * CtbSizeY, and
[0411] sps_subpic_height_minus1[spIdx] + 1) * CtbSizeY, and
[0412] where sps_subpic_ctu_top_left_y[spIdx] indicates the y coordinate of a coding tree unit located at the top left corner of said subpicture, and CtbSizeY is the width or height of a luma coding tree block or coding tree unit.
[0413] 9. The method according to any of solutions 5 to 8, wherein sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minusl[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minusl[spIdx], sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are from the original sequence parameter set and pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset and pps_scaling_win_bottom_offset are from the original picture parameter set.
[0414] 10. The method according to any of solutions 1 to 3, wherein the rule specifies that the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset and pps_scaling_win_bottom_offset in all referenced picture parameter set (PPS) network abstraction layer (NAL) units are overwritten to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset and subpicScalWinBotOffset, respectively, and
[0415] wherein pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset and pps_scaling_win_bottom_offset indicate left offset, right offset, top offset, bottom offset applied to picture size for scaling ratio calculation, respectively, and
[0416] wherein subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset and subpicScalWinBotOffset indicate the left offset, right offset, top offset and bottom offset of the scaling window applicable to the subpicture, respectively.
[0417] 11. The method according to any of solutions 1 to 4, wherein the bottom offset of the scaling window, subpicScalWinBotOffset, is derived as follows:
[0418] subpicScalWinBotOffset = (botSubpicBd >= pps_pic_height_in_luma_samples)? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - (pps_pic_height_in_luma_samples - botSubpicBd) / SubHeightC, and
[0419] wherein pps_pic_height_max_in_luma_samples indicates the maximum height of each decoded picture of the picture parameter set in luma samples, pps_scaling_win_bottom_offset indicates the bottom offset applied to the picture size for scaling ratio calculation, SubHeightC indicates the height of a video block and is obtained from a table depending on the chroma format of the picture including the video block, and
[0420] wherein botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] + sps_subpic_height_minus1[spIdx] + 1) * CtbSizeY,
[0421] wherein sps_subpic_ctu_top_left_y[spIdx] indicates the y coordinate of a coding tree unit located at the top left corner of the subpicture, and CtbSizeY is the width or height of a luma coding tree block or coding tree unit.
[0422] 12. The method according to solution 11, wherein pps_pic_height_in_luma_samples is from the original picture parameter set.
[0423] 13. The method of solution 1, wherein the rule further specifies that a reference sample enable flag in the sequence parameter set that indicates applicability of reference picture resampling is overridden.
[0424] 14. The method of solution 1, wherein the rule further specifies that a resolution change allowance flag in the sequence parameter set that indicates possible picture spatial resolution changes within a coded layer video sequence (CLVS) referring to the sequence parameter set is overridden.
[0425] 15. The method of any of solutions 1 to 14, wherein the conversion comprises encoding the video into the bitstream.
[0426] 16. The method of any of solutions 1 to 14, wherein the conversion comprises decoding the video from the bitstream.
[0427] 17. The method of any of solutions 1 to 14, wherein the conversion comprises generating the bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0428] 18. A video processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 17.
[0429] 19. A method of storing a bitstream of a video, comprising a method recited in any one of solutions 1 to 17, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0430] 20. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of solutions 1 to 17.
[0431] 21. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0432] 22. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement a method recited in any one or more of solutions 1 to 17.
[0433] The third set of solutions shows example embodiments of the techniques discussed in the preceding sections (e.g., items 2 and 7).
[0434] 1. A video processing method (e.g., as in Figure 14AThe method 1400) is shown, comprising: converting 1402, according to a rule, between a video comprising one or more layers comprising one or more video pictures comprising one or more sub-pictures and a bitstream of the video, wherein the rule defines network abstraction layer (NAL) units to be extracted from the bitstream during a sub-bitstream extraction process to output a sub-bitstream, and wherein the rule further specifies one or more inputs to the sub-bitstream extraction process and / or which combination of sub-pictures of different layers of the bitstream to use, such that the output of the sub-bitstream extraction process conforms to a predefined format.
[0435] 2. The method of solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process comprise a target output layer set (OLS) index (targetOlsldx) that identifies an OLS index of a target OLS to be decoded and is equal to an index to a list of OLSs specified by a video parameter set.
[0436] 3. The method of solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process comprise a target highest temporal identifier value (tldTarget).
[0437] 4. The method of solution 3, wherein the target highest temporal identifier value is in a range of 0 to a maximum number of temporal sub-layers allowed to be present in layers specified by a video parameter set.
[0438] 5. The method of solution 4, wherein the maximum number of temporal sub-layers is indicated by a syntax element included in a video parameter set.
[0439] 6. The method of solution 5, wherein the syntax element is vps_max_sublayers_minusl.
[0440] 7. The method of solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process comprise a target sub-picture index value (subpicIdxTarget) that is equal to a sub-picture index present in a target OLS.
[0441] 8. The method of solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process comprise a list of target sub-picture index values subpicIdxTarget[i] for i from 0 to NumLayersInOls[targetOlsldx] - 1, wherein NumLayerInOls[i] specifies a number of layers in the i-th OLS and targetOlsldx indicates a target output layer set (OLS) index.
[0442] 9. The method of solution 8, wherein the rule specifies that all layers in the targetOLsIdx-th OLS have the same spatial resolution, the same subpicture layout, and all subpictures have syntax elements that specify that corresponding subpictures of each coded picture in the coded layer video sequence are treated as a picture in a decoding process that does not include in-loop filtering operations.
[0443] 10. The method of solution 8, wherein the value of subpicIdxTarget[i] is the same for all values of i and equal to a particular value in the range of 0 to sps num subpics minusl, inclusive, where sps num subpics minusl plus 1 specifies the number of subpictures in each picture in the coded layer video sequence.
[0444] 11. The method of solution 8, wherein the rule specifies that the output sub-bitstream contains at least one VCL (video coding layer) NAL (network abstraction layer) unit with nuh layer id equal to each of the nuh layer id values in the list of LayerldlnOls[targetOlsIdx], where nuh layer id is a NAL unit header identifier, and LayerldlnOls[targetOlsIdx] specifies the nuh layer id values in the OLS with the targetOLsIdx.
[0445] 12. The method of solution 8, wherein the rule specifies that the output sub-bitstream contains at least one VCL (video coding layer) NAL (network extraction layer) unit with Temporalld equal to tldTarget, where Temporalld indicates a target highest temporal identifier, and tldTarget is a value of Temporalld provided as one or more inputs to the sub-bitstream extraction process.
[0446] 13. The method of solution 8, wherein the bitstream contains one or more coded slice NAL (network abstraction layer) units with Temporalld equal to 0 without the need to contain a coded slice NAL (network abstraction layer) unit with nuh layer id equal to 0, where nuh layer id specifies a NAL unit header identifier.
[0447] 14. The method of solution 8, wherein the rule specifies that the output sub-bitstream contains at least one VCL (Video Coding Layer) NAL (Network Abstraction Layer) unit, where for each i in the range of 0 to NumLayersInOls[targetOlsIdx] - 1, the NAL (Network Abstraction Layer) unit header identifier nuh layer id is equal to the layer identifier a layerldInOls[targetOlsIdx][i] and the subpicture identifier sh subpic id is equal to the SubpicldVal[subpicldxTarget[i]], where NumLayerInOls[i] specifies the number of layers in the i-th OLS and i is an integer.
[0448] 15. A method of video processing (e.g., method 1410 as shown in FIG. 14), comprising converting 1412, according to a rule, between a video and a bitstream of the video, wherein the bitstream comprises a first layer and a second layer, the first layer comprises pictures with multiple subpictures, and the second layer comprises pictures each with a single subpicture, and wherein the rule specifies a combination of subpictures of the first layer and the second layer that, when extracted, results in an output bitstream that conforms to a predefined format. Figure 14B
[0449] 16. The method of solution 15, wherein the rule is applied to sub-bitstream extraction to provide an output sub-bitstream that is a conforming bitstream, and wherein input to the sub-bitstream extraction process comprises: i) the bitstream, ii) a target output layer set (OLS) index (targetOlsldx) that identifies a target OLS to be decoded, and ii) a target highest temporal identifier (Temporalld) value (tldTarget).
[0450] 17. The method of solution 16, wherein rule specifies that the target output layer set (OLS) index (targetOlsldx) is equal to an index into a list of OLSs specified by a video parameter set.
[0451] 18. The method of solution 16, wherein the rule specifies that the target highest temporal identifier (Temporalld) value (tldTarget) is equal to any value in the range of 0 to a maximum number of temporal sub-layers allowed to be present in layers specified by a video parameter set.
[0452] 19. The method of solution 18, wherein the maximum number of temporal sub-layers is indicated by a syntax element included in a video parameter set.
[0453] 20. The method of solution 19, wherein the gai syntax element is vps_max_sublayer_minusl.
[0454] 21. The method of solution 17, wherein the rule specifies that for i from 0 to NumLayersInOls[targetOLsIdx] - 1, the value subpicIdxTarget[i] of the OLS list is equal to a value in the range of 0 to the number of subpictures in each picture in the i-th OLS, where NumLayerInOls[i] specifies the number of layers in the i-th OLS, and the number of subpictures is indicated by a syntax element included in the sequence parameter set referred to by the layer associated with a network abstraction layer (NAL) unit header identifier (nuh layer id) equal to LayerldInOls[targetOLsIdx][i] specifying the nuh layer id value in the i-th OLS with the targetOLsIdx, and i is an integer.
[0455] 22. The method of solution 21, wherein another syntax element sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] equal to 1 specifies that the i-th subpicture of each coded picture in the layer is treated as a picture in the decoding process without including in-loop filtering operations.
[0456] 23. The method of solution 21, wherein the rule specifies that in case the syntax element indicating the number of subpictures in each picture in the layer is equal to 0, the value of subpicIdxTarget[i] is always equal to 0.
[0457] 24. The method of solution 21, wherein the rule specifies that for any two different integer values of m and n, the syntax element is greater than 0 for two layers with nuh layer id equal to LayerldInOls[targetOLsIdx][m] and LayerldInOls[targetOLsIdx][n] respectively, and subpicIdxTarget[m] is equal to subpicIdxTarget[n].
[0458] 25. The method as described in Solution 16, wherein the rule specifies that the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit whose nuh_layer_id is equal to each nuh_layer_id value in the list LayerIdInOls[targetOlsIdx], where nuh_layer_id is the NAL unit header identifier, and LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the OLS with the targetOlsIdx.
[0459] 26. The method of solution 16, wherein the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit having a TemporalId equal to tIdTarget.
[0460] 27. The method of solution 16, wherein the consistent bitstream contains one or more encoded stripe NAL (Network Abstraction Layer) units with a TemporalId equal to 0, without the need to contain encoded stripe NAL (Network Abstraction Layer) units with a nuh_layer_id equal to 0, wherein nuh_layer_id specifies the NAL unit header identifier.
[0461] 28. The method as described in Solution 16, wherein the output sub-bitstream contains at least one VCL (Video Coding Layer) NAL (Network Abstraction Layer) unit, the VCLNAL unit having a NAL (Network Abstraction Layer) unit header identifier nuh_layer_id equal to the layer identifier LayerIdInOls[targetOlsIdx][i], and for each i in the range from 0 to NumLayersInOls[targetOlsIdx]-1 having a subpic identifier sh_subpic_id equal to SubpicIdVal[subpicIdxTarget[i]], wherein NumLayerInOls[i] specifies the number of layers in the i-th OLS and i is an integer.
[0462] 29. A video processing method (e.g., such as...) Figure 14C The method 1420 shown includes: performing a 1422 conversion between a video and a bitstream of the video, and wherein a rule specifies whether or how an output layer set having one or more layers including multiple subpictures and / or one or more layers having a single subpicture is indicated by a list of subpicture identifiers in a Scalable Nested Supplemental Enhancement Information (SEI) message.
[0463] 30. The method of solution 29 wherein the rule specifies that, in the case that sn_ols_flag and sn_subpic_flag are both equal to 1, the list of subpicture identifiers specifies subpicture identifiers of subpictures in pictures each containing a plurality of subpictures, wherein sn_ols_flag equal to 1 specifies that the scalable nesting SEI message applies to a particular OLS and sn_subpic_flag equal to 1 specifies that the scalable nesting SEI message that applies to a specified OLS or layer only applies to particular subpictures of the specified OLS or layer.
[0464] 31. The method of any of solutions 1 to 30 wherein the conversion comprises encoding the video into the bitstream.
[0465] 32. The method of any of solutions 1 to 30 wherein the conversion comprises decoding the video from the bitstream.
[0466] 33. The method of any of solutions 1 to 30 wherein the conversion comprises generating the bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0467] 34. A video processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 33.
[0468] 35. A method of storing a bitstream of a video, comprising a method recited in any one of solutions 1 to 33, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0469] 36. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of solutions 1 to 33.
[0470] 37. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0471] 38. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement a method recited in any one or more of solutions 1 to 33.
[0472] The fourth set of solutions shows example embodiments of the techniques discussed in the preceding sections (e.g., items 3, 8, and 9).
[0473] 1. A video processing method (e.g., as in Figure 15AThe method 1500 shown includes: performing a 1502 conversion between a video including one or more video images in the video layer and the bitstream of the video according to a rule, and wherein the rule specifies that: in the sub-bitstream extraction process, regardless of the availability of external components used to replace the parameter set removed during the sub-bitstream extraction, the removal of (i) Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, (ii) padding data NAL units associated with the VCLNAL units, and (iii) padding payload supplemental enhancement information (SEI) messages associated with the VCLNAL units are performed.
[0474] 2. The method as described in Solution 1, wherein the rule specifies that: for each value of i in the range of 0 to NumLayersInOls[targetOlsIdx]-1, remove from the sub-bitstream, which is the output of the sub-bitstream extraction process, all Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, their associated padding data NAL units, and their associated SEI containing padding payload SEI messages, all with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id not equal to SubpicIdVal[subpicIdxTarget[i]]. The NAL unit is defined as follows: nuh_layer_id is the NAL unit header identifier; LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the output layer set (OLS), where targetOLsIdx is the index of the target output layer set; sh_subpic_id specifies the subpick identifier containing the subpick of the stripe; subpicIdxTarget[i] indicates the target subpick index value for i; and SubpicIdVal[subpicIdxTarget[i]] is the variable used for subpicIdxTarget[i], where i is an integer.
[0475] 3. A video processing method (e.g., such as...) Figure 15B The method 1510 shown includes: performing a 1512 conversion between a video comprising one or more layers and a bitstream of the video according to a rule, the one or more layers comprising one or more video images, the one or more video images comprising one or more sub-images, and wherein the rule specifies that: in the case where a reference image is partitioned according to a sub-image layout different from the sub-image layout, the reference image is not allowed to be used as a co-bit image of the current image that is partitioned into sub-images using the sub-image layout.
[0476] 4. The method of solution 3, wherein the current picture and the reference picture are in different layers.
[0477] 5. A method of video processing (e.g., method 1520 as shown in FIG. 15), comprising converting 1522, according to a rule, between a video comprising one or more layers comprising one or more video pictures comprising one or more subpictures and a bitstream of the video, and wherein the rule specifies that a coding tool is disabled during the conversion of a current picture that is partitioned into subpictures using a subpicture layout that is different from a subpicture layout of a reference picture used for the current picture, in a case that the coding tool depends on the subpicture layout. Figure 15C
[0478] 6. The method of solution 5, wherein the current picture and the reference picture are in different layers.
[0479] 7. The method of solution 5, wherein the coding tool is bi-directional optical flow (BDOF), wherein one or more initial predictions are refined using optical flow calculations.
[0480] 8. The method of solution 5, wherein the coding tool is decoder-side motion vector refinement (DMVR), wherein motion information is refined by using a prediction block.
[0481] 9. The method of solution 5, wherein the coding tool is prediction refinement with optical flow (PROF), wherein one or more initial predictions are refined based on optical flow calculations.
[0482] 10. The method of any of solutions 1 to 9, wherein the converting comprises encoding the video into the bitstream.
[0483] 11. The method of any of solutions 1 to 9, wherein the converting comprises decoding the video from the bitstream.
[0484] 12. The method of any of solutions 1 to 9, wherein the converting comprises generating the bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0485] 13. A video processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 12.
[0486] 14. A method of storing a bitstream of a video, comprising a method recited in any one of solutions 1 to 12, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0487] 15. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of solutions 1 to 12.
[0488] 16. A computer-readable medium storing a bitstream generated according to any of the above methods.
[0489] 17. A video processing apparatus configured to implement a method recited in any one or more of solutions 1 to 12 for storing a bitstream representation.
[0490] The fifth set of solutions shows example embodiments of the techniques discussed in the preceding sections (e.g., item 4).
[0491] 1. A video processing method (e.g., method 1600 as shown in Figure 16 FIG. 1), comprising converting, according to a rule, between a video comprising one or more video pictures in a video layer and a bitstream of the video, wherein the rule specifies that a supplemental enhancement information (SEI) network abstraction layer (NAL) unit containing an SEI message having a particular payload type does not contain another SEI message having a payload type different from the particular payload type.
[0492] 2. The method of solution 1, wherein the rule specifies removal of the SEI message as removal of the SEI NAL unit containing the SEI message.
[0493] 3. The method of solution 1 or 2, wherein the SEI message having the particular payload type corresponds to a filler payload SEI message.
[0494] 4. The method of any of solutions 1 to 3, wherein the converting comprises encoding the video into the bitstream.
[0495] 5. The method of any of solutions 1 to 3, wherein the converting comprises decoding the video from the bitstream.
[0496] 6. The method of any of solutions 1 to 3, wherein the converting comprises generating the bitstream from the video, and the method further comprises storing the bitstream in a non-transitory computer-readable recording medium.
[0497] 7. A video processing apparatus comprising a processor configured to implement a method recited in any one or more of solutions 1 to 6.
[0498] 8. A method of storing a bitstream of a video, the method comprising the method recited in any one of solutions 1 to 6, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0499] 9. A computer-readable medium storing program code that, when executed, causes a processor to implement the method recited in any one or more of solutions 1 to 6.
[0500] 10. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0501] 11. A video processing apparatus configured to store a bitstream representation, wherein the video processing apparatus is configured to implement the method recited in any one or more of solutions 1 to 6.
[0502] In the solutions described herein, an encoder can conform to a format rule by producing a coded representation according to the format rule. In the solutions described herein, a decoder can parse syntax elements in a coded representation using a format rule to produce a decoded video, where the presence and absence of syntax elements is known according to the format rule.
[0503] In this document, the term “video processing” can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. A bitstream representation of a current video block can correspond to bits that are collectively located or scattered at different locations within the bitstream, as defined by syntax. For example, a macroblock can be encoded according to transformed and coded error residual values, and can also be encoded using bits in a header and other fields in the bitstream. Furthermore, during a conversion, a decoder can parse a bitstream with knowledge that certain fields can or can not be present, based on determinations described in the above solutions. Similarly, an encoder can determine that certain syntax fields are included or not included, and generate a coded representation accordingly by including or excluding syntax fields from the coded representation.
[0504] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder will use or implement the tool or mode in the processing of a video block, but does not necessarily modify the resulting bitstream based on the use of the tool or mode. In other words, when the conversion of a video block to a bitstream representation of a video will be based on a decision or determination to enable a video processing tool or mode, the conversion will use the video processing tool or mode that is enabled. In another example, when a video processing tool or mode is enabled, a decoder will use knowledge that a bitstream has been modified based on the video processing tool or mode to process the bitstream. In other words, the conversion from a bitstream representation of a video to a video block will use a video processing tool or mode that is enabled based on a decision or determination.
[0505] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use the tool or mode in the conversion of a video block to a bitstream representation of a video. In another example, when a video processing tool or mode is disabled, a decoder will use a video processing tool or mode that is disabled based on a decision or determination to process a bitstream with knowledge that the bitstream has not been modified.
[0506] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated for a purpose of encoding information for transmission to suitable receiver apparatus.
[0507] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0508] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and / or devices that are
[0509] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0510] While this patent document contains many details, these should not be construed as limiting the scope of any subject matter or potentially patentable content in any way, but as a description of features that can be particular to certain embodiments of the particular technology. Certain features described in the context of separate embodiments in this patent document can also be implemented in combination with each other. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any appropriate subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, in some cases, the features from a claimed combination can be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0511] Similarly, while operations are described in a particular, sequential order, this should not be understood as requiring that such operations be performed in the particular order described or in sequential order, or that all illustrated operations be performed, to implement desirable results. Further, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0512] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method of video processing, comprising: converting between a video comprising one or more video pictures comprising one or more subpictures and a bitstream of the video, wherein the bitstream conforms to a rule that specifies that during subpicture sub-bitstream extraction processing, one or more parameters of a scaling window applicable to a subpicture are determined by overwriting values of a first set of syntax elements included in one or more picture parameter sets (PPSs), the overwriting values of the first set of syntax elements are equal to a result of a calculation of original values of the first set of syntax elements from the original one or more PPSs and original values of a second set of syntax elements from an original one or more sequence parameter sets (SPSs).
2. The method of claim 1, wherein, the one or more parameters comprise at least one of a left offset, a right offset, a top offset, or a bottom offset of the scaling window applicable to the subpicture.
3. The method of claim 1, wherein, the first set of syntax elements comprises at least one of: pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset; wherein pps_scaling_win_left_offset indicates an original left offset applied to a picture size for scaling rate calculation, pps_scaling_win_right_offset indicates an original right offset applied to the picture size for the scaling rate calculation, pps_scaling_win_top_offset indicates an original top offset applied to the picture size for the scaling rate calculation, and pps_scaling_win_bottom_offset indicates an original bottom offset applied to the picture size for the scaling rate calculation; and wherein the second set of syntax elements comprises at least one of: sps_subpic_ctu_top_left_x[ spIdx ], sps_subpic_width_minus1[ spIdx ], sps_subpic_ctu_top_left_y[ spIdx ], sps_subpic_height_minus1[ spIdx ], sps_pic_width_max_in_luma_samples, and sps_pic_height_max_in_luma_samples; wherein sps_subpic_ctu_top_left_x[ spIdx ] specifies the original x-coordinate of a coding tree unit located at the top-left corner of the sth spIdx subpicture, sps_subpic_width_minus1[ spIdx ] specifies the original width of the sth spIdx subpicture, sps_subpic_ctu_top_left_y[ spIdx ] specifies the original y-coordinate of a coding tree unit located at the top-left corner of the sth spIdx subpicture, sps_subpic_height_minus1[ spIdx ] specifies the original height of the sth spIdx subpicture, sps_pic_width_max_in_luma_samples specifies the original maximum width of each decoded picture of the reference sequence parameter set in units of luma samples, and sps_pic_height_max_in_luma_samples specifies the original maximum height of each decoded picture of the reference sequence parameter set in units of luma samples.
4. The method of claim 1, wherein, The one or more parameters include a right offset of the scaling window applicable to the subpicture derived as follows: subpicScalWinLeftOffset = pps_scaling_win_left_offset - sps_subpic_ctu_top_left_x[ spIdx ] CtbSizeY / SubWidthC, wherein subpicScalWinRightOffset represents a right offset of the scaling window applicable to the subpicture, wherein pps_scaling_win_left_offset specifies the original left offset for scaling ratio calculation applied to picture dimensions, wherein sps_subpic_ctu_top_left_x[ spIdx ] specifies the original x-coordinate of a coding tree unit located at the top-left corner of the sth spIdx subpicture, wherein CtbSizeY is the width or height of a luma coding tree block or coding tree unit, and wherein SubWidthC specifies the width of a video block and is obtained from a table according to a chroma format of a video picture that includes the video block.
5. The method of claim 1, wherein, The one or more parameters include a right offset of the scaling window derived as follows: wherein subpicScalWinRightOffset represents a right offset of the scaling window applicable to the subpicture, wherein pps_scaling_win_left_offset specifies the original left offset for scaling ratio calculation applied to picture dimensions, wherein sps_subpic_ctu_top_left_x[ spIdx ] specifies the original x-coordinate of a coding tree unit located at the top-left corner of the sth spIdx subpicture, wherein CtbSizeY is the width or height of a luma coding tree block or coding tree unit, and wherein SubWidthC specifies the width of a video block and is obtained from a table according to a chroma format of a video picture that includes the video block. The one or more parameters include a right offset of the scaling window derived as follows: wherein subpicScalWinRightOffset represents a right offset of the scaling window applicable to the subpicture, wherein pps_scaling_win_left_offset specifies the original left offset for scaling ratio calculation applied to picture dimensions, wherein sps_subpic_ctu_top_left_x[ spIdx ] specifies the original x-coordinate of a coding tree unit located at the top-left corner of the sth spIdx subpicture, wherein CtbSizeY is the width or height of a luma coding tree block or coding tree unit, and wherein SubWidthC specifies the width of a video block and is obtained from a table according to a chroma format of a video picture that includes the video block. subpicScalWinRightOffset = ( rightSubpicBd >= sps_pic_width_max_in_luma_samples )? pps_scaling_win_right_offset : pps_scaling_win_right_offset - sps_subpic_ctu_top_left_y[ spIdx ] + sps_subpic_height_minus1[ spIdx ] + 1 ) / SubHeightC, and rightSubpicBd = ( sps_pic_width_max_in_luma_samples - leftSubpicBd ) / SubWidthC, and leftSubpicBd = ( sps_subpic_ctu_top_left_x[ spIdx ] + 1 ) CtbSizeY, wherein subpicScalWinRightOffset represents a right offset of the scaling window applicable to the subpicture, wherein sps_pic_width_max_in_luma_samples specifies the original maximum width of each decoded picture of the reference sequence parameter set in luma samples, wherein pps_scaling_win_right_offset indicates an original right offset for scaling rate calculation applied to picture dimensions, wherein SubWidthC indicates a width of a video block and is obtained from a table according to a chroma format of a video picture comprising the video block, wherein sps_subpic_ctu_top_left_x[ spIdx ] indicates an original x-coordinate of a coding tree unit located at the top-left corner of the spIdx-th subpicture, and wherein sps_subpic_width_minus1[ spIdx ] indicates an original width of the spIdx-th subpicture, and wherein CtbSizeY is a width or height of a luma coding tree block or coding tree unit.
6. The method of claim 1, wherein, The one or more parameters comprise a top offset of the scaling window derived as follows: subpicScalWinTopOffset = pps_scaling_win_top_offset - sps_subpic_ctu_top_left_y[ spIdx ] CtbSizeY / SubHeightC, wherein subpicScalWinTopOffset represents a top offset of the scaling window applicable to the subpicture, wherein pps_scaling_win_top_offset indicates an original top offset for scaling rate calculation applied to picture dimensions, wherein sps_subpic_ctu_top_left_y[ spIdx ] indicates an original y-coordinate of a coding tree unit located at the top-left corner of the spIdx-th subpicture, wherein CtbSizeY is a width or height of a luma coding tree block or coding tree unit, and wherein SubHightC indicates a height of a video block and is obtained from a table according to a chroma format of a video picture comprising the video block.
7. The method of claim 1, wherein, The one or more parameters comprise a bottom offset of the scaling window applicable to the subpicture derived as follows: subpicScalWinBottomOffset = ( bottomSubpicBd >= sps_pic_height_max_in_luma_samples )? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - wherein subpicScalWinBottomOffset represents a bottom offset of the scaling window applicable to the subpicture, wherein sps_pic_height_max_in_luma_samples specifies the original maximum height of each decoded picture of the reference sequence parameter set in luma samples, wherein pps_scaling_win_bottom_offset indicates an original bottom offset for scaling rate calculation applied to picture dimensions, wherein sps_subpic_ctu_top_left_y[ spIdx ] indicates an original y-coordinate of a coding tree unit located at the top-left corner of the spIdx-th subpicture, and wherein sps_subpic_height_minus1[ spIdx ] indicates an original height of the spIdx-th subpicture, and wherein CtbSizeY is a width or height of a luma coding tree block or coding tree unit. subpicScalWinBotOffset = ( botSubpicBd >= sps_pic_height_max_in_luma_samples )? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - ( sps_pic_height_max_in_luma_samples - botSubpicBd ) / SubHeightC, and botSubpicBd = ( sps_subpic_ctu_top_left_y[ spIdx ] + sps_subpic_height_minus1[ spIdx ] + 1 ) CtbSizeY, wherein subpicScalWinBotOffset represents a bottom offset of the scaling window applicable to the subpicture, wherein sps_pic_height_max_in_luma_samples specifies the original maximum height of each decoded picture of the reference sequence parameter set in luma samples, wherein pps_scaling_win_bottom_offset indicates an original bottom offset for scaling ratio calculation applied to picture dimensions, wherein SubHeightC indicates a height of a video block and is obtained from a table according to a chroma format of a picture comprising the video block, wherein sps_subpic_ctu_top_left_y[ spIdx ] indicates an original y-coordinate of a coding tree unit located at the top-left corner of the spIdx-th subpicture; wherein sps_subpic_height_minus1[ spIdx ] indicates an original height of the spIdx-th subpicture, and wherein CtbSizeY is a width or height of a luma coding tree block or coding tree unit.
8. The method of claim 1, wherein, The conversion comprises encoding the video into the bitstream.
9. The method of claim 1, wherein, The conversion comprises decoding the video from the bitstream.
10. An apparatus for processing video data, comprising a processor and a non-transitory memory with instructions thereon, wherein, The instructions, when executed by the processor, cause the processor to: convert between a video comprising one or more video pictures including one or more subpictures and a bitstream of the video, wherein the bitstream conforms to a rule that specifies that during subpicture sub-bitstream extraction processing, one or more parameters of a scaling window applicable to a subpicture are determined by overwriting values of a first set of syntax elements included in one or more picture parameter sets (PPSs), the overwriting values of the first set of syntax elements being equal to a result of a calculation of original values of the first set of syntax elements from the original one or more PPSs and original values of a second set of syntax elements from an original one or more sequence parameter sets (SPSs).
11. The apparatus of claim 10, wherein, the one or more parameters comprise at least one of a left offset, a right offset, a top offset or a bottom offset of the scaling window applicable to the subpicture.
12. The apparatus of claim 10, wherein, The first set of syntax elements includes at least one of: pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset; where pps_scaling_win_left_offset indicates a raw left offset for scaling window calculation applied to a picture size, pps_scaling_win_right_offset indicates a raw right offset for the scaling window calculation applied to the picture size, pps_scaling_win_top_offset indicates a raw top offset for the scaling window calculation applied to the picture size, and pps_scaling_win_bottom_offset indicates a raw bottom offset for the scaling window calculation applied to the picture size; and where the second set of syntax elements includes at least one of: sps_subpic_ctu_top_left_x[ spIdx ], sps_subpic_width_minus1[ spIdx ], sps_subpic_ctu_top_left_y[ spIdx ], sps_subpic_height_minus1[ spIdx ], sps_pic_width_max_in_luma_samples, and sps_pic_height_max_in_luma_samples; where sps_subpic_ctu_top_left_x[ spIdx ] indicates a raw x-coordinate of a coding tree unit at the top-left corner of an spIdx-th subpicture, sps_subpic_width_minus1[ spIdx ] indicates a raw width of the spIdx-th subpicture, sps_subpic_ctu_top_left_y[ spIdx ] indicates a raw y-coordinate of a coding tree unit at the top-left corner of the spIdx-th subpicture, sps_subpic_height_minus1[ spIdx ] indicates a raw height of the spIdx-th subpicture, sps_pic_width_max_in_luma_samples specifies a raw maximum width of each decoded picture of a reference sequence parameter set in luma samples, and sps_pic_height_max_in_luma_samples specifies a raw maximum height of each decoded picture of the reference sequence parameter set in luma samples.
13. The apparatus of claim 10, wherein, The one or more parameters include a left offset of the scaling window for the subpicture derived as follows: pps_scaling_win_left_offset = sps_subpic_ctu_top_left_x[ spIdx ] + sps_subpic_width_minus1[ spIdx ] + 1 subpicScalWinLeftOffset = pps_scaling_win_left_offset - sps_subpic_ctu_top_left_x[ spIdx ] CtbSizeY / SubWidthC, Wherein, subpicScalWinLeftOffset represents the left offset of the scaling window applicable to the subpicture. Here, pps_scaling_win_left_offset indicates the original left offset applied to the image size for scaling calculation. Where sps_subpic_ctu_top_left_x[ spIdx ] indicates the original x-coordinate of the encoder-decoder tree unit located at the top left corner of the spIdx-th subpicture. Where CtbSizeY is the width or height of the luma codec block or codec unit, and Wherein, SubWidthC indicates the width of the video block and is obtained from the table according to the chroma format of the video picture including the video block.
14. The apparatus of claim 10, wherein, The one or more parameters include the right offset of the scaling window derived as follows: subpicScalWinRightOffset = ( rightSubpicBd >= sps_pic_width_max_in_luma_samples ) ?pps_scaling_win_right_offset : pps_scaling_win_right_offset - ( sps_pic_width_max_in_luma_samples - rightSubpicBd ) / SubWidthC, and rightSubpicBd = ( sps_subpic_ctu_top_left_x[ spIdx ] + sps_subpic_width_minus1[ spIdx ] + 1 ) CtbSizeY, Wherein, subpicScalWinRightOffset represents the right offset of the scaling window applicable to the subpicture. Here, `sps_pic_width_max_in_luma_samples` specifies the original maximum width of each decoded image in the reference sequence parameter set, in units of luminance samples. Here, pps_scaling_win_right_offset indicates the original right offset applied to the image size for scaling calculation. Wherein, SubWidthC indicates the width of the video block and is obtained from a table based on the chroma format of the video picture including the video block. Where, sps_subpic_ctu_top_left_x[ spIdx ] indicates the original x-coordinate of the encoder-decoder tree unit located at the top left corner of the spIdx-th subpicture, and Wherein, sps_subpic_width_minus1[ spIdx ] indicates the original width of the spIdx-th subpicture, and Where CtbSizeY is the width or height of the luma codec tree block or codec tree unit.
15. The apparatus of claim 10, wherein, The one or more parameters include the top offset of the scaling window derived as follows: subpicScalWinTopOffset = pps_scaling_win_top_offset - sps_subpic_ctu_top_left_y[ spIdx ] CtbSizeY / SubHeightC, Wherein, subpicScalWinTopOffset represents the top offset of the scaling window applicable to the subpicture. Here, pps_scaling_win_top_offset indicates the original top offset used for scaling calculations when the image size is calculated. sps_subpic_ctu_top_left_y[ spIdx ] indicates the original y-coordinate of a coding tree unit at the top-left corner of the spIdx-th subpicture, where CtbSizeY is the width or height of a luma coding tree block or coding tree unit, and where SubHeightC indicates the height of a video block, and is obtained from a table according to a chroma format of a picture including the video block.
16. The apparatus of claim 10, wherein, The one or more parameters include a bottom offset of the scaling window applicable to the subpicture derived as follows: subpicScalWinBotOffset = ( botSubpicBd >= sps_pic_height_max_in_luma_samples )? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - ( sps_pic_height_max_in_luma_samples - botSubpicBd ) / SubHeightC, and botSubpicBd = ( sps_subpic_ctu_top_left_y[ spIdx ] + sps_subpic_height_minus1[ spIdx ] + 1 ) CtbSizeY, where subpicScalWinBotOffset represents a bottom offset of the scaling window applicable to the subpicture, where sps_pic_height_max_in_luma_samples specifies the original maximum height of each decoded picture of a reference sequence parameter set in luma samples, where pps_scaling_win_bottom_offset indicates an original bottom offset for scaling rate calculation applied to picture dimensions, where SubHeightC indicates the height of a video block, and is obtained from a table according to a chroma format of a picture including the video block, where sps_subpic_ctu_top_left_y[ spIdx ] indicates the original y-coordinate of a coding tree unit at the top-left corner of the spIdx-th subpicture; where sps_subpic_height_minusl[ spIdx ] indicates the original height of the spIdx-th subpicture, and where CtbSizeY is the width or height of a luma coding tree block or coding tree unit.
17. A non-transitory computer-readable storage medium storing instructions to cause a processor to: convert, between a video comprising one or more video pictures including one or more subpictures and a bitstream of the video, wherein, the bitstream conforms to a rule that specifies that, during subpicture sub-bitstream extraction process, one or more parameters applicable to a scaling window of a subpicture are determined by overwriting values of a first set of syntax elements included in one or more picture parameter sets (PPSs), the overwriting values of the first set of syntax elements are equal to a result of a calculation of original values of the first set of syntax elements from the original one or more PPSs and original values of a second set of syntax elements from an original one or more sequence parameter sets (SPSs).
18. The non-transitory computer-readable storage medium of claim 17, wherein, the one or more parameters include at least one of a left offset, a right offset, a top offset, or a bottom offset applicable to the scaling window of the subpicture. 19.A non-transitory computer-readable recording medium storing a bitstream of a video, on which a computer program is further stored, wherein, the computer program, when executed by a processor, implements the method of any one of claims 1-8 to generate the bitstream.
20. A method for storing a bitstream of a video, comprising: generating a bitstream for a video comprising one or more video pictures, the one or more video pictures comprising one or more sub-pictures; and storing the bitstream in a non-transitory computer-readable recording medium, wherein the bitstream conforms to a rule that specifies that, during subpicture sub-bitstream extraction process, one or more parameters applicable to a scaling window of a subpicture are determined by overwriting values of a first set of syntax elements included in one or more picture parameter sets (PPSs), the overwriting values of the first set of syntax elements are equal to a result of a calculation of original values of the first set of syntax elements from the original one or more PPSs and original values of a second set of syntax elements from an original one or more sequence parameter sets (SPSs).
Citation Information
Patent Citations
Real-time image processing for optimizing sub-images views
CN104685536A
Information processing apparatus and method for generating composite image
US20090080802A1