Processing of coded video in sub-bitstream extraction processing
By updating the video parameter set and removing unnecessary NAL units in VVC, the problem of inconsistent scaling windows in sub-image sub-bitstream extraction processing is resolved, improving the efficiency and quality of video encoding and decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2021-05-21
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies struggle to maintain consistency between the scaling window and the original bitstream encoding during sub-image bitstream extraction in VVC, particularly when extracting rectangular regions from image sequences, leading to scaling window mismatch.
By redefining and updating the parameter set, including the syntax elements of VPS, SPS, and PPS, during the sub-bitstream extraction process, the correctness of the scaling window is ensured, and unnecessary NAL units and SEI messages are removed, ensuring that the output bitstream conforms to the predefined format.
It achieves consistency of scaling window in sub-bitstream extraction processing, improves the efficiency and quality of video encoding and decoding, and adapts to the needs of multi-layer video encoding and decoding.
Smart Images

Figure CN115868167B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application is the Chinese national phase application of International Patent Application No. PCT / CN2021 / 095184, filed on May 21, 2021, and promptly claims priority and benefit to International Patent Application No. PCT / CN2020 / 091696, filed on May 22, 2020. The entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to image and video encoding and decoding. Background Technology
[0004] Digital video consumes the largest amount of bandwidth on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to process the codec representation of video using control information useful for decoding the codec representation.
[0006] In one example aspect, a video processing method is disclosed. The method includes: converting a video comprising one or more video images and a bitstream of the video, the one or more video images comprising one or more sub-images, wherein the conversion conforms to a rule specifying one or more parameters for determining a scaling window applicable to the sub-images from one or more syntax elements during sub-image sub-bitstream extraction processing.
[0007] In another example, a different video processing method is disclosed. This method includes: converting between a video comprising one or more layers and a bitstream of the video, the one or more layers comprising one or more video images, the one or more video images comprising one or more sub-images, according to a rule, wherein the rule defines which Network Abstraction Layer (NAL) units to extract from the bitstream to output a sub-bitstream during sub-bitstream extraction processing, and wherein the rule also specifies which combination of one or more inputs to the sub-bitstream extraction processing and / or sub-images from different layers of the bitstream is used such that the output of the sub-bitstream extraction processing conforms to a predefined format.
[0008] In another example, a different video processing method is disclosed. This method includes: converting a video between a video and a bitstream of the video according to a rule, wherein the bitstream includes a first layer and a second layer, the first layer including images having multiple sub-images, the second layer including images each having a single sub-image, and wherein the rule specifies a combination of sub-images from the first and second layers that results in an output bitstream conforming to a predefined format during extraction.
[0009] In another example, a different video processing method is disclosed. This method includes: converting between a video and a bitstream of the video, and wherein a rule specifies whether and how an output layer set is indicated by a list of subpicture identifiers in a Scalable Nested Supplemental Enhancement Information (SEI) message, having an output layer set with one or more layers comprising multiple subpictures and / or one or more layers having a single subpicture.
[0010] In another example, a different video processing method is disclosed. This method includes: converting between a video comprising one or more video images in a video layer and a bitstream of that video according to a rule, wherein the rule specifies that, during the sub-bitstream extraction process, regardless of the availability of external components used to replace the set of parameters removed during the sub-bitstream extraction, the removal of (i) Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, (ii) padding data NAL units associated with the VCL NAL units, and (iii) padding payload supplemental enhancement information (SEI) messages associated with the VCL NAL units are performed.
[0011] In another example, a different video processing method is disclosed. This method includes: converting between a video comprising one or more layers and a bitstream of the video, according to a rule, wherein the one or more layers comprise one or more video images, the one or more video images comprise one or more sub-images, and wherein the rule specifies that, in the case where a reference image is partitioned according to a sub-image layout different from the sub-image layout, it is not permitted to use the reference image as a co-bit image of the current image partitioned into sub-images using that sub-image layout.
[0012] In another example, a different video processing method is disclosed. This method includes: converting between a video comprising one or more layers and a bitstream of the video, according to a rule, where the one or more layers comprise one or more video images, the one or more video images comprise one or more sub-images, and wherein the rule specifies that the codec tool is disabled during the conversion of the current image, which is divided into sub-images using that sub-image layout, if the codec tool relies on a sub-image layout different from that of a reference image.
[0013] In another example, a different video processing method is disclosed. This method includes converting between a video comprising one or more video images in a video layer and a bitstream of that video, according to a rule specifying that a Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) unit containing an SEI message with a specific payload type does not contain another SEI message with a payload type different from that specific payload type.
[0014] In yet another example, a video encoder apparatus is disclosed. This video encoder includes a processor configured to implement the methods described above.
[0015] In yet another example, a video decoder apparatus is disclosed. This video decoder includes a processor configured to implement the methods described above.
[0016] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.
[0017] These and other features are described throughout this document. Attached Figure Description
[0018] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 patches and 3 raster scan strips.
[0019] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular strips.
[0020] Figure 3 An example of an image divided into tiles and rectangular strips is shown, where the image is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular strips.
[0021] Figure 4 The image is shown as being divided into 15 tiles, 24 strips, and 24 sub-images.
[0022] Figure 5 This illustrates a typical sub-image-based viewport-dependent 360. o Video encoding and decoding solutions.
[0023] Figure 6 This demonstrates an improved viewport-dependent 360 based on sub-pictures and spatial scalability. o Video encoding and decoding solutions.
[0024] Figure 7 This is a block diagram of an example video processing system.
[0025] Figure 8 This is a block diagram of a video processing device.
[0026] Figure 9 This is a flowchart of an example method for video processing.
[0027] Figure 10 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.
[0028] Figure 11 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0029] Figure 12 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0030] Figure 13 A flowchart is shown illustrating an example method for video processing based on some implementations of the disclosed technology.
[0031] Figures 14A to 14C A flowchart is shown illustrating an example method for video processing based on some implementations of the disclosed technology.
[0032] Figures 15A to 15C A flowchart is shown illustrating an example method for video processing based on some implementations of the disclosed technology.
[0033] Figure 16 A flowchart is shown illustrating an example method for video processing based on some implementations of the disclosed technology. Detailed Implementation
[0034] The use of chapter headings in this document is for ease of understanding and not to limit the applicability of the techniques and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs.
[0035] 1. Introduction
[0036] This document relates to video codec techniques. Specifically, it concerns subpicture sub-bitstream extraction processing, scalable nested SEI messages, and subpicture level information SEI messages. These ideas can be applied individually or in various combinations to any video codec standard or non-standard video codec that supports multi-layer video codecs, such as the Universal Video Codec (VVC) currently under development.
[0037] 2. Abbreviation
[0038] APS Adaptation Parameter Set
[0039] AU Access Unit
[0040] AUD Access Unit Delimiter
[0041] AVC Advanced Video Codec
[0042] CLVS codec layer video sequence
[0043] CPB image buffer
[0044] CRA Clean Random Access
[0045] CTU (Codec Tree Unit)
[0046] CVSC codec video sequence
[0047] DCI decoding capability information
[0048] DPB Decoding Image Buffer
[0049] EOB bitstream end
[0050] EOS sequence ends
[0051] GDR Gradual Decoding and Refresh
[0052] HEVC High Efficiency Video Encoding and Decoding
[0053] HRD Assumption Reference Decoder
[0054] IDR Instantaneous Decoding and Refresh
[0055] Interlayer prediction in ILP
[0056] ILRP interlayer reference image
[0057] JEM Joint Exploration Model
[0058] LTRP Long-Term Reference Image
[0059] MCTS Motion Constraint Plot Set
[0060] NAL Network Abstraction Layer
[0061] OLS output layer set
[0062] PH Image Header
[0063] PPS Image Parameter Set
[0064] PTLP stands for Profile, Tier, and Level.
[0065] PU Image Unit
[0066] RAP Random Access Point
[0067] RBSP raw byte sequence payload
[0068] SEI Supplemental Enhancement Information
[0069] SPS Sequence Parameter Set
[0070] STRP Short-Term Reference Image
[0071] SVC Scalable Video Codec
[0072] VCL (Video Codec Layer)
[0073] VPS Video Parameter Set
[0074] VTM VVC Test Model
[0075] VUI Video Availability Information
[0076] VVC Universal Video Codec
[0077] 3. Initial Discussion
[0078] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established in 2015 by VCEG and MPEG. Since then, JVET has adopted many new approaches and incorporated them into reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the new codec standards aim for a 50% bitrate reduction compared to HEVC. The new video codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Due to ongoing efforts to standardize VVC, new codec technologies have been adopted at each JVET meeting. The VVC working draft and the VTM test model are then updated after each meeting. The current goal of the VVC project is to achieve Technical Finalization (FDIS) at the meeting in July 2020.
[0079] 3.1. Image partitioning schemes in HEVC
[0080] HEVC includes four different image segmentation schemes: regular striping, correlated striping, tile, and wavefront parallel processing (WPP), which can be applied to maximum transmission unit (MTU) size matching, parallel processing, and reduced end-to-end latency.
[0081] The regular stripes are similar to those in H.264 / AVC. Each regular stripe is encapsulated in its own NAL unit, and intra-picture predictions (intra-sample prediction, motion information prediction, and encoding / decoding mode prediction) and entropy encoding / decoding correlations across stripe boundaries are disabled. Therefore, regular stripes can be reconstructed independently of other regular stripes within the same picture (although cross-correlation may still exist due to loop filtering operations).
[0082] Regular striping is the only tool available for parallelization, and it is available in virtually the same form in H.264 / AVC. Parallelization based on regular striping requires minimal inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predicted images, which is typically much heavier than inter-processor or inter-core data sharing due to intra-image prediction). However, for the same reason, using regular striping can incur significant encoding / decoding overhead due to the bit cost of the stripe headers and the lack of prediction at stripe boundaries. Furthermore, due to the intra-image independence of regular striping and the fact that each regular stripe is encapsulated in its own NAL unit, regular striping (as opposed to other tools mentioned below) also serves as a crucial mechanism for bitstream segmentation to match MTU size requirements. In many cases, the goals of parallelization and MTU size matching conflict with the requirements for stripe layout within the image. The implementation of such situations led to the development of the parallelization tools mentioned below.
[0083] Correlated stripes have short strip headers and allow the bitstream to be segmented at tree block boundaries without disrupting any in-picture predictions. Essentially, correlated stripes segment regular stripes into multiple NAL units to provide reduced end-to-end latency by allowing a portion of the regular stripe to be sent out before the entire regular stripe is encoded.
[0084] In WPP, the image is segmented into individual rows of codec tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other segments. Parallel processing is possible through parallel decoding of CTB rows, where the start of decoding a CTB row is delayed by two CTBs to ensure that data associated with the CTBs above and to the right of the main CTB is available before the main CTB is decoded. Using this staggered start (which looks like a wavefront when represented graphically), parallelization is possible for at most as many processors / cores as the image has CTB rows. Because intra-image prediction is allowed between adjacent tree block rows within the image, the inter-processor / inter-core communication required to enable intra-image prediction can be important. WPP segmentation does not result in additional NAL units compared to when WPP segmentation is not applied, therefore WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular stripes with some encoding / decoding overhead can be used in conjunction with WPP.
[0085] Tile definitions divide an image into tile columns and rows, forming horizontal and vertical boundaries. Tile columns extend from the top to the bottom of the image. Similarly, tile rows extend from the left to the right of the image. The number of tiles in an image can be simply derived as the number of tile columns multiplied by the number of tile rows.
[0086] Before decoding the top-left CTB of the next tile in the order of the tile raster scan of the image, the scan order of the CTBs is changed to local within the tile (in the order of the tile's CTB raster scan). Similar to regular stripes, tiles disrupt intra-image prediction correlations and entropy decoding correlations. However, they do not need to be included in separate NAL units (the same as WPP in this respect); therefore, tiles cannot be used for MTU size matching. Each tile can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-image prediction between processing units decoding adjacent tiles is limited to transmitting a shared stripe header when the stripe spans more than one tile, and loop filtering involves sharing reconstructed samples and metadata. When more than one tile or WPP segment is included in a stripe, the entry point byte offset for each tile or WPP segment other than the first tile or WPP segment in the stripe is signaled in the stripe header.
[0087] For simplicity, restrictions on the application of the four different image segmentation schemes have been specified in HEVC. A given codec video sequence cannot include tiles and wavefronts used in most of the profiles specified in HEVC. For each strip and tile, one or both of the following conditions must be met: 1) All codec tree blocks in a strip belong to the same tile; 2) All codec tree blocks in a tile belong to the same strip. Finally, a wavefront segment contains exactly one CTB line, and when using WPP, if a strip begins within a CTB line, it must end within the same CTB line.
[0088] The most recent modifications to HEVC are specified in the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)" dated October 24, 2017, available here: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this modification, HEVC specifies three types of MCTS-related SEI messages: i.e., domain-specific MCTS SEI messages, MCTS extracted information set SEI messages, and MCTS extracted information nested SEI messages.
[0089] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vector is restricted to pointing to the full sample location inside the MCTS and only the fractional sample location inside the MCTS is needed for interpolation, and motion vector candidates predicted from temporal motion vectors derived from blocks outside the MCTS are not allowed. In this way, each MCTS can be decoded independently without any tiles not included in the MCTS.
[0090] The MCTS Extraction Information Set (SEI) message provides supplementary information that can be used for MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for the MCTS set. This information consists of multiple extraction information sets, each defining multiple MCTS sets, and includes RBSP bytes for replacing VPS, SPS, and PPS used in the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) will typically need to have different values.
[0091] 3.2. Image Segmentation in VVC
[0092] In VVC, an image is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that cover a rectangular area of the image. The CTUs within a tile are scanned in raster scan order within that tile.
[0093] A stripe consists of an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of an image.
[0094] Two stripe modes are supported: raster scan stripe mode and rectangular stripe mode. In raster scan stripe mode, a stripe contains a sequence of complete tiles in a raster scan of an image. In rectangular stripe mode, a stripe contains multiple complete tiles that together form a rectangular area of the image, or multiple consecutive complete CTU rows of a single tile that together forms a rectangular area of the image. Tiles within a rectangular stripe are scanned in tile raster scan order within the rectangular area corresponding to that stripe.
[0095] A sub-image contains one or more stripes that collectively cover a rectangular area of the image.
[0096] Figure 1 An example of raster scan strip segmentation of an image is shown, where the image is divided into 12 patches and 3 raster scan strips.
[0097] Figure 2 An example of rectangular strip segmentation of an image is shown, where the image is divided into 24 tiles (6 tile columns and 4 tile rows) and 9 rectangular strips.
[0098] Figure 3An example of an image divided into tiles and rectangular strips is shown, where the image is divided into 4 tiles (2 tile columns and 2 tile rows) and 4 rectangular strips.
[0099] Figure 4 An example of sub-image segmentation of an image is shown, where the image is segmented into 18 tiles, each tile on the left covering 12 tiles of a strip of 4 x 4 CTUs, and each tile on the right covering 6 tiles of 2 vertically stacked strips of 2 x 2 CTUs, resulting in a total of 24 strips and 24 sub-images of different sizes (each strip is a sub-image).
[0100] 3.3. Changes in image resolution within a sequence
[0101] In AVC and HEVC, the spatial resolution of an IRAP image cannot be changed unless a new sequence with a new SPS begins. VVC allows for changes in image resolution at positions within a sequence without encoding the IRAP image, which is always intra-coded. This feature is sometimes called Reference Image Resampling (RPR) because it requires resampling of the reference image used for inter-frame prediction when the reference image has a different resolution than the current image being decoded.
[0102] The scaling ratio is limited to greater than or equal to 1 / 2 (2x downsampling from the reference image to the current image) and less than or equal to 8 (8x upsampling). Three sets of resampling filters with different frequency cutoffs are specified to handle various scaling ratios between the reference and current images. The three sets of resampling filters are applied to scaling ratios from 1 / 2 to 1 / 1.75, from 1 / 1.75 to 1 / 1.25, and from 1 / 1.25 to 8, respectively. Each set of resampling filters has 16 phases for luma and 32 phases for chroma, the same as in the case of motion-compensated interpolation filters. In fact, normal MC interpolation is a special case of resampling with scaling ratios ranging from 1 / 1.25 to 8. The horizontal and vertical scaling ratios are derived based on the image width and height, and the scaling offsets specified for the left, right, top, and bottom of the reference and current images.
[0103] Other aspects of the VVC design that support this feature, unlike HEVC, include: i) signaling the image resolution and corresponding conformance window in the PPS instead of the SPS, and signaling the maximum image resolution in the SPS. ii) For a single-layer bitstream, each image memory (the time slot in the DPB used to store one decoded image) occupies the buffer size required to store the decoded image with the maximum image resolution.
[0104] 3.4. Typical Scalable Video Codec (SVC) and SVC in VVC
[0105] Scalable video codec (SVC, sometimes simply referred to as scalability in video codec) refers to video codec that uses a base layer (BL) (sometimes called a reference layer (RL)) and one or more scalable enhancement layers (EL). In SVC, the base layer can carry video data with a base quality level. One or more enhancement layers can carry additional video data to support, for example, higher spatial, temporal, and / or signal-to-noise ratio (SNR) levels. Enhancement layers can be defined relative to previous encoded layers. For example, the bottom layer can be used as a BL, while the top layer can be used as an EL. Intermediate layers can be used as ELs or RLs or both. For example, an intermediate layer (e.g., a layer that is neither the lowest nor the highest layer) can be an EL of a layer below the intermediate layer (e.g., the base layer or any intervening enhancement layer) and simultaneously used as an RL of one or more enhancement layers above the intermediate layer. Similarly, in the multi-view or 3D extension of the HEVC standard, multiple views can exist, and information from one view can be used to codec (e.g., encode or decode) information from another view (e.g., motion estimation, motion vector prediction, and / or other redundancy).
[0106] In SVC, parameters used by the encoder or decoder are grouped into parameter sets based on the codec level they can utilize (e.g., video level, sequence level, picture level, stripe level, etc.). For example, parameters that can be used by one or more codec video sequences from different layers in a bitstream can be included in the Video Parameter Set (VPS), while parameters used by one or more pictures in a codec video sequence can be included in the Sequence Parameter Set (SPS). Similarly, parameters utilized by one or more stripes in a picture can be included in the Picture Parameter Set (PPS), and other parameters specific to a single strip can be included in the stripe header. Likewise, indications of the parameter sets(s) used by a particular layer at a given time can be provided at various codec levels.
[0107] Thanks to the support for Reference Picture Resampling (RPR) in VVC, it's possible to design support for bitstreams containing multiple layers (e.g., two layers with SD and HD resolutions) in VVC without requiring any additional signal processing level codec tools, as the upsampling required for spatial scalability support can be achieved using only RPR upsampling filters. However, scalability support requires high-level syntax changes (compared to a lack of scalability support). Scalability support was specified in VVC version 1. Unlike scalability support in any earlier video codec standards, including extensions to AVC and HEVC, VVC scalability was designed to be as friendly as possible to single-layer decoder designs. Decoding capabilities for multi-layer bitstreams are specified as if there were only a single layer in the bitstream. For example, decoding capabilities, such as DPB size, are specified in a way independent of the number of layers in the bitstream to be decoded. Essentially, decoders designed for single-layer bitstreams don't require many changes to decode multi-layer bitstreams. Compared to the multi-layer extensions of AVC and HEVC, the HLS aspect has been significantly simplified at the expense of some flexibility. For example, an IRAP AU is required to contain the picture for each layer present in CVS.
[0108] 3.5. Viewport-dependent 360 based on sub-images o Video streaming
[0109] In 360 o In video streaming (also known as omnidirectional video), only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is presented to the user at any given moment, while the user can turn his / her head at any time to change their viewing orientation and thus change the current viewport. While it's desirable to have at least some lower-quality representations of areas available at the client that are not covered by the current viewport and are ready to be presented to the user only if the user suddenly changes their viewing orientation to anywhere on the sphere, a high-quality representation of the omnidirectional video is required for the currently used viewport. This optimization is achieved by dividing the high-quality representation of the entire omnidirectional video into sub-pictures with appropriate granularity. Using VVC, these two representations can be encoded as two independent layers.
[0110] Figure 5 The image shows a typical sub-image-based viewport-dependent 360. o The video transmission scheme uses a higher resolution representation of the full video that consists of sub-pictures, while a lower resolution representation of the full video does not use sub-pictures and can be encoded and decoded using a higher resolution representation of less frequent random access points. The client receives the lower resolution full video, while for the higher resolution video, it only receives and decodes the sub-pictures covering the current viewport.
[0111] The latest VVC draft specification also supports, for example Figure 6 The improved 360 shown o Video encoding / decoding schemes. (and) Figure 5 The only difference between the methods shown is for... Figure 6 The method shown applies inter-layer prediction (ILP).
[0112] 3.6. Parameter Set
[0113] AVC, HEVC, and VVC specify parameter sets. Parameter set types include SPS, PPS, APS, and VPS. SPS and PPS are supported in AVC, HEVC, and VVC. VPS was introduced with HEVC and is included in both HEVC and VVC. APS is not included in AVC or HEVC, but is included in the latest VVC draft text.
[0114] SPS is designed to carry sequence-level header information, and PPS is designed to carry infrequently changing image-level header information. Using SPS and PPS, infrequently changing information does not need to be repeated for each sequence or image, thus avoiding redundant signaling. Furthermore, the use of SPS and PPS enables out-of-band transmission of important header information, thus not only avoiding the need for redundant transmission but also improving error resilience.
[0115] A VPS is introduced to carry sequence-level header information common to all layers in a multi-layer bitstream.
[0116] APS was introduced to carry image-level or stripe-level information that requires a considerable number of bits to encode and decode, can be shared by multiple images, and can have a considerable number of different variations in the sequence.
[0117] 3.7. Sub-image Sub-bitstream Extraction and Processing
[0118] Clause C.7 of the latest VVC documentation specifies the sub-image sub-bitstream extraction process as follows:
[0119] C.7 Sub-image Sub-bitstream Extraction Processing
[0120] The input to this process is a bitstream inBitstream, a target OLS index targetOlsIdx, a target highest TemporalId value tIdTarget, and an array of target subpic IdxTarget values for each layer.
[0121] The output of this process is the sub-bit stream outBitstream.
[0122] For a given input bitstream, any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0123] - The output sub-bitstream is the output of the processing specified in this clause, where for the bitstream, targetOlsIdx is equal to the index of the OLS list specified by the VPS, and subpicIdxTarget[] is equal to the index of the subpic that exists in the OLS as input.
[0124] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0125] - The output sub-bitstream contains at least one VCL NAL unit with a TemporalId equal to tIdTarget.
[0126] Note - A consistent bitstream contains one or more encoded / decoded stripe NAL units with TemporalId equal to 0, but does not necessarily contain encoded / decoded stripe NAL units with nuh_layer_id equal to 0.
[0127] - The output sub-bitstream contains at least one VCL NAL unit, for each i in the range of 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), its nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and where sh_subpic_id is equal to the value in SubpicIdVal[subpicIdxTarget[i]].
[0128] The output sub-bitstream, outBitstream, is exported as follows:
[0129] - Invoke the sub-bitstream extraction process specified in Appendix C.6 using inBitstream, targetOlsIdx, and tIdTarget as inputs, and assign the output of the process to outBitstream.
[0130] - If some external components not specified in this specification can be used to provide an alternative parameter set for the sub-bitstream outBitstream, then replace all parameter sets with the alternative parameter set.
[0131] - Otherwise, when the sub-image level information SEI message exists in inBitstream, the following applies:
[0132] - The variable subpicIdx is set to the value of subpicIdxTarget[[NumLayersInOls[targetOlsIdx]-1]].
[0133] - Rewrite the value of general_level_idc in the vps_ols_ptl_idx[targetOlsIdx]th entry of the profile_tier_level() syntax structure list in all referenced VPS NAL units to be equal to SubpicSetLevelIdc derived in Equation D.11 for the set of subpicks whose subpick index is equal to subpicIdx.
[0134] - When VCL HRD or NAL HRD parameters exist, rewrite the corresponding values of cpb_size_value_minus1[tIdTarget][j] and bit_rate_value_minus1[tIdTarget][j] of the j-th CPB in the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] ols_hrd_parameters() syntax structure in all referenced VPS NAL units and in the ols_hrd_parameters() syntax structure in all SPSNAL units referenced by the i-th layer, such that they correspond to the values of SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx] and Subp, respectively, derived from equations D.6 and D.7. icCpbSizeNal[SubpicSetLevelIdx][subpicIdx], derived from equations D.8 and D.9 respectively as SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx] and SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx], where SubpicSetLevelIdx is derived from equation D.11 for subpicks with a subpick index equal to subpicIdx, j is in the range from 0 to hrd_cpb_cnt_minus1 (inclusive), and i is in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive).
[0135] The following applies to the i-th layer where i is in the range of 0 to NumLayersInOls[targetOlsIdx]-1.
[0136] - Rewrite the value of general_level_idc in the profile_tier_level() syntax structure of all referenced SPS NAL units where sps_ptl_dpb_hrd_params_present_flag is equal to 1, to be equal to SubpicSetLevelIdc derived from Equation D.11 for the set of subpicks whose subpick index is equal to subpicIdx.
[0137] - The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are exported as follows:
[0138] subpicWidthInLumaSamples=min((sps_subpic_ctu_top_left_x[subpicIdx]+(C.24)
[0139] sps_subpic_width_minus1[subpicIdx]+1) CtbSizeY,
[0140] pps_pic_width_in_luma_samples)-
[0141] sps_subpic_ctu_top_left_x[subpicIdx] CtbSizeY
[0142] subpicHeightInLumaSamples=min((sps_subpic_ctu_top_left_y[subpicIdx]+(C.25)
[0143] sps_subpic_height_minus1[subpicIdx]+1) CtbSizeY,
[0144] pps_pic_height_in_luma_samples)-
[0145] sps_subpic_ctu_top_left_y[subpicIdx] CtbSizeY
[0146] - Rewrite the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS NAL cells and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS NAL cells to be equal to subpicWidthInLumaSamples and subpicHeightInLumaSamples, respectively.
[0147] - Rewrite the values of sps_num_subpics_minus1 in all referenced SPS NAL cells and pps_num_subpics_minus1 in all referenced PPSNAL cells to 0.
[0148] - If present, rewrite the syntax elements sps_subpic_ctu_top_left_x[subpicIdx] and sps_subpic_ctu_top_left_y[subpicIdx] in all referenced SPS NAL units to 0.
[0149] - Remove the syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j], and sps_subpic_id[j] from all referenced SPS NAL units and for each j that is not equal to subpicIdx.
[0150] - Rewrite all referenced syntax elements in PPS to signal tiles and stripes to remove all tile rows, tile columns, and stripes not associated with a subpicture whose subpicture index is equal to subpicIdx.
[0151] - The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are exported as follows:
[0152] subpicConfWinLeftOffset=sps_subpic_ctu_top_left_x[subpicIdx]==0? (C.26)
[0153] sps_conf_win_left_offset:0
[0154] subpicConfWinRightOffset=(sps_subpic_ctu_top_left_x[subpicIdx]+ (C.27)
[0155] sps_subpic_width_minus1[subpicIdx]+1) CtbSizeY>=sps_pic_width_max_in_luma_samples?sps_conf_win_right_offset:0
[0156] subpicConfWinTopOffset=sps_subpic_ctu_top_left_y[subpicIdx]==0? (C.28)
[0157] sps_conf_win_top_offset:0
[0158] subpicConfWinBottomOffset=(sps_subpic_ctu_top_left_y[subpicIdx]+ (C.29)
[0159] sps_subpic_height_minus1[subpicIdx]+1) CtbSizeY>=sps_pic_height_max_in_luma_samples?sps_conf_win_bottom_offset:0
[0160] - Rewrite the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in all referenced SPS NAL cells, as well as the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset in all referenced PPS NAL cells, to be equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively.
[0161] - Remove from outBitstream all VCL NAL units that have a nuh_layer_id equal to nuh_layer_id of layer i and a sh_subpic_id not equal to SubpicIdVal[subpicIdx].
[0162] - When sli_cbr_constraint_flag equals 1, remove all NAL units with nal_unit_type equal to FD_NUT and padding load SEI messages not associated with VCL NAL units of subpicIdTarget[]. Also, set cbr_flag[tIdTarget][j] of the j-th CPB in the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] ols_hrd_parameters() syntax structure of all referenced VPS NAL units and SPS NAL units to equal 1, where j is in the range from 0 to hrd_cpb_cnt_minus1. Otherwise (sli_cbr_constraint_flag equals 0), clear all NAL units with nal_unit_type equal to FD_NUT and padding load SEI messages, and set cbr_flag[tIdTarget][j] to equal 0.
[0163] - When outBitstream contains an SEI NAL cell with a scalable nested SEI message that is applicable to outBitstream and has sn_ols_flag equal to 1 and sn_subpic_flag equal to 1, extract the appropriate non-scalable nested SEI message with payloadType equal to 1 (PT), 130 (DUI), or 132 (decoded image hash) from the scalable nested SEI message, and put the extracted SEI message into outBitstream.
[0164] 4. Technical problems solved by the technical solution
[0165] The existing design for sub-image sub-bitstream extraction processing in the latest VVC text (in JVET-R2001-vA / v10) has the following problems:
[0166] 1) When extracting rectangular regions (where these regions cover one or more sub-images) from an image sequence, the scaling window offset parameter in PPS needs to be rewritten to maintain the same scaling window as used during encoding of the original bitstream, because the image width and / or height of the extracted images have changed. This is as follows: Figure 6 The improved 360 shown o This is particularly necessary in video encoding / decoding scenarios. However, the latest VVC text-based sub-image sub-bitstream extraction processing does not require rewriting the scaling window offset parameters.
[0167] 2) In the sub-image sub-bitstream extraction and processing of the latest VVC text, any output sub-bitstream that meets all of the following conditions must be a conforming bitstream:
[0168] - The output sub-bitstream is the output of the processing specified in this clause, where for the bitstream, targetOlsIdx is equal to the index of the OLS list specified by the VPS, and subpicIdxTarget[] is equal to the index of the subpic that exists in the OLS as input.
[0169] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0170] - The output sub-bitstream contains at least one VCL NAL unit whose TemporalId is equal to tIdTarget.
[0171] Note - A consistent bitstream contains one or more codec stripe NAL units with TemporalId equal to 0, but does not necessarily contain codec stripe NAL units with nuh_layer_id equal to 0.
[0172] - The output sub-bitstream contains at least one VCL NAL unit, for each i in the range of 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), its nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and its sh_subpic_id is equal to the value in SubpicIdVal[subpicIdxTarget[i]].
[0173] However, the above constraints have the following problems:
[0174] a. In the first item, the input tIdTarget is missing.
[0175] b. Another issue related to the first point is as follows: Only certain combinations, not all combinations, of sub-images from different layers can form a consistent bitstream.
[0176] 3) The removal of VCL NAL cells, their associated padding data NAL cells, and associated padding load SEI messages is performed only when there are no external components used to replace the parameter set. However, this removal is also necessary when there are external components used to replace the parameter set.
[0177] 4) The current removal of the filler load SEI message may involve rewriting the SEI NAL cell.
[0178] 5) The Subpicture Level Information (SLI) SEI message is specified as layer-specific. However, the SLI SEI message is used in subpicture sub-bitstream extraction processing as if the information were applied to subpictures across all layers.
[0179] 6) Scalable nested SEI messages can be used to nest SEI messages for certain extracted subpicture sequences for certain OLS. However, due to the use of the wording "in the CLVS" or "in a CLVS", the semantics of sn_num_subpics_minus1 and sn_subpic_id_len_minus1 are specified as layer-specific.
[0180] 7) Some aspects need to be changed to support bitstream extraction, such as... Figure 6In the lower part, in the extracted bitstream sent to the decoder, images that originally contained multiple sub-images now contain fewer sub-images, while images containing only one sub-image remain unchanged.
[0181] 5. Examples of solutions and implementation methods
[0182] To address the aforementioned and other issues, methods outlined below are disclosed. These items should be considered as examples for interpreting general concepts, and not interpreted in a narrow sense. Furthermore, these items can be applied individually or in any combination. In the following description, additions or modifications from the relevant specifications are highlighted in bold and italics, and some deleted portions are marked with double brackets (e.g., [[a]] indicates the deletion of the character "a"). There may be other substantially editorial changes that are therefore not highlighted.
[0183] 1) To address problem 1, the calculation and rewriting of scaling window offset parameters (such as in PPS) are specified as part of the sub-image sub-bitstream extraction process.
[0184] a. In one example, the scaling window offset parameter is calculated and overridden as follows:
[0185] - The variables subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset are exported as follows:
[0186] subpicScalWinLeftOffset=pps_scaling_win_left_offset- (C.30)
[0187] sps_subpic_ctu_top_left_x[spIdx] CtbSizeY / SubWidthC
[0188] rightSubpicBd=(sps_subpic_ctu_top_left_x[spIdx]+
[0189] sps_subpic_width_minus1[spIdx]+1) CtbSizeY
[0190] subpicScalWinRightOffset=
[0191] (rightSubpicBd>=sps_pic_width_max_in_luma_samples)? (C.31)
[0192] pps_scaling_win_right_offset:pps_scaling_win_right_offset-
[0193] (sps_pic_width_max_in_luma_samples-rightSubpicBd) /
[0194] SubWidthC
[0195] subpicScalWinTopOffset=pps_scaling_win_top_offset- (C.32)
[0196] sps_subpic_ctu_top_left_y[spIdx] CtbSizeY / SubHeightC
[0197] botSubpicBd=(sps_subpic_ctu_top_left_y[spIdx]+
[0198] sps_subpic_height_minus1[spIdx]+1) CtbSizeY
[0199] subpicScalWinBotOffset=
[0200] (botSubpicBd>=sps_pic_height_max_in_luma_samples)? (C.33)
[0201] pps_scaling_win_bottom_offset:pps_scaling_win_bottom_offset-
[0202] (sps_pic_height_max_in_luma_samples-botSubpicBd) /
[0203] SubHeightC
[0204] In the above equations, sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples, and sps_pic_height_in_luma_samples came from the original SPS before they were rewritten, and pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset came from the original PPS before they were rewritten.
[0205] - Rewrite the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in all referenced PPS NAL cells to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset, respectively.
[0206] b. In one example, sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in the above equation are replaced by pps_pic_width_in_luma_samples and pps_pic_width_in_luma_samples, respectively, and pps_pic_width_in_luma_samples are the values in the original PPS before they were overwritten.
[0207] c. Additionally, in one example, the reference sample enable flag in the SPS (e.g., sps_ref_pic_resampling_enabled_flag) is overridden when a change occurs, and the resolution change allow flag in the CLVS in the SPS (e.g., sps_res_change_in_clvs_allowed_flag) is overridden when a change occurs.
[0208] 2) To solve problems 2a and 2b, the following methods are proposed.
[0209] a. To resolve issue 2a, add the missing input tIdTarget, for example, by doing the following:
[0210] - The output sub-bitstream is the output of the processing specified in this clause, where for the bitstream, targetOlsIdx is equal to the index of the OLS list specified by the VPS, and subpicIdxTarget[] is equal to the index of the subpic that exists in the OLS as input.
[0211] Change to the following:
[0212]
[0213] b. To address problem 2b, clearly specify which combination of sub-images from different layers needs to be a consistent bitstream during extraction, for example, as follows:
[0214] The input bitstream consistency requirement is that any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0215]
[0216] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in the list LayerIdInOls[targetOlsIdx].
[0217] - The output sub-bitstream contains at least one VCL NAL unit with a TemporalId equal to tIdTarget.
[0218] Note - A consistent bitstream contains one or more encoded / decoded stripe NAL units with TemporalId equal to 0, but does not necessarily contain encoded / decoded stripe NAL units with nuh_layer_id equal to 0.
[0219] - The output sub-bitstream contains at least one VCL NAL unit, for each i in the range of 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), its nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and its sh_subpic_id is equal to SubpicIdVal[subpicIdxTarget[i]].
[0220] 3) To solve problem 3, regardless of whether there are external components for replacing the parameter set, the removal of VCL NAL units, their associated padding data NAL units, and associated padding load SEI messages is performed.
[0221] 4) To address problem 4, a constraint is added requiring that the SEI NAL cell containing the filler load SEI message does not contain other types of SEI messages, and the removal of the filler load SEI message is specified as the removal of the SEI NAL cell containing the filler load SEI message.
[0222] 5) To address issue 5, specify that the values of one or more of SubpicSetLevelIdc, SubpicCpbSizeVcl[ SubpicSetLevelIdx ][ subpicIdx ], SubpicCpbSizeNal[ SubpicSetLevelIdx ][ subpicIdx ], SubpicBitrateVcl[SubpicSetLevelIdx ][ subpicIdx ], SubpicBitrateNal[ SubpicSetLevelIdx ][ subpicIdx ], and sli_cbr_constraint_flag, derived from or found in the SLI SEI message, are identical for all layers in the target OLS of the extraction process.
[0223] a. In one example, the constraint is that, in order to be used with multi-level OLS, the SLI SEI message should be included in the scalable nested SEI message and should be indicated in the scalable nested SEI message to be applied to a specific OLS that includes at least that multi-level OLS (i.e., when sn_ols_flag equals 1).
[0224] i. In addition, the value 203 (payloadType value of SLI SEI message) is removed from the list VclAssociatedSeiList to enable SLI messages to be included in scalable nested SEI messages, where sn_ols_flag equals 1.
[0225] b. In one example, the constraint is that, in order to be used with a multi-tiered OLS, the SLI SEI message should be included in the scalable nested SEI message and should be indicated in the scalable nested SEI message to be applied to a specific list of layers consisting of all layers of that multi-tiered OLS (i.e., sn_ols_flag equals 1).
[0226] c. In one example, the SLI SEI message is specified as OLS-specific, for example, by adding an OLS index to the SLI SEI message syntax to indicate the OLS to which the information carried in the SEI message is applied.
[0227] i. Optionally, an OLS index list is added to the SLI SEI message syntax to indicate the OLS list to which the information carried in the SEI message is applied.
[0228] d. In one example, it is required that the values of one or more of the following in all SLI SEI messages used for all layers of OLS be the same: SubpicSetLevelIdc, SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx], SubpicCpbSizeNal[SubpicSetLevelIdx][subpicIdx], SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx], SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx], and sli_cbr_constraint_flag.
[0229] 6) To solve problem 6, the semantics are as follows:
[0230] Add 1 to specify the number of subpicks to which the scalable nested SEI message applies. The value of sn_num_subpics_minus1 should be less than or equal to the value of sps_num_subpics_minus1 in the SPS referenced by the picture in CLVS.
[0231] Adding 1 specifies the number of bits used to represent the syntax element sn_subpic_id[i]. The value of sn_subpic_id_len_minus1 should be in the range of 0 to 15 (inclusive).
[0232] The requirement for bitstream consistency is that the value of sn_subpic_id_len_minus1 should be the same for all scalable nested SEI messages present in CLVS.
[0233] Change to the following:
[0234]
[0235] Adding 1 specifies the number of bits used to represent the syntax element sn_subpic_id[i]. The value of sn_subpic_id_len_minus1 should be in the range of 0 to 15 (inclusive).
[0236] The requirement for bitstream consistency is that, for [[what exists in CLVS]], For all scalable nested SEI messages, the value of sn_subpic_id_len_minus1 should be the same.
[0237] 7) To solve problem 7, the following method is proposed:
[0238] a. In one example, it is clearly specified which combination of sub-images from different layers is required to be a consistent bitstream when extracted, for example, where there are at least two layers, one layer's image includes multiple sub-images and the other layer's image includes only one sub-image:
[0239] For a given input bitstream, any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0240]
[0241] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to each of the nuh_layer_id values in LayerIdInOls[targetOlsIdx].
[0242] - The output sub-bitstream contains at least one VCL NAL unit with a TemporalId equal to tIdTarget.
[0243] Note 2 - A consistent bitstream contains one or more encoded / decoded stripe NAL units with TemporalId equal to 0, but does not necessarily contain encoded / decoded stripe NAL units with nuh_layer_id equal to 0.
[0244] - The output sub-bitstream contains at least one VCL NAL unit, for each i in the range of 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), its nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and its sh_subpic_id is equal to the value in SubpicIdVal[subpicIdxTarget[i]].
[0245] b. The semantics of the scalable nested SEI message are specified such that when both sn_ols_flag and sn_subpic_flag are equal to 1, the list of sn_subpic_id[] values specifies the subpick IDs of subpicks in the applicable OLS, each of which contains multiple subpicks. This allows the OLS to which the scalable nested SEI message is applied to to have layers where each image contains only one subpick, for which the subpick ID is not indicated in the scalable nested SEI message.
[0246] 8) If the first image is divided into a sub-image in a sub-image layout that is different from the sub-image layout of the current image, then the first reference image must not be set as a co-position image of the current image.
[0247] a. In one example, the first reference image and the current image are on different layers.
[0248] 9) If codec tool X depends on a first reference image, and the first reference image is divided into sub-images in a sub-image layout that is different from the sub-image layout of the current image, then codec tool X should be disabled for the current image.
[0249] a. In one example, the first reference image and the current image are on different layers.
[0250] b. In one example, the codec tool X is a bidirectional optical flow (BDOF).
[0251] c. In one example, the codec tool X is decoder-side motion vector refinement (DMVR).
[0252] d. In one example, the codec tool X is a prediction refinement (PROF) with optical flow.
[0253] 6. Examples
[0254] The following are some example embodiments of aspects of the invention outlined in Section 5 above, which can be applied to the VVC specification. The modified text is based on the latest VVC text in JVET-R2001-vA / v10. Most of the relevant parts that have been added or modified are highlighted in bold italics, and some deleted parts are marked with double brackets (e.g., [[a]] indicates the deletion of the character "a"). There may be some other substantially editorial changes that are not highlighted.
[0255] 6.1. First Embodiment
[0256] This embodiment applies to items 1, 1.a, 2a, 2b, 3, 4, 5, 5a, 5.ai and 6.
[0257] C.7 Sub-image Sub-bitstream Extraction Processing
[0258] The input to this process is a bitstream inBitstream, a target OLS index targetOlsIdx, a target highest TemporalId value tIdTarget, and an array of target subpic IdxTarget[i] for each layer of i in the range from 0 to NumLayersInOls[targetOLsIdx]-1 (inclusive).
[0259] The output of this process is the sub-bit stream outBitstream.
[0260] For a given input bitstream, any output sub-bitstream that satisfies all of the following conditions should be a consistent bitstream:
[0261]
[0262] For use with multi-tiered OLS, the SLI SEI message should be included within a scalable nested SEI message and should be indicated in the scalable nested SEI message to be applied to a specific OLS or to all layers within a specific OLS.
[0263] - The output sub-bitstream contains at least one VCL NAL unit whose nuh_layer_id is equal to Each value in the nuh_layer_id value in LayerIdInOls[targetOlsIdx].
[0264] - The output sub-bitstream contains at least one VCL NAL unit with a TemporalId equal to tIdTarget.
[0265] Note - A consistent bitstream contains one or more encoded / decoded stripe NAL units with TemporalId equal to 0, but does not necessarily contain encoded / decoded stripe NAL units with nuh_layer_id equal to 0.
[0266] - The output sub-bitstream contains at least one VCL NAL unit, for each i in the range of 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive), its nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and its sh_subpic_id is equal to the value in SubpicIdVal[subpicIdxTarget[i]][[].
[0267] The output sub-bitstream, outBitstream, is exported as follows:
[0268]
[0269] - If some external components not specified in this specification can be used to provide an alternative parameter set for the sub-bitstream outBitstream, then replace all parameter sets with the alternative parameter set.
[0270] - Otherwise, when the sub-image level information SEI message exists in inBitstream, the following applies:
[0271] - [[The variable subpicIdx is set to the value of subpicIdxTarget[[NumLayersInOls[targetOlsIdx]-1]].]]
[0272] - Rewrite the value of general_level_idc in the vps_ols_ptl_idx[targetOlsIdx]th entry of the profile_tier_level() syntax structure list in all referenced VPS NAL units to be equal to SubpicSetLevelIdc derived in Equation D.11 for the set of subpicks whose subpick index is equal to subpicIdx.
[0273] - When VCL HRD or NAL HRD parameters exist, rewrite the corresponding values of cpb_size_value_minus1[tIdTarget][j] and bit_rate_value_minus1[tIdTarget][j] of the j-th CPB in the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] ols_hrd_parameters() syntax structure in all referenced VPS NAL units and in the ols_hrd_parameters() syntax structure in all SPS NAL units referenced by the i-th layer, such that they correspond to the values of SubpicCpbSizeVcl[SubpicSetLevelIdx][subpicIdx] and SubpicCpbSizeNal[SubpicSetLevelIdx][subpicIdx] derived from equations D.6 and D.7, respectively, respectively. The SubpicBitrateVcl[SubpicSetLevelIdx][subpicIdx] and SubpicBitrateNal[SubpicSetLevelIdx][subpicIdx] derived from D.8 and D.9, where SubpicSetLevelIdx is derived from equation D.11 for subpicks with a subpick index equal to subpicIdx, j is in the range from 0 to hrd_cpb_cnt_minus1 (inclusive), and i is in the range from 0 to NumLayersInOls[targetOlsIdx]-1 (inclusive).
[0274] - For each value of [[i-th layer]] in the range of 0 to NumLayersInOls[targetOlsIdx]-1, the following applies.
[0275] - .
[0276] - Rewrite the value of general_level_idc in the profile_tier_level() syntax structure of all referenced SPS NAL units where sps_ptl_dpb_hrd_params_present_flag is equal to 1, to be equal to SubpicSetLevelIdc derived from Equation D.11 for the set of subpicks whose subpick index is equal to subpicIdx.
[0277] - The variables subpicWidthInLumaSamples and subpicHeightInLumaSamples are exported as follows:
[0278] subpicWidthInLumaSamples=min((sps_subpic_ctu_top_left_x[spIdx]+(C.24)
[0279] sps_subpic_width_minus1[spIdx]+1) CtbSizeY,pps_pic_width_in_luma_samples)-
[0280] sps_subpic_ctu_top_left_x[spIdx] CtbSizeY
[0281] subpicHeightInLumaSamples=min((sps_subpic_ctu_top_left_y[spIdx]+(C.25)
[0282] sps_subpic_height_minus1[spIdx]+1) CtbSizeY,pps_pic_height_in_luma_samples)-
[0283] sps_subpic_ctu_top_left_y[spIdx] CtbSizeY
[0284] - Rewrite the values of sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples in all referenced SPS NAL cells and the values of pps_pic_width_in_luma_samples and pps_pic_height_in_luma_samples in all referenced PPS NAL cells to be equal to subpicWidthInLumaSamples and subpicHeightInLumaSamples, respectively.
[0285] - Rewrite the values of sps_num_subpics_minus1 in all referenced SPS NAL cells and pps_num_subpics_minus1 in all referenced PPSNAL cells to 0.
[0286] - If present, rewrite the syntax elements sps_subpic_ctu_top_left_x[spIdx] and sps_subpic_ctu_top_left_y[spIdx] in all referenced SPS NAL units to 0.
[0287] - Remove the syntax elements sps_subpic_ctu_top_left_x[j], sps_subpic_ctu_top_left_y[j], sps_subpic_width_minus1[j], sps_subpic_height_minus1[j], sps_subpic_treated_as_pic_flag[j], sps_loop_filter_across_subpic_enabled_flag[j], and sps_subpic_id[j] from all referenced SPS NAL units and for each j that is not equal to subpicIdx.
[0288] - Rewrite all referenced syntax elements in PPS to signal tiles and stripes to remove all tile rows, tile columns, and stripes not associated with a subpicture whose subpicture index is equal to subpicIdx.
[0289] - The variables subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset are exported as follows:
[0290] subpicConfWinLeftOffset=sps_subpic_ctu_top_left_x[spIdx]==0? (C.26)
[0291] sps_conf_win_left_offset:0
[0292] subpicConfWinRightOffset=(sps_subpic_ctu_top_left_x[spIdx]+ (C.27)
[0293] sps_subpic_width_minus1[spIdx]+1) CtbSizeY>=
[0294] sps_pic_width_max_in_luma_samples?sps_conf_win_right_offset:0
[0295] subpicConfWinTopOffset=sps_subpic_ctu_top_left_y[spIdx]==0? (C.28)
[0296] sps_conf_win_top_offset:0
[0297] subpicConfWinBottomOffset=(sps_subpic_ctu_top_left_y[spIdx]+ (C.29)
[0298] sps_subpic_height_minus1[spIdx]+1) CtbSizeY>=
[0299]
[0300] - Rewrite the values of sps_conf_win_left_offset, sps_conf_win_right_offset, sps_conf_win_top_offset, and sps_conf_win_bottom_offset in all referenced SPS NAL cells, as well as the values of pps_conf_win_left_offset, pps_conf_win_right_offset, pps_conf_win_top_offset, and pps_conf_win_bottom_offset in all referenced PPS NAL cells, to be equal to subpicConfWinLeftOffset, subpicConfWinRightOffset, subpicConfWinTopOffset, and subpicConfWinBottomOffset, respectively.
[0301]
[0302] - Remove all VCL NAL units from outBitstream that have a nuh_layer_id equal to the nuh_layer_id of layer i and a sh_subpic_id not equal to SubpicIdVal[subpicIdx].
[0303] - [[When]] if sli_cbr_constraint_flag equals 1[[]], [[remove all NAL units with nal_unit_type equal to FD_NUT and padding load SEI messages not associated with VCL NAL units of subpicIdTarget[], and]] set cbr_flag[tIdTarget][j] of the j-th CPB in the vps_ols_hrd_idx[MultiLayerOlsIdx[targetOlsIdx]] ols_hrd_parameters() syntax structure of all referenced VPS NAL units and SPS NAL units to equal 1, and j is in the range from 0 to hrd_cpb_cnt_minus1. Otherwise, (sli_cbr_constraint_flag equals 0), [[clear all NAL units and padding load SEI messages with nal_unit_type equal to FD_NUT, and]] set cbr_flag[tIdTarget][j] to equal 0.
[0304] - When outBitstream contains an SEI NAL cell with a scalable nested SEI message that is applicable to outBitstream and has sn_ols_flag equal to 1 and sn_subpic_flag equal to 1, extract the appropriate non-scalable nested SEI message with payloadType equal to 1 (PT), 130 (DUI), or 132 (decoded image hash) from the scalable nested SEI message, and put the extracted SEI message into outBitstream.
[0305] D2.2 General SEI Load Semantics ...
[0307] The list VclAssociatedSeiList is set to consist of payloadType values 3, 19, 45, 129, 132, 137, 144, 145, 147 to 150 (inclusive), 153 to 156 (inclusive), 168, [[203, ]], and 204.
[0308] The list PicUnitRepConSeiList is set to include payloadType values 0, 1, 19, 45, 129, 132, 133, 137, 147 to 150 (inclusive), 153 to 156 (inclusive), 168, 203, and 204.
[0309] Note 4 - VclAssociatedSeiList consists of the payloadType value of the SEI message, which, when non-scalable nested and contained within an SEI NAL unit, infers constraints on the NAL unit header of the associated VCL NAL unit based on the NAL unit header of the associated VCL NAL unit. PicUnitRepConSeiList consists of the payloadType value of the SEI message, which is limited to 4 repetitions per PU.
[0310] The requirement for bitstream consistency applies to the inclusion of SEI messages in SEI NAL units as follows:
[0311] - When the SEI NAL cell contains a non-scalable nested BP SEI message, a non-scalable nested PT SEI message, or a non-scalable nested DUI SEI message, the SEI NAL cell should not contain any other SEI message with a payloadType that is not equal to 0 (BP), 1 (PT), or 130 (DUI).
[0312] - When the SEI NAL cell contains a scalable nested BP SEI message, a scalable nested PT SEI message, or a scalable nested DUI SEI message, the SEI NAL cell should not contain any other SEI message with a payloadType that is not equal to 0 (BP), 1 (PT), 130 (DUI), or 133 (scalable nested).
[0313] ...
[0315] D.6.2 Scalable Nested SEI Message Semantics ...
[0317] Incrementing `sn_num_subpics_minus1` by 1 specifies the number of subpicks in each picture or layer (when `sn_ols_flag` equals 0) within the OLS (when `sn_ols_flag` equals 1) to which the scalable nested SEI message is applied. The value of `sps_num_subpics_minus1` should be less than or equal to the value of `sps_num_subpics_minus1` in each SPS referenced by the picture (when `sn_ols_flag` equals 1) or layer (when `sn_ols_flag` equals 0) in the OLS of [[CLVS]].
[0318] Increasing 1 in sn_subpic_id_len_minus1 specifies the number of bits used to represent the syntax element sn_subpic_id[i]. The value of sn_subpic_id_len_minus1 should be in the range of 0 to 15 (inclusive).
[0319] The requirement for bitstream consistency is that the value of sn_subpic_id_len_minus1 should be the same for all scalable nested SEI messages applied to images within CLVS [[existing in CLVS]]. ...
[0321] Figure 7 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or it may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Networking (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0322] System 1900 may include codec component 1904, which can implement the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904, used to generate a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via communication through a connection as represented by component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values or to be sent to displayable video at display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the decoder will perform a corresponding decoding tool or operation that reverses the result of the codec.
[0323] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0324] Figure 8 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors (multiple) 3602 can be configured to implement one or more methods described in this document. The memories (multiple memories) 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry.
[0325] Figure 10 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.
[0326] like Figure 10 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110, which may be referred to as a video encoding device, generates encoded video data. The destination device 120, which may be referred to as a video decoding device, can decode the encoded video data generated by the source device 110.
[0327] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0328] Video source 112 may include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.
[0329] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.
[0330] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or may be external to destination device 120 configured to interface with an external display device.
[0331] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or other standards.
[0332] Figure 11 This is a block diagram illustrating an example of a video encoder 200. The video encoder 200 can be... Figure 10 The video encoder 114 in the system 100 shown in the figure.
[0333] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 11 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0334] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.
[0335] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.
[0336] Furthermore, some components (such as motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but for interpretive purposes... Figure 11 The example is shown separately.
[0337] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0338] The mode selection unit 203 can (e.g., based on error results) select either an intra-frame or inter-frame coding / decoding mode, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 also selects the resolution (e.g., sub-pixel or integer pixel precision) for the motion vectors of the block.
[0339] To perform inter-frame prediction for the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 that are different from the images associated with the current video block.
[0340] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0341] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference images in list 0 or list 1 containing the reference video blocks, and a motion vector indicating the spatial shift between the current video block and the reference video blocks. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video blocks indicated by the motion information of the current video block.
[0342] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for reference images in list 0 for reference video blocks of the current video block, and can also search for reference images in list 1 for another reference video block of the current video block. Motion estimation unit 204 can then generate reference indices indicating the reference images of the reference video blocks contained in lists 0 and 1, as well as motion vectors indicating the spatial shift between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block of the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0343] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.
[0344] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of adjacent video blocks.
[0345] In one example, the motion estimation unit 204 may instruct the video decoder 300 in the syntax structure associated with the current video block to indicate that the current video block has the same motion information as another video block.
[0346] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntactic structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0347] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Combined Mode Signaling.
[0348] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on the decoded samples of other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.
[0349] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0350] In other examples, such as in skip mode, the current video block for the current video block may not have residual data, and the residual generation unit 207 may not perform a subtraction operation.
[0351] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0352] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0353] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples from one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0354] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0355] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0356] Figure 12 This is a block diagram illustrating an example of a video decoder 300. The video decoder 300 can be... Figure 10 The video decoder 114 in the system 100 illustrated in the figure.
[0357] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 12 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0358] exist Figure 12 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform largely the same functions as video encoder 200 (…). Figure 11 The encoding round is the opposite of the decoding round.
[0359] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-encoded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. For example, motion compensation unit 302 can determine this information by performing AMVP and merging modes.
[0360] The motion compensation unit 302 can generate motion-compensated blocks, possibly based on interpolation filters. Identifiers of the interpolation filters to be used at sub-pixel precision can be included in the syntax elements.
[0361] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of video blocks, to calculate the interpolated values of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate predictive blocks.
[0362] The motion compensation unit 302 may use some of the syntactic information in the syntactic information to determine: the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is segmented, the mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame encoded block, and other information for decoding the encoded video sequence.
[0363] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization (i.e., dequantization) on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.
[0364] The reconstruction unit 306 can sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.
[0365] The list of solutions describes some embodiments of the disclosed technology.
[0366] The first set of solutions is presented below. The following solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., Item 1).
[0367] 1. A video processing method (e.g., Figure 9 The method 900 described herein includes: (902) converting between a video comprising one or more video images and a video codec representation, the one or more video images comprising one or more sub-images, wherein the codec representation conforms to a format rule; wherein the format rule specifies one or more parameters of a scaling window applicable to the sub-images that are determined based on one or more syntax fields included in the codec representation.
[0368] 2. The method of Solution 1, wherein the one or more parameters include the left offset, right offset, top offset, or bottom offset of the scaling window applicable to the sub-image.
[0369] The following solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., Item 3).
[0370] 3. A video processing method, comprising: converting between a video and a video codec representation comprising one or more video images in a video codec layer, wherein the conversion conforms to a rule specifying the removal of network abstraction layer units, filter data, and padding enhancement information data regardless of whether the parameter set can be externally replaced.
[0371] The following solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., Item 4).
[0372] 4. A video processing method comprising: converting between a video and a video codec representation comprising one or more video images in a video codec layer, wherein the conversion conforms to a format rule specifying the exclusion of additional supplementary enhancement information (SEI) from a network abstraction layer comprising a padding SEI message payload.
[0373] 5. The method of Solution 1, wherein the transformation conforms to a rule that specifies the removal of the padding load SEI message as an SEI NAL unit.
[0374] The following solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., Item 7).
[0375] 6. A video processing method, comprising: converting between a video comprising one or more layers and a video codec representation, the one or more layers comprising one or more video images, the one or more video images comprising one or more sub-images, wherein the conversion is required to conform to a format rule based on which combination of sub-images of different layers is specified, wherein there are at least two layers such that images of one layer are each composed of multiple sub-images while images of another layer are each composed of only one sub-image.
[0376] The following solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., Item 8).
[0377] 7. A video processing method, comprising: converting between a video comprising one or more layers and a video codec representation, the one or more layers comprising one or more video images, the one or more video images comprising one or more sub-images, wherein the conversion conforms to a rule specifying that: for a first image that is divided into sub-images using a sub-image layout, if a second image is divided according to a sub-image layout different from the first image, the second image cannot be used as a co-located reference image.
[0378] 8. The method of Solution 7, wherein the first image and the second image are in different layers.
[0379] The following solutions illustrate example embodiments of the techniques discussed in previous sections (e.g., Item 9).
[0380] 9. A video processing method comprising: converting between a video comprising one or more layers and a video codec representation, the one or more layers comprising one or more video pictures, the one or more video pictures comprising one or more subpictures, wherein the conversion conforms to a rule specifying that: for a first picture divided into subpictures using a subpicture layout, if the codec tool depends on a subpicture layout different from that of a second picture used as a reference picture for the first picture, the use of the corresponding codec tool is prohibited during encoding or during decoding.
[0381] 10. As in Solution 9, wherein the first image and the second image are in different layers.
[0382] 11. The method of any one of solutions 1 to 10, wherein the conversion includes encoding the video into the codec representation.
[0383] 12. The method of any one of solutions 1 to 10, wherein the conversion includes decoding the codec representation to generate the pixel values of the video.
[0384] 13. A video decoding apparatus, comprising a processor configured to implement the methods described in one or more of solutions 1 to 12.
[0385] 14. A video encoding apparatus comprising a processor configured to implement the methods described in one or more of solutions 1 to 12.
[0386] 15. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the methods described in any one of solutions 1 to 12.
[0387] 16. A method, apparatus or system described in this document.
[0388] The second set of solutions illustrates example embodiments of the techniques discussed in the preceding sections (e.g., Project 1).
[0389] 1. A video processing method (e.g., such as...) Figure 13 The method shown (1300) includes: performing a 1302 conversion between a video comprising one or more video images and a bitstream of the video, the one or more video images comprising one or more sub-images, wherein the conversion conforms to a rule specifying one or more parameters for determining a scaling window applicable to the sub-image from one or more syntax elements during sub-image sub-bitstream extraction processing.
[0390] 2. The method according to Solution 1, wherein the one or more syntax elements are included in the original image parameter set and / or the original sequence parameter set.
[0391] 3. The method according to solution 1 or 2, wherein the original image parameter set and / or the original sequence parameter set changes as the one or more parameters are calculated and rewritten.
[0392] 4. The method according to any one of solutions 1 to 3, wherein the one or more parameters include at least one of the left offset, right offset, top offset, or bottom offset of the scaling window applicable to the sub-image.
[0393] 5. The method according to any one of solutions 1 to 4, wherein the left offset (subpicScalWinLeftOffset) of the scaling window is derived as follows:
[0394] subpicScalWinLeftOffset = pps_scaling_win_left_offset - sps_subpic_ctu_top_left_x[spIdx] CtbSizeY / SubWidthC, and
[0395] Where pps_scaling_win_left_offset indicates the left offset applied to the image size for scaling ratio calculation, sps_subpic_ctu_top_left_x[spIdx] indicates the x-coordinate of the codec tree unit located at the top left corner of the subpic, CtbSizeY is the width or height of the luminance codec tree block or codec tree unit, and SubWidthC indicates the width of the video block and is obtained from the table according to the chroma format of the picture including the video block.
[0396] 6. The method according to any one of solutions 1 to 4, wherein the right offset subpicScalWinRightOffset of the scaling window is derived as follows:
[0397] subpicScalWinRightOffset = ( rightSubpicBd >= sps_pic_width_max_in_luma_samples ) ?pps_scaling_win_right_offset : pps_scaling_win_right_offset -
[0398] ( sps_pic_width_max_in_luma_samples - rightSubpicBd ) / SubWidthC, and
[0399] Wherein, `sps_pic_width_max_in_luma_samples` specifies the maximum width of each decoded image in the reference sequence parameter set in units of luminance samples; `pps_scaling_win_right_offset` indicates the right offset applied to the image size for scaling ratio calculation; `SubWidthC` indicates the width of the video block and is obtained from a table based on the chroma format of the image including the video block; and
[0400] rightSubpicBd = ( sps_subpic_ctu_top_left_x[ spIdx ] + sps_subpic_width_minus1[ spIdx ] + 1 ) CtbSizeY, and
[0401] Where sps_subpic_ctu_top_left_x[ spIdx ] indicates the x-coordinate of the codec tree unit located at the top left corner of the subpic, sps_subpic_width_minus1[ spIdx ] indicates the width of the subpic, and CtbSizeY is the width or height of the luma codec tree block or codec tree unit.
[0402] 7. The method according to any one of solutions 1 to 4, wherein the top offset subpicScalWinTopOffset of the scaled window is derived as follows:
[0403] subpicScalWinTopOffset = pps_scaling_win_top_offset -
[0404] sps_subpic_ctu_top_left_y[spIdx] CtbSizeY / SubHeightC,
[0405] Where pps_scaling_win_top_offset indicates the top offset applied to the image size for scaling ratio calculation, sps_subpic_ctu_top_left_y [spIdx] indicates the y-coordinate of the codec tree unit located at the top left corner of the subpic, CtbSizeY is the width or height of the luminance codec tree block or codec tree unit, and SubHightC indicates the height of the video block and is obtained from the table according to the chroma format of the picture including the video block.
[0406] 8. The method according to any one of solutions 1 to 4, wherein the bottom offset ubpicScalWinBotOffset of the scaling window is derived as follows:
[0407] subpicScalWinBotOffset = ( botSubpicBd >= sps_pic_height_max_in_luma_samples ) ? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - ( sps_pic_height_max_in_luma_samples - botSubpicBd ) / SubHeightC, and
[0408] Where `sps_pic_height_max_in_luma_samples` indicates the maximum height of each decoded image in the reference sequence parameter set in units of luminance samples, `pps_scaling_win_bottom_offset` indicates the bottom offset applied to the image size for scaling ratio calculation, `SubHeightC` indicates the height of the video block and is obtained from a table based on the chroma format of the image including the video block, and
[0409] where botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] +
[0410] sps_subpic_height_minus1[ spIdx ] + 1 ) CtbSizeY, and
[0411] Where sps_subpic_ctu_top_left_y [spIdx] indicates the y-coordinate of the codec tree unit located at the top left corner of the subpic, and CtbSizeY is the width or height of the luminance codec tree block or codec tree unit.
[0412] 9. The method according to any one of solutions 5 to 8, wherein sps_subpic_ctu_top_left_x[spIdx], sps_subpic_width_minus1[spIdx], sps_subpic_ctu_top_left_y[spIdx], sps_subpic_height_minus1[spIdx], sps_pic_width_max_in_luma_samples and sps_pic_height_max_in_luma_samples are from the original sequence parameter set, and pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset and pps_scaling_win_bottom_offset are from the original image parameter set.
[0413] 10. The method according to any one of solutions 1 to 3, wherein the rule specifies that the values of pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset in all referenced Picture Parameter Set (PPS) Network Abstraction Layer (NAL) units are rewritten to be equal to subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset, respectively, and
[0414] Where pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset indicate the left, right, top, and bottom offsets applied to the image size for scaling ratio calculation, respectively.
[0415] Wherein, subpicScalWinLeftOffset, subpicScalWinRightOffset, subpicScalWinTopOffset, and subpicScalWinBotOffset indicate the left, right, top, and bottom offsets of the scaling window applicable to the subpicture, respectively.
[0416] 11. The method according to any one of solutions 1 to 4, wherein the bottom offset subpicScalWinBotOffset of the scaling window is derived as follows:
[0417] subpicScalWinBotOffset = ( botSubpicBd >= pps_pic_height_in_luma_samples ) ? pps_scaling_win_bottom_offset : pps_scaling_win_bottom_offset - ( pps_pic_height_in_luma_samples - botSubpicBd ) / SubHeightC, and
[0418] Where pps_pic_height_max_in_luma_samples indicates the maximum height of each decoded image in the reference image parameter set in units of luminance samples, pps_scaling_win_bottom_offset indicates the bottom offset applied to the image size for scaling ratio calculation, SubHeightC indicates the height of the video block and is obtained from a table based on the chroma format of the image including the video block, and
[0419] where botSubpicBd = (sps_subpic_ctu_top_left_y[spIdx] +
[0420] sps_subpic_height_minus1[ spIdx ] + 1 ) CtbSizeY,
[0421] Where sps_subpic_ctu_top_left_y [spIdx] indicates the y-coordinate of the codec tree unit located at the top left corner of the subpicture, and CtbSizeY is the width or height of the luminance codec tree block or codec tree unit.
[0422] 12. The method according to Solution 11, wherein pps_pic_height_in_luma_samples comes from the original image parameter set.
[0423] 13. The method according to Solution 1, wherein the rule further specifies a reference sample enable flag in the rewrite sequence parameter set indicating the applicability of reference image resampling.
[0424] 14. The method according to Solution 1, wherein the rule further specifies a resolution change permission flag in the rewrite sequence parameter set, the resolution change permission flag indicating possible picture spatial resolution changes within a codec layer video sequence (CLVS) referencing the sequence parameter set.
[0425] 15. The method according to any one of solutions 1 to 14, wherein the conversion includes encoding the video into the bitstream.
[0426] 16. The method according to any one of solutions 1 to 14, wherein the conversion includes decoding the video from the bitstream.
[0427] 17. The method according to any one of solutions 1 to 14, wherein the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
[0428] 18. A video processing apparatus comprising a processor configured to implement the method according to any one or more of solutions 1 to 17.
[0429] 19. A method for storing a bitstream of video, comprising the method according to any one of solutions 1 to 17, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0430] 20. A computer-readable medium storing program code that, when executed, causes a processor to perform the method according to any one or more of solutions 1 to 17.
[0431] 21. A computer-readable medium storing a bit stream generated according to any one of the above methods.
[0432] 22. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method according to any one or more of solutions 1 to 17.
[0433] The third set of solutions illustrates example embodiments of the techniques discussed in the preceding sections (e.g., items 2 and 7).
[0434] 1. A video processing method (e.g., such as...) Figure 14A The method 1400 shown includes: performing a 1402 conversion between a video comprising one or more layers and a bitstream of the video according to a rule, the one or more layers comprising one or more video pictures, the one or more video pictures comprising one or more sub-pictures, wherein the rule defines a network abstraction layer (NAL) unit to be extracted from the bitstream to output a sub-bitstream during sub-bitstream extraction processing, and wherein the rule also specifies which combination of one or more inputs to the sub-bitstream extraction processing and / or sub-pictures of different layers of the bitstream is used such that the output of the sub-bitstream extraction processing conforms to a predefined format.
[0435] 2. The method as described in Solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process include a target output layer set (OLS) index (targetOlsIdx), which identifies the OLS index of the target OLS to be decoded and is equal to the index of the list of OLS specified by the video parameter set.
[0436] 3. The method as described in Solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process include the target highest time domain identifier value (tIdTarget).
[0437] 4. The method as described in Solution 3, wherein the highest temporal identifier value of the target is in the range of 0 to the maximum number of temporal sublayers allowed to exist in the layer specified by the video parameter set.
[0438] 5. The method as described in Solution 4, wherein the maximum number of the temporal sublayer is indicated by a syntax element included in the video parameter set.
[0439] 6. The method described in Solution 5, wherein the syntax element is vps_max_sublayers_minus1.
[0440] 7. The method as described in Solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process include a target subpic index value (subpicIdxtarget) equal to the subpic index present in the target OLS.
[0441] 8. The method as described in Solution 1, wherein the rule specifies that the one or more inputs to the sub-bitstream extraction process include a list of target subpic IdxTarget[i] for i from 0 to NumLayersInOls[targetOLsIdx]-1, where NumLayerInOls[i] specifies the number of layers in the i-th OLS, and targetOLsIdx indicates the index of the target output layer set (OLS).
[0442] 9. The method as described in Solution 8, wherein the rule specifies that all layers in the targetOLsIdx OLS have the same spatial resolution, the same subpicture layout, and that all subpictures are treated as syntactic elements of a picture in decoding processing that does not include loop filtering operations, with the corresponding subpicture in the specified codec layer video sequence being treated as a picture syntactic element.
[0443] 10. The method described in Solution 8, wherein the value of subpicIdxTarget[i] is the same for all values of i and is equal to a specific value in the range of 0 to sps_num_subpics_minus1, where sps_num_subpics_minus1 plus 1 specifies the number of subpicks in each picture of the codec layer video sequence.
[0444] 11. The method as described in Solution 8, wherein the rule specifies that the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit whose nuh_layer_id is equal to each of the nuh_layer_id values in the list of LayerIdInOls[targetOlsIdx], where nuh_layer_id is the NAL unit header identifier, and LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the OLS with the targetOlsIdx.
[0445] 12. The method as described in Solution 8, wherein the rule specifies that the output sub-bitstream contains at least one VCL (Video Coding Layer) NAL (Network Extraction Layer) unit having a TemporalId equal to tIdTarget, wherein TemporalId indicates the highest temporal identifier of the target, and tIdTarget is the value of TemporalId provided as one or more inputs to the sub-bitstream extraction process.
[0446] 13. The method as described in Solution 8, wherein the bitstream contains one or more encoded stripe NAL (Network Abstraction Layer) units with TemporalId equal to 0 without requiring the inclusion of encoded stripe NAL (Network Abstraction Layer) units with nuh_layer_id equal to 0, wherein nuh_layer_id specifies the NAL unit header identifier.
[0447] 14. The method as described in Solution 8, wherein the rule specifies that the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit, wherein for each i in the range 0 to NumLayersInOls[targetOlsIdx]-1, the NAL (Network Abstraction Layer) unit header identifier nuh_layer_id is equal to the layer identifier ayerIdInOls[targetOlsIdx][i] and the subpicture identifier sh_subpic_id is equal to SubpicIdVal[subpicIdxTarget[i]], wherein NumLayerInOls[i] specifies the number of layers in the i-th OLS and i is an integer.
[0448] 15. A video processing method (e.g., such as...) Figure 14B The method 1410 shown includes: performing a 1412 conversion between a video and a bitstream of the video according to a rule, wherein the bitstream includes a first layer and a second layer, the first layer including pictures having multiple sub-pictures, and the second layer including pictures each having a single sub-picture, and wherein the rule specifies a combination of the sub-pictures of the first layer and the second layer that results in an output bitstream conforming to a predefined format during extraction.
[0449] 16. The method of solution 15, wherein the rule is applied to sub-bitstream extraction to provide an output sub-bitstream as a consistent bitstream, and wherein the input to the sub-bitstream extraction process includes: i) the bitstream, ii) the target output layer set (OLS) index (targetOlsIdx) identifying the OLS index of the target OLS to be decoded, and ii) the target highest temporal identifier (TemporalId) value (tIdTarget).
[0450] 17. The method described in Solution 16, wherein the rule specifies that the target output layer set (OLS) index (targetOlsIdx) is equal to the index of the OLS list specified by the video parameter set.
[0451] 18. The method as described in Solution 16, wherein the rule specifies that the highest temporal identifier (TemporalId) value (tIdTarget) of the target is any value ranging from 0 to the maximum number of temporal sublayers allowed to exist in the layer specified by the video parameter set.
[0452] 19. The method as described in solution 18, wherein the maximum number of the temporal sublayer is indicated by the syntax elements included in the video parameter set.
[0453] 20. The method described in Solution 19, wherein the gai syntax element is vps_max_sublayer_minus1.
[0454] 21. The method as described in Solution 17, wherein the rule specifies that for i from 0 to NumLayersInOls[targetOLsIdx]-1, the value of the OLS list subpicIdxTarget[i] is equal to a value in the range of 0 to the number of subpicks in each picture of the layer, wherein NumLayerInOls[i] specifies the number of layers in the i-th OLS, and the number of subpicks is indicated by a syntax element included in the set of sequence parameters referenced by the layer associated with the NAL (Network Abstraction Layer) unit header identifier (nuh_layer_id), wherein nuh_layer_id is equal to LayerIdInOls[targetOLsIdx][i] which specifies the nuh_layer_id value in the i-th OLS with the targetOLsIdx, and i is an integer.
[0455] 22. The method described in solution 21, wherein another syntax element sps_subpic_treated_as_pic_flag[subpicIdxTarget[i]] is equal to 1, which specifies that the i-th subpic of each encoded / decoded picture in the layer is treated as a picture in the decoding process that does not include loop filtering operations.
[0456] 23. The method described in solution 21, wherein the rule specifies that the value of subpicIdxTarget[i] is always equal to 0 when the syntax element indicating the number of subpicks in each picture in the layer is equal to 0.
[0457] 24. The method described in solution 21, wherein the rule specifies that for any two distinct integer values of m and n, for two layers where nuh_layer_id is equal to LayerIdInOls[targetOLsIdx][m] and LayerIdInOls[targetOLsIdx][n] respectively, and subpicIdxTarget[m] is equal to subpicIdxTarget[n], the syntax element is greater than 0.
[0458] 25. The method as described in Solution 16, wherein the rule specifies that the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit whose nuh_layer_id is equal to each nuh_layer_id value in the list LayerIdInOls[targetOlsIdx], where nuh_layer_id is the NAL unit header identifier, and LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the OLS with the targetOlsIdx.
[0459] 26. The method as described in Solution 16, wherein the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit having a TemporalId equal to tIdTarget.
[0460] 27. The method as described in Solution 16, wherein the consistent bitstream contains one or more encoded stripe NAL (Network Abstraction Layer) units with a TemporalId equal to 0, without the need to contain encoded stripe NAL (Network Abstraction Layer) units with a nuh_layer_id equal to 0, wherein nuh_layer_id specifies the NAL unit header identifier.
[0461] 28. The method as described in Solution 16, wherein the output sub-bitstream contains at least one VCL (Video Codec Layer) NAL (Network Abstraction Layer) unit, the VCL NAL unit having a NAL (Network Abstraction Layer) unit header identifier nuh_layer_id equal to the layer identifier LayerIdInOls[targetOlsIdx][i], and for each i in the range from 0 to NumLayersInOls[targetOlsIdx]-1 having a subpic identifier sh_subpic_id equal to SubpicIdVal[subpicIdxTarget[i]], wherein NumLayerInOls[i] specifies the number of layers in the i-th OLS and i is an integer.
[0462] 29. A video processing method (e.g., such as...) Figure 14C The method shown (1420) includes: performing a 1422 conversion between a video and a bitstream of the video, and wherein a rule specifies whether or how an output layer set having one or more layers including multiple subpictures and / or one or more layers having a single subpicture is indicated by a list of subpicture identifiers in a Scalable Nested Supplemental Enhancement Information (SEI) message.
[0463] 30. The method as described in Solution 29, wherein the rule specifies that: when both sn_ols_flag and sn_subpic_flag are equal to 1, the list of subpicture identifiers specifies the subpicture identifiers of subpictures in a picture that each contains multiple subpictures, wherein sn_ols_flag equal to 1 specifies that the scalable nested SEI message is applied to a specific OLS, and sn_subpic_flag equal to 1 specifies that the scalable nested SEI message applied to a specified OLS or layer is applied only to a specific subpicture of the specified OLS or layer.
[0464] 31. The method of any one of solutions 1 to 30, wherein the conversion includes encoding the video into the bitstream.
[0465] 32. The method of any one of solutions 1 to 30, wherein the conversion includes decoding the video from the bitstream.
[0466] 33. The method of any one of solutions 1 to 30, wherein the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
[0467] 34. A video processing apparatus comprising a processor configured to implement the methods described in any one or more of solutions 1 to 33.
[0468] 35. A method for storing a bitstream of video, comprising the method described in any one of solutions 1 to 33, and further comprising: storing the bitstream to a non-transitory computer-readable recording medium.
[0469] 36. A computer-readable medium storing program code that, when executed, causes a processor to implement the methods described in any one or more of solutions 1 to 33.
[0470] 37. A computer-readable medium that stores a bit stream generated according to any one of the methods described above.
[0471] 38. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions 1 to 33.
[0472] The fourth set of solutions shows example embodiments of the techniques discussed in the preceding sections (e.g., items 3, 8, and 9).
[0473] 1. A video processing method (e.g., such as...) Figure 15A The method 1500 shown includes: performing a 1502 conversion between a video including one or more video images in the video layer and the bitstream of the video according to a rule, and wherein the rule specifies that: in the sub-bitstream extraction process, regardless of the availability of external components used to replace the parameter set removed during the sub-bitstream extraction, the removal of (i) Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, (ii) padding data NAL units associated with the VCL NAL units, and (iii) padding payload supplemental enhancement information (SEI) messages associated with the VCL NAL units are performed.
[0474] 2. The method as described in Solution 1, wherein the rule specifies: for each value i in the range of 0 to NumLayersInOls[targetOlsIdx]-1, remove from the sub-bitstream, which is the output of the sub-bitstream extraction process, all Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, their associated padding data NAL units, and their associated SEI containing padding payload SEI messages, all with nuh_layer_id equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id not equal to SubpicIdVal[subpicIdxTarget[i]]. The NAL unit is defined as follows: nuh_layer_id is the NAL unit header identifier; LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the output layer set (OLS), where targetOLsIdx is the index of the target output layer set; sh_subpic_id specifies the subpick identifier containing the subpick of the stripe; subpicIdxTarget[i] indicates the target subpick index value for i; and SubpicIdVal[subpicIdxTarget[i]] is the variable used for subpicIdxTarget[i], where i is an integer.
[0475] 3. A video processing method (e.g., such as...) Figure 15BThe method 1510 shown includes: performing a 1512 conversion between a video comprising one or more layers and the bitstream of the video according to a rule, the one or more layers comprising one or more video images, the one or more video images comprising one or more sub-images, and wherein the rule specifies that: in the case where a reference image is partitioned according to a sub-image layout different from the sub-image layout, the reference image is not allowed to be used as a co-bit image of the current image that is partitioned into sub-images using the sub-image layout.
[0476] 4. The method described in Solution 3, wherein the current image and the reference image are in different layers.
[0477] 5. A video processing method (e.g., such as...) Figure 15C The method 1520 shown includes: performing a 1522 conversion between a video comprising one or more layers and the bitstream of the video according to a rule, the one or more layers comprising one or more video pictures, the one or more video pictures comprising one or more subpictures, and wherein the rule specifies that: if the codec tool depends on a subpicture layout different from the subpicture layout of a reference picture used for the current picture, the codec tool is disabled during the conversion of the current picture which is divided into subpictures using that subpicture layout.
[0478] 6. The method described in Solution 5, wherein the current image and the reference image are in different layers.
[0479] 7. The method described in Solution 5, wherein the encoding / decoding tool is a bidirectional optical flow (BDOF), wherein optical flow calculations are used to refine one or more initial predictions.
[0480] 8. The method described in Solution 5, wherein the encoder / decoder tool is decoder-side motion vector refinement (DMVR), wherein motion information is refined by using prediction blocks.
[0481] 9. The method described in Solution 5, wherein the encoding / decoding tool utilizes Predictive Refinement of Optical Flow (PROF), wherein one or more initial predictions are refined based on optical flow calculations.
[0482] 10. The method of any one of solutions 1 to 9, wherein the conversion includes encoding the video into the bitstream.
[0483] 11. The method of any one of solutions 1 to 9, wherein the conversion includes decoding the video from the bitstream.
[0484] 12. The method of any one of solutions 1 to 9, wherein the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
[0485] 13. A video processing apparatus comprising a processor configured to implement the methods described in any one or more of solutions 1 to 12.
[0486] 14. A method for storing a bitstream of video, comprising the method described in any one of solutions 1 to 12, and further comprising: storing the bitstream to a non-transitory computer-readable recording medium.
[0487] 15. A computer-readable medium storing program code that, when executed, causes a processor to implement the methods described in any one or more of solutions 1 to 12.
[0488] 16. A computer-readable medium storing a bit stream generated according to any one of the methods described above.
[0489] 17. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions 1 to 12.
[0490] The fifth set of solutions illustrates example embodiments of the techniques discussed in the preceding sections (e.g., Project 4).
[0491] 1. A video processing method (e.g., such as...) Figure 16 The method shown (1600) includes: converting between a video including one or more video images in a video layer and the bitstream of that video according to a rule, wherein the rule specifies that a Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) unit containing an SEI message with a specific payload type does not contain another SEI message with a payload type different from that specific payload type.
[0492] 2. The method as described in Solution 1, wherein the rule specifies the removal of the SEI message as the removal of the SEI NAL unit containing the SEI message.
[0493] 3. The method as described in Solution 1 or 2, wherein the SEI message with that specific load type corresponds to the padding load SEI message.
[0494] 4. The method of any one of solutions 1 to 3, wherein the conversion includes encoding the video into the bitstream.
[0495] 5. The method of any one of solutions 1 to 3, wherein the conversion includes decoding the video from the bitstream.
[0496] 6. The method of any one of solutions 1 to 3, wherein the conversion includes generating the bitstream from the video, and the method further includes storing the bitstream in a non-transitory computer-readable recording medium.
[0497] 7. A video processing apparatus comprising a processor configured to implement the methods described in any one or more of solutions 1 to 6.
[0498] 8. A method for storing a bitstream of video, comprising the method described in any one of solutions 1 to 6, and further comprising storing the bitstream to a non-transitory computer-readable recording medium.
[0499] 9. A computer-readable medium storing program code that, when executed, causes a processor to implement the methods described in any one or more of solutions 1 to 6.
[0500] 10. A computer-readable medium storing a bit stream generated according to any one of the methods described above.
[0501] 11. A video processing apparatus for storing a bitstream representation, wherein the video processing apparatus is configured to implement the method described in any one or more of solutions 1 to 6.
[0502] In the solution described in this paper, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described in this paper, the decoder can use the format rules to parse the syntax elements in the codec representation to produce the decoded video, where the presence and absence of syntax elements are known according to the format rules.
[0503] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. The bitstream representation of the current video block can correspond to bits that are co-located or scattered at different positions within the bitstream, as defined by the syntax. For example, a macroblock can be encoded based on the transformed and encoded / decoded error residual values, and can also be encoded using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can parse the bitstream with knowledge that certain fields may or may not be present, based on the determinations described in the solutions above. Similarly, the encoder can determine whether certain syntax fields are included or not included, and generate the encoded / decoded representation accordingly by including or excluding syntax fields from the encoded / decoded representation.
[0504] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of that tool or mode. In other words, when the conversion from video block to bitstream representation of video is based on the decision or determination to enable the video processing tool or mode, the conversion will use that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream using knowledge that the bitstream has already been modified based on the video processing tool or mode. In other words, the conversion from bitstream representation of video to video block will be performed using a video processing tool or mode enabled based on the decision or determination.
[0505] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion from video block to bitstream representation of video. In another example, when a video processing tool or mode is disabled, the decoder will use the video processing tool or mode that has been disabled based on the decision or determination to process the bitstream with knowledge that the bitstream has not yet been modified.
[0506] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that implements machine-readable propagating signals, or combinations thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, such as a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations thereof. Propagating signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.
[0507] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code sections). A computer program can be deployed on a single computer for execution or deployed on multiple computers located at a single site or distributed across multiple sites and interconnected by a communication network for execution.
[0508] The processing and logic flows described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by dedicated logic circuitry, and the devices can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0509] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to receive data from or transfer data to one or more mass storage devices for storing data, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0510] Although this patent document contains numerous details, these details should not be construed as limiting the scope of any subject matter or potentially claimed content, but rather as descriptions of features of specific embodiments that may be specific to particular technologies. Certain features described in the context of individual embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in some cases, one or more features from the claimed combination may be removed from the combination, and the claimed combination may be directed to a sub-combination or a variation of the sub-combination.
[0511] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order or sequence shown, or requiring all illustrated operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0512] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A method for processing video data, comprising: The conversion between the video and the bitstream of the video is performed according to the rules, wherein the video includes one or more video images in the video layer; as well as The rule specifies that, in the sub-picture sub-bitstream extraction process used for output sub-bitstreams, for each value i in the range of 0 to NumLayersInOls[targetOlsIdx] - 1, all Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, their associated padding data NAL units, and their associated SEI NAL units containing padding payload supplementation enhancement information (SEI) messages are removed from the output sub-bitstream if nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is not equal to SubpicIdVal[subpicIdxTarget[i]]. Here, nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which a non-VCL NAL unit applies; LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the output layer set (OLS), where targetOLsIdx is the index of the target output layer set; sh_subpic_id specifies the subpick identifier containing the subpick of the stripe; subpicIdxTarget[i] indicates the target subpick index value for i; and SubpicIdVal[subpicIdxTarget[i]] is a variable of subpicIdxTarget[i], where i is an integer.
2. The method according to claim 1, wherein, The rule also specifies that a Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) unit containing an SEI message with a specific payload type does not contain other SEI messages with payload types different from the specific payload type.
3. The method according to claim 2, wherein, The SEI message with the specific payload type corresponds to the padding payload SEI message.
4. The method according to claim 1, wherein, The rule also specifies that, when sli_cbr_constrat_flag equals 1, the cbr_flag[tIdTarget][j] of the j-th CPB in the specific parameter syntax structure associated with the output layer set (OLS) hypothesis reference decoder (HRD) in all referenced VPSs will be set to equal 1, where j is in the range of 0 to hrd_cpb_cnt_minus1.
5. The method according to claim 4, wherein, The specific parameter syntax structure associated with the OLS HRD has an index idx[MultiLayerOlsIdx[targetOlsIdx]], where idx[MultiLayerOlsIdx[targetOlsIdx]] specifies a list of parameter syntax structures associated with the OLS HRD in the Video Parameter Set (VPS).
6. The method according to claim 4, wherein, A cbr_flag[i][j] equal to 0 specifies that the bitstream is to be decoded by the HRD using the j-th CPB specification, assuming the Stream Scheduler (HSS) is operating in intermittent bit rate mode, and a cbr_flag[i][j] equal to 1 specifies that the HSS is operating in constant bit rate (CBR) mode.
7. The method according to claim 4, wherein, The rule also specifies that, when sli_cbr_constraint_flag equals 1, the output sub-bitstream is extracted without removing all NAL units with nal_unit_type equal to FD_NUT and SEI NAL units containing the padding payload SEI message.
8. The method according to claim 4, wherein, The rule also specifies that when sli_cbr_constraint_flag equals 0, cbr_flag[tIdTarget][j] is set to 0, where j is in the range of 0 to hrd_cpb_cnt_minus1.
9. The method according to claim 8, wherein, The rule also specifies that, when sli_cbr_constraint_flag equals 0, the output sub-bitstream is extracted by removing all NAL units with nal_unit_type equal to FD_NUT and SEI NAL units containing the padding payload SEI message.
10. The method according to claim 1, wherein, The conversion includes encoding the video into the bitstream.
11. The method according to claim 1, wherein, The conversion includes decoding the video from the bitstream.
12. An apparatus for processing video data, comprising a processor and a non-transitory memory thereon with instructions, wherein, When the instruction is executed by the processor, the processor: The conversion between the video and its bitstream is performed according to rules, wherein the video comprises one or more video images in the video layer; and The rule specifies that, in the sub-picture sub-bitstream extraction process used for output sub-bitstreams, for each value i in the range of 0 to NumLayersInOls[targetOlsIdx] - 1, all Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, their associated padding data NAL units, and their associated SEI NAL units containing padding payload supplementation enhancement information (SEI) messages are removed from the output sub-bitstream if nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is not equal to SubpicIdVal[subpicIdxTarget[i]]. Here, nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which a non-VCL NAL unit applies; LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the output layer set (OLS), where targetOLsIdx is the index of the target output layer set; sh_subpic_id specifies the subpick identifier containing the subpick of the stripe; subpicIdxTarget[i] indicates the target subpick index value for i; and SubpicIdVal[subpicIdxTarget[i]] is a variable of subpicIdxTarget[i], where i is an integer.
13. The apparatus according to claim 12, wherein, The rule also specifies that a Supplemental Enhancement Information (SEI) Network Abstraction Layer (NAL) unit containing an SEI message with a specific payload type does not contain other SEI messages with payload types different from the specific payload type.
14. The apparatus according to claim 13, wherein, The SEI message with the specific payload type corresponds to the padding payload SEI message.
15. A non-transitory computer-readable storage medium for storing instructions, said instructions causing a processor to: The conversion between the video and its bitstream is performed according to rules, wherein the video comprises one or more video images in the video layer; and in, The rule specifies that in the sub-picture sub-bitstream extraction process used for output sub-bitstreams, for each value i in the range of 0 to NumLayersInOls[targetOlsIdx] - 1, all Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, their associated padding data NAL units, and their associated SEI NAL units containing padding payload supplementation enhancement information (SEI) messages are removed from the output sub-bitstream if nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is not equal to SubpicIdVal[subpicIdxTarget[i]]. Here, nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which a non-VCL NAL unit applies; LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the output layer set (OLS), where targetOLsIdx is the index of the target output layer set; sh_subpic_id specifies the subpick identifier containing the subpick of the stripe; subpicIdxTarget[i] indicates the target subpick index value for i; and SubpicIdVal[subpicIdxTarget[i]] is a variable of subpicIdxTarget[i], where i is an integer.
16. A non-transitory computer-readable recording medium storing a bitstream of video, wherein a computer program is also stored thereon, When the computer program is executed by a processor, it implements the method of claim 1 to generate the bit stream.
17. A method for storing a bitstream of video, comprising: The video bitstream is generated according to the rules, and the video includes one or more video images in the video layer; as well as The bitstream is stored in a non-transitory computer-readable recording medium. The rule specifies that, in the sub-picture sub-bitstream extraction process used for output sub-bitstreams, for each value i in the range of 0 to NumLayersInOls[targetOlsIdx] - 1, all Video Codec Layer (VCL) Network Abstraction Layer (NAL) units, their associated padding data NAL units, and their associated SEI NAL units containing padding payload supplementation enhancement information (SEI) messages are removed from the output sub-bitstream if nuh_layer_id is equal to LayerIdInOls[targetOlsIdx][i] and sh_subpic_id is not equal to SubpicIdVal[subpicIdxTarget[i]]. Here, nuh_layer_id specifies the identifier of the layer to which the VCL NAL unit belongs or the identifier of the layer to which a non-VCL NAL unit applies; LayerIdInOls[targetOlsIdx] specifies the nuh_layer_id value in the output layer set (OLS), where targetOLsIdx is the index of the target output layer set; sh_subpic_id specifies the subpick identifier containing the subpick of the stripe; subpicIdxTarget[i] indicates the target subpick index value for i; and SubpicIdVal[subpicIdxTarget[i]] is a variable of subpicIdxTarget[i], where i is an integer.
Citation Information
Patent Citations
Hypothetical reference decoder model and conformance for cross-layer random access skipped pictures
US20140355692A1