Subpicture track referencing and handling
By introducing sub-picture tracks and hybrid NAL unit types into the VVC video file format, the problem of carrying VVC video bitstreams in ISOBMFF is solved, achieving efficient video data processing and transmission, and making it suitable for various video file formats and streaming systems.
Patent Information
- Application Number
- CN202111094309.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-06
- Filing Date
- 2021-09-17
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2041-09-17
AI Technical Summary
Existing video codec technologies struggle to effectively handle sub-picture transport in Multi-Video Codec (VVC) video bitstreams, especially in media files based on the ISO Basic Media File Format (ISOBMFF), resulting in inefficient bandwidth utilization and excessive encoding/decoding overhead.
The concept of sub-picture tracks is adopted to divide video data into sub-pictures and stripes. The data is stored and transmitted in ISOBMFF according to the VVC video file format specification, allowing independent encoding, decoding and extraction of sub-pictures. It supports viewport-dependent 360-degree video streaming and region of interest applications, and achieves random access through a hybrid NAL unit type.
It improves the bandwidth utilization efficiency of video data, reduces end-to-end latency and encoding/decoding overhead, supports flexible video data processing and transmission, and is suitable for various video file formats and streaming systems.
Smart Images

Figure CN114205609B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is made to timely claim priority to, and the benefit of, U.S. Provisional Patent Application No. 63 / 079,933 filed on September 17, 2020, and U.S. Provisional Patent Application No. 63 / 088,126 filed on October 6, 2020, under the applicable patent laws and / or under the rules of the Paris Convention. The entire disclosure of the foregoing applications is hereby incorporated by reference herein, in its entirety, as part of the disclosure of this application, for all purposes. TECHNICAL FIELD
[0003] This patent document relates to the generation, storage, and consumption of digital audio-visual media information in file formats. BACKGROUND
[0004] Digital video accounts for the largest bandwidth use on the internet and other digital communication networks. As the number of networked user devices capable of receiving and displaying video increases, it is expected that bandwidth demand for digital video usage will continue to grow. SUMMARY
[0005] This document discloses techniques that can be used by video encoders and decoders to process coded representations of video or images according to file formats.
[0006] In one example aspect, a method for processing visual media data is disclosed. The method comprises performing a conversion between the visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; and wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; wherein the visual media file stores the one or more tracks according to a format rule; and wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures comprises a rectangular region covering the one or more pictures.
[0007] In another example aspect, a method for processing visual media data is disclosed. The method comprises performing a conversion between the visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file comprises a base track referencing one or more sub-picture tracks, the sub-picture tracks storing coded information of one or more sub-pictures of the visual media data; and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
[0008] In yet another example aspect, a video processing apparatus is disclosed. The video processing apparatus includes a processor configured to implement a method recited above.
[0009] In yet another example aspect, a method of storing visual media data into a file including one or more bitstreams is disclosed. The method corresponds to the method recited above and further includes storing the one or more bitstreams into a non-transitory computer-readable recording medium.
[0010] In yet another example aspect, a computer-readable medium storing a bitstream is disclosed. The bitstream is generated according to the method recited above.
[0011] In yet another example aspect, a video processing apparatus storing a bitstream is disclosed, wherein the video processing apparatus is configured to implement the method recited above.
[0012] In yet another example aspect, a computer-readable medium is disclosed, on which a bitstream conforms to a file format generated according to the method recited above.
[0013] These and other features are described throughout this document. BRIEF DESCRIPTION OF DRAWINGS
[0014] Figure 1 is a block diagram of an example video processing system.
[0015] Figure 2 is a block diagram of a video processing apparatus.
[0016] Figure 3 is a flowchart of an example method of video processing.
[0017] Figure 4 is a block diagram illustrating a video coding system according to some embodiments of the present disclosure.
[0018] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0019] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0020] Figure 7 An example of an encoder block diagram is shown.
[0021] Figure 8 A picture that is partitioned into 18 tiles, 24 slices, and 24 sub-pictures is shown.
[0022] Figure 9 A typical sub-picture based viewport dependent 360-degree video delivery scheme is shown.
[0023] Figure 10An example is shown of extracting one subpicture from a bitstream containing two subpictures and four slices.
[0024] Figure 11 and Figure 12 An example method of processing visual media data based on some embodiments of the disclosed technology is shown. DETAILED DESCRIPTION
[0025] For ease of understanding, section headings are used in this document, and the teachings and embodiments disclosed in each section are not meant to be limited to that section only. Also, the use of H.266 terminology in some descriptions is merely for ease of understanding, and is not meant to limit the scope of the disclosed technology. As such, the technology described herein is applicable to other video codec protocols and designs as well. In this document, editorial changes to the text relative to the current draft of the VVC specification or ISOBMFF file format specification are shown with strikeout and highlight, strikeout indicating deleted text, and highlight indicating added text (including boldface italics).
[0026] 1. Preliminary discussion
[0027] This document is related to video file formats. In particular, it is concerned with the carriage of subpictures of Versatile Video Coding (VVC) video bitstreams in multiple tracks in ISO Base Media File Format (ISOBMFF) based media files. These ideas can be applied to video bitstreams coded by any codec (e.g., the VVC standard), and to any video file format (e.g., the VVC video file format under development), individually or in various combinations.
[0028] 2. Abbreviations
[0029] ACT adaptive colour transform
[0030] ALF adaptive loop filter
[0031] AMVR adaptive motion vector resolution
[0032] APS adaptation parameter set
[0033] AU access unit
[0034] AUD access unit delimiter
[0035] AVC advanced video coding
[0036] B bi-predictive
[0037] BCW bi-prediction with CU-level weights
[0038] BDOF bi-directional optical flow
[0039] BDPCM block-based delta pulse code modulation
[0040] BP buffering period
[0041] CABAC context-based adaptive binary arithmetic coding
[0042] CB coding block
[0043] CBR constant bit rate
[0044] CCALF cross-component adaptive loop filter
[0045] CPB coded picture buffer
[0046] CRA clean random access
[0047] CRC cyclic redundancy check
[0048] CTB coding tree block
[0049] CTU coding tree unit
[0050] CU coding unit
[0051] CVS coded video sequence
[0052] DPB decoded picture buffer
[0053] DCI decoding capability information
[0054] DRAP dependent random access point
[0055] DU decoding unit
[0056] DUI decoding unit information
[0057] EG exponential - Golomb index - Columbus
[0058] EGk k-th order exponential-Golomb k-order exponential-Columbus
[0059] EOB (End of Bitstream)
[0060] EOS end of sequence
[0061] FD filler data
[0062] FIFO (First-in, First-out)
[0063] FL fixed-length
[0064] GBR green, blue, and red
[0065] GCI general constraints information
[0066] GDR gradual decoding refresh
[0067] GPM geometric partitioning mode
[0068] HEVC high efficiency video coding
[0069] HRD hypothetical reference decoder
[0070] HSS hypothetical stream scheduler
[0071] I intra
[0072] IBC intra block copy
[0073] IDR instantaneous decoding refresh
[0074] ILRP inter-layer reference picture
[0075] IRAP intra random access point
[0076] LFNST low frequency non-separable transform
[0077] LPS least probable symbol
[0078] LSB least significant bit
[0079] LTRP long-term reference picture
[0080] LMCS luma mapping with chroma scaling
[0081] MIP matrix-based intra prediction
[0082] MPS most probable symbol
[0083] MSB most significant bit
[0084] MTS multiple transform selection
[0085] MVP motion vector prediction
[0086] NAL network abstraction layer
[0087] OLS output layer set
[0088] OP operation point
[0089] OPI operating point information
[0090] P predictive prediction
[0091] PH picture header
[0092] POC picture order count
[0093] PPS picture parameter set
[0094] PROF prediction refinement with optical flow
[0095] PT picture timing
[0096] PU picture unit
[0097] QP quantization parameter
[0098] RADL (Random Access Decodable Leading (Picture))
[0099] RASL random access skipped leading (picture)
[0100] RBSP raw byte sequence payload
[0101] RGB red, green, and blue
[0102] RPL reference picture list
[0103] SAO sample adaptive offset
[0104] SAR sample aspect ratio
[0105] SEI supplemental enhancement information
[0106] SH slice header
[0107] SLI subpicture level information
[0108] SODB string of data bits
[0109] SPS sequence parameter set
[0110] STRP short-term reference picture
[0111] STSA step-wise temporal sublayer access
[0112] TR truncated rice
[0113] VBR variable bit rate
[0114] VCL video coding layer
[0115] VPS video parameter set
[0116] VSEI versatile supplemental enhancement information (Rec. ITU-T H.274 | ISO / IEC 23002-7)
[0117] VUI video usability information
[0118] VVC versatile video coding (Rec. ITU-T H.266 | ISO / IEC 23090-3)
[0119] 3. Introduction to video coding
[0120] 3.1. Video coding standards
[0121] Video coding standards have evolved mainly through the well-known development by ITU-T and ISO / IEC. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially started, the JVET was later renamed as Joint Video Expert Team (JVET). VVC is a new coding standard targeting at 50% bitrate reduction compared to HEVC, which has been finalized by the JVET at its 19th meeting, which ended on 1st July 2020.
[0122] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the broadest range of applications, including traditional uses such as television broadcast, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple coded video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive 360-degree immersive media.
[0123] 3.2. File format standards
[0124] Media streaming applications are typically based on IP, TCP and HTTP transport methods and often rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, video format specific file format specifications such as the AVC File Format and the HEVC File Format in ISO / IEC 14496-15 ("Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format") would be needed for encapsulating video content in ISOBMFF tracks and in DASH representations and segments. Important information about the video bitstream (e.g. profile, tier and level) and many other information would need to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, e.g. selection of appropriate media segments both for initialization at the start of a streaming session and for stream adaptation during a streaming session.
[0125] Similarly, for using picture formats with ISOBMFF, picture format specific file format specifications such as the AVC Image File Format and the HEVC Image File Format in ISO / IEC 23008-12 ("Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format") would be needed.
[0126] The VVC video file format, i.e., a file format to store VVC video content based on ISOBMFF, is currently being developed by MPEG. The latest draft specification of the VVC video file format is included in the MPEG output document N19454 (“Information technology - Coding of audio-visual objects - Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format - Amendment 2: Carriage of VVC and EVC in ISOBMFF”, July 2020).
[0127] The VVC image file format, i.e., a file format to store image content coded using VVC based on ISOBMFF, is currently being developed by MPEG. The latest draft specification of the VVC image file format is included in the MPEG output document N19460 (“Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 12: Image File Format - Amendment 3: Support for VVC, EVC, slideshows and other improvements”, July 2020).
[0128] 3.3. Picture partitioning schemes in HEVC
[0129] HEVC includes four different picture partitioning schemes, i.e., regular slice, dependent slice, tile, and Wavefront Parallel Processing (WPP), which can be used for Maximum Transfer Unit (MTU) size matching, parallel processing, and reduced end-to-end delay.
[0130] A regular slice is similar to the one in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, coding mode prediction) and entropy coding dependencies across slice boundaries are disabled. Thus, a regular slice can be reconstructed independently from other regular slices within the same picture (although there can still be interdependencies due to loop filtering operations).
[0131] A regular slice is the only tool that can be used for parallelization, which is also available in H.264 / AVC in almost the same form. Parallelization based on regular slices does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is usually much more burdensome than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, using regular slices can result in a lot of coding overhead due to the bit cost of the slice header and due to the lack of prediction across slice boundaries. Furthermore, regular slices (in contrast to the other tools mentioned below) are also used as a key mechanism for bitstream segmentation to match MTU size requirements, because of the intra-picture independence of regular slices and the fact that each regular slice is encapsulated in its own NAL unit. In many cases, the goals of parallelization and MTU size matching put contradictory requirements on the slice layout in a picture. The implementation of this scenario led to the development of the parallelization tools mentioned below.
[0132] A dependent slice has a short slice header and allows for segmentation of the bitstream at tree block boundaries without breaking any intra-picture prediction. Basically, a dependent slice cuts a regular slice into multiple NAL units to provide reduced end-to-end delay by allowing a part of a regular slice to be sent out before the encoding of the entire regular slice is completed.
[0133] In WPP, a picture is partitioned into single-row coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible by decoding CTB rows in parallel, with the start of decoding of a CTB row delayed by two CTBs, ensuring that data related to CTBs above and to the right of the subject CTB is available before the subject CTB is decoded. Using this staggered start (which looks like a wavefront when represented graphically), parallelization with up to as many processors / cores as the picture contains CTB rows is possible. Because of the permission of intra-picture prediction between neighboring tree block rows within a picture, inter-processor / inter-core communication required to implement intra-picture prediction can be substantial. WPP partitioning does not create additional NAL units compared to when no WPP partitioning is applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is needed, regular slices can be used with WPP, but at some coding overhead.
[0134] Tiles define horizontal and vertical boundaries that partition a picture into tile columns and tile rows. Tile columns extend from the top of the picture to the bottom of the picture. Likewise, tile rows extend from the left side of the picture to the right side of the picture. The number of tiles in a picture can simply be derived by multiplying the number of tile columns by the number of tile rows.
[0135] The scan order of CTBs is changed to local within a tile (in the order of tile-wise CTB raster scan) before the top-left CTB of the next tile is decoded in the order of picture-wise tile raster scan. Similar to regular slices, tiles break intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included into separate NAL units (same as WPP in this regard); therefore, tiles cannot be used for MTU size matching. Each tile can be processed by one processor / core, and inter-processor / inter-core communication required for intra-picture prediction between processing units decoding neighboring tiles is limited to communicating the shared slice header in case a slice spans more than one tile, and loop filtering related to sharing of reconstructed samples and metadata. When more than one tile or WPP segment is included in a slice, the entry point byte offset for each tile or WPP segment other than the first tile or WPP segment in the slice is signaled in the slice header.
[0136] For simplicity, constraints on the application of the four different picture partitioning schemes have been specified in HEVC. A given coded video sequence cannot include both tiles and wavefronts for most profiles specified for HEVC. For each tile and slice, one or both of the following conditions must be met: 1) all coded treeblocks in a slice belong to the same tile; 2) all coded treeblocks in a tile belong to the same slice. Finally, a wavefront tile contains exactly one CTB row, and when WPP is in use, if a slice starts within a CTB row, it must end in the same CTB row.
[0137] Recent modifications to HEVC are specified in the following JCT-VC output document: JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G. J. Sullivan, A. Tourapis, Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)," Oct. 24, 2017, published for public disclosure at: http: / / phenix.int- evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Among other things, HEVC specifies three MCTS-related SEI messages, namely, the temporal MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nesting SEI message.
[0138] The temporal MCTS SEI message indicates the presence of MCTSs in the bitstream and signals the MCTSs. For each MCTS, the motion vectors are constrained to point to full-sample locations inside the MCTS and to fractional-sample locations that only require full-sample locations inside the MCTS for interpolation, and motion vector candidates are not allowed for temporal motion vector prediction derived from blocks outside the MCTS. In this way, each MCTS can be decoded independently without the presence of tiles that are not included in the MCTS.
[0139] The MCS extraction information set SEI message provides supplemental information that can be used in the MCTS sub-bitstream extraction (as part of the semantics of the SEI message) to generate a conforming bitstream of the MCTS set. The information includes multiple extraction information sets, each defining multiple MCTS sets and containing the RBSP bytes of the replacement VPS, SPS, and PPS used during the MCTS sub-bitstream extraction process. When a sub-bitstream is extracted according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated as one or all of the syntax elements related to slice address (including first_slice_segment_in_pic_flag and slice_segment_address) generally need to have different values.
[0140] 3.4. Picture partitioning and subpictures in VVC
[0141] 3.4.1. Picture partitioning in VVC
[0142] In VVC, a picture is partitioned into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular region of a picture. The CTUs in a tile are scanned in raster scan order within the tile.
[0143] A slice contains an integer number of consecutive complete CTU rows within an integer number of complete tiles or within a picture.
[0144] Two modes of slices are supported, namely, raster-scan slice mode and rectangular slice mode. In the raster-scan slice mode, a slice contains a sequence of complete tiles in the raster scan of the tiles of a picture. In the rectangular slice mode, a slice contains multiple complete tiles that collectively form a rectangular region of a picture, or contains multiple consecutive complete CTU rows of one tile that collectively form a rectangular region of a picture. The tiles within a rectangular slice are scanned in tile raster scan order within the rectangular region corresponding to the slice.
[0145] A subpicture contains one or more slices that collectively cover a rectangular region of a picture.
[0146] 3.4.2. Subpicture concept and functionality
[0147] In VVC, each subpicture consists of one or more complete rectangular slices that collectively cover a rectangular region of a picture, as illustrated in Figure 8A subpicture can be designated as extractable (i.e., coded independently of other subpictures of the same picture and independently of earlier pictures in decoding order), or non-extractable. Whether a subpicture is extractable or not, the encoder can control whether to apply in-loop filtering (including deblocking, SAO, and ALF) separately on each subpicture across subpicture boundaries.
[0148] Functionally, subpictures are similar to motion-constrained tilesets (MCTS) in HEVC. Both allow independent coding and extraction of rectangular subsets of a sequence of coded pictures for use cases such as viewport-dependent 360-degree video streaming optimization and region-of-interest (ROI) applications.
[0149] In the streaming of 360-degree video (also known as omnidirectional video), only a subset of the entire omnidirectional video sphere (i.e., the current viewport) will be presented to the user at any particular time, while the user can at any time turn his / her head to change the viewing orientation, and thus the current viewport. While it is desirable to have at least some lower-quality representation of the regions of the current viewport not covered at the client, and be ready to present it to the user in case the user suddenly changes his / her viewing orientation to anywhere on the sphere, the high-quality representation of the omnidirectional video only needs to be for the current viewport being presented to the user at any given time. Partitioning the high-quality representation of the entire omnidirectional video into subpictures at an appropriate granularity enables optimizations such as Figure 8 An optimization is shown, where the left-hand side is 12 high-resolution subpictures, and the right-hand side is the remaining 12 lower-resolution subpictures of the omnidirectional video.
[0150] Figure 9 Another typical subpicture-based viewport-dependent 360-degree video delivery scheme is shown, where only the higher-resolution representation of the full video includes subpictures, while the lower-resolution representation of the full video does not use subpictures, and can be coded with RAPs less frequently than the higher-resolution representation. The client receives the lower-resolution full video, while for the higher-resolution video, the client only receives and decodes the subpicture covering the current viewport.
[0151] 3.4.3. Differences between subpictures and MCTs
[0152] Several important design differences exist between subpictures and MCT. First, the subpicture feature in VVC allows motion vectors of the codec block to point outside the subpicture, even though the subpicture is extractable by applying sample padding at its boundaries, similar to that at the image boundaries. Second, additional changes are introduced to the selection and derivation of motion vectors during the Merge mode and decoder-side motion vector refinement in VVC. This allows for higher encoding / decoding efficiency compared to the non-standard motion constraints applied on the encoder side in MCTS. Third, when extracting one or more extractable subpictures from a picture sequence to create a subbitstream as a consistent bitstream, SH (and PH NAL units, if present) does not need to be rewritten. In HEVC MCTS-based subbitstream extraction, SH needs to be rewritten. Note that SPS and PPS need to be rewritten in both HEVC MCTS extraction and VVC subpicture extraction. However, typically only a few parameter sets exist in the bitstream, and each picture has at least one stripe, so rewriting SH can be a significant burden for the application system. Fourth, it allows stripes of different sub-pictures within an image to have different NAL unit types. This feature is often referred to as mixed NAL unit types or mixed sub-picture types within an image, which will be discussed in detail below. Fifth, VVC specifies HRD and level definitions for the sub-picture sequence, so the encoder can ensure the consistency of the sub-bitstream for each extractable sub-picture sequence.
[0153] 3.4.4. Intra-image Mixed Sub-image Type
[0154] In AVC and HEVC, all VCL NAL units within a picture must have the same NAL unit type. VVC introduces the option to mix subpictures with certain different VCL NAL unit types within a picture, thus providing random access support not only at the picture level but also at the subpicture level. In VVC, VCL NAL units within a subpicture still need to have the same NAL unit type.
[0155] The ability to randomly access sub-images from IRAP is beneficial for 360-degree video applications. In conjunction with... Figure 9 In a similar viewport-dependent 360-degree video transmission scheme, the content of spatially adjacent viewports largely overlaps. That is, during a viewport orientation change, only a small fraction of sub-pictures in the viewport are replaced by new sub-pictures, while the majority of sub-pictures remain in the viewport. The newly introduced sub-picture sequence must begin with an IRAP stripe, but by allowing inter-frame prediction of the remaining sub-pictures during viewport changes, a significant reduction in the overall transmission bit rate can be achieved.
[0156] The indication of whether a picture contains only NAL units of a single type or more than one type is provided in the PPS referred to by the picture (i.e., using a flag named pps_mixed_nalu_types_in_pic_flag). A picture can contain both sub-pictures containing IRAP slices and sub-pictures containing trailing slices. Some other combinations of different NAL unit types within a picture are also allowed, including leading picture slices of NAL unit types RASL and RADL, which allows merging sub-picture sequences with open and closed GOP coding structures extracted from different bitstreams into one bitstream.
[0157] 3.4.5. Sub-picture layout and ID signaling
[0158] The layout of sub-pictures in VVC is signaled in the SPS and thus remains constant within a CLVS. Each sub-picture is signaled by the position of its top-left CTU and its width and height in number of CTUs, ensuring that the sub-picture covers a rectangular region of the picture at CTU granularity. The order of sub-pictures signaled in the SPS determines the index of each sub-picture within the picture.
[0159] In order to enable extraction and merging of sub-picture sequences without rewriting SHs or PHs, the VVC slice addressing scheme is based on sub-picture IDs and sub-picture-specific slice indices to associate a slice with a sub-picture. In the SH, the sub-picture ID of the sub-picture containing the slice and the sub-picture-level slice index are signaled. Note that the value of the sub-picture ID of a particular sub-picture can be different from the value of its sub-picture index. The mapping between the two is signaled in the SPS or PPS (but not both) or is implicitly inferred. When present, the sub-picture ID mapping needs to be rewritten or added during the sub-picture sub-bitstream extraction process when rewriting the SPS and PPS. Together, the sub-picture ID and the sub-picture-level slice index indicate to the decoder the exact location of the first decoded CTU of a slice within the DPB Figure 10 The use of sub-picture IDs, sub-picture indices, and sub-picture-level slice indices to enable sub-picture extraction is illustrated with an example containing two sub-pictures and four slices.
[0160] Similar to subpicture extraction, the signaling of subpictures allows to merge several subpictures from different bitstreams into a single bitstream by just rewriting the SPS and PPS, provided that the different bitstreams are co-generated (e.g. using different subpicture IDs but otherwise aligned SPS, PPS and PH parameters such as CTU size, chroma format, coding tools, etc.).
[0161] Although subpictures and slices are independently signaled in SPS and PPS respectively, there are inherent mutual constraints between subpicture and slice layout in order to form a consistent bitstream. First, the presence of subpictures requires the use of rectangular slices and prohibits raster-scan slices. Second, the slices of a given subpicture shall be consecutive NAL units in decoding order, which means that the subpicture layout constrains the order of the coded slice NAL units within the bitstream.
[0162] 3.5. Some details of the VVC video file format
[0163] 3.5.1. Types of tracks
[0164] The VVC video file format specifies the following types of video tracks for carrying VVC bitstreams in ISOBMFF files:
[0165] a) VVC track:
[0166] A VVC track represents a VVC bitstream by including NAL units in its samples and sample entries, and possibly by referencing other VVC tracks containing other sublayers of the VVC bitstream, and possibly by referencing a VVC subpicture track. When a VVC track references a VVC subpicture track, it is called a VVC base track.
[0167] b) VVC non-VCL track:
[0168] APS carrying ALF, LMCS or scaling list parameters and other non-VCL NAL units can be stored and sent through a track separate from the track containing VCL NAL units; this is the VVC non-VCL track.
[0169] c) VVC subpicture track:
[0170] A VVC subpicture track contains any of the following:
[0171] A sequence of one or more VVC subpictures.
[0172] A sequence of one or more complete slices forming a rectangular region.
[0173] The samples of a VVC subpicture track contain any of the following:
[0174] one or more complete sub-pictures that are consecutive in decoding order as specified in ISO / IEC 23090-3.
[0175] one or more complete slices that form a rectangular region and are consecutive in decoding order as specified in ISO / IEC 23090-3.
[0176] the VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order.
[0177] NOTE: The VVC non-VCL tracks and VVC sub-picture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, a client can request the DASH representation containing the subset of VVC sub-picture tracks and the DASH representation containing the non-VCL tracks on a segment-by-segment basis. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided.
[0178] 3.5.2. Overview of rectangular regions carried in VVC bitstreams
[0179] This document provides support for describing rectangular regions including any of the following:
[0180] - a sequence of one or more VVC sub-pictures that are consecutive in decoding order, or
[0181] - a sequence of one or more complete slices that form a rectangular region and are consecutive in decoding order.
[0182] A rectangular region covers a rectangle without holes. Rectangular regions within a picture do not overlap each other.
[0183] A rectangular region can be described by a rectangular region visual sample group description entry (i.e., an instance of RectangularRegionGroupEntry) with rect_region_flag equal to 1.
[0184] If each sample of a track includes only one rectangular region of NAL units, a SampleToGroupBox of type "trif" can be used to associate the sample with the rectangular region, but this SampleToGroupBox of type "trif" can be omitted if the default sample grouping mechanism is used (i.e. when the version of the SampleGroupDescriptionBox of type "trif" is equal to or greater than 2). Otherwise, the samples, NAL units and rectangular regions are associated with each other by a SampleToGroupBox of type "nalm" and a SampleGroupDescriptionBox of type "nalm" with grouping_type_parameter equal to "trif". A RectangularRegionGroupEntry describes:
[0185] - a rectangular region,
[0186] - coding dependencies between this rectangular region and other rectangular regions.
[0187] Each RectangularRegionGroupEntry is assigned a unique identifier, called groupID. This identifier can be used to associate NAL units in a sample with a particular RectangularRegionGroupEntry.
[0188] The position and size of a rectangular region are identified using luma sample coordinates.
[0189] When used with movie fragments, RectangularRegionGroupEntry can be defined for the duration of a movie fragment by defining a new SampleGroupDescriptionBox in the track fragment box as defined in clause 8.9.4 of ISO / IEC 14496-12. However, there shall not be any RectangularRegionGroupEntry in a track fragment with the same groupID as a RectangularRegionGroupEntry already defined.
[0190] The underlying region used in a RectangularRegionGroupEntry is the picture to which the NAL units in the rectangular region associated with this RectangularRegionGroupEntry belong.
[0191] If there is any change in the underlying region size in consecutive samples (e.g., in case of reference picture resampling (RPR) or SPS resizing), the samples shall be associated with different RectangularRegionGroupEntry reflecting their respective underlying region size.
[0192] NAL units that map to rectangular regions can be carried in the VVC track as usual or in a separate track called VVC subpicture track.
[0193] 3.5.3. Reconstructing a picture unit from samples in the VVC track referring to the VVC subpicture track
[0194] Samples of a VVC track are parsed as access units that contain the following NAL units in the order of the items:
[0195] • AUD NAL units, if any, when present in the sample (and are the first NAL units in the sample).
[0196] • Parameter sets and SEI NAL units, if any, contained in the sample entry associated with the sample, when the sample is the first sample in the sample sequence associated with the same sample entry.
[0197] • NAL units present in the sample up to and including the PH NAL unit.
[0198] • The content of the (temporally) time-aligned parsed samples in each referred VVC subpicture track in the order specified in the "spor" sample group description entry mapped to the sample, excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units, if any. The track references are parsed as specified in the following specification.
[0199] NOTE 1: If the referred VVC subpicture track is associated with a VVC non-VCL track, the parsed samples of the VVC subpicture track contain the non-VCL NAL unit(s) of the time-aligned samples in the VVC non-VCL track, if any.
[0200] • NAL units in the sample following the PH NAL unit.
[0201] NOTE 2: NAL units in the sample following the NAL unit can include suffix SEI NAL units, suffix APS NAL units, EOS NAL units, EOB NAL units, or reserved NAL units allowed after the last VCL NAL unit. The "subp" track reference index of the "spor" sample group description entry is parsed as follows:
[0202] • If the track reference points to a track ID of a VVC subpicture track, the track reference is resolved to a VVC subpicture track.
[0203] • Otherwise (the track reference points to an "alte" track group), the track reference is resolved to any track in the "alte" track group. If a particular track reference index value was resolved to a particular track in a previous sample, it shall be resolved to any of the following in the current sample:
[0204] • The same particular track, or
[0205] • Any other track in the same "alte" track group that contains a Sync sample that is time-aligned with the current sample.
[0206] NOTE 3: VVC subpicture tracks in the same "alte" track group must be independent of any other VVC subpicture tracks referenced by the same VVC base track to avoid decoding mismatches and can therefore be constrained as follows:
[0207] • All VVC subpicture tracks contain VVC subpictures.
[0208] • Subpicture boundaries are treated like picture boundaries.
[0209] • Cross-subpicture boundary loop filtering is turned off.
[0210] If the reader selects a VVC subpicture track containing VVC subpictures, where the set of subpicture ID values is either the initial selection or different from the previous selection, the following steps can be taken:
[0211] • The "spor" sample group description entry is consulted to conclude whether a PPS or SPS NAL unit needs to be changed.
[0212] NOTE: It is only possible to change the SPS at the beginning of a CLVS.
[0213] • If the "spor" sample group description entry indicates that a start code emulation prevention byte is present before or within the subpicture ID in the contained NAL unit, the RBSP is derived from the NAL unit (i.e., the start code emulation prevention byte is removed). After rewriting in the next step, start code emulation prevention is re-applied.
[0214] • The reader uses the bit position and subpicture ID length information in the "spor" sample group entry to conclude which bits are overwritten to update the subpicture ID to the selected subpicture ID.
[0215] • When the subpicture ID value of a PPS or SPS is initially selected, the reader needs to overwrite the PPS or SPS in the reconstructed access unit with the selected subpicture ID value, respectively.
[0216] • When the subpicture ID value of a PPS or SPS changes compared to the previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value, the reader needs to include a copy of that previous PPS and SPS (if the PPS or SPS with the same PPS or SPS ID value, respectively, does not otherwise exist in the access unit) and overwrite the PPS or SPS in the reconstructed access unit with the updated subpicture ID value, respectively.
[0217] 3.5.4. Subpicture order sample group
[0218] 3.5.4.1. Definition
[0219] This sample group is used in VVC base tracks, i.e. VVC tracks with "subp" tracks referring to VVC subpicture tracks. Each sample group description entry indicates subpictures or slices of coded pictures in decoding order, where each index of a track of "subp" type refers to one or more subpictures or slices that are consecutive in decoding order.
[0220] To facilitate PPS or SPS rewriting in response to subpicture selection, each sample group description entry can contain:
[0221] - an indication whether the selected subpicture ID shall be changed in the PPS or SPS NAL unit;
[0222] - the length (in bits) of the subpicture ID syntax element;
[0223] - the bit position of the subpicture ID syntax element in the contained RBSP;
[0224] - a flag indicating whether the start code emulation prevention bytes are present before or within the subpicture ID;
[0225] - the parameter set ID of the parameter set containing the subpicture ID.
[0226] 3.5.4.2. Syntax
[0227]
[0228] 3.5.4.3. Semantics
[0229] subpic_id_info_flag equal to 0 specifies that the subpicture ID values provided in the SPS and / or PPS are correct for the indicated set of subp_track_ref_idx values and thus no rewriting of the SPS or PPS is needed. subpic_info_flag equal to 1 specifies that the SPS and / or PPS can need to be rewritten to indicate the subpictures corresponding to the set of subp_track_ref_idx values.
[0230] num_subpic_ref_idx specifies the number of reference indices of subpicture tracks or track groups of subpicture tracks referred by the VVC track.
[0231] subp_track_ref_idx, for each value of i, specifies the "subp" track reference index of the i-th list of one or more subpictures or slices to be included in the VVC bitstream reconstructed according to the VVC track.
[0232] subpic_id_len_minus1 plus 1 specifies the number of bits in the subpicture identifier syntax element in the PPS or SPS (as referred to by this structure).
[0233] subpic_id_bit_pos specifies the bit position of the first bit of the first subpicture ID syntax element in the referred PPS or SPS RBSP from 0.
[0234] start_code_emul_flag equal to 0 specifies that the start code emulation prevention bytes are not present before or within the subpicture ID of the referred PPS or SPS NAL unit.
[0235] start_code_emul_flag equal to 1 specifies that the start code emulation prevention bytes can be present before or within the subpicture ID of the referred PPS or SPS NAL unit.
[0236] pps_subpic_id_flag, when equal to 0, specifies that the PPS NAL unit applied to the samples mapped to this sample group description entry does not contain a subpicture ID syntax element.
[0237] pps_subpic_id_flag, when equal to 1, specifies that the PPS NAL unit applied to the samples mapped to this sample group description entry contains a subpicture ID syntax element.
[0238] pps_id, when present, specifies the PPS ID of the PPS applied to the samples mapped to this sample group description entry.
[0239] sps_subpic_id_flag, when present and equal to 0, specifies that the SPS NAL unit that applies to the samples mapped to this sample group description box does not contain subpicture ID syntax elements and the subpicture ID values are inferred. sps_subpic_id_flag, when present and equal to 1, specifies that the SPS NAL unit that applies to the samples mapped to this sample group description box contains subpicture ID syntax elements.
[0240] sps_id, when present, specifies the SPS ID of the SPS that applies to the samples mapped to this sample group description box.
[0241] 3.5.5. Subpicture Entity Group
[0242] 3.5.5.1. Overview
[0243] A Subpicture Entity Group is defined to provide level information indicating the level of conformance of the merged bitstream among several VVC Subpicture tracks.
[0244] NOTE: The VVC Base track provides another mechanism for merging VVC Subpicture tracks.
[0245] The implicit reconstruction process requires modification of the parameter sets. The Subpicture Entity Group gives guidance to simplify the generation of parameter sets for reconstructing the bitstream.
[0246] When the coded subpictures within a group to be jointly decoded are interchangeable (i.e., the player selects multiple active tracks from the sample-wise subpicture groups with the same level contribution), the SubpicCommonGroupBox indicates the combination rule and the level_idc of the resulting combination when jointly decoded.
[0247] When coded subpictures of different properties (e.g., different resolutions) are selected to be jointly decoded, the SubpicMultipleGroupsBox indicates the combination rule and the level_idc of the resulting combination when jointly decoded.
[0248] All entity_id values included in a Subpicture Entity Group shall identify VVC Subpicture tracks. The SubpicCommonGroupBox and SubpicMultipleGroupsBox, when present, shall be included in the GroupsListBox in the movie-level MetaBox and shall not be included in the file-level or track-level MetaBox.
[0249] 3.5.5.2. Syntax of SubpicCommonGroupBox
[0250] aligned(8) class SubpicCommonGroupBox extends EntityToGroupBox('acgl', 0, 0)
[0251] {
[0252] unsigned int(32) level_idc;
[0253] unsigned int(32) num_active_tracks;
[0254] }
[0255] 3.5.5.3. Semantics of the subpicture common group box
[0256] level_idc specifies the level to which any selection of num_active_tracks entities within the entity group conforms.
[0257] num_active_tracks specifies the number of tracks for which the value of level_idc is provided.
[0258] 3.5.5.4. Syntax of the subpicture multiple group box
[0259]
[0260]
[0261] 3.5.5.5. Semantics
[0262] level_idc specifies the level to which any combination of num_active_tracks[i] tracks selected within the subgroup with ID equal to i for all values of i in the range of 0 to num_subgroup_ids - 1, inclusive, conforms.
[0263] num_subgroup_ids specifies the number of separate subgroups, each identified by the same value of track_subgroup_id[i]. Different subgroups are identified by different values of track_subgroup_id[i].
[0264] track_subgroup_id[i] specifies the subgroup ID of the i-th track in this entity group. The subgroup ID value shall be in the range of 0 to num_subgroup_ids - 1, inclusive.
[0265] num_active_tracks[i] specifies the number of tracks in the sub-group with ID equal to i that are recorded in level_idc.
[0266] 4. Example technical problems solved by the disclosed technical solutions
[0267] The latest design of VVC video file format on carrying sub-pictures in VVC bitstream of multiple tracks has the following problems:
[0268] 1) The samples of a VVC sub-picture track contain any of the following: A) one or more complete sub-pictures that are consecutive in decoding order as specified in ISO / IEC 23090-3; B) one or more complete slices that form a rectangular region and are consecutive in decoding order as specified in ISO / IEC 23090-3.
[0269] However, there are the following problems:
[0270] a. It would make more sense to also require that a VVC sub-picture track, when containing sub-pictures, covers a rectangular region, similar to when it contains slices.
[0271] b. It would make more sense to require that the sub-pictures or slices in a VVC sub-picture track are motion-constrained (i.e., extractable or self-contained).
[0272] c. Why not allow a VVC sub-picture track to contain a set of sub-pictures that form a rectangular region but are not consecutive in decoding order in the original bitstream, but these sub-pictures are consecutive in decoding order if the track itself is decoded? For example, should this not be allowed for a field of view (FOV) of 360-degree video that is covered by some sub-pictures for the left and right boundaries of a projected picture?
[0273] 2) The order of non-VCL NAL units in a sample of a VVC base track is not explicitly specified when a PH NAL unit is not present in the sample when reconstructing a PU from the sample and the temporally aligned samples in the VVC sub-picture track list referenced by the VVC base track.
[0274] 3) The sub-picture order sample group mechanism (“spor”) enables different sub-picture orders from the sub-picture tracks in the reconstructed bitstream for different samples and enables cases where SPS and / or PPS rewriting is needed. However, it is not clear why either of these two flexibilities is needed. Therefore, the “spor” sample group mechanism is not needed and can be removed.
[0275] 4) When reconstructing a PU from a sample in the VVC base track and a time-aligned sample in the VVC subpicture track list referred by the VVC base track, all VPS, DCI, SPS, PPS, AUD, PH, EOS and EOB NAL units (if any) are excluded when NAL units in the time-aligned sample of the VVC subpicture track are added to the PU. However, what about OPI NAL units? What about SEI NAL units? Why are these non-VCL NAL units allowed to exist in a subpicture track? If they exist, can we just pass through them in bitstream reconstruction?
[0276] 5) The container of the box of two subpicture entity groups is specified as the movie level MetaBox. However, the entity group's entity id value can refer to track ID only when the box is contained in the file level MetaBox.
[0277] 6) The subpicture entity group is suitable for the case where the related subpicture information is consistent throughout the track duration. However, this is not always the case. For example, what if different CVSs have different levels for a particular subpicture sequence? In this case, sample groups should be used instead to carry basically the same information but allow some information to be different for different samples (e.g., CVS).
[0278] 7) The subpicture order (“spor”) sample group is currently mandatory to exist in each VVC base track. The “spor” sample group mechanism enables different subpicture orders from the subpicture track in the reconstructed bitstream for different samples and enables the case where SPS and / or PPS rewriting is needed. However, in the case of direct “early-binding” subpicture by referring to the ‘subp’ track in the VVC base track, the ‘spor’ sample group is not needed.
[0279] 5. List of technical solutions
[0280] To solve the above problems and other problems, the methods as outlined below are disclosed. The present invention should be considered as an example to explain the general concept and should not be interpreted in a narrow way. Furthermore, these inventions can be applied individually or in any way combined.
[0281] 1) One or more of the following items are proposed on the VVC subpicture track:
[0282] a. Require the VVC subpicture track to cover a rectangular region when containing subpictures.
[0283] b. Require that sub-pictures or slices in a VVC sub-picture track are motion constrained such that they can be extracted, decoded and presented without any sub-pictures or slices covering other areas being present.
[0284] i. Alternatively, allow sub-pictures or slices in a VVC sub-picture track to depend on sub-pictures or slices covering other areas in terms of motion compensation and thus not be extractable, decodable and presentable without any sub-pictures or slices covering other areas being present.
[0285] c. Allow a VVC sub-picture track to contain a set of sub-pictures or slices forming a rectangular area but being non-contiguous in decoding order in the original / entire VVC bitstream.
[0286] This enables, for example, the field of view (FOV) of a 360-degree video covered by sub-pictures on the left and right borders of a projected picture that are non-contiguous in decoding order in the original / entire VVC bitstream to be represented by a VVC sub-picture track.
[0287] d. Require that the order of sub-pictures or slices in each sample of a VVC sub-picture track shall be the same as their order in the original / entire VVC bitstream.
[0288] e. Add an indication whether the decoding order of sub-pictures or slices in each sample of a VVC sub-picture track is contiguous in the original / entire VVC bitstream.
[0289] i. For example, the indication is signaled in the VVC base track sample entry description or otherwise.
[0290] ii. Require that when there is no indication that the order of sub-pictures or slices in each sample of a VVC sub-picture track is contiguous in decoding order in the original / entire VVC bitstream, the sub-pictures or slices in the track shall not be merged with sub-pictures or slices in other VVC sub-picture tracks. For example, in this case, it is not allowed to reference a VVC-based track by a "subp" type track reference, including this VVC sub-picture track and another VVC sub-picture track.
[0291] f. Add a flag naluslnContiguousDecodingOrderFlag to VvcNALUConfigBox. The flag equal to 1 indicates that the NAL units in each sample are contiguous in decoding order in the original entire bitstream, so the VVC base track that references this VVC subpicture track by a track reference of "subp" type can also reference other VVC subpicture tracks by the same track reference. Value 0 indicates that the NAL units in each sample can or can not be contiguous in decoding order in the original entire bitstream, so the VVC base track that references this VVC subpicture track by a track reference of "subp" type can not reference other VVC subpicture tracks by the same track reference.
[0292] 2) When reconstructing PUs in time-aligned samples in the VVC base track and the list of VVC subpicture tracks that the VVC base track references by track references, the order of non-VCL NAL units in the samples of the VVC base track is explicitly specified regardless of whether there are PH NAL units in the samples.
[0293] a. In one example, the set of NAL units from the samples of the VVC base track that are to be placed in the PUs before the NAL units in the VVC subpicture track is specified as follows: if there are at least one NAL unit in the sample with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 (a NAL unit with such a NAL unit type cannot be before the first VCL NAL unit in a picture unit), then the NAL units in the sample up to and not including the first NAL unit of these NAL units, otherwise all the NAL units in the sample.
[0294] b. In one example, the set of NAL units from the samples of the VVC base track that are to be placed in the PUs after the NAL units in the VVC subpicture track is specified as follows: all the NAL units in the sample with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31.
[0295] 3) Allow VVC track to refer to multiple (sub-picture) tracks by using "subp" track reference, and the order of reference indicates the decoding order of sub-pictures in the bitstream reconstructed according to the referred VVC sub-picture tracks.
[0296] a. When reconstructing a PU from samples of the VVC base track and time-aligned samples in the VVC sub-picture track list referred to by the VVC base track, samples of the sub-picture tracks are processed in the order that the VVC sub-picture tracks are referred to in the "subp" track reference.
[0297] 4) Do not allow any AU-level or picture-level non-VCL NAL units in a sub-picture track, including AUD, DCI, OPI, VPS, SPS, PPS, PH, EOS and EOB NAL units, and SEI NAL units containing only AU-level and picture-level SEI messages. AU-level SEI messages apply to one or more entire AUs. Picture-level SEI messages apply to one or more entire pictures.
[0298] a. Furthermore, when reconstructing a PU from samples of the VVC base track and time-aligned samples in the VVC sub-picture track list referred to by the VVC base track, all NAL units in the time-aligned samples of the VVC sub-picture tracks are added to the PU without discarding certain non-VCL NAL units.
[0299] 5) When reconstructing a PU from samples of the VVC base track and time-aligned samples in the VVC sub-picture track list referred to by the VVC base track through track reference, remove the use of "spor" sample groups, and remove the description of parameter set override process based on "spor" sample groups.
[0300] 6) Remove the specification of "spor" sample groups.
[0301] 7) Specify that each "subp" track reference index shall refer to the track ID of a VVC sub-picture track or the track group ID of a VVC sub-picture track group, and not otherwise.
[0302] 8) To address issue 5, the container of the two sub-picture entity groups' boxes is specified as follows: When present, SubpicCommonGroupBox and SubpicMultipleGroupsBox shall be contained in GroupsListBox in the file level MetaBox, and shall not be contained in MetaBox of other levels.
[0303] 9) To address issue 6, two sample groups are added to carry similar information as those carried by the two subpicture entity groups, so that the VVC file format will support the case of inconsistent related subpicture information throughout the duration of a track, e.g., when different CVSs have different levels for a particular subpicture sequence.
[0304] 10) To address issue 7, one or more of the following items are proposed:
[0305] a. The “spor” sample group is designated as optional for each VVC base track.
[0306] b. When the “spor” sample group is not present in a VVC base track, the samples of the referenced subpicture track are processed in the order in which the VVC subpicture track is referenced in the “subp” track reference when reconstructing a PU.
[0307] 6. Embodiments
[0308] Below are some example embodiments of the invention aspects outlined in Section 5 above, which can be applied to the standard specification of the VVC video file format. The changed text is based on the latest draft of the MPEG output document N19454 (“Information technology — Coding of audio-visual objects — Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format — Amendment 2: Carriage of VVC and EVC in ISOBMFF”, July 2020). Most of the relevant parts that have been added or modified are highlighted in bold and italic, and some of the parts that have been deleted are marked with double brackets (e.g., [[a]] means the deleted character “a”). There can be other editorial nature changes that are not highlighted.
[0309] 6.1. First embodiment
[0310] This embodiment is for items 1a, 1b and 1c.
[0311] 6.1.1. Types of tracks
[0312] This specification defines the following types of video tracks for carrying VVC bitstreams:
[0313] a) VVC tracks:
[0314] A VVC track represents a VVC bitstream by including NAL units in its samples and / or sample entries, and possibly by associating other VVC tracks containing other layers and / or sub-layers of the VVC bitstream via "vopi" and "linf" sample groups or via "opeg" entity groups, and possibly by referencing a VVC sub-picture track.
[0315] When a VVC track references a VVC sub-picture track, it is also called a VVC base track. A VVC base track shall not contain VCL NAL units, and shall not be referenced by a VVC track via a "vvcN" track reference.
[0316] b) VVC non-VCL track:
[0317] A VVC non-VCL track is a track containing only non-VCL NAL units, and is referenced by a VVC track via a "vvcN" track reference.
[0318] A VVC non-VCL track can contain an APS carrying ALF, LMCS or scaling list parameters, with or without other non-VCL NAL units stored in and signaled through a track separate from the track containing VCL NAL units.
[0319] A VVC non-VCL track can also contain picture header NAL units, with or without APS NAL units, and with or without other non-VCL NAL units stored in and signaled through a track separate from the track containing VCL NAL units.
[0320] c) VVC sub-picture track:
[0321] A VVC sub-picture track contains any of the following:
[0322] A sequence of one or more VVC sub-pictures forming a rectangular region.
[0323] A sequence of one or more complete slices forming a rectangular region.
[0324] A sample of a VVC sub-picture track contains any of the following:
[0325] One or more complete sub-pictures forming an extractable rectangular region [[that are consecutive in decoding order] as specified in ISO / IEC 23090-3, where an extractable rectangular region is a rectangular region for which decoding does not use any pixel values outside the region for motion compensation.
[0326] One or more complete slices forming an extractable rectangular region [[that are consecutive in decoding order] as specified in the ISO / IEC 23090-3 standard.
[0327] [[The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order.]]
[0328] NOTE: The VVC non-VCL tracks and VVC sub-picture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, a client can request a DASH representation containing a subset of VVC sub-picture tracks and a DASH representation containing non-VCL tracks on a segment-by-segment basis. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided, and unnecessary transmission of sub-pictures can be avoided.
[0329] 6.1.2. Overview of rectangular regions carried in VVC bitstreams
[0330] This document provides support for describing rectangular regions including any of the following:
[0331] - a sequence of one or more VVC sub-pictures forming a rectangular region, [[which are consecutive in decoding order]], or
[0332] - a sequence of one or more complete slices forming a rectangular region [[the region and are consecutive in decoding order]].
[0333] A rectangular region covers a rectangle without holes. Rectangular regions within a picture do not overlap each other. ...
[0335] 6.2. Second embodiment
[0336] This embodiment is used for items 2, 2a, 2b, 3, 3a, 4, 4a and 5.
[0337] 6.2.1. Reconstructing a picture unit from samples in a VVC track referring to a VVC sub-picture track
[0338] A sample of a VVC track is parsed as a picture unit containing the following NAL units in the order of the bullets:
[0339] • AUD NAL unit, [[if any,]] when present in the sample [[(and is the first NAL unit in the sample)]].
[0340] NOTE 1: When an AUD NAL unit is present in the sample, it is the first NAL unit in the sample.
[0341] • The parameter sets and SEI NAL units contained in the sample entry, if any, when the sample is the first sample of a sample sequence associated with the same sample entry.
[0342] • [[The NAL units present in the sample up to and including the PH NAL unit]] if at least one NAL unit with nal unit type equal to EOS NUT, EOB NUT, SUFFIX APS NUT, SUFFIX SEI NUT, FD NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 (a NAL unit with such a nal unit type cannot precede the first VCL NAL unit in a picture unit) is present in the sample, the NAL units in the sample up to the first of these NAL units and not including the first NAL unit, otherwise all NAL units in the sample.
[0343] • The content of the (temporally) time-aligned decoded samples of each referenced VVC subpicture track in the order specified in the "subp" track reference in the VVC subpicture track in the "spor" sample group description entry mapped to the sample [[excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units, if any]] in the order specified in the "subp" track reference in the VVC subpicture track in the "spor" sample group description entry mapped to the sample. The track reference is parsed as specified below.
[0344] NOTE 2: If the referenced VVC subpicture track is associated with a VVC non-VCL track, the decoded samples of the VVC subpicture track contain the non-VCL NAL unit(s) of the time-aligned samples in the VVC non-VCL track, if any.
[0345] • [[The NAL units in the sample following the PH NAL unit]] all NAL units in the sample with nal unit type equal to EOS NUT, EOB NUT, SUFFIX APS NUT, SUFFIX SEI NUT, FD NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31.
[0346] [[NOTE 2: The NAL units in the sample following the PH NAL unit can include suffix SEI NAL units, suffix APS NAL units, EOS NAL units, EOB NAL units, or the reserved NAL units allowed after the last VCL NAL unit.]] The "subp" track reference index of the "spor" sample group description entry is parsed as follows:
[0347] • If the track reference points to a track ID of a VVC subpicture track, the track reference is resolved to the VVC subpicture track.
[0348] • Otherwise (the track reference points to an "alte" track group), the track reference is resolved to any track of the "alte" track group, and if the particular track reference index value is resolved to a particular track in the previous sample, it shall be resolved to any of the following in the current sample:
[0349] • The same particular track, or
[0350] • Any other track in the same "alte" track group that contains a sync sample that is time-aligned with the current sample.
[0351] NOTE 3: VVC subpicture tracks in the same "alte" track group must be independent of any other VVC subpicture tracks referenced by the same VVC base track to avoid decoding mismatches and can therefore be constrained as follows:
[0352] • All VVC subpicture tracks contain VVC subpictures.
[0353] • Subpicture boundaries behave like picture boundaries.
[0354] • Cross-subpicture boundary loop filtering is turned off.
[0355] If the reader selects a VVC subpicture track containing VVC subpictures, where the set of subpicture ID values is either the initial selection or different from the previous selection, the following steps can be taken:
[0356] • The "spor" sample group description entry is investigated to conclude whether a PPS or SPS NAL unit needs to be changed.
[0357] NOTE: It is only possible to change the SPS at the beginning of a CLVS.
[0358] • If the "spor" sample group description entry indicates that a start code emulation prevention byte exists before or within the subpicture ID in the contained NAL unit, the RBSP is derived from the NAL unit (i.e., the start code emulation prevention byte is removed). After rewriting in the next step, start code emulation prevention is re-applied.
[0359] • The reader uses the bit position and subpicture ID length information in the "spor" sample group entry to conclude which bits are overwritten to update the subpicture ID to the selected subpicture ID.
[0360] • When the subpicture ID value of a PPS or SPS is changed from the value of the previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value (respectively), the reader needs to include a copy of the previous PPS and SPS (if the PPS or SPS with the same PPS or SPS ID value, respectively, does not otherwise exist in the access unit), and overwrite the PPS or SPS (respectively) in the reconstructed access unit with the updated subpicture ID value.
[0361] • When the subpicture ID value of a PPS or SPS is changed from the value of the previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value (respectively), the reader needs to include a copy of the previous PPS and SPS (if the PPS or SPS with the same PPS or SPS ID value, respectively, does not otherwise exist in the access unit), and overwrite the PPS or SPS (respectively) in the reconstructed access unit with the updated subpicture ID value.
[0362] 6.3. Third embodiment
[0363] This embodiment is used for items 1a, 1b, 1c, 1f, 2, 2a, 2b, 4, 4a, 10.
[0364] Types of tracks:
[0365] This specification specifies the following types of video tracks for carrying VVC bitstreams:
[0366] d) VVC track:
[0367] A VVC track represents a VVC bitstream by including NAL units in its samples and / or sample entries, and possibly by associating other VVC tracks containing other layers and / or sub-layers of the VVC bitstream via "vopi" and "linf" sample groups or via "opeg" entity groups, and possibly by referencing a VVC subpicture track.
[0368] When a VVC track references a VVC subpicture track, it is also called a VVC base track.
[0369] A VVC base track shall not contain VCL NAL units, and shall not be referenced by a VVC track via a "vvcN" track reference.
[0370] e) VVC non-VCL track:
[0371] A VVC non-VCL track is a track containing only non-VCL NAL units, and is referenced by a VVC track via a "vvcN" track reference.
[0372] A VVC non-VCL track can contain APS carrying ALF, LMCS or scaling list parameters, with or without other non-VCL NAL units stored in and signaled via a track separate from the track containing VCL NAL units.
[0373] A VVC non-VCL track can also contain picture header NAL units, with or without APS NAL units and with or without other non-VCL NAL units stored in and transmitted through a track separate from the track containing VCL NAL units.
[0374] f) VVC sub-picture track:
[0375] A VVC sub-picture track contains any of the following:
[0376] A sequence of one or more VVC sub-pictures forming a rectangular region.
[0377] A sequence of one or more complete slices forming a rectangular region.
[0378] A sample of a VVC sub-picture track contains any of the following:
[0379] One or more complete sub-pictures forming an extractable rectangular region that are [[consecutive in decoding order] as specified in ISO / IEC 23090-3, where an extractable rectangular region is a rectangular region for which decoding does not use any pixel values outside the region for motion compensation.
[0380] One or more complete slices forming an extractable rectangular region [[region and are consecutive in decoding order] as specified in the ISO / IEC 23090-3 standard.
[0381] [[The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order.]]
[0382] NOTE: VVC non-VCL tracks and VVC sub-picture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, a client can request a DASH representation containing a subset of VVC sub-picture tracks and a DASH representation containing non-VCL tracks on a segment-by-segment basis. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided, and unnecessary transmission of sub-pictures can be avoided.
[0383] Overview of rectangular regions carried in a VVC bitstream:
[0384] This document provides support for describing rectangular regions including any of the following:
[0385] - a sequence of one or more VVC sub-pictures forming a rectangular region that are [[consecutive in decoding order] or
[0386] - Forming a rectangular region [[a sequence of one or more complete slices that is rectangular and contiguous in decoding order]].
[0387] A rectangular region covers a rectangle without holes. Rectangular regions within a picture do not overlap each other. ...
[0389] Reconstructing a picture unit from samples in a VVC track that reference VVC subpicture tracks:
[0390] Samples of a VVC track are parsed into picture units that contain the following NAL units in the order of items:
[0391] • AUD NAL units, [[if any,]] when present in the sample [[(and are the first NAL units in the sample)]].
[0392] NOTE 1: When an AUD NAL unit is present in a sample, it is the first NAL unit in the sample.
[0393] • Parameter sets and SEI NAL units contained in the sample entry associated with the sample, if any, when the sample is the first sample of a sample sequence associated with the same sample entry.
[0394] • [[The NAL units present in the sample up to and including the PH NAL unit]] If at least one NAL unit with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 (a NAL unit with such a NAL unit type cannot precede the first VCL NAL unit in a picture unit) is present in the sample, the NAL units in the sample up to the first NAL unit of these NAL units and not including the first NAL unit, otherwise all NAL units in the sample.
[0395] • The contents of the temporally aligned samples from each of the referenced VVC subpicture tracks in the order that the VVC subpicture tracks are referenced in the "subp" track reference (when no "spor" sample group is present in the track) or in the order specified in the "spor" sample group description entry to which the sample is mapped [[excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units, if any]]. The track references are parsed as specified below.
[0396] NOTE 2: If the referenced VVC subpicture track is associated with a VVC non-VCL track, the parsed samples of the VVC subpicture track contain the non-VCL NAL unit(s) of the time-congruent samples in the VVC non-VCL track, if any.
[0397] NOTE 3: The above steps indicate that no AU-level or picture-level non-VCL NAL units, including AUD, DCI, OPI, VPS, SPS, PPS, PH, EOS, and EOB NAL units, and SEI NAL units containing only AU-level and picture-level SEI messages, are allowed to be present in a subpicture track.
[0398] • [[the NAL units following the PH NAL unit in the sample]] all NAL units with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 in the sample.
[0399] [[NOTE 2: The NAL units following the PH NAL unit in the sample can include suffix SEI NAL units, suffix APS NAL units, EOS NAL units, EOB NAL units, or the reserved NAL units that are allowed to follow the last VCL NAL unit.]] The "subp" track reference index of the [[“spor” sample group description entry]] is parsed as follows:
[0400] • If the track reference points to a track ID of a VVC subpicture track, the track reference is resolved to the VVC subpicture track.
[0401] • Otherwise (the track reference points to an “alte” track group), the track reference is resolved to any track of the “alte” track group and, if the particular track reference index value is resolved to a particular track in the previous sample, it shall be resolved to either of the following in the current sample:
[0402] • The same particular track, or
[0403] • Any other track in the same “alte” track group that contains a sync sample that is time-congruent to the current sample.
[0404] NOTE 3: VVC subpicture tracks in the same “alte” track group must be independent of any other VVC subpicture tracks referenced by the same VVC base track to avoid decoding mismatches and can therefore be constrained as follows:
[0405] • All VVC subpicture tracks contain VVC subpictures.
[0406] • Subpicture boundaries behave like picture boundaries.
[0407] • [[Cross-subpicture boundary loop filtering is turned off.]]
[0408] If the reader selects a VVC subpicture track containing VVC subpictures, where the set of subpicture ID values is either initially selected or different from the previous selection, the following steps can be taken:
[0409] • The "spor" sample group description entry is investigated to conclude whether a PPS or SPS NAL unit needs to be changed.
[0410] NOTE: Changing the SPS is only possible at the beginning of a CLVS.
[0411] • If the "spor" sample group description entry indicates that a start code emulation prevention byte is present in the subpicture ID before or within the contained NAL unit, the RBSP is derived from the NAL unit (i.e., the start code emulation prevention byte is removed). After rewriting in the next step, start code emulation prevention is re-applied.
[0412] • The reader uses the bit position and subpicture ID length information in the "spor" sample group entry to conclude which bits are overwritten to update the subpicture ID to the selected subpicture ID.
[0413] • When the subpicture ID value of a PPS or SPS is initially selected, the reader needs to overwrite the PPS or SPS, respectively, in the reconstructed access unit with the selected subpicture ID value.
[0414] • When the subpicture ID value of a PPS or SPS changes compared to the previous PPS or SPS, respectively, having the same PPS ID value or SPS ID value, respectively, the reader needs to include a copy of that previous PPS and SPS, if the PPS or SPS with the same PPS or SPS ID value, respectively, does not otherwise exist in the access unit, and overwrite the PPS or SPS, respectively, in the reconstructed access unit with the updated subpicture ID value.
[0415] Sample entry name and format (of the VVC video stream definition):
[0416] Definition: ...
[0418] A VVC track can contain a'subp' track reference whose entry contains a trackJD value of a VVC subpicture track or a track_group_id value of an 'alte' track group of a VVC subpicture track.
[0419] [[When a VVC track contains a'subp' track reference, it is referred to as a VVC base track and the following applies:
[0420] - The samples of a VVC base track shall not contain VCL NAL units.]]
[0421] A sample group of'spor' type as specified in clause 11.7.7 [[SHOULD]] be present in each VVC base track. ...
[0423] Syntax:
[0424] class VvcConfigurationBox extends Box('vvcC'){
[0425] VvcDecoderConfigurationRecord() VvcConfig;
[0426] }
[0427] class VvcNALUConfigBox extends Box('vvcC'){
[0428] unsigned int([[6]]5) reserved = 0;
[0429] unsigned int(1) nalusInContiguousDecodingOrderFlag;
[0430] unsigned int(2) lengthSizeMinusOne;
[0431] }
[0432] class VvcSampleEntry() extends VisualSampleEntry('vvc1' or 'vvc1'){
[0433] VvcConfigurationBox config;
[0434] MPEG4ExtensionDescriptorsBox(); / / optional
[0435] }
[0436] class VvcSubpicSampleEntry() extends VisualSampleEntry('vvs1') {
[0437] VvcNALUConfigBox config;
[0438] }
[0439] Semantics:
[0440] The Compressor name in the base class VisualSampleEntry indicates the name of the compressor used, where the value "\012VVC Coding" (\012 is 10, the string length in bytes) is recommended.
[0441] VvcDecoderConfigurationRecord is defined in 11.3.3.
[0442] nalusInContiguousDecodingOrderFlag equal to 1 indicates that the NAL units in each sample are contiguous in decoding order in the original entire bitstream, so the VVC base track that references this VVC subpicture track through a track reference of type "subp" can also reference other VVC subpicture tracks through the same track reference. A value of 0 indicates that the NAL units in each sample can or can not be contiguous in decoding order in the original entire bitstream, so the VVC base track that references this VVC subpicture track through a track reference of type "subp" can not reference other VVC subpicture tracks through the same track reference.
[0443] lengthSizeMinusOne plus 1 indicates the length in bytes of the NALUnitLength field in the track containing the VvcNALUConfigBox. The value of this field shall be one of 0, 1, or 3, corresponding to lengths coded with 1, 2, or 4 bytes, respectively.
[0444] [[num_subpics_minus1 plus 1 specifies the number of subpicture sequences contained in the VVC subpicture track.
[0445] subpic_id specifies the subpicture identifier of the subpicture sequence contained in the VVC subpicture track.]]
[0446] Figure 1is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 1900. The system 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values) or can be in a compressed or encoded format. The input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi interfaces or cellular interfaces.
[0447] The system 1900 can include a codec component 1904 that can implement various coding or encoding methods described in this document. The codec component 1904 can reduce the average bitrate of video from the input 1902 to the output of the codec component 1904 to produce a coded representation of the video. Thus, the coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 can be stored or transmitted via a connected communication as represented by component 1906. The component 1908 can use the bitstream (or coded) representation of the video received at the input 1902 to generate pixel values or displayable video to be sent to a display interface 1910. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Furthermore, although certain video processing operations are referred to as “coding” operations or tools, it should be understood that the coding tools or operations are used at an encoder and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0448] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB) or High Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0449] Figure 2is a block diagram of a video processing device 3600. The device 3600 can be used to implement one or more of the methods described herein. The device 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in the present document. The memory(ies) 3604 can be used for storing data and code used during operation of the present techniques described herein. The video processing hardware 3606 can be used to implement, in hardware circuitry, some of the techniques described in the present document. In some embodiments, the video processing hardware 3606 can be included at least in part within the processor 3602 (e.g., a graphics co-processor).
[0450] Figure 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0451] As shown in Figure 4 , the video coding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110, which can be referred to as a video decoding device.
[0452] The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0453] The video source 112 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to the destination device 120 by the I / O interface 116 via the network 130a. The encoded video data can also be stored onto a storage medium / server 130b for access by the destination device 120.
[0454] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122.
[0455] I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which be configured to interface with an external display device.
[0456] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[0457] Figure 5 is a block diagram illustrating an example of a video encoder 200 that can be Figure 4 video encoder 114 in system 100 shown.
[0458] Video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 5 examples, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0459] The functional components of video encoder 200 can include partition unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.
[0460] In other examples, video encoder 200 can include more, less, or different functional components. In examples, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0461] Furthermore, some components, such as motion estimation unit 204 and motion compensation unit 205, can be highly integrated, but are represented separately for illustrative purposes. Figure 5 in examples.
[0462] The partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[0463] The mode selection unit 203 can select one of the coding modes (e.g., intra or inter) based on the error results and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-inter prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution of the motion vectors (e.g., sub-pixel or integer pixel precision) for the block.
[0464] To perform inter prediction for a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the buffer 213 to the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples from pictures of the buffer 213 other than the picture associated with the current video block.
[0465] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0466] In some examples, the motion estimation unit 204 can perform single prediction for a current video block, and the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of List 0 or List 1. The motion estimation unit 204 can then generate a reference index indicating the reference picture containing the reference video block in List 0 or List 1 and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0467] In other examples, the motion estimation unit 204 can perform bi-prediction for the current video block, the motion estimation unit 204 can search for a reference video block for the current video block in a reference picture in list 0 and also search for another reference video block for the current video block in a reference picture in list 1. The motion estimation unit 204 can then generate a reference index that indicates the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector that indicates a spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference index and the motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0468] In some examples, the motion estimation unit 204 can output a full set of motion information for the current video for use in decoding processing by the decoder.
[0469] In some examples, the motion estimation unit 204 can not output a full set of motion information for the current video. Instead, the motion estimation unit 204 can reference motion information for another video block to signal the motion information for the current video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information for a neighboring video block.
[0470] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0471] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector for the current video block and the motion vector for the indicated video block. The video decoder 300 can use the motion vector for the indicated video block and the motion vector difference to determine the motion vector for the current video block.
[0472] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of prediction signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0473] The intra prediction unit 206 can perform intra prediction for the current video block. When the intra prediction unit 206 performs intra prediction for the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0474] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0475] In other examples, e.g., in skip mode, there can be no residual data for the current video block, and residual generation unit 207 can not perform the subtraction operation for the current video block.
[0476] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0477] Quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block after transform processing unit 208 generates the transform coefficient video blocks associated with the current video block.
[0478] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0479] Loop filtering operations can be performed to reduce video block artifacts in the video blocks after reconstruction unit 212 reconstructs the video blocks.
[0480] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0481] Figure 6 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 4 the video decoder 114 in the system 100 shown.
[0482] The video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 In examples, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0483] In Figure 6 In the example of FIG. 3, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and a buffer 307. Video decoder 300 may, in some examples, perform a decoding process generally reciprocal to the encoding process described with respect to video encoder 200. Figure 5
[0484] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). Entropy decoding unit 301 can decode the entropy-encoded video data and, from the entropy-decoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. For example, motion compensation unit 302 can determine such information by performing AMVP and Merge modes.
[0485] Motion compensation unit 302 can generate a motion compensated block that can perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax elements.
[0486] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate the interpolation of sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information and use the interpolation filter to generate the prediction block.
[0487] Motion compensation unit 302 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the encoded video sequence, partitioning information describing how to partition each macroblock of the pictures of the encoded video sequence, modes indicating how to encode each partition, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence.
[0488] Intra-prediction unit 303 can use intra-prediction modes, e.g., received in the bitstream, to form a prediction block from spatially neighboring blocks. Inverse quantization unit 303 inverse quantizes, i.e., de-quantizes, quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transformation unit 303 applies an inverse transform.
[0489] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter can also be applied to the decoded block to filter out blocking artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0490] A list of preferred solutions by some embodiments is next provided.
[0491] A first set of methods is provided below. The following solutions show example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0492] 1. A method of visual media processing, comprising: performing a conversion between visual media data and a file storing a bitstream representation of the visual media data according to a format rule; wherein the file comprises a track containing data for a sub-picture of the visual media data; and wherein the format rule specifies a syntax of the track.
[0493] 2. The method of solution 1, wherein the format rule specifies that the track covers a rectangular region.
[0494] 3. The method of solution 1, wherein the format rule specifies that a sub-picture or slice included in the track is separately extractable, decodable, and presentable.
[0495] The following solutions show example embodiments of the techniques discussed in the previous section (e.g., items 3, 4).
[0496] 4. A method of visual media processing, comprising: performing a conversion between visual media data and a file storing a bitstream representation of the visual media data according to a format rule; wherein the file comprises a first track and / or one or more sub-picture tracks; wherein the format rule specifies a syntax of the first track and / or the one or more sub-picture tracks.
[0497] 5. The method of solution 4, wherein the format rule specifies that the first track comprises a reference to the one or more sub-picture tracks.
[0498] 6. The method of solution 4, wherein the format rule does not allow a non- video coding layer network abstraction layer unit at an access unit level or a picture level to be included in the one or more sub-picture tracks.
[0499] 7. The method of solution 6, wherein the disallowed unit comprises a decoding capability information structure, or a parameter set, or operation point information, or a header, or an end of stream, or an end of picture.
[0500] 8. The method according to any one of solutions 1-7, wherein converting comprises generating a bitstream representation of the visual media data according to the format rule and storing the bitstream representation into a file.
[0501] 9. The method according to any one of solutions 1-7, wherein converting comprises parsing the file according to the format rule to recover the visual media data.
[0502] 10. A video decoding apparatus comprising a processor configured to implement a method according to one or more of solutions 1 to 9.
[0503] 11. A video encoding apparatus comprising a processor configured to implement a method according to one or more of solutions 1 to 9.
[0504] 12. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement a method according to any one of solutions 1 to 9.
[0505] 13. A computer readable medium having a bitstream representation thereon conforming to a file format generated according to any one of solutions 1 to 9.
[0506] 14. A method, apparatus or system described in this document.
[0507] The second set of solutions provides example embodiments of the techniques discussed in the previous section (e.g., item 1).
[0508] 1. A method of processing visual media data (e.g., a method 110 as shown in Figure 11 Fig. 1), comprising performing 1102 a conversion between the visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data comprises one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures covers a rectangular region of the one or more pictures.
[0509] 2. The method according to solution 1, wherein the format rule specifies that the one or more sub-pictures or one or more slices comprised in the track are individually extractable, decodable and presentable in the absence of another sub-picture or another slice covering another region different from the rectangular region.
[0510] 3. The method according to clause 1, wherein the format rule specifies that one or more sub-pictures or one or more slices included in the track depend on another sub-picture or another slice covering another region different from the rectangular region in terms of motion compensation.
[0511] 4. The method according to clause 1, wherein the format rule specifies that one or more slices or one or more sub-pictures are allowed to be discontinuous in decoding order with respect to the bitstream stored in the track.
[0512] 5. The method according to clause 1, wherein a field of view of the 360-degree video covered by the one or more sub-pictures discontinuous in decoding order is represented by the track.
[0513] 6. The method according to clause 1, wherein the format rule specifies that an order of the one or more sub-pictures or the one or more slices in each sample of the track is the same as an order of the one or more sub-pictures or the one or more slices in the bitstream stored in the track.
[0514] 7. The method according to clause 1, wherein the format rule further specifies whether to include an indication indicating whether a decoding order of the one or more sub-pictures or the one or more slices in each sample of the track is continuous in the bitstream stored in the track.
[0515] 8. The method according to clause 7, wherein the indication is included in a base track sample entry description of the track.
[0516] 9. The method according to clause 7, wherein the format rule further specifies that, in response to an absence of the indication, the one or more sub-pictures or the one or more slices in the track are not allowed to be merged with another sub-picture or another slice of another track.
[0517] 10. The method according to clause 7, wherein the indication is included in a network abstraction layer (NAL) configuration box.
[0518] 11. The method according to clause 7, wherein the indication equal to 1 indicates that NAL units in each sample of the track are continuous in decoding order of the bitstream and the base track of the track is referenced with a track reference to refer to other tracks with the track reference.
[0519] 12. The method according to clause 7, wherein the indication equal to 0 indicates that NAL units in each sample of the track are allowed or not allowed to be continuous in decoding order of the bitstream and the base track of the track is not allowed to be referenced with a track reference to refer to other tracks with the track reference.
[0520] 13. The method according to any one of solutions 1-12, wherein the visual media data is processed by Versatile Video Coding (VVC) and the one or more tracks are VVC tracks.
[0521] 14. The method according to any one of solutions 1-13, wherein the converting comprises generating the visual media file according to the format rule and storing the one or more bitstreams into the visual media file.
[0522] 15. The method according to any one of solutions 1-13, wherein the converting comprises parsing the visual media file according to the format rule to reconstruct the one or more bitstreams.
[0523] 16. An apparatus for processing visual media data, comprising a processor configured to implement a method comprising: performing a conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures covers a rectangular region of the one or more pictures.
[0524] 17. The apparatus according to solution 16, wherein the format rule specifies whether to include an indication indicating whether a decoding order of the one or more sub-pictures or the one or more slices in each sample of the track is continuous in the bitstream stored in the track.
[0525] 18. A non-transitory computer-readable recording medium storing instructions causing a processor to: perform a conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures covers a rectangular region of the one or more pictures.
[0526] 19. The non-transitory computer-readable recording medium according to solution 18, wherein the format rule specifies whether to include an indication indicating whether a decoding order of the one or more sub-pictures or the one or more slices in each sample of the track is continuous in the bitstream stored in the track.
[0527] 20. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing apparatus, wherein the method comprises generating a visual media file comprising one or more tracks storing one or more bitstreams of visual media data; wherein the visual media data comprises one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more sub-pictures or one or more slices covers a rectangular region of one or more pictures.
[0528] 21. The non-transitory computer-readable recording medium according to clause 20, wherein the format rule specifies whether to include an indication of whether a decoding order of one or more sub-pictures or one or more slices in each sample of a track is consecutive in the bitstream stored in the track.
[0529] 22. A video processing apparatus comprising a processor configured to implement a method recited in any one or more of clauses 1 to 15.
[0530] 23. A method of storing visual media data into a file comprising one or more bitstreams, the method comprising a method recited in any one of clauses 1 to 15, and further comprising storing the bitstream into a non-transitory computer-readable recording medium.
[0531] 24. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of clauses 1 to 15.
[0532] 25. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0533] 26. A video processing apparatus for storing a bitstream, wherein the video processing apparatus is configured to implement a method recited in any one or more of clauses 1 to 15.
[0534] 27. A computer-readable medium having a bitstream thereon conforming to a file format generated according to any one of clauses 1 to 15.
[0535] 28. A method, apparatus or system described in this document.
[0536] The third set of clauses show example embodiments of the techniques discussed in the previous section (e.g., items 3, 5, 6, 7, and 10).
[0537] 1. A method of processing visual media data (e.g., as in Figure 12The method 1200) includes performing 1202 a conversion between visual media data and a visual media file that includes one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file includes a base track referencing one or more sub-picture tracks storing coded information of one or more sub-pictures of the visual media data, and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
[0538] 2. The method of any of solutions 1, wherein the format rule specifies that the base track includes sub-picture track references referencing the one or more sub-picture tracks, and an order of the one or more sub-picture tracks referenced in the sub-picture track references indicates an order of samples of the one or more sub-picture tracks in the video unit reconstructed from the one or more sub-picture tracks.
[0539] 3. The method of any of solutions 1, wherein the format rule further specifies that each sub-picture track reference has an index referring to a track identification of the sub-picture track or a track group identification of a group of sub-picture tracks.
[0540] 4. The method of any of solutions 1, wherein the format rule specifies that a sub-picture order sample group is optional for the base track.
[0541] 5. The method of solution 4, wherein the format rule further specifies that, in the case that the sub-picture order sample group is not present in the base track, the sub-picture track references are used in determining the order of the one or more sub-picture tracks referenced in the base track.
[0542] 6. The method of solution 4, wherein the format rule further specifies that the use of the sub-picture order sample group is removed, and a description of a parameter set override process based on the sub-picture order sample group is removed.
[0543] 7. The method of solution 4, wherein the format rule further specifies that the specification of the sub-picture order sample group is removed.
[0544] 8. The method of any of solutions 1-7, wherein the visual media data is processed by Versatile Video Coding (VVC), and the one or more tracks are VVC tracks.
[0545] 9. The method of any of solutions 1-8, wherein the conversion includes generating the visual media file and storing the one or more bitstreams into the visual media file according to the format rule.
[0546] 10. The method according to any one of solutions 1-8, wherein converting comprises parsing the visual media file according to a format rule to reconstruct the one or more bitstreams.
[0547] 11. An apparatus for processing visual media data, comprising a processor configured to implement a method comprising performing a conversion between visual media data and a visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file comprises a base track referencing one or more sub-picture tracks storing coding information of one or more sub-pictures of the visual media data, and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
[0548] 12. The apparatus according to solution 11, wherein the format rule specifies that the base track comprises a sub-picture track reference referencing the one or more sub-picture tracks, and an order of the one or more sub-picture tracks referenced in the sub-picture track reference indicates an order of samples of the one or more sub-picture tracks in the video unit reconstructed from the one or more sub-picture tracks.
[0549] 13. The apparatus according to solution 11, wherein the format rule further specifies that each sub-picture track reference has an index referring to a track identification of the sub-picture track or a track group identification of a group of sub-picture tracks.
[0550] 14. The apparatus according to solution 11, wherein the format rule specifies that a sub-picture order sample group is optional for the base track.
[0551] 15. The apparatus according to solution 14, wherein the format rule further specifies that, in case the sub-picture order sample group is not present in the base track, the sub-picture track reference is used in determining the order of the one or more sub-picture tracks referenced in the base track.
[0552] 16. The apparatus according to solution 14, wherein the format rule further specifies removing the use of the sub-picture order sample group and removing a description of a parameter set override process based on the sub-picture order sample group.
[0553] 17. The apparatus according to solution 14, wherein the format rule further specifies removing the specification of the sub-picture order sample group.
[0554] 18. A non-transitory computer-readable recording medium storing instructions causing a processor to perform a conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of visual media data according to a format rule; wherein the visual media file comprises a base track referencing one or more sub-picture tracks, the sub-picture track storing coding information of one or more sub-pictures of the visual media data, and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
[0555] 19. The non-transitory computer-readable recording medium according to clause 18, wherein the format rule specifies that the base track comprises sub-picture track references referencing the one or more sub-picture tracks, and an order of the one or more sub-picture tracks referenced in the sub-picture track references indicates an order of samples of the one or more sub-picture tracks in the video unit reconstructed from the one or more sub-picture tracks.
[0556] 20. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing device, wherein the method comprises generating a visual media file comprising one or more tracks storing one or more bitstreams of visual media data according to a format rule; wherein the visual media file comprises a base track referencing one or more sub-picture tracks, the sub-picture track storing coding information of one or more sub-pictures of the visual media data, and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
[0557] 21. A video processing device comprising a processor configured to implement a method recited in any one or more of clauses 1 to 10.
[0558] 22. A method of storing visual media data into a file comprising one or more bitstreams, the method comprising a method recited in any one or more of clauses 1 to 10, and further comprising storing the bitstream into a non-transitory computer-readable recording medium.
[0559] 23. A computer-readable medium storing program code that, when executed, causes a processor to implement a method recited in any one or more of clauses 1 to 10.
[0560] 24. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0561] 25. A video processing device for storing a bitstream, wherein the video processing device is configured to implement a method recited in any one or more of clauses 1 to 10.
[0562] 26. A computer readable medium having a bitstream thereon representing a file format generated according to any of the schemes 1 to 10.
[0563] 27. A method, apparatus or system described in this document.
[0564] In example schemes, the visual media data corresponds to video or images. In schemes described herein, an encoder can conform to a format rule by generating a coded representation according to the format rule. In schemes described herein, a decoder can parse syntax elements in a coded representation using a format rule, with knowledge of the presence and absence of syntax elements according to the format rule, to generate decoded video. In the above schemes, the visual media data corresponds to video or images.
[0565] In this document, the term “video processing” can refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm can be applied during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. A bitstream representation of a current video block, as defined by the syntax, can for example correspond to bits co-located within a bitstream or distributed at different locations within a bitstream. For example, a macroblock can be encoded according to transformed and coded error residual values and also using bits in a header and other fields in the bitstream. Furthermore, during conversion, a decoder can parse a bitstream with knowledge that some fields can be present or absent, as described in the above schemes, based on the determination. Similarly, an encoder can determine whether certain syntax fields are included and generate a coded representation accordingly by including or excluding syntax fields from the coded representation.
[0566] The disclosed and other aspects, examples, implementations, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter effecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus.
[0567] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0568] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0569] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical, or optical disks, or tape. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0570] Although the present patent document contains many details, these should not be construed as limiting the subject matter or the scope of the claims to only these specific embodiments. Some of the features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0571] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to achieve desirable results. In addition, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0572] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method of processing visual media data, comprising: performing a conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to a format rule; wherein the visual media file comprises a base track referencing one or more sub-picture tracks, the sub-picture tracks storing coding information of one or more sub-pictures of the visual media data, and wherein the format rule specifies a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track, wherein the format rule further specifies that the base track comprises sub-picture track references referencing the one or more sub-picture tracks, and an order of the one or more sub-picture tracks referenced in the sub-picture track references indicates an order of samples of the one or more sub-picture tracks in a video unit reconstructed from the one or more sub-picture tracks, wherein the format rule further specifies an order of non-video coding layer, VCL, network abstraction layer, NAL, units in samples of the base track when reconstructing a video unit from samples of the base track and samples in the one or more sub-picture tracks, regardless of whether a picture header NAL unit is present in the samples, wherein the format rule specifies that at least some non-VCL NAL units in samples of the base track are placed in the video unit before or after NAL units in the one or more sub-picture tracks.
2. The method of claim 1, wherein, the format rule further specifies that each sub-picture track reference has an index referring to a track identification of a sub-picture track or a track group identification of a group of sub-picture tracks.
3. The method of claim 1, wherein, the format rule further specifies that a sub-picture order sample group is optional for the base track.
4. The method of claim 3, wherein, the format rule further specifies that, in a case where the sub-picture order sample group is not present in the base track, the sub-picture track references are used in determining the order of the one or more sub-picture tracks referenced in the base track.
5. The method of claim 3, wherein, the format rule further specifies that a description of a parameter set rewriting process based on the sub-picture order sample group is removed from the visual media file.
6. The method of claim 3, wherein, the format rule further specifies that a specification of the sub-picture order sample group is removed from the visual media file.
7. The method of any one of claims 1-6, wherein, the visual media data is processed by Versatile Video Coding, VVC, and the one or more tracks are VVC tracks.
8. The method of any one of claims 1-6, wherein, the conversion comprises generating the visual media file and storing the one or more bitstreams into the visual media file according to the format rule.
9. The method of any one of claims 1-6, wherein, the conversion comprises parsing the visual media file to reconstruct the one or more bitstreams according to the format rule.
10. A video processing apparatus comprising a processor configured to implement a method recited in any of claims 1 to 9.
11. A method of storing visual media data into a file comprising one or more bitstreams, the method comprising the method according to any one of claims 1 to 9, and further comprising storing the one or more bitstreams into a non-transitory computer readable recording medium.
12. A computer readable medium storing program code which, when executed, causes a processor to implement the method according to any one of claims 1 to 9.
13. A video processing apparatus for storing a bitstream, wherein, The video processing apparatus is configured to implement the method according to any one of claims 1 to 9.