Sub-picture track in encoded / decoded video
By dividing the video pictures into sub-picture tracks, the bandwidth and delay problems of VVC video bitstream in 360-degree video streaming are solved, and the rapid responses related to efficient codec and viewport are achieved, improving codec efficiency and resource utilization.
Patent Information
- Application Number
- CN202111092816.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-06
- Filing Date
- 2021-09-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-09-17
AI Technical Summary
When existing video encoding and decoding technologies deal with multifunctional video encoding and decoding (VVC) video bitstreams, it is difficult to effectively use sub-picture tracks for efficient encoding and transmission, especially in 360-degree video streaming and area of interest (ROI) applications, resulting in unreasonable bandwidth usage and increased latency.
Using the concept of sub-picture track, the video picture is divided into multiple sub-pictures. Each sub-picture is composed of one or more rectangular strips, allowing independent encoding, decoding and extraction, and supporting mixed NAL unit types in the picture, ensuring the independence and merging of sub-pictures through signaling mechanisms, and supporting viewport-related 360-degree video streaming.
It realizes the reduction of bandwidth requirements and end-to-end latency in 360-degree video streaming, improves encoding and decoding efficiency, and supports fast response and efficient resource utilization when viewport changes.
Smart Images

Figure CN114205607B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is filed to timely claim priority to and the benefit of U.S. Provisional Patent Application No. 63 / 079,933, filed on September 17, 2020, and U.S. Provisional Patent Application No. 63 / 088,126, filed on October 6, 2020, under applicable patent laws and / or under the rules of the Paris Convention. The entire disclosures of the foregoing applications are incorporated herein by reference as a part of the disclosure of this application for all purposes under law. Technical Field
[0003] This patent document relates to the generation, storage, and consumption of digital audio-visual media information in file format. Background Art
[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of networked user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to process coded representations of video or images according to file formats.
[0006] In one example aspect, a method for processing visual media data is disclosed. The method includes performing conversion between the visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; wherein the visual media file stores the one or more tracks according to a format rule; and wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures comprises a rectangular area covering the one or more pictures.
[0007] In another example aspect, a method for processing visual media data is disclosed. The method includes performing conversion between the visual media data and a visual media file, the visual media file including one or more tracks storing one or more bitstreams of the visual media data according to format rules; wherein the visual media file includes a base track that references one or more sub-picture tracks, the sub-picture tracks storing codec information for one or more sub-pictures of the visual media data; and wherein the format rules specify a process for reconstructing a video unit from samples in the one or more sub-picture tracks and the base track.
[0008] In yet another exemplary aspect, a video processing device is disclosed, wherein the video processing device includes a processor configured to implement the above method.
[0009] In yet another exemplary aspect, a method of storing visual media data in a file comprising one or more bitstreams is disclosed. The method corresponds to the above method and further includes storing the one or more bitstreams in a non-transitory computer-readable recording medium.
[0010] In yet another exemplary aspect, a computer-readable medium storing a bitstream is disclosed, wherein the bitstream is generated according to the above method.
[0011] In yet another exemplary aspect, a video processing device storing a bitstream is disclosed, wherein the video processing device is configured to implement the above method.
[0012] In yet another example aspect, a computer-readable medium is disclosed, on which a bitstream conforms to a file format generated according to the above method.
[0013] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram of an example video processing system.
[0015] Figure 2 It is a block diagram of a video processing device.
[0016] Figure 3 is a flow chart of an example method of video processing.
[0017] Figure 4 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.
[0018] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0019] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.
[0020] Figure 7 An example of an encoder block diagram is shown.
[0021] Figure 8 A picture is shown partitioned into 18 slices, 24 slices, and 24 sub-pictures.
[0022] Figure 9 A typical sub-picture-based viewport-dependent 360-degree video delivery scheme is shown.
[0023] Figure 10An example of extracting one sub-picture from a bitstream containing two sub-pictures and four slices is shown.
[0024] Figure 11 and Figure 12 Example methods for processing visual media data based on some implementations of the disclosed technology are shown. DETAILED DESCRIPTION
[0025] Section headings are used in this document to facilitate understanding and do not limit the applicability of the techniques and embodiments disclosed in each section to only that section. In addition, the use of H.266 terminology in some descriptions is intended solely to facilitate understanding and is not intended to limit the scope of the disclosed techniques. As such, the techniques described herein are also applicable to other video codec protocols and designs. In this document, editorial changes to text relative to the current draft of the VVC specification or the ISOBMFF file format specification are shown with strikethrough indicating deleted text and highlighting indicating added text (including bold italics).
[0026] 1. Preliminary Discussion
[0027] This document is related to video file formats. Specifically, it relates to the carriage of sub-pictures of Versatile Video Codec (VVC) video bitstreams in multiple tracks in a media file based on the ISO Base Media File Format (ISOBMFF). These concepts can be applied individually or in various combinations to video bitstreams encoded or decoded by any codec (e.g., the VVC standard), and to any video file format (e.g., the VVC video file format under development).
[0028] 2. Abbreviation
[0029] ACT adaptive color transform adaptive color transform
[0030] ALF adaptive loop filter adaptive loop filter
[0031] AMVR adaptive motion vector resolution
[0032] APS adaptation parameter set
[0033] AU access unit access unit
[0034] AUD access unit delimiter access unit delimiter
[0035] AVC advanced video coding (Rec.ITU-T H.264|ISO / IEC14496-10)
[0036] B bi-predictive
[0037] BCW bi-prediction with CU-level weights
[0038] BDOF bi-directional optical flow
[0039] BDPCM block-based delta pulse code modulation
[0040] BP buffering period
[0041] CABAC context-based adaptive binary arithmetic coding
[0042] CB coding block
[0043] CBR constant bit rate
[0044] CCALF cross-component adaptive loop filter CPBcoded picture buffer
[0045] CRA clean random access
[0046] CRC cyclic redundancy check
[0047] CTB coding tree block
[0048] CTU coding tree unit
[0049] CU coding unit
[0050] CVS coded video sequence
[0051] DPB decoded picture buffer decoded picture buffer
[0052] DCI decoding capability information
[0053] DRAP dependent random access point
[0054] DU decoding unit
[0055] DUI decoding unit information
[0056] EG exponential-Golomb index-Columbus
[0057] EGk k-th order exponential-Golomb k-th order exponential-Golomb
[0058] EOB end of bitstream
[0059] EOS end of sequence
[0060] FD filler data filter data
[0061] FIFO first-in, first-out
[0062] FL fixed-length fixed length
[0063] GBR green,blue,and red
[0064] GCI general constraints information
[0065] GDR gradual decoding refresh
[0066] GPM geometric partitioning mode
[0067] HEVC high efficiency video coding (Rec.ITU-T H.265|ISO / IEC 23008-2)
[0068] HRD hypothetical reference decoder
[0069] HSS hypothetical stream scheduler
[0070] I intraframe
[0071] IBC intra block copy
[0072] IDR instantaneous decoding refresh
[0073] ILRP inter-layer reference picture
[0074] IRAP intra random access point intra frame random access point
[0075] LFNST low frequency non-separable transform low frequency non-separable transform
[0076] LPS least probable symbol
[0077] LSB least significant bit
[0078] LTRP long-term reference picture long-term reference picture
[0079] LMCS luma mapping with chroma scaling
[0080] MIP matrix-based intra prediction matrix-based intra prediction
[0081] MPS most probable symbol
[0082] MSB most significant bit
[0083] MTS multiple transform selection
[0084] MVP motion vector prediction motion vector prediction
[0085] NAL network abstraction layer
[0086] OLS output layer set output layer set
[0087] OP operation point
[0088] OPI operating point information
[0089] P predictive
[0090] PH picture header
[0091] POC picture order count picture order count
[0092] PPS picture parameter set
[0093] PROF prediction refinement with optical flow
[0094] PT picture timing
[0095] PU picture unit
[0096] QP quantization parameter quantization parameter
[0097] RADL random access decodable leading (picture)
[0098] RASL random access skipped leading (picture) Random access skipped leading (picture)
[0099] RBSP raw byte sequence payload
[0100] RGB red,green,and blue
[0101] RPL reference picture list
[0102] SAO sample adaptive offset
[0103] SAR sample aspect ratio
[0104] SEI supplemental enhancement information
[0105] SH slice header
[0106] SLI subpicture level information
[0107] SODB string of data bits
[0108] SPS sequence parameter set
[0109] STRP short-term reference picture
[0110] STSA step-wise temporal sublayer access
[0111] TR truncated rice
[0112] VBR variable bit rate
[0113] VCL video coding layer
[0114] VPS video parameter set video parameter set
[0115] VSEI versatile supplemental enhancement information (Rec.ITU-T H.274|ISO / IEC 23002-7)
[0116] VUI video usability information
[0117] VVC versatile video coding (Rec.ITU-T H.266|ISO / IEC23090-3)
[0118] 3. Video Codec Introduction
[0119] 3.1. Video Codec Standards
[0120] Video codec standards have evolved primarily through the well-known ITU-T and ISO / IEC developments. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially launched, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a new codec standard that aims to reduce bit rate by 50% compared to HEVC. The standard was finalized by JVET at its 19th meeting, which ended on July 1, 2020.
[0121] The Versatile Video Codec (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, compositing and merging content from multiple coded video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360-degree immersive media.
[0122] 3.2. File format standards
[0123] Media streaming applications are typically based on IP, TCP, and HTTP transport methods, and often rely on file formats such as the ISO base media file format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In order to use video formats with ISOBMFF and DASH, file format specifications specific to the video format will be required, such as the AVC file format and the HEVC file format in ISO / IEC 14496-15 (“Information technology—Coding of audio-visual objects—Part 15: Carriage of network abstraction layer (NAL) unitstructured video in the ISO base media file format”) for encapsulating video content in ISOBMFF tracks and in DASH representations and fragments. Important information about the video bitstream (e.g., profile, tier, and level) and many other information will need to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, e.g., selection of appropriate media segments for both initialization at the start of a streaming session and stream adaptation during a streaming session.
[0124] Similarly, for using a picture format with ISOBMFF, a file format specification specific to the image format would be required, such as the AVC image file format and the HEVC image file format in ISO / IEC 23008-12 (“Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 12: Image File Format”).
[0125] The VVC video file format, a file format for storing VVC video content based on ISOBMFF, is currently under development by MPEG. The latest draft specification for the VVC video file format is included in MPEG output document N19454 ("Information technology—Coding of audio-visual objects—Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media file format—Amendment 2: Carriage of VVC and EVC in ISOBMFF," July 2020).
[0126] The VVC image file format, a file format based on ISOBMFF for storing image content encoded and decoded using VVC, is currently under development by MPEG. The latest draft specification for the VVC image file format is included in MPEG output document N19460 ("Information technology—High efficiency coding and media delivery inheterogeneous environments—Part 12: Image File Format—Amendment 3: Support for VVC, EVC, slideshows and other improvements," July 2020).
[0127] 3.3. Image Segmentation Scheme in HEVC
[0128] HEVC includes four different picture segmentation schemes, namely, regular slice, dependent slice, tile, and Wavefront Parallel Processing (WPP), which can be used for Maximum Transfer Unit (MTU) size matching, parallel processing, and reduced end-to-end latency.
[0129] Regular slices are similar to those in H.264 / AVC. Each regular slice is encapsulated in its own NAL unit, and intra-picture prediction (intra-sample prediction, motion information prediction, codec mode prediction) and entropy codec dependencies are disabled across slice boundaries. Therefore, regular slices can be reconstructed independently of other regular slices in the same picture (although mutual dependencies may still exist due to loop filtering operations).
[0130] Regular strips are the only tool that can be used for parallelization, and they are also available in H.264 / AVC in almost the same form. Parallelization based on regular strips does not require much inter-processor or inter-core communication (except for inter-processor or inter-core data sharing for motion compensation when decoding predictively coded pictures, which is generally much more onerous than inter-processor or inter-core data sharing due to intra-picture prediction). However, for the same reasons, the use of regular strips may result in a large amount of codec overhead due to the bit cost of the slice header and the lack of prediction across slice boundaries. In addition, regular strips (in contrast to the other tools mentioned below) are also used as a key mechanism for bitstream segmentation to match MTU size requirements because of the intra-picture independence of regular strips and the fact that each regular strip is encapsulated in its own NAL unit. In many cases, the goals of parallelization and MTU size matching place conflicting demands on the stripe layout in the picture. The realization of this situation led to the development of the parallelization tools mentioned below.
[0131] Dependent slices have short slice headers and allow the bitstream to be split at treeblock boundaries without breaking any intra-picture prediction. Essentially, dependent slices cut a regular slice into multiple NAL units to provide reduced end-to-end latency by allowing part of a regular slice to be sent before coding of the entire regular slice is complete.
[0132] In WPP, a picture is partitioned into a single row of coding tree blocks (CTBs). Entropy decoding and prediction are allowed to use data from CTBs in other partitions. Parallel processing is possible by decoding CTB rows in parallel, where the start of decoding a CTB row is delayed by two CTBs, ensuring that data associated with CTBs above and to the right of the subject CTB is available before decoding the subject CTB. Using this staggered start (which, when represented graphically, looks like a wavefront), parallelization is possible using as many processors / cores as the picture contains CTB rows. Because intra-picture prediction is permitted between adjacent treeblock rows within a picture, the inter-processor / inter-core communication required to implement intra-picture prediction can be extensive. WPP partitioning does not generate additional NAL units compared to when WPP partitioning is not applied, so WPP is not a tool for MTU size matching. However, if MTU size matching is required, regular striping can be used with WPP, but with some codec overhead.
[0133] Slices define the horizontal and vertical boundaries that divide an image into slice columns and slice rows. Slice columns extend from the top to the bottom of the image. Similarly, slice rows extend from the left to the right. The number of slices in an image can be simply derived by multiplying the number of slice columns by the number of slice rows.
[0134] Before decoding the top left CTB of the next slice in the order of the slice raster scan of the picture, the scanning order of the CTBs is changed to local within the slice (in the order of the CTB raster scan of the slice). Similar to regular slices, slices destroy intra-picture prediction dependencies and entropy decoding dependencies. However, they do not need to be included into separate NAL units (same as WPP in this respect); therefore, slices cannot be used for MTU size matching. Each slice can be processed by one processor / core, and the inter-processor / inter-core communication required for intra-picture prediction between processing units decoding adjacent slices is limited to conveying a shared slice header and loop filtering related to sharing of reconstructed samples and metadata when a slice spans more than one slice. When more than one slice or WPP fragment is included in a slice, the entry point byte offset of each slice or WPP fragment except the first slice or WPP fragment in the slice is signaled in the slice header.
[0135] For simplicity, HEVC specifies constraints for four different picture partitioning schemes. A given coded video sequence cannot contain both slices and wavefronts for most profiles specified by HEVC. For each slice and slice, one or both of the following conditions must be met: 1) all coded treeblocks in a slice belong to the same slice; 2) all coded treeblocks in a slice belong to the same slice. Finally, a wavefront segment contains exactly one CTB row, and when WPP is in use, if a slice starts within a CTB row, it must end within the same CTB row.
[0136] The latest amendments to HEVC are specified in the following JCT-VC output document: JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Skupin, G.J. Sullivan, A. Tourapis, Y.-K. Wang (eds.), "HEVC Additional Supplemental Enhancement Information (Draft 4)", October 24, 2017, publicly available here: http: / / phenix.int-evry.fr / jct / doc_end_user / documents / 29_Macau / wg11 / JCTVC-AC1005-v2.zip. Including this amendment, HEVC specifies three MCTS-related SEI messages: the temporal MCTS SEI message, the MCTS extraction information set SEI message, and the MCTS extraction information nesting SEI message.
[0137] The temporal MCTS SEI message indicates the presence of MCTS in the bitstream and signals the MCTS. For each MCTS, the motion vector is constrained to point to the full-sample location inside the MCTS and the fractional-sample location that only requires the full-sample location inside the MCTS for interpolation, and the motion vector candidate is not allowed to be used for temporal motion vector prediction derived from blocks outside the MCTS. In this way, each MCTS can be decoded independently without the existence of slices that are not included in the MCTS.
[0138] The MCS Extraction Information Set SEI message provides supplementary information that can be used in MCTS sub-bitstream extraction (specified as part of the semantics of the SEI message) to generate a conforming bitstream for the MCTS set. This information includes multiple extraction information sets, each of which defines multiple MCTS sets and contains RBSP bytes that replace the VPS, SPS, and PPS used during the MCTS sub-bitstream extraction process. When extracting a sub-bitstream according to the MCTS sub-bitstream extraction process, the parameter sets (VPS, SPS, and PPS) need to be rewritten or replaced, and the slice header needs to be slightly updated because one or all of the syntax elements related to the slice address (including first_slice_segment_in_pic_flag and slice_segment_address) usually need to have different values.
[0139] 3.4. Image Segmentation and Sub-Images in VVC
[0140] 3.4.1. Image Segmentation in VVC
[0141] In VVC, a picture is divided into one or more slice rows and one or more slice columns. A slice is a sequence of CTUs covering a rectangular area of the picture. The CTUs in a slice are scanned in raster scan order within the slice.
[0142] A slice includes an integer number of complete slices or an integer number of consecutive complete CTU rows within a slice of a picture.
[0143] Two modes of slices are supported: raster scan slice mode and rectangular slice mode. In raster scan slice mode, a slice contains a sequence of complete slices from a slice raster scan of a picture. In rectangular slice mode, a slice contains multiple complete slices that together form a rectangular area of the picture, or multiple consecutive complete CTU rows that together form a slice of a rectangular area of the picture. Slices within a rectangular slice are scanned in slice raster scan order within the rectangular area corresponding to the slice.
[0144] A sub-picture consists of one or more strips that together cover a rectangular area of the picture.
[0145] 3.4.2. Sub-image concept and function
[0146] In VVC, each sub-image consists of one or more complete rectangular strips that together cover the rectangular area of the image, such as Figure 8As shown. Sub-pictures can be designated as extractable (i.e., coded and decoded independently of other sub-pictures of the same picture and independent of earlier pictures in decoding order) or non-extractable. Regardless of whether a sub-picture is extractable, the encoder can control whether in-loop filtering (including deblocking, SAO, and ALF) is applied to each sub-picture individually across sub-picture boundaries.
[0147] Functionally, sub-pictures are similar to motion-constrained tilesets (MCTS) in HEVC. They both allow independent encoding and decoding and the extraction of rectangular subsets of a coded picture sequence for use cases such as viewport-dependent 360-degree video streaming optimization and region of interest (ROI) applications.
[0148] In streaming 360-degree video (also called omnidirectional video), at any particular moment, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) will be presented to the user, while the user can change the viewing orientation and thus the current viewport at any time by turning his / her head. A high-quality representation of the omnidirectional video is only required for the current viewport presented to the user at any given moment, although it is desirable to have at least some lower-quality representation of the area not covered by the current viewport at the client, and ready to be presented to the user in case the user suddenly changes his / her viewing orientation to anywhere on the sphere. Dividing the high-quality representation of the entire omnidirectional video into sub-pictures at an appropriate granularity enables such Figure 8 Optimization shown, where on the left hand side are the 12 high resolution sub-pictures and on the right hand side are the remaining 12 lower resolution sub-pictures of the omnidirectional video.
[0149] Figure 9 Another typical sub-picture-based viewport-dependent 360-degree video delivery scheme is shown, in which only the higher-resolution representation of the full video includes sub-pictures, while the lower-resolution representation of the full video does not use sub-pictures and can be encoded and decoded with less frequent RAPs than the higher-resolution representation. The client receives the lower-resolution full video, while for the higher-resolution video, the client only receives and decodes the sub-picture covering the current viewport.
[0150] 3.4.3. Differences between sub-images and MCT
[0151] There are several important design differences between sub-pictures and MCT. First, the sub-picture feature in VVC allows the motion vectors of a codec block to point outside the sub-picture, even in this case the sub-picture is extractable by applying sample padding at the sub-picture boundary, similar to at the picture boundary. Second, additional changes are introduced to the selection and derivation of motion vectors in VVC's Merge mode and decoder-side motion vector refinement process. This allows for higher codec efficiency compared to the non-standard motion constraints applied on the encoder side of MCTS. Third, when extracting one or more extractable sub-pictures from a picture sequence to create a sub-bitstream as a conforming bitstream, the SH (and PH NAL unit, when present) does not need to be rewritten. In HEVC MCTS-based sub-bitstream extraction, the SH needs to be rewritten. Note that in both HEVC MCTS extraction and VVC sub-picture extraction, the SPS and PPS need to be rewritten. However, typically there are only a few parameter sets in the bitstream, and each picture has at least one slice, so rewriting the SH can be a significant burden for the application system. Fourth, slices of different sub-pictures within a picture are allowed to have different NAL unit types. This feature is often referred to as mixed NAL unit types or mixed sub-picture types within a picture and is discussed in detail below. Fifth, VVC specifies HRD and level definitions for sub-picture sequences, so the encoder can ensure the consistency of the sub-bitstream for each extractable sub-picture sequence.
[0152] 3.4.4. Mixing sub-image types within an image
[0153] In AVC and HEVC, all VCL NAL units in a picture need to have the same NAL unit type. VVC introduces the option of mixing sub-pictures with certain different VCL NAL unit types within a picture, thus providing support for random access not only at the picture level but also at the sub-picture level. In VVC, VCL NAL units within a sub-picture still need to have the same NAL unit type.
[0154] The ability to randomly access sub-pictures from IRAPs is beneficial for 360-degree video applications. Figure 9 In a similar viewport-dependent 360-degree video delivery scheme as shown, the contents of spatially adjacent viewports largely overlap, i.e., during a viewport orientation change, only a fraction of sub-pictures in the viewport are replaced by new sub-pictures, while most sub-pictures remain in the viewport. The sequence of sub-pictures newly introduced into the viewport must start with an IRAP slice, but when the remaining sub-pictures are allowed to perform inter-frame prediction during viewport changes, a significant reduction in the overall transmission bitrate can be achieved.
[0155] An indication of whether a picture contains only a single type of NAL unit or more than one type is provided in the PPS to which it refers (i.e., using a flag named pps_mixed_nalu_types_in_pic_flag). A picture can contain both sub-pictures containing IRAP slices and sub-pictures containing trailing slices. Some other combinations of different NAL unit types within a picture are also allowed, including leading picture slices of NAL unit types RASL and RADL, which allow sub-picture sequences with open GOP and closed GOP codec structures extracted from different bitstreams to be merged into a single bitstream.
[0156] 3.4.5. Sub-image layout and ID signaling
[0157] The layout of sub-pictures in VVC is signaled in the SPS and therefore remains unchanged within the CLVS. Each sub-picture is signaled by the position of its top-left CTU and its width and height (in units of CTUs), ensuring that the sub-picture covers a rectangular area of the picture at CTU granularity. The order in which the sub-pictures are signaled in the SPS determines the index of each sub-picture within the picture.
[0158] In order to be able to extract and merge sub-picture sequences without rewriting SH or PH, the slice addressing scheme in VVC is based on sub-picture ID and sub-picture specific slice index to associate slices with sub-pictures. In SH, the sub-picture ID and sub-picture level slice index of the sub-picture containing the slice are signaled. Note that the value of the sub-picture ID of a particular sub-picture can be different from the value of its sub-picture index. The mapping between the two is signaled in the SPS or PPS (but not both) or implicitly inferred. When present, the sub-picture ID mapping needs to be rewritten or added when the SPS and PPS are rewritten during the sub-picture sub-bitstream extraction process. The sub-picture ID and sub-picture level slice index together indicate to the decoder the exact position of the first decoded CTU of the slice within the DPB time slot of the decoded picture. After sub-bitstream extraction, the sub-picture ID of the sub-picture remains unchanged, while the sub-picture index may change. Even when the raster scan CTU address of the first CTU in a slice in a sub-picture has changed compared to the value in the original bitstream, the unchanged values of the sub-picture ID and sub-picture level slice index in the corresponding SH will still correctly determine the position of each CTU in the decoded picture of the extracted sub-bitstream. Figure 10 An example including two sub-pictures and four slices is used to illustrate the use of sub-picture ID, sub-picture index and sub-picture level slice index to achieve sub-picture extraction.
[0159] Similar to sub-picture extraction, sub-picture signaling allows merging several sub-pictures from different bitstreams into a single bitstream simply by rewriting the SPS and PPS, provided that the different bitstreams are generated collaboratively (e.g., using different sub-picture IDs but otherwise mostly aligned SPS, PPS, and PH parameters such as CTU size, chroma format, codec tools, etc.).
[0160] Although sub-pictures and slices are signaled independently in the SPS and PPS, respectively, there are inherent mutual constraints between the sub-picture and slice layouts in order to form a consistent bitstream. First, the presence of sub-pictures requires the use of rectangular slices and prohibits raster scan slices. Second, the slices of a given sub-picture should be consecutive NAL units in decoding order, which means that the sub-picture layout constrains the order of the coded slice NAL units in the bitstream.
[0161] 3.5. Some details of the VVC video file format
[0162] 3.5.1. Track types
[0163] The VVC video file format specifies the following types of video tracks for carrying VVC bitstreams in ISOBMFF files:
[0164] a) VVC track:
[0165] A VVC track represents a VVC bitstream by including NAL units in its samples and sample entries, and possibly by referencing other VVC tracks containing other sub-layers of the VVC bitstream, and possibly by referencing a VVC sub-picture track. When a VVC track references a VVC sub-picture track, it is called a VVC base track.
[0166] b) VVC non-VCL track:
[0167] APSs carrying ALF, LMCS, or scaling list parameters, and other non-VCL NAL units may be stored in and sent over a track separate from the track containing the VCL NAL units; this is the VVC non-VCL track.
[0168] c) VVC sub-picture track:
[0169] A VVC sub-picture track contains any of the following:
[0170] A sequence of one or more VVC sub-pictures.
[0171] A sequence of one or more complete strips forming a rectangular area.
[0172] A sample of a VVC sub-picture track includes any of the following:
[0173] One or more complete sub-pictures, consecutive in decoding order, as specified in ISO / IEC 23090-3.
[0174] One or more complete slices, as specified in ISO / IEC 23090-3, forming a rectangular area and contiguous in decoding order.
[0175] The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order.
[0176] Note: VVC non-VCL tracks and VVC sub-picture tracks enable optimal delivery of VVC video in streaming applications, as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, the client can request a DASH representation containing a subset of VVC sub-picture tracks along with a DASH representation containing non-VCL tracks on a fragment-by-fragment basis. This way, redundant transmission of APS and other non-VCL NAL units can be avoided.
[0177] 3.5.2. Overview of rectangular regions carried in VVC bitstreams
[0178] This document provides support for describing rectangular regions that include any of the following:
[0179] - a sequence of one or more VVC sub-pictures that are consecutive in decoding order, or
[0180] - A sequence of one or more complete strips forming a rectangular area and consecutive in decoding order.
[0181] Rectangular regions cover rectangles without holes. Rectangular regions within the image do not overlap each other.
[0182] A rectangular region may be described by a rectangular region visual sample group description entry (ie, an instance of RectangularRegionGroupEntry), where rect_region_flag is equal to 1.
[0183] If each sample of the track consists of only one NAL unit of a rectangular region, a SampleToGroupBox of type "trif" can be used to associate the sample with the rectangular region, but if the default sample grouping mechanism is used (i.e., when the version of the SampleGroupDescriptionBox of type "trif" is equal to or greater than 2), the SampleToGroupBox of type "trif" can be omitted. Otherwise, the sample, NAL unit, and rectangular region are associated with each other through a SampleToGroupBox of type "nalm" and a grouping_type_parameter equal to "trif" and a SampleGroupDescriptionBox of type "nalm". The RectangularRegionGroupEntry describes:
[0184] - rectangular area,
[0185] - The codec correlation between this rectangular area and other rectangular areas.
[0186] Each RectangularRegionGroupEntry is assigned a unique identifier, called groupID. This identifier can be used to associate NAL units in a sample with a specific RectangularRegionGroupEntry.
[0187] The location and size of the rectangular region are identified using the luminance sample coordinates.
[0188] When used with movie fragments, a RectangularRegionGroupEntry can be defined for the duration of a movie fragment by defining a new SampleGroupDescriptionBox in the TrackFragmentBox as defined in clause 8.9.4 of ISO / IEC 14496-12. However, there should not be any RectangularRegionGroupEntry in the track fragment with the same groupID as an already defined RectangularRegionGroupEntry.
[0189] The base region used in a RectangularRegionGroupEntry is the picture to which the NAL units in the rectangular region associated with the rectangular region group entry belong.
[0190] If there is any change in the base region size in consecutive samples (eg in case of reference picture resampling (RPR) or SPS resizing), the samples should be associated with different RectangularRegionGroupEntry reflecting their respective base region sizes.
[0191] NAL units mapped to rectangular regions can be carried in the VVC track as usual, or in a separate track called the VVC sub-picture track.
[0192] 3.5.3. Reconstructing a picture unit from samples in the VVC track of the reference VVC sub-picture track
[0193] The samples of a VVC track are parsed into access units, which contain the following NAL units in bulleted order:
[0194] • The AUD NAL unit, if any, when present in the sample (and when it is the first NAL unit in the sample).
[0195] When the sample is the first sample in a sequence of samples associated with the same sample entry, the parameter set and SEI NAL unit contained in that sample entry, if any.
[0196] • The NAL units present in the sample up to and including the PH NAL unit.
[0197] The contents of the time-aligned (in decoding time) parsed samples from each referenced VVC sub-picture track, in the order specified in the "spor" sample group description entry that maps to this sample, excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units, if any. Track references are parsed as per the following specification.
[0198] NOTE 1: If the referenced VVC sub-picture track is associated with a VVC non-VCL track, the parsed samples of the VVC sub-picture track contain the non-VCL NAL unit(s) of the temporally aligned samples in the VVC non-VCL track, if any.
[0199] The NAL unit following the PH NAL unit in the sample.
[0200] NOTE 2: NAL units following a NAL unit in a sample may include a suffix SEI NAL unit, a suffix APS NAL unit, an EOS NAL unit, an EOB NAL unit, or a reserved NAL unit that is allowed to follow the last VCL NAL unit.
[0201] The "subp" track reference index of the "spor" sample group description entry is parsed as follows:
[0202] If the track reference points to the track ID of a VVC sub-picture track, the track reference is resolved to a VVC sub-picture track.
[0203] Otherwise (the track reference points to the "alte" track group), the track reference is resolved to any track in the "alte" track group. If a specific track reference index value was resolved to a specific track in the previous sample, it should be resolved to any of the following in the current sample:
[0204] the same specific track, or
[0205] Any other tracks in the same "alte" track group that contain Sync samples that are time-aligned with the current sample.
[0206] NOTE 3: VVC sub-picture tracks in the same "alte" track group must be independent of any other VVC sub-picture tracks referenced by the same VVC base track to avoid decoding mismatches, and can therefore be constrained as follows:
[0207] All VVC sub-picture tracks contain VVC sub-pictures.
[0208] Sub-picture boundaries are like picture boundaries.
[0209] Turn off loop filtering across sub-picture boundaries.
[0210] If the reader selects a VVC sub-picture track containing VVC sub-pictures, where the set of sub-picture ID values is the initial selection or different from a previous selection, the following steps may be taken:
[0211] Study the "spor" sample group description entry to conclude whether changes to the PPS or SPS NAL units are required.
[0212] Note: Changing the SPS is only possible when CLVS is started.
[0213] If the "spor" sample group description entry indicates that a start code emulation prevention byte is present before or within the sub-picture ID in the containing NAL unit, the RBSP is derived from the NAL unit (i.e., the start code emulation prevention byte is removed). After rewriting in the next step, start code emulation prevention is re-performed.
[0214] • The reader uses the bit position and sub-picture ID length information in the "spor" sample group entry to figure out which bits are overwritten to update the sub-picture ID to the selected sub-picture ID.
[0215] • When the sub-picture ID value of a PPS or SPS is initially selected, the reader needs to rewrite the PPS or SPS, respectively, with the selected sub-picture ID value in the reconstructed access unit.
[0216] When the sub-picture ID value of a PPS or SPS changes compared to a previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value, the reader needs to include a copy of the previous PPS and SPS (if a PPS or SPS (respectively) with the same PPS or SPS ID value is not otherwise present in the access unit) and rewrite the PPS or SPS (respectively) with the updated sub-picture ID value in the reconstructed access unit.
[0217] 3.5.4. Sub-image order sample group
[0218] 3.5.4.1. Definition
[0219] This sample group is used in the VVC base track, that is, in a VVC track with a "subp" track that references a VVC sub-picture track. Each sample group description entry indicates a sub-picture or slice of a coded picture in decoding order, where each index of a track reference of type "subp" indicates one or more sub-pictures or slices that are consecutive in decoding order.
[0220] To facilitate PPS or SPS rewriting in response to sub-picture selection, each sample group description entry can contain:
[0221] - an indication of whether the selected sub-picture ID should be changed in the PPS or SPS NAL unit;
[0222] - length of the sub-picture ID syntax element (in bits);
[0223] - the bit position of the sub-picture ID syntax element in the included RBSP;
[0224] - A flag indicating whether the start code emulation prevention byte exists before or within the sub-picture ID;
[0225] - Parameter set ID of the parameter set containing the sub-picture ID.
[0226] 3.5.4.2. Syntax
[0227]
[0228] 3.5.4.3. Semantics
[0229] subpic_id_info_flag equal to 0 specifies that the sub-picture ID values provided in the SPS and / or PPS are correct for the indicated set of subp_track_ref_idx values, and therefore the SPS or PPS does not need to be rewritten. subpic_info_flag equal to 1 specifies that the SPS and / or PPS may need to be rewritten to indicate the sub-pictures corresponding to the set of subp_track_ref_idx values.
[0230] num_subpic_ref_idx specifies the number of reference indices of the sub-picture track or track group of sub-picture tracks referenced by the VVC track.
[0231] subp_track_ref_idx, for each value of i, specifies the “subp” track reference index of the i-th list of one or more sub-pictures or slices to be included in the VVC bitstream reconstructed according to the VVC track.
[0232] subpic_id_len_minus1 plus 1 specifies the number of bits in the sub-picture identifier syntax element in the PPS or SPS (whichever the structure refers to).
[0233] subpic_id_bit_pos specifies the bit position, starting from 0, of the first bit of the first sub-picture ID syntax element in the referenced PPS or SPS RBSP.
[0234] start_code_emul_flag equal to 0 specifies that the start code emulation prevention byte is not present before or within the sub-picture ID of the referenced PPS or SPS NAL unit.
[0235] start_code_emul_flag equal to 1 specifies that a start code emulation prevention byte may be present before or within the sub-picture ID of the referenced PPS or SPS NAL unit.
[0236] pps_subpic_id_flag, when equal to 0, specifies that the PPS NAL units applicable to the samples mapped to this sample group description entry do not contain a sub-picture ID syntax element.
[0237] pps_subpic_id_flag, when equal to 1, specifies that the PPS NAL units applicable to the samples mapped to this sample group description entry contain a sub-picture ID syntax element.
[0238] pps_id, when present, specifies the PPSID of the PPS that applies to the sample mapped to this sample group description entry.
[0239] sps_subpic_id_flag, when present and equal to 0, specifies that the SPS NAL units applicable to the samples mapped to this sample group description entry do not contain a sub-picture ID syntax element and that the sub-picture ID value is inferred. sps_subpic_id_flag, when present and equal to 1, specifies that the SPS NAL units applicable to the samples mapped to this sample group description entry contain a sub-picture ID syntax element.
[0240] sps_id, when present, specifies the SPSID of the SPS that applies to the sample mapped to this sample group description entry.
[0241] 3.5.5. Sub-image entity group
[0242] 3.5.5.1. Overview
[0243] A sub-picture entity group is defined to provide level information indicating the consistency of the merged bitstream among several VVC sub-picture tracks.
[0244] Note: The VVC base track provides an alternative mechanism for merging VVC sub-picture tracks.
[0245] The implicit reconstruction process requires modification of parameter sets. The sub-picture entity group provides guidance to simplify the generation of parameter sets used to reconstruct the bitstream.
[0246] When the coded sub-pictures to be jointly decoded within a group are interchangeable (ie the player selects multiple active tracks from a per-sample sub-picture group with the same level contribution), SubpicCommonGroupBox indicates the combining rule and the resulting combined level_idc upon joint decoding.
[0247] When there are coded sub-pictures of different attributes (eg, different resolutions) selected for joint decoding, SubpicMultipleGroupsBox indicates a combination rule and level_idc of the combination obtained in joint decoding.
[0248] All entity_id values included in the sub-picture entity group should identify a VVC sub-picture track. When present, SubpicCommonGroupBox and SubpicMultipleGroupsBox should be included in the GroupsListBox in the movie-level MetaBox, and should not be included in the file-level or track-level MetaBox.
[0249] 3.5.5.2. Syntax of the Sub-Image Common Group Box
[0250]
[0251] 3.5.5.3. Semantics of the Sub-Image Common Group Box
[0252] level_idc specifies the level within the entity group to which any selection of num_active_tracks entities qualifies.
[0253] num_active_tracks specifies the number of tracks that provide the value of level_idc.
[0254] 3.5.5.4. Syntax of sub-image multi-group boxes
[0255]
[0256]
[0257] 3.5.5.5. Semantics
[0258] level_idc specifies the level to which any combination of num_active_tracks[i] tracks in the subgroup with ID equal to i is selected for all values of i in the range 0 to num_subgroup_ids-1, inclusive.
[0259] num_subgroup_ids specifies the number of separate subgroups, each identified by the same value of track_subgroup_id[i]. Different subgroups are identified by different values of track_subgroup_id[i].
[0260] track_subgroup_id[i] specifies the subgroup ID of the i-th track in this entity group. The subgroup ID value should range from 0 to num_subgroup_ids–1 (inclusive).
[0261] num_active_tracks[i] specifies the number of tracks in the subgroup with ID equal to i recorded in level_idc.
[0262] 4. Example technical problems solved by the disclosed technical solution
[0263] The latest design of the VVC video file format for carrying sub-pictures in VVC bitstreams of multiple tracks has the following problems:
[0264] 1) The samples of a VVC sub-picture track contain any of the following: A) one or more complete sub-pictures specified in ISO / IEC 23090-3 that are consecutive in decoding order; B) one or more complete slices specified in ISO / IEC 23090-3 that form a rectangular area and are consecutive in decoding order.
[0265] However, there are the following problems:
[0266] a. It would make more sense to also require the VVC sub-picture track to cover a rectangular area when it contains sub-pictures, similarly to when it contains strips.
[0267] b. It would make more sense to require sub-pictures or slices in a VVC sub-picture track to be motion constrained (ie, extractable or self-contained).
[0268] c. Why isn't it allowed for a VVC sub-picture track to contain a set of sub-pictures that form a rectangular area but are not contiguous in decoding order in the original bitstream, but are contiguous in decoding order if the track itself is decoded? For example, shouldn't this be allowed for the field of view (FOV) of a 360-degree video to be covered by some sub-pictures at the left and right borders of a projected picture?
[0269] 2) When reconstructing a PU from samples of the VVC base track and time-aligned samples in a VVC sub-picture track list referenced by the VVC base track, the order of non-VCL NAL units in samples of the VVC base track is not explicitly specified when the PH NAL unit is not present in the sample.
[0270] 3) The sub-picture order sample group mechanism ("spor") enables different sub-picture orders from the sub-picture track in the reconstructed bitstream for different samples and enables situations where SPS and / or PPS rewriting is required. However, it is not clear why either of these two flexibilities is needed. Therefore, the "spor" sample group mechanism is not needed and can be removed.
[0271] 4) When reconstructing a PU from samples of the VVC base track and the time-aligned samples in the VVC sub-picture track list referenced by the VVC base track, all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units (if any) are excluded when the NAL units in the time-aligned samples of the VVC sub-picture track are added to the PU. However, what about OPI NAL units? What about SEI NAL units? Why are these non-VCL NAL units allowed to exist in the sub-picture track? If they exist, can they just be passed through in the bitstream reconstruction?
[0272] 5) The container of the boxes of the two sub-picture entity groups is specified as a movie-level MetaBox. However, the entity_id value of the entity group can refer to the track ID only when the box is contained in a file-level MetaBox.
[0273] 6) Sub-picture entity groups are suitable for situations where the relevant sub-picture information remains consistent throughout the duration of the track. However, this is not always the case. For example, what if different CVSs have different levels for a particular sub-picture sequence? In this case, sample groups should be used instead to carry essentially the same information, but allow some information (e.g., CVS) to be different for different samples.
[0274] 7) Currently, it is mandatory that sub-picture order ("spor") sample groups exist in each VVC base track. The "spor" sample group mechanism enables different sub-picture orders from the sub-picture track in the reconstructed bitstream for different samples, and enables cases where SPS and / or PPS rewriting is required. However, in the case of direct "early-binding" sub-picture references via the 'subp' track in the VVC base track, the 'spor' sample group is not required.
[0275] 5. Technical solution list
[0276] To solve the above problems and other problems, the following methods are disclosed. The present invention should be regarded as an example to explain the general concept and should not be interpreted narrowly. In addition, these inventions can be applied alone or in combination in any way.
[0277] 1) One or more of the following items are proposed on the VVC sub-picture track:
[0278] a. VVC sub-picture tracks are required to cover a rectangular area when containing sub-pictures.
[0279] b. Sub-pictures or slices in a VVC sub-picture track are required to be motion constrained so that they can be extracted, decoded, and presented in the absence of any sub-pictures or slices covering other areas.
[0280] i. Alternatively, sub-pictures or slices in a VVC sub-picture track are allowed to depend on sub-pictures or slices covering other areas for motion compensation, and therefore cannot be extracted, decoded, and presented in the absence of any sub-pictures or slices covering other areas.
[0281] c. A VVC sub-picture track is allowed to contain a set of sub-pictures or slices that form a rectangular area but are not consecutive in decoding order in the original / whole VVC bitstream.
[0282] This enables the field of view (FOV) of a 360-degree video covered by sub-pictures that are not consecutive in decoding order in the original / whole VVC bitstream, e.g. on the left and right borders of a projected picture, to be represented by a VVC sub-picture track.
[0283] d. It is required that the order of sub-pictures or slices in each sample of a VVC sub-picture track should be the same as their order in the original / entire VVC bitstream.
[0284] e. Add an indication of whether the decoding order of sub-pictures or slices in each sample of the VVC sub-picture track is continuous in the original / entire VVC bitstream.
[0285] i. For example, the indication is signaled in the VVC base track sample entry description or in other ways.
[0286] ii. Requires that when there is no indication that the order of sub-pictures or slices in each sample of a VVC sub-picture track is continuous in decoding order in the original / whole VVC bitstream, sub-pictures or slices in a track shall not be merged with sub-pictures or slices in other VVC sub-picture tracks. For example, in this case, it is not allowed to refer to a VVC-based track, including this VVC sub-picture track and another VVC sub-picture track, through a track reference of type "subp".
[0287] f. Add the flag nalusInContiguousDecodingOrderFlag to VvcNALUConfigBox. This flag equal to 1 indicates that the NAL units in each sample are continuous in decoding order in the original entire bitstream, so the VVC base track that references this VVC sub-picture track through a track reference of type "subp" may also reference other VVC sub-picture tracks through the same track reference. A value of 0 indicates that the NAL units in each sample may or may not be continuous in decoding order in the original entire bitstream, so the VVC base track that references this VVC sub-picture track through a track reference of type "subp" may not reference other VVC sub-picture tracks through the same track reference.
[0288] 2) When reconstructing a PU from samples of the VVC base track and temporally aligned samples in a VVC sub-picture track list referenced by the VVC base track via a track reference, the order of non-VCL NAL units in samples of the VVC base track is well-specified, regardless of the presence of PH NAL units in the samples.
[0289] a. In one example, the set of NAL units of samples from the VVC base track in the PU to be placed before the NAL unit in the VVC sub-picture track is specified as follows: if there is at least one NAL unit with nal_unit_type equal to EOS_NUT EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 in the sample (NAL units with such NAL unit types must not precede the first VCL NAL unit in a picture unit), then the NAL units in the sample up to and not including the first of these NAL units, otherwise all NAL units in the sample.
[0290] b. In one example, the set of NAL units of samples from the VVC base track to be placed in the PU following the NAL unit in the VVC sub-picture track is specified as follows: all NAL units in the sample with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31.
[0291] 3) A VVC track is allowed to reference multiple (sub-picture) tracks by using a "subp" track reference, and the order of the references indicates the decoding order of the sub-pictures in the bitstream reconstructed from the referenced VVC sub-picture tracks.
[0292] a. When reconstructing a PU from samples of the VVC base track and the time-aligned samples in the VVC sub-picture track list referenced by the VVC base track, the samples of the reference sub-picture track are processed in the order in which the VVC sub-picture tracks are referenced in the "subp" track reference.
[0293] 4) No AU-level or picture-level non-VCL NAL units are allowed in the sub-picture track, including AUD, DCI, OPI, VPS, SPS, PPS, PH, EOS, and EOB NAL units, as well as SEI NAL units that only contain AU-level and picture-level SEI messages. AU-level SEI information applies to one or more entire AUs. Picture-level SEI messages apply to one or more entire pictures.
[0294] a. In addition, when reconstructing a PU from samples of the VVC base track and the time-aligned samples in the VVC sub-picture track list referenced by the VVC base track, all NAL units in the time-aligned samples of the VVC sub-picture track are added to the PU without discarding certain non-VCL NAL units.
[0295] 5) Remove the use of the “spor” sample group when reconstructing a PU from samples of the VVC base track and time-aligned samples in the VVC sub-picture track list referenced by the VVC base track through track references, and remove the description of the parameter set rewriting process based on the “spor” sample group.
[0296] 6) Remove the specification of the "spor" sample group.
[0297] 7) Specifies that each "subp" track reference index shall refer to the track ID of a VVC sub-picture track or the track group ID of a VVC sub-picture track group, and not otherwise.
[0298] 8) To solve problem 5, the containers of the boxes of the two sub-picture entity groups are specified as file-level MetaBox as follows: When present, SubpicCommonGroupBox and SubpicMultipleGroupsBox should be contained in the GroupsListBox in the file-level MetaBox, and should not be contained in MetaBoxes at other levels.
[0299] 9) To address issue 6, two sample groups are added to carry information similar to that carried by the two sub-picture entity groups, so that the VVC file format will support situations where the relevant sub-picture information is inconsistent throughout the duration of the track, for example, when different CVSs have different levels for a particular sub-picture sequence.
[0300] 10) To address issue 7, one or more of the following items were proposed:
[0301] a. The “spor” sample group is designated as optional for each VVC base track.
[0302] b. When reconstructing a PU, when the “spor” sample group does not exist in the VVC base track, the samples of the referenced sub-picture track are processed in the order in which the VVC sub-picture tracks are referenced in the “subp” track reference.
[0303] 6. Examples
[0304] The following are some example embodiments of some of the inventive aspects outlined in Section 5 above, which can be applied to the standard specification of the VVC video file format. The changed text is based on the latest draft specification in MPEG output document N19454 ("Information technology—Coding of audio-visual objects—Part 15: Carriage of network abstraction layer (NAL) unit structured video in the ISO base media fileformat—Amendment 2: Carriage of VVC and EVC in ISOBMFF", July 2020). Most of the relevant parts that have been added or modified are highlighted in bold and italics, and some deleted parts are marked with double brackets (for example, [[a]] means the character "a" is deleted). There may be some other changes of an editorial nature, so they are not highlighted.
[0305] 6.1. First embodiment
[0306] This embodiment is used for items 1a, 1b, and 1c.
[0307] 6.1.1. Track Types
[0308] This specification specifies the following types of video tracks for carrying VVC bitstreams:
[0309] a) VVC track:
[0310] A VVC track represents a VVC bitstream by including NAL units in its samples and / or sample entries, and possibly by associating other VVC tracks of other layers and / or sub-layers containing VVC bitstreams via the "vopi" and "linf" sample groups or via the "opeg" entity group, and possibly by referencing a VVC sub-picture track.
[0311] When a VVC track references a VVC sub-picture track, it is also called a VVC base track. The VVC base track should not contain VCL NAL units and should not be referenced by VVC tracks via the "vvcN" track reference.
[0312] b) VVC non-VCL track:
[0313] A VVC non-VCL track is a track that contains only non-VCL NAL units and is referenced by a VVC track through a "vvcN" track reference.
[0314] A VVC non-VCL track may contain APSs carrying ALF, LMCS, or scaling list parameters, with or without other non-VCL NAL units stored in and sent over a separate track from the track containing the VCL NAL units.
[0315] VVC non-VCL tracks may also contain picture header NAL units, with or without APS NAL units, and with or without other non-VCL NAL units stored in and sent over a separate track from the track containing the VCL NAL units.
[0316] c) VVC sub-picture track:
[0317] A VVC sub-picture track contains any of the following:
[0318] A sequence of one or more VVC sub-pictures forming a rectangular area.
[0319] A sequence of one or more complete strips forming a rectangular area.
[0320] A sample of a VVC sub-picture track contains any of the following:
[0321] One or more complete sub-pictures [[contiguous in decoding order]] forming an extractable rectangular area as specified in ISO / IEC 23090-3, where an extractable rectangular area is a rectangular area whose decoding does not use any pixel values outside the area for motion compensation.
[0322] One or more complete slices forming an extractable rectangular region [[region and continuous in decoding order]] as specified in the ISO / IEC 23090-3 standard.
[0323] [[The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order. ]]
[0324] Note: VVC non-VCL tracks and VVC sub-picture tracks enable optimal delivery of VVC video in streaming applications, as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, the client can request a DASH representation containing a subset of VVC sub-picture tracks along with a DASH representation containing non-VCL tracks on a fragment-by-fragment basis. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided, and unnecessary transmission of sub-pictures can be avoided.
[0325] 6.1.2. Overview of rectangular regions carried in VVC bitstreams
[0326] This document provides support for describing rectangular regions that include any of the following:
[0327] - a sequence of one or more VVC sub-pictures [[contiguous in decoding order]] forming a rectangular region, or
[0328] - A sequence of one or more complete strips forming a rectangular area [[area and contiguous in decoding order]]
[0329] Rectangular regions cover rectangles without holes. Rectangular regions within the image do not overlap each other. ...
[0331] 6.2. Second embodiment
[0332] This embodiment is used for items 2, 2a, 2b, 3, 3a, 4, 4a and 5.
[0333] 6.2.1. Reconstructing a picture unit from samples in a VVC track that references a VVC sub-picture track
[0334] The samples of a VVC track are parsed into [[access units]] picture units containing the following NAL units in bulleted order:
[0335] AUD NAL unit, [[if any,]] when present in the sample [[(and is the first NAL unit in the sample)]].
[0336] NOTE 1: When the AUD NAL unit is present in a sample, it is the first NAL unit in the sample.
[0337] When the sample is the first sample of a sequence of samples associated with the same sample entry, the parameter set and SEI NAL unit contained in that sample entry, if any.
[0338] [[NAL units up to and including the PH NAL unit present in the sample]] If at least one NAL unit with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 is present in the sample (NAL units with such NAL unit types must not precede the first VCL NAL unit in a picture unit), then the NAL units in the sample up to and including the first of these NAL units; otherwise, all NAL units in the sample.
[0339] The contents of the time-aligned (in decoding time) parsed samples from each referenced VVC sub-picture track in the order in which the VVC sub-picture track is referenced in the “subp” track reference [[as specified in the “spor” sample group description entry mapped to that sample]] [[excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units, if any]]. Track references are parsed as per the following specification.
[0340] NOTE 2: If the referenced VVC sub-picture track is associated with a VVC non-VCL track, the parsed samples of the VVC sub-picture track contain the non-VCL NAL unit(s) of the temporally aligned samples in the VVC non-VCL track, if any.
[0341] [[NAL units following the PH NAL unit in the sample]] All NAL units in the sample whose nal_unit_type is equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31.
[0342] [[NOTE 2: The NAL units following the PH NAL unit in a sample may include a suffix SEI NAL unit, a suffix APS NAL unit, an EOS NAL unit, an EOB NAL unit, or a reserved NAL unit that is allowed to follow the last VCL NAL unit. ]] The "subp" track reference index of the ["spor" sample group descriptor entry]] is parsed as follows:
[0343] If the track reference points to the track ID of a VVC sub-picture track, the track reference is resolved to a VVC sub-picture track.
[0344] Otherwise (the track reference points to the "alte" track group), the track reference is resolved to any track of the "alte" track group, and if a particular track reference index value was resolved to a particular track in the previous sample, it shall be resolved in the current sample to any of the following:
[0345] the same specific track, or
[0346] Any other tracks in the same "alte" track group that contain sync samples that are time-aligned with the current sample.
[0347] NOTE 3: A VVC sub-picture track in the same "alte" track group must be independent of any other VVC sub-picture track referenced by the same VVC base track to avoid decoding mismatches and can therefore be constrained as follows:
[0348] All VVC sub-picture tracks contain VVC sub-pictures.
[0349] Sub-picture boundaries are like picture boundaries.
[0350] [[Turn off loop filtering across sub-picture boundaries.
[0351] If the reader selects a VVC sub-picture track containing VVC sub-pictures, where the set of sub-picture ID values is the initial selection or different from a previous selection, the following steps may be taken:
[0352] The “spor” sample group description entry will be investigated to determine whether changes to the PPS or SPS NAL units are required.
[0353] Note: Changing the SPS is only possible when CLVS is started.
[0354] If the "spor" sample group description entry indicates that a start code emulation prevention byte is present before or within the sub-picture ID in the containing NAL unit, the RBSP is derived from the NAL unit (i.e., the start code emulation prevention byte is removed). After rewriting in the next step, start code emulation prevention is re-performed.
[0355] • The reader uses the bit position and sub-picture ID length information in the “spor” sample group entry to figure out which bits are overwritten to update the sub-picture ID to the selected sub-picture ID.
[0356] • When the sub-picture ID value of a PPS or SPS is initially selected, the reader needs to rewrite the PPS or SPS, respectively, with the selected sub-picture ID value in the reconstructed access unit.
[0357] When the sub-picture ID value of a PPS or SPS changes compared to a previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value (respectively), the reader needs to include a copy of the previous PPS and SPS (if a PPS or SPS (respectively) with the same PPS or SPS ID value is not otherwise present in the access unit) and rewrite the PPS or SPS (respectively) with the updated sub-picture ID value in the reconstructed access unit.
[0358] 6.3. Third embodiment
[0359] This embodiment is used for items 1a, 1b, 1c, 1f, 2, 2a, 2b, 4, 4a, and 10.
[0360] Type of track:
[0361] This specification specifies the following types of video tracks for carrying VVC bitstreams:
[0362] d) VVC track:
[0363] A VVC track represents a VVC bitstream by including NAL units in its samples and / or sample entries, and possibly by associating other VVC tracks of other layers and / or sub-layers containing VVC bitstreams via the "vopi" and "linf" sample groups or via the "opeg" entity group, and possibly by referencing a VVC sub-picture track.
[0364] When a VVC track references a VVC sub-picture track, it is also called a VVC base track. The VVC base track should not contain VCL NAL units and should not be referenced by VVC tracks via the "vvcN" track reference.
[0365] e) VVC non-VCL track:
[0366] A VVC non-VCL track is a track that contains only non-VCL NAL units and is referenced by a VVC track through the "vvcN" track reference.
[0367] A VVC non-VCL track may contain APSs carrying ALF, LMCS, or scaling list parameters, with or without other non-VCL NAL units stored in and sent over a separate track from the track containing the VCL NAL units.
[0368] A VVC non-VCL track may also contain picture header NAL units, with or without APS NAL units and with or without other non-VCL NAL units stored in and sent over a separate track from the track containing the VCL NAL units.
[0369] f) VVC sub-picture track:
[0370] A VVC sub-picture track contains any of the following:
[0371] A sequence of one or more VVC sub-pictures forming a rectangular area.
[0372] A sequence of one or more complete strips forming a rectangular area.
[0373] A sample of a VVC sub-picture track includes any of the following:
[0374] One or more complete sub-pictures [[contiguous in decoding order]] forming an extractable rectangular area as specified in ISO / IEC 23090-3, where an extractable rectangular area is a rectangular area whose decoding does not use any pixel values outside the area for motion compensation.
[0375] One or more complete slices forming an extractable rectangular region [[region and continuous in decoding order]] as specified in the ISO / IEC 23090-3 standard.
[0376] [[The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order. ]]
[0377] Note: VVC non-VCL tracks and VVC sub-picture tracks enable optimal delivery of VVC video in streaming applications, as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, the client can request a DASH representation containing a subset of VVC sub-picture tracks along with a DASH representation containing non-VCL tracks on a fragment-by-fragment basis. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided, and unnecessary transmission of sub-pictures can be avoided.
[0378] Overview of rectangular regions carried in the VVC bitstream:
[0379] This document provides support for describing rectangular regions that include any of the following:
[0380] - a sequence of one or more VVC sub-pictures [[contiguous in decoding order]] forming a rectangular region, or
[0381] - A sequence of one or more complete strips forming a rectangular area [[area and contiguous in decoding order]]
[0382] Rectangular regions cover rectangles without holes. Rectangular regions within the image do not overlap each other. ...
[0384] Reconstruct a picture unit from samples in the VVC track of the reference VVC sub-picture track:
[0385] The samples of a VVC track are parsed into [[access units]] picture units containing the following NAL units in bulleted order:
[0386] AUD NAL unit, [[if any,]] when present in the sample [[(and is the first NAL unit in the sample)]].
[0387] NOTE 1: When the AUD NAL unit is present in a sample, it is the first NAL unit in the sample.
[0388] When the sample is the first sample of a sequence of samples associated with the same sample entry, the parameter set and SEI NAL unit contained in that sample entry, if any.
[0389] [[NAL units up to and including the PH NAL unit present in the sample]] If at least one NAL unit with nal_unit_type equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 is present in the sample (NAL units with such NAL unit types must not precede the first VCL NAL unit in a picture unit), then the NAL units in the sample up to and including the first of these NAL units; otherwise, all NAL units in the sample.
[0390] The contents of the (in decoding time) time-aligned parsed samples from each referenced VVC sub-picture track in the order in which the VVC sub-picture track is referenced in the “subp” track reference (when no “spor” sample group is present in the track), or in the order specified in the “spor” sample group description entry mapped to that sample [[excluding all VPS, DCI, SPS, PPS, AUD, PH, EOS, and EOB NAL units, if any]]. The track reference is parsed as per the following specification.
[0391] NOTE 2: If the referenced VVC sub-picture track is associated with a VVC non-VCL track, the parsed samples of the VVC sub-picture track contain the non-VCL NAL unit(s) of the temporally aligned samples in the VVC non-VCL track, if any.
[0392] NOTE 3: The above steps indicate that no AU-level or picture-level non-VCL NAL units are allowed in the sub-picture track, including AUD, DCI, OPI, VPS, SPS, PPS, PH, EOS, and EOB NAL units, as well as SEI NAL units that only contain AU-level and picture-level SEI messages. AU-level SEI information applies to one or more entire AUs. Picture-level SEI messages apply to one or more entire pictures.
[0393] [[NAL units following the PH NAL unit in the sample]] All NAL units in the sample whose nal_unit_type is equal to EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31.
[0394] [[NOTE 2: The NAL units following the PH NAL unit in a sample may include a suffix SEI NAL unit, a suffix APS NAL unit, an EOS NAL unit, an EOB NAL unit, or a reserved NAL unit that is allowed to follow the last VCL NAL unit. ]] The "subp" track reference index of the ["spor" sample group descriptor entry]] is parsed as follows:
[0395] If the track reference points to the track ID of a VVC sub-picture track, the track reference is resolved to a VVC sub-picture track.
[0396] Otherwise (the track reference points to the "alte" track group), the track reference is resolved to any track of the "alte" track group, and if a particular track reference index value was resolved to a particular track in the previous sample, it shall be resolved in the current sample to any of the following:
[0397] the same specific track, or
[0398] Any other tracks in the same "alte" track group that contain sync samples that are time-aligned with the current sample.
[0399] NOTE 3: A VVC sub-picture track in the same "alte" track group must be independent of any other VVC sub-picture track referenced by the same VVC base track to avoid decoding mismatches and can therefore be constrained as follows:
[0400] All VVC sub-picture tracks contain VVC sub-pictures.
[0401] Sub-picture boundaries are like picture boundaries.
[0402] [[Disable loop filtering across sub-picture boundaries.]]
[0403] If the reader selects a VVC sub-picture track containing VVC sub-pictures, where the set of sub-picture ID values is the initial selection or different from a previous selection, the following steps may be taken:
[0404] The “spor” sample group description entry will be investigated to determine whether changes to the PPS or SPS NAL units are required.
[0405] Note: Changing the SPS is only possible when CLVS is started.
[0406] If the "spor" sample group description entry indicates that a start code emulation prevention byte is present before or within the sub-picture ID in the containing NAL unit, the RBSP is derived from the NAL unit (i.e., the start code emulation prevention byte is removed). After rewriting in the next step, start code emulation prevention is re-performed.
[0407] • The reader uses the bit position and sub-picture ID length information in the “spor” sample group entry to figure out which bits are overwritten to update the sub-picture ID to the selected sub-picture ID.
[0408] • When the sub-picture ID value of a PPS or SPS is initially selected, the reader needs to rewrite the PPS or SPS, respectively, with the selected sub-picture ID value in the reconstructed access unit.
[0409] When the sub-picture ID value of a PPS or SPS changes compared to a previous PPS or SPS (respectively) with the same PPS ID value or SPS ID value, the reader needs to include a copy of the previous PPS and SPS (if a PPS or SPS with the same PPS or SPS ID value, respectively, does not otherwise exist in the access unit) and rewrite the PPS or SPS (respectively) with the updated sub-picture ID value in the reconstructed access unit.
[0410] Sample entry name and format (defined by VVC video stream):
[0411] definition: ...
[0413] A VVC track may contain a “subp” track reference, whose entries contain the track_ID value of a VVC sub-picture track or the track_group_id value of an “alte” track group of a VVC sub-picture track.
[0414] [[When a VVC track contains a 'subp' track reference, it is called a VVC base track, and the following applies:
[0415] - Samples of a VVC track shall not contain VCL NAL units. ]]
[0416] Sample groups of type "spor" as specified in clause 11.7.7 [[SHOULD]] be present in each VVC base track. ...
[0418] grammar:
[0419]
[0420] Semantics:
[0421] Compressorname in the base class VisualSampleEntry indicates the name of the compressor used, where the value "\012VVC Coding" is recommended (\012 is 10, the string length in bytes).
[0422] VvcDecoderConfigurationRecord is defined in 11.3.3.
[0423] nalusInContiguousDecodingOrderFlag equal to 1 indicates that the NAL units in each sample are consecutive in decoding order in the original entire bitstream, so the VVC base track that refers to this VVC sub-picture track through a track reference of type "subp" may also refer to other VVC sub-picture tracks through the same track reference. A value of 0 indicates that the NAL units in each sample may or may not be consecutive in decoding order in the original entire bitstream, so the VVC base track that refers to this VVC sub-picture track through a track reference of type "subp" may not refer to other VVC sub-picture tracks through the same track reference.
[0424] lengthSizeMinusOne plus 1 indicates the length in bytes of the NALUnitLength field in the track containing the VvcNALUConfigBox. The value of this field should be one of 0, 1, or 3 corresponding to the length encoded with 1, 2, or 4 bytes, respectively.
[0425] [[num_subpics_minus1 plus 1 specifies the number of sub-picture sequences contained in the VVC sub-picture track.
[0426] subpic_id specifies the sub-picture identifier of the sub-picture sequence contained in the VVC sub-picture track. ]]
[0427] Figure 1 1 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values), or may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, a passive optical network (PON), etc.) and wireless interfaces (such as a Wi-Fi interface or a cellular interface).
[0428] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 can be stored or transmitted via a connected communication, as represented by component 1906. Component 1908 can use the bitstream (or codec) representation of the stored or transmitted video received at input 1902 to generate pixel values or displayable video to be sent to display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the results of the codec will be performed by the decoder.
[0429] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0430] Figure 2 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor(s) 3602 can be configured to implement one or more methods described in this document. Memory(s) 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing hardware 3606 may be at least partially included in processor 3602 (e.g., a graphics coprocessor).
[0431] Figure 4 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.
[0432] like Figure 4As shown, the video encoding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.
[0433] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0434] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be sent directly to the destination device 120 via the network 130a via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0435] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0436] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 and configured to interface with an external display device.
[0437] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or future standards.
[0438] Figure 5 is a block diagram illustrating an example of a video encoder 200, which may be Figure 4 The video encoder 114 in the system 100 is shown.
[0439] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 5 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0440] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0441] In other examples, the video encoder 200 may include more, fewer, or different functional components. In an example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.
[0442] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the purpose of explanation, they are shown in FIG. Figure 5 are represented separately in the examples.
[0443] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0444] The mode selection unit 203 can select one of the coding modes (e.g., intra or inter) based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).
[0445] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0446] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0447] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0448] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures containing the reference video blocks in list 0 and list 1, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0449] In some examples, motion estimation unit 204 may output a complete motion information set for use in a decoding process by a decoder.
[0450] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0451] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0452] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0453] As described above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0454] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0455] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0456] In other examples, such as in skip mode, for the current video block, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0457] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0458] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0459] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0460] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0461] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0462] Figure 6 is a block diagram illustrating an example of a video decoder 300, which may be Figure 4 The video decoder 114 in the system 100 is shown.
[0463] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 6 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0464] exist Figure 6 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations generally related to the video encoder 200 ( Figure 5 ) is a decoding process that is the inverse of the encoding process described.
[0465] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and Merge modes.
[0466] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0467] The motion compensation unit 302 may calculate interpolated values of sub-integer pixels of a reference block using the interpolation filter used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 based on received syntax information and use the interpolation filter to generate a prediction block.
[0468] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence.
[0469] The intra prediction unit 303 can form a prediction block based on spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0470] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces decoded video for presentation on a display device.
[0471] A list of solutions preferred by some embodiments is provided below.
[0472] A first set of methods is provided below.The following scenarios illustrate example embodiments of the techniques discussed in the previous section (eg, item 1).
[0473] 1. A method for processing visual media, comprising: performing conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the file includes tracks containing data for sub-pictures of the visual media data; and wherein the format rules specify a syntax for the tracks.
[0474] 2. The method of claim 1 , wherein the format rule specifies that the track covers a rectangular area.
[0475] 3. The method of claim 1, wherein the format rules specify that sub-pictures or slices included in a track are individually extractable, decodable, and presentable.
[0476] The following scenarios illustrate example embodiments of the techniques discussed in the previous section (eg, items 3, 4).
[0477] 4. A method for processing visual media, comprising: performing conversion between visual media data and a file storing a bitstream representation of the visual media data according to format rules; wherein the file includes a first track and / or one or more sub-picture tracks; wherein the format rules specify the syntax of the track and / or one or more sub-picture tracks.
[0478] 5. The method of claim 4, wherein the format rules specify that the track includes references to one or more sub-picture tracks.
[0479] 6. The method according to claim 4, wherein the format rules do not allow inclusion of non-video codec layer network abstraction layer units at the access unit level or the picture level in one or more sub-picture tracks.
[0480] 7. The method according to scheme 6, wherein the disallowed unit includes a decoding capability information structure, or a parameter set, or an operation point information, or a header, or an end of a stream, or an end of a picture.
[0481] 8. The method according to any one of schemes 1-7, wherein converting comprises generating a bitstream representation of the visual media data according to a format rule and storing the bitstream representation in a file.
[0482] 9. The method of any one of claims 1 to 7, wherein converting comprises parsing the file according to format rules to recover the visual media data.
[0483] 10. A video decoding device comprising a processor, the processor being configured to implement the method according to one or more of schemes 1 to 9.
[0484] 11. A video encoding device comprising a processor, the processor being configured to implement the method according to one or more of schemes 1 to 9.
[0485] 12. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of schemes 1 to 9.
[0486] 13. A computer-readable medium having a bitstream representation thereon conforming to a file format generated according to any one of schemes 1 to 9.
[0487] 14. A method, apparatus or system as described in this document.
[0488] The second set of scenarios provides example embodiments of the techniques discussed in the previous section (eg, item 1).
[0489] 1. A method for processing visual media data (e.g., Figure 11 The method 110 shown includes: performing 1102 conversion between visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data includes one or more pictures, the one or more pictures including one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track including a sequence of one or more slices or one or more sub-pictures covers a rectangular area of the one or more pictures.
[0490] 2. A method according to claim 1, wherein the format rule specifies that one or more sub-pictures or one or more slices included in the track are individually extractable, decodable and presentable in the absence of another sub-picture or another slice covering another area different from the rectangular area.
[0491] 3. The method of claim 1, wherein the format rule specifies that one or more sub-pictures or one or more slices included in the track are dependent on another sub-picture or another slice covering another area different from the rectangular area in terms of motion compensation.
[0492] 4. The method of claim 1, wherein the format rule specifies that one or more slices or one or more sub-pictures are allowed to be non-contiguous in decoding order for the bitstream stored in the track.
[0493] 5. The method of claim 1 , wherein the field of view of the 360-degree video covered by one or more sub-pictures that are discontinuous in decoding order is represented by a track.
[0494] 6. The method of claim 1 , wherein the format rule specifies that an order of the one or more sub-pictures or the one or more slices in each sample of the track is the same as an order of the one or more sub-pictures or the one or more slices in the bitstream stored in the track.
[0495] 7. The method according to claim 1, wherein the format rule further specifies whether to include an indication indicating whether a decoding order of one or more sub-pictures or one or more slices in each sample of the track is continuous in a bitstream stored in the track.
[0496] 8. The method of claim 7, wherein the indication is included in a base track sample entry description of the track.
[0497] 9. The method according to claim 7, wherein the format rule further specifies that, in response to the absence of the indication, one or more sub-pictures or one or more slices in the track are not allowed to be merged with another sub-picture or another slice of another track.
[0498] 10. The method of claim 7, wherein the indication is included in a network abstraction layer (NAL) configuration box.
[0499] 11. The method of claim 7, wherein the indication being equal to 1 indicates that the NAL units in each sample of the track are consecutive in decoding order of the bitstream, and a base track that references a track with a track reference references other tracks with a track reference.
[0500] 12. The method of claim 7, wherein the indication being equal to 0 indicates that the NAL units in each sample of the track are allowed or not allowed to be consecutive in decoding order of the bitstream, and a base track that references a track with a track reference is not allowed to reference other tracks with a track reference.
[0501] 13. The method according to any one of schemes 1-12, wherein the visual media data is processed by Versatile Video Codec (VVC), and the one or more tracks are VVC tracks.
[0502] 14. The method according to any one of the schemes 1-13, wherein converting comprises generating a visual media file according to a format rule and storing the one or more bitstreams in the visual media file.
[0503] 15. The method of any one of claims 1 to 13, wherein converting comprises parsing the visual media file according to format rules to reconstruct one or more bitstreams.
[0504] 16. An apparatus for processing visual media data, comprising a processor configured to implement a method comprising: performing conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures covers a rectangular area of the one or more pictures.
[0505] 17. The apparatus of claim 16, wherein the format rule specifies whether to include an indication indicating whether a decoding order of one or more sub-pictures or one or more slices in each sample of the track is continuous in a bitstream stored in the track.
[0506] 18. A non-transitory computer-readable recording medium storing instructions that cause a processor to: perform conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more slices or one or more sub-pictures covers a rectangular area of the one or more pictures.
[0507] 19. The non-transitory computer-readable recording medium of claim 18, wherein the format rule specifies whether to include an indication indicating whether a decoding order of one or more sub-pictures or one or more slices in each sample of a track is continuous in a bitstream stored in the track.
[0508] 20. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing device, wherein the method comprises: generating a visual media file, the visual media file comprising one or more tracks of one or more bitstreams storing visual media data; wherein the visual media data comprises one or more pictures, the one or more pictures comprising one or more sub-pictures or one or more slices; and wherein the visual media file stores the one or more tracks according to a format rule; wherein the format rule specifies that a track comprising a sequence of one or more sub-pictures or one or more slices covers a rectangular area of the one or more pictures.
[0509] 21. The non-transitory computer-readable recording medium of scheme 20, wherein the format rule specifies whether to include an indication indicating whether a decoding order of one or more sub-pictures or one or more slices in each sample of a track is continuous in a bitstream stored in the track.
[0510] 22. A video processing device comprising a processor, the processor being configured to implement the method according to any one or more of schemes 1 to 15.
[0511] 23. A method of storing visual media data in a file comprising one or more bitstreams, the method comprising the method according to any one of schemes 1 to 15, and further comprising storing the bitstreams in a non-transitory computer-readable recording medium.
[0512] 24. A computer-readable medium storing program code, which, when executed, causes a processor to implement the method according to any one or more of schemes 1 to 15.
[0513] 25. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0514] 26. A video processing device for storing a bitstream, wherein the video processing device is configured to implement the method according to any one or more of schemes 1 to 15.
[0515] 27. A computer-readable medium, wherein a bitstream thereon complies with a file format generated according to any one of schemes 1 to 15.
[0516] 28. A method, apparatus or system as described in this document.
[0517] The third set of scenarios illustrates example embodiments of the techniques discussed in the previous section (eg, items 3, 5, 6, 7, and 10).
[0518] 1. A method for processing visual media data (e.g., Figure 12 1200 ), comprising: performing 1202 a conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to format rules; wherein the visual media file comprises a base track that references one or more sub-picture tracks that store codec information for one or more sub-pictures of the visual media data, and wherein the format rules specify a process for reconstructing a video unit based on samples in the one or more sub-picture tracks and the base track.
[0519] 2. The method of claim 1 , wherein the format rule specifies that the base track includes a sub-picture track reference that references one or more sub-picture tracks, and an order of the one or more sub-picture tracks referenced in the sub-picture track reference indicates an order of samples of the one or more sub-picture tracks in a video unit reconstructed based on the one or more sub-picture tracks.
[0520] 3. The method according to claim 1, wherein the format rule further specifies that each sub-picture track reference has an index that refers to a track identifier of a sub-picture track or a track group identifier of a sub-picture track group.
[0521] 4. The method of claim 1 , wherein the format rule specifies that the sub-picture order sample group is optional for the base track.
[0522] 5. The method of claim 4, wherein the format rule further specifies that, in the event that a sub-picture order sample group does not exist in the base track, a sub-picture track reference is used in determining the order of one or more sub-picture tracks referenced in the base track.
[0523] 6. The method according to claim 4, wherein the format rule further specifies removing the use of sub-picture order sample groups and removing the description of the parameter set rewriting process based on the sub-picture order sample groups.
[0524] 7. The method according to claim 4, wherein the format rule further specifies a specification for removing sub-picture order sample groups.
[0525] 8. The method according to any one of schemes 1-7, wherein the visual media data is processed by Versatile Video Codec (VVC), and the one or more tracks are VVC tracks.
[0526] 9. The method according to any one of solutions 1 to 8, wherein converting comprises generating a visual media file according to a format rule and storing the one or more bitstreams in the visual media file.
[0527] 10. The method of any one of claims 1 to 8, wherein converting comprises parsing the visual media file according to format rules to reconstruct one or more bitstreams.
[0528] 11. An apparatus for processing visual media data, comprising a processor configured to implement a method comprising: performing conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to format rules; wherein the visual media file comprises a base track that references one or more sub-picture tracks, the sub-picture tracks storing codec information for one or more sub-pictures of the visual media data, and wherein the format rules specify a process for reconstructing a video unit based on samples in the one or more sub-picture tracks and the base track.
[0529] 12. An apparatus according to claim 11, wherein the format rule specifies that the base track includes a sub-picture track reference that references one or more sub-picture tracks, and an order of the one or more sub-picture tracks referenced in the sub-picture track reference indicates an order of samples of the one or more sub-picture tracks in a video unit reconstructed based on the one or more sub-picture tracks.
[0530] 13. The apparatus according to claim 11, wherein the format rule further specifies that each sub-picture track reference has an index that refers to a track identifier of the sub-picture track or a track group identifier of the sub-picture track group.
[0531] 14. The apparatus of claim 11, wherein the format rule specifies that the sub-picture order sample group is optional for the base track.
[0532] 15. The apparatus of claim 14, wherein the format rule further specifies that, in the event that a sub-picture order sample group does not exist in the base track, a sub-picture track reference is used in determining the order of the one or more sub-picture tracks referenced in the base track.
[0533] 16. The apparatus according to claim 14, wherein the format rule further specifies removing the use of sub-picture order sample groups and removing the description of the parameter set rewriting process based on the sub-picture order sample groups.
[0534] 17. The apparatus of claim 14, wherein the format rule further specifies a specification for removing sub-picture order sample groups.
[0535] 18. A non-transitory computer-readable recording medium storing instructions that cause a processor to: perform conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data according to format rules; wherein the visual media file comprises a base track that references one or more sub-picture tracks that store codec information for one or more sub-pictures of the visual media data, and wherein the format rules specify a process for reconstructing a video unit based on samples in the one or more sub-picture tracks and the base track.
[0536] 19. A non-transitory computer-readable recording medium according to scheme 18, wherein the format rule specifies that the base track includes a sub-picture track reference that references one or more sub-picture tracks, and the order of the one or more sub-picture tracks referenced in the sub-picture track reference indicates the order of samples of the one or more sub-picture tracks in a video unit reconstructed based on the one or more sub-picture tracks.
[0537] 20. A non-transitory computer-readable recording medium storing a bitstream generated by a method performed by a video processing device, wherein the method comprises: generating a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of visual media data according to format rules; wherein the visual media file comprises a base track that references one or more sub-picture tracks, the sub-picture tracks storing codec information of one or more sub-pictures of the visual media data, and wherein the format rules specify a process for reconstructing a video unit based on samples in the one or more sub-picture tracks and the base track.
[0538] 21. A video processing device comprising a processor, the processor being configured to implement the method according to any one or more of schemes 1 to 10.
[0539] 22. A method of storing visual media data in a file comprising one or more bitstreams, the method comprising the method according to any one of schemes 1 to 10, and further comprising storing the bitstreams in a non-transitory computer-readable recording medium.
[0540] 23. A computer-readable medium storing program code, which, when executed, causes a processor to implement the method according to any one or more of schemes 1 to 10.
[0541] 24. A computer-readable medium storing a bitstream generated according to any one of the above methods.
[0542] 25. A video processing device for storing a bitstream, wherein the video processing device is configured to implement the method according to any one or more of solutions 1 to 10.
[0543] 26. A computer-readable medium having a bitstream representation thereon conforming to a file format generated according to any one of schemes 1 to 10.
[0544] 27. A method, apparatus or system as described in this document.
[0545] In example embodiments, the visual media data corresponds to a video or image. In the embodiments described herein, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the embodiments described herein, a decoder can parse syntax elements in the codec representation using the format rules, knowing the presence or absence of syntax elements according to the format rules, to produce a decoded video. In the embodiments described above, the visual media data corresponds to a video or image.
[0546] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits that are co-located within the bitstream or distributed across different locations within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on a transformed and coded error residual value and also using bits in the header and other fields in the bitstream. Furthermore, during conversion, the decoder may, based on this determination, parse the bitstream with the knowledge that some fields may or may not be present, as described in the above scheme. Similarly, the encoder may determine whether certain syntax fields are included and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.
[0547] The disclosed and other schemes, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition that implements a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0548] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form (including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment). A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program may be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communications network.
[0549] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0550] By way of example, processors suitable for executing computer programs include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include, or be operatively coupled to, one or more mass storage devices for storing data (e.g., magnetic, magneto-optical, or optical disks), to receive data from or send data to, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example: semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0551] Although this patent document contains many details, these details should not be interpreted as limitations on any subject matter or the scope of the claims, but rather as descriptions of features that may be specific to a particular embodiment of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as working in certain combinations, and even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the combination, and the claimed combination may point to a variant of a sub-combination or sub-combination.
[0552] Similarly, while operations may be described in a particular order in the drawings, this should not be understood as requiring that the operations be performed in the particular order shown or in the sequential order shown, or that all illustrated operations be performed, in order to achieve the desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0553] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing visual media data, comprising: performing conversion between visual media data and a visual media file, the visual media file comprising one or more tracks storing one or more bitstreams of the visual media data; The visual media data includes one or more pictures, and the one or more pictures include one or more sub-pictures or one or more strips; wherein the visual media file stores the one or more tracks according to a format rule, the format rule being specified to comply with the Versatile Video Codec (VVC) standard, and the visual media file includes a base track that references one or more sub-picture tracks, the one or more sub-picture tracks storing codec information of one or more sub-pictures of the visual media data; wherein the format rule specifies that a track comprising a sequence of the one or more sub-pictures or the one or more slices covers a rectangular area of the one or more pictures; and The format rule further specifies that the base track includes a sub-picture track reference to reference the one or more sub-picture tracks, and the order in which the one or more sub-picture tracks in the sub-picture track reference are referenced indicates the order of samples of the one or more sub-picture tracks in a video unit reconstructed from the one or more sub-picture tracks.
2. The method according to claim 1, wherein The format rules specify that one or more sub-pictures or one or more slices included in the track are individually extractable, decodable and presentable in the absence of another sub-picture or another slice covering another area different from the rectangular area.
3. The method according to claim 1, wherein The format rule specifies that one or more sub-pictures or one or more slices included in the track are dependent on another sub-picture or another slice covering another area different from the rectangular area in terms of motion compensation.
4. The method according to claim 1, wherein The format rule specifies that the one or more slices or the one or more sub-pictures are allowed to be non-contiguous in decoding order for a bitstream stored in the track.
5. The method according to claim 1, wherein The field of view of the 360-degree video covered by one or more sub-pictures that are not consecutive in decoding order is represented by the track.
6. The method according to claim 1, wherein The format rules specify that the order of sub-pictures or slices in each sample of the track is the same as the order of sub-pictures or slices in the bitstream stored in the track.
7. The method according to claim 1, wherein The format rule further specifies whether to include an indication indicating whether the decoding order of one or more sub-pictures or one or more slices in each sample of the track is continuous in the bitstream stored in the track.
8. The method according to claim 7, wherein: The indication is included in the description of the base track sample entry for the track.
9. The method according to claim 7, wherein: The format rules further specify that, in response to the absence of the indication, one or more sub-pictures or one or more slices in the track are not allowed to be merged with another sub-picture or another slice of another track.
10. The method according to claim 7, wherein: The indication is included in a Network Abstraction Layer (NAL) configuration box.
11. The method according to claim 7, wherein: The base track references tracks with track references, the indication being equal to 1 indicates that NAL units in each sample of the track are consecutive in decoding order of the bitstream, and the base track references other tracks with the track references.
12. The method according to claim 7, wherein: The base track references a track using a track reference, the indication being equal to 0 indicates that NAL units in each sample of the track are allowed or not allowed to be consecutive in decoding order of the bitstream, and the base track is not allowed to reference other tracks using the track reference.
13. The method according to any one of claims 1 to 12, wherein: The visual media data is processed by a versatile video codec (VVC), and the one or more tracks are VVC tracks.
14. The method according to any one of claims 1 to 12, wherein: The converting includes generating the visual media file according to the format rules and storing the one or more bitstreams in the visual media file.
15. The method according to any one of claims 1 to 12, wherein: The conversion includes parsing the visual media file according to format rules to reconstruct the one or more bitstreams.
16. A video processing device comprising a processor configured to implement the method according to any one of claims 1 to 15.
17. A method for storing a video bitstream, comprising: Generating a bitstream of a video according to the method of any one of claims 1 to 15, and The bitstream is stored in a non-transitory computer-readable recording medium.
Citation Information
Patent Citations
An apparatus, a method and a computer program for video coding and decoding
WO2020141248A1