Subpicture Track in Encoded Video
The VVC standard's sub-picture feature addresses the inefficiencies in encoding and decoding video data by enabling independent encoding and decoding of rectangular areas within a picture, improving encoding efficiency and reducing computational complexity for viewport-dependent streaming.
Patent Information
- Application Number
- JP2023196800
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-06
- Filing Date
- 2023-11-20
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2041-09-17
AI Technical Summary
Existing video coding standards face challenges in efficiently encoding and decoding video data, particularly in applications like 360° immersive media, where viewport-dependent streaming requires high bandwidth and efficient data extraction of sub-pictures, leading to increased computational complexity and overhead.
The introduction of sub-pictures in the VVC standard allows for independent encoding and extraction of rectangular areas within a picture, enabling viewport-dependent streaming by allowing sub-pictures to be encoded and decoded independently, with features like in-loop filtering and motion vector fine-tuning, reducing computational overhead and improving encoding efficiency.
This approach enhances encoding efficiency by optimizing bandwidth usage and reducing computational complexity, allowing for seamless viewport changes in 360° video streaming and other applications with reduced bitstream redundancy.
Smart Images

Figure 0007798846000024 
Figure 0007798846000025 
Figure 0007798846000026
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on Japanese Patent Application No. 2021-152668, filed on September 17, 2021. Timely claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 079933, filed September 17, 2020, and U.S. Provisional Patent Application No. 63 / 088126, filed October 6, 2020. do. For all purposes under law, the entire disclosure of the above application is incorporated by reference as part of the disclosure of this specification.
[0002] This patent document describes a method for generating and recording digital audiovisual media information in a file format. memory, and consumption. [Background technology]
[0003] Digital video is the largest bandwidth used on the Internet and other digital communication networks. The number of connected users who can receive and display video is used. As the number of devices increases, the bandwidth demands for digital video usage will continue to grow. is predicted. Summary of the Invention
[0004] This specification describes how a video encoder and decoder can encode video or The present invention discloses techniques that can be used to process coded representations of images.
[0005] In one exemplary embodiment, a method for processing visual media data is disclosed. The method comprises: visual media data; and one or more bitstreams of said visual media data. performing a conversion between a visual media file including one or more tracks storing the The visual media data includes one or more sub-pictures or slices. The visual media file includes one or more pictures, and the visual media file is formatted according to a format rule. the one or more tracks are stored, and the formatting rules are a track containing a sequence of one or more sub-pictures is included in the sequence of one or more pictures Specifies that the rectangular area covered by the
[0006] In another exemplary aspect, another method of processing visual media data is disclosed. The method may include the creation of visual media data and one or more components of visual media data in accordance with the formatting rules. A visual media file containing one or more tracks storing the above bitstreams. and converting the visual media file to one or more of the visual media data. one or more subpicture tracks storing coded information for the subpictures of the the format rule includes a base track that references the base track, Used to reconstruct a video unit from samples and one or more subpicture tracks. It specifies the process to be used.
[0007] In yet another exemplary aspect, a video processing device is disclosed, the video processing device comprising: A processing device configured to implement the above-described method is provided.
[0008] In yet another exemplary embodiment, a file containing one or more bitstreams may be A method for storing media data is disclosed, the method corresponding to the method described above, and storing the one or more bitstreams on a non-transitory computer-readable recording medium; Further includes:
[0009] In yet another exemplary aspect, a computer-readable medium storing a bitstream The bitstream is generated according to the method described above.
[0010] In yet another exemplary aspect, a video processing device for storing a bitstream comprises: A video processing device is disclosed, configured to implement the above-mentioned method.
[0011] In yet another exemplary embodiment, the bitstream is generated according to the method described above. A computer readable medium conforming to the file format is disclosed.
[0012] These and other features are described throughout this specification. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a block diagram illustrating an example of a video processing system. [Figure 2] FIG. 2 is a block diagram of the video processing device. [Figure 3] FIG. 3 is a flowchart showing an example of a video processing method. [Figure 4] FIG. 4 is a block diagram illustrating a video encoding system according to some embodiments of the present disclosure. [Figure 5] FIG. 5 is a block diagram illustrating an encoder according to some embodiments of this disclosure. [Figure 6] FIG. 6 is a block diagram illustrating a decoder according to some embodiments of the present disclosure. [Figure 7] FIG. 7 shows an example of an encoder block diagram. [Figure 8] FIG. 8 shows a picture divided into 18 tiles, 24 slices, and 24 sub-pictures. [Figure 9] FIG. 9 shows a typical sub-picture-based viewport-dependent 360° video distribution scheme. [Figure 10]FIG. 10 shows an example of extracting one sub-picture from a bitstream containing two sub-pictures and four slices. [Figure 11] FIG. 11 illustrates an exemplary method for visual media data processing according to some implementations of the disclosed technology. [Figure 12] FIG. 12 illustrates an exemplary method for visual media data processing according to some implementations of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION
[0014] This specification uses section headings to facilitate understanding, and the technology and each The applicability of an embodiment described in a section is not limited to that section alone. The term .266 is used in certain descriptions for ease of understanding only and is not intended to be a substitute for disclosure. It is not intended to limit the scope of the technology described herein. The techniques presented are applicable to other video codec protocols and designs. In the specification, editing changes are made in accordance with the VVC standard or the ISOBMFF file format standard. The current draft of the case has a strikethrough indicating the struck out text and the added text. This is indicated in the text by highlighting (including bold italics) to indicate the
[0015] 1. Initial consultations This specification relates to video file formats. Specifically, the present invention relates to ISO-based Media files are based on the ISOBMFF media file format. Subpictures in multi-track Versatile Video Coding (VVVC) video bitstreams This idea applies to any codec, e.g., video encoded by the VVC standard. Image bitstreams, and any video file formats, e.g., May be applied individually or in various combinations to the VVC video file format . 2. Abbreviation ACT Adaptive Color Transform ALF Adaptive Loop Filter AMVR Adaptive Motion Vector Resolution APS adaptive parameter set AU Access Unit AUD Access Unit Delimiter AVC Advanced Video Coding (Rec.ITU-T H.264|ISO / IEC14496 -10) B. Bidirectional prediction BCW CU-level weighted bidirectional prediction BDOF Bidirectional Optical Flow BDPCM Block-Based Delta Pulse Code Modulation BP Buffering Time CABAC: Context-based adaptive binary arithmetic coding. CB coded block CBR Constant Bitrate CCALF Cross-Component Adaptive Loop Filter CPB Coded Picture Buffer CRA Clean Random Access CRC Cyclic Redundancy Check CTB coding tree block CTU Coding Tree Unit CU Coding Unit CVS coded video sequence DPB Decoded Picture Buffer DCI Decoding Capability Information DRAP Dependent Random Access Point DU Decoding Unit DUI Decoding Unit Information EG Exponential Golomb EGk kth exponential Golomb EOB End of bitstream EOS End of sequence FD Filler Data FIFO First In First Out FL fixed length GBR Green, Blue, Red GCI General Restrictions Information GDR Gradual Decryption Update GPM Geometry Division Mode HEVC High Efficiency Video Coding (Rec.ITU-T H.265|ISO / IEC230 08-2) HRD Hypothetical Reference Decoder HSS Virtual Stream Scheduler I Intra IBC Intra Block Copy IDR Instant Decoding Update ILRP Interlayer Reference Image IRAP Intra Random Access Point LFNST Low Frequency Non-separable Transform LPS Minimum Probability Symbol LSB least significant bit LTRP Long-Term Reference Image LMCS Luminance Mapping with Saturation Scaling Intra prediction based on MIP matrix MPS Maximum Probability Symbol MSB Most Significant Bit MTS Multiple Conversion Selection MVP Motion Vector Prediction NAL Network Abstraction Layer OLS output layer set OP operating point OPI operating point information P prediction PH Picture Header POC Picture Order Count PPS Picture Parameter Set PROF Optical Flow Prediction Fine Tuning PT Picture Timing PU Picture Unit QP quantization parameter RADL Random Access Decodable Read (Picture) RASL Random Access Skip Read (Picture) RBSP Raw Byte Sequence Payload RGB Red, Green, Blue RPL Reference Picture List SAO Sample Adaptive Offset SAR Sample Aspect Ratio SEI Supplemental Enhancement Information SH slice header SLI subpicture level information SODB String of Data Bits SPS Sequence Parameter Set STRP Short-Term Reference Picture STSA stepwise temporal sublayer access TR shortened rice VBR Variable Bit Rate VCL video encoding layer VPS video parameter set VSEI General Supplemental Enhancement Information (Rec.ITU-T H.274 | ISO / IEC2 3002-7) VUI video availability information VVC(Rec.ITU-T H.266 | ISO / IEC23090-3)H.2 65|ISO / IEC23008-2)
[0016] 3. Introduction of video coding 3.1. Video Coding Standards Video coding standards have emerged primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T has developed H.261 and H.263, and ISO / IEC has developed MPEG- 1 and MPEG-4 Visual, and both organizations are working on H.262 / MPEG-2 Video o and H.264 / MPEG-4 AVC(Advanced Video Coding ) and co-created the H.265 / HEVC standard. Since H.262, video coding standards have It is based on a hybrid video coding structure that utilizes inter-prediction and transform coding. In 2015, VCEG and MPEG jointly launched the We established the Joint Video Exploration Team (JVET). Since then, many new methods have been adopted by JVET and JEM (Joint Exp. It has been incorporated into the reference software called JVE (Informational Mode). T was later appointed as the Joint Video Expo'er when the Universal Video Coding (VVC) project was formally launched. The name was changed to JVET (Joint Video Engineering Test Team). VVC is a new coding standard, and HEVC The 19th J It was completed at the VET General Assembly.
[0017] General purpose video coding (VVC) standard (ITU-T H.266|ISO / IEC 2309 0-3) and the General Purpose Supplementary Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IE C 23002-7) is used for traditional purposes such as television broadcasting, video conferencing, and playback from storage media. In addition to the above, adaptive bitrate streaming, video region extraction, and multiple coded video bitstreaming are also supported. Composition and merging of content from streams, multi-view video, scalable layered content Newer, more advanced features include 360° immersive media that adapts to the viewport, It is designed for use in the widest range of applications, including high-speed applications.
[0018] 3.2. File Format Standards Media streaming applications typically use IP, TCP, and HTTP. It is based on the P transport method and is commonly used for ISO-based media file formats. Such streams depend on file formats such as ISOBMFF. One streaming system is Dynamic Adaptive Streaming over HTTP (DASH). To use ISOBMFF and DASH video formats, you must comply with ISO / IEC 1 4496-15 ("Information Technology - Coding of Audiovisual Objects - Part 1" 5: ISO-based media file structured in Network Abstraction Layer (NAL) units AVC file format and HEVC file format The file format specifications specific to video formats, such as the ISO 14001 standard, are To encapsulate video content into BMFF tracks and DASH representations and segments Important information about the video bitstream, e.g., profile, layer , levels, and many more for content selection, e.g., streaming sessions for both initialization at the start of a streaming session and stream adaptation during the streaming session. File format level metadata and / or DASH media presentation The MPI must be published as a Multi-Profile Component Description (MPD).
[0019] Similarly, to use an image format using ISOBMFF, AVC image file in ISO / IEC23008-12, for example, formats and file formats such as HEVC image file format Specification ("Information technology - High-efficiency coding and media delivery in heterogeneous environments - Part 12: image file format) is required.
[0020] A file format for storing VVC video content based on ISOBMFF The VVC video file format is currently being developed by MPEG. The latest draft specification for the VC video file format is MPEG Output Document N19454 ("Information Information Technology - Audiovisual Object Coding - Part 15: Network Abstraction ISO-based media file format video structured in NAL (Non-Alternative Layer) units Carriage Correction 2: VVC and EVC carriage correction in ISOBMFF ", July 2020).
[0021] A file based on ISOBMFF for storing image content coded using VVC The VVC image file format is currently supported by MPEG The latest draft specification for the VVC image file format is available at the MPEG Output Document N19460 ("Information Technology - Highly Efficient Coding and Media Delivery in Heterogeneous Environments - Part 12 Part: Image File Formats - Correction 3: VVC, EVC, Slideshows and Others "Support for Improvements to the Salesforce Platform," July 2020.
[0022] 3.3. Picture Partitioning Scheme in HEVC HEVC supports regular slices, dependent slices, tiles, WPP (Waveform Presets) Four different picture division processes are available: It includes a scheme that allows you to apply it to match the maximum transmission unit (MTU) size. This allows for faster packetization, parallel processing, and reduced end-to-end latency.
[0023] Regular slices are similar to H.264 / AVC. Each regular slice is its own NAL units and performs in-picture prediction across slice boundaries (intra- sample prediction, motion information prediction, coding mode prediction) and entropy coding dependency are not This way, regular slices can be distinguished from other regular slices in the same picture. can be independently reconfigured (but still requires a loop filtering operation) (There may be interdependencies between them).
[0024] Regular slicing is the only tool available for parallelization, even in H.264 / AVC. It is available in almost the same format. Regular slice-based parallelism requires less processing. No inter-device or inter-core communication is required (decoding predictively coded pictures In general, there is no inter-processor or inter-core data sharing within a picture, except for motion compensation. (Much heavier than inter-processor or inter-core data sharing for prediction). However, For the same reason, regular slices allow for a bit cost for the slice header and slice The lack of prediction across the boundary of the data stream results in a large coding overhead. Furthermore, regular slicing (as opposed to other tools discussed below) is Intra-picture independence of slices and each regular slice is encapsulated in its own NAL unit Due to the celling, the bitstream must be split to fit the MTU size requirements. It also serves as a key mechanism for dividing the The goal of TU size matching is to resolve conflicting requirements on slice layout in a picture. This situation has led to the development of the following parallelization tools:
[0025] Dependent slices have short slice headers and do not suspend intra-picture prediction at all. This allows splitting the bitstream at treeblock boundaries rather than at the A dependent slice is a regular slice fragmented into multiple NAL units, and the entire regular slice is By allowing some of the regular slices to be sent before the field encoding is complete, to reduce end-to-end latency.
[0026] In WPP, a picture is divided into single-row coding tree blocks (CTBs). Entropy decoding and prediction can be performed using data from the CTB in other partitions. Parallel processing is possible by parallel decoding of CTB rows, and the number of The start of decoding is delayed by two CTBs, so that the target CTB is decoded before This ensures that data on the CTB to the right of the target CTB is available. By using the start of the wave (which looks like a wave front when represented graphically), It is possible to parallelize the processor using up to the number of processors / cores that the processor contains in the CTB row. Intra-picture prediction between neighboring treeblock rows within a picture is allowed, so The inter-processing unit / inter-core communication required to enable in-core prediction may be sufficient. The allocation does not result in the generation of additional NAL units compared to the case where it is not applied, and therefore WPP is not a tool for MTU size matching. If matching is required, WPP can generate regular slices with some coding overhead. can be used.
[0027] Tiles define horizontal and vertical boundaries that divide the picture into columns and rows of tiles. Columns of tiles run from top to bottom of the picture. Similarly, rows of tiles run from left to right of the picture. The number of tiles in a picture is simply the number of tile columns multiplied by the number of tile rows. It can be derived by calculating
[0028] The scan order of the CTB is local within one tile (the CTB of one tile TB raster scan order) and then tile raster scan of one picture The next tile's top left CTB is decoded according to the order of the slices. However, the rule impairs intra-picture prediction dependency and entropy decoding dependency. They do not need to be contained in individual NAL units (as in WPP) and are therefore Therefore, tiles cannot be used for MTU size matching. Each tile is allocated to one processing unit / The inter-processing unit / inter-processor unit required for intra-picture prediction may be processed by the core. Inter-core communication allows decoding neighboring tiles, which means that one slice can span two or more tiles. If so, it carries a shared slice header and the reconstructed samples and metadata sharing related to loop filtering. If a slice contains more than one tile or WPP segment, the first The entry point byte offset of each tile or WPP segment other than the one Signaled in the rice header.
[0029] For ease of explanation, HEVC supports four different picture partitioning schemes: Restrictions on application are specified. A given coded video sequence must conform to the HEVC For most of the profiles, it is not possible to include both tiles and wavefronts. For each slice and tile, one or both of the following conditions must be met: 1) All coding tree blocks in one slice belong to the same tile. 2) All coding tree blocks in one tile belong to the same slice. Finally, one wavefront segment contains exactly one CTB row, and WPP is used. When using , if a slice starts within a CTB line, it must end on the same CTB line. It must be.
[0030] Recent HEVC amendments include the JCT-VC output document JCTVC-AC1005, J. Boyce, A. Ramasubramonian, R. Sukupin, G.J. Suri, A. Tulapis ,Y.-K. Wang (editors), "HEVC Additional Supplemental Enhancement Information (Draft 4)," October 24, 2017. Available at: http: / / phenix.int-ev ry.fr / jct / doc_end_user / documents / 29_Maca u / wg11 / JCTVC-AC1005-v2.zip. Including this correction, HEVC ,Three MCTS-related SEI messages, namely, temporal MCTS SEI messages, M CTS Extraction Information Set SEI message and MCTS Extraction Information Nest SEI message This defines the
[0031] The temporal MCTS SEI message indicates the presence of an MCTS in the bitstream. In each MCTS, the motion vector is It refers to the full sample position inside the MCTS, and only the full sample position inside the MCTS is used for interpolation. The block is limited to point to the fractional sample location you need and is external to the MCTS. The use of motion vector candidates derived from blocks for temporal motion vector prediction is permitted. In this way, each MCTS is independent and there are no tiles that are not included in the MCTS. The encoded data may be decoded as follows.
[0032] The MCTS Extraction Information Set (SEI) message is used to extract the MCTS sub-bitstream (S Provides supplemental information that can be used in the EI message (specified as part of the semantics of the message). , generate a conforming bitstream for the MCTS set. This information is extracted from multiple Each extracted information set defines multiple MCTS sets, and the MCTS sub-sets are Alternate VPS, SPS, and PPS RBs used in the bitstream extraction process Contains SP bytes. Sub-bitstream extraction is performed according to the MCTS sub-bitstream extraction process. When extracting a program, you can rewrite the parameter set (VPS, SPS, PPS) or It must be replaced because one or more of the syntax elements related to the slice address are all (first_slice_segment_in_pic_flag and sl ice_segment_address) should generally be different values. Therefore, the slice header needs to be updated slightly.
[0033] 3.4. Picture Partitioning and Subpictures in VVC 3.4.1. Picture Splitting in VVC In VVC, a picture consists of one or more tile rows and one or more tile columns. A tile is a sequence of one CTU that covers a rectangular area of one picture. The CTUs in a tile are sequenced in raster scan order within that tile. It will be canceled.
[0034] A slice is an integer number of complete tiles or It consists of an integer number of consecutive complete CTU rows.
[0035] Two modes of slicing: raster scan slicing mode and rectangular slicing mode. In raster scan slice mode, one slice is one A tile in a picture contains one complete sequence of tiles in a raster scan. In the shape slice mode, a slice forms a set of rectangular areas of the picture. A set of complete tiles or a single tile that together form a rectangular area of the picture. Contains any of multiple contiguous complete CTU rows. The tiles are scanned in the order of raster scan within a rectangular area corresponding to the device.
[0036] A subpicture is one or more sliders that together cover a rectangular area of a picture. Includes
[0037] 3.4.2. Subpicture Concept and Functionality In VVC, each sub-picture is a rectangular area of a picture, as shown in Figure 8. A subpicture consists of one or more complete rectangular slices that together cover the entire image area. It may be specified as possible (i.e., other subpictures of the same picture and the previous picture) may be coded independently in the decoding order of the image) and are designated as non-extractable. Whether or not a sub-picture is extractable, the encoder may In-loop filtering (non-blocking) across sub-picture boundaries for each picture individually You can control whether or not to apply any of the following features (including locking, SAO, and ALF).
[0038] Functionally, subpictures correspond to Motion Constrained Tile Sets (MCTS) in HEVC. They are similar. They both provide viewport-dependent 360° video streaming. For use cases such as coding optimization and region of interest (ROI) applications, It allows for the independent encoding and extraction of rectangular subsets of a sequence of pictures.
[0039] In 360° video streaming, also known as omnidirectional video streaming, any At any given moment, only a subset of the entire omnidirectional video sphere (i.e., the current viewport) is rendered to the user, while the user can turn their head at any time to change the direction of their gaze. It can be used on the client side to change the current viewport. Represent areas not covered by the current viewport with at least some reduced quality, where possible, and ready to render to the user, but the user may suddenly In case the line direction is changed to any location on the sphere, a high-quality representation of the omnidirectional image is possible. It is only needed for the current viewport being rendered to the user at that moment. By dividing the high-quality representation of the entire image into sub-pictures of appropriate granularity, as shown in Figure 8, On the left are 12 high-resolution sub-pictures, and on the right are the remaining 12 low-resolution omnidirectional This allows for optimization by placing sub-pictures of the video.
[0040] Another typical subpicture-based viewport-dependent 360° video streaming scheme A system like this is shown in Figure 9, where only the higher resolution representation of the full image is shown as subpictures. while a lower resolution representation of the full image does not use subpictures and is more It can be encoded with less frequent RAP than high resolution representation. If the video is received at a lower resolution and a higher resolution video is received, the client will Only the subpictures that cover the image are received and decoded.
[0041] 3.4.3. Differences between Subpicture and MCTS There are several important design differences between Subpicture and MCTS. First, The feature of subpictures in VVC is that in this case, samples are By applying padding, even if subpictures can be extracted, the picture motion vectors of coding blocks pointing outside the sub-picture, as in the case at the boundaries of Second, VVC merge mode and decoder-side motion vector fine-tuning processing In this paper, we introduce additional modifications for the selection and derivation of motion vectors. Higher coding efficiency compared to non-prescriptive motion constraints applied at the encoder side for CTS Third, it allows for the extraction of one or more extractable subpictures from a sequence of pictures. When extracting the data and generating a sub-bitstream that is a conforming bitstream, SH( and PH NAL units, if present, do not need to be rewritten. C Sub-bitstream extraction based on MCTS requires rewriting of SH. In both HEVC MCTS extraction and VVC subpicture extraction, SP However, in general, the bitstream There are only a few parameter sets in the image, and each picture has at least one slice. Therefore, rewriting SH can be a heavy burden on application systems. Fourth, slices of different subpictures within one picture may have different NAL unit types. This often involves one pin, as explained in more detail below. This is a feature called mixed NAL unit types or mixed subpicture types within a picture. Fifth, VVC specifies HRD and level definition for subpicture sequences. Therefore, the sub-bitstream conformance of each extractable sub-picture sequence is evaluated. This can be guaranteed by the encoder.
[0042] 3.4.4. Mixing Subpicture Types Within a Picture In AVC and HEVC, all VCL NAL units in one picture Each picture must have the same NAL unit type. Options for mixing subpictures with different VCL NAL unit types within a This allows random access not only at the picture level but also at the sub-picture level. In the VVC VCL, NAL units within one subpicture are supported. must still have the same NAL unit type.
[0043] The random access capability from IRAP subpictures is ideal for 360° video applications. This is useful for viewport-dependent 360° video distribution similar to that shown in Figure 9. In this communication scheme, the contents of spatially adjacent viewports overlap significantly, i.e., Only part of the subpicture in the viewport is displayed in the new viewport while changing the viewport orientation. The subpicture is replaced with the viewport subpicture, leaving most of the subpicture in the viewport. Any new sub-picture sequence introduced into a video frame must start with an IRAP slice. However, it ensures that the remaining subpictures perform inter prediction when the viewport changes. If permitted, the overall transmission bit rate can be significantly reduced.
[0044] A picture may contain only one type of NAL unit, or may contain more than one type. An indication of the mix is provided for the PPS to which the picture refers (i.e., pps_mixed_n alu_types_in_pic_flag). The picture consists of a subpicture containing the IRAP slice and a subpicture containing the last slice. The NAL unit type RASL can be configured simultaneously within one picture. and a few others of different NAL unit types, including the first picture slice of the RADL This allows for a combination of open Subpicture sequences with GOP and close GOP coding structures are coded as one bit Can be merged into a stream.
[0045] 3.4.5. Subpicture Layout and ID Signaling The layout of sub-pictures in VVC is signaled in the SPS, and therefore: Each subpicture has a fixed position of its top left CTU and a fixed number of CTUs. , and thus one subpicture is signaled by its width and height in CT It covers a rectangular area of a picture with U granularity exactly. The order in which they are known determines the index of each sub-picture within the picture.
[0046] Extraction and merging of subpicture sequences is possible without rewriting SH or PH. To make it more efficient, the slice addressing scheme in VVC uses subpicture IDs. and subpicture-specific slice indexes to assign slices to subpictures. In SH, the subpicture ID of the subpicture containing the slice and the subpicture ID of the subpicture containing the slice are The picture-level slice index is signaled. The value of a subpicture ID may be different from the value of its subpicture index. The mapping between is signaled in the SPS or PPS (but not both), or If present, the sub-picture sub-bitstream extraction process Whether to rewrite subpicture ID mapping when rewriting SPS and PPS during or need to add subpicture ID and subpicture level slice index. The index and the slice index together represent the number of the slice within the DPB slot of the decoded picture. Indicates to the decoder the exact location of the decoded CTU of 1. After sub-bitstream extraction , the subpicture ID of the subpicture does not change, but the subpicture index Even if the raster scan CTU address of the first CTU in the subpicture is Subpicture I is a bitstream, even if its address has changed compared to the value in the original bitstream. The unchanged value of D and the subpicture level slice index of each SH The method still determines the position of each CTU in the decoded picture of the extracted sub-bitstream. Figure 10 shows the subpicture ID, subpicture index, and and subpicture level slice indexes to create two subpictures and four This example includes a slice of , which demonstrates enabling subpicture extraction.
[0047] Similar to subpicture extraction, the signaling for subpictures is done in a different bitstream. If the programs are generated cooperatively (e.g., using different subpicture IDs), but not In the case of almost aligned SPS, PPS, and PH parameters, e.g., CTU size size, chroma format, encoding tools, etc.), rewrite SPS and PPS Just combine several subpictures from different bitstreams into one bitstream. Allows a new version to be merged into the system.
[0048] In SPS and PPS, subpictures and slices are transmitted independently. However, to form a compliant bitstream, the There are inherent mutual constraints on the use of subpictures. First, the existence of subpictures requires the use of rectangular slices, Raster scan slicing must be prohibited. Second, the slices of a given subpicture should be consecutive NAL units in decoding order, which means that The layout of the slice controls the order of coded slice NAL units in the bitstream. It means to promise. 3.5. VVC Video File Format Details 3.5.1. Track Types The VVC video file format is the VVC bitstream in an ISOBMFF file. The following types of video tracks are defined for carriage of the beam: a) VVC Truck: A VVC track may contain NAL units in its samples and sample entries. and possibly other sublayers of the VVC bitstream. By referencing VVC tracks and possibly VVC subpicture tracks Represents a VVC bitstream by referencing a track. One VVC track When a references a VVC subpicture track, it is called a VVC base track. b) VVC non-VCL tracks: ALF, LMCS, or APS carrying scaling list parameters, and others Non-VCL NAL units are stored in tracks separate from tracks containing VCL NAL units. may be stored in and transmitted over the track, which is a VVC non-VCL track It is. c) VVC Subpicture Track: A VVC sub-picture track contains either: A sequence of one or more VVC subpictures. A sequence of one or more complete slices that form a single rectangular area. One sample of a VVC sub-picture track contains either: Consecutive 1s in decoding order as specified in ISO / IEC23090-3 More than one complete subpicture. Form a rectangular area as specified in ISO / IEC23090-3, One or more complete slices consecutive in decoding order. Any VVC subpicture or sample in any sample of the VVC subpicture track The rice is consecutive in decoding order. Note: VVC non-VCL tracks and VVC subpicture tracks are not supported for streaming. These enable optimal delivery of VVC video in applications as follows: Each track may be carried in its own DASH representation, and a subset of tracks may be A subset of the VVC subpicture tracks is used to decode and render the video. DASH representations containing non-VCL tracks, and DASH representations containing non-VCL tracks, are In this way, APS and other non-VCL NAs can be Redundant transmission of L units can be avoided. 3.5.2. Overview of rectangular regions carried in VVC bitstreams This specification helps to describe a rectangular area that consists of either: - a sequence of one or more VVC sub-pictures consecutive in decoding order, or - A sequence of one or more complete slices that form a rectangular region and are consecutive in decoding order. Sequence. Rectangular areas cover rectangles without holes. Rectangular areas within a picture do not overlap each other. The rectangular region is a rectangular visual sampling group with rect_region_flag equal to 1. Loop description entries (i.e., RectangularRegionGroupEnt ry) may be described as Each sample in a track consists of a single rectangular NAL unit. If so, use SampleToGroupBox of type 'trif' to group the samples into a rectangle. can be associated with a shape region, but the default sample grouping mechanism is used. (i.e., SampleGroupDescript of type 'trif') If your ionBox version is 2 or higher, you can use the Sample eToGroupBox is optional. Otherwise, SampleToGroupBox pBoxes(type 'nalm') and grouping_type_parame ter is 'trif' and SampleGroupDescriptionBox( Associates samples, NAL units, and rectangular regions via the A RectangularRegionGroupEntry describes the following: - one rectangular area, - Coding dependencies between this rectangle and other rectangles. Each RectangularRegionGroupEntry has a groupID and This identifier is used to identify the NAs in the sample. Associates an L unit with a particular RectangularRegionGroupEntry. It can be done. The luminance sample coordinates are used to identify the location and size of the rectangular region. RectangularRegion when used with a movie fragment GroupEntry is defined in Section 8.9.4 of ISO / IEC14496-12. Add a new SampleGroupDescrip to the track fragment box, like so: Define the duration of the movie fragment by defining a tionBox. However, if a RectangularRegionGroupEn is already defined, Track fragments with the same groupID as the try will have a Rectangular rRegionGroupEntry does not exist. The base region used in a RectangularRegionGroupEntry is The NAL units in the rectangular area associated with this rectangular area group entry belong to This is a picture that If there is any change in the size of the base region in successive samples (e.g., In the case of reference picture resampling (RPR) or SPS resizing), the samples are , which have different RectangularRegions that reflect the size of their respective base regions. onGroupEntry Should be associated with an entry. NAL units mapped to a single rectangular region are included in a VVC track as usual. It may be included in the VVC video stream or in a separate track called the VVC subpicture track. That's fine. 3.5.3. Samples in VVC Tracks Referencing VVC Subpicture Tracks How to reconstruct a picture unit from A sample of a VVC track is shown in the order of bullets, with access units including the following NAL units: It is decomposed into ● The AUD NAL unit (if any) in the sample (and the first NAL unit). The first sample in a series of samples associated with the same sample entry If so, the parameter set and SEI NA included in the sample entry L unit (if any). • NAL units present in the sample and up to the PH NAL unit. -Specified in the 'spor' sample group description entry that is mapped to this sample Temporally aligned (decoded) subpictures from each referenced VVC subpicture track in the order they were In the contents of the resolved sample (within the time of the implementation), VPS, DCI, SPS, PPS, AUD, P Excluding all H, EOS, and EOB NAL units, if any. Track references are It is decomposed as follows. Note 1: The referenced VVC subpicture track is associated with a VVC non-VCL track. If VVC subpicture track is configured, the decomposed samples of the VVC subpicture track are Contains the non-VCL NAL units (if any) of the track's time-aligned samples. • The NAL unit that follows the PH NAL unit in the sample. NOTE 2: The NAL units following the PH NAL unit in the sample are suffixed. Suffix SEI NAL unit, Suffix APS NAL unit, EOS NAL unit, EOB NAL unit, or after the last VCL NAL unit. It may contain reserved NAL units that are used for 'spor' sample group description entry's 'subp' track reference index is decomposed as follows: If the track reference points to the track ID of a VVC subpicture track, The block references are resolved into VVC subpicture tracks. ● Otherwise (the track reference points to the 'alte' track group), the track Resolves a track reference to one of the tracks in the 'alte' track group. If the track reference index value resolves to a specific track in the previous sample, In the current sample, it can be decomposed into one of the following: ● The same specific track, or The same 'alte' track containing a sync sample time-aligned with the current sample Any other track in the group. Note 3: VVC subpicture tracks in the same 'alte' track group ,To avoid decoding inconsistencies, other It is necessarily independent of the VVC subpicture track and is therefore subject to the following constraints: There is a match. ●All VVC subpicture tracks contain VVC subpictures. ● Subpicture boundaries are similar to picture boundaries. ● Turn off loop filtering across subpicture boundaries. The reader selects a set of subpicture IDs that are either the first selection or different from the previous selection. If you select a VVC subpicture track that contains a VVC subpicture with a value: The steps of: Check the 'spor' sample group description entry and verify the PPS or SPS NAL Determine whether the unit needs to be changed. Note: SPS can only be changed at the start of CLVS. NALs containing the 'spor' sample group description entry Start code emulation before, after, or inside the subpicture ID in the unit If it indicates the presence of a prevention byte, derive the RBSP from the NAL unit (i.e., (This will remove the start code emulation prevention byte.) The next step is to override After downloading, start code emulation prevention is performed again. The reader reads the bit positions and sub-bits in the 'spor' sample group entry. The length of the picture ID is used to determine which bits to overwrite, and the subpicture ID is Update to selected. When you first select a PPS or SPS subpicture ID value, the reader In the constructed access unit, the selected subpicture ID value is used for PPS or SPS It is necessary to rewrite each of them. ● The subpicture ID value of a PPS or SPS is the same as the PPS ID value or SPS ID When compared with a previous PPS or SPS (respectively) with a D value, the reader Copies of PPSs and SPSs (PPSs or SPSs with the same PPS or SPS ID value) Updated subpictures I, I, and PS, respectively (if they are not present in the access unit) Rewrite the PPS or SPS (respectively) with the D value into a reconstructed access unit It is necessary to 3.5.4. Subpicture Order Sample Groups 3.5.4.1. Definition This sample group is a VVC base track, i.e., a VVC subpicture track. Used in VVC tracks that have a 'subp' track that references each sample. A group description entry describes a sub-picture or slide of a coded picture. The index of a track reference of type 'subp' is the indicates one or more consecutive subpictures or slices in the order To easily rewrite the PPS or SPS in response to the selection of a subpicture, A pull group description entry may include: - Change the selected subpicture ID in a PPS or SPS NAL unit Instructions on whether to do it or not. - The length of the Subpicture ID syntax element (in bits). - The bit position of the Subpicture ID syntax element in the included RBSP. - Start code emulation before or within the subpicture ID Flag indicating whether the anti-spam byte is present. - The parameter set ID of the parameter set that contains the subpicture ID. Syntax aligned(8) class VvcSubpicOrderEntry()ex tends VisualSampleGroupEntry('spor') { unsigned int(1)subpic_id_info_flag; unsigned int(15)num_subpic_ref_idx; for(i=0;i <num_subpic_ref_idx;i++) unsigned int(16)subp_track_ref_idx; if(subpic_id_info_flag){ unsigned int(4)subpic_id_len_minus1; unsigned int(12)subpic_id_bit_pos; unsigned int(1)start_code_emul_flag; unsigned int(1)pps_subpic_id_flag; if(pps_subpic_id_flag) unsigned int(6)pps_id; else { unsigned int(1)sps_subpic_id_flag; unsigned int(4)sps_id; bit(1) reserved=0; } } } Semantics If subpic_id_info_flag is 0, SPS and / or PP The subpicture ID value provided in S is the subpicture ID value indicated in p_track_ref_idx. accurate for a collection of , so no rewriting of the SPS or PPS is required If subpic_info_flag is 1, it indicates SPS and / or PPS indicates the subpicture that corresponds to the set of subp_track_ref_idx values. This indicates that it should be rewritten as num_subpic_ref_idx is the number of subpicture tracks that the VVC track references. Indicates the number of reference indices for the track group of a track or sub-picture track. subp_track_ref_idx is the VVC track for each value of i. One or more subpictures or subframes to be included in the VVC bitstream reconstructed from the block. Specifies the 'subp' track reference index of the ith list of slices or slices. subpic_id_len_minus1+1 is the subpicture ID of a PPS or SPS Indicates the number of bits in the Header ID syntax element, either of which may be referenced by this structure. subpic_id_bit_pos is the bit in the referenced PPS or SPS RBSP. indicates the zero-based bit position of the first bit of the first subpicture ID syntax element in . If start_code_empul_flag is 0, the referenced PPS or The start code error occurs before or within the subpicture ID in an SPS NAL unit. Indicates that no emulation prevention bytes are present. If flag is 1, the sub- PPS or SPS NAL unit There may be a start code emulation prevention byte before or within the picture ID This indicates that... If pps_subpic_id_flag is 0, this sample group description error The PPS NAL unit applied to the sample mapped to the entry is a subpicture. Indicates that the ID syntax element is not included. pps_subpic_id_flag is 1. If so, the PP that applies to the samples mapped to this Sample Group Description Entry The S NAL unit contains a subpicture ID syntax element. The pps_id (if present) is the ID that is mapped to this sample group description entry. This indicates the PPS ID of the PPS that applies to the sample being analyzed. If pps_subpic_id_flag is present and is 0, this sample The PPS NAL unit that applies to the sample mapped to the loop description entry is Indicates that the subpicture ID syntax element is not included and a subpicture ID value is inferred. If ps_subpic_id_flag is present and is 1, this sample group The SPS NAL unit that applies to the sample mapped to the sample description entry is Contains the Picture ID syntax element. The sps_id (if present) is mapped to this sample group description entry. This indicates the SPS ID of the SPS that applies to the sample being analyzed. 3.5.5. Subpicture Entity Group 3.5.5.1. General Indicates the conformance of merged bitstreams from multiple VVC subpicture tracks. A sub-picture entity group is defined that provides level information for the Note: The VVC base track is a separate track for merging VVC subpicture tracks. Provide a mechanism. The implicit reconstruction process requires modification of the parameter set. The group easily generates parameter sets for the reconstructed bitstream. Give guidelines so that you can do so. The coded sub-pictures to be jointly decoded within a group are replaced with each other. It is possible, i.e. the player can select sub-pictures per sample with the same level contribution. If you want to select multiple active tracks from a group of chats, use SubpicCommo When nGroupBox is jointly decoded, the resulting combination rules and lev Indicates el_idc. Encoded sub-pictures with different characteristics, e.g. different resolutions, are jointly decoded. If selected to be included, the SubpicMultipleGroupSBox When decoded with , the resulting combination rule and level_idc are shown. All entity_id values contained in a subpicture entity group are VV C Identifies the subpicture track. If present, the SubpicCommonGroup pBox and SubpicMultipleGroupSBox are movie-level It is included in the GroupsListBox in MetaBox, and is a file Not included in Bell or Track level MetaBoxes. 3.5.5.2. Subpicture common group box syntax aligned(8) class SubpicCommonGroupBox ex tends EntityToGroupBox('acgl',0,0) { unsigned int(32)level_idc; unsigned int(32)num_active_tracks; } 3.5.5.3. Subpicture common group box semantics level_idc is the num_active_track from the entity group If an entity is selected, indicate the level to which the entity conforms. num_active_tracks is the number of tracks that specify the value of level_idc Indicates the number of 3.5.5.4. Syntax for multiple subpicture group boxes aligned(8) class SubpicMultipleGroupsBox extends EntityToGroupBox('amgl',0,0) { unsigned int(32)level_idc; unsigned int(32)num_subgroup_ids; subgroupIdLen=(num_subgroup_ids>=(1<<24 )) ?32: (num_subgroup_ids>=(1<<16))?24: (num_subgroup_ids>=(1<<8))?16:8; for(i=0;i <num_entities_in_group;i++) unsigned int(subgroupIdLen)track_subgr oup_id[i]; for(i=0;i <num_subgroup_ids;i++) unsigned int(32)num_active_tracks[i]; } Semantics level_idc is an i in the range 0 to num_subgroup_ids-1 For all values of , select any num_activ from the subgroup with ID i. e_tracks[i] indicates the level at which the combination of selecting tracks is suitable. num_subgroup_id indicates the number of distinct subgroups, and each subgroup is They are identified by the same value of track_subgroup_id[i]. Different values of group_id[i] identify different subgroups. track_subgroup_id[i] is the i-th subgroup of this entity group Indicates the subgroup ID of the track. The subgroup ID value ranges from 0 to num_subg The range is from group_ids to 1 (inclusive). num_active_tracks[i] is listed in level_idc Indicates the number of tracks in the subgroup with ID i.
[0049] 4. Exemplary Technical Problems Addressed by the Disclosed Technical Solutions Regarding carriage of subpictures in multi-track VVC bitstreams The current design of the VVC video file format has the following problems: 1) One sample of the VVC subpicture track contains either: A) IS One or more consecutive decoded sequences as specified in O / IEC 23090-3 B) A complete subpicture, as specified in ISO / IEC 23090-3. One or more complete slices that form a rectangular region of the image and are consecutive in decoding order. However, the following problems exist. a. Also, VVC subpicture tracks, like tracks containing slices, are rectangular. It makes more sense to say that you have to cover the area. b. Subpictures and slices in a VVC subpicture track are motion constrained. It makes more sense to require that it be extractable or self-contained. It has become. c. The VVC sub-picture track contains the images that are consecutively decoded in the original bitstream. However, if this track itself is decoded, these subpictures will be It can contain a set of subpictures that form a rectangular area, such that the subpictures are sequentially For example, if the field of view (FOV) of a 360° image is too large for the left and right sides of the projected image, If the image is covered by some subpictures at the border of the Is it not allowed? 2) Samples of the VVC base track and the VVs referenced by the VVC base track C Reconstruct a PU from the time-aligned samples in the list of subpicture tracks. When a PH NAL unit is not present in the sample, the VVC base track The order of non-VCL NAL units in a sample is not explicitly specified. 3) The sub-picture order sample group mechanism ('spor') allows for the For the reconstructed bitstream, the subpictures from the subpicture track are If you want to enable a different order of SPS and / or PPS rewrites, However, it is unclear why any of this flexibility is needed. The 'spor' sample group mechanism is no longer needed and sample groups can be deleted. It is possible. 4) Samples of the VVC base track and the VVs referenced by the VVC base track C. Reconstructing a PU from time-aligned samples in a list of subpicture tracks , adding NAL units in time-aligned samples of the VVC subpicture track to the PU All VPS, DCI, SPS, PPS, AUD, PH, EOS, EOB NA If there are L units, exclude them. But what about OPI NAL units? What about EI NAL units? These non-VCL NAL units are subpicture tracks. Why is it allowed to exist in the block? If so, is it allowed to exist in the block? In your configuration, should you just let them pass through? 5) Add the container for the two subpicture entity group boxes to the movie level. However, if the file-level MetaBox contains a box, The entity_id value of the entity group references the track ID only if It is possible. 6) The Subpicture Entity Group is a group of entities in which the associated subpicture information is stored in the time domain of the track. This works if the length of the For example, different CVSs for a particular subpicture sequence may not be What happens if you have levels? If so, use sample groups instead. should carry essentially the same information, but for different samples (e.g., CVS) It should be possible for certain information to differ. 7) Each VVC base track currently contains a sample of subpicture order ('spor'). The existence of a sample group is mandatory. The algorithm uses subpicture tracking in the reconstructed bitstream for different samples. Enables different ordering of subpictures from the block and SPS and / or PPS rewriting However, if you want to use the 'subp' track of the VVC base track, Straightforward "early binding" of subpictures via a reference In this case, the 'spor' sample group is not needed.
[0050] 5. List of technical solutions In order to solve the above-mentioned problems, the following method is disclosed. These should be seen as illustrative examples of general concepts and should not be interpreted in a narrow sense. Furthermore, the present invention may be applied individually or in any combination. stomach. 1) In the VVC sub-picture track, one or more of the following items are proposed: a. If subpictures are included, one VVC subpicture track covers one rectangular area. It is necessary to cover it. b. Subpictures or slices in a VVC subpicture track do not share other regions. It can be extracted, decoded, and presented even if there are no overlying subpictures or slices. It is necessary for the movement to be restricted. i. Alternatively, a sub-picture or slice in a VVC sub-picture track , which allows relying on motion compensation of subpictures or slices that cover other regions, As a result, if there is no subpicture or slice that covers another area, The picture or slice cannot be extracted, decoded, or presented. c. One VVC subpicture track forms one rectangular area, but the original / whole One subpicture that is not consecutive in decoding order in a VVC bitstream It may contain a set of cha or slices. This allows you to see, for example, the left and right edges of the projection image in the original / entire VVC bitstream. At the right border, 360 is covered by sub-pictures that are not consecutive in decoding order. o picture The field of view (FOV) can be expressed by the VVC subpicture track. become. d. Subpicture or slice in each sample of the VVC subpicture track The order must be the same as their order in the original / entire VVC bitstream. It is necessary. e. Subpicture or slice in each sample of the VVC subpicture track Whether the decoding order is consecutive in the original / whole VVC bitstream Add instructions to show i. This indication may be, for example, in a VVC Base Track Sample Entry description, or or somewhere else. ii. In the original / entire VVC bitstream, the VVC subpicture track The order of subpictures or slices in each sample is consecutive in decoding order. If not indicated, the subpictures or slices in this track are It must not be merged with subpictures or slices in the VVC subpicture track. For example, in this example, the VVC base track reference is in track reference format. The 'subp' option allows you to select this VVC subpicture track and another VVC subpicture track. It is not permitted to reference both the . f.Add the flag nalusInContiguous to VvcNALUConfigBox Add DecodingOrderFlag. This flag is set to 1 when NAL units in a packet are consecutive in decoding order throughout the original bitstream. Therefore, track references of type 'subp' refer to VVC subpictures. A VVC base track that references a chat track can be used to connect to other VVCs through the same track reference. A value of 0 means that the NAL units in each sample are Note that they may or may not be consecutive in decoding order throughout the original bitstream. indicates a VVC subpicture track, and thus a track reference of type 'subp' A VVC base track that references a VVC base track can also reference other VVC subpictures through the same track reference. There is no need to refer to the track. 2) Samples of the VVC base track and the VVs referenced by the VVC base track C: Time-aligned samples in a list of subpicture tracks, from which track references When reconstituting PU by Regardless of the order of non-VCL NAL units in a sample of a VVC-based track The number is clearly identified. In one example, a NAL unit in a VVC subpicture track is preceded by a PU Set NAL units from the VVC base track samples to be placed in the following In the sample, nal_unit_type is EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD _NUT, RSV_NVCL_27, UNSPEC_30, or UNSPEC_31 If there is at least one NAL unit that is A NAL unit with a hexadecimal number precedes the first VCL NAL unit in a picture unit. (It is not possible to do this up to the first of these NAL units in a sample) Excludes NAL units, otherwise all NAL units in the sample. b. In one example, after the NAL units in the VVC subpicture track, Set NAL units from the VVC base track samples to be placed in the following It is specified as follows. nal_unit_type is EOS_NUT, EOB_NUT, SUFFIX_APS_NUT, SUFFIX_SEI_NUT, FD_NUT, RSV In samples that are _NVCL_27, UNSPEC_30, or UNSPEC_31 All NAL units. 3) Use 'subp' track references to allow VVC tracks to contain multiple (subpicture) tracks. The reference order is restarted from the referenced VVC subpicture track. 1 shows the decoding order of sub-pictures in a constructed bitstream. a.VVC Base Track Samples and VVC Base Track References Reconstructing PUs from time-aligned samples in the VC subpicture track list If so, the samples of the reference subpicture track are referenced in the 'subp' track reference. The VVC subpicture tracks are processed in the order of the VVC subpicture tracks they reference. 4) Subpicture track contains AU-level or picture-level non-VCL NAL units. NAL for AUD, DCI, OPI, VPS, SPS, PPS, PH, EOS, EOB SEI NA, which contains only unit, AU-level and picture-level SEI messages The presence of an AU-level SEI message is prohibited. Picture-level SEI messages apply to the entire U. Picture-level SEI messages apply to one or more entire pictures. Applies. a. In addition, samples of the VVC base track and the VVC base track referenced Recreates a PU from a list of time-aligned samples in a VVC subpicture track. When composing, all the time-aligned samples in the VVC subpicture track are All NAL units are sent to the PU without discarding specific non-VCL NAL units. will be added. 5) VVC base track samples and track references from VVC base tracks Temporally offset samples in a list of VVC subpicture tracks referenced via Removed the use of the 'spor' sample group when reconfiguring the PU from the or' Delete the description of the parameter set rewriting process based on the sample group. 6) Remove the 'spor' sample group specification. 7) Each 'subp' track reference index is a track reference index for a VVC subpicture track. Either the track ID of the VVC subpicture track group or the track group ID of the VVC subpicture track group. It specifies that the reference is made to the following and not to anything else. 8) To solve problem 5, we use the boxes of two subpicture entity groups. The container is defined as a file-level MetaBox as follows: cCommonGroupBox and SubpicMultipleGroupSBo x is the GroupsListB in the file-level MetaBox, if it exists It should be included in the MetaBox and not in any other level of MetaBox. 9) To solve problem 6, add two sample groups and two subpictures. This allows the VVC file to convey the same information as the entity group. The format is such that the associated subpicture information is not consistent throughout the duration of the track. In this case, for example, different CVSs may have different levels for a particular subpicture sequence. This will enable us to respond to situations such as 10) To solve issue 7, suggest one or more of the following: a. One "spor" sample group for each VVC base track It is specified to be selectable. b. When reconfiguring the PU, the 'spor' sample group is If not present in the 'subp' track, the reference subpicture track samples are The references are processed in the order of the VVC sub-picture tracks referenced in the references. 6. Implementation Below are some exemplary implementations for some of the inventive aspects summarized in Section 5 above. This is an embodiment that can be applied to the standard specifications of the VVC video file format. The test is based on the MPEG Output Document N19454 Final Draft Specification (Information Technology - Audio-Visual Encoding of Media Objects - Part 15: ISO Base Media File Formats Video carriage structured by Network Abstraction Layer (NAL) units, correction 2: ISO Based on the carriage of VVC and EVC in BMFF, July 2020. The most relevant additions or modifications are highlighted in bold and italics. , and the deleted part is marked with double brackets (e.g., [[a]] is (This indicates the deletion of the letter 'a'.) Because it is inherently editable, it is not highlighted. There may be some changes. 6.1. First embodiment This embodiment is items 1a, 1b, and 1c. 6.1.1. Track Types In this specification, the following types of video transports are used to carry VVC bitstreams: Specify the rack. a) VVC Truck: A VVC track contains NAL units in its samples and / or sample entries. By including, and possibly including, the 'vopi' and 'linf' sampling VVC bitstream via loop or via 'opeg' entity group Associating other VVC tracks with other layers and / or sublayers of the and possibly by referencing a VVC subpicture track, Represents a VVC bitstream. If a VVC track references a VVC subpicture track, this is called a VVC base track. A VVC Base Track shall not contain any VCL NAL units. , which are not referenced by a VVC track via a 'vvcN' track reference Let's say. b) VVC non-VCL tracks: A VVC non-VCL track is a track that contains only non-VCL NAL units and is vvcN" track reference by VVC track. Non-VCL tracks in VVC can have ALF, LMCS, or scaling list parameters. If the APS carrying the meter is not present along with other non-VCCL NAL units or other non-VC L NAL units and separate tracks from tracks containing VCL NAL units. may contain APS that may be stored in and transmitted via that track. . Non-VCL tracks of VVC may also contain APS NAL units, either with or without them. without, together with, or with other non-VCL NAL units The picture header NAL unit is recorded in a track separate from the track containing the picture header NAL unit, without any accompanying bits. It may also contain an APS that may be stored and transmitted over that track. c) VVC Subpicture Track: A VVC sub-picture track contains either:
[0051] [ka]
[0052] A sequence of one or more complete slices that form a single rectangular area. One sample of a VVC sub-picture track contains either:
[0053] [ka]
[0054] [[VVC subpicture or The slices are consecutive in decoding order.
[0055] [ka]
[0056] 6.1.2. Overview of rectangular regions carried in VVC bitstreams This specification helps to describe a rectangular area that consists of either:
[0057] [ka]
[0058] ... 6.2. Second embodiment This embodiment relates to items 2, 2a, 2b, 3, 3a, 4, 4a, and 5. 6.2.1. Samples in a VVC track referencing a VVC subpicture track How to reconstruct a picture unit from
[0059] [ka]
[0060] The AUD NAL unit and the first NAL unit present in the sample, if any L unit]].
[0061] [ka]
[0062] The first sample in a series of samples associated with the same sample entry If so, the parameter set and SEI NAL unit.
[0063] [ka]
[0064] [ka]
[0065] Note 2: The referenced VVC subpicture track is associated with a VVC non-VCL track. If VVC subpicture track is configured, the decomposed samples of the VVC subpicture track are If there are non-VCL NAL unit(s) of time-aligned samples in the track, Contains non-VCL NAL units.
[0066] [ka]
[0067] [Note 2: The NAL units following the PH NAL unit in the sample are Fix SEI NAL unit, Suffix APS NAL unit, EOS N AL unit, EOB NAL unit, or the last VCL NAL unit. It may contain reserved NAL units that are permitted. 'spor' sample group description entry's 'subp' track criteria entry Dex is broken down as follows: If the track reference points to the track ID of a VVC subpicture track, The block references are resolved into VVC subpicture tracks. ● Otherwise (the track reference points to the 'alte' track group), the track The track reference resolves to one of the tracks in the 'alte' track group, and the specific track If the track reference index resolves to a specific track in the previous sample, In the rule, it is decomposed into one of the following: ● The same specific track, or The same 'alte' track containing a sync sample time-aligned with the current sample Any other track in the group. Note 3: VVC subpicture tracks in the same 'alte' track group To avoid decoding inconsistencies, other VVCs referenced by the same VVC base track It is necessarily independent of the VC subpicture track and is therefore constrained as follows: There is. ●All VVC subpicture tracks contain VVC subpictures. ● Subpicture boundaries are similar to picture boundaries. ●[[Turn off loop filtering at subpicture boundaries. The reader selects a set of subpicture IDs that are either the first selection or different from the previous selection. If you select a VVC subpicture track that contains a VVC subpicture with a value: The steps of: Check the 'spor' sample group description entry and verify the PPS or SPS NAL Determine whether the unit needs to be changed. Note: SPS can only be changed at the start of CLVS. NALs containing the 'spor' sample group description entry Start code emulation before, after, or inside the subpicture ID in the unit If it indicates the presence of a prevention byte, derive the RBSP from the NAL unit (i.e., (This will remove the start code emulation prevention byte.) The next step is to override After downloading, start code emulation prevention is performed again. The reader uses the bit position and sub-bits in the "spor" sample group entry. The length of the picture ID is used to determine which bits to overwrite, and the subpicture ID is Update to selected. When you first select a PPS or SPS subpicture ID value, the reader In the constructed access unit, the selected subpicture ID value is used for PPS or SPS It is necessary to rewrite each of them. ● The subpicture ID value of a PPS or SPS is the same as the PPS ID value or SPS ID When compared with a previous PPS or SPS (respectively) with a D value, the reader Copies of PPSs and SPSs (PPSs or SPSs with the same PPS or SPS ID value) Updated subpictures I, I, and PS, respectively (if they are not present in the access unit) Rewrite the PPS or SPS (respectively) with the D value into a reconstructed access unit It is necessary to 6.3. Third embodiment This embodiment is items 1a, 1b, 1c, 1f, 2, 2a, 2b, 4, 4a, and 10. Truck type In this specification, the following types of video transports are used to carry VVC bitstreams: Specify the rack. d) VVC Track: A VVC track contains NAL units in its samples and / or sample entries. By including, and possibly including, the 'vopi' and 'linf' sampling VVC bitstream via loop or via 'opeg' entity group Associating other VVC tracks with other layers and / or sublayers of the and possibly by referencing a VVC subpicture track, Represents a VVC bitstream.
[0068] [ka]
[0069] e) VVC non-VCL tracks: A VVC non-VCL track is a track that contains only non-VCL NAL units, vvcN' track referenced by VVC track. Non-VCL tracks in VVC can have ALF, LMCS, or scaling list parameters. If the APS carrying the meter is not present along with other non-VCCL NAL units or other non-VC L NAL units and separate tracks from tracks containing VCL NAL units. may contain APS that may be stored in and transmitted via that track. . Non-VCL tracks of VVC may also contain APS NAL units, either with or without them. without, together with, or with other non-VCL NAL units The picture header NAL unit is recorded in a track separate from the track containing the picture header NAL unit, without any accompanying bits. It may also contain an APS that may be stored and transmitted over that track. f) VVC Subpicture Track: A VVC sub-picture track contains either:
[0070] [ka]
[0071] A sequence of one or more complete slices that form a single rectangular area. One sample of a VVC sub-picture track contains either:
[0072] [ka]
[0073] [[VVC subpicture or The slices are consecutive in decoding order.
[0074] [ka]
[0075] Overview of rectangular regions carried in VVC bitstreams This specification helps to describe a rectangular area that consists of either:
[0076] [ka]
[0077] Rectangular areas cover rectangles without holes. Rectangular areas within a picture do not overlap each other. ... Sample to Picture in a VVC track that references a VVC subpicture track How to reconfigure the unit
[0078] [ka]
[0079] The AUD NAL unit (if any) present in the sample and the first N AL unit)]].
[0080] [ka]
[0081] The first sample in a series of samples associated with the same sample entry If so, the parameter set and SEI NAL unit.
[0082] [ka]
[0083] [ka]
[0084] Note 2: The referenced VVC subpicture track is associated with a VVC non-VCL track. If VVC subpicture track is configured, the decomposed samples of the VVC subpicture track are If there are non-VCL NAL unit(s) of time-aligned samples in the track, Contains non-VCL NAL units.
[0085] [ka]
[0086] [ka]
[0087] [Note 2: The NAL units following the PH NAL unit in the sample are Fix SEI NAL unit, Suffix APS NAL unit, EOS N AL unit, EOB NAL unit, or the last VCL NAL unit. It may contain reserved NAL units that are permitted. 'spor' sample group description entry's 'subp' track criteria index The hex is decomposed as follows: If the track reference points to the track ID of a VVC subpicture track, The block references are resolved into VVC subpicture tracks. ● Otherwise (the track reference points to the 'alte' track group), the track The track reference resolves to one of the tracks in the 'alte' track group, and the specific track If the track reference index resolves to a specific track in the previous sample, In the rule, it is decomposed into one of the following: ● The same specific track, or The same 'alte' track containing a sync sample time-aligned with the current sample Any other track in the group. Note 3: VVC subpicture tracks in the same "alte" track group To avoid decoding inconsistencies, other VVCs referenced by the same VVC base track It is necessarily independent of the VC subpicture track and is therefore constrained as follows: There is. ●All VVC subpicture tracks contain VVC subpictures. ● Subpicture boundaries are similar to picture boundaries. Turn off loop filtering at subpicture boundaries. The reader selects a set of subpicture IDs that are either the first selection or different from the previous selection. If you select a VVC subpicture track that contains a VVC subpicture with a value: The steps of: Check the 'spor' sample group description entry and verify the PPS or SPS NAL Determine whether the unit needs to be changed. Note: SPS can only be changed at the start of CLVS. NALs containing the 'spor' sample group description entry Start code emulation before, after, or inside the subpicture ID in the unit If it indicates the presence of a prevention byte, derive the RBSP from the NAL unit (i.e., (This will remove the start code emulation prevention byte.) The next step is to override After downloading, start code emulation prevention is performed again. The reader reads the bit positions and sub-bits in the 'spor' sample group entry. The length of the picture ID is used to determine which bits to overwrite, and the subpicture ID is Update to selected. When you first select a PPS or SPS subpicture ID value, the reader In the constructed access unit, the selected subpicture ID value is used for PPS or SPS It is necessary to rewrite each of them. ● The subpicture ID value of a PPS or SPS is the same as the PPS ID value or SPS ID When compared with a previous PPS or SPS (respectively) with a D value, the reader Copies of PPSs and SPSs (PPSs or SPSs with the same PPS or SPS ID value) Updated subpictures I, I, and PS, respectively (if they are not present in the access unit) Rewrite the PPS or SPS (respectively) with the D value into a reconstructed access unit It is necessary to Sample entry name and format (VVC video stream definition) definition ... A VVC track may contain a 'subp' track reference, the entry track_ID value of VVC subpicture track or 'a' of VVC subpicture track lte' track_group_id value of the track group. [[If a VVC track contains a 'subp' track reference, it is a VVC base track is called, and the following applies: - Samples of a VVC track shall not contain VCL NAL units.
[0088] [ka]
[0089] ... syntax
[0090] [ka]
[0091] Semantics Compressor class in VisualSampleEntry e is the compressor used when the value "\012VVC Coding" is recommended (\012 is 10, the length of the string is bytes). VvcDecoderConfigurationRecord is defined in 11.3.3 It is defined as follows.
[0092] [ka]
[0093] lengthSizeMinusOne plus 1 is VvcNALUConf The byte length of the NALUnitLength field in the track containing the igBox. The value of this field corresponds to the encoded length in 1, 2 or 4 bytes, respectively. The value is one of 0, 1 or 3. [num_subpics_minus1+1] VVC subpicture tracks Specifies the number of subpicture sequences to be included. subpic_id, the sequence of the subpicture contained in the VVC subpicture track Specifies the subpicture identifier of the image.
[0094] FIG. 1 illustrates an exemplary video processing system 1 in which various techniques disclosed herein may be implemented. 900. Various implementations may be implemented using one of the modules of the system 1900. The system 1900 may include an input for receiving video content. The video content may include a raw or uncompressed format, For example, it may be received as 8 or 10 bit multi-module pixel values, or may be compressed or The input unit 1902 may receive the image data in a coded format. It may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces are Ethernet, Passive Optical Networks (PO N), and Wi-Fi or cellular interfaces. This includes wireless interfaces such as wireless LAN.
[0095] The system 1900 may implement various encoding or encoding methods described herein. The encoding module 1904 may include an encoding module 1904 that can encode an input The average bit rate of the video from the input unit 1902 is output to the encoding module 1904. This encoding technique may thus be referred to as video compression or This is sometimes called video transcoding technology. The output of the encoding module 1904 is The data may be stored in the memory or transmitted via a connected communication channel, as represented by module 1906. The data received, stored or communicated at the input unit 1902 may be transmitted. The bitstream (or coded) representation of the video is used by module 1908. to generate pixel values or displayable images that are sent to the display interface unit 1910. The process of generating a user-viewable video from a bitstream representation is Furthermore, certain video processing operations are sometimes called "encoding." " operations or tools, while the encoding tools or operations are the encoder and its corresponding It is understood that a decoding tool or operation is performed by the decoder that reverses the result of the decoding. Let's do it.
[0096] An example of a peripheral bus interface unit or display interface unit is the Universal USB Serial Bus (USB) or High-Definition Multimedia Interface (HDMI (registered) Examples of storage interfaces include DisplayPort, DisplayPort, etc. Serial Advanced Technology Attachment (SATA), PCI, IDE The technology described herein is applicable to mobile phones, laptops, smartphones, etc. smartphone, or other device capable of digital data processing and / or video display. It may be implemented in a variety of electronic devices.
[0097] 2 is a block diagram of a video processing device 3600. The device 3600 is The device 3600 may be used to implement one or more of the methods described above. It may be implemented on tablets, computers, Internet of Things (IoT) receivers, etc. The device 3600 includes one or more processing units 3602, one or more memories 3604, and a video and image processing hardware 3606. The one or more processing units 3602 may include: , may be configured to implement one or more of the methods described herein. Numerical value 3604 is a device used to implement the methods and techniques described herein. The video processing hardware 3606 may be used to store data and codes. The techniques described herein may also be used to implement in hardware circuits. In some embodiments, video processing hardware 3606 includes a processing unit 3602, e.g. It may be at least partially contained in a graphics co-processor.
[0098] FIG. 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. Figure.
[0099] As shown in FIG. 4, the video encoding system 100 includes a source device 110 and a destination device 111. The source device 110 may include a video encoding device. The destination device 120 generates encoded video data that can be transmitted to the source device 110. The encoded video data generated by the device may be referred to as a video decoding device. stomach.
[0100] The source device 110 includes a video source 112, a video encoder 114, and an input / output (I and a 116 (I / O) interface.
[0101] The video source 112 may be a source such as a video capture device, a video content provider, or the like. Interface for receiving video data from and / or generating video data a computer graphics system for The video data may include one or more pictures. 4 encodes the video data from the video source 112 and generates a bitstream. The bit stream may include a sequence of bits that form a coded representation of the video data. A bitstream may contain coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data is stored in a sequence It may contain parameter sets, picture parameter sets, and other syntax structures. The O interface 116 may include a modulator-demodulator (modem) and / or a transmitter. The encoded video data is transmitted to the I / O interface 11 via the network 130a. 6 to the destination device 120. It may be stored in the storage medium / server 130b for access by the destination device 120. good.
[0102] The destination device 120 includes an I / O interface 126, a video decoder 124, and A display device 122 may be included.
[0103] The I / O interface 126 may include a receiver and / or a modem. The interface 126 receives the signal from the source device 110 or the storage medium / server 130b. The video decoder 124 may obtain the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be an external display. may be external to the destination device 120 configured to interface with the device .
[0104] The video encoder 114 and the video decoder 124 are compliant with the High Efficiency Video Coding (HEVC) standard. standards, the Universal Video Coding (VVVM) standard, and other current and / or future standards. It may operate in accordance with the video compression standard.
[0105] FIG. 5 is a block diagram showing an example of a video encoder 200. 00 may be the video encoder 114 in the system 100 shown in FIG.
[0106] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the embodiment of FIG. 5, the video encoder 200 includes several functional modules. The techniques described in this disclosure are shared among various modules of video encoder 200. In some examples, the processor may include any of the techniques described in this disclosure. Or it may be configured to perform all of them.
[0107] The functional modules of the video encoder 200 include a splitting unit 201, a predication The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206 , residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit a buffer 213, an inverse transform unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an It may also include an entropy encoding unit 214 .
[0108] In other examples, video encoder 200 may have more, less, or different functionality. In one example, the prediction unit 202 may include an IBC (Integrated Block Code) module. An IBC unit may contain at least one In IBC mode, one of the reference pictures is the picture in which the current video block is located. Predication can be performed.
[0109] In addition, some modules such as a motion estimation unit 204 and a motion compensation unit 205 The modules may be highly integrated, but for purposes of illustration, are represented separately in the example of FIG. It is being done.
[0110] The division unit 201 may divide a picture into one or more video blocks. The video encoder 200 and the video decoder 300 support a variety of video block sizes. It can be done.
[0111] The mode selection unit 203 may, for example, select an intra-coding mode based on the error result. or select one of the inter-coding modes to encode the resulting intra-coded block or or inter-coded blocks to a residual generation unit 207, and the residual block data is and provides it to the reconstruction unit 212, and uses the coding block as a reference picture. In some examples, the mode select unit 203 may be reconfigured for inter-prediction. Combined intra and inter prediction based on measured signal and intra prediction signal The mode selection unit 203 may select the CIIP mode. For inter-motion, the resolution of the motion vectors of the blocks (e.g., sub-pixel or integer) You may also select a resolution of 1000x1000 pixels.
[0112] To perform inter prediction on the current video block, motion estimation unit 204 By comparing the current video block with one or more reference frames from buffer 213 , may generate motion information for the current video block. The movement of pictures from buffer 213 other than the picture associated with the current video block. A prediction video block for the current video block based on the received information and the decoded samples. may be determined.
[0113] The motion estimation unit 204 and the motion compensation unit 205 determine whether the current video block is an I-sequence. Depending on whether it is a slice, P slice, or B slice, for example, Different operations may be performed on the current video block.
[0114] In some examples, motion estimation unit 204 may perform unidirectional prediction for the current video block. and motion estimation unit 204 estimates the motion of the reference video block relative to the current video block. The reference picture in list 0 or list 1 for the current video block is searched for. Then, the motion estimation unit 204 selects the reference video block in list 0 or list 1. and a reference video block including a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may generate a reference index that points to the reference picture. The motion vector, the prediction direction indicator, and the motion vector of the current video block are used to The motion compensation unit 205 outputs the motion information of the current video block as A prediction video block for the current block may be generated based on the reference video block.
[0115] In another example, motion estimation unit 204 may bidirectionally predict the current video block. Preferably, the motion estimation unit 204 selects a reference picture from the list 0 for the current video block. A reference video block for locking may be searched for, and the reference in Listing 1 Searching for another reference video block in the picture to find the current video block Next, motion estimation unit 204 divides List 0 and List 1 into the reference video blocks. a reference index indicating the reference picture in the current video block and a reference video block in the current video block. The motion estimation unit 204 may generate a motion vector that indicates the spatial displacement of the current image. The reference index and motion vector of the current video block are calculated based on the motion information of the current video block. The motion compensation unit 205 outputs the reference motion information of the current video block. A prediction video block for the current video block is generated based on the reference video block.
[0116] In some examples, the motion estimation unit 204 may calculate the motion vector for the decoding process of the decoder. The full set of information may be output.
[0117] In some examples, the motion estimation unit 204 may generate a full set of motion information for the current picture. Rather, motion estimation unit 204 may output a motion vector of another video block. The motion information of the current video block may be signaled by referring to the motion information. The estimation unit 204 calculates the motion information of the current video block by comparing it with the motion information of the neighboring video blocks. may be determined to be sufficiently similar to
[0118] In one example, the motion estimation unit 204 calculates the syntax associated with the current video block. In the structure, the current video block is shown to have the same motion information as another video block. The value shown to the image decoder 300 may be:
[0119] In another example, the motion estimation unit 204 may estimate the structure associated with the current video block. In the sentence structure, different video blocks and motion vector differences (MVDs) may be identified. The motion vector difference is the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 calculates the motion vector of the indicated video block. The vector and the difference of the motion vector are used to determine the motion vector of the current video block. That's fine.
[0120] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 are: Includes AMVP and merge mode signaling.
[0121] Intra prediction unit 206 may perform intra prediction on the current video block. If intra prediction unit 206 intra predicts the current video block, The traprediction unit 206 uses the decoded samples of other video blocks in the same picture. The prediction data may be generated for the current video block based on the The prediction data for a block may include predicted video blocks and various syntax elements. stomach.
[0122] The residual generation unit 207 generates a predicted residual of the current video block from the current video block. By subtracting the video block(s) (e.g., indicated by a minus sign) (which is the case for the current video block), residual data for the current video block may be generated. The residual data of the block corresponds to different sample components of the samples in the current video block. The residual video block may include a residual video block.
[0123] In another example, for example in skip mode, the residual for the current video block The data may be missing and the residual generation unit 207 may not perform the subtraction operation.
[0124] Transform processing unit 208 performs a transform on the residual video block associated with the current video block. Applying one or more transformations to generate one or more transformation functions for the current video block Several video blocks may be generated.
[0125] The transform processing unit 208 generates the transform coefficient video block associated with the current video block. After generating , the quantization unit 209 generates one or more The transformation coefficient associated with the current video block is calculated based on the quantization parameter (QP) value of the Several image blocks may be quantized.
[0126] Inverse quantization unit 210 and inverse transform unit 211 apply inverse quantization to the transform coefficient video block. and applying the inverse transform to the residual video block. The reconstruction unit 212 may generate the predication signal generated by the predication unit 202. a residual video block reconstructed from corresponding samples from one or more prediction video blocks; In addition, a reconstructed image block associated with the current block is generated and stored in buffer 213. It may be memorized.
[0127] After reconstruction unit 212 reconstructs the video blocks, To reduce filtering artifacts, a loop filtering operation may be performed.
[0128] The entropy coding unit 214 is a separate module from other functional modules of the video encoder 200. The entropy encoding unit 214 may receive data from and the entropy encoding unit 214 performs one or more entropy encoding operations; Generating entropy-encoded data and a bitstream including the entropy-encoded data may be output.
[0129] FIG. 6 is a block diagram showing an example of a video decoder 300. may be the video decoder 114 in the system 100 shown in FIG.
[0130] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. In the embodiment of FIG. 6, the video decoder 300 includes several functional modules. The techniques described in this disclosure may be shared among various modules of the video decoder 300. In some examples, the processor may implement any or all of the techniques described in this disclosure. It may be configured to perform all of the above.
[0131] In the embodiment of FIG. 6, the video decoder 300 includes an entropy decoding unit 301; A motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 306, an inverse transformation unit 308, an inverse transformation unit 309, an inverse transformation unit 310, an inverse transformation unit 311, an inverse transformation unit 312, an inverse transformation unit 313, an inverse transformation unit 314, an inverse transformation unit 315, an inverse transformation unit 3 The video data includes a conversion unit 305, a reconstruction unit 306, and a buffer 307. The coder 300 may, in some examples, be any of the codes described with respect to the video encoder 200 (FIG. 5). A decryption pass may be performed that is roughly the reverse of the encryption pass.
[0132] The entropy decoding unit 301 may retrieve the encoded bitstream. The coded bitstream is entropy coded video data (e.g., video data The entropy decoding unit 301 may include an entropy - Decode the coded video data, and extract the motion from the entropy-decoded video data. The compensation unit 302 receives the motion vector, the motion vector precision, the reference picture list index, and the The motion compensation unit 302 may determine motion information including the frame rate, frame rate, and other motion information. For example, AMVP and merge mode may be performed to determine such information.
[0133] The motion compensation unit 302 may generate motion compensated blocks, possibly The interpolation is based on an interpolation filter. The syntax elements include the An identifier for the interpolation filter to be used may also be included.
[0134] Motion compensation unit 302 is used by video encoder 200 during the encoding of a video block. Interpolated values for sub-integer pixels of the reference block using an interpolation filter such as The motion compensation unit 302 may calculate the motion vectors according to the received syntax information. The interpolation filter used by the decoder 200 is determined, and the interpolation filter is used to generate the predicted block. may be generated.
[0135] The motion compensation unit 302 uses a portion of the syntax information to generate a motion compensation signal for the encoded video sequence. used to encode the frame(s) and / or slice(s) of Block size, how each macroblock of a picture in a coded video sequence is partition information that describes how the image is divided, and a mode that indicates how each partition is coded. , each of one or more reference frames (and reference frame lists) between inter-coded blocks , and other information for decoding the encoded video sequence.
[0136] The intra prediction unit 303 may, for example, predict the intra frames received in the bitstream. Prediction modes may be used to form a prediction block from spatially adjacent blocks. The inverse quantization unit 303 is provided to the bitstream, and the entropy decoding unit The quantized video block coefficients decoded by 301 are dequantized, i.e., dequantized. The inverse transform unit 303 applies the inverse transform.
[0137] The reconstruction unit 306 combines the residual block with the motion compensation unit 202 or intra prediction. and the corresponding predicted block generated by the measurement unit 303 to obtain the decoded block. If desired, the decoded blocks may be filtered to remove block artifacts. A deblocking filter may be applied to filter the blocks. The video blocks are stored in a buffer 307, which is used for subsequent motion compensation / inversion. providing reference blocks for intra-prediction and decoding the decoded data for display on a display device; Generate video.
[0138] Next, preferred solutions are listed in some embodiments.
[0139] A first set of solutions is provided below. The following solutions are based on the previous chapter (e.g., item 1). 1 illustrates an exemplary embodiment of the discussed techniques.
[0140] 1. A system for storing visual media data and a bitstream representation of this visual media data. and converting between files in accordance with format rules, said files being a track containing data of sub-pictures of the visual media data, A visual media processing method, wherein the track rules define the syntax of the tracks.
[0141] 2. The formatting rules specify that the track covers a rectangular area; The method described in Solution 1.
[0142] 3. The format rule specifies that the subpictures or slices included in the track are Solution 1, which specifies that the data must be independently extractable, decodable, and presentable. The method described.
[0143] The following solutions represent exemplary implementations of the techniques discussed in previous sections (e.g., items 3 and 4). .
[0144] 4. A method for storing visual media data and a bitstream representation of this visual media data. and converting between files in accordance with format rules, said files being , a first track and / or one or more sub-picture tracks, The bit rules define the syntax of said track and / or said one or more subpicture tracks. A method for processing visual media is provided.
[0145] 5. The formatting rules define whether the track is to be included in the one or more subpicture tracks. The method of Solution 4, providing that the method includes a reference to
[0146] 6. The format rule assigns access units to the one or more subpicture tracks. Include non-video coding layer Network Abstraction Layer units at the block or picture level. Do not allow this, as described in Solution 4.
[0147] 7. The disallowed unit is a decoding capability information structure or a parameter set , or operation point information, or header, or end of stream, or end of picture The method described in Solution 6, including:
[0148] 8. The transforming generates a bitstream representation of the visual media data; and storing said bitstream representation into said file according to said formatting rules; and the method according to any one of Solutions 1 to 7, comprising:
[0149] 9. The conversion parses the file according to the formatting rules and A method according to any of Solutions 1-7, including restoring media data.
[0150] 10. A processing device configured to implement the method described in one or more of solutions 1 to 9. A video decoding device.
[0151] 11. A method for manufacturing a computer system, comprising: A video decoding device.
[0152] 12. A computer program product having computer code stored therein, said computer code being The code, when executed by a processing device, causes the processing device to perform the processing described in any one of Solutions 1 to 9. A computer program product implementing the method of
[0153] 13. Create a video file that conforms to the file format generated according to any of solutions 1-9. A computer-readable medium carrying a data stream representation.
[0154] 14. A method, apparatus or system as described herein.
[0155] The second set of solutions incorporates exemplary embodiments of the techniques discussed in the previous section (e.g., item 1). provide.
[0156] 1. visual media data and one or more bitstreams of said visual media data 1102, storing and converting a visual media file including one or more tracks; wherein the visual media data comprises one or more sub-pictures or a plurality of slices. and the visual media file comprises one or more pictures including and the formatting rules are used to store the one or more tracks. a track containing a sequence of said one or more sub-pictures is A method of processing video media data (e.g., Figure 1) that specifies the covering of a rectangular area of a picture. 11) Method 110).
[0157] 2. The formatting rules define one or more sub-pictures or One or more slices cover another subpicture or another area different from the rectangular area. It specifies that there are no other slices and that they are independently extractable, decodable and presentable. Determine the method described in Solution 1.
[0158] 3. The formatting rules define one or more sub-pictures or One or more slices may be a sub-picture covering a different area than the rectangular area, or The method according to Solution 1, which defines a motion compensation dependency on another slice.
[0159] 4. The formatting rules specify that the one or more slices or sub-pictures are: The bitstreams stored in the tracks do not have to be consecutive in decoding order. The method described in Solution 1 is defined.
[0160] 5. 360 degrees covered by one or more sub-pictures that are not consecutive in decoding order The method according to Solution 1, wherein the field of view of the video is represented by this track.
[0161] 6. The format rule defines one or more subpictures in each sample of the track. or the order of one or more slices is different from that in the bitstream stored in said track. The interpretation specifies that the order of one or more subpictures or slices in the The method described in Solution 1.
[0162] 7. The format rule defines the one or more subpictures in each sample of the track. The decoding order of the track or the one or more slices is determined by the bits stored in the track. further specifying whether it includes an indication of whether it is contiguous in the stream; The method described in Solution 1.
[0163] 8. The instruction is included in a base track sample entry description of the track. The method described in Solution 7.
[0164] 9. The formatting rules, in response to the absence of the instruction, The one or more subpictures or slices can be transferred to another subpicture or another track. Additionally specify that merging into another slice of the block is not allowed, as described in Solution 7 How to do it.
[0165] 10. The instructions are included in a Network Abstraction Layer (NAL) configuration box. 7. The method described in Item 7.
[0166] 11. The indication being 1 indicates that the NAL unit are consecutive in the decoding order of the bitstream, and the track is referred to as a track reference. indicates that the base track referenced by , the method described in Solution 7.
[0167] 12. The indication being 0 indicates that the NAL unit indicates whether or not the are allowed to be consecutive in bitstream decoding order, and A base track that references the track by a track reference has the track reference. The method in Solution 7, which shows that you do not need to refer to the track.
[0168] 13. The visual media data is processed by Universal Video Coding (VVC), One or more tracks are VVC tracks, as described in any one of Solutions 1 to 12 Law.
[0169] 14. The conversion includes generating the visual media file and conforming to the formatting rules. storing said one or more bitstreams in said visual media file according to a rule; The method according to any one of Solutions 1 to 13, comprising:
[0170] 15. The conversion includes decrypting the visual media file according to the formatting rules. any of Solutions 1 to 13, comprising analyzing the one or more bitstreams and reconstructing the one or more bitstreams. 1. The method according to claim 1.
[0171] 16. Visual media data and one or more bitstreams of said visual media data 1102 to convert a visual media file including one or more tracks to and from a wherein the visual media data includes one or more sub-pictures or multiple slides. a processing device configured to implement a method including one or more pictures including a The visual media file stores the one or more tracks according to a format rule. The format rule may be a track containing a sequence of pictures covering a rectangular area of said one or more pictures. A processing device for video media data.
[0172] 17. The format rule defines the one or more sub-pictures in each sample of the track. The order of decoding of the image or said one or more slices is determined by the bit order stored in said track. The resolution specifies whether the resolution includes an indication of whether the resolution is continuous in the Measure 16. The device according to measure 16.
[0173] 18. A processing device comprising: a processor; A visual media file containing one or more tracks that store a video stream and is converted to and from a non-transitory computer-readable storage medium storing instructions for causing the visual media device to The data consists of one or more pictures containing one or more subpictures or slices. wherein the visual media file is encoded in the one or more tracks according to a formatting rule. and the formatting rules are stored in the one or more slices or the one or more sub-segments. A track containing a sequence of pictures covers a rectangular area of said one or more pictures. and a non-transitory computer-readable recording medium,
[0174] 19. The format rule defines the one or more sub-pictures in each sample of a track. The order of decoding of the image or said one or more slices is determined by the bit order stored in said track. The resolution specifies whether the resolution includes an indication of whether the resolution is continuous in the 19. The non-transitory computer-readable storage medium according to claim 18.
[0175] 20. Storing a bitstream generated by a method performed by a video processing device a non-transitory computer-readable recording medium for performing the method, the method comprising: Generate visual media files containing one or more tracks storing more than one bitstream wherein the visual media data includes one or more sub-pictures or a plurality of the visual media file includes one or more pictures containing slices of the format storing said one or more tracks according to a formatting rule, said formatting rule being A track containing a sequence of one or more subpictures or slices of said first a non-transitory computer-readable storage medium that defines a rectangular area of one or more pictures to be covered; body.
[0176] 21. The format rule defines the one or more sub-pictures in each sample of a track. The order of decoding of the image or said one or more slices is determined by the bit order stored in said track. The resolution specifies whether the resolution includes an indication of whether the resolution is continuous in the 19. The non-transitory computer-readable storage medium according to claim 18.
[0177] 22. A process configured to implement the method described in any one or more of solutions 1 to 15. A video processing device comprising a processing device.
[0178] 23. Storing visual media data in a file containing one or more bitstreams A method comprising the method according to any one of Solutions 1 to 15, and The method further includes storing the stream on a non-transitory computer-readable recording medium.
[0179] 24. When executed, the method according to any one or more of solutions 1 to 15 is performed on a processing device. A computer-readable medium having stored thereon program code for implementing the present invention.
[0180] 25. A computer storing a bitstream generated according to any of the above methods. Computer-readable medium.
[0181] 26. A method for implementing the method described in any one or more of solutions 1 to 15. A video processing device for storing the bitstream.
[0182] 27. Compliant with the file format generated according to any of solutions 1 to 15 A computer-readable medium carrying a bitstream.
[0183] 28. A method, apparatus or system as described herein.
[0184] The third set of solutions is based on the techniques discussed in previous chapters (e.g., items 3, 5, 6, 7, and 10). 1 illustrates an exemplary embodiment of the technique.
[0185] 1. A visual media data processing method (method 1200 shown in FIG. 12), comprising: visual media data and one or more bitstreams of said visual media data according to a bitstream rule; converting to and from a visual media file containing one or more tracks storing a stream; 1202, wherein the visual media file includes one or more sub-ports of the visual media data. A vector references one or more sub-picture tracks that store coding information for a vector picture. the format rule includes a sample in the base track and A program used to reconstruct a video unit from one or more subpicture tracks. A method for defining the process.
[0186] 2. The format rule is such that the base track is a subpicture track reference for referencing a track, The order of the one or more subpicture tracks referenced in a subpicture track reference is The sub-pictures in the video unit reconstructed from one or more sub-picture tracks The method described in Solution 1, showing the order of samples in the Kucha track.
[0187] 3. The format rule states that each subpicture track reference is a subpicture track reference. Track identification of a rack or track group identification of one subpicture track group The method of Solution 1, further specifying having an index pointing to either
[0188] 4. The format rule is such that a subpicture order sample group is included in the base track. The method according to Solution 1, which specifies that the
[0189] 5. The format rule is such that a sub-picture order sample group is included in the base track. If the sub-picture track referenced in the base track is not included in It is furthermore possible to use one or more subpicture track references when determining the order of the tracks. Define the method described in Solution 4.
[0190] 6. The formatting rules eliminate the use of sub-picture order sample groups; and A parameter set rewriting process is performed based on the sub-picture order sample group. 5. The method of claim 4, further comprising removing the statement:
[0191] 7. The formatting rules delete the specification of the sub-picture order sample group. 5. The method of claim 4, further comprising:
[0192] 8. The visual media data is processed by Universal Video Coding (VVC), and The method according to any one of Solutions 1 to 7, wherein one or more tracks are VVC tracks.
[0193] 9. The conversion generates the visual media file and the format rules. storing the one or more bitstreams in the visual media file according to and the method according to any one of Solutions 1 to 8, comprising:
[0194] 10. The conversion decodes the visual media file according to the formatting rules. Any of Solutions 1 to 8, comprising analyzing the one or more bitstreams and reconstructing the one or more bitstreams. 1. The method according to claim 1.
[0195] 11. A visual media data processing device that processes visual media data according to format rules. and one or more bitstreams of said visual media data. and converting the visual media file containing the track of The file contains encoded information for one or more sub-pictures of the visual media data. a base track that references one or more sub-picture tracks that store the The matte rules are based on the samples in the base track and one or more subpicture tracks. An apparatus that defines the process used to reconstruct a video unit from a rack.
[0196] 12. The format rule is such that the base track is a subpicture track reference for referencing a subpicture track; The order of the one or more subpicture tracks referenced in a subpicture track reference is The sub-picture tracks in the video unit reconstructed from the one or more sub-picture tracks 12. The apparatus of solution 11, further comprising: an apparatus for indicating an order of samples of a picture track.
[0197] 13. The format rule specifies that each subpicture track reference is a subpicture track reference. Track identification of a track or track group identification of one subpicture track group The apparatus of Solution 11, further specifying that the device has an index pointing to either .
[0198] 14. The format rule includes a sub-picture order sample group that is included in the base track. The device according to solution 11, wherein the device is optional for the command.
[0199] 15. The format rule includes a sub-picture order sample group that is included in the base track. If the sub-picture track is not included in the base track, It also supports the use of one or more subpicture track references when determining the order of the tracks. 15. The device according to solution 14, as defined in
[0200] 16. The formatting rules eliminate the use of sub-picture order sample groups, and and a parameter set rewriting process based on the sub-picture order sample group. The apparatus of solution 14, further providing for removing the description.
[0201] 17. The formatting rules delete the specification of the sub-picture order sample group. The apparatus according to solution 14, further providing that
[0202] 18. A non-transitory computer-readable recording medium, which is capable of being read by a processing device according to a format rule. Thus, visual media data and one or more bitstreams of said visual media data a visual media file including one or more tracks storing the visual The media file includes an encoding for one or more sub-pictures of the visual media data. The base track includes one or more sub-picture tracks that store information related to the base track. The format rule defines a sample in the base track and one or more subpictures. A non- prescribes the process used to reconstruct video units from the chatracks. A temporary computer-readable storage medium.
[0203] 19. The format rule is such that the base track is a subpicture track reference for referencing a subpicture track; The order of the one or more subpicture tracks referenced in a subpicture track reference is The sub-picture tracks in the video unit reconstructed from the one or more sub-picture tracks 19. The non-transitory computer program of solution 18, showing the order of samples in the picture track A readable recording medium.
[0204] 20. Storing a bitstream generated by a method performed by a video processing device a non-transitory computer-readable recording medium for performing a method according to a format rule; and one or more bitstreams of said visual media data. generating a visual media file including one or more tracks for storing said visual media file; A visual media file may include one or more sub-files of said visual media data that store encoded information. Refers to one or more sub-picture tracks that store coded information for a picture. the format rule includes a bass track containing a sample in the bass track; It is used to reconstruct a video unit from a video track and one or more subpicture tracks. A non-transitory computer-readable storage medium that defines a process to be performed.
[0205] 21. A process configured to implement the method described in any one or more of solutions 1 to 10. A video processing device comprising a processing device.
[0206] 22. Storing visual media data in a file containing one or more bitstreams A method comprising the method according to any one of solutions 1 to 10, and The method further includes storing the stream on a non-transitory computer-readable recording medium.
[0207] 23. When executed, a processing device performs the method described in any one or more of solutions 1 to 10. A computer-readable medium having stored thereon program code for implementing the present invention.
[0208] 24. A computer storing a bitstream generated according to any of the above methods. Computer-readable medium.
[0209] 25. A system configured to implement the method described in any one or more of solutions 1 to 10. A video processing device for storing the bitstream.
[0210] 26. Compliant with the file format generated according to any of solutions 1 to 10 A computer-readable medium embodying a bitstream representation.
[0211] 27. A method, apparatus or system as described herein.
[0212] In an exemplary solution, the visual media data corresponds to a video or image. In the solution described in the document, the encoder generates the coded representation according to the formatting rules. This allows for compliance with formatting rules. The decoder then interprets the coded representation according to the formatting rules, knowing whether or not a syntax element is present. This format is used to generate the decoded video by parsing the syntax elements in In the above solution, visual media data may be video or Corresponding to the image.
[0213] As used herein, the term "video processing" refers to video encoding, video decoding, video compression, or can refer to image decompression. For example, a video compression algorithm converts a pixel representation of an image into a It may be applied during conversion to a corresponding bitstream representation or vice versa. The bitstream representation of the current video block, as specified by the syntax, may be, for example: This may correspond to bits spread at the same or different locations in the bitstream. For example, one macroblock can be divided into two parts in terms of transformed and coded error residual values: and coded using bits in the header and other fields in the bitstream. Furthermore, during the conversion, the decoder may Based on the decision, you have the knowledge that some fields may or may not be present. Similarly, an encoder may parse a bitstream using a particular syntax. Determines whether a field should be included or not, and By including or excluding fields from the coded representation, The encoded representation may be generated according to:
[0214] The disclosed and other solutions, examples, embodiments, modules, The implementation of the rules and functional operations is within the scope of the structures disclosed herein and their structural equivalents. any digital electronic circuit, including computer software, firmware, or may be implemented in hardware, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products. , i.e., to be implemented by or to control the operation of a data processing apparatus. one of computer program instructions encoded on a computer-readable medium for controlling The computer-readable medium may be implemented as one or more modules. -readable storage device, machine-readable storage substrate, memory device, material that provides a machine-readable propagated signal The term "data processing device" may be a combination of the above, or one or more of these. The term may refer to, for example, a programmable processing device, a computer, or a plurality of processing devices, or any apparatus or device for processing data, including a computer; This device includes a computer program execution environment in addition to the hardware. code that creates the system, e.g., processor firmware, protocol stacks, database management the physical system, operating system, or a combination of one or more of these. A propagated signal may include an artificially generated signal, e.g., a machine-generated signal. A digital signal is an electrical, optical, or electromagnetic signal that encodes information for transmission to an appropriate receiving device. It is generated to
[0215] Computer programs (programs, software, software applications) , script, or code) is a language that is written in a compiled or interpreted It can be written in any form of programming language, including standard Modules suitable for use as standalone programs or in any computing environment Deploy in any form, including as modules, components, subroutines, or other units. A computer program does not necessarily have to access a file in a file system. A program may not respond to requests from other programs or files that hold data. recorded in a part (e.g., one or more scripts stored in a markup language document) The program may be stored in a single file dedicated to that program, or in multiple files. A coordination file (e.g., one or more modules, subprograms, or pieces of code) A computer program may be stored in a file (a file storing the program). A single computer located at a site or a communication network distributed across multiple sites It can also be deployed to run on multiple computers interconnected by do.
[0216] The processes and logic flows described herein operate on input data and produce output. execute one or more computer programs to perform functions by creating The processing and logic flow can be performed by one or more programmable processing units. It also includes application-specific logic circuits, such as FPGAs (Field Programmable Gate Arrays). This can be done by a chip array (chip array) or an ASIC (application specific integrated circuit), and the device can also It can also be implemented as special purpose logic circuitry.
[0217] Suitable processing devices for the execution of a computer program include, for example, both general purpose and special purpose microprocessors. Both the processor and any one or more processors of any kind of digital computer Generally, a processing unit may have a read-only memory or a random access memory. The essential elements of a computer are the a processing unit for executing instructions and one or more memory devices for storing instructions and data Generally, a computer has one or more mass storage devices for storing data. The device may include, or may be a magnetic, magneto-optical, or optical disk. to receive data from or transfer data to these mass storage devices. However, a computer may be operatively coupled to such a device. It is not necessary to have a computer suitable for storing computer program instructions and data. Computer readable media includes all forms of non-volatile memory, media, and memory devices. Includes, for example, EPROM, EEPROM, flash storage, magnetic disks, e.g. Internal hard disk or removable disk, magneto-optical disk, and CD-ROM and semiconductor storage devices such as DVD-ROM disks. may be supplemented by or incorporated into special purpose logic circuits. It may be included.
[0218] This patent specification contains many details, which may not be sufficient to encompass the scope of any subject matter or the scope of any claim. The present invention should not be construed as limiting the scope of the present invention, but rather as being specific to particular embodiments of particular technologies. The description of the features that may be present in the present patent document should be interpreted as a description of the features that may be present in the present patent document. Certain features described in the context may be implemented in combination in a single example. Conversely, various features that are described in the context of one example may be used in multiple embodiments. Further, the features may be implemented separately or in any suitable subcombination. The compounds described above as acting in specific combinations and originally claimed as such Although one or more features from a claimed combination may be combined, The claimed combination may be extracted from the set, and the claimed combination may be a subcombination or sub-combination. It may also be directed to variations of the combination.
[0219] Similarly, although operations may be shown in a particular order in the figures, this is not to be construed as a guarantee that a desired result will be achieved. that such actions be performed in the particular order or sequence shown, in order to It should not be understood as requiring that all actions be performed. Also, the separation of the various system components in the examples described in this patent specification It should not be understood that all embodiments require such separation.
[0220] Only some implementations and examples are described and illustrated in this patent document. Other embodiments, extensions, and variations are possible based on the content provided.
Claims
1. 1. A method of processing visual media data comprising converting between visual media data and a visual media file including one or more tracks storing one or more bitstreams of the visual media data, the method comprising: the visual media data comprises one or more pictures including one or more sub-pictures or one or more slices; the visual media file stores the one or more tracks according to a formatting rule; the visual media file includes a base track that references one or more sub-picture tracks; the format rules specify that the one or more slices or the track containing the sequence of one or more sub-pictures covers a rectangular area of the one or more pictures; the format rules specify an order of non-VCL (Video Coding Layer) NAL (Network Abstraction Layer) units within samples of the base track in reconstructing a video unit from samples of the base track and samples in the one or more sub-picture tracks; the format rules specify that at least some of the non-VCL NAL units in the base track are placed in the video units before or after NAL units in the one or more sub-picture tracks; 10. A method for processing visual media data, wherein the format rules specify an order of the non-VCL NAL units, regardless of the presence or absence of a picture header NAL within the sample.
2. 2. The method of claim 1, wherein the formatting rules specify that the one or more subpictures or one or more slices included in the track are independently extractable, decodable, and presentable without the presence of another subpicture or another slice covering another area different from the rectangular area.
3. The method of claim 1 , wherein the format rules specify that the one or more sub-pictures or one or more slices included in the track are motion-compensated dependent on another sub-picture or another slice covering another area different from the rectangular area.
4. The method of claim 1 , wherein the format rules specify that the one or more slices or the one or more sub-pictures do not have to be consecutive in decoding order of the bitstream stored in the track.
5. The method described in claim 1, wherein the track represents a 360-degree field of view of the image covered by the one or more sub-pictures that are not consecutive in decoding order.
6. 2. The method of claim 1, wherein the format rules specify that the order of the one or more subpictures or the one or more slices in each sample of the track is the same as the order of the one or more subpictures or the one or more slices in the bitstream stored in the track.
7. 2. The method of claim 1, wherein the format rules further specify whether the format rules include an indication indicating whether the decoding order of the one or more subpictures or the one or more slices in each sample of a track is consecutive in the bitstream stored in the track.
8. The method of claim 7 , wherein the indication is included in a description of a base track sample entry for the track.
9. 8. The method of claim 7, wherein the format rules further specify, in response to the absence of the instruction, that merging of the one or more subpictures or the one or more slices in the track with another subpicture or another slice in another track is not permitted.
10. The method of claim 7 , wherein the instructions are contained in a Network Abstraction Layer (NAL) configuration box.
11. 8. The method of claim 7, wherein the indication being 1 indicates that the NAL units in each sample of the track are consecutive in the decoding order of the bitstream, and that a base track that references the track with a track reference points to another track that has the track reference.
12. 8. The method of claim 7, wherein the indication being 0 indicates whether or not NAL units in each sample of the track are allowed to be consecutive in the decoding order of the bitstream, and indicates that a base track that references the track with a track reference may not reference other tracks that have the track reference.
13. The method of any one of claims 1 to 12, wherein the visual media data is processed by Universal Video Coding (VVC) and the one or more tracks are VVC tracks.
14. The method of any one of claims 1 to 13, wherein the converting comprises generating the visual media file and storing the one or more bitstreams in the visual media file according to the formatting rules.
15. The method of any one of claims 1 to 13, wherein the converting comprises parsing the visual media file in accordance with the formatting rules and reconstructing the one or more bitstreams.
16. Video processing device comprising a processing device configured to implement a method according to any one or more of the preceding claims.
17. A computer readable medium storing instructions which, when executed by a processing device, implement the method of any one of claims 1 to 15.
Citation Information
Patent Citations
Method, device, and computer program for generating timed media data
US20200245041A1
An apparatus, a method and a computer program for video coding and decoding
WO2020141248A1