Decoding capability information storage in video coding

By optimizing the track definition and signaling mechanism of the VVC video file format, the problems of storage and transmission efficiency of the VVC video file format were solved, and flexible multi-layer bitstream storage and accurate transmission of decoder configuration information were realized, thereby improving the transmission efficiency of video files and the adaptability of the decoder.

CN114205610BActive Publication Date: 2026-01-06FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111095947.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-17
Filing Date
2021-09-17
Publication Date
2026-01-06
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

The existing VVC video file format has problems with signaling notification and storage, including unclear definitions of VVC reference tracks and non-VCL tracks, overly simplistic storage limitations of APS NAL units, and imperfect signaling mechanisms for DCI and OPI NAL units, resulting in low efficiency in video file storage and transmission.

Method used

By redefining the storage method of VVC non-VCL tracks and VVC reference tracks, APS, image headers, and other non-VCL NAL units are allowed to be stored independently in different tracks. DCI and OPI NAL units are signaled at the track level or in sample entries, optimizing the definition of video elementary streams and non-VCL elementary streams and ensuring the proper use of DCI and OPI NAL units.

Benefits of technology

It achieves more efficient video file storage and transmission, supports flexible storage of multi-layer bitstreams and accurate transmission of decoder configuration information, and improves the transmission efficiency of video files and the adaptability of decoders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114205610B_ABST
    Figure CN114205610B_ABST
Patent Text Reader

Abstract

This application relates to decoding capability information storage in video coding, describes systems, methods, and apparatuses for encoding or decoding a file format storing one or more images. One example method includes performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, where the format rule specifies that a type of a sample entry determines whether a decoding capability information network abstraction layer unit is included in a sample entry of a video track in the visual media file or in a sample of a video track and a sample entry of the video track in the visual media file.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is made to promptly claim priority and interest in U.S. Provisional Patent Application No. 63 / 079,869, filed September 17, 2020, in accordance with applicable patent law and / or the rules of the Paris Convention. For all purposes under the law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to the generation, storage, and consumption of digital audio and video media information in file formats. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to process video or image representations according to file formats.

[0006] In one example aspect, a video processing method is disclosed. The method includes performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the format rules specify conditions for whether control information items are included in non-video codec layer tracks of the visual media file, and wherein the presence of non-video codec layer tracks in the visual media file is indicated by a specific track reference in the video codec layer tracks of the visual media file.

[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the format rules specify the type of sample entries that determines whether a decoding capability information network abstraction layer unit is included in the sample entries of the video track in the visual media file or in both the sample entries of the video track and the sample entries of the video track in the visual media file.

[0008] In another example, a video processing method is disclosed. The method includes performing a conversion between visual media data and a file storing information corresponding to the visual media data according to format rules; wherein the format rules specify a first condition for identifying the non-video codec layer (VCL) track of the file and / or a second condition for identifying the VCL track of the file.

[0009] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0010] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0011] In yet another example, a computer-readable medium storing code is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0012] In yet another example, a computer-readable medium storing a bitstream is disclosed. This bitstream is generated or processed using the methods described in this document.

[0013] These and other features are described throughout this document. Attached Figure Description

[0014] Figure 1 This is a block diagram of an example video processing system.

[0015] Figure 2 This is a block diagram of a video processing device.

[0016] Figure 3 This is a flowchart of an example method for video processing.

[0017] Figure 4 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0018] Figure 5 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0019] Figure 6 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0020] Figure 7 An example of an encoder block diagram is shown.

[0021] Figures 8 to 9 This is a flowchart of an example method for video processing. Detailed Implementation

[0022] For ease of understanding, chapter headings are used in this document, and the applicability of the techniques and embodiments disclosed in each chapter is not limited to that chapter. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Thus, the techniques described herein are also applicable to other video codec protocols and designs. In this document, edits to text are indicated relative to the current draft of the VVC specification or the ISOBMFF file format specification by double brackets (indicating that the text between the brackets is deleted text) (e.g., [[]]) and by bold italic text (indicating that the added text).

[0023] 1. Brief discussion

[0024] This document relates to video file formats. Specifically, it addresses the signaling notification and storage of the picture header (PH), adaptation parameter set (APS), decoding capability information (DCI), and operating point information (OPI) network abstraction layer (NAL) units in a Universal Video Codec (VVC) video bitstream within a media file based on the ISO-based Media File Format (ISOBMFF). These ideas can be applied individually or in various combinations to video bitstreams encoded and decoded by any codec (e.g., the VVC standard) and to any video file format (e.g., the VVC video file format under development).

[0025] 2. Abbreviation

[0026] ACT Adaptive Color Transformation

[0027] ALF Adaptive Loop Filter

[0028] AMVR Adaptive Motion Vector Resolution

[0029] APS Adaptive Parameter Set

[0030] AU Access Unit

[0031] AUD Access Unit Separator

[0032] AVC Advanced Video Codec (Rec.ITU-T H.264|ISO / IEC 14496-10)

[0033] B Two-way prediction

[0034] BCW utilizes bidirectional prediction with CU-level weights.

[0035] BDOF bidirectional optical flow

[0036] BDPCM (Block-based Incremental Pulse Codec Modulation)

[0037] BP buffer period

[0038] CABAC Context-Based Adaptive Binary Arithmetic Encoding and Decoding

[0039] CB codec block

[0040] CBR Constant Bit Rate

[0041] CCALF Cross-Component Adaptive Loop Filter

[0042] CPB image buffer

[0043] CRA (Clean Random Access)

[0044] CRC Cyclic Redundancy Check

[0045] CTB codec tree block

[0046] CTU (Codec Tree Unit)

[0047] CU encoding / decoding unit

[0048] CVS codec video sequence

[0049] DPB Decoding Image Buffer

[0050] DCI decoding capability information

[0051] DRAP relies on random access points

[0052] DU decoding unit

[0053] DUI Decoding Unit Information

[0054] EG Index Columbus

[0055] EGk k-order exponent Columbus

[0056] EOB bitstream end

[0057] EOS sequence ends

[0058] FD fill data

[0059] FIFO (First In First Out)

[0060] FL fixed length

[0061] GBR Green, Blue and Red

[0062] GCI General Constraint Information

[0063] GDR gradually decoded and refreshed

[0064] GPM Geometric Segmentation Mode

[0065] HEVC High-Efficiency Video Codec (Rec.ITU-T H.265|ISO / IEC 23008-2)

[0066] HRD Assumption Reference Decoder

[0067] HSS Assumption Flow Scheduler

[0068] I within the frame

[0069] IBC Intra-block Copy

[0070] IDR Instant Decoding and Refresh

[0071] ILRP interlayer reference image

[0072] Intra-IRAP Random Access Point

[0073] LFNST Low-Frequency Inseparable Transform

[0074] LPS least likely symbol

[0075] LSB (Least Significant Bit)

[0076] LTRP Long-Term Reference Image

[0077] LMCS Luminance Map with Chroma Scaling

[0078] MIP (Matrix-Based Intra-Frame Prediction)

[0079] MPS Maximum Possible Symbol

[0080] MSB Most significant bit

[0081] MTS Multiple Transformation Selection

[0082] MVP Motion Vector Prediction

[0083] NAL Network Abstraction Layer

[0084] OLS Output Layer Set

[0085] OP operation point

[0086] OPI Operation Point Information

[0087] P prediction

[0088] PH Image Header

[0089] POC Image Sequential Counting

[0090] PPS Image Parameter Set

[0091] PROF refines predictions using optical flow.

[0092] PT Image Time Sequence

[0093] PU Image Unit

[0094] QP Quantization Parameters

[0095] RADL Random Access Decodable Bootstrap (Image)

[0096] RASL Random Access Skip to Bootstrap (Image)

[0097] RBSP raw byte sequence payload

[0098] RGB red, green and blue

[0099] RPL Reference Image List

[0100] SAO Sample Adaptive Offset

[0101] SAR sample aspect ratio

[0102] SEI Supplemental Enhancement Information

[0103] SH strip head

[0104] SLI sub-image level information

[0105] SODB data bit string

[0106] SPS Sequence Parameter Set

[0107] STRP Short-Term Reference Image

[0108] STSA Stepwise Temporal Sublayer Access

[0109] TR (truncated rice)

[0110] VBR Variable Bit Rate

[0111] VCL (Video Codec Layer)

[0112] VPS Video Parameter Set

[0113] VSEI General Supplemental Enhancement Information (Rec.ITU-T H.274|ISO / IEC 23002-7)

[0114] VUI Video Availability Information

[0115] VVC Universal Video Codec (Rec.ITU-T H.266|ISO / IEC 23090-3)

[0116] 3. Introduction to Video Encoding and Decoding

[0117] 3.1. Video codec standards

[0118] Video codec standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Group (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software called the Joint Exploration Model (JEM). When the Universal Video Codec (VVC) project was officially launched, JVET was later renamed the Joint Video Experts Group (JVET). VVC is a new codec standard finalized by JVET at its 19th meeting, which concluded on July 1, 2020. Its goal is to reduce the bit rate by 50% compared to HEVC.

[0119] The Universal Video Coding (VVC) standard (ITU-T H.266|ISO / IEC 23090-3) and the associated Universal Supplemental Enhancement Information (VSEI) standard (ITU-T H.274|ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcasting, video conferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple codec video bitstreams, multi-view video, scalable layered codecs, and viewport-adaptive 360° immersive media.

[0120] 3.2. File Format Standards

[0121] Media streaming applications are typically based on IP, TCP, and HTTP transport methods and often rely on file formats such as ISO-based Media File Format (ISOBMFF). One such streaming system is HTTP-based Dynamic Adaptive Streaming (DASH). To use video formats with both ISOBMFF and DASH, a video-specific file format specification (such as AVC and HEVC) is required to encapsulate the video content within ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream (e.g., quality, layers, and levels, and much more) will need to be presented as file format-level metadata and / or a DASH media presentation description (MPD) for content selection purposes, such as selecting appropriate media segments for initialization at the start of a streaming session and for stream adaptation during the session.

[0122] Similarly, in order to use an image format with ISOBMFF, a file format specification specific to the image format will be required, such as the AVC image file format and the HEVC image file format.

[0123] The VVC video file format (a file format for storing VVC video content based on ISOBMFF) is currently being developed by MPEG.

[0124] The VVC image file format (a file format for storing image content encoded and decoded using VVC based on ISOBMFF) is currently being developed by MPEG.

[0125] 3.3. PH, APS, DCI, and OPI NAL units in VVC

[0126] Several new types of NAL units have been introduced into VVC, including PH, APS, DCI and OPI NAL units.

[0127] 3.3.1. Adaptive Parameter Set (APS)

[0128] The Adaptive Parameter Set (APS) conveys image-level and / or stripe-level information that can be shared across multiple stripes of an image and / or stripes from different images. However, this information can change frequently across images, and the total number of variations can be high, making it unsuitable for inclusion in the PPS. The APS includes three types of parameters: Adaptive Loop Filter (ALF) parameters, Luminance Map with Chroma Scaling (LMCS) parameters, and scaling list parameters. The APS can be carried in two different NAL unit types, as a prefix or suffix before or after the associated stripe. The latter can be helpful in ultra-low latency scenarios, for example, allowing the encoder to send the image stripe before generating ALF parameters based on the image, which will be used by subsequent images in the decoding order.

[0129] 3.3.2. Image Header (PH)

[0130] Each PU has a Picture Header (PH) structure. The PH exists in a separate PH NAL unit or is included in a stripe header (SH). If the PU consists of only one stripe, the PH can only be included in the SH. To simplify design, within a CLVS, the PH can only be entirely in a PH NAL unit or entirely in the SH. When the PH is in the SH, there is no PH NAL unit in the CLVS.

[0131] The PH is designed for two objectives. First, it helps reduce the signaling overhead of SHs for images containing multiple stripes by carrying all parameters with the same values ​​for all stripes of the image, thus avoiding the repetition of the same parameters in each SH. These include IRAP / GDR image indication, inter-frame / intra-frame stripe allow flags, and information related to POC, RPL, deblocking filter, SAO, ALF, LMCS, scaling list, QP increment, weighted prediction, codec block splitting, virtual boundary, and juxtaposed images. Second, it helps the decoder identify the first stripe of each codec image containing multiple stripes. Since there is one and only one PH for each PU, when the decoder receives a PH NAL unit, it can easily know that the next VCL NAL unit is the first stripe of the image.

[0132] 3.3.3. Decoding Capability Information (DCI)

[0133] A DCI NAL unit contains bitstream-level PTL information. A DCI NAL unit includes one or more PTL syntax structures that can be used during session negotiation between the sender and receiver of a VVC bitstream. When a DCI NAL unit is present in a VVC bitstream, each Output Layer Set (OLS) in the bitstream's CVS should conform to the PTL information carried in at least one PTL structure within the DCI NAL unit.

[0134] In AVC and HEVC, session negotiation PTL information is available in the SPS (for HEVC and AVC) and VPS (for HEVC layered extensions). This design of conveying session negotiation PTL information in HEVC and AVC has drawbacks because the scope of the SPS and VPS is within the CVS, not the entire bitstream. Therefore, sender-receiver session initiation may be re-initiated during each new CVS bitstream streaming. DCI solves this problem because it carries bitstream-level information, thus guaranteeing compliant decoding capabilities until the end of the bitstream.

[0135] 3.3.4. Operation Point Information (OPI)

[0136] Both HEVC and VVC decoding processes have similar input variables to set the decoding operation point via the decoder API: the target OLS and highest sublayer of the bitstream to be decoded. However, in scenarios where layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder API to the application, the decoder may not be correctly informed of the operation point for processing a given bitstream. Therefore, the decoder may fail to draw conclusions about the properties of images within the bitstream, such as the correct buffer allocation for the decoded image and whether to output a separate image. To address this issue, VVC adds a pattern indicating these two variables within the bitstream through a newly introduced Operation Point Information (OPI) NAL unit. The OPI NAL unit informs the decoder of the target OLS and highest sublayer of the bitstream to be decoded in the AU at the beginning of the bitstream and in its individual CVS.

[0137] In cases where OPI NAL units are present and operation points are also provided to the decoder via decoder API information (e.g., the application may have more updated information about the target OLS and sublayers), decoder API information takes precedence. In cases where there is no decoder API and no OPI NAL units in the bitstream, appropriate fallback options are specified in VVC to allow for correct decoder operation.

[0138] 3.4. Some details about the VVC video file format

[0139] 3.4.1. Types of Tracks

[0140] The VVC video file format specifies the following types of video tracks that carry VVC bitstreams in ISOBMFF files:

[0141] a) VVC track:

[0142] A VVC track represents a VVC bitstream by including NAL cells in its samples and sample entries, and possibly by referencing other VVC tracks that contain other sub-layers of the VVC bitstream, and possibly by referencing VVC subpicture tracks. When a VVC track references a VVC subpicture track, it is called a VVC reference track.

[0143] b) VVC non-VCL track:

[0144] APS and other non-VCL NAL units carrying ALF, LMCS, or scaling list parameters can be stored in and transmitted through a separate track from the track containing VCL NAL units; this is the VVC non-VCL track.

[0145] c) VVC sub-image track:

[0146] VVC sub-image tracks contain any of the following:

[0147] A sequence of one or more VVC sub-images.

[0148] A sequence of one or more complete stripes forming a rectangular region.

[0149] The sample points of the VVC sub-image track include any of the following:

[0150] One or more complete sub-images in succession according to the decoding order as specified in ISO / IEC 23090-3.

[0151] As specified in ISO / IEC 23090-3, it forms a rectangular area and consists of one or more complete stripes in the order of decoding.

[0152] VVC subpicks or stripes included in any sample point of a VVC subpick track are consecutive in the decoding order.

[0153] Note: VVC non-VCL tracks and VVC subpicture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each carry their own DASH representation, and for decoding and rendering subsets of tracks, the client can request the DASH representation containing the VVC subpicture track subset and the DASH representation containing the non-VCL tracks segment by segment. This avoids redundant transmission of APS and other non-VCL NAL units.

[0154] 3.4.2. VVC Basic Flow Structure

[0155] Three basic stream types are defined for storing VVC content:

[0156] The video basic stream does not contain any parameter set; all parameter sets are stored in one or more sample entries.

[0157] A video and parameter set basic stream can contain a parameter set, and can also store the parameter set in its sample entries or multiple sample entries;

[0158] Non-VCL elementary stream, containing non-VCL NAL units synchronized with the elementary stream carried in the video track.

[0159] Note: VVC non-VCL tracks do not include parameter sets in their sample entries.

[0160] 3.4.3. Decoder Configuration Information Sample Group

[0161] 3.4.3.1. Definition

[0162] The sample group description entry for this sample group contains DCI NAL units. All samples mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream.

[0163] This sample group indicates whether the same DCI NAL unit is used for different sample entries in a VVC track, i.e., whether samples belonging to different sample entries belong to the same VVC bitstream. When samples from two sample entries are mapped to the same decoder configuration information sample group description entry, the player can switch sample entries without reinitializing the decoder.

[0164] If any sample entry or memory is present in any DCI NAL unit, then that unit should be identical to the DCI NAL unit included in the decoder configuration information sample group.

[0165] 3.4.3.2. Syntax

[0166] class DecoderConfigurationInformation extends VisualSampleGroupEntry

[0167] ('dcfi'){

[0168] unsigned int(16)dciNalUnitLength;

[0169] bit(8*nalUnitLength)dciNalUnit;

[0170] }

[0171] 3.4.3.3. Semantics

[0172] dciNalUnitLength indicates the byte length of a DCI NAL unit.

[0173] dciNalUnit contains DCI NAL units as specified in ISO / IEC 23090-3.

[0174] 4. Example technical problems solved by the disclosed technical solutions

[0175] The latest design of the VVC video file format for signaling notifications regarding PH, APS, DCI, and OPI NAL units has the following issues:

[0176] 1) Neither the VVC reference track nor the VVC non-VCL track should contain VCL NAL elements. However, the current definition of the VVC non-VCL track will also apply to the VVC reference track. Furthermore, by the current definition, the VVC non-VCL track always contains APSNAL elements. However, this will not allow non-VCL NAL elements to contain image header NAL elements and possibly other non-VCL NAL elements, but not APS NAL elements.

[0177] Allowing such VVC non-VCL tracks would enable optimal storage of a single-layer bitstream based on extractable subpictures in the file for late-banding of subpicture tracks when different subpictures use different sets of APS, for example, by having a PH track (as a non-VCL track, although it contains the same information as the VVC baseline track), multiple APS tracks (as VVC non-VCL tracks), and multiple VVC subpicture tracks, each containing a sequence of subpictures.

[0178] 2) All APS NAL units are stored in a single VVC non-VCL track or VVC track. In other words, APS NAL units cannot be stored in more than one track. This applies to APS NAL units containing LMCS parameters (i.e., LMCS APS) or APS NAL units containing scaling list (SL) parameters (i.e., SL APS), but not to APS NAL units containing ALF parameters (i.e., ALF APS). Because different VVC subpicture tracks can use different sets of ALF APS, it is desirable to enable multiple VVC non-VCL tracks to carry ALF APS of the VVC bitstream.

[0179] 3) The definitions of video elementary streams and video and parameter set elementary streams do not consider DCI NAL units. Therefore, video elementary streams do not contain parameter sets, but may contain DCI NAL units.

[0180] 4) The definition of a non-VCL basic stream does not exclude the possibility of including VCL NAL units in a non-VCL basic stream.

[0181] 5) The decoder configuration information sample group provides a mechanism for signaling notification to the DCI NAL unit. However, the following problems exist:

[0182] a. In the most common use case, all samples of a track will belong to the same bitstream (or share the same DCI, regardless of the number of bitstreams). In this case, calculating the applicable DCI from the sample group signaling is complex.

[0183] b. It is said that all samples mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream. However, this does not allow samples belonging to multiple VVC bitstreams (e.g., determined by the EOB NAL unit) but in the same track to share the same DCI NAL unit, even when they could share.

[0184] 6) OPI NAL units are not allowed to be included in the sample entry description. However, in many cases, when OPI NAL units exist in the VVC bitstream, they should be treated similarly as parameter sets and therefore should be allowed to be included in the sample entry description.

[0185] 5. Example Solutions and Implementation Examples

[0186] To address the above and other issues, the following summarized methods are disclosed. These items should be considered as examples for explaining general concepts and should not be interpreted narrowly. Furthermore, these items can be applied individually or in any combination.

[0187] 1) To address problems 1 and 2, one or more of the following items are proposed:

[0188] a. VVC non-VCL orbitals are defined as orbitals containing only non-VCL NAL cells, and are referred to by VVC orbitals via the "vvcN" orbital reference.

[0189] b. Specifies that a VVC non-VCL track may contain an APS stored in a separate track from the track containing VCL NAL units and transmitted through that track, carrying ALF, LMCS, or scaling list parameters, with or without other non-VCL NAL units.

[0190] c. It is specified that VVC non-VCL tracks may also contain image header NAL units stored in a track separate from the track containing VCL NAL units and transmitted through that track, with or without APS NAL units, and with or without other non-VCL NAL units.

[0191] d. It is specified that the NAL unit of the video stream image header can be stored in samples of the VVC track or samples of the VVC non-VCL track, but not in both at the same time.

[0192] 2) To address problem 3, one or more of the following options were proposed:

[0193] a. A video elementary stream is defined as an elementary stream that contains VCL NAL units but no parameter sets, DCI, or OPI NAL units; all parameter sets, DCI, and OPI NAL units are stored in sample entries.

[0194] i. Alternatively, a video elementary stream is defined as an elementary stream that contains VCL NAL units and no parameter sets or DCINAL units; all parameter sets and DCI NAL units are stored in sample entries.

[0195] b. Treat the DCI NAL unit as exactly the same as the parameter set, that is, the DCI NAL unit may be in only the sample entry of the video track (e.g., when the sample entry type name is "vvc1"), or may be in one or both of the samples and sample entries of the video track (e.g., when the sample entry type name is "vvi1").

[0196] 3) To solve problem 4, it is specified that a non-VCL elementary stream is an elementary stream that contains only non-VCL NAL units, and these non-VCL NAL units are synchronized with the elementary stream carried in the video track.

[0197] 4) To solve problem 5, one or more of the following items were proposed:

[0198] a. If all samples of a track belong to the same bitstream (or share the same DCI, regardless of the number of bitstreams), the DCI NAL unit can be signaled in the track level box (e.g., track head box, track level meta box, or another track level box).

[0199] b. Samples belonging to multiple VVC bitstreams (e.g., determined by the EOB NAL unit) but in the same track are allowed to belong to the same decoder configuration information sample group and therefore share the same decoder configuration information sample group description entry.

[0200] 5) To address issue 6, OPI NAL units are allowed to be included in the sample entry description, for example, as one of the non-VCL NAL unit arrays in the decoder configuration record.

[0201] a. Alternatively, the OPI NAL unit can be considered exactly the same as the parameter set, i.e., the OPI NAL unit can be in only the sample entry of the video track (e.g., when the sample entry type name is "vvc1"), or it can be in one or both of the samples and sample entries of the video track (e.g., when the sample entry type name is "vvi1").

[0202] 6. Example

[0203] The following are some example embodiments of the inventions summarized in Section 5 above, which can be applied to the standard specification of the VVC video file format. The changed text is based on the latest draft specification. Most relevant sections that have been added or modified are indicated by bold italic text, and some deleted sections are indicated by double brackets (e.g., [[]]), where the deleted text between the brackets indicates the text that was deleted or canceled. There may be some other editable changes, which are therefore not highlighted.

[0204] 6.1. First Embodiment

[0205] This embodiment is for item 1.

[0206] 6.1.1. Types of Tracks

[0207] This specification defines the following types of video tracks for carrying VVC bitstreams:

[0208] a) VVC track:

[0209] A VVC track represents a VVC bitstream by including NAL units in its samples and / or sample entries, and possibly by being associated with other VVC tracks containing other layers and / or sublayers of the VVC bitstream by the “vopi” and “linf” sample groups or by the “opeg” entity group, and possibly by referring to a VVC subpicture track.

[0210] When a VVC orbital references a VVC sub-image orbital, it is also called a VVC reference orbital. A VVC reference orbital should not contain VCL NAL cells and should not be referred to by a VVC orbital via the "vvcN" orbital reference.

[0211] b) VVC non-VCL track:

[0212] A VVC non-VCL orbital is an orbital that contains only non-VCL NAL cells and is referred to by the VVC orbital via the "vvcN" orbital reference.

[0213] VVC non-VCL tracks may contain APSs stored in and transmitted through a track separate from the track containing VCL NAL units, carrying ALF, LMCS, or scaling list parameters, with or without other non-VCL NAL units.

[0214] VVC non-VCL tracks may also contain image header NAL units stored in and transmitted through a separate track from the track containing VCL NAL units, with or without APS NAL units, and with or without other non-VCL NAL units.

[0215] c) VVC sub-image track:

[0216] VVC sub-image tracks contain any of the following:

[0217] A sequence of one or more VVC sub-images.

[0218] A sequence of one or more complete stripes forming a rectangular region.

[0219] The sample points of the VVC sub-image track include any of the following:

[0220] One or more complete sub-images in succession according to the decoding order as specified in ISO / IEC 23090-3.

[0221] As specified in ISO / IEC 23090-3, it forms a rectangular area and consists of one or more complete stripes in the order of decoding.

[0222] VVC subpicks or stripes included in any sample point of a VVC subpick track are consecutive in the decoding order.

[0223] Note: VVC non-VCL tracks and VVC subpicture tracks enable optimal delivery of VVC video in streaming applications as follows: These tracks can each carry their own DASH representation, and for decoding and rendering of track subsets, the client can request the DASH representation containing the VVC subpicture track subset and the DASH representation containing the non-VCL tracks segment by segment. This avoids redundant transmission of APS and other non-VCL NAL units, and also avoids unnecessary subpicture transmission.

[0224] 6.2. Second Embodiment

[0225] This embodiment is for item 4.b.

[0226] 6.2.1. Decoder [[Configuration]] Capability Information Sample Group

[0227] 6.2.1.1. Definition

[0228] The sample group description entry for this sample group contains DCI NAL units. [[All samples mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream.]]

[0229] This sample group indicates whether the same DCI NAL unit is used for different sample entries in a VVC track [[, i.e., whether samples belonging to different sample entries belong to the same VVC bitstream]]. When samples from two sample entries are mapped to the same decoder configuration information sample group description entry, the player can switch sample entries without reinitializing the decoder.

[0230] If any sample entry or memory is in any DCI NAL unit, then that unit should be exactly the same as the DCI NAL unit included in the corresponding decoder configuration information sample group entry.

[0231] 6.2.1.2. Syntax

[0232] class DecoderConfigurationInformation extends VisualSampleGroupEntry

[0233] ('dcfi'){

[0234] unsigned int(16)dciNalUnitLength;

[0235] bit(8*nalUnitLength)dciNalUnit;

[0236] }

[0237] 6.2.1.3. Semantics

[0238] dciNalUnitLength indicates the byte length of a DCI NAL unit.

[0239] dciNalUnit contains DCI NAL units as specified in ISO / IEC 23090-3.

[0240] 6.3. Third Embodiment

[0241] This embodiment is for item 5.

[0242] 6.3.1. Definition of VVC decoder configuration record

[0243] This section specifies the decoder configuration information for ISO / IEC 23090-3 video content.

[0244] This record contains the dimensions of a length field in each sample point, indicating the lengths of its contained NAL cells, as well as the lengths of the parameter sets, DCI, OPI, and SEINAL cells, if stored in the sample point entry. This record is externally defined (its dimensions are provided by the structure containing it).

[0245] This record contains a version field. The specification defines version 1 for this record. Incompatible changes to the record will be indicated by changes to the version number. If the version number is not recognized, the reader should not attempt to decode the record or its applicable stream.

[0246] A compatible extension to this record will extend it without changing the configuration version code. Readers should be prepared to ignore unrecognized data that exceeds their understanding of the data definition.

[0247] When a track itself or through parsing a “subp” track reference contains a VVC bitstream, the VvcPtlRecord should exist in the decoder configuration record, and in this case, the specific set of output layers for the VVC bitstream is indicated by the field output_layer_set_idx. If ptl_present_flag is equal to zero in the track’s decoder configuration record, then the track should have an “oref” track reference. ...

[0249] Array sets exist to carry initialization non-VCL NAL cells. NAL cell types are restricted to indicating only DCI, OPI, VPS, SPS, PPS, prefix APS, and prefix SEI NAL cells. NAL cell types reserved in ISO / IEC 23090-3 and this specification may be defined in the future, and readers should disregard arrays with reserved or unpermitted values ​​for NAL cell types.

[0250] Note 2: This "tolerance" behavior is designed to prevent errors and allow for backward compatibility expansion of these arrays in future specifications.

[0251] Note 3: NAL cells carried in a sample entry are included in the access cell reconstructed from the first sample of the reference sample entry, immediately following the AUD and OPI NAL cells (if any), or are included at the beginning of that access cell.

[0252] It is recommended that the array be arranged in the following order: DCI, OPI, VPS, SPS, PPS, prefix APS, prefix SEI. ...

[0254] 6.3.2. Semantics of VVC Decoder Configuration Records ...

[0256] numArrays indicates the number of arrays of NAL cells of type (multiple).

[0257] array_completeness, when equal to 1, indicates that all NAL cells of a given type are in the following array and none are in the stream; when equal to 0, it indicates that additional NAL cells of the indicated type may be in the stream; [[default and]] allowed values ​​are constrained by the sample entry name.

[0258] NAL_unit_type indicates the type of NAL unit in the following array (which should be all of that type); it takes a value as defined in ISO / IEC 23090-3; it is restricted to taking one of the values ​​indicating DCI, OPI, VPS, SPS, PPS, prefix APS or prefix SEI[[, or suffix SEI]] NAL units.

[0259] `numNalus` indicates the number of NAL units of the indication type included in the configuration record of the stream to which this configuration record applies. The SEI array should contain only SEI messages of a "declarative" nature, i.e., those that provide overall information about the stream. An example of such an SEI could be a user data SEI.

[0260] nalUnitLength indicates the byte length of a NAL unit.

[0261] nalUnit includes DCI, OPI, VPS, SPS, PPS, APS or declarative SEI NAL units as specified in ISO / IEC 23090-3.

[0262] Figure 1 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0263] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Codec techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via a communication connection, as represented by component 1906. The bitstream (or codec) representation of the video received at input 1902, whether stored or communicated, can be used by component 1908 to generate pixel values ​​or transmit as displayable video to display interface 1910. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it will be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the codec results will be performed by the decoder.

[0264] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0265] Figure 2 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors(multiple) 3602 may be configured to implement one or more methods described herein. The memories(multiple) 604 may be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 606 may be used to implement some of the techniques described herein in a hardware circuit system. In some embodiments, the video processing hardware 3606 may be at least partially included in the processor 3602 (e.g., a graphics coprocessor).

[0266] Figure 4 This is a block diagram illustrating an example video codec system 100 that can utilize the techniques disclosed herein.

[0267] like Figure 4 As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, and this source device 110 may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110, and the target device 120 may be referred to as a video decoding device.

[0268] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0269] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec pictures and related data. A codec picture is a codec representation of a picture. Related data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via I / O interface 116 through network 130a. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0270] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0271] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interface with an external display device.

[0272] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or additional standards.

[0273] Figure 5 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 4The video encoder 114 in the system 100 shown.

[0274] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 5 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0275] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0276] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0277] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 5 The examples are shown separately.

[0278] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0279] The mode selection unit 203 can select one of the encoding / decoding modes (e.g., intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction modes (CIIP), where the prediction is based on the inter-frame prediction signal and the intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the block's motion vector (e.g., sub-pixel or integer pixel precision).

[0280] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0281] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0282] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image in list 0 or list 1, which contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0283] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference images in lists 0 and 1 containing the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0284] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process.

[0285] In some examples, the motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, the motion estimation unit 204 may refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0286] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0287] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0288] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling notification techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling Notification.

[0289] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0290] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0291] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0292] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0293] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0294] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in buffer 213.

[0295] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce the video block effect in the video block.

[0296] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0297] Figure 6 This is a block diagram illustrating an example of a video decoder 300, which can be... Figure 4 The video decoder 114 in the system 100 shown.

[0298] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 6 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0299] exist Figure 6 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 5 The encoding process described is the opposite of the decoding process.

[0300] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and based on the entropy-coded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by executing AMVP and Merge modes.

[0301] The motion compensation unit 302 can generate motion compensation blocks and can perform interpolation based on an interpolation filter. The identifier of the interpolation filter to be used at sub-pixel precision can be included in the syntax element.

[0302] The motion compensation unit 302 can use an interpolation filter, such as that used by the video encoder 200 during the encoding of a video block, to calculate the interpolation of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0303] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0304] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0305] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in the buffer 307 to provide a reference block for subsequent motion compensation / intra-frame prediction, and also generates the decoded video for presentation on the display device.

[0306] The following is a list of preferred solutions for some embodiments.

[0307] The following solutions illustrate example embodiments of the techniques discussed in the previous sections (e.g., items 1 through 4).

[0308] 1. A method for processing visual media (e.g., Figure 3 The method described in the text (3000) includes: performing a conversion between visual media data and a file storing information corresponding to the visual media data according to format rules (3002); wherein the format rules specify a first condition for identifying the non-video codec layer (VCL) track of the file and / or a second condition for identifying the VCL track of the file.

[0309] 2. The method according to Solution 1, wherein the first condition specifies that the non-VCL track contains only non-VCL network abstraction layer units and is identified in the VCL track by a specific track reference.

[0310] 3. The method according to solutions 1-2, wherein the first condition specifies that the non-VCL orbit contains an adaptive parameter set (APS) corresponding to the VCL orbit.

[0311] 4. The method according to any one of solutions 1-3, wherein the second condition of the VCL track specifies that the VCL track is not allowed to include decoding capability information (DCI) or operation point information (OPI) network abstraction units.

[0312] 5. The method according to Solution 1, wherein the first condition specifies that the non-VCL track includes one or more basic streams containing non-VCL network abstraction layer units, and wherein the non-VCL network abstraction layer units are synchronized with the basic streams in the VCL track.

[0313] 6. The method according to any one of solutions 1-5, wherein the conversion includes generating a bitstream representation of visual media data and storing the bitstream representation in a file according to format rules.

[0314] 7. The method according to any one of solutions 1-5, wherein the conversion includes parsing the file according to format rules to recover visual media data.

[0315] 8. A video decoding apparatus, comprising a processor configured to implement the method according to one or more of solutions 1 to 7.

[0316] 9. A video encoding apparatus, comprising a processor configured to implement the method according to one or more of solutions 1 to 7.

[0317] 10. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of solutions 1 to 7.

[0318] 11. A computer-readable medium on which a bitstream representation conforms to a file format generated according to any one of solutions 1 to 7.

[0319] 12. A method, apparatus or system described in this document.

[0320] In the solution described in this paper, the encoder conforms to the format rules by generating a codec representation based on those rules. In the solution described in this paper, the decoder can parse the syntax elements in the codec representation using the format rules, knowing whether or not syntax elements exist according to the format rules used to generate the decoded video.

[0321] Technology 1. A method for processing visual media data (e.g., Figure 8 The method described in the image (8000) includes: performing a conversion between a visual media file and a bitstream of visual media data according to format rules (8002), wherein the format rules specify conditions for whether control information items are included in non-video codec layer tracks of the visual media file, and wherein the presence of non-video codec layer tracks in the visual media file is indicated by a specific track reference in the video codec layer tracks of the visual media file.

[0322] Technique 2. The method according to Technique 1, wherein the condition specifies that the non-video codec layer track includes only non-video codec layer network abstraction layer units as information items.

[0323] Technique 3. The method according to any one of Techniques 1-2, wherein the condition specifies that the non-video codec layer track includes an adaptive parameter set as an information item, wherein the adaptive parameter set includes adaptive loop filter parameters, luminance mapping parameters with chroma scaling, or scaling list parameters, and wherein the condition specifies that the adaptive parameter set is stored in and transmitted through the track, which is separate from another track including a video codec layer network abstraction layer unit.

[0324] Technique 4. The method according to Technique 3, wherein the condition allows the non-video codec layer track to additionally include other types of non-video codec layer network abstraction layer units.

[0325] Technique 5. The method according to Technique 3, wherein the condition does not allow the non-video codec layer track to additionally include other types of non-video codec layer network abstraction layer units.

[0326] Technique 6. The method according to any one of Techniques 1-2, wherein the condition specifies that a non-video codec layer track includes a picture header network abstraction layer unit as an information item, and wherein the condition specifies that the picture header network abstraction layer unit is stored in and transmitted through the track, which is separate from another track that includes a video codec layer network abstraction layer unit.

[0327] Technique 7. The method according to Technique 6, wherein the condition allows the non-video codec layer track to additionally include other types of non-video codec layer network abstraction layer units.

[0328] Technique 8. The method according to Technique 6, wherein the condition does not allow the non-video codec layer track to additionally include other types of non-video codec layer network abstraction layer units.

[0329] Technique 9. The method according to Technique 6, wherein the condition allows non-video codec layer tracks to additionally include adaptive parameter set network abstraction layer units.

[0330] Technique 10. The method according to Technique 6, wherein the condition does not allow non-video codec layer tracks to additionally include adaptive parameter set network abstraction layer units.

[0331] Technical 11. The method according to any one of Technical 1-2, wherein the condition specifies that the image header network abstraction layer unit of the video stream is an information item stored in a first sample set of a track containing the video codec layer network abstraction layer unit or a second sample set of a non-video codec layer track, but not simultaneously stored in both.

[0332] Technique 12. The method according to any one of Techniques 1-11, wherein the conversion includes generating a visual media file according to format rules and storing a bitstream into the visual media file.

[0333] Technique 13. The method according to any one of Techniques 1-11, wherein the conversion includes generating a visual media file, and the method further includes storing the visual media file in a non-transitory computer-readable recording medium.

[0334] Technique 14. The method according to any one of Techniques 1-11, wherein the conversion includes parsing a visual media file according to format rules to reconstruct a bitstream.

[0335] Technique 15. The method according to any one of Techniques 1 to 14, wherein the visual media file is processed by a universal video codec (VVC), and wherein a non-video codec layer track or a video codec layer track is a VVC track.

[0336] Technology 16. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one or more of technologies 1 to 15.

[0337] Technique 17. A non-transitory computer-readable storage medium for storing instructions that cause a processor to perform the method according to any one or more of Techniques 1 to 15.

[0338] Technology 18. A video decoding apparatus, including a processor configured to implement the method according to one or more of technologies 1 to 15.

[0339] Technology 16. A video encoding apparatus, including a processor configured to implement the method according to one or more of technologies 1 to 15.

[0340] Technology 17. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of technologies 1 to 15.

[0341] Technology 18. A computer-readable medium on which a visual media file conforms to a file format generated according to any one of technologies 1 to 15.

[0342] Technology 19. A non-transitory computer-readable recording medium for storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is described in any one of technologies 1 to 15.

[0343] Technique 20. A method for generating a visual media file, comprising: generating a visual media file according to any one of Techniques 1 to 15, and storing the visual media file on a computer-readable program medium.

[0344] Implementation 1. A method for processing visual media data (e.g., Figure 9 The method described in the text (9000) includes: performing a conversion between a visual media file and a bitstream of visual media data according to format rules (9002), wherein the format rules specify the type of sample entry to determine whether the decoding capability information network abstraction layer unit is included in the sample entry of the video track in the visual media file or in the sample of both the video track and the video track in the visual media file.

[0345] Implementation Method 2. The method according to claim 1, wherein the format rule specifies that, in response to the sample entry type being vvc1, the decoding capability information network abstraction layer unit is included in the sample entry of the video track.

[0346] Implementation Method 3. According to the method of Implementation Method 1, wherein the format rule specifies that, in response to the sample entry type being vvi1, the decoding capability information network abstraction layer unit is included in the sample of the video track and the sample entry of the video track.

[0347] Implementation 4. According to the method of Implementation 1, wherein the format rules specify that the video basic stream in the visual media file includes a video codec layer network abstraction layer unit, wherein the format rules specify that the video basic stream in the visual media file is not allowed to include a parameter set or decoding capability information network abstraction unit, and wherein the format rules specify that the sample entries in the visual media file store the parameter set and the decoding capability information network abstraction unit.

[0348] Implementation 5. According to the method of implementation 4, wherein the format rules stipulate that the video basic stream in the visual media file is not allowed to include a parameter set, a decoding capability information network abstraction unit, or an operation point information network abstraction unit, and wherein the format rules stipulate that the sample entries in the visual media file store the parameter set, the decoding capability information network abstraction unit, and the operation point information network abstraction unit.

[0349] Implementation 6. A method for processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that, in response to a sample in the visual media file belonging to multiple common video codec bitstreams and in response to the sample being included in the same track, the sample is allowed to belong to the same decoder capability information sample group, and wherein the format rule specifies that all samples belonging to the same decoder capability information sample group share the same decoder capability information sample group description entry. In some embodiments, the format rule specifies that, in response to a sample in the visual media file belonging to multiple common video codec bitstreams and in response to the sample being included in the same track, the sample is allowed to belong to the same decoder capability information sample group, and wherein the format rule specifies that all samples belonging to the same decoder capability information sample group share the same decoder capability information sample group description entry.

[0350] Implementation 7. According to the method of Implementation 6, wherein the format rule specifies that, in response to all samples of a track belonging to the same bitstream or in response to all samples sharing the same decoding capability information regardless of the number of bitstreams, a decoding capability information network abstraction layer unit is indicated in the track-level box of the visual media file. In some embodiments, in response to all samples of a track belonging to the same bitstream or in response to all samples sharing the same decoding capability information regardless of the number of bitstreams, the format rule specifies that a decoding capability information network abstraction layer unit is indicated in the track-level box of the visual media file.

[0351] Implementation method 8. According to the method described in implementation method 7, wherein the track level box is a track head box, a track level element box, or another track level box.

[0352] Implementation 9. A method for processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that operation point information network abstraction layer units (NPIS) are allowed to be included in the visual media file as one of a plurality of non-video codec layer network abstraction layer unit arrays in a decoder configuration record. In some embodiments, the format rule specifies that operation point information network abstraction layer units (NPIS) are allowed to be included in the visual media file as one of a plurality of non-video codec layer network abstraction layer unit arrays in a decoder configuration record.

[0353] Implementation 10. A visual media processing method, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies the type of sample entry to determine whether an Operation Point Information Network Abstraction Layer (OPI) unit is included in: (1) a sample entry of a video track in the visual media file, or (2) a sample of a video track in the visual media file or a sample entry of a video track in the visual media file, or both. In some embodiments, the format rule specifies a second type of a second sample entry to determine whether an OPI unit is included in: (1) a second sample entry of a video track in the visual media file, or (2) a sample of a video track in the visual media file or a second sample entry of a video track in the visual media file, or both.

[0354] Implementation 11. The method according to Implementation 10, wherein the format rule specifies that, in response to the type of the sample entry being vvc1, the operation point information network abstraction layer unit is included in the sample entry of the video track. In some embodiments, the format rule specifies that, in response to the second type of the second sample entry being vvc1, the operation point information network abstraction layer unit is included in the second sample entry of the video track.

[0355] Implementation 12. The method according to Implementation 10, wherein the format rule specifies that: in response to the type of the sample entry being vvi1, the operation point information network abstraction layer unit is included in the sample of the video track or the sample entry of the video track, or both. In some embodiments, the format rule specifies that: in response to the second type of the second sample entry being vvi1, the operation point information network abstraction layer unit is included in the sample of the video track or the second sample entry of the video track, or both.

[0356] Implementation 13. A method for processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to format rules, wherein the format rules specify that non-video codec layer elementary streams in the visual media file are not allowed to include video codec layer network abstraction layer units (VCNs), and wherein the format rules specify that the non-video codec layer network abstraction layer units are synchronized with the elementary stream carried in a video track. In some embodiments, the format rules specify that non-video codec layer elementary streams in the visual media file are not allowed to include video codec layer network abstraction layer units (VCNs), and wherein the format rules specify that the non-video codec layer network abstraction layer units are synchronized with the elementary stream carried in a video track.

[0357] Implementation Method 14. The method according to any one of Implementation Methods 1-13, wherein the conversion includes generating a visual media file according to format rules and storing the bitstream into the visual media file.

[0358] Implementation 15. The method according to any one of Implementations 1-13, wherein the conversion includes generating a visual media file, and the method further includes storing the visual media file in a non-transitory computer-readable recording medium.

[0359] Implementation 16. The method according to any one of Implementations 1-13, wherein the conversion includes parsing the visual media file according to format rules to reconstruct the bitstream.

[0360] Implementation 17. The method according to any one of Implementations 1 to 16, wherein the visual media file is processed by a universal video codec (VVC), and the video track is a VVC track.

[0361] Embodiment 18. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform one or more of the methods according to Embodiments 1 to 17.

[0362] Embodiment 19. A non-transitory computer-readable storage medium for storing instructions that cause a processor to perform the method according to any one of Embodiments 1 to 17.

[0363] Embodiment 20. A video decoding apparatus, including a processor configured to implement one or more of the methods according to Embodiments 1 to 17.

[0364] Embodiment 21. A video encoding apparatus, including a processor configured to implement the method according to one or more of Embodiments 1 to 17.

[0365] Implementation 22. A computer program product storing computer code, which, when executed by a processor, causes the processor to perform the method according to any one of Implementations 1 to 17.

[0366] Embodiment 23. A computer-readable medium on which a visual media file conforms to a file format generated according to any one of Embodiments 1 to 17.

[0367] Embodiment 24. A non-transitory computer-readable recording medium for storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is described in any one of Embodiments 1 to 17.

[0368] Implementation 25. A method for generating a visual media file, comprising: generating a visual media file according to any one of Implementations 1 to 17, and storing the visual media file on a computer-readable program medium.

[0369] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to its corresponding bitstream representation, and vice versa. The bitstream representation of the current video block can, for example, correspond to bits juxtaposed or scattered in different places within the bitstream, as defined by the syntax. For example, a macroblock can be encoded based on the error residual values ​​after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude certain syntax fields and generate the codec representation accordingly by including or excluding syntax fields from the codec representation.

[0370] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for use by a data processing apparatus to operate or control the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting machine-readable propagation signals, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an operating environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.

[0371] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code sections). Computer programs can be deployed to run on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected through a communications network.

[0372] The processes and logic described in this document can be executed by one or more programmable processors running one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0373] Processors suitable for running computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0374] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0375] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0376] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A method of processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a type of a sample entry determines whether a decoding capability information network abstraction layer unit is included in a sample entry of a video track or in a sample of the video track and a sample entry of the video track in the visual media file, and wherein the format rule specifies that an operation point information network abstraction layer unit is allowed to be included in the visual media file in a sample entry description as one of a plurality of non-video coded layer network abstraction layer unit arrays in a decoder configuration record.

2. The method of claim 1, wherein the format rule specifies that the decoding capability information network abstraction layer unit is included in a sample entry of the video track in response to the type of the sample entry being vvc1.

3. The method of claim 1, wherein the format rule specifies that the decoding capability information network abstraction layer unit is included in a sample of the video track and a sample entry of the video track in response to the type of the sample entry being vvi1.

4. The method of claim 1, wherein the format rule specifies that a video elementary stream in the visual media file includes a video coded layer network abstraction layer unit, wherein the format rule specifies that a video elementary stream in the visual media file is not allowed to include a parameter set or the decoding capability information network abstraction layer unit, and wherein the format rule specifies that a sample entry in the visual media file stores the parameter set and the decoding capability information network abstraction layer unit.

5. The method of claim 4, wherein the format rule specifies that a video elementary stream in the visual media file is not allowed to include the parameter set, the decoding capability information network abstraction layer unit, or an operation point information network abstraction layer unit, and wherein the format rule specifies that a sample entry in the visual media file stores the parameter set, the decoding capability information network abstraction layer unit, and the operation point information network abstraction layer unit.

6. The method of claim 1, wherein the format rule specifies that, in response to samples in the visual media file belonging to a plurality of versatile video coding bitstreams and in response to the samples being included in a same track, the samples are allowed to belong to a same decoder capability information sample group, and wherein the format rule specifies that all samples belonging to the same decoder capability information sample group share a same decoder capability information sample group description entry.

7. The method of claim 6, wherein, the format rule specifies that, in response to all samples of a track belonging to a same bitstream or in response to all samples sharing a same decoding capability information, the decoding capability information network abstraction layer unit is indicated in a track level box in the visual media file regardless of a number of bitstreams.

8. The method of claim 7, wherein, the track level box is a track header box, a track level meta box, or another track level box.

9. The method of claim 1, wherein, The format rule specifies that a second type of second sample entry determines whether operation point information network abstraction layer units are included in: (1) the second sample entry of the video track in the visual media file, or (2) the sample of the video track in the visual media file or the second sample entry of the video track in the visual media file or both.

10. The method of any one of claims 1-9, wherein, The conversion includes generating the visual media file according to the format rule and storing the bitstream to the visual media file.

11. The method of any one of claims 1-9, wherein, The conversion includes parsing the visual media file according to the format rule to reconstruct the bitstream.

12. The method of any one of claims 1-9, wherein, The visual media file is processed by Versatile Video Coding (VVC), and the video track is a VVC track.

13. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-12.

14. A non-transitory computer-readable storage medium storing instructions, which cause a processor to perform the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Design of tracks and operation point signaling in layered HEVC file format

    US20160373771A1