Adaptive parameter set storage in video coding
By redefining the VVC track type and signaling mechanism, the storage and signaling notification problems in the VVC video file format are solved, more efficient multi-layer bitstream storage and decoder configuration information sharing are achieved, and streaming efficiency is improved.
Patent Information
- Application Number
- CN202111145077.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-07
- Filing Date
- 2021-09-28
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2041-09-28
AI Technical Summary
The existing VVC video file format has problems in signaling and storage, including unclear definitions of VVC reference tracks and non-VCL tracks, APS NAL units cannot be stored in multiple tracks, DCI NAL units are not considered in video elementary streams, non-VCL elementary streams may contain VCL NAL units, decoder configuration information sample group signaling is complex, and OPI NAL units are not allowed to be included in sample entries.
By redefining the VVC track type, VVC non-VCL tracks are allowed to contain APFs such as ALF, LMCS, or scaling list parameters in separate tracks. Picture header NAL units can be stored in VVC tracks or non-VCL tracks. DCI NAL units are signaled in track-level boxes. OPI NAL units are allowed to be included in sample entry descriptions. Decoder configuration information sample groups simplify the DCI sharing mechanism.
It achieves more optimized signaling and storage in the VVC video file format, supports efficient storage of multi-layer bitstreams and simplification of decoder configuration information, avoids redundant transmission of non-VCL NAL units and unnecessary sub-picture transmission, and improves the efficiency of streaming transmission.
Smart Images

Figure CN114302143B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is filed under applicable patent law and / or in accordance with the Paris Convention to claim priority to and the benefit of U.S. Provisional Patent Application No. 63 / 088,809, filed on October 7, 2020. The entire disclosure of the foregoing application is incorporated by reference as a part of the disclosure of this application for all purposes under the law. Technical Field
[0003] This patent document relates to the generation, storage and consumption of digital audio and video media information in file format. Background Art
[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses techniques that can be used by video encoders and decoders to process codec representations of videos or images according to file formats.
[0006] In one example aspect, a video processing method is disclosed. The method includes performing conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule stipulates that a first adaptation parameter set network abstraction layer unit is not allowed to be stored simultaneously in the visual media file in any one or both of (1) samples of a video codec layer track or sample entries of a video codec layer track, and (2) samples of a non-video codec layer track, wherein the video codec layer track is a track containing the video codec layer network abstraction layer unit, and wherein the first adaptation parameter set network abstraction layer unit includes luma mapping parameters with chroma scaling for the video stream and scaling list parameters for the video stream.
[0007] In another example aspect, a video processing method is disclosed. The method includes performing conversion between visual media data and a file storing information corresponding to the visual media data according to a format rule, wherein the format rule specifies a first condition for identifying a non-Video Codec Layer (VCL) track of the file and / or a second condition for identifying a VCL track of the file.
[0008] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.
[0009] In yet another exemplary aspect, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the above method.
[0010] In yet another example aspect, a computer readable medium having a bitstream stored thereon is disclosed. The bitstream is generated or processed using the methods described in this document.
[0011] In yet another example aspect, a computer readable medium having a bitstream stored thereon is disclosed. The bitstream is generated or processed using the methods described in this document.
[0012] These, additional, and other aspects are described throughout the present document. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 is a block diagram of an example video processing system.
[0014] Figure 2 is a block diagram of a video processing apparatus.
[0015] Figure 3 is a flowchart of an example method of video processing.
[0016] Figure 4 is a block diagram illustrating a video coding system according to some embodiments of the disclosure.
[0017] Figure 5 is a block diagram illustrating an encoder according to some embodiments of the disclosure.
[0018] Figure 6 is a block diagram illustrating a decoder according to some embodiments of the disclosure.
[0019] Figure 7 An example of an encoder block diagram is shown.
[0020] Figure 8 is a flowchart of an example method of video processing. DETAILED DESCRIPTION
[0021] For ease of understanding, section headings are used in this document, and the teachings and embodiments disclosed in each section are not meant to be limited to that section only. Also, H.266 terminology is used in some descriptions merely for ease of understanding, and not for limiting the scope of the disclosed technology. As such, the technology described herein is applicable to other video codec protocols and designs. In this document, editorial changes to text are shown by left and right double brackets, which indicate that the text between the double brackets is deleted text (e.g., [[ ]]), and by bold italic text, which indicates added text.
[0022] 1. Brief Discussion
[0023] This document relates to video file formats. In particular, this document relates to signaling and storage of picture header (PH), adaptation parameter set (APS), decoding capability information (DCI), and operating point information (OPI) network abstraction layer (NAL) units for Versatile Video Coding (VVC) video bitstreams in media files based on the ISO Base Media File Format (ISOBMFF). These ideas can be applied to video bitstreams coded by any codec (e.g., the VVC standard) and to any video file format (e.g., the VVC video file format under development), either individually or in various combinations.
[0024] 2. Abbreviations
[0025] ACT Adaptive Color Transform
[0026] ALF Adaptive Loop Filter
[0027] AMVR Adaptive Motion Vector Resolution
[0028] APS Adaptation Parameter Set
[0029] AU Access Unit
[0030] AUD Access Unit Delimiter
[0031] AVC Advanced Video Coding (Rec. ITU-T H.264 | ISO / IEC 14496-10)
[0032] B Bi-prediction
[0033] BCW Bi-prediction with CU-level weights
[0034] BDOF Bi-directional Optical Flow
[0035] BDPCM Block-based Delta Pulse Code Modulation
[0036] BP Buffer Period
[0037] CABAC Context-based Adaptive Binary Arithmetic Coding
[0038] CB Coded Block
[0039] CBR Constant Bit Rate
[0040] CCALF cross-component adaptive loop filter
[0041] CPB coded picture buffer
[0042] CRA clean random access
[0043] CRC cyclic redundancy check
[0044] CTB coded tree block
[0045] CTU coded tree unit
[0046] CU coding unit
[0047] CVS coded video sequence
[0048] DPB decoded picture buffer
[0049] DCI decoding capability information
[0050] DRAP dependent random access point
[0051] DU decoding unit
[0052] DUI decoding unit information
[0053] EG exponential Golomb
[0054] EGk k-th order exponential Golomb
[0055] EOB end of bitstream
[0056] EOS end of sequence
[0057] FD filler data
[0058] FIFO first in, first out
[0059] FL fixed length
[0060] GBR green, blue, and red
[0061] GCI general constraint information
[0062] GDR gradual decoding refresh
[0063] GPM geometric partition mode
[0064] HEVC high efficiency video coding (Rec. ITU-T H.265 | ISO / IEC 23008-2)
[0065] HRD hypothetical reference decoder
[0066] HSS hypothetical stream scheduler
[0067] I intra
[0068] IBC intra block copy
[0069] IDR instantaneous decoding refresh
[0070] ILRP inter-layer reference picture
[0071] IRAP intra random access point
[0072] LFNST low frequency non-separable transform
[0073] LPS least probable symbol
[0074] LSB least significant bit
[0075] LTRP long-term reference picture
[0076] LMCS luma mapping with chroma scaling
[0077] MIP matrix-based intra prediction
[0078] MPS most probable symbol
[0079] MSB most significant bit
[0080] MTS multiple transform selection
[0081] MVP motion vector prediction
[0082] NAL network abstraction layer
[0083] OLS output layer set
[0084] OP operation point
[0085] OPI operation point information
[0086] P prediction
[0087] PH picture header
[0088] POC picture order count
[0089] PPS picture parameter set
[0090] PROF prediction refinement with optical flow
[0091] PT picture timing
[0092] PU picture unit
[0093] QP quantization parameter
[0094] RADL random access decodable leading (picture)
[0095] RASL random access skipped leading (picture)
[0096] RBSP raw byte sequence payload
[0097] RGB red, green, and blue
[0098] RPL reference picture list
[0099] SAO sample adaptive offset
[0100] SAR sample aspect ratio
[0101] SEI supplemental enhancement information
[0102] SH slice header
[0103] SLI subpicture level information
[0104] SODB data bit string
[0105] SPS sequence parameter set
[0106] STRP short-term reference picture
[0107] STSA stepping time sublayer access
[0108] TR truncated rice
[0109] VBR variable bit rate
[0110] VCL video coding layer
[0111] VPS video parameter set
[0112] VSEI versatile supplemental enhancement information (Rec. ITU-T H.274 | ISO / IEC 23002-7)
[0113] VUI video usability information
[0114] VVC versatile video coding (Rec. ITU-T H.266 | ISO / IEC 23090-3)
[0115] 3. Introduction to video coding
[0116] 3.1. Video coding standards
[0117] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, the video coding standards are based on the hybrid video coding structure, where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project was officially started, the JVET was later renamed as the Joint Video Team (JVT). VVC is the new coding standard that has been finalized by the JVET at its 19th meeting on July 1, 2020, with the goal of 50% bitrate reduction compared to HEVC.
[0118] The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) have been designed for the widest range of applications, including traditional uses such as television broadcast, video conferencing or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding and viewport-adaptive 360° immersive media.
[0119] 3.2. File format standards
[0120] Media streaming applications are typically based on IP, TCP and HTTP transport methods and typically rely on file formats such as ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). In order to use video formats with ISOBMFF and DASH, a video format specific file format specification (such as AVC File Format and HEVC File Format) is needed for encapsulating video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream (e.g. profiles, tiers and levels, and many other information) will need to be exposed as file format level metadata and / or DASH media presentation description (MPD) for content selection purposes, e.g. for selecting appropriate media segments for initialization at the start of a streaming session and for stream adaptation during the streaming session.
[0121] Similarly, in order to use image formats with ISOBMFF, image format specific file format specifications such as AVC Image File Format and HEVC Image File Format will be needed.
[0122] The VVC Video File Format (a file format for storing VVC video content based on ISOBMFF) is currently being developed by MPEG.
[0123] The VVC Image File Format (a file format for storing image content coded using VVC based on ISOBMFF) is currently being developed by MPEG.
[0124] 3.3. PH, APS, DCI and OPI NAL units in VVC
[0125] Some new types of NAL units have been introduced in VVC, including PH, APS, DCI and OPI NAL units.
[0126] 3.3.1. Adaptation Parameter Set (APS)
[0127] Adaptive parameter set (APS) conveys picture-level and / or slice-level information that can be shared across multiple slices of a picture and / or slices of different pictures, but can change frequently across pictures, and the total number of variants can be high, thus not suitable to be included into PPS. Three types of parameters are included in APS: adaptive loop filter (ALF) parameters, luma mapping with chroma scaling (LMCS) parameters, and scaling list parameters. APS can be carried in two different NAL unit types, either in front of or after the associated slices as a prefix or suffix. The latter can be helpful in ultra-low delay scenarios, e.g., allowing the encoder to send slices of a picture before generating ALF parameters for the picture that will be used by subsequent pictures in decoding order.
[0128] 3.3.2. Picture header (PH)
[0129] A picture header (PH) structure exists for each PU. The PH is either in a separate PH NAL unit or included in the slice header (SH). If a PU consists of only one slice, the PH can only be included in the SH. To simplify the design, within a CLVS, the PH can only be all in PH NAL units or all in SHs. When the PH is in the SH, there is no PH NAL unit in the CLVS.
[0130] The PH is designed for two purposes. First, to help reduce the signaling overhead of the SHs of a picture that contains multiple slices by carrying all the parameters that have the same value for all slices of the picture, thus not repeating the same parameters in each SH. These include IRAP / GDR picture indication, inter / intra slice allowed flags, and information related to POC, RPL, deblocking filter, SAO, ALF, LMCS, scaling list, QP delta, weighted prediction, coding block partitioning, virtual boundary, collocated picture, etc. Second, to help the decoder identify the first slice of each coded picture that contains multiple slices. Since there is one and only one PH for each PU, when the decoder receives a PH NAL unit, it knows easily that the next VCL NAL unit is the first slice of the picture.
[0131] 3.3.3. Decoding capability information (DCI)
[0132] The DCI NAL unit contains bitstream-level PTL information. The DCI NAL unit includes one or more PTL syntax structures, which can be used during session negotiation between the sender and receiver of a VVC bitstream. When a DCI NAL unit is present in a VVC bitstream, each output layer set (OLS) in the CVS of the bitstream shall conform to the PTL information carried in at least one of the PTL structures in the DCI NAL unit.
[0133] In AVC and HEVC, the PTL information for session negotiation is available in SPS (for HEVC and AVC) and VPS (for HEVC layered extension). This design of conveying PTL information for session negotiation in HEVC and AVC has drawbacks, as the scope of SPS and VPS is within a CVS, not the whole bitstream. Therefore, sender-receiver session initiation can suffer from re-initiation during bitstream streaming of each new CVS. DCI solves this problem as it carries bitstream level information, thus, it can guarantee the decoding capability as indicated until the end of the bitstream.
[0134] 3.3.4. Operation point information (OPI)
[0135] The decoding process of HEVC and VVC has similar input variables to set the decoding operation point, i.e. the target OLS and the highest sublayer of the bitstream to be decoded, through the decoder API. However, in scenarios where the layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder API to the application, it can happen that the decoder cannot be correctly informed of the operation point of the decoder processing a given bitstream. Therefore, the decoder can not be able to conclude on the properties of the pictures in the bitstream, e.g. the correct buffer allocation of the decoded pictures and whether to output individual pictures or not. To solve this problem, VVC adds a mode in the bitstream to indicate these two variables through a newly introduced operation point information (OPI) NAL unit. In the AUs at the beginning of the bitstream and their individual CVSs, the OPI NAL unit informs the decoder of the target OLS and the highest sublayer of the bitstream to be decoded.
[0136] In case there is an OPI NAL unit and the operation point is also provided to the decoder via the decoder API information (e.g. the application can have more updated information on the target OLS and sublayers), the decoder API information takes precedence. In case there is no decoder API and any OPI NAL unit in the bitstream, a suitable fallback is specified in VVC to allow correct decoder operation.
[0137] 3.4. Some details on the VVC video file format
[0138] 3.4.1. Type of track
[0139] The VVC video file format specifies the following types of video tracks that carry VVC bitstreams in an ISOBMFF file:
[0140] a) VVC track:
[0141] A VVC track represents a VVC bitstream by including NAL units in its sample and sample entries, and possibly by referencing other VVC tracks containing other sublayers of the VVC bitstream, and possibly by referencing VVC subpicture tracks. When a VVC track references a VVC subpicture track, it is called a VVC reference track.
[0142] b) VVC non-VCL track:
[0143] APS carrying ALF, LMCS or scaling list parameters and other non-VCL NAL units can be stored in and transported through a track separate from the track containing VCL NAL units; this is the VVC non-VCL track.
[0144] c) VVC subpicture track:
[0145] A VVC subpicture track contains any of the following:
[0146] A sequence of one or more VVC subpictures.
[0147] A sequence of one or more complete slices forming a rectangular region.
[0148] A sample of a VVC subpicture track contains any of the following:
[0149] One or more complete subpictures consecutive in decoding order as specified in ISO / IEC 23090-3.
[0150] One or more complete slices forming a rectangular region and consecutive in decoding order as specified in ISO / IEC 23090-3.
[0151] The VVC subpictures or slices included in any sample of a VVC subpicture track are consecutive in decoding order.
[0152] NOTE: The VVC non-VCL track and the VVC subpicture track enable optimal delivery of VVC video in streaming applications as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, the client can request the DASH representation containing the subset of VVC subpicture tracks and the DASH representation containing the non-VCL track piecewise. In this way, redundant transport of APS and other non-VCL NAL units can be avoided.
[0153] 3.4.2. VVC elementary stream structure
[0154] Three types of elementary streams are defined for storing VVC content:
[0155] Video elementary stream, does not contain any parameter sets; all parameter sets are stored in the sample entry or sample entries;
[0156] Video and parameter set elementary stream, can contain parameter sets and can also store parameter sets in its sample entry or sample entries;
[0157] Non-VCL elementary stream, contains non-VCL NAL units that are synchronized with the elementary streams carried in the video track.
[0158] Note: VVC non-VCL tracks do not contain parameter sets in their sample entries.
[0159] 3.4.3. Decoder configuration information sample group
[0160] 3.4.3.1. Definition
[0161] The sample group description entry of this sample group contains the DCI NAL unit. All samples mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream.
[0162] This sample group indicates whether the same DCI NAL unit is used for different sample entries in the VVC track, i.e. whether the samples belonging to different sample entries belong to the same VVC bitstream. When the samples of two sample entries are mapped to the same decoder configuration information sample group description entry, the player can switch sample entries without reinitializing the decoder.
[0163] If any DCI NAL unit is present in any sample entry or in-band, this unit shall be exactly the same as the DCI NAL unit included in the decoder configuration information sample group.
[0164] 3.4.3.2. Syntax
[0165] class DecoderConfigurationInformation extends VisualSampleGroupEntry
[0166] ('dcfi'){
[0167] unsigned int(16) dciNalUnitLength;
[0168] bit(8 * nalUnitLength) dciNalUnit;
[0169] }
[0170] 3.4.3.3. Semantics
[0171] dciNalUnitLength indicates the byte length of the DCI NAL unit.
[0172] dciNalUnit contains a DCI NAL unit as specified in ISO / IEC 23090-3.
[0173] 4. Example technical problems solved by the disclosed technical solution
[0174] The latest design of the signaling of PH, APS, DCI and OPI NAL units in VVC video file format has the following problems:
[0175] 1) Neither VVC base track nor VVC non-VCL track shall contain VCL NAL units. However, the current definition of VVC non-VCL track would also apply to VVC base track. Moreover, according to the current definition, a VVC non-VCL track always contains APS NAL units. However, this would not allow a non-VCL NAL unit to contain picture header NAL units and possibly other non-VCL NAL units, but not including APS NAL units. Allowing such a VVC non-VCL track would enable optimal storage of extractable sub-picture based single-layer bitstreams in a file for late-banding of sub-picture tracks when different sub-pictures use different APS sets, e.g., by having one PH track (as a non-VCL track, although it contains the same information as a VVC base track), multiple APS tracks (as VVC non-VCL tracks) and multiple VVC sub-picture tracks each containing a sequence of sub-pictures.
[0176] 2) APS NAL units are all stored in one VVC non-VCL track or VVC track. In other words, APS NAL units cannot be stored in more than one track. This applies to APS NAL units containing LMCS parameters (i.e., LMCS APS) or APS NAL units containing scaling list (SL) parameters (i.e., SL APS), but not to APS NAL units containing ALF parameters (i.e., ALF APS). Since different VVC sub-picture tracks can use different sets of ALF APS, it is desirable to enable multiple VVC non-VCL tracks to carry the ALF APS of a VVC bitstream.
[0177] 3) DCI NAL units are not considered in the definition of video elementary stream as well as video and parameter set elementary stream. Therefore, a video elementary stream does not contain parameter sets, but can contain DCI NAL units.
[0178] 4) The definition of non-VCL elementary stream does not exclude the possibility of containing VCL NAL units in a non-VCL elementary stream.
[0179] 5) The decoder configuration information sample group provides a mechanism for signaling of the DCI NAL unit. However, there are the following problems:
[0180] a. In the most common use case, all samples of a track will belong to the same bitstream (or share the same DCI regardless of the number of bitstreams). For this case, it is complex to figure out the applicable DCI through the sample group signaling.
[0181] b. It is stated that all samples that map to the same decoder configuration information sample group description entry belong to the same VVC bitstream. However, this does not allow samples that belong to multiple VVC bitstreams (e.g., determined by EOB NAL units) but are in the same track to share the same DCI NAL unit even when they can.
[0182] 6) OPI NAL units are not allowed to be included in the sample entry description. However, in many cases, when OPI NAL units are present in a VVC bitstream, the OPI NAL units should be similarly considered as parameter sets, and thus they should be allowed to be included in the sample entry description.
[0183] 5. Example solutions and embodiments
[0184] To solve the above problems and other problems, the following methods summarized below are disclosed. These items should be considered as examples to explain the general concept and should not be interpreted in a narrow sense. Furthermore, these items can be applied individually or in any combination.
[0185] 1) To solve problems 1 and 2, one or more of the following items are proposed:
[0186] a. A VVC non-VCL track is defined as a track containing only non-VCL NAL units and is referred to by a VVC track through the "vvcN" track reference.
[0187] b. It is specified that a VVC non-VCL track can contain APSs carrying ALF, LMCS, or scaling list parameters stored in and transported through a separate track from the track containing VCL NAL units, with or without other non-VCL NAL units.
[0188] c. It is specified that a VVC non-VCL track can also contain picture header NAL units stored in and transported through a separate track from the track containing VCL NAL units, with or without APS NAL units, and with or without other non-VCL NAL units.
[0189] d. It is specified that picture header NAL units of a video stream can be stored in a sample of a VVC track or in a sample of a VVC non-VCL track, but not in both.
[0190] e. It is specified that LMCS APS NAL units (i.e. APS NAL units containing LMCS parameters) and scaling list APS NAL units (i.e. APS NAL units containing scaling list parameters) of a video stream can be stored in a sample and / or sample entry of a VVC track or in a sample of a VVC non-VCL track, but not in both.
[0191] f. It is specified that ALF APS NAL units (i.e. APS NAL units containing ALF parameters) of a video stream can be stored in a sample and / or sample entry of a VVC track, in a sample of a VVC non-VCL track, or in both.
[0192] 2) To solve problem 3, one or more of the following is proposed:
[0193] a. A video elementary stream is defined as an elementary stream containing VCL NAL units and not containing any parameter set, DCI or OPI NAL units; all parameter set, DCI and OPI NAL units are stored in sample entries.
[0194] i. Alternatively, a video elementary stream is defined as an elementary stream containing VCL NAL units and not containing any parameter set or DCI NAL units; all parameter set and DCI NAL units are stored in sample entries.
[0195] b. DCI NAL units are considered exactly the same as parameter sets, i.e. DCI NAL units can be only in sample entries of a video track (e.g. when the sample entry type name is “vvc1”) or can be in one or both of a sample and a sample entry of a video track (e.g. when the sample entry type name is “vvi1”).
[0196] 3) To solve problem 4, it is specified that a non-VCL elementary stream is an elementary stream containing only non-VCL NAL units and these non-VCL NAL units are synchronized with the elementary stream carried in a video track.
[0197] 4) To solve problem 5, one or more of the following is proposed:
[0198] a. For the case that all samples of a track belong to the same bitstream (or share the same DCI regardless of the number of bitstreams), the DCI NAL unit can be signaled in a track level box (e.g., track header box, track level meta box, or another track level box).
[0199] b. Allow samples belonging to multiple VVC bitstreams (e.g., determined by EOB NAL units) but in the same track to belong to the same decoder configuration information sample group, and thus share the same decoder configuration information sample group description entry.
[0200] 5) To address issue 6, allow OPI NAL units to be included in the sample entry description, e.g., as one of the non-VCL NAL unit arrays in the decoder configuration record.
[0201] a. Alternatively, treat OPI NAL units as exactly the same as parameter sets, i.e., OPI NAL units can be only in the sample entries of a video track (e.g., when the sample entry type name is “vvc1”), or can be in one or both of the samples and sample entries of a video track (e.g., when the sample entry type name is “vvi1”).
[0202] 6. Embodiments
[0203] Below are some example embodiments of the inventive aspects summarized in Section 5 above, which can be applied to the standard specification of the VVC video file format. The changed text is based on the latest draft specification. Most of the relevant parts that have been added or modified are indicated by bold italic text, and some deleted parts are indicated by double square brackets (e.g., [ [ ] ), where the deleted text between the double square brackets indicates deleted or cancelled text. Some other changes can be of editorial nature, thus not highlighted.
[0204] 6.1. First embodiment
[0205] This embodiment is for items 1a, 1b, 1c.
[0206] 6.1.1. Type of track
[0207] This specification specifies the following types of video tracks for carrying VVC bitstreams:
[0208] a) VVC track:
[0209] A VVC track is associated with other VVC tracks containing other layers and / or sub-layers of a VVC bitstream, possibly through the inclusion of NAL units in its samples and / or sample entries, and possibly through "vopi" and "linf" sample groups or through "opeg" entity groups, and possibly by referencing VVC sub-picture tracks, to represent a VVC bitstream.
[0210] A VVC track is also called a VVC reference track when it references a VVC sub-picture track. A VVC reference track shall not contain VCL NAL units and shall not be referred to by a VVC track through a "vvcN" track reference.
[0211] b) VVC non-VCL track:
[0212] A VVC non-VCL track is a track containing only non-VCL NAL units and is referred to by a VVC track through a "vvcN" track reference.
[0213] A VVC non-VCL track can contain APS, carrying ALF, LMCS or scaling list parameters, stored in and transported through a track separate from the track containing VCL NAL units, with or without other non-VCL NAL units.
[0214] A VVC non-VCL track can also contain picture header NAL units, with or without APS NAL units, and with or without other non-VCL NAL units, stored in and transported through a track separate from the track containing VCL NAL units.
[0215] c) VVC sub-picture track:
[0216] A VVC sub-picture track contains any of the following:
[0217] A sequence of one or more VVC sub-pictures.
[0218] A sequence of one or more complete slices forming a rectangular region.
[0219] A sample of a VVC sub-picture track contains any of the following:
[0220] One or more complete sub-pictures consecutive in decoding order as specified in ISO / IEC 23090-3.
[0221] One or more complete slices forming a rectangular region and consecutive in decoding order as specified in ISO / IEC 23090-3.
[0222] The VVC sub-pictures or slices included in any sample of a VVC sub-picture track are consecutive in decoding order.
[0223] NOTE: The VVC non-VCL tracks and VVC subpicture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each carry in their own DASH representation, and for decoding and rendering of a subset of tracks, the client can request piece-wise the DASH representation containing the subset of VVC subpicture tracks and the DASH representation containing the non-VCL tracks. In this way, redundant transmission of APS and other non-VCL NAL units can be avoided, and also unnecessary transmission of subpictures.
[0224] 6.2. Second embodiment
[0225] This embodiment is for item 4.b.
[0226] 6.2.1. Decoder [[configuration]] capability information sample group
[0227] 6.2.1.1. Definition
[0228] The sample group description entry of this sample group contains a DCI NAL unit. [[All samples that are mapped to the same decoder configuration information sample group description entry belong to the same VVC bitstream.]]
[0229] This sample group indicates whether the same DCI NAL unit is used for different sample entries in a VVC track [[i.e. whether samples belonging to different sample entries belong to the same VVC bitstream]]. When the samples of two sample entries are mapped to the same decoder configuration information sample group description entry, the player can switch sample entries without re-initializing the decoder.
[0230] If any DCI NAL unit is present in-band or in any sample entry, this unit shall be identical to the DCI NAL unit included in the corresponding decoder configuration information sample group entry.
[0231] 6.2.1.2. Syntax
[0232] class DecoderConfigurationInformation extends VisualSampleGroupEntry
[0233] ('dcfi'){
[0234] unsigned int(16) dciNalUnitLength;
[0235] bit(8 * nalUnitLength) dciNalUnit;
[0236] }
[0237] 6.2.1.3. Semantics
[0238] dciNalUnitLength indicates the byte length of the DCI NAL unit.
[0239] dciNalUnit contains a DCI NAL unit as specified in ISO / IEC 23090-3.
[0240] 6.3. Third embodiment
[0241] This embodiment is for clause 5.
[0242] 6.3.1. Definition of VVC decoder configuration record
[0243] This subclause specifies the decoder configuration information for ISO / IEC 23090-3 video content.
[0244] This record contains the size of the length fields in each sample for indicating the length of the NAL units and parameter sets, DCI, OPI and SEI NAL units it contains, if stored in the sample entry. This record is externally scoped (its size is provided by the structure containing it).
[0245] This record contains a version field. This version of the specification defines version 1 of this record. Incompatible changes to the record will be indicated by a change in the version number. If the version number is not recognized, the reader shall not attempt to decode the record or the stream to which it applies.
[0246] Compatible extensions to this record shall extend it and shall not change the configuration version code. The reader should be prepared to ignore unrecognized data that extends beyond the data definition they understand.
[0247] When a track itself or by parsing a "subp" track reference contains a VVC bitstream, a VvcPtlRecord shall be present in the decoder configuration record and, in this case, the particular output layer set of the VVC bitstream is indicated by the field output_layer_set_idx. If ptl_present_flag is equal to zero in the decoder configuration record of a track, the track shall have an "oref" track reference. ...
[0249] There is an array set to carry initialization non-VCL NAL units. The NAL unit types are restricted to indicate only DCI, OPI, VPS, SPS, PPS, prefix APS and prefix SEI NAL units. NAL unit types reserved in ISO / IEC 23090-3 and this specification can be taken for definition in the future and the reader should ignore the array with NAL unit types of reserved or not permitted values.
[0250] NOTE 2: This "tolerant" behavior is designed so as not to generate errors, allowing the possibility of backward-compatible extensions of these arrays in future specifications.
[0251] NOTE 3: The NAL units carried in the sample entry are included in the access unit reconstructed from the first sample of the reference sample entry, immediately after the AUD and OPINAL units (if any), or at the beginning of the access unit.
[0252] The suggested array order is DCI, OPI, VPS, SPS, PPS, prefix APS, prefix SEI. ...
[0254] 6.3.2. Semantics of VVC decoder configuration record ...
[0256] numArrays indicates the number of arrays of NAL units of the indicated type.
[0257] array_completeness indicates, when equal to 1, that all NAL units of the given type are in the following array and none is in the stream; when equal to 0, that additional NAL units of the indicated type can be in the stream; [[default and]] permitted values are constrained by the sample entry name.
[0258] NAL_unit_type indicates the type of the NAL units in the following array (which should be all of that type); it takes values as defined in ISO / IEC 23090-3; it is constrained to take one of the values indicating DCI, OPI, VPS, SPS, PPS, prefix APS or prefix SEI [[, or suffix SEI]] NAL units.
[0259] numNalus indicates the number of NAL units of the indicated type included in the configuration record of the stream for which this configuration record applies. SEI arrays shall contain only SEI messages of "declarative" nature, i.e. those providing information about the stream as a whole. An example of such SEI can be the user data SEI.
[0260] nalUnitLength indicates the byte length of the NAL unit.
[0261] nalUnit contains a DCI, OPI, VPS, SPS, PPS, APS or declarative SEI NAL unit as specified in ISO / IEC 23090-3.
[0262] 6.4. Fourth embodiment
[0263] This example is for items 1a, 1b, 1c, 1d, 1e and 1f.
[0264] 6.4.1. Background: Characteristics of VVC (informative)
[0265] The storage of VVC content uses existing capabilities of the ISO base media file format, but also defines extensions to support the following characteristics of the VVC codec:
[0266] d) Parameter sets and DCI and OPI NAL units:
[0267] The VPS, SPS and PPS mechanisms decouple the transmission of infrequent changing information from the transmission of coded block data. Each coded picture references a PPS that contains its decoding parameters. In turn, a PPS references an SPS that contains sequence level decoding parameter information, and an SPS references a VPS that contains cross-layer global decoding parameter information. When the referenced VPS ID has a zero value, the bitstream contains only one layer, and there is effectively no reference to a VPS.
[0268] The APS mechanism allows efficient storage and transmission of information that typically has a large amount of variation within the bitstream and is frequently updated across pictures, such as the parameters of the adaptive loop filter (ALF), the parameters of the adaptive loop reshaper (also known as luma mapping with chroma scaling (LMCS)), and the parameters of the scaling list that associates each frequency index with a scaling factor of the scaling process specified in ISO / IEC 23090-3. The APS carrying the parameters of ALF, LMCS and scaling list are also referred to as ALF APS, LMCS APS and scaling list APS, respectively. Each slice containing coded block data can reference one or more APS containing the ALF parameters of that slice. In addition, each picture header can reference any type of APS.
[0269] A VVC bitstream can contain a DCI NAL unit that contains parameters describing the maximum capabilities required to decode the entire bitstream.
[0270] A VVC bitstream can also contain an OPI NAL unit that contains an indication of an operation point determined by a target output layer set and a target highest temporal ID.
[0271] e) Picture header
[0272] The picture header (PH) structure contains parameters that are the same for all slices of a picture, is present for each picture, and is either present in its own NAL unit or directly in the slice header (SH). If a picture has only one slice, the PH can only be included in the SH. Within a CLVS, the PH can only be all in PH NAL units or all in SHs.
[0273] f) Subpicture
[0274] A VVC subpicture is a rectangular region of one or more slices within a picture. An encoder can treat a subpicture boundary like a picture boundary and can turn off in-loop filtering across subpicture boundaries. Thus, a subpicture can be coded so that a selected subpicture can be extracted from the VVC bitstream(s) or merged into a target VVC bitstream. Furthermore, such VVC bitstream extraction or merging operation can be performed without modifying the VCL NAL units. The subpicture identifier (ID) of a subpicture present in the bitstream can be indicated in the SPS(s) or PPS(s).
[0275] 6.4.2. Types of tracks
[0276] This specification specifies the following types of video tracks that carry VVC bitstreams:
[0277] g) VVC track:
[0278] A VVC track represents a VVC bitstream by including NAL units in its samples and / or sample entries, and can associate other VVC tracks containing other layers and / or sublayers of the VVC bitstream by "vopi" and "linf" sample groups or by "opeg" entity groups, and can reference VVC subpicture tracks.
[0279] When a VVC track references a VVC subpicture track, it is also called a VVC reference track. A VVC reference track shall not contain VCL NAL units and shall not be referred to by a VVC track through a "vvcN" track reference.
[0280] h) VVC non-VCL track:
[0281] A VVC non-VCL track is a track that contains only non-VCL NAL units and is referred to by a VVC track through a "vvcN" track reference.
[0282] A VVC non-VCL track can contain APS that carry ALF, LMCS, or scaling list parameters, with or without other non-VCL NAL units, stored in and transported through a track separate from the track containing VCL NAL units.
[0283] A VVC non-VCL track can also contain picture header NAL units stored in and transported through a track separate from the track containing the VCL NAL units, with or without APS NAL units, and with or without other non-VCL NAL units.
[0284] i) A VVC subpicture track:
[0285] A VVC subpicture track contains any of the following:
[0286] A sequence of one or more VVC subpictures.
[0287] A sequence of one or more complete slices forming a rectangular region.
[0288] A sample of a VVC subpicture track contains any of the following:
[0289] One or more complete subpictures that are consecutive in decoding order as specified in ISO / IEC 23090-3.
[0290] One or more complete slices forming a rectangular region and being consecutive in decoding order as specified in ISO / IEC 23090-3.
[0291] The VVC subpictures or slices included in any sample of a VVC subpicture track are consecutive in decoding order.
[0292] NOTE: VVC non-VCL tracks and VVC subpicture tracks enable optimal delivery of VVC video in streaming applications as follows. These tracks can each be carried in their own DASH representation, and for decoding and rendering of a subset of tracks, a client can request a DASH representation containing a subset of VVC subpicture tracks and a DASH representation containing non-VCL tracks piece-wise. In this way, redundant transport of APS and other non-VCL NAL units can be avoided, and unnecessary transport of subpictures can also be avoided.
[0293] 6.4.3. Typical order and restrictions
[0294] In addition to the general conditions in 4.3.2, a typical stream format is a VVC elementary stream that also satisfies the following conditions:
[0295] • Access unit delimiter NAL units: The constraints that access unit delimiter NAL units adhere to are defined in ISO / IEC 23090-3.
[0296] • DCI NAL units, VPS, SPS and PPS: The VPS, SPS or PPS to be used for decoding a picture must be sent either before the samples of the picture or in the samples of the picture. For a video stream for which the video stream for which the sample entry name is "vvc1", DCI NAL units, VPS, SPS and PPS shall be stored only in the sample entry and, for a video stream for which the sample entry name is "vvi1", they can be stored in the sample entry and in the samples.
[0297] NOTE 1: Storing DCI NAL units, VPS, SPS and PPS in the sample entry of a video stream provides a simple and static way to provide decoding capability information and parameter sets. On the other hand, storing these NAL units in the samples is more complex but allows more dynamism in case of parameter set update (the content of a particular parameter set is changed but the same ID is used) and in case of addition of additional parameter sets. From any sample marked as synchronization sample, the decoder is initialized with these parameter sets in the sample entry and then updated when these parameter sets occur in the stream. Such an update can replace these parameter sets with a new definition using the same identifier. Each time the sample entry changes, the decoder is reinitialized with these parameter sets included in the sample entry.
[0298] • APS: A prefix APS NAL unit is constrained to be sent before the first VCL NAL unit of a picture unit (PU). An APS NAL unit including a prefix APS NAL unit and a suffix APS NAL unit and having a particular value of aps_adaptation_parameter_set_id and a particular value of aps_params_type within a PU is allowed to exist but is required to have the same content. A suffix APS NAL unit is constrained to be sent after the last VCL NAL unit of a PU. For a video stream for which the video stream for which the sample entry name is "vvc1", the following applies:
[0299] o [[If the track has a track reference of type "vvcN", APS shall be stored only in the VVC non-VCL track referred to by the track reference of type "vvcN".]]
[0300] o [[Else if the sample entry name is "vvc1", APS shall be stored only in the sample entry.]]
[0301] o [[Else (the track does not have a track reference of type "vvcN" and the sample entry name is "vvi1"), APS shall be stored only in the sample entry and in the samples.]]
[0302] o LMCS APS NAL units and scaling list APS NAL units for a video stream can be stored in samples and / or sample entries of a VVC track or in samples of a VVC non-VCL track, but not in both.
[0303] o ALF APS NAL units for a video stream can be stored in samples and / or sample entries of a VVC track, in samples of a VVC non-VCL track, or in both.
[0304] • Picture header NAL units: Picture header NAL units for a video stream can be stored in samples of a VVC track or in samples of a VVC non-VCL track, but not in both.
[0305] • SEI messages: SEI messages of declarative nature can be stored in sample entries; there is no provision for removing such SEI messages from samples.
[0306] • Padding data: Video data is naturally represented in variable bit rate in file formats, and padding should be provided for transport if needed. When sample entries do not permit in- stream parameter sets either, padding data NAL units and padding payload SEI messages shall not be present in file format stored streams.
[0307] NOTE 2: When operating an HRD in CBR mode as specified in ISO / IEC 23090-3 Annex C, removing or adding padding data NAL units, start codes, SEI messages, or padding payload SEI messages can change bitstream characteristics with respect to conformance with the HRD.
[0308] Figure 1 FIG. 19 is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 1900. The system 1900 can include an input 1902 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be in a compressed or encoded format. The input 1902 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0309] The system 1900 can include a codec component 1904 that can implement various coding or encoding methods described in this document. The codec component 1904 can reduce the average bitrate of video from the input 1902 to the output of the codec component 1904 to produce a coded representation of the video. The codec techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 1904 can be stored, or transmitted via a communication connection as represented by the component 1906. The stored or communicated bitstream (or coded) representation of the video received at the input 1902 can be used by the component 1908 to generate pixel values or a displayable video to the display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Also, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0310] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices that are capable of performing digital data processing and / or video display.
[0311] Figure 2 is a block diagram of a video processing apparatus 3600. The apparatus 3600 can be used to implement one or more methods described herein. The apparatus 3600 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 3600 can include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processor(s) 3602 can be configured to implement one or more methods described in this document. The memory(ies) 604 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing hardware 606 can be used to implement, in hardware circuitry, some of the techniques described in this document. In some embodiments, the video processing hardware 3606 can be included at least in part in the processor 3602 (e.g., a graphics co-processor).
[0312] Figure 4 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure.
[0313] like Figure 4 As shown, the video encoding and decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data, wherein the source device 110 may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110, wherein the target device 120 may be referred to as a video decoding device.
[0314] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .
[0315] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. Associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.
[0316] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0317] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain coded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the coded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120, or may be external to target device 120 configured to interface with an external display device.
[0318] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or additional standards.
[0319] Figure 5 is a block diagram illustrating an example of a video encoder 200, which may be Figure 4The video encoder 114 in the illustrated system 100.
[0320] The video encoder 200 can be configured to perform any or all of the techniques of this disclosure. In Figure 5 In an example, the video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0321] The functional components of the video encoder 200 can include a partition unit 201, a prediction unit 202 (which can include a mode select unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206), a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214.
[0322] In other examples, the video encoder 200 can include more, less, or different functional components. In an example, the prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0323] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but are represented separately for explanatory purposes. Figure 5
[0324] The partition unit 201 can partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0325] The mode select unit 203 can select one of the coding modes (e.g., intra or inter) based on the error results and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 for generation of residual block data and to the reconstruction unit 212 for reconstruction of the encoded block for use as a reference picture. In some examples, the mode select unit 203 can select a combination of intra and inter prediction modes (CIIP), where the prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode select unit 203 can also select a resolution for the motion vector of the block (e.g., sub-pixel or integer pixel precision).
[0326] To perform inter prediction on a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.
[0327] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0328] In some examples, motion estimation unit 204 can perform uni-prediction on a current video block, and motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of list 0 or list 1. Motion estimation unit 204 can then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0329] In other examples, motion estimation unit 204 can perform bi-prediction on a current video block, motion estimation unit 204 can search for a reference video block for the current video block in a reference picture of list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 204 can then generate a reference index indicating the reference pictures in list 0 and list 1 that contain the reference video blocks and a motion vector indicating a spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and the motion vector for the current video block as the motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0330] In some examples, motion estimation unit 204 can output a full set of motion information for a current video block for decoding processing at a decoder.
[0331] In some examples, motion estimation unit 204 can not output a full set of motion information for a current video block. Instead, motion estimation unit 204 can signal the motion information for the current video block with reference to motion information of another video block. For example, motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.
[0332] In one example, the motion estimation unit 204 can indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0333] In another example, the motion estimation unit 204 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0334] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0335] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0336] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0337] In other examples, such as in skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit 207 can not perform the subtraction operation.
[0338] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0339] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0340] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated from prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.
[0341] After reconstruction unit 212 reconstructs a video block, in-loop filtering operations can be performed to reduce video block artifacts in the video block.
[0342] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0343] Figure 6 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 4 the video decoder 114 in the system 100 shown.
[0344] Video decoder 300 can be configured to perform any or all of the techniques of this disclosure. In Figure 6 In examples where video decoder 300 includes multiple functional components, the techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0345] In Figure 6 In examples where video decoder 300 includes multiple functional components, the techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure. Figure 5 ) described for video encoder 200.
[0346] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data, and from the entropy decoded video data, motion compensation unit 302 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes.
[0347] The motion compensation unit 302 may generate a motion compensated block and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0348] Motion compensation unit 302 may calculate interpolated values of sub-integer pixels of a reference block using interpolation filters such as those used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.
[0349] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence.
[0350] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0351] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307 to provide reference blocks for subsequent motion compensation / intra-frame prediction and also to generate decoded video for presentation on a display device.
[0352] The following provides a list of preferred solutions for some embodiments.
[0353] The following solutions illustrate example embodiments of the techniques discussed in the previous section (eg, items 1 through 4).
[0354] 1. A method for processing visual media (e.g., Figure 3 ), comprising: performing conversion (3002) between visual media data and a file storing information corresponding to the visual media data according to a format rule; wherein the format rule specifies a first condition for identifying a non-video codec layer (VCL) track of the file and / or a second condition for identifying a VCL track of the file.
[0355] 2. The method according to solution 1, wherein the first condition specifies that the non-VCL track contains only non-VCL network abstraction layer units and is identified in the VCL track by a specific track reference.
[0356] 3. The method according to solutions 1-2, wherein the first condition specifies that the non-VCL track contains an adaptation parameter set (APS) corresponding to the VCL track.
[0357] 4. The method according to any of solutions 1-3, wherein the second condition for the VCL track specifies that the VCL track is not allowed to include decoding capability information (DCI) or operation point information (OPI) network abstraction units.
[0358] 5. The method according to solution 1, wherein the first condition specifies that the non-VCL track includes one or more elementary streams containing non-VCL network abstraction layer units, and wherein the non-VCL network abstraction layer units are synchronized with the elementary streams in the VCL track.
[0359] 6. The method according to any of solutions 1-5, wherein the conversion comprises generating a bitstream representation of the visual media data, and storing the bitstream representation into a file according to the format rule.
[0360] 7. The method according to any of solutions 1-5, wherein the conversion comprises parsing a file according to the format rule to recover the visual media data.
[0361] 8. A video decoding apparatus comprising a processor configured to implement a method according to one or more of solutions 1 to 7.
[0362] 9. A video encoding apparatus comprising a processor configured to implement a method according to one or more of solutions 1 to 7.
[0363] 10. A computer program product having computer code stored therein, the code, which when executed by a processor, causes the processor to implement a method according to any of solutions 1 to 7.
[0364] 11. A computer readable medium, a bitstream representation on the computer readable medium being in accordance with a file format generated according to any of solutions 1 to 7.
[0365] 12. A method, apparatus or system described in the present document.
[0366] In the solutions described herein, an encoder can conform to a format rule by generating a coded representation according to the format rule. In the solutions described herein, a decoder can use a format rule to parse syntax elements in a coded representation knowing whether a syntax element is present or not according to the format rule to generate a decoded video.
[0367] TECHNIQUE 1. A method of processing visual media data (e.g., method 8000 depicted in FIG. 8), comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule (8002), wherein the format rule specifies that an adaptation parameter set network abstraction layer unit is not allowed to be simultaneously stored in (1) any or both of a sample of a video coding layer track or a sample entry of a video coding layer track, and (2) the visual media file in a sample of a non-video coding layer track, wherein the video coding layer track is a track containing video coding layer network abstraction layer units, and wherein the adaptation parameter set network abstraction layer unit includes a luma mapping with chroma scaling parameter of a video stream and a scaling list parameter of the video stream. Figure 8 TECHNIQUE 2. The method of technique 1, wherein the format rule specifies that the adaptation parameter set network abstraction layer unit is stored in the visual media file in any or both of the sample of the video coding layer track or the sample entry of the video coding layer track.
[0368] TECHNIQUE 3. The method of technique 1, wherein the format rule specifies that the adaptation parameter set network abstraction layer unit is stored in the visual media file in the sample of the non-video coding layer track.
[0369] TECHNIQUE 4. The method of technique 1, wherein the format rule specifies that the adaptation parameter set network abstraction layer unit is stored in the visual media file in a sample of a non-video coding layer track.
[0370] TECHNIQUE 4. The method according to TECHNIQUE 1, wherein the format rule further specifies that the second adaptation parameter set network abstraction layer unit is allowed to be simultaneously stored in the visual media file in (1) any one or both of a sample of a video coding layer track or a sample entry of a video coding layer track, and (2) a sample of a non-video coding layer track, wherein the video coding layer track is a track containing video coding layer network abstraction layer units, and wherein the second adaptation parameter set network abstraction layer unit includes an adaptive loop filter parameter of a video stream. In some embodiments, a method of processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of the visual media data according to a format rule, wherein the format rule specifies that an adaptation parameter set network abstraction layer unit is allowed to be simultaneously stored in the visual media file in (1) any one or both of a sample of a video coding layer track or a sample entry of a video coding layer track, and (2) a sample of a non-video coding layer track, wherein the video coding layer track is a track containing video coding layer network abstraction layer units, and wherein the adaptation parameter set network abstraction layer unit includes an adaptive loop filter parameter of a video stream.
[0371] TECHNIQUE 5. The method according to TECHNIQUE 4, wherein the format rule specifies that the second adaptation parameter set network abstraction layer unit is stored in the visual media file in any one or both of a sample of a video coding layer track or a sample entry of a video coding layer track. In some embodiments, the format rule specifies that the adaptation parameter set network abstraction layer unit is stored in the visual media file in any one or both of a sample of a video coding layer track or a sample entry of a video coding layer track.
[0372] TECHNIQUE 6. The method according to TECHNIQUE 4, wherein the format rule specifies that the second adaptation parameter set network abstraction layer unit is stored in the visual media file in a sample of a non-video coding layer track. In some embodiments, the format rule specifies that the adaptation parameter set network abstraction layer unit is stored in the visual media file in a sample of a non-video coding layer track.
[0373] TECHNIQUE 7. The method according to any one of TECHNIQUES 1-6, wherein the conversion comprises generating the visual media file and storing the bitstream to the visual media file according to the format rule.
[0374] TECHNIQUE 8. The method according to any one of TECHNIQUES 1-6, wherein the conversion comprises generating the visual media file, and the method further comprises storing the visual media file in a non-transitory computer-readable recording medium.
[0375] TECHNIQUE 9. The method according to any one of TECHNIQUES 1-6, wherein the conversion comprises parsing the visual media file to reconstruct the bitstream according to the format rule.
[0376] Technique 10. The method of any of techniques 1 to 9, wherein the visual media file is processed by Versatile Video Coding (VVC), and wherein the non-video coded layer track or the video coded layer track is a VVC track.
[0377] Technique 11. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method of one or more of techniques 1 to 10.
[0378] Technique 11. A non-transitory computer-readable storage medium storing instructions that cause a processor to implement the method of one or more of techniques 1 to 10.
[0379] Technique 12. A non-transitory computer-readable storage medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method comprises determining a format rule according to the method of any one or more of techniques 1 to 10, and generating the visual media file based on the determination. In some embodiments, a non-transitory computer-readable storage medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method comprises generating the visual media file based on visual media data according to a format rule, wherein the format rule specifies that a first adaptation parameter set network abstraction layer unit is not allowed to be simultaneously stored in (1) any or both of a sample of a video coded layer track or a sample entry of a video coded layer track, and (2) the visual media file in a sample of a non-video coded layer track, wherein the video coded layer track is a track containing video coded layer network abstraction layer units, and wherein the first adaptation parameter set network abstraction layer unit includes a luma mapping with chroma scaling parameter of a video stream and a scaling list parameter of the video stream.
[0380] Technique 13. A video decoding apparatus comprising a processor configured to implement the method of one or more of techniques 1 to 10.
[0381] Technique 14. A video encoding apparatus comprising a processor configured to implement the method of one or more of techniques 1 to 10.
[0382] Technique 15. A computer program product having computer code stored thereon, the code, when executed by a processor, causing the processor to implement the method of any of techniques 1 to 10.
[0383] Technique 16. A computer readable medium, a visual media file on the computer readable medium conforming to a file format generated according to any of techniques 1 to 10.
[0384] TECHNIQUE 17. A non-transitory computer readable storage medium storing a bitstream of a visual media file generated by a method performed by a video processing apparatus, wherein the method is recited in any of TECHNIQUES 1 to 10.
[0385] TECHNIQUE 18. A method of visual media file generation, comprising: generating a visual media file according to the method recited in any of TECHNIQUES 1 to 10, and storing the visual media file on a computer readable program medium.
[0386] In this document, the term “video processing” can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block can for example correspond to collocated or scattered bits within the bitstream as defined by the syntax. For example, a macroblock can be encoded according to a transform and coded error residual values and also using bits in headers and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination, knowing that some fields can or can not be present, as described in the above solutions. Similarly, an encoder can determine to include or not include certain syntax fields and generate the coded representation accordingly by including the syntax fields or excluding the syntax fields from the coded representation.
[0387] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of one or more of them, or a combination of one or more of them and other like articles. The term “data processing apparatus” encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated for the purpose of encoding information in a modulated data signal for transmission to an appropriate receiver device.
[0388] A computer program (which can also be referred to or referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to operate on one computer or on multiple computers that are located by one site or distributed across multiple sites and interconnected by a communication network.
[0389] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and / or devices that are
[0390] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0391] While this patent document contains many details, these should not be construed as limiting the scope of any subject matter or of any embodiment, but as merely describing features that are specific to certain embodiments of the specific technology described in this patent document. Certain features described in the context of separate embodiments in this patent document can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any appropriate subcombination. Moreover, although features can be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0392] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order nor that all illustrated operations be performed to achieve desirable results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0393] Only a few implementations and examples are described and other implementations, enhancements and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method of processing visual media data, comprising: performing a conversion between a visual media file and a bitstream of visual media data according to a format rule, wherein the format rule specifies that a first adaptation parameter set network abstraction layer unit is not allowed to be simultaneously stored in the visual media file in (1) any one or both of a sample of a video coding layer track or a sample entry of the video coding layer track, and (2) a sample of a non-video coding layer track, wherein the video coding layer track is a track that includes video coding layer network abstraction layer units, and wherein the first adaptation parameter set network abstraction layer unit includes a luma mapping with chroma scaling parameter of a video stream and a scaling list parameter of the video stream, wherein the format rule further specifies that a second adaptation parameter set network abstraction layer unit is allowed to be simultaneously stored in the visual media file in (1) any one or both of a sample of a video coding layer track or a sample entry of the video coding layer track, and (2) a sample of a non-video coding layer track, wherein the second adaptation parameter set network abstraction layer unit includes an adaptive loop filter parameter of a video stream, wherein the format rule specifies that the second adaptation parameter set network abstraction layer unit is stored in the visual media file in any one or both of a sample of a video coding layer track or a sample entry of the video coding layer track, wherein the format rule specifies that the second adaptation parameter set network abstraction layer unit is stored in the visual media file in a sample of a non-video coding layer track, wherein the format rule specifies that a sample entry type determines whether an operation point information network abstraction layer unit is included in (1) a sample entry of a video track of the visual media file, or (2) a sample of a video track of the visual media file, a sample entry of a video track of the visual media file, or a sample of a video track of the visual media file and a sample entry of a video track of the visual media file.
2. The method of claim 1, wherein, the format rule specifies that the first adaptation parameter set network abstraction layer unit is stored in the visual media file in any one or both of a sample of a video coding layer track or a sample entry of the video coding layer track.
3. The method of claim 1, wherein, the format rule specifies that the first adaptation parameter set network abstraction layer unit is stored in the visual media file in a sample of a non-video coding layer track.
4. The method of any one of claims 1-3, wherein, the conversion includes generating the visual media file and storing the bitstream to the visual media file according to the format rule.
5. The method of any one of claims 1-3, wherein, the conversion includes generating the visual media file, and the method further includes storing the visual media file in a non-transitory computer readable storage medium.
6. The method of any one of claims 1-3, wherein, the conversion includes parsing the visual media file to reconstruct the bitstream according to the format rule.
7. The method of any of claims 1-3, wherein the visual media file is processed by Versatile Video Coding (VVC), and the format rule specifies that the first adaptation parameter set network abstraction layer unit is stored in the visual media file in any one or both of a sample of a video coding layer track or a sample entry of the video coding layer track. The video coded layer track or the non-video coded layer track is a VVC track.
8. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium storing instructions that cause a processor to implement the method according to any one of claims 1 to 7.