Video availability information in a common media application format
By specifying the SPS field and DCI NAL unit in the VVC CMAF track to convey decoding capability information, the problem of inconsistent video availability information in the VVC codec standard is resolved, improving the efficiency and flexibility of streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FACE CUTE CO LTD
- Filing Date
- 2022-04-18
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, the video codec standard VVC lacks effective video availability information interleaving and progressive scan indication in the CMAF track, which leads to frequent transcoding and signaling notifications during content preparation, affecting the flexibility and efficiency of streaming.
By specifying the video availability information field in the Sequence Parameter Set (SPS) in the VVC CMAF track, explicitly indicating whether the video is encoded and decoded progressively or interleaved, and conveying the decoding capability information of the bitstream in the DCI NAL unit, consistency is ensured throughout the bitstream, and the frequency of signaling notifications is reduced.
It enables more flexible content preparation in VVC CMAF tracks, reduces transcoding processes, improves the efficiency and flexibility of streaming, and adapts to different network conditions and device capabilities.
Smart Images

Figure CN115225907B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] Pursuant to the applicable patent law and / or rules of the Paris Convention, this application aims to promptly claim priority and benefit to U.S. Provisional Patent Application No. 63 / 176,315, filed April 18, 2021. For all purposes required by law, the entire disclosure of the aforementioned application is incorporated by reference as part of the disclosure of this application. Technical Field
[0003] This patent document relates to the generation, storage, and consumption of digital audio and video media information in file formats. Background Technology
[0004] Digital video consumes the largest share of bandwidth on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands of digital video are likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining information in a Sequence Parameter Set (SPS) in a Multi-Functional Video Codec (VVC) Universal Media Application Format (CMAF) track, wherein a rule specifies that the value of the video availability information progressive source flag (vui_progressive_source_flag) field in the SPS should be equal to 1; and performing a conversion between visual media data and media data files based on the SPS.
[0006] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that vui_progressive_source_flag is included in the video availability information payload (vui_payload) structure.
[0007] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that vui_progressive_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is encoded and decoded according to progressive scan.
[0008] Alternatively, in any of the foregoing aspects, another implementation of this aspect provides that the rule specifies that the value of the video availability information interlaced source flag (vui_interlaced_source_flag) field in the SPS should be equal to 1.
[0009] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides that vui_interlaced_source_flag is included in the video availability information payload (vui_payload) structure.
[0010] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that vui_interlaced_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is coded according to interlaced.
[0011] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the SPS is contained in a VVC elementary stream, and wherein the VVC elementary stream is contained in the CMAF track.
[0012] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the converting includes encoding the visual media data into the media data file.
[0013] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the converting includes decoding the visual media data from the media data file.
[0014] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determine information in a sequence parameter set (SPS) in a versatile video coding (VVC) common media application format (CMAF) track, wherein a rule specifies that a value of a video usability information progressive source flag (vui_progressive_source_flag) field in the SPS should be equal to 1; and perform a conversion between visual media data and a media data file based on the SPS.
[0015] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that vui_progressive_source_flag is included in a video usability information payload (vui_payload) structure.
[0016] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that vui_progressive_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is coded according to progressive.
[0017] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the rule specifies that a value of a video usability information interlaced source flag (vui_interlaced_source_flag) field in the SPS should be equal to 1.
[0018] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that vui_interlaced_source_flag is included in a video usability information payload (vui_payload) structure.
[0019] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that vui_interlaced_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is coded according to interlaced
[0020] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the SPS is contained in a VVC elementary stream, and wherein the VVC elementary stream is contained in the CMAF track.
[0021] A third aspect relates to a non-transitory computer readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to: determine information in a sequence parameter set (SPS) in a versatile video coding (VVC) common media application format (CMAF) track, wherein a rule specifies that a value of a video usability information progressive source flag (vui_progressive_source_flag) field in the SPS should be equal to 1; perform a conversion between visual media data and a media data file based on the SPS.
[0022] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the vui_progressive_source_flag is included in a video usability information payload (vui_payload) structure.
[0023] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that vui_progressive_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is coded according to progressive
[0024] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the rule specifies that a value of a video usability information interlaced source flag (vui_interlaced_source_flag) field in the SPS should be equal to 1.
[0025] For purposes of clarity, any of the preceding embodiments can be combined with any one or more of the other preceding embodiments, to create new embodiments within the scope of the present disclosure.
[0026] These and other features will become more apparent from the following detailed description, taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF DRAWINGS
[0027] For a more complete understanding of the present disclosure, reference is now made to the following brief description of the drawings and detailed description, taken in connection with the accompanying figures and in which like numerals represent like parts.
[0028] Figure 1 is a schematic diagram illustrating an example common media application format (CMAF) track.
[0029] Figure 2 is a block diagram of an example video processing system.
[0030] Figure 3 is a block diagram of an example video processing device.
[0031] Figure 4 is a flowchart of an example method of video processing.
[0032] Figure 5 is a block diagram of an example video coding system.
[0033] Figure 6 is a block diagram of an example encoder.
[0034] Figure 7 is a block diagram of an example decoder.
[0035] Figure 8 is a schematic diagram of an example encoder. DETAILED DESCRIPTION
[0036] It should be understood at the outset that, although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or in existence or developed later. The disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but can be modified in any manner within the scope of the appended claims and their equivalents.
[0037] This patent document relates to video streaming. In particular, this document relates to specifying constraints on video encoding and encapsulation into media tracks and segments in file formats. Such file formats can include the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF). Such file formats can also include adaptive streaming representation formats, such as Dynamic Adaptive Streaming over Hypertext Transfer Protocol (HTTP) (DASH) and / or Common Media Application Format (CMAF). For media streaming systems, the ideas described herein can be applied individually or in various combinations, such systems being based on the DASH standard and related extensions and / or based on the CMAF standard and related extensions.
[0038] The disclosure includes the following acronyms. Adaptive color transform (ACT), adaptive loop filter (ALF), adaptive motion vector resolution (AMVR), adaptive parameter set (APS), access unit (AU), access unit delimiter (AUD), advanced video coding ((Rec. ITU-T H.264 | ISO / IEC 14496-10) (AVC), bi-prediction (B), bi-prediction with coding unit level weights (BCW), bi-directional optical flow (BDOF), block-based delta pulse code modulation (BDPCM), buffering period (BP), context-based adaptive binary arithmetic coding (CABAC), coding block (CB), constant bit rate (CBR), cross-component adaptive loop filter (CCALF), coded picture buffer (CPB), clean random access (CRA), cyclic redundancy check (CRC), coding tree block (CTB), coding tree unit (CTU), coding unit (CU), coded video sequence (CVS), decoding capability information (DCI), decoding initialization information (DII), decoded picture buffer (DPB), dependent random access point (DRAP), decoding unit (DU), decoding unit information (DUI), exponential-golomb (EG), k-th order exponential-golomb (EGk), end of bitstream (EOB), end of sequence (EOS), filler data (FD), first-in first-out (FIFO), fixed length (FL), green, blue, and red (GBR), general constraint information (GCI), gradual decoding refresh (GDR), geometry partition mode (GPM), high-efficiency video coding (also referred to as Rec. ITU-T H.265|ISO / IEC 23008-2)(HEVC), hypothetical reference decoder (HRD), hypothetical stream scheduler (HSS), intra (I), intra block copy (IBC), instantaneous decoding refresh (IDR), inter-layer reference picture (ILRP), intra random access point (IRAP), low-frequency non-separable transform (LFNST), least probable symbol (LPS), least significant bit (LSB), long-term reference picture (LTRP), luma mapping and chroma scaling (LMCS), matrix-based intra prediction (MIP), most probable symbol (MPS), most significant bit (MSB), multiple transform selection (MTS), motion vector prediction (MVP), network abstraction layer (NAL), output layer set (OLS), operation point (OP), operation point information (OPI), prediction (P), picture header (PH), picture order count (POC), picture parameter set (PPS), prediction refinement with optical flow (PROF), picture timing (PT), picture unit (PU), quantization parameter (QP), random access decodable leading pictures (RADL), random access skipped leading pictures (RASL), raw byte sequence payload (RBSP), red, green, and blue (RGB), reference picture list (RPL), sample adaptive offset (SAO), sample aspect ratio (SAR), supplemental enhancement information (SEI), slice header (SH), subpicture level information (SLI), string of data bits (SODB), sequence parameter set (SPS), short-term reference picture (STRP), subpicture-wise temporal sub-layer access (STSA), truncated rice (TR), variable bit rate (VBR), video coding layer (VCL), video parameter set (VPS), versatile supplemental enhancement information (also referred to as Rec. ITU-T H. 274 | ISO / IEC 23002-7) (VSEI), video usability information (VUI), and versatile video coding (also referred to as Rec. ITU-T H. 266 | ISO / IEC 23090-3) (VCC).
[0039] Video coding standards have evolved primarily through the development of international telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and ISO / International Electrotechnical Commission (IEC) standards. The ITU-T produced H.261 and H.263, ISO / IEC produced Motion Pictures Expert Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC. Since H.262, the video coding standards are based on hybrid video coding structure where temporal prediction plus transform coding is utilized. To explore video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was formed by the Video Coding Experts Group (VCEG) and MPEG. The JVET adopted a number of methods and put them into a reference software named Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project officially started, the JVET was later renamed as Joint Video Expert Team (JVET). VVC is a coding standard targeting at 50% bitrate reduction compared to HEVC. VVC has been finalized by the JVET.
[0040] The VVC standard (also known as ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplemental Enhancement Information (VSEI) standard (also known as ITU-T H.274 | ISO / IEC 23002-7) are designed for a wide range of applications such as television broadcasting, video conferencing, playback from storage media, adaptive bitrate streaming, video region extraction, compositing and merging of content from multiple coded video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive 360° immersive media. The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard developed by MPEG.
[0041] File format standards will be discussed below. Media streaming applications are typically based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hyper Text Transfer Protocol (HTTP) transport methods, and typically rely on file formats such as ISOBMFF. One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). Video can be coded in video formats such as AVC and / or HEVC. The coded video can be encapsulated in ISOBMFF tracks and included in DASH representations and segments. For content selection purposes, important information about the video bitstream, such as profiles, tiers, and levels, etc., can be exposed as file format level metadata and / or in the DASH Media Presentation Description (MPD). Such information can be used, for example, to select appropriate media segments, for initialization at the beginning of a streaming session and for stream adaptation during the streaming session.
[0042] Similarly, when using image formats with ISOBMFF, file format specifications specific to the image format can be employed, such as AVC image file format and HEVC image file format. MPEG is developing a VVC video file format, which is a file format based on ISOBMFF for storing VVC video content. MPEG is also developing a VVC image file format based on ISOBMFF, which is a file format for storing image content coded using VVC.
[0043] File format standards are now discussed. Media streaming applications can be based on Internet Protocol (IP), Transmission Control Protocol (TCP), and HyperText Transfer Protocol (HTTP) transport mechanisms. Such media streaming applications can also rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). To use a video format with ISOBMFF and DASH, a file format specification specific to the video format can be employed to encapsulate video content in ISOBMFF tracks and DASH representations and segments. Such file format specifications can include the AVC file format and the HEVC file format. For the purpose of content selection, important information about the video bitstream, such as profiles, layers, and levels, can be exposed in file format level metadata and / or in a DASH Media Presentation Description (MPD). For example, content selection can include selecting appropriate media segments for initialization at the beginning of a streaming session and for stream adaptation during the streaming session. Similarly, to use an image format with ISOBMFF, a file format specification specific to the image format can be employed, such as the AVC image file format and the HEVC image file format. The VVC video file format is a file format based on ISOBMFF for storing VVC video content. The VVC video file format is developed by MPEG. The VVC image file format is a file format based on ISOBMFF for storing image content coded using VVC. The VVC image file format is also developed by MPEG.
[0044] DASH is now discussed. In DASH, video and / or audio data of multimedia content can have multiple representations. Different representations can correspond to different coding characteristics, such as different profiles or levels of a video coding standard, different bitrates, different spatial resolutions, and so on. A manifest of such representations can be defined in a Media Presentation Description (MPD) data structure. A media presentation can correspond to a structured collection of data accessible to a DASH streaming client device. A DASH streaming client device can request and download media data information to present a streaming service to a user of the client device. A media presentation can be described in an MPD data structure, which can include updates to the MPD.
[0045] A media presentation can contain a sequence of one or more periods. Each period can extend to the beginning of the next period, or to the end of the media presentation in the case of the last period. Each period can contain one or more representations of the same media content. A representation can be one of a plurality of alternative encoded versions of audio, video, timed text, or other such data. The representations can differ according to encoding type, such as bitrate, resolution, codec for video data, and bitrate, language, and / or codec for audio data. The term representation can be used to refer to a portion of encoded audio or video data that corresponds to a particular period of multimedia content and is encoded in a particular manner.
[0046] Representations of a particular period can be assigned to groups indicated by attributes in the MPD, which indicate the adaptation set to which the representation belongs. Representations in the same adaptation set are generally considered to be mutually substitutable. Thus, a client device can dynamically and seamlessly switch between the representations, such as to perform bandwidth adaptation. For example, each representation of video data for a particular period can be assigned to the same adaptation set, such that any representation can be selected for decoding to present media data, such as video data or audio data, of the multimedia content for the corresponding period. In some examples, media content within a period can be represented by a combination of one representation from group 0, if present, or at most one representation from each non-zero group. Timing data for each representation of a period can be expressed relative to the start time of the period.
[0047] A representation can include one or more segments. Each representation can include an initialization segment, or each segment of a representation can be self-initializing. When present, an initialization segment can contain initialization information for accessing the representation. Typically, an initialization segment does not contain media data. Segments can be uniquely referenced by an identifier, such as a uniform resource locator (URL), uniform resource name (URN), or uniform resource identifier (URI). An MPD can provide an identifier for each segment. In some examples, an MPD can also provide a byte range in the form of a range attribute, which can correspond to data of a segment within a file that is accessible by a URL, URN, or URI.
[0048] For different types of media data, different representations can be selected for substantially simultaneous retrieval. For example, a client device can select an audio representation, a video representation, and a timed text representation from which to retrieve segments. In some examples, a client device can select a particular adaptation set to perform bandwidth adaptation. For example, a client device can select an adaptation set that includes video representations, an adaptation set that includes audio representations, and / or an adaptation set that includes timed text. In examples, a client device can select an adaptation set for certain types of media, such as video, and directly select a representation for other types of media, such as audio and / or timed text.
[0049] An example DASH streaming process can be shown by the following steps. A client obtains an MPD. Then, the client estimates the downlink bandwidth and selects video and audio representations based on the estimated downlink bandwidth, codecs, decoding capabilities, display size, audio language settings, etc. Until the end of the media presentation is reached, the client requests media segments of the selected representations and presents the streaming content to the user. The client continuously estimates the downlink bandwidth. When the bandwidth changes significantly, e.g., by becoming lower or becoming higher, the client selects a different video representation to match the newly estimated bandwidth and continues to download segments with the updated downlink bandwidth.
[0050] CMAF is now discussed. CMAF specifies a set of constraints on media encoding and encapsulation into ISOBMFF tracks, ISOBMFF segments, ISOBMFF fragments, DASH representations, and / or CMAF tracks, CMAF fragments, etc. Such constraints are per encapsulation for each interoperability point defined as a media profile. The main goal of the CMAF development is to enable the reuse of the same media content encoded using a specific codec (e.g., AVC for video) and encapsulated in a specific format (e.g., ISOBMFF) through both the DASH and Apple HTTP Live Streaming (HLS) two separate media streaming worlds.
[0051] Decoding capability information (DCI) in VVC is now discussed. A DCI NAL unit contains bitstream level profile, tier, and level (PTL) information. The DCI NAL unit includes one or more PTL syntax structures that can be used during session negotiation between a sender and a receiver of a VVC bitstream. When a DCI NAL unit is present in a VVC bitstream, each output layer set (OLS) in the CVS of the bitstream shall conform to the PTL information carried in at least one PTL structure in the DCI NAL unit. In AVC and HEVC, the PTL information for session negotiation is available in the SPS (for HEVC and AVC) and VPS (for HEVC layered extension). This design of conveying PTL information for session negotiation in HEVC and AVC has drawbacks because the scope of SPS and VPS is within a CVS, not the entire bitstream. This can cause sender-receiver session initiation to suffer from re-initiation during bitstream streaming of each new CVS. DCI solves this problem because DCI carries bitstream level information, thus it can guarantee adherence to the indicated decoding capabilities until the end of the bitstream.
[0052] Video parameter set (VPS) in VVC is now discussed. A VVC bitstream can contain a video parameter set (VPS) that contains information for the description of the operation of the decoding process of a scalable bitstream and the set of output layers (OLS). An OLS is a set of layers in a bitstream where one or more layers are designated to be output from the decoder. Other layers identified in the OLS can also be decoded in order to decode the output layers, although such layers are not designated to be output. Much of the information contained in the VPS can be used in the system for purposes such as session negotiation and content selection. The VPS is introduced to handle multi-layer bitstreams. For single-layer VVC bitstreams, the presence of a VPS in the CVS is optional. This is because the information contained in the VPS is not necessary for the operation of the decoding process of the bitstream. The absence of a VPS in the CVS is indicated by referring to the VPS identifier (ID) in the SPS being equal to 0, in which case default values are inferred for the VPS parameters.
[0053] Sequence parameter set (SPS) in VVC is now discussed. The SPS conveys sequence-level information that is shared by all pictures in an entire coded layer video sequence (CLVS). This includes PTL indicators, picture formats, feature and / or tool control flags, coding, prediction, and / or pr transform block structures and hierarchies, candidate RPLs that the encoder can refer to, etc. The picture formats can include color sampling format, maximum picture width, maximum picture height, and bit depth. In most applications, only one or a few SPSs are employed for the entire bitstream. Therefore, there is no need to update the SPSs within the bitstream. Updating an SPS can include sending a new SPS using the SPS ID of an existing SPS but with different values for certain parameters. Pictures from a particular layer that refer to an SPS with a different SPS ID or with the same SPS ID but with different SPS content belong to different CLVSs. Like AVC and HEVC, the SPS can be transmitted in-band or using a mix of in-band and out-of-band signaling. In-band signaling indicates that data such as the SPS is transmitted with the coded pictures, and out-of-band signaling indicates that data such as the SPS is not transmitted with the coded pictures.
[0054] Picture Parameter Set (PPS) in VVC is now discussed. PPS conveys picture-level information shared by all slices of a picture. Such information can also be shared across multiple pictures. This includes features and / or tool on / off flags, picture width and height, default RPL size, tile and slice configuration, etc. By design, two consecutive pictures can refer to two different PPSs. This can result in a large number of PPSs used within a CLVS. In practice, the number of PPSs for the entire bitstream can not be high, as PPS is designed to carry parameters that do not change frequently and can be applicable to multiple pictures. Therefore, there can be no need to update PPSs within a CLVS or even the entire bitstream. Adaptive Parameter Set (APS) can be used for parameters that can be applicable to multiple pictures but are expected to change frequently between different pictures. Like SPS, PPS can be transmitted in-band, out-of-band, or using a mix of in-band and out-of-band signaling. One basic design principle regarding which picture-level parameters should be included in PPS and which should be in APS is the frequency at which the parameters can change. Therefore, parameters that change frequently are not included in PPS to avoid requiring PPS update, which in typical use cases would not allow out-of-band transmission of PPS.
[0055] Adaptive Parameter Set (APS) in VVC is now discussed. APS conveys picture- and / or slice-level information that can be shared by multiple slices of a picture and / or slices of different pictures but can change frequently across pictures. APS supports information with a large number of variables that are not suitable to be included in PPS. Three types of parameters are included in APS, including Adaptive Loop Filter (ALF) parameters, Luma Mapping and Chroma Scaling (LMCS) parameters, and scaling list parameters. APS can be carried in two different NAL unit types, which can precede or follow the associated slice as a prefix or suffix. The latter can be helpful in ultra-low latency scenarios, such as allowing an encoder to send slices of a picture before generating ALF parameters based on the picture, which will be used by subsequent pictures in decoding order.
[0056] Picture header (PH) is now discussed. For each PU, there is a picture header (PH) structure. The PH is either present in a separate PH NAL unit or is included in the slice header (SH). If a PU includes only one slice, the PH can only be included in the SH. To simplify the design, within a CLVS, the PH can only be all in PH NAL units or all in SHs. When the PH is in the SH, there is no PH NAL unit in the CLVS. There are two purposes for designing the PH. First, the PH helps to reduce the signaling overhead of the SHs for pictures that contain multiple slices. The PH does this by carrying all the parameters that have the same value for all slices of a picture, thus preventing the repetition of the same parameters in each SH. These include the IRAP and / or GDR picture indication, the inter and / or intra slice allowed flags, and information related to POC, RPL, deblocking filter, SAO, ALF, LMCS, scaling list, QP delta, weighted prediction, coding block partitioning, virtual boundary, collocated picture, etc. Second, the PH helps the decoder to identify the first slice of each coded picture that contains multiple slices. Since there is one and only one PH for each PU, when the decoder receives a PH NAL unit, the decoder knows that the next VCL NAL unit is the first slice of the picture.
[0057] Operation point information (OPI) is now discussed. The decoding process of HEVC and VVC has similar input variables to set the decoding operation point. These include the target OLS and the highest sublayer of the bitstream to be decoded through the decoder application programming interface (API). However, in scenarios where the layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder application programming interface (API) to the application, the decoder can not be able to correctly determine the operation point for processing the bitstream. As a result, the decoder can not be able to infer the properties of the pictures in the bitstream, such as the proper buffer allocation of the decoded pictures and whether to output individual pictures. To address this issue, VVC includes a mode to indicate these two variables within the bitstream through the OPI NAL unit. At the beginning of the bitstream and in the individual CVS of the bitstream, the OPI NAL unit informs the decoder of the target OLS and the highest sublayer of the bitstream to be decoded. In the case where the OPI NAL unit is present and the operation point is also provided to the decoder via the decoder API information, the decoder API information takes precedence. For example, the application can have more updated information related to the target OLS and the sublayer. In the case where there is no decoder API and any OPI NAL unit in the bitstream, a suitable fallback selection is specified in VVC to allow proper decoder operation.
[0058] An example CMAF specification is now discussed. A VVC video CMAF track can be described as follows. A VVC CMAF track shall conform to the requirements of a NAL structured video CMAF track. In addition, the CMAF track can conform to all other requirements described herein. If the CMAF track conforms to these requirements, the CMAF track is referred to as a VVC video CMAF track and can use the brand "cvvc". VVC video track constraints are also discussed. In an example, the VVC video CMAF switching set constraints are as follows. Each CMAF track in the CMAF switching set shall conform to the VVC video CMAF track defined herein. The VVC video CMAF switching set shall conform to the constraints of a NAL structured video CMAF switching set.
[0059] Visual sample entries are now discussed. The syntax and values of a VVC video track visual sample entry shall conform to the VVCSampleEntry('vvc1') or VVCSampleEntry('vvci') sample entry. Constraints on a VVC elementary stream are now discussed. With respect to VPS, each VVC video media sample in a CMAF track shall refer to an SPS with sps_video_parameter_set_id equal to 0 (in this case, there is no VPS in the elementary stream), or shall refer to a VPS in the CMAF header sample entry. If present, the following additional constraints apply. For each profile_tier_level() structure in the VPS, the values of the following fields shall not change throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i].
[0060] SPS NAL units that occur within a CMAF VVC track shall conform to the constraints herein and have the following additional constraints. The following fields shall have the following predetermined values: first, vui_parameters_present_flag shall be set to 1; second, if a profile_tier_level() structure is present in the SPS, the following fields shall not change throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i].
[0061] Now discuss the picture crop parameters. The SPS and PPS crop parameters conf win top offset and conf win left offset shall be set to 0. The SPS and PPS crop parameters conf win bottom offset and conf win right offset can be set to a value other than 0. If set to a non-zero value, such syntax elements are intended to be used by CMAF players to remove video spatial sample points that are not intended to be displayed.
[0062] Now discuss the video codec parameters. VVC signaling of codec parameters (information) is described below. The presentation application shall use parameter signaling to signal the video codec profile and level for each VVC track and CMAF switching set. Encryption is also discussed. Encryption of CMAF VVC tracks and CMAF VVC switching sets shall use the "cenc" AES-CTR scheme or the "cbcs" AES-CBC sub-sample mode encryption scheme. Furthermore, if the "cbcs" mode of general encryption uses pattern encryption, the pattern block length 10 and encryption: skip mode 1 :9 shall be applied.
[0063] The following are example technical problems solved by the disclosed technical solutions. For example, in the example VVC CMAF design, it can be desirable to signal the profile, tier, and level in the VPS and SPS so that it can not change throughout the VVC elementary stream. However, for VVC bitstreams, the DCI NAL unit can instead be used to convey the decoding capability required for the entire bitstream, while allowing the profile, tier, and level to differ between CVS within the bitstream. This would allow for greater flexibility, resulting in less transcoding and other processes required in content preparation.
[0064] Disclosed herein are mechanisms to address one or more of the above-listed issues. For example, a VVC elementary stream, also referred to as a VVC bitstream, can be included in a VVC CMAF track. The VVC elementary stream can include one or more CVSs. The profile, tier, and level (PTL) information of the bitstream can vary among the CVSs in the same bitstream. To allow this functionality, the PTL information can be signaled in the DCI NAL unit, VPS, and / or SPS, as long as the corresponding constraints are maintained. In an example, the DCI NAL unit needs to be included in the CMAF track. In an example, when multiple DCI NAL units are included in the CMAF track, all DCI NAL units can need to include the same content. In another example, the CMAF track can include only a single DCI NAL unit. In an example, the DCI NAL unit can need to be included in the CMAF header sample entry. In various examples, the DCI NAL unit can contain a dci num ptls minusl (dci num ptls minusl) field of the PTL, a ptl frame only constraint flag (ptl frame only constraint flag) field, and a ptl multilayer enabled flag (ptl multilayer enabled flag) of the PTL, which need to be equal to 0, 1, and 0, respectively. In an example, the CMAF track is constrained to contain a single VPS. In an example, when there is no DCI NAL unit, various PTL-related information in the VPS needs to be set to a predetermined value and / or needs to be kept the same among the CVSs, as discussed further below. In an example, various PTL-related information in the SPS needs to be set to a predetermined value and / or needs to be kept the same among the CVSs, as discussed further below. In an example, the timing-related hypothetical reference decoder (HRD) parameters can also need to be kept the same among the CVSs.
[0065] Figure 1 FIG. 1 is a diagram illustrating an example CMAF track 100. The CMAF track 100 is a track of video data that has been encapsulated based on the constraints specified in the CMAF standard. The CMAF track 100 is constrained to support delivery and decoding by a wide range of client devices according to adaptive streaming. In adaptive streaming, a media profile describes multiple different interchangeable representations, which allows a client device to select a desired representation based on decoder capabilities and / or current network conditions. The CMAF track 100 can support such functionality by containing representations that are constrained to be decodable by a client that is capable of decoding a corresponding profile, tier, and level (PTL), is capable of using a corresponding codec tool, and / or satisfies other predetermined constraints.
[0066] The CMAF track 123 can contain many types of decodable video streams. In the present example, the CMAF track 123 contains a VVC stream 121. The VVC stream 121, also referred to as a bitstream, is a stream of video data that has been coded according to the VVC standard. For example, the VVC stream 121 can include a stream of coded pictures and associated syntax that describes the coding process and / or other data useful to a decoder. The VVC stream 121 can contain one or more CVSs 117. A CVS 117 is a sequence of access units (AUs) in decoding order. An AU is a set of one or more pictures with a corresponding output / display time. As such, the CVS 117 contains a series of related pictures and corresponding syntax to support decoding and / or describe the pictures.
[0067] The CVS 117 can include a DCI NAL unit 115, a VPS 111, an SPS 113, and / or coded video 119. The DCI NAL unit 115 contains information that describes the requirements for decoding the video data in the CVS 117 and / or the entire VVC stream 121. The DCI NAL unit 115 is optional and can be omitted in some VVC streams 121 and / or CVSs 117. It should be noted that although depicted as part of the VVC stream 121, the DCI NAL unit 115 can also be included in a CMAF header sample entry in the CMAF track 123 in some examples. The VPS 111 can contain data related to the entire VVC stream 121. For example, the VPS 111 can contain data related output layer sets (OLSs), layers, and / or sub-layers used in the VVC stream 121. The VPS 111 is optional and can be omitted in some VVC streams 121 and / or CVSs 117. The SPS 113 contains sequence data common to all pictures in the CVS 117 contained in the VVC stream 121. Parameters in the SPS 113 can include picture size, bit depth, coding tool parameters, bit rate limits, etc. The SPS 113 should be contained in at least one CVS 117. However, multiple CVSs 117 can refer to the same SPS 113. Thus, the VVC stream 121 should contain one or more SPSs 113. The coded video 119 includes pictures that are coded according to VVC and corresponding syntax.
[0068] The present disclosure relates to constraints applied to syntax elements contained in DCI NAL units 115, VPS 111, and / or SPS 113. In an example, a DCI NAL unit 115 can be required to be present in a CMAF track 123. In an example, when more than one DCI NAL unit 115 is present in a single CMAF track 123, all such DCI NAL units 115 can be required to include the same content. In some examples, a CMAF 123 track can be constrained to include one and only one DCI NAL unit 115. In this case, multiple CVS 117 of video content can be described by a single DCI NAL unit 115. When present, a DCI NAL unit 115 can contain a DCI number of PTLs minus 1 (dci_num_ptls_minus1) 132 and / or a PTL syntax (profile_tier_level) structure 130. The dci_num_ptls_minus1 132 can specify, in a minus 1 format, the number of profile_tier_level structures 130 contained in the DCI NAL unit 115. The minus 1 format indicates that the value contained by the syntax element is 1 less than the actual value, so 1 is added to the value contained in the syntax element to determine the actual value. In an example, the dci_num_ptls_minus1 132 can be constrained to be equal to 0, which indicates a single profile_tier_level structure 130. This indicates that the CMAF track 123 contains video that adheres to a single set of PTL information. Depending on the example, the profile_tier_level structure 130 can be contained in the DCI NAL unit 115, VPS 111, and / or SPS 113, and will be discussed in more detail below.
[0069] In an example, the CMAF track 123 is constrained to contain one and only one VPS 111. In this case, multiple CVSs 117 can be described by a single VPS 111. The VPS 111 can contain a vps max layers minus 1 (vps max layers minus 1) field 134, a vps num ptls minus 1 (vps num ptls minus 1) field 133, a general timing hrd parameters (general timing hrd parameters) structure 131, and a profile tier level structure 130. The vps max layers minus 1 field 134 indicates the number of layers specified by the VPS 111 in a minus 1 format. In an example, the vps max layers minus 1 field 134 can be constrained to contain the value 0, which indicates that the VPS 111 describes a single layer. The vps num ptls minus 1 field 133 can specify the number of profile tier level structures 130 contained in the VPS 111 in a minus 1 format. In an example, the vps num ptls minus 1 field 133 can be constrained to contain the value 0, which indicates to the VPS 111 a single PTL information set.
[0070] Depending on the example, general timing hrd parameters 131 can be included in VPS 111 and / or SPS 113. For example, when VPS 111 is included, VPS 111 can include general timing hrd parameters 131. When VPS 111 is not included, SPS can include general timing hrd parameters 131. General timing hrd parameters 131 include timing related parameters used by a HRD operating at the encoder. Generally, a HRD can use the HRD parameters to check that VVC stream 121 conforms to the VVC standard. General timing hrd parameters 131 indicate to the encoder timing parameters related to coded video 119. For example, general timing hrd parameters 131 can indicate how fast each picture should be decoded and reconstructed by the decoder for accurate display. In an example, general timing hrd parameters 131 can include a time scale field and a num units in tick field in ticks. The time scale field indicates the number of time units that pass in one second, where a time unit corresponds to the picture rate frequency of the video signal. Num units in tick indicates the number of time units of a clock operating at the frequency of time scale in Hertz (Hz) that corresponds to one increment called a clock tick. In an example, the values of num units in tick and time scale in general timing hrd parameters 131 are constrained to remain constant between CVSs 117 in the same VVC stream 121. In an example, the values of num units in tick and time scale in general timing hrd parameters 131 are constrained to remain constant for the entire CMAF track 123.
[0071] The SPS 113 can include the general timing hrd parameters 131 as discussed above, e.g., when the VPS 111 is not included. The SPS 113 can also include the profile tier level structure 130, e.g., when the DCI NAL unit 115 and / or the VPS 111 is not included. The SPS 113 can also include a video usability information payload (vui_payload) structure 135. The vui_payload structure 135 includes information describing how a decoder should use the coded video 119. For example, the vui_payload structure 135 can include a video usability information progressive source flag (vui_progressive_source_flag) field 139 and a video usability information interlaced source flag (vui_interlaced_source_flag) field 138. The vui_progressive_source_flag field 139 can be set to indicate whether the video in the CMAF track 123 is coded according to progressive scanning. The vui_interlaced_source_flag field 138 can be set to indicate whether the video in the CMAF track 123 is coded according to interlaced. In an example, the vui_interlaced_source_flag field 138, the vui_progressive_source_flag field 139, or both, can need to be set to 1. This indicates that the coded video 119 is coded according to interlaced, progressive, or both.
[0072] As discussed above, the DCI NAL unit 115, the VPS 111, and / or the SPS 113 can include the profile tier level structure 130. The profile tier level structure 130 includes information about the profile, tier, and level used to code the coded video. The profile indicates the profile used to code the coded video. Different profiles have different coding characteristics (e.g., availability of different coding tools), such as different bit depths, different chroma sampling formats, availability of cross-component prediction, availability of intra smoothing disabling. The tier indicates whether the coded video 119 is coded according to a high tier or a main tier, thus a requirement for general application requirements for the coded video to be coded. The level indicates constraints on the coded video 119, such as a maximum bit rate, a maximum picture size, a maximum sample rate, a resolution at a highest frame rate, a maximum number of tiles, a maximum number of slices per picture, etc. Thus, the PTL information in the profile tier level structure 130 describes the capabilities a decoder must have in order to decode and display the coded video 119.
[0073] The profile_tier_level structure 130 can include a PTL frame only constraint flag (ptl_frame_only_constraint_flag) field 141, a PTL multi-layer enabled flag (ptl_multilayer_enabled_flag) 143, a general profile identification code (general_profile_idc) 145, a general tier flag (general_tier_flag) 147, a general level identification code (general_level_idc) 149, a number of sub-profiles (num_sub_profiles) 142, and / or a general sub-profile identification code for each ith interoperability indicator (general_sub_profile_idc[i]) 144. The ptl_frame_only_constraint_flag field 141 specifies whether the CVS 117 conveys pictures that represent frames (e.g., complete screen images) or fields (e.g., partial screen images intended to be combined to fill a screen). In an example, a constraint can require the ptl_frame_only_constraint_flag field 141 to be set to 1, which indicates that the coded video 119 includes pictures that are coded as frames. The ptl_multilayer_enabled_flag 143 indicates whether the coded video 119 is coded in multiple layers. In an example, the ptl_multilayer_enabled_flag 143 is set to 0, which indicates that the coded video 119 is coded in a single layer.
[0074] general_profile_idc 145, general_tier_flag 147, and general_level_idc 149 indicate the profile, tier, and level of the coded video 119, respectively. general_sub_profile_idc[i] 144 indicates the 0 to i value of the interoperability indicator. num_sub_profiles 142 indicates the number of syntax elements contained in general_sub_profile_idc[i] 144. In an example, the values contained in general_profile_idc 145, general_tier_flag 147, general_level_idc 149, num_sub_profiles 142, and general_sub_profile_idc[i] 144 need to remain constant across CVSs 117 in the same VVC stream. In another example, the values contained in general_profile_idc 145, general_tier_flag 147, general_level_idc 149, num_sub_profiles 142, and general_sub_profile_idc[i] 144 need to remain constant across CMAF tracks 123.
[0075] To address the above issues and other issues, methods summarized as follows are disclosed. These items should be considered examples of explaining general concepts and should not be construed in a narrow way. Furthermore, these items can be applied individually or in any combination.
[0076] Example 1
[0077] In one example, the rule can specify that a DCI NAL unit should be present in a VVC CMAF track.
[0078] Example 2
[0079] In one example, the rule can specify that a DCI NAL unit should be present in a VVC CMAF track.
[0080] Example 3
[0081] In one example, the rule can specify that all DCI NAL units present in a VVC CMAF track should have the same content.
[0082] Example 4
[0083] In one example, the rule can specify that one and only one DCI NAL unit should be present in a VVC CMAF track.
[0084] Example 5
[0085] In one example, the rule can specify that a DCI NAL unit shall be present in the CMAF header sample entry when the DCI NAL unit is present in the VVC CMAF track.
[0086] Example 6
[0087] In one example, the rule can specify that the value of the dci num ptls minusl field in the DCI NAL unit in the VVC CMAF track shall be equal to 0.
[0088] Example 7
[0089] In one example, the rule can specify that the value of the ptl frame only constraint flag field in the profile tier level( ) structure in the DCI NAL unit in the VVC CMAF track shall be equal to 1.
[0090] Example 8
[0091] In one example, the rule can specify that the value of the ptl multilayer enabled flag field in the profile tier level( ) structure in the DCI NAL unit in the VVC CMAF track shall be equal to 0.
[0092] Example 9
[0093] In one example, the rule can specify that one and only one VPS unit shall be present in the VVC CMAF track.
[0094] Example 10
[0095] In one example, the rule can specify that when no DCI NAL unit is present in the VVC CMAF track and one or more VPS is present in the VVC CMAF track, one or more of the following constraints apply. The constraints include that for each VPS, the value of the vps max layers minusl field shall be equal to 0 and for each VPS, the value of the vps num ptls minusl field shall be equal to 0.
[0096] In an example, the following constraint applies to the profile_tier_level( ) structure in each VPS. Such constraints include the value of the ptl_frame_only_constraint_flag field shall be equal to 1; and the value of the ptl_multilayer_enabled_flag field shall be equal to 0.
[0097] In an example, the value of each of the following fields in the profile_tier_level( ) structure of the referenced VPS shall not change from one coded video sequence to another coded video sequence throughout a VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[ i ] for each i value. In an example, the rule can require that the value of each of these fields be the same for all VPSs present in a VVC CMAF track.
[0098] Example 11
[0099] In an example, the rule can specify that the value of the vui_progressive_source_flag field in the vui_payload( ) structure in the SPS in a VVC CMAF track shall be equal to 1.
[0100] Example 12
[0101] In an example, the rule can specify that the value of the vui_interlaced_source_flag field in the vui_payload( ) structure in the SPS in a VVC CMAF track shall be equal to 1.
[0102] Example 13
[0103] In one example, the rule can specify that the value of each of the following fields in the profile_tier_level() structure of the referenced SPS shall not change from one coded video sequence to another coded video sequence throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i] for each i value when no DCI NAL unit is present and no VPS is present in the VVC CMAF track. In an example, the rule can require that the value of each of these fields be the same for all SPSs present in the VVC CMAF track.
[0104] Example 14
[0105] In one example, the rule can specify that the value of each of the following fields in the general_timing_hrd_parameters() structure (when present) in the referenced VPS or SPS shall not change from one coded video sequence to another coded video sequence throughout the VVC elementary stream: num_units_in_tick; and time_scale. In an example, the rule can require that the value of each of these fields be the same for all general_timing_hrd_parameters() structures in the VPS or SPS present in the VVC CMAF track.
[0106] An embodiment of the previous example is now described. This embodiment can be applied to CMAF. Most of the relevant parts that have been added or modified with respect to the VVC CMAF specification are shown in bold underlined font, and some deleted parts are shown in bold italic font. There can be some other changes that are editorial in nature and thus not highlighted.
[0107] X.1 VVC video CMAF track. A VVC CMAF track shall conform to the requirements of a NAL-structured video CMAF track. In addition, it shall conform to all the remaining requirements in this annex. If a CMAF track conforms to these requirements, it is referred to as a VVC video CMAF track and can use the brand "cvvc".
[0108] X.2.1 VVC video CMAF switch constraint. Each CMAF track in a CMAF switch set shall conform to the VVC video CMAF track as defined in clause X.1. The VVC video CMAF switch set shall conform to the constraints of NAL structured video CMAF switch set.
[0109] X.2.2 Visual sample entry. The syntax and values of the visual sample entry of a VVC video track shall conform to the VVC Sample Entry ("vvc1") or VVC Sample Entry ("vvc2") as defined in ISO / IEC 14496-15. Sample Entry.
[0110]
[0111] X.3.2 Video parameter set (VPS). Each VVC video media sample in a CMAF track shall refer to an SPS with sps_video_parameter_set_id equal to 0, in which case there is no VPS in the elementary stream, or shall refer to The following additional constraints apply: profile_tier_level() structure throughout the VVC elementary stream general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, general_sub_profile_idc[i].
[0112] X.3.3 Sequence parameter set (SPS). Sequence parameter set NAL units occurring within a CMAF VVC track shall conform to the following additional constraints: The following fields shall have the following predetermined values: vui_parameters_present_flag shall be set equal to 1. throughout the VVC elementary stream The video sequence to another codec shall not change: general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] for each i value.
[0113]
[0114] X.3.5 Picture Cropping Parameters. SPS and PPS Cropping Parameters conf_win_top_offset and conf_win_left_offset / shall be set to 0. SPS and PPS Cropping Parameters conf_win_bottom_offset and conf_win_right_offset may be set to a value other than 0. If set to a non-zero value, it is intended to be used by a CMAF player to remove video spatial samples that are not intended to be displayed.
[0115] Figure 2 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein can be implemented. Various implementations can include some or all of the components of the system 4000. The system 4000 can include an input 4002 for receiving video content. The video content can be received in a raw or uncompressed format, e.g., 8 or 10 bit multi-component pixel values, or can be in a compressed or encoded format. The input 4002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.
[0116] The system 4000 can include a codec component 4004 that can implement various coding or encoding methods described in this document. The codec component 4004 can reduce the average bitrate of video from the input 4002 to the output of the codec component 4004 to produce a coded representation of the video. The codec techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 can be stored, or transmitted via a communication connection as represented by the component 4006. The stored or communicated bitstream (or coded) representation of the video received at the input 4002 can be used by the component 4008 to generate pixel values or displayable video to a display interface 4010. The process of generating user-viewable video from a bitstream representation is sometimes referred to as video decompression. Also, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that the coding tools or operations are used at an encoder, and corresponding decoding tools or operations that reverse the results of the coding will be performed by a decoder.
[0117] Examples of peripheral bus interfaces or display interfaces can include Universal Serial Bus (USB), or High Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The techniques described in this document can be embodied in various electronic devices such as mobile telephones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0118] Figure 3 is a block diagram of an example video processing device 4100. The device 4100 can be used to implement one or more of the methods described herein. The device 4100 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 4100 can include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processor(s) 4102 can be configured to implement one or more methods described in this document. The memory(ies) 4104 can be used for storing data and code used for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement, in hardware circuitry, some of the techniques described in this document. In some embodiments, the video processing circuitry 4106 can be included at least in part in the processor 4102 (e.g., a graphics co-processor).
[0119] Figure 4FIG. 42 is a flowchart of an example method 4200 of video processing. The method 4200 includes determining, at step 4202, information in an SPS in a VVC CMAF track. In an example, a rule specifies that a value of a vui_progressive_source_flag field in the SPS shall be equal to 1. In an example, the vui_progressive_source_flag is included in a vui_payload structure. In an example, the vui_progressive_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is coded according to progressive scanning. In an example, the rule specifies that a value of a vui_interlaced_source_flag field in the SPS shall be equal to 1. In an example, the vui_interlaced_source_flag is included in the vui_payload structure. In an example, the vui_interlaced_source_flag is equal to 1 to indicate that the video in the VVC CMAF track is coded according to interlaced scanning. In an example, the SPS is contained in a VVC elementary stream, and wherein the VVC elementary stream is contained in a CMAF track.
[0120] At step 4204, a conversion is performed between the visual media data and the media data file based on the SPS. When the method 4200 is performed on an encoder, the conversion includes generating the media data file from the visual media data. The conversion includes determining the SPS and encoding it as a bitstream contained in the CMAF track. When the method 4200 is performed on a decoder, the conversion includes parsing and decoding the bitstream in the CMAF track according to the SPS to obtain the visual media data.
[0121] It should be noted that the method 4200 can be implemented in an apparatus for processing video data, the apparatus comprising a processor and a non-transitory memory having stored thereon instructions, such as the video encoder 4400, the video decoder 4500, and / or the encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform the method 4200. Furthermore, the method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use by a video coding device. The computer program product includes computer executable instructions stored on the non-transitory computer-readable medium such that when executed by a processor cause the video coding device to perform the method 4200.
[0122] Figure 5is a block diagram illustrating an example video coding system 4300 that can utilize the techniques of this disclosure. Video coding system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where the source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, and the destination device 4320 can be referred to as a video decoding device.
[0123] Source device 4310 can include a video source 4312, a video encoder 4314, and an input / output (VO) interface 4316. Video source 4312 can include a source such as a video capture device, an interface to receive video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data can comprise one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream can include a sequence of bits that form a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. VO interface 4316 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be transmitted directly to destination device 4320 via VO interface 4316 and network 4330. The encoded video data can also be stored onto a storage medium / server 4340 for access by destination device 4320.
[0124] Destination device 4320 can include an VO interface 4326, a video decoder 4324, and a display device 4322. VO interface 4326 can include a receiver and / or a modem. VO interface 4326 can acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 can decode the encoded video data. Display device 4322 can display the decoded video data to a user. Display device 4322 can be integrated with destination device 4320, or can be external to destination device 4320 which can be configured to interface with an external display device.
[0125] Video encoder 4314 and video decoder 4324 can operate according to a video compression standard, such as High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVM) standard, and other current and / or further standards.
[0126] Figure 6 is a block diagram illustrating an example of a video encoder 4400, which can be Figure 5The video encoder 4314 in the illustrated system 4300. The video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. The video encoder 4400 includes a number of functional components. The techniques described in this disclosure can be shared amongst the various components of the video encoder 4400. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0127] The functional components of the video encoder 4400 can include a partitioning unit 4401, a prediction unit 4402 (which can include a mode select unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy encoding unit 4414.
[0128] In other examples, the video encoder 4400 can include more, less, or different functional components. In examples, the prediction unit 4402 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0129] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405, can be highly integrated, but are represented separately in the example of the video encoder 4400 for explanatory purposes.
[0130] The partitioning unit 4401 can partition a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0131] The mode select unit 4403 can select one of the coding modes (e.g., intra or inter) based on the error results and provide the resulting intra-coded or inter-coded block to the residual generation unit 4407 for generation of residual block data and to the reconstruction unit 4412 for reconstruction of the encoded block for use as a reference picture. In some examples, the mode select unit 4403 can select a combination of intra and inter prediction modes (CIIP), where the prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode select unit 4403 can also select a resolution for the motion vector of the block (e.g., sub-pixel or integer pixel precision).
[0132] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.
[0133] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0134] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for reference images in list 0 or list 1 for reference video blocks of the current video block. Motion estimation unit 4404 can then generate a reference index and motion vector indicating the reference image in list 0 or list 1, which contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0135] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images in list 0, and can also search for another reference video block for the current video block in list 1. Motion estimation unit 4404 can then generate a reference index and a motion vector, where the reference index indicates the reference images in lists 0 and 1 containing the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0136] In some examples, the motion estimation unit 4404 can output a complete set of motion information for the decoder's decoding process. In other examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can refer to motion information signaling from another video block to inform the motion information of the current video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0137] In one example, the motion estimation unit 4404 can indicate, in a syntax structure associated with the current video block, a value that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.
[0138] In another example, the motion estimation unit 4404 can identify, in a syntax structure associated with the current video block, another video block and a motion vector difference (MVD). The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 4500 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0139] As discussed above, the video encoder 4400 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0140] The intra prediction unit 4406 can perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0141] The residual generation unit 4407 can generate residual data for the current video block by subtracting the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0142] In other examples, such as in skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit 4407 can not perform the subtraction operation.
[0143] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0144] After the transform processing unit 4408 generates the transform coefficient video blocks associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0145] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated from the prediction unit 4402 to produce a reconstructed video block associated with the current block for storage in the buffer 4413.
[0146] After the reconstruction unit 4412 reconstructs the video block, loop filtering operations can be performed to reduce video block artifacts in the video block.
[0147] The entropy encoding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy encoding unit 4414 receives data, the entropy encoding unit 4414 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0148] Figure 7 is a block diagram illustrating an example of a video decoder 4500 that can be Figure 5 the system 4300 shown. The video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, the video decoder 4500 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0149] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 4400.
[0150] The entropy decoding unit 4501 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy encoded video data and, from the entropy decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precisions, reference picture list indices, and other motion information. The motion compensation unit 4502 can determine such information, for example, by performing AMVP and Merge modes.
[0151] Motion compensation unit 4502 can generate a motion compensated block, and can perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision can be included in the syntax elements.
[0152] Motion compensation unit 4502 can calculate the interpolation of sub-integer pixels of the reference block using the interpolation filter as used by video encoder 4400 during encoding of the video block. Motion compensation unit 4502 can determine the interpolation filter used by video encoder 4400 from the received syntax information and use the interpolation filter to generate the prediction block.
[0153] Motion compensation unit 4502 can use some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence was partitioned, modes indicating how each partition was encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information for decoding the coded video sequence.
[0154] Intra prediction unit 4503 can use, for example, intra prediction modes received in the bitstream to form a prediction block from spatially adjacent blocks. Dequantization unit 4504 dequantizes quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501, i.e., inverse quantizes the quantized video block coefficients. Inverse transform unit 4505 applies an inverse transform.
[0155] Reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by motion compensation unit 4502 or intra prediction unit 4503 to form a decoded block. If desired, a deblocking filter can also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video blocks are then stored in buffer 4507 to provide reference blocks for subsequent motion compensation / intra prediction, and also to produce decoded video for presentation on a display device.
[0156] Figure 8 is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing the techniques of VVC. Encoder 4600 includes three in-loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses a pre-defined filter, SAO 4604 and ALF 4606 reduce the mean square error between original and reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, with coded side information signaling the offset and filter coefficients, exploiting original samples of the current picture. ALF 4606 is located at the last processing stage of each picture and can be viewed as a tool that attempts to capture and fix artifacts created by previous stages.
[0157] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive the input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction with reference to reference pictures held in the reference picture buffer 4612. Residual blocks from either inter or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding component 4618. The entropy coding component 4618 entropy codes the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). Quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting pictures to the DF 4602, the SAO 4604, and the ALF 4606 for filtering before the pictures are stored in the reference picture buffer 4612.
[0158] Next, a list of some example preferred solutions is provided.
[0159] The following solutions show examples of the techniques discussed herein.
[0160] 1. A method of media data processing (e.g., the method 4200 depicted in FIG. 4B), comprising performing a conversion between visual media information and a digital representation of the visual media information according to a rule, wherein the rule specifies whether or how decoding capability information (DCI) network abstraction layer (NAL) units are included in tracks of coded elementary streams in the digital representation. Figure 4
[0161] 2. The method of solution 1, wherein the rule specifies that a DCI NAL unit is included in each track of coded elementary streams.
[0162] 3. The method of any of solutions 1-2, wherein the rule specifies that, in a case where multiple DCI NAL units are included in a track of coded elementary streams, the multiple DCI NAL units have a same content.
[0163] 4. The method of any of solutions 1-2, wherein the rule specifies that only one DCI NAL unit is included in a track of coded elementary streams.
[0164] 5. The method of solution 1, wherein the rule specifies that a DCI NAL unit, when present in a track of coded elementary streams, is constrained to be in a header sample entry of the track.
[0165] 6. The method according to any of solutions 1-5, wherein the rule specifies that a DCINAL unit conforms to a constraint that a value of a field in the DCINAL unit is constrained to be equal to a predetermined value.
[0166] 7. The method according to solution 6, wherein the field indicates a tier, a layer, a number of layer structures minus 1, and wherein the predetermined value is equal to 0.
[0167] 8. The method according to solution 6, wherein the field indicates whether tier-layer-level multi-layer indication is enabled, and wherein the predetermined value is equal to 1.
[0168] 9. A method of media data processing, comprising performing a conversion between visual media information and a digital representation of the visual media information according to a rule that specifies whether or how a video parameter set (VPS) unit is included in a track of a coded elementary stream in the digital representation.
[0169] 10. The method according to solution 9, wherein the rule specifies that only one VPS unit is included in a track of a coded elementary stream.
[0170] 11. The method according to any of solutions 9-10, wherein the rule specifies that the digital representation satisfies a constraint in case a track of a coded elementary stream includes a VPS unit but does not include a decoding capability information (DCI) network abstraction layer (NAL) unit.
[0171] 12. The method according to any of solutions 9-11, wherein the rule specifies that a VPS conforms to a constraint that a value of a field in the VPS is constrained to be equal to a predetermined value.
[0172] 13. The method according to any of solutions 9-12, wherein the rule specifies that the digital representation satisfies a constraint in case a track of a coded elementary stream does not include a VPS unit and a decoding capability information (DCI) network abstraction layer (NAL) unit.
[0173] 14. A method of media data processing, comprising performing a conversion between visual media information and a digital representation of the visual media information according to a rule that specifies whether or how a value of a field included in a hypothetical reference decoder structure referenced by a video parameter set of a sequence parameter set is allowed to change from one coded video sequence to a second coded video sequence in a coded video elementary stream in the digital representation.
[0174] 15. The method according to solution 14, wherein the value indicates a time scale.
[0175] 16. The method of any of solutions 14-15, wherein the rule specifies that a value of a field is the same in each hypothetical reference decoder structure in a numerical representation.
[0176] 17. A method of media data processing, comprising: obtaining a digital representation of visual media information, wherein the digital representation is generated according to the method of any of solutions 1-16; and streaming the digital representation.
[0177] 18. A method of media data processing, comprising: receiving a digital representation of visual media information, wherein the digital representation is generated according to the method of any of solutions 1-16; and generating the visual media information from the digital representation.
[0178] 19. The method of any of solutions 1-18, wherein the conversion comprises generating a bitstream representation of the visual media data, and storing the bitstream representation to a file according to a format rule.
[0179] 20. The method of any of solutions 1-18, wherein the conversion comprises parsing a file according to a format rule to recover the visual media data.
[0180] 21. A video decoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 20.
[0181] 22. A video encoding apparatus comprising a processor configured to implement a method recited in one or more of solutions 1 to 20.
[0182] 23. A computer program product having computer code stored thereon, the code, which when executed by a processor, causes the processor to implement a method recited in any of solutions 1 to 20.
[0183] 24. A computer readable medium having a bitstream representation conforming to a file format generated according to any of solutions 1 to 20 thereon.
[0184] 25. A method, apparatus or system described in this document. In the solutions described herein, an encoder can conform to a format rule by producing a coded representation according to the format rule. In the solutions described herein, a decoder can conform to a format rule by parsing syntax elements in a coded representation using the format rule, with knowledge of the presence and absence of syntax elements, to produce decoded video according to the format rule.
[0185] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during a conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block can for example correspond to collocated or scattered bits within a bitstream, as defined by the syntax. For example, a macroblock can be encoded according to a transform and coded error residual values and also using bits in headers and other fields in the bitstream. Furthermore, during the conversion, a decoder can parse the bitstream based on the determination, knowing that some fields can or can not be present, as described by the above solutions. Similarly, an encoder can determine to include or not include certain syntax fields and generate the coded representation accordingly by including the syntax fields or excluding the syntax fields from the coded representation.
[0186] The disclosed and other solutions, examples, embodiments, modules and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer readable medium for execution by, or to control the operation of, data processing apparatus. The computer readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of one or more of them, or a combination of one or more of them and other computer program products. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. The propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated for the purpose of encoding information for transmission to suitable receiver apparatus.
[0187] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be run on one computer or on multiple computers that are located at one site or distributed across multiple sites and are interconnected by a communication network.
[0188] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, and / or devices that are
[0189] Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0190] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular art. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[0191] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0192] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
[0193] When there is no intermediate component between the first and second components other than a line, trace, or other medium, the first component is directly coupled to the second component. When there is an intermediate component between the first and second components other than a line, trace, or other medium, the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. Unless otherwise stated, the use of the term "about" means including a range of ±10% of the following figures.
[0194] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. These examples are intended to be illustrative rather than restrictive and are not limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0195] Also, techniques, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate can be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as separate from the other items or illustrative of other embodiments can be, in some embodiments, merged with the other items or each made into a single item. Other examples of changes, substitutions, and alterations are ascertainable by one skilled in the art and could be made without departing from the spirit and scope disclosed herein.
Claims
1. A method for processing video data, comprising: obtaining, based on a rule, information in a sequence parameter set (SPS) in a versatile video coding (VVC) common media application format (CMAF) track, wherein the rule specifies that a value of a video usability information progressive source flag (vui_progressive_source_flag) field in the SPS is constrained to be equal to 1; and performing a conversion between visual media data and a media data file based on the SPS; wherein the rule specifies that a value of a video usability information interlaced source flag (vui_interlaced_source_flag) field in the SPS is constrained to be equal to 1; wherein the method further comprises: determining, based on the rule, information in a sequence parameter set (SPS) in a versatile video coding (VVC) basic stream carried in a VVC common media application format (CMAF) track and information in a decoding capability information (DCI) network abstraction layer (NAL) unit; wherein the rule specifies that a value of a syntax element num_units_in_tick of a number of units in a specified tick in a general hypothetical reference decoder (HRD) parameters (general_timing_hrd_parameters) structure and a value of a syntax element time_scale of a specified time scale, when present in the SPS, must not change between video sequences in the VVC basic stream; wherein the rule further specifies that a value of a syntax element dci_num_ptls_minusl in the DCI NAL unit must be equal to 0; and performing the conversion between the visual media data and the media data file based on the information in the SPS and the information in the DCI NAL unit.
2. The method of claim 1, wherein, vui_progressive_source_flag is included in a video usability information payload (vui_payload) structure.
3. The method of claim 2, wherein, vui_progressive_source_flag is equal to 1 to indicate that video in a VVC CMAF track is coded according to progressive scanning.
4. The method of claim 1, wherein, vui_interlaced_source_flag is included in a video usability information payload (vui_payload) structure.
5. The method of claim 4, wherein, vui_interlaced_source_flag is equal to 1 to indicate that video in a VVC CMAF track is coded according to interlaced.
6. The method of claim 1, wherein, the SPS is contained in a VVC basic stream, and wherein the VVC basic stream is contained in a CMAF track.
7. The method of any one of claims 1-6, wherein, the conversion comprises encoding visual media data into a media data file.
8. The method of any one of claims 1-6, wherein, the conversion comprises decoding visual media data from a media data file.
9. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: based on the rule, obtaining information in a sequence parameter set (SPS) in a versatile video coding (VVC) common media application format (CMAF) track, wherein the rule specifies that a value of a video usability information progressive source flag (vui_progressive_source_flag) field in the SPS is constrained to be equal to 1; and performing a conversion between visual media data and a media data file based on the SPS; wherein the rule specifies that a value of a video usability information interlaced source flag (vui_interlaced_source_flag) field in the SPS is constrained to be equal to 1; based on the rule, determining information in a sequence parameter set (SPS) in a versatile video coding (VVC) basic stream carried in a VVC common media application format (CMAF) track and information in a decoding capability information (DCI) network abstraction layer (NAL) unit; wherein the rule specifies that a value of a syntax element num_units_in_tick of a unit number in a specified tick and a value of a syntax element time_scale of a specified time scale in a general hypothetical reference decoder (HRD) parameter (general_timing_hrd_parameters) structure, when present in the SPS, must not change between video sequences in the VVC basic stream; wherein the rule further specifies that a value of a syntax element dci_num_ptls_minus1 in the DCI NAL unit must be equal to 0; and performing a conversion between visual media data and a media data file based on the information in the SPS and the information in the DCI NAL unit.
10. The apparatus of claim 9, wherein, vui_progressive_source_flag is included in a video usability information payload (vui_payload) structure.
11. The apparatus of claim 10, wherein, vui_progressive_source_flag is equal to 1 to indicate that video in a VVC CMAF track is coded according to progressive scanning.
12. The apparatus of claim 9, wherein, vui_interlaced_source_flag is included in a video usability information payload (vui_payload) structure.
13. The apparatus of claim 12, wherein, vui_interlaced_source_flag is equal to 1 to indicate that video in a VVC CMAF track is coded according to interlacing.
14. The apparatus of claim 9, wherein, the SPS is contained in a VVC basic stream, and wherein the VVC basic stream is contained in a CMAF track.
15. A non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product comprising computer executable instructions stored on the non-transitory computer readable medium such that when executed by a processor cause the video coding device to: based on the SPS, performing a conversion between the visual media data and a media data file; based on the SPS, performing a conversion between the visual media data and a media data file; wherein based on the SPS, performing a conversion between the visual media data and a media data file; based on the SPS, performing a conversion between the visual media data and a media data file; based on the SPS, performing a conversion between the visual media data and a media data file.
16. The non-transitory computer-readable medium of claim 15, wherein, vui_progressive_source_flag is included in a video usability information payload (vui_payload) structure.
17. The non-transitory computer-readable medium of claim 16, wherein vui_progressive_source_flag is equal to 1 to indicate that video in a Versatile Video Coding (VVC) Common Media Application Format (CMAF) track is coded according to progressive scanning.
18. A non-transitory computer-readable storage medium storing instructions that cause a processor to: based on a rule, obtain information in a sequence parameter set (SPS) in a Versatile Video Coding (VVC) Common Media Application Format (CMAF) track, wherein the rule specifies that a value of a video usability information progressive source flag (vui_progressive_source_flag) field in the SPS is constrained to be equal to 1; and based on the SPS, perform a conversion between visual media data and a media data file; based on the SPS, performing a conversion between the visual media data and a media data file; based on the SPS, performing a conversion between the visual media data and a media data file; wherein determining, based on the rule, information in a sequence parameter set (SPS) in a Versatile Video Coding (VVC) elementary stream carried in a Versatile Media Application Format (CMAF) track and information in a Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit; wherein the rule specifies that a value of a syntax element num_units_in_tick of a number of units in a specified tick in a general hypothetical reference decoder (HRD) parameters (general_timing_hrd_parameters) structure and a value of a syntax element time_scale of a specified time scale, when present in the SPS, must not change between video sequences in the VVC elementary stream; wherein the rule further specifies that a value of a syntax element dci_num_ptls_minus1 in the DCI NAL unit must be equal to 0; and performing a conversion between the visual media data and the media data file based on the information in the SPS and the information in the DCI NAL unit.
19. A non-transitory computer-readable recording medium storing a visual media file generated by a method performed by a video processing apparatus, wherein, The method comprises: obtaining, based on a rule, information in a sequence parameter set (SPS) in a Versatile Video Coding (VVC) elementary stream carried in a Versatile Media Application Format (CMAF) track, wherein the rule specifies that a value of a video usability information progressive source flag (vui_progressive_source_flag) field in the SPS is constrained to be equal to 1; and generating, for visual media data, a media data file based on the SPS; wherein the rule specifies that a value of a video usability information interlaced source flag (vui_interlaced_source_flag) field in the SPS is constrained to be equal to 1; wherein the method further comprises: determining, based on the rule, information in a sequence parameter set (SPS) in a Versatile Video Coding (VVC) elementary stream carried in a Versatile Media Application Format (CMAF) track and information in a Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit; wherein the rule specifies that a value of a syntax element num_units_in_tick of a number of units in a specified tick in a general hypothetical reference decoder (HRD) parameters (general_timing_hrd_parameters) structure and a value of a syntax element time_scale of a specified time scale, when present in the SPS, must not change between video sequences in the VVC elementary stream; wherein the rule further specifies that a value of a syntax element dci_num_ptls_minus1 in the DCI NAL unit must be equal to 0; and generating, for the visual media data, the media data file based on the information in the SPS and the information in the DCI NAL unit.
20. A method of storing a visual media file, comprising: based on a rule that specifies a value of a video usability information vui_progressive_source_flag field in a sequence parameter set (SPS) in a versatile video coding (VVC) common media application format (CMAF) track is constrained to be equal to 1; generating, for the visual media data, a media data file based on the SPS; and storing the video media file in a non-transitory computer-readable recording medium; wherein the rule specifies a value of a video usability information vui_interlaced_source_flag field in the SPS is constrained to be equal to 1; wherein the method further comprises: based on the rule, determining information in a sequence parameter set (SPS) in a versatile video coding (VVC) basic stream carried in a VVC common media application format (CMAF) track and information in a decoding capability information (DCI) network abstraction layer (NAL) unit; wherein the rule specifies a value of a syntax element num_units_in_tick of a number of units in a specified tick in a general hypothetical reference decoder (HRD) parameters (general_timing_hrd_parameters) structure and a value of a syntax element time_scale of a specified time scale, when present in the SPS, must not change between video sequences in the VVC basic stream; wherein the rule further specifies a value of a syntax element dci_num_ptls_minus1 in the DCI NAL unit must be equal to 0; and generating, for the visual media data, the media data file based on the information in the SPS and the information in the DCI NAL unit.