Parameter Sets in the Generic Media Application Format

By introducing rule management num_units_in_tick and time_scale in the VVC CMAF track, the problem of inconsistent transmission of information of video parameter sets and sequence parameter sets in the VVC bitstream is solved, which improves the flexibility and efficiency of the video encoding and decoding process and reduces the transcoding process.

CN115225909BActive Publication Date: 2025-07-11FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210406241.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-18
Filing Date
2022-04-18
Publication Date
2025-07-11
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

In VVC bitstreams, it is difficult for the prior art to effectively manage and transmit key information in video parameter sets (VPS) and sequence parameter sets (SPS), resulting in inconsistency and unnecessary transcoding processes in video encoding and decoding.

Method used

By introducing rules in the VVC CMAF track, ensure that num_units_in_tick and time_scale are consistent across video sequences in the VVC elementary stream and use DCI NAL units, VPS and SPS to communicate grade, layer and level information when necessary, so that the decoder can correctly decode and transmit.

Benefits of technology

It realizes a more flexible content preparation process, reduces the need for transcoding, ensures the stability and efficiency of the video encoding and decoding process, and adapts to the decoding needs under different network conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115225909B_ABST
    Figure CN115225909B_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. Information in a sequence parameter set (SPS) in a versatile video coding (VVC) elementary stream is determined. Rules specify that the number of units in a tick (num_units_in_tick) and the time scale should not change between video sequences in the VVC elementary stream when present in the SPS. A conversion is performed between visual media data and a media data file based on the SPS.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] Under the applicable patent laws and / or rules of the Paris Convention, this application is intended to claim the priority and benefit of U.S. Provisional Patent Application No. 63 / 176,315, filed on April 18, 2021, in a timely manner. For all purposes provided by law, the entire disclosure of the above - mentioned application is incorporated by reference as part of the disclosure of this application. Technical field

[0003] This patent document relates to the generation, storage, and consumption of digital audio - visual media information in a file format. Background art

[0004] Digital video occupies the largest bandwidth used on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video use may continue to grow. Summary of the invention

[0005] A first aspect relates to a method for processing video data, including: determining information in a sequence parameter set (SPS) in a VVC elementary stream carried in a multi - functional video coding (VVC) common media application format (CMAF) track, where a rule specifies that the number of units in a tick (num_units_in_tick) and the time scale (time_scale) should not change between video sequences in the VVC elementary stream when present in the SPS; and performing a conversion between visual media data and a media data file based on the SPS.

[0006] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that the VVC elementary stream includes a video parameter set (VPS), and where the rule further specifies that num_units_in_tick and time_scale should not change between video sequences in the VVC elementary stream when present in the VPS.

[0007] Optionally, in any of the foregoing aspects, another implementation of this aspect provides that num_units_in_tick and time_scale are included in a general hypothetical reference decoder (HRD) parameter (general_timing_hrd_parameters) structure.

[0008] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that the values of num_units_in_tick and time_scale should be the same for all general_timing_hrd_parameters structures in the VVC CMAF track.

[0009] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that there should be one and only one Video Parameter Set (VPS) unit in the VVC CMAF track.

[0010] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that when there is no Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit in the VVC CMAF track and when there is no Video Parameter Set (VPS) in the VVC CMAF track, the values of general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] for each i-th interoperability indicator should not change from one video sequence to another in the VVC elementary stream.

[0011] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that when there is no Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit in the VVC CMAF track and when there is one or more Video Parameter Sets (VPS) in the VVC CMAF track, one or more constraints apply, and one or more of the constraints include: the value of the vps_max_layers_minus1 field should be equal to 0 for each VPS, the value of vps_num_ptls_minus1 for the Profile, Tier, and Level (PTL) of each VPS should be equal to 0, the value of the ptl_fram_only_constraint_flag in the profile_tier_level structure in each VPS should be equal to 1, and the value of the ptl_multilayer_enabled_flag in the profile_tier_level structure in each VPS should be equal to 0.

[0012] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that the values of the general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] of the i-th interoperability indicator should not change from one video sequence to another in the VVC elementary stream.

[0013] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the conversion includes encoding visual media data into a media data file.

[0014] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the conversion includes decoding visual media data from a media data file.

[0015] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determine information in a sequence parameter set (SPS) in a VVC elementary stream carried in a versatile video coding (VVC) common media application format (CMAF) track, wherein the rule specifies that the number of units in a tick (num_units_in_tick) and the time scale (time_scale) should not change between video sequences in the VVC elementary stream when present in the SPS; and perform a conversion between visual media data and a media data file based on the SPS.

[0016] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the VVC elementary stream includes a video parameter set (VPS), and wherein the rule further specifies that num_units_in_tick and time_scale should not change between video sequences in the VVC elementary stream when present in the VPS.

[0017] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that num_units_in_tick and time_scale are included in a general hypothetical reference decoder (HRD) parameter (general_timing_hrd_parameters) structure.

[0018] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that the values of num_units_in_tick and time_scale shall be the same for all general_timing_hrd_parameters structures in the VVC CMAF track.

[0019] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that there shall be one and only one video parameter set (VPS) unit in the VVC CMAF track.

[0020] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that when there is no decoding capability information (DCI) network abstraction layer (NAL) unit in the VVC CMAF track and when there is no video parameter set (VPS) in the VVC CMAF track, the values of general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] for each i-th interoperability indicator shall not change from one video sequence to another in the VVC elementary stream.

[0021] A third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium such that when executed by a processor, cause the video codec device to: determine information in a sequence parameter set (SPS) in a VVC elementary stream carried in a Versatile Video Coding (VVC) Common Media Application Format (CMAF) track, where a rule specifies that the number of units in a tick (num_units_in_tick) and a time scale (time_scale) shall not change between video sequences in the VVC elementary stream when present in the SPS; and perform a conversion between visual media data and a media data file based on the SPS.

[0022] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the VVC elementary stream includes a video parameter set (VPS), where the rule further specifies that num_units_in_tick and time_scale shall not change between video sequences in the VVC elementary stream when present in the VPS.

[0023] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that num_units_in_tick and time_scale are included in a general_timing_hrd_parameters structure of general Hypothetical Reference Decoder (HRD) parameters.

[0024] Optionally, in any of the foregoing aspects, another implementation of the aspect provides that the rule specifies that the values of num_units_in_tick and time_scale shall be the same for all general_timing_hrd_parameters structures in a VVC CMAF track.

[0025] For the purposes of clarity, any of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0026] These and other features will be more clearly understood from the following detailed description in conjunction with the drawings and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To more fully understand the present disclosure, reference is now made to the following brief description taken in conjunction with the drawings and detailed description, in which like reference numerals represent like parts.

[0028] Figure 1 is a schematic diagram showing an example Common Media Application Format (CMAF) track.

[0029] Figure 2 is a block diagram showing an example video processing system.

[0030] Figure 3 is a block diagram of an example video processing apparatus.

[0031] Figure 4 is a flowchart of an example method of video processing.

[0032] Figure 5 is a block diagram showing an example video codec system.

[0033] Figure 6 is a block diagram showing an example encoder.

[0034] Figure 7 is a block diagram showing an example decoder.

[0035] Figure 8 is a schematic diagram of an example encoder. DETAILED DESCRIPTION

[0036] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or to be developed. The present disclosure should not in any way be limited to the illustrative implementations, figures, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the full scope of the appended claims and their equivalents.

[0037] This patent document relates to video streams. Specifically, this document relates to specifying constraints for video encoding and encapsulation into media tracks and segments of a file format. Such file formats may include the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF). Such file formats may also include adaptive streaming representation formats such as Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) and / or Common Media Application Format (CMAF). For media streaming systems, the ideas described herein may be applied individually or in various combinations, such systems being based on the DASH standard and related extensions and / or on the CMAF standard and related extensions.

[0038] This disclosure includes the following abbreviations. Adaptive Color Transformation (ACT), Adaptive Loop Filter (ALF), Adaptive Motion Vector Resolution (AMVR), Adaptive Parameter Set (APS), Access Unit (AU), Access Unit Delimiter (AUD), Advanced Video Coding ((Rec. ITU-T H.264|ISO / IEC 14496-10) (AVC), Bidirectional Prediction (B), Bidirectional Prediction with Coding Unit Level Weights (BCW), Bidirectional Optical Flow (BDOF), Block-based Differential Pulse Coding Modulation (BDPCM), Buffer Period (BP), Context-based Adaptive Binary Arithmetic Coding (CABAC), Coding Block (CB), Constant Bit Rate (CBR), Cross-component Adaptive Loop Filter (CCALF), Coding Picture Buffer (CPB), Clean Random Access (CRA), Cyclic Redundancy Check (CRC), Coding Tree Block (CTB), Coding Tree Unit (CTU), Coding Unit (CU), Coding Video Sequence (CVS), Decoding Capability Information (DCI), Decoding Initialization Information (DII), Decoded Picture Buffer (DPB), Dependent Random Access Point (DRAP), Decoding Unit (DU), Decoding Unit Information (DUI), Exponential-Golomb (EG), k-th Exponential-Golomb (EGk), End of Bitstream (EOB), End of Sequence (EOS), Fill Data (FD), First In First Out (FIFO), Fixed Length (FL), Green, Blue, and Red (GBR), General Constraint Information (GCI), Progressive Decoding Refresh (GDR), Geometric Partitioning Mode (GPM), High Efficiency Video Coding (also known as Rec. ITU-T H.265|ISO / IEC 23008-2)(High Efficiency Video Coding (HEVC)), Hypothetical Reference Decoder (HRD), Hypothetical Stream Scheduler (HSS), Intra (I), Intra Block Copy (IBC), Instantaneous Decoding Refresh (IDR), Inter Layer Reference Picture (ILRP), Intra Random Access Point (IRAP), Low Frequency Non-Separable Transform (LFNST), Least Probable Symbol (LPS), Least Significant Bit (LSB), Long Term Reference Picture (LTRP), Luma Mapping and Chroma Scaling (LMCS), Matrix-Based Intra Prediction (MIP), Most Probable Symbol (MPS), Most Significant Bit (MSB), Multiple Transform Selection (MTS), Motion Vector Prediction (MVP), Network Abstraction Layer (NAL), Output Layer Set (OLS), Operation Point (OP), Operation Point Information (OPI), Prediction (P), Picture Header (PH), Picture Order Count (POC), Picture Parameter Set (PPS), Prediction Refinement using Optical Flow (PROF), Picture Timing (PT), Picture Unit (PU), Quantization Parameter (QP), Random Access Decodable Leading Picture (RADL), Random Access Skipped Leading Picture (RASL), Raw Byte Sequence Payload (RBSP), Red, Green, and Blue (RGB), Reference Picture List (RPL), Sample Adaptive Offset (SAO), Sample Aspect Ratio (SAR), Supplemental Enhancement Information (SEI), Slice Header (SH), Sub-Picture Level Information (SLI), String of Data Bits (SODB), Sequence Parameter Set (SPS), Short Term Reference Picture (STRP), Successive Temporal Sub-Layer Access (STSA), Truncated Rice (TR), Variable Bit Rate (VBR), Video Coding Layer (VCL), Video Parameter Set (VPS), Versatile Supplemental Enhancement Information (also known as Rec. ITU-T H.274|ISO / IEC 23002-7) (VSEI), Video Usability Information (VUI), and Versatile Video Coding (also known as Rec. ITU-T H.266|ISO / IEC 23090-3) (VCC).

[0039] Video coding standards have evolved mainly through the development of standards by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the ISO / International Electrotechnical Commission (IEC). ITU-T developed H.261 and H.263, ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, which utilizes temporal prediction plus transform coding. To explore video coding technologies beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many methods and incorporated them into a reference software called the Joint Exploration Model (JEM). When the Versatile Video Coding (VVC) project officially started, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a coding standard aiming to reduce the bitrate by 50% compared to HEVC. VVC has been completed by JVET.

[0040] The VVC standard (also known as ITU-T H.266|ISO / IEC 23090-3) and the associated Versatile Supplementary Enhancement Information (VSEI) standard (also known as ITU-T H.274|ISO / IEC 23002-7) are designed for a wide range of applications, such as television broadcasting, video conferencing, playback from storage media, adaptive bitrate streaming, video region extraction, synthesis and merging of content from multiple coded video bitstreams, multi-view video, scalable hierarchical coding, and viewport-adaptive 360-degree (360°) immersive media. The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video coding standard developed by MPEG.

[0041] The file format standards will be discussed below. Media streaming applications typically rely on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport methods and often depend on file formats such as ISOBMFF. One such streaming system is HTTP-based Dynamic Adaptive Streaming over HTTP (DASH). Videos can be encoded in video formats such as AVC and / or HEVC. The encoded videos can be encapsulated in ISOBMFF tracks and included in DASH representations and segments. For content selection purposes, important information about the video bitstream, such as profile, tier, and level, can be presented as file format level metadata and / or in the DASH Media Presentation Description (MPD). For example, such information can be used to select appropriate media segments for initialization at the start of a streaming session and for stream adaptation during the streaming session.

[0042] Similarly, when using an image format with ISOBMFF, file format specifications specific to that image format can be adopted, such as the AVC image file format and the HEVC image file format. MPEG is developing the VVC video file format, which is a file format based on ISOBMFF for storing VVC video content. MPEG is also developing a VVC image file format based on ISOBMFF, which is a file format for storing image content encoded and decoded using VVC.

[0043] Now discuss the file format standards. Media streaming applications can be based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport mechanisms. Such media streaming applications can also rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is HTTP-based Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, a file format specification dedicated to that video format can be adopted to encapsulate video content in ISOBMFF tracks as well as DASH representations and segments. Such file format specifications can include the AVC file format and the HEVC file format. For content selection purposes, important information about the video bitstream, such as profile, tier, and level, etc., can be presented in the file format level metadata and / or the DASH Media Presentation Description (MPD). For example, content selection can include selecting appropriate media segments for initialization at the start of a streaming session and stream adaptation during the streaming session. Similarly, to use image formats with ISOBMFF, a file format specification dedicated to that image format can be adopted, such as the AVC image file format and the HEVC image file format. The VVC video file format is a file format based on ISOBMFF for storing VVC video content. The VVC video file format was developed by MPEG. The VVC image file format is a file format based on ISOBMFF for storing image content encoded and decoded using VVC. The VVC image file format was also developed by MPEG.

[0044] Now discuss DASH. In DASH, the video and / or audio data of multimedia content may have multiple representations. Different representations can correspond to different codec characteristics, such as different profiles or levels of a video codec standard, different bitrates, different spatial resolutions, etc. A list of such representations can be defined in the Media Presentation Description (MPD) data structure. A media presentation can correspond to a structured set of data accessible to a DASH streaming client device. The DASH streaming client device can request and download media data information to present a streaming service to the user of the client device. The media presentation can be described in the MPD data structure, which can include updates to the MPD.

[0045] A media presentation may include a sequence of one or more periods. Each period may extend to the start of the next period or, in the case of the last period, to the end of the media presentation. Each period may include one or more representations of the same media content. A representation may be one of multiple alternative encoded versions of audio, video, timed text, or other such data. These representations may differ according to the encoding type, such as the bitrate, resolution, codec of video data, and the bitrate, language, and / or codec of audio data. The term representation may be used to refer to a portion of the encoded audio or video data corresponding to a particular period of multimedia content and encoded in a particular way.

[0046] Representations of a particular period may be assigned to groups indicated by attributes in the MPD, which indicate the adaptation sets to which the representations belong. Representations within the same adaptation set are generally considered to be mutually substitutable. Thus, a client device may switch dynamically and seamlessly between these representations, such as performing bandwidth adaptation. For example, each representation of the video data for a particular period may be assigned to the same adaptation set such that any representation may be selected for decoding to render the media data of the corresponding period of multimedia content, such as video data or audio data. In some examples, the media content within a period may be represented by one representation (if any) from group 0 or a combination of at most one representation from each non-zero group. The timing data for each representation of a period may be expressed relative to the start time of the period.

[0047] A representation may include one or more segments. Each representation may include an initialization segment, or each segment of a representation may be self-initializing. When present, the initialization segment may contain initialization information for accessing the representation. Generally, the initialization segment does not contain media data. Segments may be uniquely referenced by identifiers, such as Uniform Resource Locators (URLs), Uniform Resource Names (URNs), or Uniform Resource Identifiers (URIs). The MPD may provide an identifier for each segment. In some examples, the MPD may also provide byte ranges in the form of range attributes, which may correspond to the data of the segments within a file accessible by a URL, URN, or URI.

[0048] For different types of media data, different representations may be selected for substantially simultaneous retrieval. For example, a client device may select an audio representation, a video representation, and a timed text representation from which to retrieve segments. In some examples, a client device may select a particular adaptation set to perform bandwidth adaptation. For example, a client device may select an adaptation set that includes a video representation, an adaptation set that includes an audio representation, and / or an adaptation set that includes timed text. In an example, a client device may select an adaptation set for some types of media such as video and directly select representations for other types of media such as audio and / or timed text.

[0049] The example DASH streaming process can be shown through the following steps. The client obtains the MPD. Then, the client estimates the downlink bandwidth and selects video and audio representations based on the estimated downlink bandwidth, codec, decoding capabilities, display size, audio language settings, etc. Until the end of the media presentation is reached, the client requests media segments of the selected representations and presents the streaming content to the user. The client continuously estimates the downlink bandwidth. When the bandwidth changes significantly, e.g., by becoming lower or higher, the client selects different video representations to match the newly estimated bandwidth and continues to download segments with the updated downlink bandwidth.

[0050] Now discuss CMAF. CMAF specifies a set of constraints for encoding and encapsulating media into ISOBMFF tracks, ISOBMFF segments, ISOBMFF fragments, DASH representations, and / or CMAF tracks, CMAF segments, etc. Such constraints are for the encapsulation at each interoperability point defined as a media profile. The main goal of CMAF development is to enable the reuse of the same media content encoded with a specific codec (e.g., AVC for video) and encapsulated in a specific format (e.g., ISOBMFF) across two separate media streaming worlds of DASH and Apple HTTP Live Streaming (HLS).

[0051] Now discuss the Decoding Capability Information (DCI) in VVC. The DCI NAL unit contains bitstream-level profile, tier, and level (PTL) information. The DCI NAL unit includes one or more PTL syntax structures that can be used during session negotiation between the sender and receiver of the VVC bitstream. When a DCI NAL unit is present in the VVC bitstream, each output layer set (OLS) in the CVS of the bitstream shall conform to the PTL information carried in at least one PTL structure in the DCI NAL unit. In AVC and HEVC, the PTL information for session negotiation is available in the SPS (for HEVC and AVC) and VPS (for HEVC hierarchical extension). This design of conveying PTL information for session negotiation in HEVC and AVC has drawbacks because the scope of the SPS and VPS is within the CVS, rather than within the entire bitstream. This may cause the sender-receiver session initiation to suffer re-initiation during the bitstream streaming of each new CVS. DCI solves this problem because DCI carries bitstream-level information and thus can guarantee compliance with the indicated decoding capabilities until the end of the bitstream.

[0052] Now discuss the Video Parameter Set (VPS) in VVC. The VVC bitstream can contain a Video Parameter Set (VPS), which contains information about the description layer and the Output Layer Set (OLS) for the operations of the decoding process of the scalable bitstream. The OLS is a set of layers in the bitstream, where one or more layers are designated to be output from the decoder. Other layers identified in the OLS can also be decoded in order to decode the output layer, although such layers are not designated for output. Much of the information contained in the VPS can be used in the system for purposes such as session negotiation and content selection. The VPS was introduced to handle multi-layer bitstreams. For a single-layer VVC bitstream, the presence of the VPS in the CVS is optional. This is because the information contained in the VPS is not necessary for the operations of the decoding process of the bitstream. The absence of the VPS in the CVS is indicated by referring to the VPS identifier (ID) equal to 0 in the SPS, in which case default values are inferred for the VPS parameters.

[0053] Now discuss the Sequence Parameter Set (SPS) in VVC. The SPS conveys sequence-level information shared by all pictures in the entire Coding Layer Video Sequence (CLVS). This includes PTL indicators, picture format, feature and / or tool control flags, coding, prediction, and / or transform block structure and hierarchy, candidate RPLs that the encoder can refer to, etc. The picture format can include color sampling format, maximum picture width, maximum picture height, and bit depth. In most applications, the entire bitstream uses only one or a few SPSs. Therefore, it is not necessary to update the SPS within the bitstream. Updating the SPS can include sending a new SPS with the SPSID of the existing SPS but with different values for some parameters. Pictures from a specific layer that refer to an SPS with a different SPSID or with the same SPS ID but different SPS content belong to different CLVSs. Similar to AVC and HEVC, the SPS can be transmitted in-band, or using a hybrid of in-band and out-of-band signaling. In-band signaling indicates that data such as the SPS is transmitted together with the coded pictures, and out-of-band signaling indicates that data such as the SPS is not transmitted together with the coded pictures.

[0054] Now discuss the Picture Parameter Set (PPS) in VVC. The PPS conveys picture-level information shared by all slices of a picture. Such information can also be shared across multiple pictures. This includes feature and / or tool on / off flags, picture width and height, default RPL size, slice and strip configurations, etc. By design, two consecutive pictures can refer to two different PPSs. This may result in a large number of PPSs being used within the CLVS. In fact, the number of PPSs in the entire bitstream may not be high because the PPS is designed to carry parameters that do not change frequently and may apply to multiple pictures. Therefore, it may not be necessary to update the PPS within the CLVS or even within the entire bitstream. The Adaptive Parameter Set (APS) can be used for parameters that can apply to multiple pictures but are expected to change frequently between different pictures. Like the SPS, the PPS can be transmitted in-band, out-of-band, or using a hybrid of in-band and out-of-band signaling. A fundamental design principle regarding which picture-level parameters should be included in the PPS and which should be in the APS is the frequency at which such parameters may change. Therefore, frequently changing parameters are not included in the PPS to avoid requiring PPS updates, which would not allow out-of-band transmission of the PPS in typical use cases.

[0055] Now discuss the Adaptive Parameter Set (APS) in VVC. The APS conveys picture and / or slice-level information that can be shared by multiple slices of a picture and / or slices of different pictures but can change frequently across pictures. The APS supports information with a large number of variables that are not suitable for inclusion in the PPS. Three types of parameters are included in the APS, including Adaptive Loop Filter (ALF) parameters, Luminance Mapping and Chrominance Scaling (LMCS) parameters, and Scaling List parameters. The APS can be carried in two different NAL unit types, which can be prefixed or suffixed before or after the associated slice. The latter can be helpful in ultra-low latency scenarios, such as allowing the encoder to send the slices of a picture before generating the ALF parameters for that picture, which will be used by subsequent pictures in decoding order.

[0056] Now discuss the Picture Header (PH). For each PU, there is a Picture Header (PH) structure. The PH exists as a separate PH NAL unit or is included in the Slice Header (SH). If a PU contains only one slice, the PH can only be included in the SH. To simplify the design, within the CLVS, the PH can only be entirely in the PH NAL unit or entirely in the SH. When the PH is in the SH, there is no PH NAL unit in the CLVS. The PH is designed for two purposes. First, the PH helps reduce the signaling overhead of the SH for pictures containing multiple slices. The PH achieves this by carrying all parameters that have the same value for all slices of a picture, thus preventing the repetition of the same parameters in each SH. These include IRAP and / or GDR picture indication, inter and / or intra slice enable flags, and information related to POC, RPL, deblocking filter, SAO, ALF, LMCS, scaling list, QP delta, weighted prediction, codec block partitioning, virtual boundary, collocated pictures, etc. Second, the PH helps the decoder identify the first slice of each decoded picture containing multiple slices. Since there is one and only one PH for each PU, when the decoder receives a PH NAL unit, the decoder knows that the next VCL NAL unit is the first slice of the picture.

[0057] Now discuss the Operation Point Information (OPI). The decoding processes of HEVC and VVC have similar input variables to set the decoding operation point. These include the target OLS and the highest sublayer of the bitstream to be decoded through the decoder API. However, in scenarios where the layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder application programming interface (API) to the application, the decoder may not be able to correctly determine the operation point for processing the bitstream. Therefore, the decoder may not be able to infer the attributes of the pictures in the bitstream, such as the appropriate buffer allocation for decoded pictures and whether to output individual pictures. To address this issue, VVC includes a mode to indicate these two variables within the bitstream through the OPI NAL unit. In the AU at the start of the bitstream and in the separate CVS of the bitstream, the OPI NAL unit notifies the decoder of the target OLS and the highest sublayer of the bitstream to be decoded. In the case where the OPI NAL unit exists and the operation point is also provided to the decoder via the decoder API information, the decoder API information takes precedence. For example, the application may have more up-to-date information related to the target OLS and sublayer. In the case where there is no decoder API and any OPI NAL unit in the bitstream, appropriate fallback options are specified in VVC to allow correct decoder operation.

[0058] Now discuss the example CMAF specification. The VVC video CMAF track can be described as follows. The VVC CMAF track shall comply with the requirements of the NAL-structured video CMAF track. In addition, the CMAF track can comply with all other requirements described in this document. If the CMAF track complies with these requirements, the CMAF track is called a VVC video CMAF track and can use the brand "cvvc". The VVC video track constraints are also discussed. In the example, the VVC video CMAF switch set constraints are as follows. Each CMAF track in the CMAF switch set shall comply with the VVC video CMAF track defined in this document. The VVC video CMAF switch set shall comply with the constraints of the NAL-structured video CMAF switch set.

[0059] Now discuss the visual sample entry. The syntax and values of the visual sample entry of the VVC video track shall comply with the VVCSampleEntry("vvc1") or VVCSampleEntry("vvci") sample entry. Now discuss the constraints on the VVC elementary stream. Regarding the VPS, each VVC video media sample in the CMAF track shall refer to an SPS with a sps_video_parameter_set_id equal to 0 (in this case, there is no VPS in the elementary stream), or shall refer to the VPS in the CMAF header sample entry. If present, the following additional constraints apply. For each profile_tier_level() structure in the VPS, the values of the following fields shall not change throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_subc_profile_idc[i].

[0060] The SPS NAL unit that appears within the CMAF VVC track shall comply with the constraints here and have the following additional constraints. The following fields shall have the following predetermined values: First, the vui_parameters_present_flag shall be set to 1; Second, if the profile_tier_level() structure exists in the SPS, the conditions of the following fields shall not change throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i].

[0061] Now let's discuss the image crop parameters. The SPS and PPS crop parameters conf_win_top_offset and conf_win_left_offset should be set to 0. The SPS and PPS crop parameters conf_win_bottom_offset and conf_win_right_offset can be set to values ​​other than 0. If set to a non-zero value, such syntax elements are expected to be used by CMAF players to remove video spatial samples that are not intended to be displayed.

[0062] Video codec parameters are now discussed. VVC signaling of codec parameters (information) is described below. The demo application SHOULD use parameter signaling to signal the video codec profile and level for each VVC track and CMAF switch set. Encryption is also discussed. Encryption of CMAF VVC tracks and CMAF VVC switch sets SHOULD use the "cenc" AES-CTR scheme or the "cbcs" AES-CBC subsample mode encryption scheme. Additionally, if mode encryption is used with the "cbcs" mode of common encryption, mode block length 10 and Encryption: Skip Mode 1:9 SHOULD be applied.

[0063] The following are example technical problems solved by the disclosed technical solutions. For example, in the example VVC CMAF design, the profile, layer, and level may need to be signaled in the VPS and SPS, so they may not change throughout the VVC elementary stream. However, for VVC bitstreams, DCI NAL units can instead be used to convey the decoding capabilities required for the entire bitstream, while allowing the profiles, layers, and levels within the bitstream to differ between CVSs. This will allow greater flexibility, resulting in fewer transcoding and other processes required in content preparation.

[0064] This document discloses mechanisms for solving one or more of the problems listed above. For example, a VVC elementary stream, also known as a VVC bitstream, can be included in a VVC CMAF track. The VVC elementary stream can include one or more CVSs. The profile, tier, and level (PTL) information of the bitstream can vary between CVSs in the same bitstream. To allow this functionality, the PTL information can be signaled in DCI NAL units, VPS, and / or SPS as long as the corresponding constraints are maintained. In an example, the DCI NAL unit needs to be included in the CMAF track. In an example, when multiple DCI NAL units are included in the CMAF track, all DCI NAL units may need to include the same content. In another example, the CMAF track can include only a single DCI NAL unit. In an example, the DCI NAL unit may need to be included in the CMAF header sample entry. In various examples, the DCI NAL unit can contain the DCI number of PTL minus 1 (dci_num_ptls_minus1) field, the PTL frame only constraint flag (ptl_frame_only_constraint_flag) field, and the PTL multi-layer enabled flag (ptl_multilayer_enabled_flag), which need to be equal to 0, 1, and 0 respectively. In an example, the CMAF track is restricted to containing a single VPS. In an example, when there is no DCI NAL unit, various PTL-related information in the VPS needs to be set to a predetermined value and / or needs to be the same between CVSs, as discussed further below. In an example, various PTL-related information in the SPS needs to be set to a predetermined value and / or needs to be the same between CVSs, as discussed further below. In an example, the timing-related hypothetical reference decoder (HRD) parameters may also need to be the same between CVSs.

[0065] Figure 1 FIG. is a schematic diagram showing an example CMAF track 100. The CMAF track 100 is a track of video data that has been encapsulated based on the constraints specified in the CMAF standard. The CMAF track 100 is constrained to support the delivery and decoding of a wide range of client devices according to adaptive streaming. In adaptive streaming, the media profile describes multiple different interchangeable representations, which allows client devices to select the desired representation based on decoder capabilities and / or current network conditions. The CMAF track 100 can support such functionality by including representations that are constrained to be decodable by the client, which is capable of decoding at the corresponding profile, tier, and level (PTL), capable of using the corresponding codec tools, and / or capable of meeting other predetermined constraints.

[0066] The CMAF track 123 may include many types of decodable video streams. In this example, the CMAF track 123 includes a VVC stream 121. The VVC stream 121, also known as a bitstream, is a video data stream that has been encoded and decoded according to the VVC standard. For example, the VVC stream 121 may include a stream of encoded and decoded pictures and associated syntax that describes the encoding and decoding process and / or other data useful to the decoder. The VVC stream 121 may contain one or more CVSs 117. The CVS 117 is a sequence of access units (AUs) in decoding order. An AU is a collection of one or more pictures with corresponding output / display times. Thus, the CVS 117 contains a series of related pictures and corresponding syntax for supporting decoding and / or describing the pictures.

[0067] The CVS 117 may include DCI NAL units 115, VPS 111, SPS 113, and / or decoded video 119. The DCI NAL unit 115 contains information describing the requirements for decoding the video data in the CVS 117 and / or the entire VVC stream 121. The DCI NAL unit 115 is optional and may be omitted in some VVC streams 121 and / or CVSs 117. It should be noted that although depicted as part of the VVC stream 121, in some examples, the DCU NAL unit 115 may also be included in the CMAF header sample entry in the CMAF track 123. The VPS 111 may contain data related to the entire VVC stream 121. For example, the VPS 111 may contain data-related output layer sets (OLSs), layers, and / or sub-layers used in the VVC stream 121. The VPS 111 is optional and may be omitted in some VVC streams 121 and / or CVSs 117. The SPS 113 contains sequence data common to all pictures in the CVS 117 contained in the VVC stream 121. The parameters in the SPS 113 may include picture size, bit depth, encoding and decoding tool parameters, bit rate limits, etc. The SPS 113 should be included in at least one CVS 117. However, multiple CVSs 117 may refer to the same SPS 113. Thus, the VVC stream 121 should contain one or more SPS 113s. The decoded video 119 includes pictures encoded and decoded according to VVC and the corresponding syntax.

[0068] The present disclosure relates to constraints applied to syntax elements included in DCI NAL unit 115, VPS 111, and / or SPS 113. In an example, DCI NAL unit 115 may be required to be present in CMAF track 123. In an example, when more than one DCI NAL unit 115 is present in a single CMAF track 123, all such DCI NAL units 115 may be required to include the same content. In some examples, CMAF 123 track may be restricted to include one and only one DCI NAL unit 115. In such a case, multiple CVSs 117 of the video content may be described by a single DCI NAL unit 115. When present, DCI NAL unit 115 may contain the DCI number minus 1 (dci_num_ptls_minus1) 132 of PTL and / or the PTL syntax (profile_tier_level) structure 130. The dci_num_ptls_minus1 132 may specify, in a minus 1 format, the number of profile_tier_level structures 130 included in DCI NAL unit 115. The minus 1 format indicates that the value included in the syntax element is 1 less than the actual value, so 1 is added to the value included in the syntax element to determine the actual value. In an example, the dci_num_ptls_minus1 132 may be restricted to be equal to 0, which indicates a single profile_tier_level structure 130. This indicates that CMAF track 123 contains video that complies with a single set of PTL information. Depending on the example, the profile_tier_level structure 130 may be included in DCI NAL unit 115, VPS 111, and / or SPS 113, and will be discussed in more detail below.

[0069] In the example, the CMAF track 123 is restricted to contain one and only one VPS 111. In this case, multiple CVSs 117 can be described by a single VPS 111. The VPS 111 can include a vps_max_layers_minus1 field 134, a vps_num_ptls_minus1 field 133 for the number of VPSs of the PTL minus 1, a general_timing_hrd_parameters structure 131 of general Hypothetical Reference Decoder (HRD) parameters, and a profile_tier_level structure 130. The vps_max_layers_minus1 field 134 indicates the number of layers specified by the VPS 111 in a minus 1 format. In the example, the vps_max_layers_minus1 field 134 can be restricted to contain the value 0, which indicates that the VPS 111 describes a single layer. The vps_num_ptls_minus1 field 133 can specify the number of profile_tier_level structures 130 included in the VPS 111 in a minus 1 format. In the example, the vps_num_ptls_minus1 field 133 can be restricted to contain the value 0, which indicates to the VPS 111 a single set of PTL information.

[0070] Depending on the example, the general_timing_hrd_parameters 131 can be included in the VPS 111 and / or the SPS 113. For example, when the VPS 111 is included, the VPS 111 can contain the general_timing_hrd_parameters 131. When the VPS 111 is not included, the SPS can contain the general_timing_hrd_parameters 131. The general_timing_hrd_parameters 131 includes timing-related parameters used by the HRD operating at the encoder. Generally, the HRD can use the HRD parameters to check that the VVC stream 121 conforms to the VVC standard. The general_timing_hrd_parameters 131 indicates to the encoder the timing parameters related to the coded and decoded video 119. For example, the general_timing_hrd_parameters 131 can indicate how fast each picture should be decoded and reconstructed by the decoder for accurate display. In the example, the general_timing_hrd_parameters 131 can contain a time_scale field and a num_units_in_tick field. The time_scale field indicates the number of time units elapsed in one second, where the time unit corresponds to the picture rate frequency of the video signal. The num_units_in_tick indicates the number of time units of a clock operating at the frequency of the time_scale in Hertz (Hz), which corresponds to an increment called a clock tick. In the example, the values of num_units_in_tick and time_scale in the general_timing_hrd_parameters 131 are restricted to remain the same between the CVSs 117 in the same VVC stream 121. In the example, the values of num_units_in_tick and time_scale in the general_timing_hrd_parameters 131 are restricted to remain the same for the entire CMAF track 123.

[0071] The SPS 113 may contain the general_timing_hrd_parameters 131 as discussed above, for example when the VPS 111 is not included. For example, when the DCI NAL unit 115 and / or the VPS 111 are not included, the SPS 113 may further contain the profile_tier_level structure 130. The SPS 113 may also contain the video usability information payload (vui_payload) structure 135. The vui_payload structure 135 contains information describing how the decoder should use the coded video 119. For example, the vui_payload structure 135 may include a video usability information progressive source flag (vui_progressive_source_flag) field 139 and a video usability information interlaced source flag (vui_interlaced_source_flag) field 138. The vui_progressive_source_flag field 139 may be set to indicate whether the video in the CMAF track 123 is coded according to progressive scanning. The vui_interlaced_source_flag field 138 may be set to indicate whether the video in the CMAF track 123 is coded according to interlacing. In an example, the vui_interlaced_source_flag field 138, the vui_progressive_source_flag field 139, or both may need to be set to 1. This indicates that the coded video 119 is coded according to interlacing, progressive scanning, or both.

[0072] As described above, the DCI NAL unit 115, the VPS 111, and / or the SPS 113 may contain the profile_tier_level structure 130. The profile_tier_level structure 130 contains information related to the profile, tier, and level used for coding the coded video. The profile indicates the profile used for coding the coded video. Different profiles have different coding characteristics (e.g., the availability of different coding tools), such as different bit depths, different chroma sampling formats, the availability of cross-component prediction, the availability of intra-frame smoothing disabled. The tier indicates whether the coded video 119 is coded according to the high tier or the main tier, and thus is coded according to the requirements of general applications. The level indicates the constraints on the coded video 119, such as the maximum bit rate, the maximum picture size, the maximum sampling rate, the resolution at the highest frame rate, the maximum number of slices, the maximum number of stripes per picture, etc. Therefore, the PTL information in the profile_tier_level structure 130 describes the capabilities that the decoder must have to decode and display the coded video 119.

[0073] The profile_tier_level structure 130 may include a PTL frame only constraint flag (ptl_frame_only_constraint_flag) field 141, a PTL multi-layer enabled flag (ptl_multilayer_enabled_flag) 143, a general profile identifier code (general_profile_idc) 145, a general tier flag (general_tier_flag) 147, a general level identifier code (general_level_idc) 149, a number of sub-profile tiers (num_sub_profiles) 142, and / or a general sub-profile identifier code for each i-th interoperability indicator (general_sub_profile_idc[i]) 144. The ptl_frame_only_constraint_flag field 141 specifies whether the CVS 117 conveys a picture representing a frame (e.g., a full screen image) or a field (e.g., a partial screen image intended to be combined to fill the screen). In an example, the constraint may require the ptl_frame_only_constraint_flag field 141 to be set to 1, which indicates that the coded video 119 includes pictures coded as frames. The ptl_multilayer_enabled_flag 143 indicates whether the coded video 119 is coded in multiple layers. In an example, the ptl_multilayer_enabled_flag 143 is set to 0, which indicates that the coded video 119 is coded in a single layer.

[0074] The general_profile_idc 145, general_tier_flag 147, and general_level_idc 149 indicate the profile, tier, and level of the coded video 119, respectively. The general_sub_profile_idc[i] 144 indicates the values of the interoperability indicator from 0 to i. The num_sub_profiles 142 indicates the number of syntax elements included in the general_sub_profile_idc[i] 144. In an example, the values included in the general_profile_idc 145, general_tier_flag 147, general_level_idc 149, num_sub_profiles 142, and general_sub_profile_idc[i] 144 need to be kept the same among the CVSs 117 in the same VVC stream. In another example, the values included in the general_profile_idc 145, general_tier_flag 147, general_level_idc 149, num_sub_profiles 142, and general_sub_profile_idc[i] 144 need to be kept the same in the CMAF track 123.

[0075] To address the above problems and other problems, the following summarized methods are disclosed. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow way. In addition, these items can be applied individually or in any combined way.

[0076] Example 1

[0077] In one example, the rule can specify that the DCI NAL unit should be present in the VVC CMAF track.

[0078] Example 2

[0079] In one example, the rule can specify that the DCI NAL unit should be present in the VVC CMAF track.

[0080] Example 3

[0081] In one example, the rule can specify that all DCI NAL units present in the VVC CMAF track should have the same content.

[0082] Example 4

[0083] In one example, the rule can specify that there should be one and only one DCI NAL unit in the VVC CMAF track.

[0084] Example 5

[0085] In one example, the rule may specify that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit shall be present in the CMAF header sample entry.

[0086] Example 6

[0087] In one example, the rule may specify that the value of the dci_num_ptls_minus1 field in the DCI NAL unit in a VVC CMAF track shall be equal to 0.

[0088] Example 7

[0089] In one example, the rule may specify that the value of the ptl_frame_only_constraint_flag field in the profile_tier_level() structure in the DCI NAL unit in a VVC CMAF track shall be equal to 1.

[0090] Example 8

[0091] In one example, the rule may specify that the value of the ptl_multilayer_enabled_flag field in the profile_tier_level() structure in the DCI NAL unit in a VVC CMAF track shall be equal to 0.

[0092] Example 9

[0093] In one example, the rule may specify that there shall be one and only one VPS unit in a VVC CMAF track.

[0094] Example 10

[0095] In one example, the rule may specify that when no DCI NAL unit is present in a VVC CMAF track and one or more VPSs are present in the VVC CMAF track, one or more of the following constraints apply. The constraints include that for each VPS, the value of the vps_max_layers_minus1 field shall be equal to 0, and for each VPS, the value of the vps_num_ptls_minus1 field shall be equal to 0.

[0096] In the example, the following constraints apply to the profile_tier_level() structure in each VPS. Such constraints include that the value of the ptl_frame_only_constraint_flag field shall be equal to 1; and the value of the ptl_multilayer_enabled_flag field shall be equal to 0.

[0097] In the example, the value of each of the following fields in the profile_tier_level() structure of the referenced VPS shall not change from one coded video sequence to another coded video sequence throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i] for each i value. In the example, the rule may require that the value of each of these fields be the same for all VPSs present in the VVC CMAF track.

[0098] Example 11

[0099] In one example, the rule may specify that the value of the vui_progressive_source_flag field in the vui_payload() structure in the SPS in the VVC CMAF track shall be equal to 1.

[0100] Example 12

[0101] In one example, the rule may specify that the value of the vui_interlaced_source_flag field in the vui_payload() structure in the SPS in the VVC CMAF track shall be equal to 1.

[0102] Example 13

[0103] In one example, the rules may specify that when no DCI NAL unit is present and no VPS is present in the VVCCMAF track, the value of each of the following fields in the profile_tier_level() structure of the referenced SPS should not change from one coded video sequence to another coded video sequence throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i] for each i value. In the example, the rules may require that the value of each of these fields be the same for all SPSs present in the VVC CMAF track.

[0104] Example 14

[0105] In one example, the rules may specify that the value of each of the following fields in the general_timing_hrd_parameters() structure in the referenced VPS or SPS (when present) should not change from one coded video sequence to another coded video sequence throughout the VVC elementary stream: num_units_in_tick; and time_scale. In the example, the rules may require that the value of each of these fields be the same for all general_timing_hrd_parameters() structures in the VPS or SPS present in the VVC CMAF track.

[0106] An embodiment of the previous example is now described. This embodiment may be applied to CMAF. Most of the relevant parts that have been added or modified are shown in bold underlined font with respect to the VVC CMAF specification, and some of the deleted parts are shown in bold italic font. There may be some other changes that are essentially editorial and thus not highlighted.

[0107] X.1 VVC Video CMAF Track. The VVC CMAF track shall comply with the requirements of the NAL structured video CMAF track. In addition, it shall comply with all remaining requirements in this appendix. If the CMAF track complies with these requirements, it is called a VVC video CMAF track and may use the brand “cvvc”.

[0108] X.2 VVC Video Track Constraints. X.2.1 VVC Video CMAF Switching Set Constraints. Each CMAF track in the CMAF switching set shall conform to the VVC video CMAF track as defined in Article X.1. The VVC video CMAF switching set shall conform to the constraints of the NAL structured video CMAF switching set.

[0109] X.2.2 Visible Sample Entry. The syntax and values of the visible sample entry of the VVC video track shall conform to the VVCSampleEntry(“vvc1”) or VVCSampleEntry as defined in ISO / IEC 14496-15 Sample Entry.

[0110]

[0111] X.3.2 Video Parameter Set (VPS). Each VVC video media sample in the CMAF track shall refer to the SPS with sps_video_parameter_set_id equal to 0, in which case there is no VPS in the elementary stream, or shall refer to the VPS in the CMAF header sample entry. The following additional constraints apply: The profile_tier_level() structure Throughout the VVC elementary stream shall not change: general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, general_sub_profile_idc[i].

[0112] X.3.3 Sequence Parameter Set (SPS). The sequence parameter set NAL units that appear within the CMAF VVC track shall conform to the following additional constraints: The following fields shall have the following predetermined values: vui_parameters_present_flag shall be set to 1. Throughout the VVC elementary stream To another coded video sequence should not change: general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] for each i value.

[0113]

[0114] X.3.5 Picture cropping parameters. SPS and PPS cropping parameters conf_win_top_offset and conf_win_left_offset should be set to 0. SPS and PPS cropping parameters conf_win_bottom_offset and conf_win_right_offset can be set to a value other than 0. If set to a non-zero value, it is expected to be used by the CMAF player to remove video spatial samples that are not intended to be displayed.

[0115] Figure 2 is a block diagram showing an example video processing system 4000 in which various techniques disclosed herein can be implemented. Various embodiments may include some or all components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8- or 10-bit multi-component pixel values, or may be in a compressed or coded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), etc., and wireless interfaces such as Wi-Fi or cellular interfaces.

[0116] System 4000 may include a codec component 4004 that may implement various encoding or decoding methods described in this document. The codec component 4004 may reduce the average bit rate of the video from the input 4002 to the output of the codec component 4004 to produce a coded representation of the video. Coding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the codec component 4004 may be stored or transmitted via a communication connection represented by component 4006. The stored or communicated bitstream (or coded) representation of the video received at the input 4002 may be used by component 4008 to generate pixel values or a displayable video for transmission to the display interface 4010. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Additionally, while certain video processing operations are referred to as "coding" operations or tools, it will be understood that coding tools or operations are used at the encoder and the corresponding decoding tools or operations that reverse the coded result will be performed by the decoder.

[0117] Examples of a peripheral bus interface or a display interface may include Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI), or Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0118] Figure 3 is a block diagram of an example video processing apparatus 4100. The apparatus 4100 may be used to implement one or more methods described herein. The apparatus 4100 may be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The (multiple) processors 4102 may be configured to implement one or more methods described in this document. The memory (multiple memories) 4104 may be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 may be used to implement some of the techniques described in this document in a hardware circuit system. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102 (e.g., a graphics coprocessor).

[0119] Figure 4It is a flowchart of an example method 4200 for video processing. Method 4200 includes determining information in the SPS in the VVC elementary stream at step 4202. In the example, the rule specifies that num_units_in_tick and time_scale should not change between video sequences in the VVC elementary stream when present in the SPS. In the example, the VVC elementary stream includes a VPS, and the rule further specifies that num_units_in_tick and time_scale should not change between video sequences in the VVC elementary stream when present in the VPS. In the example, num_units_in_tick and time_scale are included in the general_timing_hrd_parameters structure. In the example, the rule specifies that the value of num_units_in_tick and the value of time_scale should be the same for all general_timing_hrd_parameters structures in the VVC CMAF track. In the example, the rule specifies that there should be one and only one VPS unit in the VVC CMAF track.

[0120] In the example, the rule specifies that when there is no DCI NAL unit in the VVC CMAF track and when there is no VPS in the CMAF track, the values of general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] should not change from one video sequence to another in the VVC elementary stream. In the example, the rule specifies that when there is no DCI NAL unit in the VVC CMAF track and when there is one or more VPS in the CMAF track, one or more constraints apply. The one or more constraints include: the value of the vps_max_layers_minus1 field should be equal to 0 for each VPS, the value of vps_num_ptls_minus1 should be equal to 0 for each VPS, the value of the ptl_fram_only_constraint_flag in the profile_tier_level structure in each VPS should be equal to 1, and the value of the ptl_multilayer_enabled_flag in the profile_tier_level structure in each VPS should be equal to 0. In the example, the rule specifies that the values of general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] should not change from one video sequence to another in the VVC elementary stream.

[0121] In step 4204, perform a conversion between the visual media data and the media data file based on the SPS. When method 4200 is executed on the encoder, the conversion includes generating the media data file according to the visual media data. The conversion includes determining the SPS and encoding it into the bitstream included in the VVC elementary stream. When method 4200 is executed on the decoder, the conversion includes parsing and decoding the VVC elementary stream according to the SPS to obtain the visual media data.

[0122] It should be noted that the method 4200 can be implemented in a device for processing video data, which includes a processor and a non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to execute the method 4200. Additionally, the method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use in a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium such that, when executed by a processor, cause the video codec device to execute the method 4200.

[0123] Figure 5 FIG. is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of the present disclosure. The video codec system 4300 can include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, where the source device 4310 can be referred to as a video encoding device. The destination device 4320 can decode the encoded video data generated by the source device 4310, and the destination device 4320 can be referred to as a video decoding device.

[0124] The source device 4310 can include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 can include sources such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data can include one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream can include a sequence of bits forming a codec representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a codec representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 4316 can include a modulator / demodulator (modem) and / or a transmitter. The encoded video data can be directly sent to the destination device 4320 via the I / O interface 4316 over a network 4330. The encoded video data can also be stored on a storage medium / server 4340 for access by the destination device 4320.

[0125] The target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from a source device 4310 or a storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the target device 4320 or may be external to the target device 4320 that may be configured to interface with an external display device.

[0126] The video encoder 4314 and the video decoder 4324 may operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or additional standards.

[0127] Figure 6 is a block diagram showing an example of a video encoder 4400, which may be Figure 5 the video encoder 4314 in the system 4300 shown. The video encoder 4400 may be configured to perform any or all of the techniques of the present disclosure. The video encoder 4400 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0128] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.

[0129] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In an example, the prediction unit 4402 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in the IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0130] In addition, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405, may be highly integrated, but are shown separately in the example of the video encoder 4400 for purposes of explanation.

[0131] The splitting unit 4401 can split an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.

[0132] The mode selection unit 4403 can select one of the encoding and decoding modes (e.g., intra or inter) based on the error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a combination of intra and inter prediction modes (CIIP), where the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 can also select the resolution of the motion vector of the block (e.g., sub-pixel or integer pixel precision).

[0133] To perform inter prediction on the current video block, the motion estimation unit 4404 can generate motion information of the current video block by comparing one or more reference frames from the buffer 4413 with the current video block. The motion compensation unit 4405 can determine the predicted video block of the current video block based on the motion information and the decoded samples of the pictures from the buffer 4413 other than the picture associated with the current video block.

[0134] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice or a B-slice.

[0135] In some examples, the motion estimation unit 4404 can perform uni-directional prediction on the current video block, and the motion estimation unit 4404 can search for the reference pictures in list 0 or list 1 of the reference video blocks for the current video block. The motion estimation unit 4404 can then generate a reference index and a motion vector indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 4404 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 4405 can generate the predicted video block of the current block based on the reference video block indicated by the motion information of the current video block.

[0136] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block. The motion estimation unit 4404 may search for a reference video block of the current video block in the reference pictures in list 0, and may also search for another reference video block of the current video block in list 1. The motion estimation unit 4404 may then generate a reference index and a motion vector. The reference index indicates the reference pictures in list 0 and list 1 that contain the reference video block, and the motion vector indicates the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and the motion vector of the current video block as the motion information of the current video block. The motion compensation unit 4405 may generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.

[0137] In some examples, the motion estimation unit 4404 may output a complete set of motion information for the decoding process of the decoder. In some examples, the motion estimation unit 4404 may not output a complete set of motion information of the current video. Instead, the motion estimation unit 4404 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is similar enough to the motion information of an adjacent video block.

[0138] In one example, the motion estimation unit 4404 may indicate a value in the syntax structure associated with the current video block, and this value indicates to the video decoder 4500 that the current video block has the same motion information as another video block.

[0139] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0140] As discussed above, the video encoder 4400 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0141] The intra prediction unit 4406 may perform intra prediction on the current video block. When the intra prediction unit 4406 performs intra prediction on the current video block, the intra prediction unit 4406 may generate prediction data of the current video block based on the decoded samples of other video blocks in the same picture. The prediction data of the current video block may include a predicted video block and various syntax elements.

[0142] The residual generation unit 4407 may generate residual data for a current video block by subtracting a (plurality of) predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0143] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform the subtraction operation.

[0144] The transform processing unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0145] After the transform processing unit 4408 generates the transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0146] The inverse quantization unit 4410 and the inverse transform unit 4411 may respectively apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.

[0147] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0148] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives the data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy coded data and output a bitstream including the entropy coded data.

[0149] Figure 7 is a block diagram showing an example of a video decoder 4500, and the video decoder 4500 may be Figure 5 the video decoder 4324 in the system 4300 shown. The video decoder 4500 may be configured to perform any or all of the techniques of the present disclosure. In the example shown, the video decoder 4500 includes a plurality of functional components. The techniques described in the present disclosure may be shared among various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0150] In the illustrated example, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process generally opposite to the encoding process described for video encoder 4400.

[0151] Entropy decoding unit 4501 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., coded blocks of video data). Entropy decoding unit 4501 may decode the entropy-coded video data, and based on the entropy-decoded video data, motion compensation unit 4502 may determine motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 4502 may determine such information, for example, by performing AMVP and Merge modes.

[0152] Motion compensation unit 4502 may generate a motion-compensated block and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.

[0153] Motion compensation unit 4502 may calculate interpolation of sub-integer pixels of a reference block using the same interpolation filter as used by video encoder 4400 during encoding of a video block. Motion compensation unit 4502 may determine the interpolation filter used by video encoder 4400 based on received syntax information and may use that interpolation filter to generate a prediction block.

[0154] Motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of an encoded video sequence, partitioning information that describes how each macroblock of a picture describing the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0155] Intra prediction unit 4503 may form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. Inverse quantization unit 4504 inverse quantizes the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501, i.e., dequantizes. Inverse transform unit 4505 applies an inverse transform.

[0156] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block in order to remove block artifacts. The decoded video block is then stored in the buffer 4507, providing a reference block for subsequent motion compensation / intra prediction, and also generating decoded video for presentation on a display device.

[0157] Figure 8 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing the techniques of VVC. The encoder 4600 includes three loop filters, namely, the deblocking filter (DF) 4602, the sample adaptive offset (SAO) 4604, and the adaptive loop filter (ALF) 4606. Different from the DF 4602 that uses predefined filters, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets and applying finite impulse response (FIR) filters with the coded side information signaling the offsets and filter coefficients respectively. The ALF 4606 is located at the last processing stage of each picture and can be regarded as a tool for trying to capture and fix the artifacts created by the previous stages.

[0158] The encoder 4600 also includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive the input video. The intra prediction component 4608 is configured to perform intra prediction, while the ME / MC component 4610 is configured to perform inter prediction using the reference pictures obtained from the reference picture buffer 4612. The residual blocks from the inter prediction or intra prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy coding / decoding component 4618. The entropy coding / decoding component 4618 performs entropy coding / decoding on the prediction results and the quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 is capable of outputting the image to the DF 4602, the SAO 4604, and the ALF 4606 for filtering before these images are stored in the reference picture buffer 4612.

[0159] Next, a list of some example preferred solutions is provided.

[0160] The following solutions illustrate examples of the techniques discussed herein.

[0161] 1. A method for media data processing (e.g., Figure 4The method depicted in (4200) includes: performing a conversion between visual media information and a digital representation of the visual media information according to a rule, where the rule specifies whether or how a decoding capability information (DCI) network abstraction layer (NAL) unit is included in an elementary stream track of the digital representation.

[0162] 2. The method according to solution 1, wherein the rule specifies that a DCI NAL unit is included in each elementary stream track.

[0163] 3. The method according to any one of solutions 1-2, wherein the rule specifies that when multiple DCI NAL units are included in an elementary stream track, the multiple DCI NAL units have the same content.

[0164] 4. The method according to any one of solutions 1-2, wherein the rule specifies that only one DCI NAL unit is included in an elementary stream track.

[0165] 5. The method according to solution 1, wherein the rule specifies that a DCI NAL unit is constrained in a header sample entry of the track when present in an elementary stream track.

[0166] 6. The method according to any one of solutions 1-5, wherein the rule specifies that a DCI NAL unit complies with a constraint that a value of a field in the DCI NAL unit is constrained to be equal to a predetermined value.

[0167] 7. The method according to solution 6, wherein the field indicates the profile, layer, and one less than the number of layer structures, and wherein the predetermined value is equal to 0.

[0168] 8. The method according to solution 6, wherein the field indicates whether to enable multi-layer indication of profile-layer-level, and wherein the predetermined value is equal to 1.

[0169] 9. A method for media data processing, including: performing a conversion between visual media information and a digital representation of the visual media information according to a rule, where the rule specifies whether or how a video parameter set (VPS) unit is included in an elementary stream track of the digital representation.

[0170] 10. The method according to solution 9, wherein the rule specifies that only one VPS unit is included in an elementary stream track.

[0171] 11. The method according to any one of solutions 9-10, wherein the rule specifies that when an elementary stream track includes a VPS unit but does not include a decoding capability information (DCI) network abstraction layer (NAL) unit, the digital representation satisfies a constraint.

[0172] 12. The method according to any one of Solutions 9-11, wherein the rule specifies a constraint that a value of a field in a VPS conforming to the VPS is constrained to be equal to a predetermined value.

[0173] 13. The method according to any one of Solutions 9-12, wherein the rule specifies that a digital representation satisfies a constraint in a case where a track for encoding / decoding a basic stream does not include a VPS unit and a decoding capability information (DCI) network abstraction layer (NAL) unit.

[0174] 14. A method for media data processing, comprising: performing a conversion between visual media information and a digital representation of the visual media information according to a rule, wherein the rule specifies whether or how a value of a field included in a hypothetical reference decoder structure referred to by a video parameter set of a sequence parameter set is allowed to change from one encoded video sequence to a second encoded video sequence in an encoded video basic stream in the digital representation.

[0175] 15. The method according to Solution 14, wherein the value indicates a time scale.

[0176] 16. The method according to any one of Solutions 14-15, wherein the rule specifies that the value of the field is the same in each hypothetical reference decoder structure in the digital representation.

[0177] 17. A method for media data processing, comprising: obtaining a digital representation of visual media information, wherein the digital representation is generated according to the method according to any one of Solutions 1-16; and streaming the digital representation.

[0178] 18. A method for media data processing, comprising: receiving a digital representation of visual media information, wherein the digital representation is generated according to the method according to any one of Solutions 1-16; and generating visual media information from the digital representation.

[0179] 19. The method according to any one of Solutions 1-18, wherein the conversion includes generating a bitstream representation of visual media data and storing the bitstream representation into a file according to format rules.

[0180] 20. The method according to any one of Solutions 1-18, wherein the conversion includes parsing a file according to format rules to recover visual media data.

[0181] 21. A video decoding device, comprising a processor configured to implement the method according to one or more of Solutions 1 to 20.

[0182] 22. A video encoding device includes a processor configured to implement the method according to one or more of Solutions 1 to 20.

[0183] 23. A computer program product storing computer code that, when executed by a processor, causes the processor to implement the method according to any one of Solutions 1 to 20.

[0184] 24. A computer-readable medium having a bitstream representation conforming to a file format generated according to any one of Solutions 1 to 20.

[0185] 25. A method, device, or system described in this document. In the solutions described herein, the encoder may conform to format rules by generating an encoded / decoded representation according to the format rules. In the solutions described herein, the decoder may use the format rules to parse the syntax elements in the encoded / decoded representation, knowing the presence and absence of the syntax elements, according to the format rules, to produce a decoded video.

[0186] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, and vice versa. The bitstream representation of the current video block may correspond, for example, to bits juxtaposed or scattered in different places within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on the transformed and encoded / decoded error residual values and also using bits in the header and other fields in the bitstream. In addition, during the conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may be present or absent, as described in the above solutions. Similarly, the encoder may determine whether to include or exclude certain syntax fields and generate the encoded / decoded representation accordingly by including the syntax fields or excluding the syntax fields from the encoded / decoded representation.

[0187] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for running by, or controlling the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus can also include code that creates an execution environment for the computer programs being discussed, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A propagated signal is an artificially generated signal that is generated to encode information for transmission to a suitable receiver apparatus, e.g., a machine-generated electrical, optical, or electromagnetic signal.

[0188] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to run on one computer or on multiple computers distributed across one site or multiple sites and interconnected by a communication network.

[0189] The processes and logical flows described in this document can be performed by one or more programmable processors running one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0190] Processors suitable for running a computer program include, for example, any one or more processors of general-purpose and special-purpose microprocessors, as well as any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operably coupled to receive data from or transfer data to or receive data from and transfer data to the one or more mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read-only memory (CD ROM) and digital versatile disc read-only memory (DVD-ROM) disks. The processor and memory may be supplemented by, or incorporated in, special logic circuitry.

[0191] Although this patent document contains many details, these details should not be construed as limitations on any subject or the scope that may be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features described in the context of separate embodiments in this patent document may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Moreover, although features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be excluded from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0192] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve a desired result. Additionally, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0193] Only some embodiments and examples have been described, and other embodiments, enhancements, and variations may be made based on what is described and illustrated in this patent document.

[0194] When there is no intermediate component between the first component and the second component other than a wire, a trace, or another medium, the first component is directly coupled to the second component. When there is an intermediate component between the first component and the second component other than a wire, a trace, or another medium, the first component is indirectly coupled to the second component. The term "coupled" and its variants include direct coupling and indirect coupling. Unless otherwise specified, the use of the term "about" means a range of ±10% of the subsequent number.

[0195] Although several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. This example is considered illustrative rather than restrictive and is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated in another system, or some features may be omitted or not implemented.

[0196] In addition, without departing from the scope of the disclosure, the techniques, systems, subsystems, and methods described and shown as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as being coupled may be directly connected or may be indirectly coupled or communicate through some interface, device, or intermediate component in an electrical, mechanical, or other manner. Other examples of changes, substitutions, and alterations may be determined by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method for processing video data, comprising: Determining, based on rules, information in a sequence parameter set (SPS) in a VVC elementary stream carried in a multi-functional video coding (VVC) common media application format (CMAF) track and information in a decoding capability information (DCI) network abstraction layer (NAL) unit, where the rules specify that the value of a syntax element num_units_in_tick representing the number of units at a specified punctual point in a general hypothetical reference decoder (HRD) parameter general_timing_hrd_parameters structure and the value of a syntax element time_scale representing a specified time scale should not change between video sequences in the VVC elementary stream when present in the SPS; and where the rules further specify that the value of a syntax element dci_num_ptls_minus1 in the DCI NAL unit should be equal to 0; Performing a conversion between visual media data and a media data file based on the information in the SPS and the information in the DCI NAL unit.

2. The method according to claim 1, wherein, The VVC elementary stream includes a video parameter set (VPS), and where the rules further specify that the value of num_units_in_tick and the value of time_scale should not change between video sequences in the VVC elementary stream when present in the VPS.

3. The method according to claim 1, wherein The rules also specify that the value of num_units_in_tick and the value of time_scale should be the same for all general_timing_hrd_parameters structures of general hypothetical reference decoders in the VVC CMAF track.

4. The method according to claim 1, wherein The rules also specify that there should be one and only one video parameter set (VPS) unit in the VVC CMAF track.

5. The method according to claim 1, wherein The rules also specify that when there is no decoding capability information DCI network abstraction layer NAL unit in the VVC CMAF track and when there is no video parameter set VPS in the VVC CMAF track, the value of a general profile identifier general_profile_idc, the value of a general tier flag general_tier_flag, the value of a general level identifier general_level_idc, the value of a number of sub-layer profiles num_sub_profiles, and the value of a general sub-profile identifier general_sub_profile_idc[i] for each i-th interoperability indicator should not change from one video sequence to another in the VVC elementary stream.

6. The method according to claim 1, wherein, The rules specify that when there is no Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit in the VVC CMAF track and when there are one or more Video Parameter Sets (VPSs) in the VVC CMAF track, one or more constraints apply, and wherein the one or more constraints include: the value of the vps_max_layers_minus1 field for each VPS should be equal to 0, the value of the vps_max_layers_minus1 of the number of VPSs of Profile, Tier, and Level (PTLs) should be equal to 0 for each VPS, the value of the ptl_fram_only_constraint_flag in the profile_tier_level structure of PTL frames in each VPS should be equal to 1, and the value of the ptl_multilayer_enabled_flag in the profile_tier_level structure of each VPS should be equal to 0.

7. The method according to claim 1, wherein, The rules also specify that the values of the general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, and general_sub_profile_idc[i] of the i-th interoperability indicator should not change from one video sequence to another in the VVC elementary stream.

8. The method according to any one of claims 1-7, wherein The transformation includes encoding the visual media data into the media data file.

9. The method according to any one of claims 1-7, wherein, The transformation includes decoding the visual media data from the media data file.

10. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: Based on rules, determine information in the Sequence Parameter Set (SPS) in the VVC elementary stream carried in the VVC Common Media Application Format (CMAF) track and information in the Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit, wherein the rules specify that the value of the num_units_in_tick syntax element, which is the number of units in a specified tick in the general_timing_hrd_parameters structure of the General Hypothetical Reference Decoder (HRD) parameters, and the value of the time_scale syntax element of the specified time base should not change between video sequences in the VVC elementary stream when present in the SPS; and wherein the rules also specify that the value of the dci_num_ptls_minus1 syntax element in the DCI NAL unit should be equal to 0; Perform the conversion between the visual media data and the media data file based on the information in the SPS and the information in the DCI NAL unit.

11. The device according to claim 10, wherein, The VVC elementary stream includes a Video Parameter Set (VPS), and the rule further specifies that the value of num_units_in_tick and the value of time_scale should not change between video sequences in the VVC elementary stream when they are present in the VPS.

12. The device according to claim 10, wherein, The rule also specifies that the value of num_units_in_tick and the value of time_scale should be the same for all General Hypothetical Reference Decoder (HRD) parameter general_timing_hrd_parameters structures in the VVC CMAF track.

13. The apparatus according to claim 10, wherein, The rule also specifies that there should be one and only one VPS unit in the VVC CMAF track.

14. The apparatus according to claim 10, wherein, The rule also specifies that when there is no Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit in the VVC CMAF track and when there is no VPS in the VVC CMAF track, the value of the general_profile_idc, the value of the general_tier_flag, the value of the general_level_idc, the value of num_sub_profiles, and the value of general_sub_profile_idc[i] for each i-th interoperability indicator should not change from one video sequence to another in the VVC elementary stream.

15. A non-transitory computer-readable medium, comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, cause the video codec device to: Determine, based on a rule, information in a Sequence Parameter Set (SPS) in a VVC elementary stream carried in a VVC Common Media Application Format (CMAF) track and information in a Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit, where the rule specifies that the value of the syntax element num_units_in_tick, which represents the number of units in a specified tick in the General Hypothetical Reference Decoder (HRD) parameter general_timing_hrd_parameters structure, and the value of the syntax element time_scale, which represents a specified time scale, should not change between video sequences in the VVC elementary stream when they are present in the SPS; and where the rule also specifies that the value of the syntax element dci_num_ptls_minus1 in the DCI NAL unit should be equal to 0; Perform the conversion between the visual media data and the media data file based on the information in the SPS and the information in the DCI NAL unit.

16. The non-transitory computer-readable medium according to claim 15, wherein, The VVC elementary stream includes a Video Parameter Set (VPS), and wherein the rule further specifies that the values of num_units_in_tick and time_scale, when present in the VPS, shall not change between video sequences of the VVC elementary stream.

17. The non-transitory computer-readable medium according to claim 15, wherein, The rule also specifies that the values of num_units_in_tick and time_scale shall be the same for all General Hypothetical Reference Decoder (HRD) parameter general_timing_hrd_parameters structures in the VVC CMAF track.