Encapsulation and constraints of adaptive video streaming

By introducing rule constraints to manage DCI NAL units in the VVC CMAF track, the inflexible VVC media data transmission problem in the existing technology is solved, achieving more efficient video data processing and a high-quality user experience.

CN115225910BActive Publication Date: 2025-09-16FACE CUTE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210406258.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-18
Filing Date
2022-04-18
Publication Date
2025-09-16
Estimated Expiration
2042-04-18

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively managing and transmitting Versatile Video Codec (VVC) media data when processing video data, especially in adaptive streaming systems, resulting in inflexible and inefficient use of Decoding Capability Information (DCI) Network Abstraction Layer (NAL) units.

Method used

By introducing rule constraints into the Versatile Video Codec (VVC) Common Media Application Format (CMAF) track, the presence and content consistency of DCI NAL units in the VVC CMAF track are ensured, thereby achieving effective management and transmission of decoding capability information.

Benefits of technology

This method improves the flexibility and efficiency of video data processing, enables the adaptive streaming system to better adapt to the decoding capabilities of different devices, and improves the transmission quality of media data and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115225910B_ABST
    Figure CN115225910B_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. Information in a Decoding Capability Information (DCI) Network Abstraction Layer (NAL) unit is determined. Rules governing the use of DCI NAL units in a Versatile Video Codec (VVC) Common Media Application Format (CMAF) track are also provided. Conversion between visual media data and media data files is performed based on the DCI NAL units.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is intended to timely claim priority to and the benefit of U.S. Provisional Patent Application No. 63 / 176,315, filed on April 18, 2021, under applicable patent laws and / or rules of the Paris Convention. The entire disclosure of the above application is incorporated by reference as a part of the disclosure of this application for all purposes prescribed by law. Technical Field

[0003] This patent document relates to the generation, storage and consumption of digital audio and video media information in file format. Background Art

[0004] Digital video occupies the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is likely to continue to grow. Summary of the Invention

[0005] A first aspect relates to a method for processing video data, comprising: determining information in a decoding capability information (DCI) network abstraction layer (NAL) unit, wherein rules govern use of the DCI NAL unit in a versatile video codec (VVC) common media application format (CMAF) track; and performing conversion between visual media data and a media data file based on the DCI NAL unit.

[0006] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that the DCI NAL unit should be present in the VVC CMAF track.

[0007] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that the DCI NAL unit should be present in the VVC CMAF track.

[0008] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that all DCI NAL units present in the VVC CMAF track should have the same content.

[0009] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the rule specifies that there is one and only one DCI NAL unit in the VVC CMAF track.

[0010] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit should be present in a CMAF header sample entry.

[0011] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the DCI NAL unit includes bitstream level profile, layer and level (PTL) information.

[0012] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the converting includes encoding the visual media data into a media data file.

[0013] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the converting includes decoding the visual media data from a media data file.

[0014] A second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determine information in a decoding capability information (DCI) network abstraction layer (NAL) unit, wherein rules constrain use of the DCI NAL unit in a versatile video codec (VVC) common media application format (CMAF) track; and perform conversion between visual media data and a media data file based on the DCI NAL unit.

[0015] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that the DCI NAL unit should be present in the VVC CMAF track.

[0016] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that the DCI NAL unit should be present in the VVC CMAF track.

[0017] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that all DCI NAL units present in the VVC CMAF track should have the same content.

[0018] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the rule specifies that there is one and only one DCI NAL unit in the VVC CMAF track.

[0019] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit should be present in a CMAF header sample entry.

[0020] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, which, when executed by a processor, causes the video codec device to: determine information in a decoding capability information (DCI) network abstraction layer (NAL) unit, wherein rules constrain the use of the DCI NAL unit in a versatile video codec (VVC) common media application format (CMAF) track; and perform conversion between visual media data and a media data file based on the DCI NAL unit.

[0021] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that the DCI NAL unit should be present in the VVC CMAF track.

[0022] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that the DCI NAL unit should be present in the VVC CMAF track.

[0023] Optionally, in any of the aforementioned aspects, another embodiment of this aspect provides that the rule specifies that all DCI NAL units present in the VVC CMAF track should have the same content.

[0024] Optionally, in any of the preceding aspects, another embodiment of this aspect provides that the rule specifies that there is one and only one DCI NAL unit in the VVC CMAF track.

[0025] For purposes of clarity, any of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0026] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in connection with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0028] Figure 1 is a diagram illustrating an example Common Media Application Format (CMAF) track.

[0029] Figure 2 is a block diagram illustrating an example video processing system.

[0030] Figure 3 is a block diagram of an example video processing device.

[0031] Figure 4is a flow chart of an example method of video processing.

[0032] Figure 5 is a block diagram illustrating an example video encoding and decoding system.

[0033] Figure 6 is a block diagram illustrating an example encoder.

[0034] Figure 7 is a block diagram illustrating an example decoder.

[0035] Figure 8 is a schematic diagram of an example encoder. DETAILED DESCRIPTION

[0036] It should be understood at the outset that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques shown below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.

[0037] This patent document relates to video streaming. Specifically, the document relates to specifying constraints on video encoding and encapsulation into media tracks and segments in file formats. Such file formats may include the International Organization for Standardization (ISO) Base Media File Format (ISOBMFF). Such file formats may also include adaptive streaming media representation formats, such as Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) and / or Common Media Application Format (CMAF). For media streaming systems, the ideas described herein can be applied alone or in various combinations, such systems being based on the DASH standard and related extensions and / or based on the CMAF standard and related extensions.

[0038] This disclosure includes the following abbreviations: Adaptive Color Transform (ACT), Adaptive Loop Filter (ALF), Adaptive Motion Vector Resolution (AMVR), Adaptive Parameter Set (APS), Access Unit (AU), Access Unit Delimiter (AUD), Advanced Video Codec ((Rec. ITU-T H.264 | ISO / IEC 14496-10)) (AVC), Bidirectional Prediction (B), Bidirectional Prediction with Codec Level Weights (BCW), Bidirectional Optical Flow (BDOF), Block-Based Delta Pulse Codec Modulation (BDPCM), Buffer Period (BP), Context-Based Adaptive Binary Arithmetic Codec (CABAC), Codec Block (CB), Constant Bit Rate (CBR), Cross-Component Adaptive Loop Filter (CCALF), Codec Picture Buffer (CPB), Clean Random Access (CRA), Cyclic Redundancy Check (CRC), Codec Tree Block (CTB), Codec Tree Unit (CTU), Codec Unit (CU) ), Coded Video Sequence (CVS), Decoding Capability Information (DCI), Decoding Initialization Information (DII), Decoded Picture Buffer (DPB), Dependent Random Access Point (DRAP), Decoding Unit (DU), Decoding Unit Information (DUI), Exponential-Golomb (EG), Exponential-Golomb k (EGk), End of Bitstream (EOB), End of Sequence (EOS), Padding Data (FD), First In First Out (FIFO), Fixed Length (FL), Green, Blue, and Red (GBR), General Constraint Information (GCI), Progressive Decoding Refresh (GDR), Geometric Partitioning Mode (GPM), High Efficiency Video Codec (also known as Rec. ITU-T H.265|ISO / IEC 23008-2)(HEVC), Hypothetical Reference Decoder (HRD), Hypothetical Stream Scheduler (HSS), Intra (I), Intra Block Copy (IBC), Instantaneous Decoding Refresh (IDR), Inter-layer Reference Picture (ILRP), Intra Random Access Point (IRAP), Low-Frequency Non-separable Transform (LFNST), Least Probable Symbol (LPS), Least Significant Bit (LSB), Long-Term Reference Picture (LTRP), Luma Mapping and Chroma Scaling (LMCS), Matrix-based Intra Prediction (MIP), Most Probable Symbol (MPS), Most Significant Bit (MSB), Multiple Transform Selection (MTS), Motion Vector Prediction (MVP), Network Abstraction Layer (NAL), Output Layer Set (OLS), Operation Point (OP), Operation Point Information (OPI), Prediction (P), Picture Header (PH), Picture Order Count (POC), Picture Parameter Set (P

[0014] The present invention relates to the following video codes:

[0014] ,

[0015] ,

[0016] ,

[0017] ,

[0018] ,

[0019] ,

[0020] ,

[0021] ,

[0022] ,

[0023] ,

[0024] ,

[0025] ,

[0026] ,

[0027] ,

[0028] ,

[0029] ,

[0030] ,

[0031] ,

[0032] ,

[0033] ,

[0034] ,

[0035] ,

[0036] ,

[0037] ,

[0038] ,

[0039] ,

[0040] ,

[0041] ,

[0042] ,

[0043] ,

[0044] ,

[0045] ,

[0046] ,

[0047] ,

[0048] ,

[0049] ,

[0050] ,

[0051] ,

[0052] ,

[0053] ,

[0054] ,

[0055] ,

[0056] ,

[0057] ,

[0058] ,

[0059] .

[0039] Video codec standards have evolved primarily through the development of standards within the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the ISO / International Electrotechnical Commission (IEC). ITU-T developed H.261 and H.263, while ISO / IEC developed the Moving Picture Experts Group (MPEG)-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore video codec technologies beyond HEVC, the Video Codec Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). JVET adopted many approaches and incorporated them into reference software called the Joint Exploration Model (JEM). When the Versatile Video Codec (VVC) project officially began, JVET was later renamed the Joint Video Experts Team (JVET). VVC is a codec standard that aims to reduce bit rate by 50% compared to HEVC. VVC has been completed by JVET.

[0040] The VVC standard (also known as ITU-T H.266 | ISO / IEC 23090-3) and the associated Versatile Supplemental Enhancement Information (VSEI) standard (also known as ITU-T H.274 | ISO / IEC 23002-7) are designed for a wide range of applications, such as television broadcasting, video conferencing, playback from storage media, adaptive bitrate streaming, video region extraction, compositing and merging of content from multiple codec video bitstreams, multi-view video, scalable layered coding, and viewport-adaptive three hundred sixty degree (360°) immersive media. The Essential Video Codec (EVC) standard (ISO / IEC 23094-1) is another video codec standard developed by MPEG.

[0041] File format standards are discussed below. Media streaming applications are typically based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport methods, and often rely on file formats such as ISOBMFF. One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). Video can be encoded in a video format such as AVC and / or HEVC. The encoded video can be encapsulated in ISOBMFF tracks and included in DASH representations and segments. For content selection purposes, important information about the video bitstream, such as profile, tier, and level, can be presented as file format level metadata and / or in the DASH Media Presentation Description (MPD). For example, such information can be used to select appropriate media segments for initialization at the start of a streaming session and for stream adaptation during a streaming session.

[0042] Similarly, when using an image format with ISOBMFF, file format specifications specific to that image format can be adopted, such as the AVC image file format and the HEVC image file format. MPEG is developing the VVC video file format, which is a file format based on ISOBMFF for storing VVC video content. MPEG is also developing the VVC image file format based on ISOBMFF, which is a file format for storing image content using the VVC codec.

[0043] Now let's discuss file format standards. Media streaming applications can be based on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport mechanisms. Such media streaming applications can also rely on file formats such as the ISO Base Media File Format (ISOBMFF). One such streaming system is Dynamic Adaptive Streaming over HTTP (DASH). To use video formats with ISOBMFF and DASH, file format specifications specific to that video format can be adopted to encapsulate video content in ISOBMFF tracks and DASH representations and segments. Such file format specifications can include the AVC file format and the HEVC file format. For content selection purposes, important information about the video bitstream, such as profile, tier, and level, can be presented in file format-level metadata and / or the DASH Media Presentation Description (MPD). For example, content selection can include selecting appropriate media segments for initialization at the start of a streaming session and for stream adaptation during the streaming session. Similarly, to use image formats with ISOBMFF, file format specifications specific to that image format can be adopted, such as the AVC image file format and the HEVC image file format. The VVC video file format is based on ISOBMFF and is used to store VVC video content. The VVC video file format was developed by MPEG. The VVC image file format is based on ISOBMFF and is used to store image content using the VVC codec. The VVC image file format was also developed by MPEG.

[0044] Now let's discuss DASH. In DASH, there may be multiple representations of video and / or audio data of multimedia content. Different representations may correspond to different codec characteristics, such as different profiles or levels of a video codec standard, different bit rates, different spatial resolutions, etc. A list of such representations may be defined in a media presentation description (MPD) data structure. A media presentation may correspond to a structured set of data accessible to a DASH streaming client device. A DASH streaming client device may request and download media data information to present a streaming service to a user of the client device. A media presentation may be described in an MPD data structure, which may include updates to the MPD.

[0045] Media presentation can comprise a sequence of one or more cycles. Each cycle can extend to the beginning of the next cycle, or in the case of the last cycle, to the end of media presentation. Each cycle can comprise one or more representations of the same media content. A representation can be one of a plurality of alternative coded versions of audio, video, timed text or other such data. These representations can be different according to the coding type, such as the bit rate, resolution, codec of video data and the bit rate, language and / or codec of audio data. The term representation can be used to refer to a part of coded audio or video data that is encoded in a particular manner and corresponding to a specific cycle of multimedia content.

[0046] Representations of a particular period may be assigned to a group indicated by an attribute in the MPD, which indicates the adaptation set to which the representation belongs. Representations in the same adaptation set are generally considered to be mutually interchangeable. Thus, a client device can dynamically and seamlessly switch between these representations, for example to perform bandwidth adaptation. For example, each representation of video data for a particular period may be assigned to the same adaptation set, so that any representation may be selected for decoding to present media data of the multimedia content of the corresponding period, such as video data or audio data. In some examples, the media content within a period may be represented by a combination of one representation from group 0 (if any) or at most one representation from each non-zero group. The timing data for each representation of a period may be expressed relative to the start time of the period.

[0047] A representation may include one or more segments. Each representation may include an initialization segment, or each segment of a representation may be self-initializing. When present, the initialization segment may contain initialization information for accessing the representation. Typically, the initialization segment does not contain media data. Segments may be uniquely referenced by an identifier, such as a Uniform Resource Locator (URL), Uniform Resource Name (URN), or Uniform Resource Identifier (URI). The MPD may provide an identifier for each DD, which may correspond to the data of the segment within a file accessible by the URL, URN, or URI.

[0048] For different types of media data, different representations can be selected for substantially simultaneous retrieval. For example, a client device can select an audio representation, a video representation, and a timed text representation from which to retrieve a segment. In some examples, the client device can select a specific adaptation set to perform bandwidth adaptation. For example, the client device can select an adaptation set that includes a video representation, an adaptation set that includes an audio representation, and / or an adaptation set that includes timed text. In an example, the client device can select an adaptation set for certain types of media, such as video, and directly select representations for other types of media, such as audio and / or timed text.

[0049] An example DASH streaming process can be illustrated by the following steps. The client obtains the MPD. The client then estimates the downlink bandwidth and selects a video representation and an audio representation based on the estimated downlink bandwidth, codec, decoding capabilities, display size, audio language settings, etc. Until the end of the media presentation is reached, the client requests media segments of the selected representation and presents the streaming content to the user. The client continuously estimates the downlink bandwidth. When the bandwidth changes significantly, such as by becoming lower or higher, the client selects a different video representation to match the new estimated bandwidth and continues downloading the segments at the updated downlink bandwidth.

[0050] Now let's talk about CMAF. CMAF specifies a set of constraints on the encoding and encapsulation of media into ISOBMFF tracks, ISOBMFF segments, ISOBMFF fragments, DASH representations, and / or CMAF tracks, CMAF segments, etc. Such constraints are encapsulated for each interoperability point defined as a media profile. The main goal of CMAF development was to enable the reuse of the same media content encoded with a specific codec (e.g., AVC for video) and encapsulated in a specific format (e.g., ISOBMFF) across the two separate media streaming worlds of DASH and Apple HTTP Live Streaming (HLS).

[0051] Now let’s discuss Decoding Capability Information (DCI) in VVC. DCI NAL units contain bitstream-level profile, layer, and level (PTL) information. DCI NAL units include one or more PTL syntax structures that can be used during session negotiation between the sender and receiver of a VVC bitstream. When a DCI NAL unit is present in a VVC bitstream, each output layer set (OLS) in the CVS of the bitstream should conform to the PTL information carried in at least one PTL structure in the DCI NAL unit. In AVC and HEVC, the PTL information used for session negotiation is available in the SPS (for HEVC and AVC) and VPS (for HEVC layered extensions). This design of conveying PTL information for session negotiation in HEVC and AVC has a disadvantage because the scope of the SPS and VPS is within the CVS, not within the entire bitstream. This can cause the sender-receiver session initiation to be subject to re-initiation during bitstream streaming of each new CVS. DCI solves this problem because DCI carries bitstream-level information, so the indicated decoding capabilities can be guaranteed to be adhered to until the end of the bitstream.

[0052] Let's now discuss the Video Parameter Set (VPS) in VVC. A VVC bitstream may contain a Video Parameter Set (VPS) which contains information describing layers and the Output Layer Set (OLS) for the operation of the decoding process of a scalable bitstream. An OLS is a set of layers in a bitstream where one or more layers are designated to be output from the decoder. Other layers identified in the OLS may also be decoded in order to decode the output layers, although such layers are not designated to be output. Much of the information contained in the VPS can be used in the system for purposes such as session negotiation and content selection. The VPS was introduced to handle multi-layer bitstreams. For a single-layer VVC bitstream, the presence of a VPS in the CVS is optional. This is because the information contained in the VPS is not required for the operation of the decoding process of the bitstream. The absence of a VPS in the CVS is indicated by reference to a VPS identifier (ID) equal to 0 in the SPS, in which case default values ​​are inferred for the VPS parameters.

[0053] Now let's discuss the Sequence Parameter Set (SPS) in VVC. The SPS conveys sequence-level information shared by all pictures in the entire Codec Layer Video Sequence (CLVS). This includes PTL indicators, picture formats, feature and / or tool control flags, codec, prediction and / or PR transform block structure and hierarchy, candidate RPLs that the encoder can reference, etc. Picture formats can include color sampling format, maximum picture width, maximum picture height, and bit depth. In most applications, the entire bitstream uses only one or a few SPSs. Therefore, there is no need to update the SPS within the bitstream. Updating the SPS can include sending a new SPS using the SPS ID of an existing SPS, but with different values ​​for certain parameters. Pictures from a particular layer that reference an SPS with a different SPS ID or an SPS with the same SPS ID but different SPS content belong to different CLVSs. As with AVC and HEVC, the SPS can be transmitted in-band or using a mix of in-band and out-of-band signaling. In-band signaling indicates that data such as the SPS is transmitted with the codec picture, and out-of-band signaling indicates that data such as the SPS is not transmitted with the codec picture.

[0054] Let's now discuss the Picture Parameter Set (PPS) in VVC. The PPS conveys picture-level information shared by all slices of a picture. Such information can also be shared across multiple pictures. This includes feature and / or tool on / off flags, picture width and height, default RPL size, slice and slice configuration, etc. By design, two consecutive pictures can refer to two different PPSs. This may result in a large number of PPSs being used within the CLVS. In practice, the number of PPSs for the entire bitstream may not be high because the PPS is designed to carry parameters that do not change frequently and may apply to multiple pictures. Therefore, there may be no need to update the PPS within the CLVS or even the entire bitstream. The Adaptive Parameter Set (APS) can be used for parameters that may apply to multiple pictures but are expected to change frequently between different pictures. Like the SPS, the PPS can be transmitted in-band, out-of-band, or using a mix of in-band and out-of-band signaling. A basic design principle regarding which picture-level parameters should be included in the PPS and which should be in the APS is the frequency with which such parameters are likely to change. Therefore, frequently changing parameters are not included in the PPS to avoid requiring PPS updates, which in typical use cases would not allow out-of-band transmission of the PPS.

[0055] Now let's discuss the Adaptive Parameter Set (APS) in VVC. The APS conveys picture and / or slice level information that can be shared by multiple slices of a picture and / or slices of different pictures, but can change frequently across pictures. The APS supports information with a large number of variables that are not suitable for inclusion in the PPS. Three types of parameters are included in the APS, including adaptive loop filter (ALF) parameters, luma mapping and chroma scaling (LMCS) parameters, and scaling list parameters. The APS can be carried in two different NAL unit types, which can be prefixed or suffixed before or after the associated slice. The latter can be helpful in ultra-low latency scenarios, such as allowing the encoder to send the slices of a picture before generating ALF parameters based on the picture, which will be used by subsequent pictures in decoding order.

[0056] Now let's discuss the picture header (PH). For each PU, there is a picture header (PH) structure. The PH exists as a separate PH NAL unit or is included in the slice header (SH). If the PU only includes one slice, the PH can only be included in the SH. To simplify the design, within the CLVS, the PH can only be in the PH NAL unit or in the SH. When the PH is in the SH, there is no PH NAL unit in the CLVS. The design of the PH has two purposes. First, the PH achieves this by carrying all parameters that have the same value for all slices of the picture, thereby preventing the same parameters from being repeated in each SH. These include IRAP and / or GDR picture indication, inter-frame and / or intra-frame slice enable flags, and information related to POC, RPL, deblocking filter, SAO, ALF, LMCS, scaling lists, QP increment, weighted prediction, codec block partitioning, virtual boundaries, collocated pictures, etc. Second, the PH helps the decoder identify the first slice of each codec picture that contains multiple slices. Since there is one and only one PH for each PU, when the decoder receives a PH NAL unit, the decoder knows that the next VCL NAL unit is the first slice of the picture.

[0057] Now let's discuss Operation Point Information (OPI). The decoding processes of HEVC and VVC have similar input variables to set the decoding operation point. These include the target OLS and highest sublayer of the bitstream to be decoded via the decoder API. However, in scenarios where layers and / or sublayers of the bitstream are removed during transmission or the device does not expose the decoder application programming interface (API) to the application, the decoder may not be able to correctly determine the operation point for processing the bitstream. As a result, the decoder may not be able to infer properties of the pictures in the bitstream, such as the appropriate buffer allocation for decoding pictures and whether to output individual pictures. To address this issue, VVC includes a mode for indicating these two variables within the bitstream via the OPI NAL unit. In the initial AU of the bitstream and in the individual CVS of the bitstream, the OPI NAL unit informs the decoder of the target OLS and highest sublayer of the bitstream to be decoded. In the case where the OPI NAL unit is present and the operation point is also provided to the decoder via decoder API information, the decoder API information takes precedence. For example, the application may have more updated information related to the target OLS and sublayers. In the absence of a decoder API and any OPI NAL units in the bitstream, appropriate fallback selections are specified in VVC to allow correct decoder operation.

[0058] An example CMAF specification is now discussed. A VVC video CMAF track can be described as follows. A VVC CMAF track SHOULD conform to the requirements for a NAL structured video CMAF track. In addition, a CMAF track can conform to all other requirements described herein. If a CMAF track conforms to these requirements, the CMAF track is referred to as a VVC video CMAF track and can use the brand "cvvc". VVC video track constraints are also discussed. In the example, the VVC video CMAF switch set constraints are as follows. Each CMAF track in a CMAF switch set SHOULD conform to the VVC video CMAF track defined herein. A VVC video CMAF switch set SHOULD conform to the constraints of the NAL structured video CMAF switch set.

[0059] Now let's discuss the visual sample entries. The syntax and values ​​of the VVC video track visual sample entries shall conform to the VVCSampleEntry("vvc1") or VVCSampleEntry("vvci") sample entries. Now let's discuss the constraints on the VVC elementary stream. With respect to the VPS, each VVC video media sample in the CMAF track shall reference an SPS with a sps_video_parameter_set_id aCMAF header sample entry. If present, the following additional constraints apply. For each profile_tier_level() structure in the VPS, the values ​​of the following fields shall not change throughout the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_subc_profile_idc[i].

[0060] The SPS NAL units appearing within a CMAF VVC track shall conform to the constraints herein, with the following additional constraints. The following fields shall have the following predetermined values: First, the vui_parameters_present_flag shall be set to 1; Second, if the profile_tier_level() structure is present in the SPS, the conditions of the following fields shall not change across the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i].

[0061] Now let's discuss the image crop parameters. The SPS and PPS crop parameters conf_win_top_offset and conf_win_left_offset should be set to 0. The SPS and PPS crop parameters conf_win_bottom_offset and conf_win_right_offset can be set to values ​​other than 0. If set to non-zero values, such syntax elements are expected to be used by CMAF players to remove video spatial samples that are not intended to be displayed.

[0062] Now let's discuss the video codec parameters. VVC signaling of codec parameters (information) is described below. The demo application SHOULD use parameter signaling to signal the video codec profile and level for each VVC track and CMAF switch set. Encryption is also discussed. Encryption of CMAF VVC tracks and CMAF VVC switch sets SHOULD use the "cenc" AES-CTR scheme or the "cbcs" AES-CBC subsample mode encryption scheme. Additionally, if the "cbcs" mode of the general encryption is used for mode encryption, then mode block length 10 and Encryption: Skip Mode 1:9 SHOULD be applied.

[0063] The following are example technical issues addressed by the disclosed technical solutions. For example, in the example VVC CMAF design, the profile, tier, and level may need to be signaled in the VPS and SPS, so they may not change across the entire VVC elementary stream. However, for VVC bitstreams, DCI NAL units can instead be used to convey the required decoding capabilities for the entire bitstream, while allowing the profile, tier, and level within the bitstream to differ between CVSs. This will allow for greater flexibility, resulting in less transcoding and other processes required in content preparation.

[0064] This document discloses a mechanism for solving one or more of the problems listed above. For example, a VVC elementary stream, also known as a VVC bitstream, can be included in a VVC CMAF track. A VVC elementary stream may include one or more CVSs. The profile, layer, and level (PTL) information of the bitstream can vary between CVSs in the same bitstream. To allow this functionality, the PTL information can be signaled in the DCI NAL unit, VPS, and / or SPS as long as the corresponding constraints are maintained. In an example, the DCI NAL unit needs to be included in the CMAF track. In an example, when multiple DCI NAL units are included in the CMAF track, all DCI NAL units may need to include the same content. In another example, the CMAF track may include only a single DCI NAL unit. In an example, the DCI NAL unit may need to be included in the CMAF header sample entry. In various examples, the DCI NAL unit may contain a DCI number of PTL minus 1 (dci_num_ptls_minus1) field, a PTL frame only constraint flag (ptl_frame_only_constraint_flag) field, and a PTL multilayer enabled flag (ptl_multilayer_enabled_flag), which need to be equal to 0, 1, and 0, respectively. In the example, the CMAF track is restricted to containing a single VPS. In the example, when the DCI NAL unit is not present, various PTL related information in the VPS needs to be set to predetermined values ​​and / or needs to remain the same between CVSs, as further discussed below. In the example, various PTL related information in the SPS needs to be set to predetermined values ​​and / or needs to remain the same between CVSs, as further discussed below. In the example, timing-related hypothetical reference decoder (HRD) parameters may also need to remain the same between CVSs.

[0065] Figure 1 is a schematic diagram illustrating an example CMAF track 100. A CMAF track 100 is a track of video data that has been encapsulated based on constraints specified in the CMAF standard. The CMAF track 100 is constrained to support delivery and decoding by a wide range of client devices in accordance with adaptive streaming. In adaptive streaming, a media profile describes multiple different interchangeable representations, which allows a client device to select a desired representation based on decoder capabilities and / or current network conditions. A CMAF track 100 can support such functionality by including representations that are constrained to be decodable by a client that is capable of decoding at the corresponding profile, tier, and level (PTL), is capable of using the corresponding codec tools, and / or is capable of meeting other predetermined constraints.

[0066] A CMAF track 123 may contain many types of decodable video streams. In this example, the CMAF track 123 contains a VVC stream 121. A VVC stream 121, also known as a bitstream, is a stream of video data that has been encoded and decoded according to the VVC standard. For example, a VVC stream 121 may include a stream of coded and decoded pictures and associated syntax that describes the encoding and decoding process and / or other data useful to the decoder. A VVC stream 121 may contain one or more CVSs 117. A CVS 117 is a sequence of access units (AUs) in decoding order. An AU is a collection of one or more pictures with corresponding output / display times. Thus, a CVS 117 contains a series of related pictures and corresponding syntax for supporting decoding and / or describing the pictures.

[0067] CVS 117 may include DCI NAL units 115, VPS 111, SPS 113, and / or codec video 119. DCI NAL units 115 contain information describing the requirements for decoding the video data in CVS 117 and / or the entire VVC stream 121. DCI NAL units 115 are optional and may be omitted in some VVC streams 121 and / or CVS 117. It should be noted that although depicted as part of VVC stream 121, in some examples, DCI NAL units 115 may also be included in the CMAF header sample entry in CMAF track 123. VPS 111 may contain data related to the entire VVC stream 121. For example, VPS 111 may contain data related to the output layer set (OLS), layers, and / or sub-layers used in the VVC stream 121. VPS 111 is optional and may be omitted in some VVC streams 121 and / or CVS 117. SPS 113 contains sequence data common to all pictures in CVS 117 included in VVC stream 121. Parameters in SPS 113 may include picture size, bit depth, codec tool parameters, bit rate limits, etc. SPS 113 should be included in at least one CVS 117. However, multiple CVSs 117 can refer to the same SPS 113. Therefore, VVC stream 121 should contain one or more SPSs 113. Codec video 119 includes pictures coded and decoded according to VVC and the corresponding syntax.

[0068] The present disclosure relates to constraints applied to syntax elements contained in ( ) that are required to be present in a CMAF track 123. In an example, when more than one DCI NAL unit 115 is present in a single CMAF track 123, all such DCI NAL units 115 may need to include the same content. In some examples, a CMAF 123 track may be restricted to include one and only one DCI NAL unit 115. In this case, multiple CVSs 117 of video content may be described by a single DCI NAL unit 115. When present, the DCI NAL unit 115 may include a DCI number of PTLs minus 1 (dci_num_ptls_minus1) 132 and / or a PTL syntax (profile_tier_level) structure 130. dci_num_ptls_minus1 132 may specify the number of profile_tier_level structures 130 contained in the DCI NAL unit 115 in a minus-1 format. The minus-1 format indicates that the value contained in the syntax element is one less than the actual value, so 1 is added to the value contained in the syntax element to determine the actual value. In an example, dci_num_ptls_minus1 132 may be constrained to be equal to 0, indicating a single profile_tier_level structure 130. This indicates that the CMAF track 123 contains video that adheres to a single PTL information set. Depending on the example, the profile_tier_level structure 130 may be included in a DCI NAL unit 115, a VPS 111, and / or an SPS 113, and will be discussed in more detail below.

[0069] In an example, a CMAF track 123 is restricted to containing one and only one VPS 111. In this case, multiple CVSs 117 may be described by a single VPS 111. The VPS 111 may contain a VPS maximum layers minus 1 (vps_max_layers_minus1) field 134, a number of VPSs for PTL minus 1 (vps_num_ptls_minus1) field 133, a general hypothetical reference decoder (HRD) parameters (general_timing_hrd_parameters) structure 131, and a profile_tier_level structure 130. The vps_max_layers_minus1 field 134 indicates the number of layers specified by the VPS 111 in a minus-1 format. In an example, the vps_max_layers_minus1 field 134 may be restricted to contain a value of 0, which indicates that the VPS 111 describes a single layer. The vps_num_ptls_minus1 field 133 may specify the number of profile_tier_level structures 130 contained in the VPS 111 in a minus 1 format. In an example, the vps_num_ptls_minus1 field 133 may be restricted to contain a value of 0, which indicates to the VPS 111 a single PTL information set.

[0070] Depending on the example, general_timing_hrd_parameters 131 may be included in VPS 111 and / or SPS 113. For example, when VPS 111 is included, VPS 111 may include general_timing_hrd_parameters 131. When VPS 111 is not included, the SPS may include general_timing_hrd_parameters 131. general_timing_hrd_parameters 131 include timing-related parameters used by an HRD operating at an encoder. Typically, the HRD may use the HRD parameters to check that the VVC stream 121 conforms to the VVC standard. general_timing_hrd_parameters 131 indicate timing parameters related to encoding and decoding video 119 to the encoder. For example, general_timing_hrd_parameters 131 may indicate how quickly each picture should be decoded and reconstructed by the decoder for accurate display. In an example, general_timing_hrd_parameters 131 may include a time scale (time_scale) field and a number of units in a tick (num_units_in_tick) field. The time_scale field indicates the number of time units that pass in one second, where the time unit corresponds to the picture rate frequency of the video signal. num_units_in_tick indicates the number of time units of a clock operating at the frequency of time_scale in Hertz (Hz), which corresponds to one increment called a clock tick. In an example, the values ​​of num_units_in_tick and time_scale in general_timing_hrd_parameters 131 are constrained to remain constant across CVSs 117 in the same VVC stream 121. In an example, the values ​​of num_units_in_tick and time_scale in general_timing_hrd_parameters 131 are constrained to remain constant for the entire CMAF track 123.

[0071] The SPS 113 may include the general_timing_hrd_parameters 131 as discussed above, for example, when the VPS 111 is not included. For example, when the DCI NAL unit 115 and / or the VPS 111 are not included, the SPS 113 may also include the profile_tier_level structure 130. The SPS 113 may also include a video usability information payload (vui_payload) structure 135. The vui_payload structure 135 contains information that describes how the decoder should use the coded video 119. For example, the vui_payload structure 135 may include a video usability information progressive source flag (vui_progressive_source_flag) field 139 and a video usability information interlaced source flag (vui_interlaced_source_flag) field 138. The vui_progressive_source_flag field 139 may be set to indicate whether the video in the CMAF track 123 is coded or decoded according to progressive scanning. The vui_interlaced_source_flag field 138 may be set to indicate whether the video in the CMAF track 123 is coded according to interlaced. In an example, the vui_interlaced_source_flag field 138, the vui_progressive_source_flag field 139, or both may need to be set to 1. This indicates that the codec video 119 is coded according to interlaced, progressive, or both.

[0072] As described above, the DCI NAL unit 115, VPS 111, and / or SPS 113 may include a profile_tier_level structure 130. The profile_tier_level structure 130 contains information related to the profile, layer, and level used to encode the codec video. The profile indicates the profile used to encode the codec video. Different profiles have different codec characteristics (e.g., the availability of different codec tools), such as different bit depths, different chroma sampling formats, the availability of cross-component prediction, and the availability of disabling intra-frame smoothing. The layer indicates whether the codec video 119 is encoded according to a higher layer or a main layer, thus being encoded for general application requirements. The level indicates the constraints on the codec video 119, such as the maximum bit rate, maximum picture size, maximum sampling rate, resolution at the highest frame rate, maximum number of slices, maximum number of slices per picture, etc. Therefore, the PTL information in the profile_tier_level structure 130 describes the capabilities that a decoder must possess in order to decode and display the codec video 119.

[0073] The profile_tier_level structure 130 may include a PTL frame only constraint flag (ptl_frame_only_constraint_flag) field 141, a PTL multilayer enabled flag (ptl_multilayer_enabled_flag) 143, a general profile identification code (general_profile_idc) 145, a general layer flag (general_tier_flag) 147, a general level identification code (general_level_idc) 149, a number of sub-layer profiles (num_sub_profiles) 142, and / or a general sub-layer profile identification code for each i-th interoperability indicator (general_sub_profile_idc[i]) 144. The ptl_frame_only_constraint_flag field 141 specifies whether the CVS 117 conveys pictures representing frames (e.g., full screen images) or fields (e.g., partial screen images intended to be combined to fill the screen). In an example, the constraint may require that the ptl_frame_only_constraint_flag field 141 be set to 1, indicating that the codec video 119 includes pictures that are coded as frames. The ptl_multilayer_enabled_flag 143 indicates whether the codec video 119 is coded in multiple layers. In an example, the ptl_multilayer_enabled_flag 143 is set to 0, indicating that the codec video 119 is coded in a single layer.

[0074] general_profile_idc 145, general_tier_flag 147, and general_level_idc 149 indicate the profile, tier, and level, respectively, of the codec video 119. general_sub_profile_idc[i] 144 indicates a value from 0 to i for the interoperability indicator. num_sub_profiles 142 indicates the number of syntax elements included in general_sub_profile_idc[i] 144. In this example, the values ​​included in general_profile_idc 145, general_tier_flag 147, general_level_idc 149, num_sub_profiles 142, and general_sub_profile_idc[i] 144 need to remain constant across CVSs 117 in the same VVC stream. In another example, the values ​​contained in the general_profile_idc 145 , the general_tier_flag 147 , the general_level_idc 149 , the num_sub_profiles 142 , and the general_sub_profile_idc[i] 144 need to remain unchanged in the CMAF track 123 .

[0075] To address the above and other issues, the following methods are disclosed. These items should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these items can be applied alone or in any combination.

[0076] Example 1

[0077] In one example, a rule may specify that a DCI NAL unit should be present in a VVC CMAF track.

[0078] Example 2

[0079] In one example, a rule may specify that a DCI NAL unit should be present in a VVC CMAF track.

[0080] Example 3

[0081] In one example, a rule may specify that all DCI NAL units present in a VVC CMAF track should have the same content.

[0082] Example 4

[0083] In one example, a rule may specify that there should be one and only one DCI NAL unit in a VVC CMAF track.

[0084] Example 5

[0085] In one example, a rule may specify that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit should be present in the CMAF header sample entry.

[0086] Example 6

[0087] In one example, a rule may specify that the value of the dci_num_ptls_minus1 field in the DCI NAL unit in the VVC CMAF track should be equal to 0.

[0088] Example 7

[0089] In one example, a rule may specify that the value of the ptl_frame_only_constraint_flag field in the profile_tier_level() structure in the DCI NAL unit in the VVC CMAF track should be equal to 1.

[0090] Example 8

[0091] In one example, a rule may specify that the value of the ptl_multilayer_enabled_flag field in the profile_tier_level() structure in the DCI NAL unit in the VVC CMAF track should be equal to 0.

[0092] Example 9

[0093] In one example, a rule may specify that there should be one and only one VPS unit in a VVC CMAF track.

[0094] Example 10

[0095] In one example, the rules may specify that when no DCI NAL units are present in a VVC CMAF track and one or more VPSs are present in the VVC CMAF track, one or more of the following constraints apply: The constraints include that for each VPS, the value of the vps_max_layers_minus1 field shall be equal to 0, and for each VPS, the value of the vps_num_ptls_minus1 field shall be equal to 0.

[0096] In an example, the following constraints apply to the profile_tier_level() structure in each VPS: Such constraints include that the value of the ptl_frame_only_constraint_flag field should be equal to 1; and the value of the ptl_multilayer_enabled_flag field should be equal to 0.

[0097] In an example, the value of each of the following fields in the profile_tier_level() structure of the referenced VPS should not change from one codec's video sequence to another across the entire VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i] for each value of i. In an example, a rule may require that the value of each of these fields be the same for all VPSs present in a VVC CMAF track.

[0098] Example 11

[0099] In one example, a rule may specify that the value of the vui_progressive_source_flag field in the vui_payload() structure in the SPS in the VVC CMAF track should be equal to 1.

[0100] Example 12

[0101] In one example, a rule may specify that the value of the vui_interlaced_source_flag field in the vui_payload() structure in the SPS in the VVC CMAF track should be equal to 1.

[0102] Example 13

[0103] In one example, a rule may specify that when no DCI NAL unit is present and no VPS is present in a VVC CMAF track, the value of each of the following fields in the profile_tier_level() structure of the referenced SPS should not change from one codec's video sequence to another across the VVC elementary stream: general_profile_idc; general_tier_flag; general_level_idc; num_sub_profiles; and general_sub_profile_idc[i] for each value of i. In an example, the rule may require that the value of each of these fields be the same for all SPSs present in a VVC CMAF track.

[0104] Example 14

[0105] In one example, a rule may specify that the value of each of the following fields in the general_timing_hrd_parameters() structure (when present) in the referenced VPS or SPS should not change from one codec's video sequence to another across the VVC elementary stream: num_units_in_tick; and time_scale. In an example, a rule may require that the value of each of these fields should be the same for all general_timing_hrd_parameters() structures present in the VPS or SPS in the VVC CMAF track.

[0106] The previously described embodiment will now be described. This embodiment can be applied to CMAF. Most of the relevant sections that have been added or modified relative to the VVC CMAF specification are shown in bold underlined font, and some deleted sections are shown in bold italic font. There may be some other changes that are editorial in nature and therefore not highlighted.

[0107] X.1 VVC Video CMAF Track. A VVC CMAF track should conform to the requirements for a NAL structured video CMAF track. In addition, it should conform to all remaining requirements in this appendix. If a CMAF track conforms to these requirements, it is called a VVC Video CMAF track and may use the brand "cvvc".

[0108] X.2 VVC video track constraints. X.2.1 VVC video CMAF switch set constraints. Each CMAF track in a CMAF switch set shall conform to the VVC video CMAF track as defined in clause X.1. The VVC video CMAF switch set shall conform to the constraints of the NAL structured video CMAF switch set.

[0109] X.2.2 Visual Sample Entry. The syntax and values ​​of the visual sample entry for a VVC video track shall conform to VVCSampleEntry("vvc1") or VVCSampleEntry("vvci") as defined in ISO / IEC 14496-15. ) sample entry.

[0110]

[0111] X.3.2 Video Parameter Set (VPS). Each VVC video media sample in the CMAF track shall refer to an SPS with sps_video_parameter_set_id equal to 0, in which case there is no VPS in the elementary stream, or shall refer to the VPS in the CMAF header sample entry. VPS. The following additional constraints apply: profile_tier_level() structure In the entire VVC basic stream Should not be changed: general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles, general_sub_profile_idc[i].

[0112] X.3.3 Sequence Parameter Set (SPS). Sequence parameter set NAL units appearing within a CMAF VVC track shall conform to the following additional constraints: The following fields shall have the following predetermined values: vui_parameters_present_flag shall be set to 1. In the entire VVC basic stream A video sequence going to another codec should not change: general_profile_idc, general_tier_flag, general_level_idc, num_sub_profiles and general_sub_profile_idc[i] for each value of i.

[0113]

[0114] X.3.5 Image cropping parameters. SPS and PPS cropping parameters conf_win_top_offset and conf_win_left_offset Should be set to 0. SPS and PPS clipping parameters conf_win_bottom_offset and conf_win_right_offset May be set to a value other than 0. If set to a non-zero value, it is expected to be used by CMAF players to remove video spatial samples that are not intended to be displayed.

[0115] Figure 2 is a block diagram illustrating an example video processing system 4000 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0116] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this document. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to produce a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or sent via a communication connection such as represented by component 4006. The bitstream (or codec) representation of the video received at input 4002 or the communication transmission can be used by component 4008 to generate pixel values ​​or transmit to a displayable video of display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0117] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smart phones, or other devices capable of performing digital data processing and / or video display.

[0118] Figure 3 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in this document. Memory(s) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this document in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102 (e.g., a graphics coprocessor).

[0119] Figure 44200 is a flow chart of an example method 4200 of video processing. The method 4200 includes determining information in a DCI NAL unit at step 4202. For example, a rule constrains the use of DCI NAL units in a VVC CMAF track. In the example, the rule specifies that the DCI NAL unit should be present in the VVC CMAF track. In the example, the rule specifies that the DCI NAL unit shall be present in the VVC CMAF track. In the example, the rule specifies that all DCI NAL units present in the VVC CMAF track shall have the same content. In the example, the rule specifies that there is one and only one DCI NAL unit in the VVC CMAF track. In the example, the rule specifies that when the DCI NAL unit is present in the VVC CMAF track, the DCI NAL unit shall be present in the CMAF header sample entry. In the example, the DCI NAL unit contains bitstream level PTL information.

[0120] At step 4204, conversion is performed between the visual media data and the media data file based on the DCI NAL unit. When method 4200 is performed on an encoder, the conversion includes generating a media data file based on the visual media data. The conversion includes determining the DCI NAL unit and encoding it into a bitstream contained in a CMAF track. When method 4200 is performed on a decoder, the conversion includes parsing and decoding the bitstream in the CMAF track based on the DCI NAL unit to obtain the visual media data.

[0121] It should be noted that method 4200 can be implemented in an apparatus for processing video data, the apparatus comprising a processor and non-transitory memory having instructions thereon, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use with a video codec device. The computer program product comprises computer-executable instructions stored on a non-transitory computer-readable medium, such that, when executed by the processor, the video codec device performs method 4200.

[0122] Figure 5 4 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of this disclosure. Video codec system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, which can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, which can be referred to as a video decoding device.

[0123] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include codec pictures and associated data. A codec picture is a codec representation of a picture. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be sent directly to target device 4320 via network 4330 via I / O interface 4316. The encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.

[0124] Target device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320, or may be external to target device 4320 and configured to interface with an external display device.

[0125] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVM) standard, and other current and / or additional standards.

[0126] Figure 6 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 5 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0127] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413 and an entropy coding unit 4414.

[0128] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In an example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is a picture in which the current video block is located.

[0129] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated, but are represented separately in the example of the video encoder 4400 for purposes of explanation.

[0130] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.

[0131] The mode selection unit 4403 can select one of the coding modes (e.g., intra or inter) based on the error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a combination of intra and inter prediction mode (CIIP), where prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 4403 can also select the resolution of the motion vector of the block (e.g., sub-pixel or integer pixel precision).

[0132] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 4413 other than the picture associated with the current video block.

[0133] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0134] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block and can search for a reference picture in list 0 or list 1 for a reference video block of the current video block. Motion estimation unit 4404 can then generate a reference index and a motion vector indicating a reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0135] In other examples, the motion estimation unit 4404 may perform bidirectional prediction on the current video block. The motion estimation unit 4404 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in list 1. The motion estimation unit 4404 may then generate a reference index indicating the reference pictures in list 0 and list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 4404 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0136] In some examples, motion estimation unit 4404 may output a complete set of motion information for use in a decoding process at a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0137] In one example, the motion estimation unit 4404 may indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 4500 that the current video block has the same motion information as another video block.

[0138] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0139] As discussed above, the video encoder 4400 can predictively signal motion vectors.Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0140] Intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0141] The residual generation unit 4407 can generate residual data for the current video block by subtracting the predicted video blocks of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0142] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform a subtraction operation.

[0143] Transform processing unit 4408 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[0144] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0145] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.

[0146] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0147] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives the data, it may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0148] Figure 7 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 5 Video decoder 4324 in system 4300 is shown. Video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0149] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally the reverse of the encoding process described for video encoder 4400.

[0150] The entropy decoding unit 4501 can retrieve a coded bitstream. The coded bitstream can include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 can determine such information, for example, by performing AMVP and Merge modes.

[0151] The motion compensation unit 4502 may generate a motion compensated block and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in a syntax element.

[0152] The motion compensation unit 4502 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters as used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information and use the interpolation filters to generate a prediction block.

[0153] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) slices of the coded video sequence, partitioning information describing how each macroblock of the pictures of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information used to decode the coded video sequence.

[0154] The intra-frame prediction unit 4503 can form a prediction block from spatially adjacent blocks using, for example, an intra-frame prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0155] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 4507 to provide a reference block for subsequent motion compensation / intra-frame prediction and also to generate a decoded video for presentation on a display device.

[0156] Figure 8 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques for VVC. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter IT (ALF) 4606. Unlike the DF 4602, which uses predefined filters, the SAO 4604 and the ALF 4606 utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, using coded side information that signals the offset and filter coefficients. The ALF 4606 is located at the last processing stage for each picture and can be seen as a tool that attempts to capture and repair artifacts created by previous stages.

[0157] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture buffer 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy codec component 4618. The entropy codec component 4618 performs entropy coding and decoding on the prediction results and quantized transform coefficients and sends them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before the images are stored in the reference picture buffer 4612 .

[0158] A list of some example preferred solutions is provided next.

[0159] The following solutions illustrate examples of the techniques discussed herein.

[0160] 1. A method for processing media data (e.g., Figure 4 ), comprising: performing conversion between visual media information and a digital representation of the visual media information according to a rule, wherein the rule specifies whether or how a decoding capability information (DCI) network abstraction layer (NAL) unit is included in a track of a codec elementary stream in the digital representation.

[0161] 2. The method of solution 1, wherein the rule specifies that the DCI NAL unit is included in each track of the codec elementary stream.

[0162] 3. The method according to any of solutions 1-2, wherein the rule specifies that, in case multiple DCI NAL units are included in a track of a codec elementary stream, the multiple DCI NAL units have the same content.

[0163] 4. The method according to any of solutions 1-2, wherein the rule specifies that only one DCI NAL unit is included in a track of a codec elementary stream.

[0164] 5. The method of solution 1, wherein the rule specifies that the DCI NAL unit, when present in a track of a codec elementary stream, is constrained to be in the head sample entry of the track.

[0165] 6. The method of solution 1-5, wherein the rule specifies that the DCI NAL unit conforms to a constraint that the value of a field in the DCI NAL unit is constrained to be equal to a predetermined value.

[0166] 7. The method of solution 6, wherein the field indicates the number of levels, layers, or layer structures minus 1, and wherein the predetermined value is equal to 0.

[0167] 8. The method of solution 6, wherein the field indicates whether multi-layer indication of profile-tier-level is enabled, and wherein the predetermined value is equal to 1.

[0168] 9. A method for media data processing, comprising: performing conversion between visual media information and a digital representation of the visual media information according to a rule, wherein the rule specifies whether or how a video parameter set (VPS) unit is included in a track of a codec elementary stream in the digital representation.

[0169] 10. The method of solution 9, wherein the rule specifies that only one VPS unit is included in a track of a codec elementary stream.

[0170] 11. A method according to any of solutions 9-10, wherein the rule specifies that the digital representation satisfies the constraint when a track of a codec elementary stream includes a VPS unit but does not include a decoding capability information (DCI) network abstraction layer (NAL) unit.

[0171] 12. The method according to any of solutions 9-11, wherein the rule specifies that the VPS complies with a constraint that the value of a field in the VPS is constrained to be equal to a predetermined value.

[0172] 13. A method according to any of solutions 9-12, wherein the rule specifies that the digital representation satisfies the constraint when the track of the codec elementary stream does not include a VPS unit and a decoding capability information (DCI) network abstraction layer (NAL) unit.

[0173] 14. A method of media data processing, comprising: performing conversion between visual media information and a digital representation of the visual media information according to a rule, wherein the rule specifies whether or how values ​​of fields included in a hypothetical reference decoder structure referenced by a video parameter set of a sequence parameter set are allowed to change from one coded video sequence to a second coded video sequence in a coded video elementary stream in the digital representation.

[0174] 15. The method of solution 14, wherein the value indicates a time scale.

[0175] 16. A method according to any of solutions 14-15, wherein the value of the rule specifies that the field is the same in each hypothetical reference decoder structure in the digital representation.

[0176] 17. A method of media data processing, comprising: obtaining a digital representation of visual media information, wherein the digital representation is generated according to the method of any one of solutions 1-16; and streaming the digital representation.

[0177] 18. A method of media data processing, comprising: receiving a digital representation of visual media information, wherein the digital representation is generated according to the method of any one of solutions 1-16; and generating visual media information from the digital representation.

[0178] 19. The method according to any of solutions 1-18, wherein the converting comprises generating a bitstream representation of the visual media data and storing the bitstream representation to a file according to format rules.

[0179] 20. The method according to any of solutions 1-18, wherein the conversion comprises parsing the file according to format rules to recover the visual media data.

[0180] 21. A video decoding device comprising a processor configured to implement the method according to one or more of solutions 1 to 20.

[0181] 22. A video encoding device comprising a processor configured to implement the method according to one or more of solutions 1 to 20.

[0182] 23. A computer program product storing computer code, which, when executed by a processor, causes the processor to implement the method according to any one of solutions 1 to 20.

[0183] 24. A computer readable medium having thereon a bitstream representation conforming to a file format generated according to any one of solutions 1 to 20.

[0184] 25. A method, apparatus, or system as described herein. In the solutions described herein, an encoder can conform to a format rule by generating a codec representation according to the format rule. In the solutions described herein, a decoder can use the format rule to parse syntax elements in the codec representation according to the format rule, knowing the presence or absence of syntax elements, to produce decoded video.

[0185] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits that are collocated or dispersed in different places within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on an error residual value after transformation and encoding and also using bits from headers and other fields in the bitstream. Furthermore, during conversion, the decoder may, based on this determination, parse the bitstream knowing that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include or not include certain syntax fields and generate the codec representation accordingly by including the syntax fields or excluding the syntax fields from the codec representation.

[0186] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, which are used to be executed by a data processing device or to control the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that affect a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing device" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, a device may also include code that creates an operating environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical signal, an optical signal, or an electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0187] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to run on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0188] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0189] Processors suitable for running computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or be operatively coupled to receive data from or transfer data to or from the one or more mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0190] Although this patent document contains many details, these details should not be interpreted as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features specified for particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, although features may be described above as working in certain combinations and even initially claimed as such, one or more features from the claimed combination may be excluded from the combination in some cases, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0191] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0192] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

[0193] A first component is directly coupled to a second component when there are no intervening components between the first and second components other than a wire, trace, or another medium. A first component is indirectly coupled to a second component when there are intervening components between the first and second components other than a wire, trace, or another medium. The term "coupled" and its variations encompass both direct and indirect couplings. Unless otherwise specified, the use of the term "approximately" is intended to encompass a range of ±10% of the subsequent figure.

[0194] Although several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered illustrative rather than restrictive, and the present invention is not intended to be limited to the details given herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.

[0195] In addition, without departing from the scope of the present disclosure, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods. Other items shown or discussed as coupled may be directly connected, or may be indirectly coupled or communicated through some interface, device, or intermediate component in an electrical, mechanical, or other manner. Other examples of changes, substitutions, and modifications may be determined by those skilled in the art and may be made without departing from the spirit and scope of the present disclosure.

Claims

1. A method for processing video data, comprising: determining information in a decoding capability information (DCI) network abstraction layer (NAL) unit based on rules, wherein the rules constrain use of the DCI NAL unit in a Versatile Video Codec (VVC) Common Media Application Format (CMAF) track; and performing conversion between visual media data and media data files based on information in the DCI NAL unit, The rule specifies that in the profile_tier_level() structure in the DCI NAL unit, the value of the PTL frame only constraint flag (ptl_frame_only_constraint_flag) should be restricted to 1; wherein the rule specifies that the DCI NAL unit should be present in a VVC CMAF track; The rule specifies that all DCI NAL units present in the VVC CMAF track should have the same content.

2. The method according to claim 1, wherein The rules specify that there is one and only one DCI NAL unit in a VVC CMAF track.

3. The method according to claim 1, wherein The rule specifies that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit should be present in the CMAF header sample entry.

4. The method according to claim 1, wherein The DCI NAL unit contains bitstream level profile, layer and level (PTL) information.

5. The method according to any one of claims 1 to 4, wherein The converting includes encoding the visual media data into a media data file.

6. The method according to any one of claims 1 to 4, wherein The converting includes decoding the visual media data from the media data file.

7. An apparatus for processing video data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determining information in a decoding capability information (DCI) network abstraction layer (NAL) unit based on rules, wherein the rules constrain use of the DCI NAL unit in a Versatile Video Codec (VVC) Common Media Application Format (CMAF) track; and performing conversion between visual media data and media data files based on information in the DCI NAL unit, The rule specifies that in the profile_tier_level() structure in the DCI NAL unit, the value of the PTL frame only constraint flag (ptl_frame_only_constraint_flag) should be restricted to 1; wherein the rule specifies that the DCI NAL unit should be present in a VVC CMAF track; The rule specifies that all DCI NAL units present in the VVC CMAF track should have the same content.

8. The device according to claim 7, wherein The rules specify that there is one and only one DCI NAL unit in a VVC CMAF track.

9. The device according to claim 7, wherein The rule specifies that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit should be present in the CMAF header sample entry.

10. The device according to claim 7, wherein The DCI NAL unit contains bitstream level profile, layer and level (PTL) information.

11. A non-transitory computer-readable medium comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device: determining information in a decoding capability information (DCI) network abstraction layer (NAL) unit based on rules, wherein the rules constrain use of the DCI NAL unit in a Versatile Video Codec (VVC) Common Media Application Format (CMAF) track; and performing conversion between visual media data and media data files based on information in the DCI NAL unit, in, The rule specifies that in the profile_tier_level() structure in the DCI NAL unit, the value of the PTL frame only constraint flag (ptl_frame_only_constraint_flag) should be restricted to 1; wherein the rule specifies that the DCI NAL unit should be present in a VVC CMAF track; The rule specifies that all DCI NAL units present in the VVC CMAF track should have the same content.

12. The non-transitory computer-readable medium of claim 11, wherein: The rules specify that there is one and only one DCI NAL unit in a VVC CMAF track.

13. The non-transitory computer-readable medium of claim 11, wherein: The rule specifies that when a DCI NAL unit is present in a VVC CMAF track, the DCI NAL unit should be present in the CMAF header sample entry.

14. The non-transitory computer-readable medium of claim 11, wherein: The DCI NAL unit contains bitstream level profile, layer and level (PTL) information.

Citation Information

Patent Citations

  • Systems and methods for signaling scalable video in media application format

    CN110506421A