Improved Extended Dependent Random Access Point Support in ISO Base Media File Format

By specifying that EDRAP samples can decode all subsequent samples with a preceding SAP or EDRAP sample available, the solution addresses the inefficiency in EDRAP-based video coding, storage, and streaming by clearly defining required samples, enhancing decoding and streaming efficiency.

JP2025515738APending Publication Date: 2025-05-20BYTEDANCE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024566345
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-10
Filing Date
2023-05-09
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The existing EDRAP specification lacks a clear definition of the set of consecutive samples required for random access from an EDRAP picture, leading to potential inefficiencies in video coding, storage, and streaming processes.

Method used

Specify that an EDRAP sample is a sample that can correctly decode all subsequent samples in both decoding and output order if a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference, and ensure that associated tracks contain only one sample with the same decoding time as the EDRAP sample, including any necessary preceding SAP or EDRAP samples.

Benefits of technology

This solution enables efficient decoding and streaming by clearly defining the samples required for random access from EDRAP pictures, improving coding efficiency and reducing the need for additional samples in associated tracks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515738000001_ABST
    Figure 2025515738000001_ABST
Patent Text Reader

Abstract

A mechanism for processing visual media data is disclosed. An Extended Dependent Random Access Point (EDRAP) sample is determined. The EDRAP sample is a sample for which all subsequent samples can be correctly decoded, both in decoding order and output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples. Based on the EDRAP sample, a conversion is performed between the visual media data and the media data file.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 340,167, filed May 10, 2022. All of the foregoing patent applications are incorporated herein by reference in their entireties.

[0002] This patent document relates to the creation, storage and consumption of digital audiovisual media information in file formats. [Background technology]

[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communications networks, and the bandwidth demands for digital video use will continue to grow as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention

[0004] A first aspect relates to a method for processing visual media data, the method including determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample that allows all subsequent samples to be correctly decoded, both in decoding order and output order, provided that a required preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and the subsequent samples, and performing a conversion between the visual media data and a media data file based on the EDRAP sample.

[0005] A second aspect relates to a method for processing visual media data, the method including: determining Extended Dependent Random Access Point (EDRAP) samples; if a media track has a track reference of type “aest” referencing an associated track, then for each EDRAP sample, denoted as sampleA in the media track, there is exactly one sample, denoted as sampleB, in the associated track that has the same decoding time as sampleA; and performing conversion between the visual media data and a media data file based on the EDRAP samples.

[0006] A third aspect is an apparatus for processing visual media data, comprising one or more processors and one or more non-transitory memories having instructions stored thereon that, when executed by the processor, cause the processor to perform any of the aspects discussed above.

[0007] A fourth aspect is a non-transitory computer-readable medium comprising a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by one or more processors of the video coding device, causes the video coding device to perform a method of any of the preceding aspects.

[0008] A fifth aspect is a non-transitory computer-readable recording medium storing a media data file generated by a method executed by a media processing device, the method including: determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample that can correctly decode all subsequent samples in both decoding order and output order when a required preceding Streaming Access Point (SAP) or EDRAP sample is referenceable when decoding the EDRAP sample and subsequent samples, and generating the media data file based on determining the media data file.

[0009] For clarity, any one of the above-described embodiments may be combined with any other one of the above-described embodiments to create a new embodiment that is within the scope of the present disclosure.

[0010] These and other features will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings and claims.

[0011] For a more complete understanding of this disclosure, reference should be made to the following brief description taken in conjunction with the accompanying drawings and detailed description, where like reference numbers represent like parts. [Brief description of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram illustrating an example mechanism for random access when decoding a bitstream with Intra Random Access Point (IRAP) pictures. [Diagram 2] FIG. 2 is a schematic diagram illustrating an example mechanism for random access when decoding a bitstream using Dependent Random Access Point (DRAP) pictures. [Diagram 3] FIG. 3 is a schematic diagram illustrating an exemplary mechanism for random access when decoding a bitstream using EDRAP pictures. [Figure 4] FIG. 4 is a schematic diagram of an example mechanism for signaling an external bitstream to support EDRAP-based random access. [Diagram 5] FIG. 5 is a diagram illustrating an example of EDRAP-based random access. [Figure 6] FIG. 6 is a block diagram illustrating an example of a video processing system. [Figure 7] FIG. 7 is a block diagram illustrating an example of a video processing device. [Figure 8] FIG. 8 is a flow chart illustrating an example of a video processing method. [Figure 9] FIG. 9 is a block diagram illustrating an example of a video coding system. [Figure 10] FIG. 10 is a block diagram illustrating an example of an encoder. [Figure 11] FIG. 11 is a block diagram illustrating an example of a decoder. [Figure 12] FIG. 12 is a circuit diagram of an example of an encoder. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0013] Although exemplary implementations of one or more embodiments are provided below, it should be understood at the outset that the disclosed systems and / or methods may be implemented using any number of technologies, whether now known or later developed. The present disclosure should in no way be limited to the exemplary implementations, drawings, and technologies described herein, but includes the exemplary designs and designs illustrated and described herein, which may be modified within the scope of the appended claims, along with their full scope of equivalents.

[0014] Section headings are used in this document for ease of understanding and do not imply a limitation that the techniques and embodiments disclosed in each section apply only to that section. In addition, some descriptions use H.266 terminology for ease of understanding and do not limit the scope of the disclosed techniques. Thus, the techniques described herein are applicable to other video codec protocols and designs. In this document, editorial changes to the VVC specification or ISOBMFF file format specification drafts are indicated with bold italics to indicate canceled text and bold underline to indicate added text.

[0015] 1. Initial discussion This document relates to media file formats, and in particular to support for Extended Dependent Random Access Point (EDRAP) signaling in the International Organization for Standardization (ISO) based media file format (ISOBMFF). These ideas may be applied, individually or in various combinations, to media files conforming to any media file format, such as ISOBMFF and file formats derived from ISOBMFF.

[0016] 2. Introduction to video coding 2.1 Video Coding Standards Video coding standards have evolved primarily through standards development by the International Telecommunication Union (ITU) Telecommunication Standardization Sector (ITU-T) and the ISO / International Electrotechnical Commission (IEC). ITU-T developed H.261 and H.263, while ISO / IEC developed Motion picture Experts Group (MPEG)-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / High Efficiency Video Coding (HEVC)[1] standards. Video coding standards after H.262 are based on a hybrid video coding structure that utilizes temporal prediction + transform coding. To explore future video coding techniques beyond HEVC, the Video Coding Experts Group (VCEG) and MPEG jointly established the Joint Video Exploration Team (JVET). Many methods have been adopted by the JVET and summarized in a reference software called the Joint Exploration Model (JEM)[2]. JVET was renamed the Joint Video Experts Team (JVET) when the Generic Video Coding (VVC) project was officially launched. VVC [3] is a coding standard that aims to achieve a 50% bitrate reduction compared to HEVC.

[0017] The Generic Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3)[3][4] and the related Generic Supplementary Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7)[5][6] are designed for use in the widest range of applications, including both traditional uses such as television broadcast, videoconferencing, or playback from storage media, as well as newer and more advanced use cases such as adaptive bitrate streaming, video region extraction, composition and merging of content from multiple coded video bitstreams, multiview video, scalable layered coding, and viewport-adaptive 360° immersive media.

[0018] 2.2 File Format Standards Media streaming applications are based on IP, TCP, and HTTP transport methods and rely on file formats such as the ISO Base Media File Format (ISOBMFF) [7]. One such streaming system is dynamic adaptive streaming over HTTP (DASH) [8]. When using video formats over ISOBMFF and DASH, video format-specific file format specifications are required, such as the AVC and HEVC file formats in [9], for the encapsulation of video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream, such as profile, tier, level, and many other information, needs to be exposed as file format level metadata and / or DASH Media Presentation Description (MPD) for content selection purposes, e.g., selection of appropriate media segments for initialization at the beginning of a streaming session and stream adaptation during a streaming session.

[0019] Similarly, the use of a picture format in ISOBMFF requires a file format specification specific to the picture format, such as the AVC picture file format and the HEVC picture file format in

[10] .

[0020] 2.3 Random Access and its Support in HEVC and VVC Random access refers to accessing and decoding a bitstream starting from a picture that is not the first picture of the bitstream in decoding order. To support tuning and channel switching in broadcast, multicast, and multi-party video conferencing, local playback and searching in streaming, stream adaptation in streaming, etc., a bitstream needs to contain frequent random access points. Such random access points may be intra-coded pictures, but also inter-coded pictures, e.g., in the case of gradual decoding updates.

[0021] HEVC includes signaling of Intra Random Access Point (IRAP) pictures in the NAL unit header through NAL unit types. Three types of IRAP pictures are supported in HEVC. These are Instantaneous Decoding Refresh (IDR), Clean Random Access (CRA), and Broken Link Access (BLA) pictures. IDR pictures constrain the inter-picture prediction structure to not reference pictures prior to the current group of pictures (GOP). Reference pictures in the current GOP are sometimes referred to as closed GOP random access points. CRA pictures are less restricted by allowing a particular picture to reference pictures prior to the current GOP, and in the case of random access, all pictures prior to the current GOP are discarded. CRA pictures are sometimes referred to as open GOP random access points. BLA pictures usually result from the splicing of two bitstreams or parts of them at CRA pictures, for example, when switching streams. To enable better system utilization of IRAP pictures, six different NAL units are defined to signal the characteristics of IRAP pictures. Such properties are used to support random access in Dynamic Adaptive Streaming over Hypertext Transfer Protocol (DASH) [8] and to better match the stream access point types defined in ISOBMFF [7].

[0022] VVC supports three types of IRAP pictures, two types of IDR pictures (one with and one without an associated Random Access Decodable Leading (RADL) picture), and one type of CRA picture. In HEVC, they are used in a similar manner. The HEVC BLA picture type is not included in VVC for two reasons. First, the basic functionality of a BLA picture can be realized by a CRA picture and an end of sequence NAL unit, the presence of which indicates that the following picture starts a new CVS in a single layer bitstream. Second, there is a desire during the development of VVC to specify fewer NAL unit types than HEVC, as indicated by the use of 5 bits instead of 6 bits for the NAL unit type field in the NAL unit header.

[0023] Another difference in random access support between VVC and HEVC is that GDR is supported in a more normative way in VVC. In GDR, bitstream decoding can start from an inter-coded picture. At the beginning of the access, the entire picture area cannot be correctly decoded. However, after several pictures, the entire picture area is correctly decoded. AVC and HEVC also support GDR by using the recovery point supplemental enhancement information (SEI) message to signal GDR random access points and recovery points. In VVC, a NAL unit type is specified to indicate a GDR picture, and recovery points are signaled in the picture header syntax structure. A coded video sequence (CVS) and bitstream can start using a GDR picture. This means that the entire bitstream can contain only inter-coded pictures, not one intra-coded picture. The main advantage of specifying GDR support in this way is to provide compliant operation of GDR. GDR allows encoders to smooth the bitstream bitrate by distributing intra-coded slices or blocks across multiple pictures, rather than intra-coding the entire picture. This allows for a significant reduction in end-to-end delay, which may become even more important in many cases as ultra-low latency applications like wireless display, online gaming, and drone-based applications become more popular.

[0024] Another GDR-related feature in VVC is virtual boundary signaling. In pictures between the GDR picture and the corresponding recovery point, the boundary between the updated and non-updated regions that are correctly decoded regions can be signaled as a virtual boundary. When signaled, no in-loop filtering is applied across the boundary. Thus, no decoding inconsistency occurs for some samples at or near the boundary. This is useful when an application decides to display correctly decoded regions during GDR processing. IRAP and GDR pictures can be collectively referred to as Random Access Point (RAP) pictures.

[0025] 2.4 Enhanced Dependent Random Access Point (EDRAP) based video coding, storage and streaming 2.4.1 Concepts and Standards Support Here, we will describe the concept of EDRAP-based video coding, storage and streaming. As shown in Figure 1, an application (e.g., adaptive streaming) determines the frequency of random access points (RAPs), e.g., RAP period 1s or 2s. In one example, the RAPs are provided by coding IRAP pictures. Note that inter-prediction references of non-key pictures between RAP pictures are not shown, and the output order is from left to right. In case of random access from CRA4, the decoder receives CRA4, CRA5, etc. and the associated inter-prediction pictures and decodes them correctly.

[0026] Figure 2 shows the DRAP approach, which provides improved coding efficiency by allowing DRAP pictures (and subsequent pictures) to refer to previous IRAP pictures for inter prediction. Note that inter prediction of non-key pictures between IRAP pictures is not shown, output order is from left to right. In case of random access from DRAP4, the decoder receives and correctly decodes IDR0, DRAP4, DRAP5, etc. and the associated inter predicted pictures.

[0027] Figure 3 shows the EDRAP approach, which offers a bit more flexibility by allowing EDRAP pictures (and subsequent pictures) to reference some of the previous RAP pictures (IRAP or EDRAP). Note that inter-prediction of non-key pictures between RAP pictures is not shown, output order is from left to right. In case of random access from EDRAP4, the decoder receives and correctly decodes IDR0, EDRAP2, EDRAP4, EDRAP5, etc. and the associated inter-predicted pictures.

[0028] Figure 4 shows an example of the EDRAP approach with MSR and ESR segments. Figure 5 shows an example of random access from EDRAP4. When random accessing from a segment starting from EDRAP4 or switching segments, the decoder receives and decodes segments including IDR0, EDRAP2, EDRAP4, EDRAP5, etc. and associated inter-predicted pictures.

[0029] EDRAP-based video coding is supported by the EDRAP Indication SEI message contained in the VSEI standard amendment

[11] , the storage part is supported by the EDRAP Sample Group and associated External Stream Track References contained in the ISOBMFF standard amendment

[12] , and the streaming part is supported by the Main Stream Representation (MSR) and External Stream Representation (ESR) Descriptors contained in the DASH standard amendment

[13] . The support for these standards is described below.

[0030] 2.4.2 EDRAP Indication SEI Message An amendment to the VSEI standard is currently under development. A draft example specification for this amendment is included in

[11] , including the specification of the EDRAP indication SEI message.

[0031] The syntax and semantics of the EDRAP indication SEI message are as follows:

[0032] [Table 1]

[0033] A picture associated with an Extended DRAP (EDRAP) indication SEI message is referred to as an EDRAP picture.

[0034] The presence of an EDRAP indication SEI message indicates that the constraints on picture order and picture referencing specified in this section apply. These constraints allow a decoder to properly decode an EDRAP picture and pictures that follow it in the same layer, in both decoding order and output order, without the need to decode any other pictures in the same layer, except for the list of referenceable Pictures, which contains a list of IRAP or EDRAP pictures in the same CLVS and in decoding order identified by the edrap_ref_rap_id[i] syntax element.

[0035] The constraints, indicated by the presence of the EDRAP indication SEI message, which all apply, are as follows: - The EDRAP picture is a subsequent picture. - An EDRAP picture has a temporal sub-layer identifier equal to 0. - An EDRAP picture does not include pictures of the same layer in the active entries of its reference picture list, except for referenceablePictures. - Pictures that are in the same layer and follow the EDRAP picture in both decoding order and output order shall not include in the active entries of the referenceable picture list any pictures that are in the same layer and precede the EDRAP picture in decoding order or output order, except for referenceablePictures. - The pictures in the referenceablePictures list do not include in the active entry of the reference picture list a picture that is not a picture that is in the same layer and is in a previous position in the referenceablePictures list. NOTE - As a result, the first picture in referenceablePictures does not contain a picture from the same layer in the active entry of the reference picture list, even if it is an EDRAP picture instead of an IRAP picture.

[0036] The value of edrap_rap_id_minus1 plus 1 specifies the RAP picture identifier of the EDRAP picture, denoted as RapPicId.

[0037] Each IRAP or EDRAP picture has an associated RapPicId value. The RapPicId value of an IRAP picture is presumed to be equal to 0. Two EDRAP pictures associated with the same IRAP picture have different RapPicId values.

[0038] Specifies that if edrap_leading_pictures_decodable_flag is equal to 1, then the following constraints apply: A picture that is in the same layer and that follows an EDRAP picture in decoding order must follow, in output order, a picture that is in the same layer and that precedes the EDRAP picture in decoding order. - Pictures that are in the same layer, follow the EDRAP picture in decoding order, and precede the EDRAP picture in output order must not include pictures in the active entries of the reference picture list that are in the same layer and precede the EDRAP picture in decoding order, except for referenceablePictures.

[0039] edrap_leading_pictures_decodable_flag equal to 0 imposes no such constraint.

[0040] edrap_reserved_zero_12bits MUST be equal to 0 in bitstreams conforming to this version of this specification. Other values ​​of edrap_reserved_zero_12bits are reserved for use by ITU-T|ISO / IEC. Decoders MUST ignore values ​​of edrap_reserved_zero_12bits.

[0041] The value of edrap_num_ref_rap_pics_minus1 plus 1 indicates the number of IRAP pictures or EDRAP pictures that are in the same CLVS as the EDRAP picture and may be included in the active entries of the reference picture list of the EDRAP picture.

[0042] edrap_ref_rap_id[i] indicates the RapPicId of the i-th RAP picture that may be included in the active entries of the reference picture list of the EDRAP picture. The i-th RAP picture must be either an IRAP picture associated with the current EDRAP picture or an EDRAP picture associated with an IRAP picture similar to the current EDRAP picture.

[0043] 2.4.3 EDRAP Sample Groups and Associated External Stream Track References An amendment to the ISOBMFF standard is under development. The draft amendment includes the specification of the EDRAP sample group and associated external stream track references.

[0044] The specifications of these two ISOBMFF functions are as follows:

[0045] 3.1 Definition ... 3.2 Abbreviations EDRAP extended dependent random access point

[0046] 8.3.3.4 Related external stream track references Track references of type 'aest' (meaning 'associated external stream track') may be contained in video tracks and refer to associated video tracks. If present, a TrackReferenceTypeBox with reference_type equal to 'aest' must only contain a track identifier and must not contain a track group identifier.

[0047] If a video track has a track reference of type 'aest' the following applies: -A video track must have at least one sample containing an EDRAP picture. - for each sample A in a video track containing an EDRAP picture, there is exactly one sample B in the associated video track that has the same decoding time as sample A, and the set of consecutive samples in the associated video track starting from sample B exclusively includes all pictures that are not included in the video track containing sample A and that are required for random access from the EDRAP picture contained in sample A.

[0048] All samples in the referenced track must be identified as sync samples. The referenced track header flags must have track_in_movie and track_in_preview both set to 0.

[0049] Each track referenced must use the following restricted scheme: 1) At least one sample entry type of each sample entry in the track must be equal to "resv". NOTE 1: "resv" does not necessarily have to be the sample entry type of the SampleEntry directly contained in the SampleDescriptionBox if the track has undergone some transformation. 2) The unconverted sample entry type is stored in an OriginalFormatBox contained in a RestrictedSchemeInfoBox. 3) The scheme_type field of the SchemeTypeBox contained in the RestrictedSchemeInfoBox is equal to "aest", indicating that the samples in the track may contain one or more coded pictures. 4) Bit 0 of the flags field of the SchemeTypeBox is equal to 0 and the value of (flags&0x000001) is equal to 0.

[0050] 10.11 Extended DRAP (EDRAP) Sample Group 10.11.1 Definition This sample group is similar to the DRAP sample group defined in Section 10.8, but allows more flexible inter-RAP referencing.

[0051] An EDRAP sample is a sample from which all samples after it in decoding and output order can be correctly decoded if the closest SAP sample of type 1, 2, or 3 that precedes the EDRAP sample and zero or more other identified EDRAP samples that are earlier in decoding order than the EDRAP sample are available for reference.

[0052] 10.11.2 Syntax class VisualEdrapEntry() extends VisualSampleGroupEntry('edrp'){ unsigned int(3) edrap_type; unsigned int(3) num_ref_sap_or_edrap_samples_minus1; unsigned int(26) reserved=0; for(i=0;i<=num_ref_sap_or_edrap_samples_minus1;i++){ unsigned int(16) ref_sap_or_edrap_idx_delta[i]; } }

[0053] 10.11.3 Semantics edrap_type is a non-negative integer. If edrap_type is in the range 1 to 3, it indicates the SAP_type (as specified in Annex I) that the EDRAP sample would have corresponded to if it had not depended on the closest preceding SAP sample or on other EDRAP samples. Other type values ​​are reserved.

[0054] num_ref_sap_or_edrap_samples_minus1 plus 1 indicates the number of preceding SAP or EDRAP samples that are earlier in the decoding order than the EDRAP sample and that are required as a reference to correctly decode the EDRAP sample and all samples that follow it in both decoding order and output order when decoding begins with the EDRAP sample. Note that for an EDRAP sample that is also a DRAP sample, the value of num_ref_sap_or_edrap_samples_minus1 is equal to 0.

[0055] The reserved value MUST be equal to 0. The semantics of this section apply only to sample group description entries with the reserved value equal to 0. A parser MAY tolerate and ignore sample group description entries with the reserved value greater than 0 when parsing this sample group.

[0056] ref_sap_or_edrap_idx_delta[i] indicates the i-th required preceding SAP or EDRAP sample of the current EDRAP sample. The list of SAP or EDRAP samples associated with a SAP sample of type 1, 2 or 3 is the SAP sample and all EDRAP samples that follow the SAP sample and, if any, precede the next SAP sample. The SAP_or_EDRAP sample index is defined as the index of this list of SAP or EDRAP samples. The value of ref_sap_or_edrap_idx_delta[i] is equal to the difference between the SAP_or_EDRAP sample index of the current EDRAP sample and the SAP_or_EDRAP sample index of the i-th required preceding SAP or EDRAP sample. A value of 1 indicates that the i-th required SAP or EDRAP sample is the last SAP or EDRAP sample preceding this EDRAP sample in decoding order, a value of 2 indicates that the i-th required SAP or EDRAP sample is the second-last EDRAP sample preceding this EDRAP sample in decoding order, and so on.

[0057] 3. The technical problem solved by the disclosed solution There is a problem related to the design of the storage part of EDRAP-based video coding, storage and streaming. The EDRAP specification specifies that for each sample denoted as sampleA in a video track containing an EDRAP picture, there is exactly one sample denoted as sampleB in the associated video track that has the same decoding time as sampleA. Furthermore, the set of consecutive samples in the associated video track starting from sampleB shall exclusively include all pictures not included in the video track containing sampleA that are required for random access from the EDRAP picture in sampleA. However, no set of consecutive samples that satisfies this condition is specified. As a result, to randomly access the video track from the EDRAP picture, the file parser may have to provide sampleB and all subsequent samples in the associated video track to the file player.

[0058] 4. List of solutions and implementations. In order to solve the above-mentioned problems, the methods summarized below are disclosed. The present invention should be considered as an example to illustrate the general concept and should not be interpreted narrowly. Moreover, these inventions can be applied individually or combined in any way.

[0059] Example 1 In one example, this specification can specify that an EDRAP sample is a sample for which all subsequent samples, in both decoding order and output order, can be correctly decoded, provided that a necessary preceding streaming access point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples. In one example, the necessary preceding SAP or EDRAP sample consists of one or more of a set of samples starting with the closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order and including all EDRAP samples between the closestSapSample and the sample in decoding order.

[0060] Example 2 In one example, this specification may specify that when a video track has a track reference of type "aest" referencing an associated track, for each EDRAP sample sampleA in the video track in the associated track, there should be only one sample sampleB with the same decoding time as sampleA, and sampleB should contain all pictures contained in closestSapSample of sampleA and any necessary preceding SAP or EDRAP samples of sampleA. If a TrackReferenceTypeBox with reference_type equal to "aest" is present, it should contain only a track identifier and no track group identifiers.

[0061] 5. Embodiment Below are examples of some embodiments of the disclosed subject matter summarized in Section 4 above. The most relevant portions that have been added or changed are underlined and bold, and some portions that have been deleted are italicized and bold. Due to the nature of the editorial, there may be other changes that are not highlighted.

[0062] First embodiment This embodiment is for item 1, item 2 and all their subitems.

[0063] 3.1 Definition ...

[0064] [ka]

[0065] ...

[0066] 3.2 Abbreviations ... EDRAP extended dependent random access point ...

[0067] 8.3.3.4 Related external stream track references Track references of type 'aest' (meaning 'associated external stream track') may be contained in a video track. If present, a TrackReferenceTypeBox with reference_type equal to 'aest' must only contain a track identifier and not a track group identifier. If a video track has a track reference of type 'aest' referencing an associated track, the following applies:

[0068] A video track should have at least one EDRAP sample indicated by an EDRAP sample group.

[0069] [ka]

[0070] Each sample in the associated track must be identified as a sync sample. The associated track must have both header flags track_in_movie and track_in_preview equal to 0.

[0071] The associated track MUST use the following restricted scheme: At least one sample entry type of each sample entry of the track MUST be equal to "resv". NOTE 1: "resv" does not have to be the sample entry type of the SampleEntry directly contained in the SampleDescriptionBox if the track has undergone some transformations. The untransformed sample entry type is stored in the OriginalFormatBox contained in the RestrictedSchemeInfoBox. The scheme_type field of the SchemeTypeBox contained in the RestrictedSchemeInfoBox is equal to "aest", indicating that the samples in the track may contain one or more coded pictures. Bit 0 of the flags field of the SchemeTypeBox is equal to 0 and the value of (flags&0x000001) is equal to 0.

[0072] 10.11 Extended DRAP (EDRAP) Sample Group 10.11.1 Definition This sample group is similar to the DRAP sample group specified in subclause 10.8, but allows enabling more flexible inter-prediction references for pictures of EDRAP samples and pictures of subsequent samples, which improves coding efficiency for these pictures. NOTE 1: Similarly for DRAP samples, EDRAP samples can only be used in combination with SAP samples of types 1, 2 and 3. NOTE 2: DRAP samples are always EDRAP samples.

[0073] 10.11.2 Syntax class VisualEdrapEntry() extends VisualSampleGroupEntry('edrp'){ unsigned int(3) edrap_type; unsigned int(3) num_ref_edrap_pics; unsigned int(26) reserved=0; for(i=0;i <num_ref_edrap_pics;i++) unsigned int(16) ref_edrap_idx_delta[i]; }

[0074] 10.11.3 Semantics edrap_type is a non-negative integer. If edrap_type is in the range 1 to 3, it indicates the SAP_type (as specified in Annex I) that the EDRAP sample would have corresponded to if it had not depended on the closest preceding SAP sample or on other EDRAP samples. Other type values ​​are reserved.

[0075] num_ref_edrap_pics indicates the number of other EDRAP samples that are earlier in decoding order than the EDRAP sample and that must be referenced in order to correctly decode the EDRAP sample and all samples that follow it in both decoding order and output order if decoding were to start with the EDRAP sample. The reserved value MUST be equal to 0. The semantics of this section apply only to sample group description entries with reserved value equal to 0. Parsers tolerate and ignore sample group description entries with reserved value greater than 0 when parsing this sample group.

[0076] ref_edrap_idx_delta[i] indicates the difference between the EDRAP sample index of this EDRAP sample (i.e., the index into the list of decoding order of all EDRAP samples contained in this sample group) and the EDRAP sample index of the i-th EDRAP sample that is earlier in decoding order than the EDRAP sample and that needs to be referenced in order to be able to correctly decode the EDRAP sample and all samples that follow it in both decoding order and output order when decoding starts from this EDRAP sample. A value of 1 indicates that the i-th EDRAP sample is the last EDRAP sample in the sample group and precedes this EDRAP sample in decoding order, a value of 2 indicates that the i-th EDRAP sample is the second-last EDRAP sample in the sample group and precedes this EDRAP sample in decoding order, and so on.

[0077] 7.References [1] ITU-T and ISO / IEC, High efficiency video coding”, Rec. ITU-T H.265|ISO / IEC 23008-2 (current version). [2] J.Chen, E.Alshina, GJSullivan, J.-R.Ohm, J.Boyce, “Algorithm description of Joint Exploration Test Model 7(JEM7),” JVET-G1001, Aug.2017. [3] Rec.ITU-T H.266|ISO / IEC 23090-3, “Versatile Video Coding”, 2020. [4] BBBross, J. Chen, S. Liu, Y.-K. Wang (editors), “Versatile Video Coding (Draft 10),” JVET-S2001. [5] Rec.ITU-T Rec.H.274|ISO / IEC 23002-7, “Versatile Supplemental Enhancement Information Messages for Coded Video Bitstream”, 2020. [6] J.Boyce, V.Drugeon, GJSullivan, Y.-K.Wang(editors), “Versatile supplemental enhancement information messages for coded video bitstreams(Draft 5),” JVET-S2007. [7] ISO / IEC 14496-12: "Information technology - Coding of audiovisual objects - Part 12: ISO base media file format". [8] ISO / IEC 23009-1: "Information technology - Dynamic Adaptive Streaming over HTTP (DASH) - Part 1: Media presentation description and segment formats". [9] ISO / IEC 14496-15: "Information technology - Coding of audiovisual objects - Part 15: Transport of structured video units (Network Abstraction Layer (NAL)) in the ISO Base Media File Format".

[10] ISO / IEC 23008-12: "Information technology - Efficiency coding and delivery of media in heterogeneous environments - Part 12: Picture file formats".

[11] J.Boyce, GJSullivan, Y.-K.Wang(editors), “Additional SEI messages for VSEI(Draft 6)”, JVET-Y2006.

[12] ISO / IEC JTC 1 / SC 29 / WG 03 Output Document N0471, "Text of CDAM ISO / IEC 14496-12:2021 AMD 1 Improved brand documentation and other improvements", January 2022.

[13] ISO / IEC JTC 1 / SC 29 / WG 03 Output Document N0486, "WD EDRAP Streaming and Other Extensions to ISO / IEC 23009-1 5th Edition AMD2", January 2022.

[0078] 6 is a block diagram illustrating an example of a video processing system 4000 in which various techniques disclosed in this disclosure can be implemented. Various embodiments can include some or all of the components of the system 4000. The system 4000 can include an input 4002 for receiving video content. The video content can be received in raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or in compressed or encoded format. The input 4002 can represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0079] The system 4000 may include a coding component 4004 that may implement various coding or encoding methods disclosed in this disclosure. The coding component 4004 may reduce the average bit rate of the video from the input 4002 to the output of the coding component 4004 to generate a coded representation of the video. Thus, the coding techniques may be referred to as video compression techniques or video transcoding techniques. The output of the coding component 4004 is either stored or transmitted over a communication connection, as represented by component 4006. The stored or communicated bitstream (or coded) representation of the video received at the input 4002 may be used by component 4008 to generate pixel values ​​or displayable video that are sent to the display interface 4010. The process of generating a user viewable video from the bitstream representation may be referred to as decompressing the video. Furthermore, while certain video processing operations are referred to as "coding" operations or tools, it will be understood that the coding tools or operations are used in an encoder and corresponding decoding tools or operations that reverse the results of the coding are performed in a decoder.

[0080] Examples of peripheral bus interfaces or display interfaces include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described in this disclosure may be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0081] FIG. 7 is a block diagram of an example of a video processing device 4100. The device 4100 may be used to perform one or more of the methods described herein. The device 4100 may be implemented in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 4100 may include one or more processors 4102, one or more memories 4104, and a video processing circuit 4106. The processor(s) 4102 may be configured to perform one or more methods described in this disclosure. The memory(s) 4104 may be used to store data and coding used to perform the methods and techniques described in this disclosure. The video processing circuit 4106 may be used to perform some techniques described in this disclosure in a hardware circuit. In some embodiments, the video processing circuit 4106 may be at least partially included in the processor 4102, for example, a graphics co-processor.

[0082] 8 is a flow chart of an example method 4200 of visual media processing. The method 4200 determines an EDRAP sample at step 4202. The EDRAP sample is a sample that can correctly decode all subsequent samples, both in decoding order and output order, provided that necessary preceding Streaming Access Points (SAPs) or EDRAP samples are available for reference when decoding the EDRAP sample and subsequent samples. At step 4204, conversion is performed between visual media data based on the EDRAP sample and a media data file.

[0083] It should be noted that method 4200 may be implemented in an apparatus for processing visual media data comprising a processor and a non-transitory memory having instructions, such as video encoder 4400, video decoder 4500, and / or encoder 4600. In such a case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 may be executed by a non-transitory computer-readable medium including a computer program product for use by a video coding device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by the processor, cause the video coding device to perform method 4200.

[0084] 9 is a block diagram illustrating an example of a video coding system 4300 that may utilize techniques of this disclosure. The video coding system 4300 may include a source device 4310 and a destination device 4320. The source device 4310 generates encoded video data, which may be referred to as a video encoding device. The destination device 4320 may decode the encoded video data generated by the source device 4310, which may be referred to as a video decoding device.

[0085] The source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. The video source 4312 may include a source, such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may consist of one or more pictures. The video encoder 4314 encodes the video data from the video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include a coded picture and associated data. A coded picture is a coded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to the destination device 4320 via the I / O interface 4316 through the network 4330. The encoded video data may be stored in a storage medium / server 4340 for access by the destination device 4320 .

[0086] The destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from the source device 4310 or the storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320 or may be external to the destination device 4320 that may be configured to interface with an external display device.

[0087] The video encoder 4314 and the video decoder 4324 may operate in accordance with a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or future standards.

[0088] 10 is a block diagram illustrating an example of a video encoder 4400, which may be the video encoder 4314 in the system 4300 shown in FIG. 9. The video encoder 4400 may be configured to perform some or all of the techniques of this disclosure. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure may be shared among various components of the video encoder 4400. In some examples, a processor may be configured to perform some or all of the techniques described in this disclosure.

[0089] Functional components of the video encoder 4400 may include a splitting unit 4401, a prediction unit 4402 which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, an intra prediction unit 4406, a residual generation unit 4407, a transformation unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transformation unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.

[0090] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the predictor 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0091] Additionally, some components, such as the motion estimator 4404 and the motion compensator 4405, may be highly integrated, but are depicted separately in the example video encoder 4400 for purposes of illustration.

[0092] The divider 4401 can divide a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support a variety of video block sizes.

[0093] The mode selector 4403 may select either intra or inter coding mode, for example based on the error result, and provide the resulting intra or inter coded block to the residual generator 4407, which generates residual block data, and to the reconstruction unit 4412, which reconstructs the coded block for use as a reference picture. In some examples, the mode selector 4403 may select a combined intra and inter prediction (CIIP) mode, where prediction is based on an inter prediction signal and an intra prediction signal. The mode selector 4403 may also similarly select the resolution of the motion vector of the block (e.g., sub-pixel precision or integer pixel precision) in the case of inter prediction.

[0094] To perform inter prediction on a current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing one or more reference frames from buffer 4413 to the current video block. The motion compensation unit 4405 may determine a prediction video block for the current video block based on the motion information and decoded picture samples of pictures from buffer 4413 other than the picture associated with the current video block.

[0095] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different processing on the current video block depending on whether the current video block is an I slice, a P slice, or a B slice, for example.

[0096] In some examples, the motion estimator 4404 may perform unidirectional prediction on the current video block, and the motion estimator 4404 may search reference pictures in list 0 or list 1 for a reference block for the current video block. The motion estimator 4404 may then generate a reference index indicating a reference picture in list 0 or list 1 that contains the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. The motion estimator 4404 may output the reference index, the prediction direction indicator, and the motion vector as motion information for the current video block. The motion compensation unit 4405 may generate a prediction video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0097] In another example, the motion estimator 4404 may perform bidirectional prediction on the current video block, and the motion estimator 4404 may search the reference pictures of list 0 for a reference video block for the current video block and may also search the reference pictures of list 1 for another reference video block for the current video block. The motion estimator 4404 may then generate reference indexes indicating the reference pictures of list 0 and list 1 that contain the reference video blocks, and motion vectors indicating spatial displacements between the reference video blocks and the current video block. The motion estimator 4404 may output the reference indexes and the motion vector of the current video block as motion information for the current video block. The motion compensation unit 4405 may generate a prediction video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0098] In some examples, the motion estimator 4404 may output a full set of motion information for a decoder's decoding process. In some examples, the motion estimator 4404 may not output a full set of motion information for a video. Rather, the motion estimator 4404 may signal motion information for a video block by reference to motion information of another video block. For example, the motion estimator 4404 may determine that the motion information of a current video block is sufficiently similar to the motion information of a neighboring video block.

[0099] In one example, the motion estimator 4404 may indicate to the video decoder 4500 a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block.

[0100] In another example, the motion estimator 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 can use the difference between the motion vector of the video block and the motion vector to determine the motion vector of the current video block.

[0101] As mentioned above, the video encoder 4400 can predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 4400 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0102] The intra predictor 4406 may perform intra prediction on the current video block. When the intra predictor 4406 performs intra prediction on the current video block, the intra predictor 4406 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.

[0103] The residual generator 4407 may generate residual data for the current video block by subtracting a prediction video block(s) of the current video block from the current video block. The residual video blocks of the current video block may include residual video blocks corresponding to different sample components of the samples of the current video block.

[0104] In other examples, for example in skip mode, residual data for the current video block may not be present and residual generator 4407 may not perform a subtraction operation.

[0105] The transform unit 4408 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to a residual video block related to the current video block.

[0106] After the transform unit 4408 generates a transform coefficient image block associated with the current image block, the quantization unit 4409 may quantize the transform coefficient image block associated with the current image block based on one or more quantization parameter (QP) values ​​associated with the current image block.

[0107] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples of one or more prediction video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.

[0108] After the reconstructor 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0109] The entropy encoder 4414 may receive data from other functional components of the video encoder 4400. Once the entropy encoder 4414 receives the data, it may perform one or more entropy encoding operations to generate entropy coded data and output a bitstream that includes the entropy coded data.

[0110] 11 is a block diagram illustrating an example of a video decoder 4500, which may be the video decoder 4324 in the system 4300 illustrated in FIG. 9. The video decoder 4500 may be configured to perform some or all of the techniques of this disclosure. In the illustrated example, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure may be shared among various components of the video decoder 4500. In some examples, a processor may be configured to perform some or all of the techniques described in this disclosure.

[0111] In the shown example, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. The video decoder 4500 may, in some examples, perform a decoding pass that is generally inverse to the encoding pass described with respect to the video encoder 4400.

[0112] The entropy decoding unit 4501 may obtain an encoded bitstream. The encoded bitstream may include entropy coded video data (e.g., coded blocks of video data). The entropy decoding unit 4501 may decode the entropy coded video data, and from the entropy decoded video data, the motion compensation unit 4502 may determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 may determine the information, for example, by implementing AMVP and merge mode.

[0113] The motion compensation unit 4502 generates motion compensated blocks and may perform interpolation based on an interpolation filter. An identifier of the interpolation filter used with sub-pixel accuracy may be included in the syntax element.

[0114] The motion compensation unit 4502 may calculate sub-integer pixel interpolated values ​​of the reference block using an interpolation filter used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filter used by the video encoder 4400 according to the received syntax information and generate a prediction block using the interpolation filter.

[0115] The motion compensation unit 4502 may use a portion of the syntax information to determine the size of the blocks used to code a frame(s) and / or slice(s) of the coded video sequence, partition information describing how each macroblock of a picture of the coded video sequence is divided, a mode indicating how each partition is coded, one or more reference frames (and reference frame lists) between each inter-coded block, and other information for decoding the coded video sequence.

[0116] The intra prediction unit 4503 may form a prediction block from spatially neighboring blocks, for example using an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0117] The reconstruction unit 4506 may sum the residual blocks with corresponding prediction blocks generated by the motion compensation unit 4502 or intra prediction unit 4503 to form decoded blocks. If necessary, a deblocking filter may be applied to filter the decoded blocks to remove blocking artifacts. The decoded video blocks are stored in a buffer 4507 to provide reference blocks for subsequent motion compensation / intra prediction and to generate decoded video for display on a display device.

[0118] FIG. 12 is a schematic diagram of an example of an encoder 4600. The encoder 4600 is suitable for implementing the technology of VVC. The encoder 4600 includes three in-loop filters, namely, a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 utilize the original samples of the picture to add an offset and reduce the mean square error between the original samples and the reconstructed samples by applying a Finite Impulse Response (FIR) filter with coded side information signaling the offset and the filter coefficients, respectively. The ALF 4606 is located at the final processing stage of each picture and can be considered as a tool that tries to catch and correct the artifacts created in the previous stage.

[0119] The encoder 4600 further includes an intra prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive an input video. The intra prediction component 4608 is configured to perform intra prediction, and the ME / MC component 4610 is configured to perform inter prediction using a reference picture obtained from a reference picture buffer 4612. Residual blocks from the inter prediction or intra prediction are provided to a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients and provided to an entropy coding component 4618. The entropy coding component 4618 entropy codes the prediction results and the quantized transform coefficients and transmits them towards a video decoder (not shown). The quantization components output from the quantization component 4616 may be provided to an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output pictures to the DF 4602, SAO 4604, and ALF 4606 for filtering before the pictures are stored in the reference picture buffer 4612.

[0120] Below, we list some examples of possible solutions.

[0121] The following solutions provide examples of the techniques discussed in this disclosure.

[0122] The following solutions represent example implementations of the techniques described in the previous section (eg, item 1).

[0123] 1. A method for processing visual media data (e.g., method 4200 shown in FIG. 8), comprising: determining (4202) an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample that can correctly decode all subsequent samples, in both decoding order and output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples; and performing a conversion between the visual media data and a media data file based on the EDRAP sample (4204).

[0124] 2. The method according to solution 1, wherein the closest predecessor SAP sample and zero or more predecessor EDRAP samples of type 1, 2 or 3 are referred to as the necessary predecessor SAP samples and the necessary EDRAP samples of the EDRAP sample.

[0125] The following solutions represent example implementations of the techniques described in the previous section (eg, item 2).

[0126] 3. A method for processing visual media data, comprising: determining an Extended Dependent Random Access Point (EDRAP) sample; and, if a video track has a track reference of type "aest" that references an associated track, for each EDRAP sample denoted as sampleA in the video track, there is only one sample denoted as sampleB in the associated track that has the same decoding time as sampleA; and performing conversion between the visual media data and a media data file based on the EDRAP sample.

[0127] 4. The method according to any of Solutions 1 to 3, wherein sampleB contains all pictures in the closestSapSample of sampleA and the necessary preceding SAP or EDRAP samples of sampleA.

[0128] 5. An apparatus for processing visual media data, comprising a processor and a non-transitory memory having instructions stored thereon, the instructions, when executed by the processor, causing the processor to perform a method according to any one of solutions 1 to 4.

[0129] 6. A non-transitory computer-readable medium including a computer program product for use by a video coding device, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium that, when executed by a processor of the video coding device, causes the video coding device to perform a method according to any one of Solutions 1 to 4.

[0130] 7. A non-transitory computer-readable recording medium storing a media data file generated by a method performed by a media processing device, the method including: determining an Enhanced Dependent Random Access Point (EDRAP) sample; and generating the media data file based on determining that the EDRAP sample is a sample that can correctly decode all subsequent samples, in both decoding order and output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and the subsequent samples.

[0131] 8. A method for storing a media data file for a video, the method comprising: determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample that will correctly decode all subsequent samples, in both decoding order and output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and the subsequent samples; generating a media data file based on the determination; and storing the media data file on a non-transitory computer-readable recording medium.

[0132] 9. A method, apparatus or system as described herein.

[0133] In the solution, an encoder can comply with the format rules by generating a coded representation according to the format rules. In the solution described in this disclosure, a decoder can parse syntax elements in the coded representation with knowledge of the presence or absence of syntax elements according to the format rules and generate decoded video using the format rules.

[0134] In this specification, the term "video processing" may refer to video encoding, video decoding, video compression or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation or vice versa. The bitstream representation of a current video block corresponds to bits spread in the same or different locations in the bitstream, for example as defined by a syntax. For example, a macroblock is coded in terms of transformed and coded error residual values, and using bits of headers and other fields in the bitstream. Furthermore, during conversion, the decoder may analyze the bitstream with the knowledge that some fields may or may not be present based on a decision, as described in the solution. Similarly, the encoder may determine that a particular syntax field is or is not included and generate the coded representation accordingly by including or excluding the syntax field in the coded representation.

[0135] The present disclosure and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. The present disclosure and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or for controlling the operation of a data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter effecting a machine-readable propagated signal, or one or more combinations thereof. The term "data processing device" encompasses all apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that establishes an execution environment for the computer program, such as code that constitutes the firmware of the processor, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, for example a mechanically generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiving device.

[0136] A computer program (also referred to as a program, software, software application, script, or coding) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other part suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in several coordinated files (e.g., files that store one or more modules, subprograms, or parts of code). A computer program can be arranged to be executed on one computer, on several computers located at one site, or on several computers distributed across several sites and interconnected by a communication network.

[0137] The processes and logic flows described in this disclosure may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, or an apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).

[0138] Processors suitable for executing computer programs include, by way of example, both general purpose and special purpose microprocessors, and one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer is also operatively coupled to receive data from, transfer data to, or both of, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, and optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks; and compact disk read-only memory (CD ROM) and digital versatile disk read-only memory (DVD-ROM) disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0139] Although this patent document contains many specifics, these should not be interpreted as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to certain embodiments of a particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as acting in a particular combination, and may even be initially claimed as such, one or more features from a claimed combination may in some cases be excluded from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0140] Similarly, although operations are depicted in the figures in a particular order, this should not be understood as requiring such operations to be performed in the particular order shown, or in any sequential order, or that all of the illustrated operations be performed, to achieve desirable results. Further, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0141] Although only a few implementations and examples have been described, other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

[0142] A first component is directly coupled to a second component when there are no intervening components, other than lines, traces, or other media, between the first and second components. A first component is indirectly coupled to a second component when there are intervening components, other than lines, traces, or other media, between the first and second components. The term "coupled" and variations thereof include both direct and indirect couplings. Use of the term "about" means a range that includes ±10% of the succeeding numerical value, unless otherwise specified.

[0143] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of the disclosure. The examples are to be considered as illustrative and not restrictive, and the intent is not to be limited to the details given herein. For example, various elements or components may be combined or integrated in other systems, or certain features may be omitted or not implemented.

[0144] Furthermore, in each embodiment, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of the present disclosure. Other items shown or discussed as being coupled may be directly connected, or indirectly coupled or in communication through some interface, device, or intermediate component, whether electrical, mechanical, or otherwise. Other examples of changes, substitutions, and alterations are ascertainable by those skilled in the art and may be made without departing from the spirit and scope disclosed herein.

Claims

1. 1. A method for processing visual media data, comprising: determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample from which all subsequent samples can be correctly decoded, both in decoding order and in output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples; and performing a conversion between visual media data and a media data file based on the EDRAP sample. method.

2. 2. The method of claim 1 , wherein the required preceding SAP or EDRAP samples include one or more of a sample set starting with a closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order and including any EDRAP samples between the closestSapSample and the EDRAP sample in decoding order.

3. 2. The method of claim 1, wherein if a media track containing the EDRAP sample has a track reference of type "aest" that references an associated track, then for each EDRAP sample denoted as sampleA in the media track, there is exactly one sample denoted as sampleB in the associated track that has the same decoding time as sampleA.

4. The method of claim 3 , wherein if the media track contains a TrackReferenceTypeBox with reference_type equal to “aest”, then the TrackReferenceTypeBox contains only a track identifier and does not contain a track group identifier.

5. 4. The method of claim 3, wherein sampleB includes the closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order of sampleA and all media data of the required preceding SAP or EDRAP samples of sampleA.

6. The method of claim 1 , wherein the converting comprises generating the media data file from the visual media data.

7. The method of claim 1 , wherein the converting comprises parsing the media data file into the visual media data.

8. 1. An apparatus for processing visual media data, comprising: one or more processors; one or more non-transitory memories having instructions stored thereon, the instructions, when executed by the one or more processors, causing the one or more processors to: determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample from which all subsequent samples can be correctly decoded, both in decoding order and in output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples; performing a conversion between visual media data and a media data file based on the EDRAP sample; Device.

9. 9. The apparatus of claim 8, wherein the required preceding SAP or EDRAP samples include one or more of a sample set starting with a closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order and including any EDRAP samples between the closestSapSample and the EDRAP sample in decoding order.

10. 9. The apparatus of claim 8, wherein if a media track containing the EDRAP sample has a track reference of type "aest" that references an associated track, then for each EDRAP sample denoted as sampleA in the media track, there is exactly one sample denoted as sampleB in the associated track that has the same decoding time as sampleA.

11. The apparatus of claim 10 , wherein if the media track includes a TrackReferenceTypeBox with reference_type equal to “aest”, then the TrackReferenceTypeBox includes only a track identifier and does not include a track group identifier.

12. 11. The apparatus of claim 10, wherein sampleB includes the closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order of sampleA and all media data of the required preceding SAP or EDRAP samples of sampleA.

13. 1. A non-transitory computer-readable medium containing a computer program product for use by a video coding apparatus, the computer program product, when executed by one or more processors of the video coding apparatus, causing the video coding apparatus to: determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample from which all subsequent samples can be correctly decoded, both in decoding order and in output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples; performing a conversion between visual media data and a media data file based on the EDRAP sample; and Non-transitory computer-readable medium.

14. 14. The non-transitory computer-readable medium of claim 13, wherein the required preceding SAP or EDRAP samples include one or more of a sample set starting with a closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order and including any EDRAP samples between the closestSapSample and the EDRAP sample in decoding order.

15. 14. The non-transitory computer-readable medium of claim 13, wherein if a media track containing the EDRAP sample has a track reference of type "aest" that references an associated track, then for each EDRAP sample denoted as sampleA in the media track, there shall be exactly one sample denoted as sampleB in the associated track that has the same decoding time as sampleA.

16. 16. The non-transitory computer-readable medium of claim 15, wherein if the media track includes a TrackReferenceTypeBox with reference_type equal to "aest", then the TrackReferenceTypeBox includes only a track identifier and does not include a track group identifier.

17. 16. The non-transitory computer-readable medium of claim 15, wherein sampleB includes a closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order of sampleA and all media data of the required preceding SAP or EDRAP samples of sampleA.

18. 1. A non-transitory computer-readable recording medium storing a media data file generated by a method executed by a media processing device, the method comprising: determining an Extended Dependent Random Access Point (EDRAP) sample, the EDRAP sample being a sample from which all subsequent samples can be correctly decoded, both in decoding order and in output order, provided that a necessary preceding Streaming Access Point (SAP) or EDRAP sample is available for reference when decoding the EDRAP sample and subsequent samples; generating the media data file based on said determining. A non-transitory computer-readable recording medium.

19. 20. The non-transitory computer-readable storage medium of claim 18, wherein the required preceding SAP or EDRAP samples include one or more of a sample set starting with a closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order and including any EDRAP samples between the closestSapSample and the EDRAP sample in decoding order.

20. 20. The non-transitory computer-readable storage medium of claim 18, wherein if a media track containing the EDRAP sample has a track reference of type "aest" that references an associated track, then for each EDRAP sample denoted as sampleA in the media track, there is exactly one sample denoted as sampleB in the associated track that has the same decoding time as sampleA.

21. 21. The non-transitory computer-readable medium of claim 20, wherein if the media track includes a TrackReferenceTypeBox with reference_type equal to "aest", then the TrackReferenceTypeBox includes only a track identifier and does not include a track group identifier.

22. 21. The non-transitory computer-readable storage medium of claim 20, wherein sampleB includes the closest preceding SAP sample (closestSapSample) of type 1, 2, or 3 in decoding order of sampleA and all media data of the required preceding SAP or EDRAP samples of sampleA.

Citation Information

Patent Citations

  • Method and apparatus for smooth stream switching in MPEG / 3GPP-DASH

    JP2015518350A

  • Method, device, and computer program for improving random picture access in video streaming

    WO2021204827A1