Methods, apparatus, and media for video processing

By introducing MSR and ESR descriptors, the method addresses inefficiencies in identifying mainstream representations and signaling EDRAP-based video streaming, ensuring efficient and adaptive video streaming.

JP7839271B2Active Publication Date: 2026-04-01BYTEDANCE INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Current methods for identifying mainstream representations in video streaming systems, such as DASH, are inefficient and lack the ability to effectively signal Stream Access Points (SAPs) for Extended Dependent Random Access Points (EDRAP) based video streaming.

Method used

Introduce descriptors like MSR and ESR to identify mainstream and external stream representations, specifying their association and ensuring that EDRAP pictures are the first in segments, allowing efficient identification and access to EDRAP-based video content.

Benefits of technology

Enhances the efficiency of identifying mainstream representations and signaling Stream Access Points, enabling seamless and adaptive video streaming by ensuring EDRAP pictures are correctly identified and accessed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839271000009
    Figure 0007839271000009
  • Figure 0007839271000010
    Figure 0007839271000010
  • Figure 0007839271000011
    Figure 0007839271000011
Patent Text Reader

Abstract

An embodiment of the present disclosure provides a solution for video processing. A video processing method is proposed, the method including: receiving, at a first device, a metadata file from a second device; and determining a descriptor in a dataset in the metadata file, the presence of the descriptor indicating that a representation in the dataset is a mainstream stream representation (MSR).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0002]

[0001] [Cross - Reference to Related Applications] This application claims the benefit of priority of U.S. Provisional Application No. 63 / 251,336, filed on October 1, 2021, the entire content of which is incorporated herein by reference.

[0002] Embodiments of the present disclosure generally relate to video encoding techniques, and more particularly, to main - stream representation descriptors.

Background Art

[0003] Media streaming applications typically rely on Internet Protocol (IP), Transmission Control Protocol (TCP), and Hypertext Transfer Protocol (HTTP) transport methods and usually depend on file formats such as the ISO - based Media File Format (ISOBMFF). One such streaming system is HTTP - based Dynamic Adaptive Streaming over HTTP (DASH). In DASH, there may be multiple representations of the video and / or audio data of multimedia content, and different representations may correspond to different encoding characteristics (e.g., different profiles or levels of video encoding standards, different bitrates, different spatial resolutions, etc.). Also, video encoding and streaming based on Extended Dependent Random Access Point (EDRAP) pictures have been proposed. Therefore, there is a value in researching a mechanism for identifying the main - stream representation.

Summary of the Invention

[0004] Embodiments of the present disclosure provide a solution for video processing.

[0005] In a first aspect, a video processing method is proposed. The method includes: receiving a metadata file from a second device on a first device; determining a descriptor in a dataset within the metadata file, wherein the presence of the descriptor indicates that the representation in the dataset is a mainstream representation (MSR).

[0006] Descriptors are used to identify MSRs based on the method according to the first aspect of this disclosure. Compared with conventional methods of identifying MSRs using attributes, the proposed method has the advantage of being able to identify MSRs more efficiently.

[0007] In a second embodiment, another video processing method is proposed. The method includes the steps of: determining a descriptor in a dataset within a metadata file on a second device, wherein the presence of the descriptor indicates that the representation in the dataset is an MSR; and transmitting the metadata file to a first device.

[0008] Descriptors are used to identify MSRs based on the method according to a second aspect of this disclosure. Compared with conventional methods of identifying MSRs using attributes, the proposed method has the advantage of being able to identify MSRs more efficiently.

[0009] In a third aspect, a device for processing video data is proposed. The device for processing video data includes a processor and a non-temporary memory having instructions. When the instructions are executed by the processor, the processor causes the processor to perform a method according to the first or second aspect of the present disclosure.

[0010] In a fourth aspect, a non-temporary computer-readable storage medium is proposed, which stores instructions causing a processor to perform a method according to the first or second aspect of the present disclosure.

[0011] This summary of the invention is provided to introduce, in a simplified form, a selection of concepts that will be further described in the detailed description below. The content of this invention is not intended to identify the main or essential features of the subject matter described in the claims, nor is it intended to be used to limit the scope of the subject matter described in the claims. [Brief explanation of the drawing]

[0012] The above and other purposes, features, and advantages of the exemplary embodiments of this disclosure will become more apparent through the following detailed description with reference to the attached drawings. In the exemplary embodiments of this disclosure, the same reference numerals generally refer to the same components. [Figure 1] The following are block diagrams illustrating exemplary video coding systems according to several embodiments of this disclosure. [Figure 2] The following are block diagrams of exemplary video encoders according to some embodiments of the present disclosure. [Figure 3] The following are block diagrams of exemplary video decoders according to several embodiments of the present disclosure. [Figure 4] This illustrates the concept of a Random Access Point (RAP). [Figure 5] This illustrates the concept of a Random Access Point (RAP). [Figure 6] This illustrates the concept of dependent random access points (DRAPs). [Figure 7] This illustrates the concept of dependent random access points (DRAPs). [Figure 8] This illustrates the concept of Extended Dependent Random Access Points (EDRAPs). [Figure 9] This illustrates the concept of Extended Dependent Random Access Points (EDRAPs). [Figure 10] This demonstrates EDRAP-based video streaming. [Figure 11] This demonstrates EDRAP-based video streaming. [Figure 12]Flowcharts of video processing methods according to several embodiments of this disclosure are shown. [Figure 13] Flowcharts of video processing methods according to several embodiments of this disclosure are shown. [Figure 14] A block diagram of a computing device capable of carrying out various embodiments of this disclosure is shown. Throughout the drawings, the same or similar reference numerals generally refer to the same or similar elements. [Modes for carrying out the invention]

[0013] Next, the principles of this disclosure will be described with reference to several embodiments. These embodiments are provided for illustrative purposes only and are intended to help those skilled in the art understand and implement this disclosure; they should not be interpreted as implying any limitation on the scope of this disclosure. The disclosures described herein can be implemented in a variety of ways other than those described below.

[0014] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art to which this disclosure belongs.

[0015] References in this disclosure such as “one embodiment,” “one example,” and “exemplary embodiment” indicate that the described embodiments may include certain features, structures, or characteristics, but not all embodiments necessarily include such features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same embodiment. In addition, if certain features, structures, or characteristics are described in relation to an exemplary embodiment, it should be noted that, whether explicitly stated or not, the influence of such features, structures, or characteristics in relation to other embodiments is within the knowledge of those skilled in the art.

[0016] Terms such as "first" and "second" may be used herein to describe various elements, but it should be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0017] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are to be construed to include the plural forms as well, unless the context clearly dictates otherwise. It will be further understood that the terms "comprises," "comprising," "includes," "including," and / or "having," when used herein, specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0018] Exemplary environment FIG. 1 is a block diagram showing an exemplary video encoding system 100 that may utilize the technology of the present disclosure. As shown, the video encoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0019] The video source 112 may include sources such as a video capture device. Examples of video capture devices include, but are not limited to, interfaces that receive video data from a video content provider, computer graphics systems that generate video data, and / or combinations thereof.

[0020] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a set of bits that form the encoded representation of the video data. The bitstream may include the encoded picture and associated data. The encoded picture is the encoded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or transmitter. The encoded video data may be transmitted directly to the destination device 120 via the network 130A through the I / O interface 116. The encoded video data may be stored in a storage medium / server 130B for access by the destination device 120.

[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120, or it may be outside the destination device 120 configured to interface with an external display device.

[0022] The video encoder 114 and video decoder 124 may operate in accordance with video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or further standards.

[0023] Figure 2 is a block diagram showing an example of a video encoder 200, which may be an example of a video encoder 114 in the system 100 shown in Figure 1, according to some embodiments of the present disclosure.

[0024] The video encoder 200 may be configured to implement any or all of the technologies of this disclosure. In the example in Figure 2, the video encoder 200 includes several functional components. The technologies described in this disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to implement any or all of the technologies described in this disclosure.

[0025] In some embodiments, the video encoder 200 may include a splitting unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-prediction unit 206, a residual generation unit 207, a conversion unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse conversion unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intrablock copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture on which the current video block is located.

[0027] Furthermore, several components, such as the motion estimation unit 204 and the motion compensation unit 205, can be integrated, but in the example in Figure 2, they are shown separately for illustrative purposes.

[0028] The splitting unit 201 can divide the picture into one or more video blocks. The video encoder 200 and video decoder 300 can support a variety of video block sizes.

[0029] The mode selection unit 203 may, for example, select one of the intra or inter encoding modes based on the error result, provide the resulting intra-encoded or inter-encoded block to the residual generation unit 207 to generate residual block data, and provide the encoded block to the reconstruction unit 212 to reconstruct it and use it as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter-prediction (CIIP) mode in which the prediction is based on an inter-prediction signal and an intra-prediction signal. In the case of inter-prediction, the mode selection unit 203 may select the resolution of the block motion vector (e.g., sub-pixel or integer pixel precision).

[0030] To perform interpretation on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. The motion compensation unit 205 may determine the predicted video block for the current video block based on the motion information of pictures from buffer 213 other than the picture associated with the current video block and the decoded samples.

[0031] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block depending, for example, whether the video block is currently in an I-slice, P-slice, or B-slice. As used herein, “I-slice” may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some embodiments, “P-slice” and “B-slice” may refer to portions of a picture composed of macroblocks that are independent of macroblocks within the same picture.

[0032] In some examples, the motion estimation unit 204 may perform a unidirectional prediction on the current video block and may search for a reference picture in List 0 or List 1 for the reference video block of the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in List 0 or List 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0033] As an alternative, in another example, the motion estimation unit 204 may perform bidirectional prediction for the current video block. The motion estimation unit 204 may look for a reference picture in List 0 for a reference video block of the current video block, or it may look for a reference picture in List 1 for another reference video block of the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in Lists 0 and 1 containing the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0034] In some examples, the motion estimation unit 204 may output a full set of motion information for the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0035] In one example, the motion estimation unit 204 may indicate a value in the syntax structure currently associated with the video block that tells the video decoder 300 that the current video block has the same motion information as another video block.

[0036] In another example, the motion estimation unit 204 may identify another video block and its motion vector difference (MVD) in the syntax structure associated with the currently associated video block. The motion vector difference represents the difference between the motion vector of the currently associated video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and its motion vector difference to determine the motion vector of the currently associated video block.

[0037] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merged mode signaling.

[0038] The intra-prediction unit 206 can perform intra-prediction on the current video block. When the intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks within the same picture. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0039] The residual generation unit 207 may generate residual data for the current video block by subtracting the predicted video block of the current video block from the current video block (e.g., indicated by a minus sign). The residual data for the current video block may include residual video blocks corresponding to different sample components of the sample within the current video block.

[0040] In other examples, for instance, in skip mode, residual data for the current video block may not exist, and the residual generation unit 207 may not perform a subtraction operation.

[0041] The conversion processing unit 208 may generate one or more conversion coefficient video blocks for the current video block by applying one or more conversions to the residual video block associated with the current video block.

[0042] After the conversion processing unit 208 generates a conversion coefficient video block associated with the current video block, the quantization unit 209 may quantize the conversion coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0043] The inverse quantization unit 210 and the inverse transform unit 211 can reconstruct a residual video block from the transformed coefficient video block by applying inverse quantization and inverse transform, respectively, to the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the currently active video block for storage in the buffer 213.

[0044] After the reconstruction unit 212 reconfigures the video block, a loop filtering operation may be performed to reduce video blocking artifacts within the video block.

[0045] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, it may perform one or more entropy coding operations to generate entropy coded data and output a bitstream containing the entropy coded data.

[0046] Figure 3 is a block diagram showing an example of a video decoder 300, which may be an example of a video decoder 124 in the system 100 shown in Figure 1, according to some embodiments of the present disclosure.

[0047] The video decoder 300 may be configured to perform any or all of the technologies of this disclosure. In the example in Figure 3, the video decoder 300 includes several functional components. The technologies described in this disclosure may be shared among the various components of the video decoder 300. In some examples, the processor may be configured to perform any or all of the technologies described in this disclosure.

[0048] In the example shown in Figure 3, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 may perform a decoding path that is generally the reverse of the encoding path described for the video encoder 200.

[0049] The entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may contain entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-encoded video data, and from the entropy-decoded video data, the motion compensation unit 302 may determine motion information, including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and merge mode. AMVP is used and involves deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of prediction regions within a B slice, identification of which reference picture list is associated with each index. As used herein, in some embodiments, “merge mode” may refer to deriving motion information from spatially or temporally adjacent blocks.

[0050] The motion compensation unit 302 may generate motion-compensated blocks while performing interpolation, possibly based on an interpolation filter. The identifier of the interpolation filter used for subpixel precision may be included in the syntax element.

[0051] The motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using an interpolation filter used by the video encoder 200 during video block encoding. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information and use that interpolation filter to generate a predicted block.

[0052] The motion compensation unit 302 may use at least a portion of the syntax information to determine the size of the blocks used to encode frames and / or slices of the encoded video sequence, partition information describing how each macroblock of the picture in the encoded video sequence is divided, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each interencoded block, and other information for decoding the encoded video sequence. As used herein, in some embodiments, a “slice” may refer to a data structure that can be decoded independently from other slices of the same picture with respect to entropy coding, signal prediction, and residual signal reconstruction. A slice may be either an entire picture or a region of a picture.

[0053] The intra-prediction unit 303 can form prediction blocks from spatially adjacent blocks, for example, using an intra-prediction mode received in a bitstream. The inverse quantization unit 304 inverse quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies the inverse transform.

[0054] The reconstruction unit 306 may obtain the decoded blocks, for example, by adding the residual blocks to the corresponding predicted blocks generated by the motion compensation unit 302 or the intra-prediction unit 303. If necessary, a deblocking filter may be applied to filter the decoded blocks to remove block noise artifacts. The decoded video blocks are then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra-prediction and also generates the decoded video for presentation on a display device.

[0055] Several exemplary embodiments of this disclosure will be described in detail below. Section headings are used in this specification for ease of understanding, but it should be understood that the embodiments disclosed in a section are not limited to that section alone. Furthermore, while certain embodiments are described with reference to a general-purpose video encoding or other specific video codecs, the disclosed techniques are also applicable to other video encoding techniques. Furthermore, while some embodiments describe the video encoding step in detail, it will be understood that the corresponding decoding step, which reverses the encoding, is performed by a decoder. Furthermore, the term video processing includes video encoding or compression, video decoding (decoding) or decompression, and video transcoding, which represents video pixels from one compressed format to another compressed format or a different compressed bitrate. 1. Overview This disclosure relates to video streaming. Specifically, it relates to the design of mainstream and external stream representation descriptors for extended dependent random access point (EDRAP) based video streaming, and to the signaling of stream access points (SAPs) in the mainstream representation. This idea can be applied individually or in various combinations to media streaming systems, for example, based on the Dynamic Adaptive Streaming over HTTP (DASH) standard or its extensions. 2.Background 2.1. Video Encoding Standards Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and these two organizations jointly created the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes time prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established in 2015 by VCEG and MPEG. Since then, many new methods have been adopted by JVET and incorporated into reference software called JEM (Joint Exploration Model). Subsequently, when the Versatile Video Coding (VVC) project was officially launched, JVET was renamed Joint Video Experts Team (JVET). VVC is a new encoding standard that aims for a 50% bitrate reduction compared to HEVC, and was finalized by JVET at its 19th meeting, which concluded on July 1, 2020. The Versatile Video Coding (VVC) standard (ITU-T H.266 | ISO / IEC 23090-3) and the related Versatile Supplemental Enhancement Information (VSEI) standard (ITU-T H.274 | ISO / IEC 23002-7) are designed for use in the widest possible range of applications, including both traditional uses such as television broadcasting, video conferencing, or playback from storage media, and newer, more advanced uses such as adaptive bitrate streaming, video region extraction, content synthesis and merging from multiplexed video bitstreams, multiview video, scalable layered coding, and viewport-adaptive 360-degree immersive media. The Essential Video Coding (EVC) standard (ISO / IEC 23094-1) is another video encoding standard recently developed by MPEG. 2.2. File Format Standards Media streaming applications are typically based on IP, TCP, and HTTP transport methods and depend on file formats such as ISO-based media file formats (ISOBMFF). One such streaming system is HTTP-based Dynamic Adaptive Streaming (DASH). When using video formats with ISOBMFF and DASH, file format specifications specific to the video format, such as AVC or HEVC file formats, may be required for encapsulating video content in ISOBMFF tracks and DASH representations and segments. Important information about the video bitstream, such as profile, hierarchy, level, and much more, may be exposed as file format-level metadata and / or DASH media presentation description (MPD) for content selection purposes, such as selecting appropriate media segments for both initialization at the start of a streaming session and stream adaptation during a streaming session. Similarly, when using image formats with ISOBMFF, specific file format specifications may be required, such as AVC and HEVC image file formats. The VVC video file format, a file format for storing VVC video content based on ISOBMFF, is currently being developed by MPEG. The VVC image file format, a file format for storing image content encoded using VVC and based on ISOBMFF, is currently being developed by MPEG. 2.3. DASH In Dynamic Adaptive Streaming over HTTP (DASH), multiple representations of video and / or audio data of multimedia content may exist, and different representations may correspond to different encoding characteristics (e.g., different profiles or levels of video encoding standards, different bitrates, different spatial resolutions, etc.). A manifest for such representations can be defined in a Media Presentation Description (MPD) data structure. A media presentation can correspond to a structured collection of data accessible to a DASH streaming client device. A DASH streaming client device may request and download media data information to provide a streaming service to the user of the client device. A media presentation can be described in an MPD data structure, including updates to the MPD. A media presentation may consist of one or more periods. Each period may extend until the start of the next period, or, in the case of the final period, until the end of the media presentation. Each period may contain one or more representations of the same media content. A representation may be audio, video, timed text, or one of many alternative encoded versions of such data. A representation may differ by encoding type, e.g., bitrate, resolution, and / or codec for video data, and bitrate, language, and / or codec for audio data. The term "representation" may be used to refer to a section of encoded audio or video data that corresponds to a particular period of multimedia content and is encoded in a particular manner. Representations for a specific period can be assigned to a group indicated by an attribute in the MPD that identifies the adaptation set to which the representation belongs. Representations within the same adaptation set are generally considered interchangeable with each other, in that a client device can dynamically and seamlessly switch between these representations to perform, for example, bandwidth adaptation. For example, each representation of video data for a particular period may be assigned to the same adaptation set, but one of the representations may be selected for decoding to present media data, such as video or audio data, of the multimedia content for the corresponding period. Media content within a period may, in some examples, be represented by either one representation from group 0 (if any), or a combination of up to one representation from each non-zero group. The timing data for each representation of a period may be expressed relative to the start time of the period. A representation may contain one or more segments. Each representation may contain an initialization segment, and each segment of a representation may be self-initializing. If present, the initialization segment may contain initialization information for accessing that representation. Generally, initialization segments do not contain media data. Segments may be uniquely referred to by identifiers such as a Uniform Resource Locator (URL), Uniform Resource Name (URN), or Uniform Resource Identifier (URI). The MPD may provide identifiers for each segment. In some examples, the MPD may provide a range attribute that represents a byte range that may correspond to the data of a segment in a file accessible by a URL, URN, or URI. Different representations may be selected to search for different types of media data substantially simultaneously. For example, a client device may select audio, video, and timed text representations to search for a segment. In some examples, a client device may select a specific set of adaptations to perform bandwidth adaptation; that is, a client device may select an adaptation set containing video representations, an adaptation set containing audio representations, and / or an adaptation set containing timed text. Alternatively, a client device may select an adaptation set for a specific type of media (e.g., video) and directly select representations of other types of media (e.g., audio and / or timed text). The general procedure for DASH streaming is outlined in the following steps. 1) The client obtains the MPD. 2) The client estimates the downlink bandwidth and selects the video and audio representations according to the estimated downlink bandwidth, codec, decoding capability, display size, and audio language settings. 3) Unless the end of the media presentation is reached, the client will request the media segment of the selected presentation and present the streaming content to the user. 4) The client continues to estimate the downlink bandwidth. If the bandwidth changes significantly in any direction (for example, decreases), the client selects a different video representation that matches the newly estimated bandwidth and proceeds to step 3. 2.4. Extended Dependent Random Access Point (EDRAP) Picture-Based Video Coding and Streaming Signaling of EDRAP pictures using Auxiliary Extended Information (SEI) messages was proposed in JVET-U0084 and adopted into the VSEI specification at the 21st JVET meeting in January 2021. At the 133rd MPEG meeting in January 2021, the EDRAP sample group was agreed upon based on the proposal in MPEG input document m56020. Regarding support for EDRAP-based video streaming, at the 134th MPEG meeting in April 2021, MPEG input document m56675 proposed an external stream track (EST) design for ISOBMFF. MPEG input document m57430 proposed an external stream representation (ESR) design for DASH. Figures 4 and 5 illustrate the existing concept of a Random Access Point (RAP). The application (e.g., adaptive streaming) determines the frequency of the Random Access Point (RAP) (e.g., RAP duration of 1 second or 2 seconds). Traditionally, RAPs are provided by encoding IRAP pictures, as shown in Figure 4. Note that interpredictive referencing of non-key pictures between RAP pictures is not shown, and the output order is left to right. When random access is performed from CRA6, the decoder receives and correctly decodes the picture, as shown in Figure 5. Figures 6 and 7 illustrate the concept of dependent random access point (DRAP). The DRAP approach provides improved coding efficiency by allowing DRAP pictures (and subsequent pictures) to reference the previous IRAP picture for interpretation, as shown in Figure 6. Note that interpretation of non-key pictures between RAP pictures is not shown, and the output order is left to right. When random access is performed from DRAP6, the decoder receives the picture and decodes it correctly, as shown in Figure 7. Figures 8 and 9 illustrate the concept of Extended Dependent Random Access Point (EDRAP). The EDRAP approach offers greater flexibility by allowing EDRAP pictures (and subsequent pictures) to reference several previous RAP pictures (IRAP or EDRAP), as shown in Figure 8, for example. Note that interpretation of non-key pictures between RAP pictures is not shown, and the output order is left to right. When random access is performed from EDRAP6, the decoder receives and correctly decodes the picture, as shown in Figure 9. Figures 10 and 11 illustrate EDRAP-based video streaming. When random access is performed from the segment starting with EDRAP6, or when switching to that segment, the decoder receives and decodes the segment, as shown in Figure 11. The ESR design proposed in the MPEG input document m57430 is as follows: 2.1.1 Overview The External Stream Representation (ESR) is time-synchronized with the associated Mainstream Representation (MSR), i.e., the "normal" representation. The ESR contains only the Random Access Point (RAP) pictures that are additionally required when randomly accessing time-synchronized Extended Dependent Random Access Point (EDRAP) pictures / samples within the MSR. The design can be summarized as follows: 1) Five definitions have been proposed for the terms EDRAP picture, external elementary stream, external picture, external stream representation (ESR), and mainstream stream representation (MSR). 2) An optional adaptation set level attribute named @esasFlag has been proposed to indicate whether the representation within the adaptation set is ESR or MSR. 3) As part of the semantics of the @esasFlag attribute, the following has been proposed: a. The association between ESRs and MSRs via the existing representation attributes @associationId and @associationType is based on the newly specified association type value "aest" ("associated external stream track", same 4CC as the ISOBMFF track reference type). b. A new “EssentialProperty” descriptor is proposed to be included in an adaptation set containing an ESR, indicating that a representation in such an adaptation set cannot be consumed or played independently without other video representations. c. Several constraints for simplifying EDRAP-based streaming operations: i. Each EDRAP picture within MSR is assumed to be the first picture within the segment. ii. The following constraints apply to interrelated MSRs and ESRs: 1. For each segment in the MSR that begins with an EDRAP picture, it is assumed that there exists a segment in the ESR that has the same segment start time derived from the MPD as a segment in the MSR, and the segment in the ESR carries the external picture necessary for decoding that EDRAP picture and the subsequent picture in the decoding order in the bitstream carried by the MSR. 2. For each segment within an MSR that does not start with an EDRAP picture, it is assumed that there are no segments within an ESR that have the same segment start time derived from the same MPD as the segment within the MSR. 2.1.2 Definition Extended Dependent Random Access Point (EDRAP) Picture Pictures in samples that are members of the EDRAP or DRAP sample group within the ISOBMFF track External Elemental Stream Elementary stream containing access units with external pictures External picture When an external elementary stream within an ESR is randomly accessed from a specific EDRAP picture within an MSR, the picture required for interpredictive referencing in decoding the elementary stream within the MSR is located in the external elementary stream within the ESR. External stream representation (ESR) Expressions including external elemental streams Mainstream Representation (MSR) Expressions including video elementary stream

[0056] 2.1.3 Semantics of the AdaptationSet element [Table 1] TIFF0007839271000002.tif238166

[0057] 2.1.4 XML Syntax

number

[0058] 3.Problems The design proposed in MPEG input document m57430 has the following problem: For Main Streaming Representation (MSR), the current definitions of different Stream Access Point (SAP) types cannot be applied to EDRAP-based Random Access Points because external pictures from different tracks or representations are required. This makes it impossible to signal whether a segment starts with an SAP, and what type of SAP it is. 4. Detailed plan To solve the above problems, a method summarized below is disclosed. The embodiments should be considered as examples to illustrate a general concept and should not be interpreted narrowly. Furthermore, these embodiments can be applied individually or in any combination. 1) A mainstream representation (MSR) descriptor is specified to identify the MSR. a. In one example, an MSR descriptor is defined as an “EssentialProperty” descriptor with a specific value for @schemeIdUri (e.g., urn:mpeg:dash:msr:2021). i. In one example, an MSR descriptor is specified to be included in an adaptation set, i.e., to be at the adaptation set level. When included in an adaptation set, it indicates that all representations within the adaptation set are MSRs. ii. In one example, an MSR descriptor is specified to be included in a representation, i.e., to be at the representation level. When included in a representation, it indicates that the representation is an MSR. iii. In one example, an MSR descriptor is specified to be included in either an adaptation set or a representation, i.e., to be either an adaptation set level or a representation level. 1. If included in an adaptation set, indicate that all representations within the adaptation set are MSRs. a. As an alternative, indicate that some or all of the representations within an adaptation set may be MSRs if they are included in the adaptation set. 2. If included in a representation, indicate that the representation is an MSR. b. In one example, an MSR descriptor is defined as a “SupplementalProperty” descriptor with a specific value for @schemeIdUri (e.g., urn:mpeg:dash:msr:2021). 2) Each Stream Access Point (SAP) within the MSR specifies that it can be used to access the content within the representation only if a time-synchronized sample exists in the track being transported by the associated ESR and is available to the client. 3) Optionally, specify that each EDRAP picture in the MSR is the first picture in the segment (i.e., each EDRAP picture starts the segment). 4) An External Stream Representation (ESR) descriptor is specified to identify the ESR. a. In one example, an ESR descriptor is defined as an “EssentialProperty” descriptor with a specific value for @schemeIdUri (e.g., equal to urn:mpeg:dash:esr:2021). i. In one example, an ESR descriptor is specified to be included in an adaptation set, i.e., to be at the adaptation set level. When included in an adaptation set, it indicates that all representations within the adaptation set are ESRs. ii. In one example, an ESR descriptor is specified to be included in a representation, i.e., to be at the representation level. If it is included in a representation, it indicates that the representation is an ESR. iii. In one example, the ESR descriptor is specified to be included in either an adaptation set or a representation, i.e., to be at the adaptation set level or / or representation level. 1. If included in an adaptation set, indicate that all representations within the adaptation set are ESRs. a. As an alternative, if included in an adaptation set, indicate that some or all of the representations within the adaptation set may be ESRs. 2. If included in a representation, indicate that the representation is ESR. b. In one example, an ESR descriptor is defined as a “SupplementalProperty” descriptor with a specific value for @schemeIdUri (e.g., urn:mpeg:dash:msr:2021). 5) Specify that each ESR shall be associated with an MSR through the (existing) representation level attributes @associationId and @associationType within the MSR, as follows: The @id of the associated ESR shall be referenced by the value contained in the attribute @associationId, where the corresponding value of the attribute @associationType is equal to "aest". 5. Embodiments The following are some exemplary embodiments relating to all of the plan items and some of their sub-items summarized above in Section 4. These embodiments may be applied to DASH. Changes are marked in relation to the text of the design in Clause 2.4. Most of the added or modified parts are underline Some parts of the deleted sections are marked with a checkmark and have a strikethrough. There may also be several other changes that are not highlighted due to their editing nature. 5.1.1 Definition Extended Dependent Random Access Point (EDRAP) Picture Pictures in samples that are members of the EDRAP or DRAP sample group within the ISOBMFF track External Elemental Stream Elementary stream containing access units with external pictures External picture When an external elementary stream within an ESR is randomly accessed from a specific EDRAP picture within an MSR, the picture required for interpredictive referencing in decoding the elementary stream within the MSR is located in the external elementary stream within the ESR. External stream representation (ESR) Expressions including external elemental streams Mainstream Representation (MSR) Expressions including video elementary stream 5.1.2 MSR and ESR Descriptors An adaptation set may have an "EssentialProperty" descriptor where @schemeIdUri is equal to urn:mpeg:dash:msr:2021. This descriptor is called an MSR descriptor. The presence of this "EssentialProperty" indicates that each representation in this adaptation set is an MSR. The following applies to MSR: - Each SAP within an MSR representation in an adaptation set can be used to access the content within the representation, only if a time-synchronized sample exists in the truck being transported by the associated ESR and is available to the client. - Each EDRAP picture within the MSR is assumed to be the first picture within the segment (i.e., each EDRAP picture is assumed to be the start of the segment). An adaptation set may have an "EssentialProperty" descriptor where @schemeIdUri is equal to urn:mpeg:dash:esr:2021. This descriptor is called an ESR descriptor. The presence of this "EssentialProperty" indicates that each representation within this adaptation set is an ESR. An ESR is not meant to be consumed or played on its own without other video representations. Each MSR shall be associated with the MSR through the (existing) representation level attributes @associationId and @associationType within the MSR, as follows: The @id of the associated ESR shall be referenced by the value contained in the attribute @associationId, where the corresponding value of the attribute @associationType is equal to "aest". Optionally, the following constraints apply to MSRs and ESRs that are linked to each other through the representation attributes @associationId and @associationType within the MSR: - For each segment in the MSR that begins with an EDRAP picture, it is assumed that there exists a segment in the ESR with the same segment start time derived from the same MPD as the segment in the MSR, and the segment in the ESR carries the external picture necessary for decoding that EDRAP picture and the subsequent picture in the decoding order within the bitstream carried by the MSR. - For each segment within an MSR that does not start with an EDRAP picture, it is assumed that there are no segments within an ESR that have the same segment start time derived from the same MPD as the segment within the MSR.

[0059] 5.1.3 Semantics of the AdaptationSet element [Table 2] TIFF0007839271000006.tif231166

[0060] 5.1.4 XML Syntax

number

[0061] Embodiments of this disclosure relate to mainstream representation descriptors.

[0062] Figure 12 shows a flowchart of Method 1200 for video processing according to some embodiments of the present disclosure. Method 1200 may be embodied in a first device. For example, Method 1200 may be embedded in a client or receiver. As used herein, the term “client” may refer to computer hardware or software that accesses services made available by a server as part of a client-server model of a computer network. As mere examples, a client may be a smartphone or a tablet. In some embodiments, the first device may be embodied in the destination device 120 shown in Figure 1.

[0063] In block 1210, the first device receives a metadata file from the second device. The metadata file may contain important information about the video bitstream, such as profiles, hierarchies, levels, etc. For example, the metadata file may be a DASH Media Presentation Description (MPD). It should be understood that the above example is provided for illustrative purposes only. The scope of this disclosure is not limited thereto.

[0064] In block 1220, the first device determines a descriptor in the dataset within the metadata file. The presence of the descriptor indicates that the representation in the dataset is a mainstream representation (MSR). In other words, if the dataset contains the descriptor, it means that the representation in the dataset is an MSR.

[0065] According to Method 1200, descriptors are used to identify MSRs. Compared to conventional methods that use attributes to identify MSRs, the proposed method has the advantage of being able to identify MSRs more efficiently.

[0066] In some embodiments, a descriptor may be defined as a data structure having an attribute equal to a Uniform Resource Name (URN) string. For example, the metadata file may be a Media Presentation Description (MPD), and the data structure may be an EssentialProperty within the MPD. Furthermore, the attribute may be a schemeIdUri attribute, and the URN string may be "urn:mpeg:dash:msr:2022". That is, the descriptor may be defined as an EssentialProperty descriptor having an @schemeIdUri value equal to a specific URN string (e.g., "urn:mpeg:dash:msr:2022"). It should be understood that the possible implementations of URN strings described herein are merely descriptive and should not be construed as limiting this disclosure in any way.

[0067] In another example, the metadata file may be an MPD, and the data structure may be a SupplementalProperty within the MPD. Similarly, the attribute may be a schemeIdUri attribute, and the URN string may be "urn:mpeg:dash:msr:2022". That is, the descriptor may be defined as a SupplementalProperty descriptor with a @schemeIdUri value equal to a specific URN string (e.g., "urn:mpeg:dash:msr:2022"). It should be understood that the possible implementations of the URN strings described herein are merely descriptive and should therefore not be interpreted as limiting this disclosure in any way.

[0068] In some embodiments, the dataset may be an adaptation set. In this case, all representations within the adaptation set may be MSRs. Alternatively, some representations within the adaptation set may be MSRs.

[0069] In some embodiments, the dataset may be a representation. In this case, the representation may be an MSR.

[0070] In some embodiments, an Extended Dependent Random Access Point (EDRAP) sample within the MSR may include an instruction for the Start Access Unit (SAU) of a Stream Access Point (SAP). In one example, the first byte position of the EDRAP sample may be an index of the SAU. It should be understood that the above example is provided for illustrative purposes only. The scope of this disclosure is not limited thereto. This has the advantage that the proposed method can improve compatibility between the MSR and the Stream Access Point (SAP).

[0071] In some additional embodiments, the EDRAP sample may be provided to the decoder after the external stream representation (ESR) sample associated with the EDRAP sample has been provided to the decoder. That is, the first byte position of each EDRAP sample in the MSR may be the ISAU of the SAP, thereby enabling playback of the media stream in the MSR, provided that the corresponding ESR media sample is provided to the media decoder immediately before the EDRAP sample. This makes it possible for the proposed method to signal whether a segment begins with an SAP and what type of SAP it is.

[0072] In some embodiments, the metadata file may be an MDP, and the segments within the MDP begin with an EDRAP picture in the MSR. In one example, each EDRAP picture in the MSR is the first picture in the segment.

[0073] Figure 13 shows a flowchart of Method 1300 for video processing according to some embodiments of the present disclosure. Method 1300 may be embodied in a second device. For example, Method 1300 may be embedded in a server or transmitter. As used herein, the term “server” may refer to a computeable device in which a client accesses the service over a network. The server may be a physical computing device or a virtual computing device. In some embodiments, the second device may be embodied in the source device 110 shown in Figure 1.

[0074] In block 1310, the second device determines the descriptor in the dataset within the metadata file. The metadata file may contain important information about the video bitstream, such as profiles, hierarchies, levels, etc. For example, the metadata file may be a DASH Media Presentation Description (MPD). The presence of the descriptor indicates that the representation in the dataset is a mainstream representation (MSR). In other words, if the dataset contains the descriptor, it means that the representation in the dataset is an MSR.

[0075] In block 1320, the second device sends the metadata file to the first device.

[0076] According to Method 1300, descriptors are used to identify MSRs. Compared to conventional methods that use attributes to identify MSRs, the proposed method has the advantage of being able to identify MSRs more efficiently.

[0077] In some embodiments, a descriptor may be defined as a data structure having an attribute equal to a Uniform Resource Name (URN) string. For example, the metadata file may be a Media Presentation Description (MPD), and the data structure may be an EssentialProperty within the MPD. Furthermore, the attribute may be a schemeIdUri attribute, and the URN string may be "urn:mpeg:dash:msr:2022". That is, the descriptor may be defined as an EssentialProperty descriptor having an @schemeIdUri value equal to a specific URN string (e.g., "urn:mpeg:dash:msr:2022"). It should be understood that the possible embodiment of URN strings described herein is merely descriptive and should therefore not be interpreted as limiting this disclosure in any way.

[0078] In another example, the metadata file may be an MPD, and the data structure may be a SupplementalProperty within the MPD. Similarly, the attribute may be a schemeIdUri attribute, and the URN string may be "urn:mpeg:dash:msr:2022". That is, the descriptor may be defined as a SupplementalProperty descriptor having a @schemeIdUri value equal to a specific URN string (e.g., "urn:mpeg:dash:msr:2022"). It should be understood that the possible embodiment of URN strings described herein is merely descriptive and should therefore not be interpreted as limiting this disclosure in any way.

[0079] In some embodiments, the dataset may be an adaptation set. In this case, all representations within the adaptation set may be MSRs. Alternatively, some representations within the adaptation set may be MSRs.

[0080] In some embodiments, the dataset may be a representation. In this case, the representation may be an MSR.

[0081] In some embodiments, an Extended Dependent Random Access Point (EDRAP) sample within the MSR may include an instruction for the Start Access Unit (SAU) of a Stream Access Point (SAP). In one example, the first byte position of the EDRAP sample may be an index of the SAU. It should be understood that the above example is provided for illustrative purposes only. The scope of this disclosure is not limited thereto. This has the advantage that the proposed method can improve compatibility between the MSR and the Stream Access Point (SAP).

[0082] In some additional embodiments, the EDRAP sample may be provided to the decoder after the external stream representation (ESR) sample associated with the EDRAP sample has been provided to the decoder. That is, the first byte position of each EDRAP sample in the MSR may be the ISAU of the SAP, thereby enabling playback of the media stream in the MSR, provided that the corresponding ESR media sample is provided to the media decoder immediately before the EDRAP sample. This makes it possible for the proposed method to signal whether a segment begins with an SAP and what type of SAP it is.

[0083] In some embodiments, the metadata file may be an MDP, and the segments within the MDP begin with an EDRAP picture in the MSR. In one example, each EDRAP picture in the MSR is the first picture in the segment.

[0084] Embodiments of this disclosure can be embodied individually. Alternatively, embodiments of this disclosure can be embodied in any suitable combination. Embodiments of this disclosure can be described with consideration to the following clauses, and their features can be combined in any reasonable manner.

[0085] Clause 1. A video processing method comprising: a first device receiving a metadata file from a second device; determining a descriptor in a dataset within the metadata file, wherein the presence of the descriptor indicates that the representation in the dataset is a mainstream representation (MSR);

[0086] Clause 2. A video processing method comprising: a step of determining a descriptor in a dataset in a metadata file on a second device, wherein the presence of the descriptor indicates that the representation in the dataset is an MSR; and a step of transmitting the metadata file to a first device.

[0087] Clause 3. The descriptor is defined as a data structure having attributes equal to a uniform resource name (URN) string, as described in any one of Clauses 1 to 2.

[0088] Clause 4. The method according to Clause 3, wherein the metadata file is a media presentation description (MPD) and the data structure is an EssentialProperty in the MPD.

[0089] Clause 5. The method according to Clause 3, wherein the metadata file is a Media Presentation Description (MPD) and the data structure is a SupplementalProperty in the MPD.

[0090] Clause 6. The method according to any one of Clauses 4 to 5, wherein the attribute is the schemeIdUri attribute and the URN string is "urn:mpeg:dash:msr:2022".

[0091] Clause 7. The dataset is an adaptation set or representation, as described in any one of Clauses 1 to 6.

[0092] Clause 8. The method according to any one of Clauses 1 to 6, wherein the dataset is an adaptation set, and all or part of the representations in the adaptation set are MSRs.

[0093] Clause 9. An extended dependent random access point (EDRAP) sample in the MSR is provided for in any one of Clauses 1 to 8, including instructions for a starting access unit (SAU) of a stream access point (SAP).

[0094] Clause 10. The method according to Clause 9, wherein the EDRAP sample is provided to the decoder after the external stream representation (ESR) sample associated with the EDRAP sample has been provided to the decoder.

[0095] Clause 11. The method according to any one of Clauses 9 to 10, wherein the first byte position of the EDRAP sample is the index of the SAU.

[0096] Clause 12. The method according to any one of Clauses 1 to 11, wherein the metadata file is an MDP, and the segments within the MDP begin with an EDRAP picture within the MSR.

[0097] Clause 13. An apparatus for processing video data, comprising a processor and non-temporary memory equipped with instructions, wherein, when the instructions are executed by the processor, the apparatus causes the processor to perform the method described in any one of Clauses 1 to 12.

[0098] Clause 14. A non-temporary, computer-readable storage medium that stores instructions causing a processor to perform the method described in any one of Clauses 1 through 12.

[0099] Example device Figure 14 shows a block diagram of a computing device 1400 that can embody various embodiments of the present disclosure. The computing device 1400 may be embodied as or included in a source device 110 (or a video encoder 114 or 200) or a destination device 120 (or a video decoder 124 or 300).

[0100] It will be understood that the computing device 1400 shown in Figure 14 is for illustrative purposes only and is not intended to imply any limitation in any way to the functionality and scope of the embodiments of this disclosure.

[0101] As shown in Figure 14, the computing device 1400 includes a general-purpose computing device 1400. The computing device 1400 may include at least one or more processors or processing units 1410, memory 1420, storage unit 1430, one or more communication units 1440, one or more input devices 1450, and one or more output devices 1460.

[0102] In some embodiments, the computing device 1400 may be embodied as any user terminal or server terminal having computing capabilities. The server terminal may be a server or large-scale computing device provided by a service provider. The user terminal may be any type of mobile, fixed, or portable terminal, including, for example, a mobile phone, station, unit, device, multimedia computer, multimedia tablet, internet node, communicator, desktop computer, laptop computer, notebook computer, netbook computer, tablet computer, personal communication system (PCS) device, personal navigation device, personal digital assistant (PDA), audio / video player, digital camera / video camera, positioning device, television receiver, radio receiver, e-book device, game device, or any combination thereof (including accessories and peripherals for these devices, or any combination thereof). The computing device 1400 may support any type of interface to the user (such as a “wearable” circuit).

[0103] The processing unit 1410 may be a physical or virtual processor and can implement various processes based on a program stored in memory 1420. In a multiprocessor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the computing device 1400. The processing unit 1410 may be called a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0104] The computing device 1400 typically includes various computer storage media. Such media may be any media accessible by the computing device 1400, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 1420 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or any combination thereof. Storage unit 1430 may be any removable or non-removable medium, which can be used to store information and / or data and is accessible by the computing device 1400, and may include machine-readable media such as memory, flash memory drives, magnetic disks, or other media.

[0105] The computing device 1400 may further include additional removable / non-removable, volatile / non-volatile memory media. Although not shown in Figure 14, it is possible to provide a magnetic disk drive for reading and writing removable non-volatile magnetic disks, and an optical disk drive for reading and writing removable non-volatile optical disks. In such cases, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0106] The communication unit 1440 communicates with further computing devices via a communication medium. Furthermore, the functionality of the components within the computing device 1400 can be embodied by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 1400 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or further general network nodes.

[0107] The input device 1450 may be one or more of various input devices such as a mouse, keyboard, tracking ball, or voice input device. The output device 1460 may be one or more of various output devices such as a display, speaker, or printer. The communication unit 1440 allows the computing device 1400 to communicate further with one or more external devices (not shown), such as a storage device and a display device, which may enable a user to interact with the computing device 1400, or, if necessary, enable the computing device 1400 to communicate with one or more other computing devices via any device (such as a network card or modem). Such communication may be performed via an input / output (I / O) interface (not shown).

[0108] In some embodiments, instead of being integrated into a single device, some or all components of computing device 1400 may be located in a cloud computing architecture. In the cloud computing architecture, components may be delivered remotely and work together to embody the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services, and the end user does not need to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing provides services over a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network that can be accessed through a web browser or other computing component. Software or components of the cloud computing architecture and corresponding data may be stored on servers located remotely. Computing resources in the cloud computing environment may be merged or distributed across remote data center locations. The cloud computing infrastructure may act as a single access point for the user but may provide services through a shared data center. Thus, the components and functions described herein may be provided from a remote service provider using a cloud computing architecture. Alternatively, they may be delivered from conventional servers or installed directly or otherwise on client devices.

[0109] The computing device 1400 may be used to embody video coding / decoding in embodiments of this disclosure. The memory 1420 may include one or more video coding modules 1425 having one or more program instructions. These modules are accessible and executable by the processing unit 1410 to perform the functions of the various embodiments described herein.

[0110] In an exemplary embodiment of video encoding, input device 1450 may receive video data to be encoded as input 1470. The video data may be processed, for example, by video encoding module 1425 to generate an encoded bitstream. The encoded bitstream may be provided as output 1480 via output device 1460.

[0111] In an exemplary embodiment of video decoding, input device 1450 may receive an encoded bitstream as input 1470. The encoded bitstream may be processed, for example, by a video encoding module 1425 to generate decoded video data. The decoded video data may be provided as output 1480 via output device 1460.

[0112] While this disclosure has been illustrated and described in particular with reference to its preferred embodiments, it will be understood by those skilled in the art that various modifications in form and detail can be made without departing from the spirit and scope of this application as defined by the appended claims. Such modifications are to be considered within the scope of this application. Accordingly, the foregoing description relating to embodiments of this application is not intended to be limiting.

Claims

1. A method for media streaming, The first device receives a metadata file from the second device, A method for determining whether a main stream representation (MSR) descriptor exists in a dataset within the metadata file, wherein the presence of the MSR descriptor indicates that the representation in the dataset is a mainstream representation (MSR).

2. The method according to claim 1, wherein the MSR descriptor is defined as a data structure having attributes equal to a uniform resource name (URN) string.

3. The method according to claim 2, wherein the metadata file is a media presentation description (MPD), and the data structure is an EssentialProperty in the MPD.

4. The method according to claim 3, wherein the attribute is the schemeIdUri attribute and the URN string is "urn:mpeg:dash:msr:2022".

5. The method according to claim 1, wherein the dataset is an adaptation set.

6. The method according to claim 1, wherein the extended dependent random access point (EDRAP) sample in the MSR includes an instruction for a starting access unit (SAU) of a stream access point (SAP).

7. The method according to claim 6, wherein the EDRAP sample is provided to the decoder after an external stream representation (ESR) sample associated with the EDRAP sample has been provided to the decoder.

8. The method according to claim 6, wherein the first byte position of the EDRAP sample is the index of the SAU.

9. A method for media streaming, The second device determines whether to include the MSR descriptor in the dataset within the metadata file, wherein the presence of the MSR descriptor indicates that the representation in the dataset is an MSR. Based on the above decision, the metadata file is generated, A method comprising transmitting the metadata file to a first device.

10. The method according to claim 9, wherein the MSR descriptor is defined as a data structure having attributes equal to a uniform resource name (URN) string.

11. The method according to claim 10, wherein the metadata file is a media presentation description (MPD), and the data structure is an EssentialProperty in the MPD.

12. The method according to claim 11, wherein the attribute is the schemeIdUri attribute and the URN string is "urn:mpeg:dash:msr:2022".

13. The method according to claim 9, wherein the dataset is an adaptation set.

14. The method according to claim 9, wherein the Extended Dependent Random Access Point (EDRAP) sample in the MSR includes instructions from the Start Access Unit (SAU) of a Stream Access Point (SAP).

15. The method according to claim 14, wherein the EDRAP sample is provided to the decoder after an external stream representation (ESR) sample associated with the EDRAP sample has been provided to the decoder.

16. The method according to claim 14, wherein the first byte position of the EDRAP sample is the index of the SAU.

17. A device for media streaming comprising a processor and non-temporary memory containing instructions, An apparatus that, when the instruction is executed by the processor, causes the processor to perform the method according to any one of claims 1 to 8.

18. A non-temporary computer-readable storage medium that stores instructions for a processor to perform the method according to any one of claims 1 to 8.

19. A media streaming apparatus comprising a processor and non-temporary memory containing instructions, wherein, when the instructions are executed by the processor, the processor causes the processor to perform the method according to any one of claims 9 to 16.

20. A non-temporary computer-readable storage medium that stores instructions causing a processor to perform the method described in any one of claims 9 to 16.

Citation Information

Patent Citations

  • External stream representation properties

    JP2022170709A

  • Method and apparatus for providing free viewpoint video

    US20200177929A1

  • Session-based information for dynamic adaptive streaming over HTTP

    US20210099508A1

  • Information processing apparatus, information recording medium, and information processing method, and program

    WO2017199743A1