Method, apparatus, and system for signaling preselection

The method addresses the lack of preselection signaling in media file formats by identifying and processing preselection-related boxes, ensuring efficient and reliable format-agnostic handling of media streams for tailored user experiences.

JP7798301B2Active Publication Date: 2026-01-14DOLBY INTERNATIONAL AB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023573086
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-01-07
Filing Date
2022-06-28
Publication Date
2026-01-14
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

Existing media file formats like ISOBMFF lack the ability to signal preselection of media content components, which are essential for providing a tailored user experience, especially in Next-Generation Audio and video technologies.

Method used

A method for processing media streams that includes identifying and signaling preselection-related boxes within the media stream, allowing for efficient determination and processing of tracks contributing to a specific media presentation, using metadata to indicate how these tracks should be processed, and enabling format-agnostic implementations.

Benefits of technology

Enables efficient and flexible signaling of preselection information in media streams, reducing implementation effort and increasing reliability by allowing format-agnostic processing and avoiding computationally expensive operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798301000007
    Figure 0007798301000007
  • Figure 0007798301000008
    Figure 0007798301000008
  • Figure 0007798301000009
    Figure 0007798301000009
Patent Text Reader

Abstract

A method for processing a media stream is described, the method comprising: receiving a packetized media stream according to a predefined transport format, the packetized media stream including a plurality of hierarchical boxes, each associated with a respective box type identifier, the plurality of boxes including one or more track boxes referencing respective tracks indicating media components of the media stream; determining whether the media stream includes a preselection-related box of a predefined type indicating a preselection, the preselection corresponding to a media presentation to a user; if it is determined that the media stream includes the preselection-related box: parsing metadata information corresponding to the preselection-related box, the metadata information indicating a characteristic of the preselection; identifying one or more tracks in the packetized media stream that contribute to the preselection based on the metadata information; and providing the one or more tracks for downstream processing according to the given preselection.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of the following priority applications: U.S. Provisional Application No. 63 / 216,029 (Docket No. D21064USP1), filed June 29, 2021, and U.S. Provisional Application No. 63 / 297,473 (Docket No. D21064USP2), filed January 7, 2022, which are incorporated herein by reference.

[0002] Technical Field The present disclosure is directed to the general field of audio and / or video coding (encoding / decoding), and more particularly to methods, apparatus and systems for signaling a preselection corresponding to a media presentation to a user, and processes therefor. [Background technology]

[0003] Generally speaking, with the advent of Next-Generation Audio (NGA) or similar video technologies, the entire audio experience is no longer transmitted as a pre-generated, single-instance media component, but rather as individual semantic objects delivered separately, providing end users with an efficient way to tailor the content to their preferences.

[0004] For example, dialogue may be provided in multiple languages ​​as additional selectable audio components, and language selection can be implemented by combining different audio components or through different balances between components.

[0005] In a broader sense, modern video compression schemes generally take advantage of the possibility of spanning the entire available media data across several streams for a variety of reasons, including the possibility of conserving transmission bandwidth for users who do not require certain portions of the media asset.

[0006] In any event, as will be understood and appreciated by those skilled in the art, media players typically rely on accompanying metadata for guidance on how to render components to result in a predefined user experience.

[0007] For transmission, several of these content components (or CCs for short) may be multiplexed into one single elementary stream, or the components may be spread across several elementary streams.

[0008] Typically, not all available components should be presented simultaneously; only certain combinations of these components may provide the desired user experience.

[0009] Metadata in the various multiplex and transport layers provides the client (or end user) with the necessary knowledge of all available components, and in some cases this data may also be needed to determine which elementary streams to download and decode.

[0010] The International Organization for Standardization (ISO) includes and specifies the Base Media File Format, commonly known as ISOBMFF. Specifically, ISOBMFF is specified by ISO / IEC 14496-12 MPEG-4 Part 12. Generally speaking, it defines a general structure for time-based multimedia files such as video and / or audio. While most existing multiplex formats already provide a means for annotating components within a file with their respective properties, the ISO Base Media File Format (ISOBMFF) multiplex appears to somewhat lack the ability to signal the overall experience composed of a combination of content components. In short, there appears to be a gap compared to some other standards that appear to have introduced the concept of preselection, such as MPEG-DASH (ISO / IEC 23009-1).

[0011] In light of this, it appears that, generally speaking and particularly within the context of ISOBMFF, techniques are needed to signal to users information indicating such preselection (and possibly also information indicating how such preselection should be handled). Summary of the Invention [Problem to be solved by the invention]

[0012] In view of the above, the present disclosure generally provides a method for processing a media stream, a media stream processing device, a program, and a computer-readable storage medium, having the features of the respective independent claims. [Means for solving the problem]

[0013] According to a first aspect of the present disclosure, there is provided a method for processing a media stream. The media stream may be an audio stream, a video stream, or a combination thereof. The method may be performed at a user end or, in some cases, in a user end (decode end) environment, which may include, but is not limited to, a TV, a sound bar, a web browser, a media player, a plug-in, etc., depending on various implementations.

[0014] In particular, the method may include receiving a packetized media stream according to a predefined transport format. The predefined transport format may be the Base Media File Format (ISOBMFF), as defined by ISO / IEC 14496-12 MPEG-4 Part 12, or any other suitable transport format. The packetized media stream may include multiple hierarchical boxes, each associated with a respective box type identifier. In particular, as used herein, the term "box" may be used generally to refer to an object-oriented building block, defined in some possible cases by a unique box type identifier (and possibly a respective length), as described in ISO 14496-12, for example. Of course, the term "box" used throughout this disclosure should not be understood to be limited solely to such specifications. Rather, the term "box" shall generally be understood as any suitable data structure that can serve as a placeholder for media data or other data of the packetized media stream. Furthermore, as will also be understood and appreciated by those skilled in the art, such "boxes" may be referred to using any other suitable term. As an example, in some possible specifications (including the first definition of MP4), a "box" may, in some cases, be alternatively referred to as an "atom." Furthermore, as shown ("hierarchical"), multiple boxes may be at the same or different levels (or positions), nested (child / sub-boxes and parent box), etc., depending on various implementations and / or requirements, as will be understood and appreciated by those skilled in the art. More specifically, multiple boxes may contain, among other possibilities, one or more track boxes that reference (or in other words, indicate) respective tracks that indicate media (content) components of a media stream.Broadly speaking, a media (content) component can generally refer to a single / individual continuous component of media content (and can also be associated with a corresponding (e.g., assigned) media content component type, which may typically include, but is not limited to, audio, video, text, etc.). Examples for understanding the concept of "media (content) component" can be found, for example, as defined / explained in MPEG-DASH (ISO / IEC 23009-1).

[0015] The method may further include determining whether the media stream includes a preselection-related box of a predefined type indicating a preselection, which may correspond to a media presentation to a user. More specifically, as used herein, the term / phrase "preselection" is generally used to refer to a set of media content components (of a media stream) that are intended to be consumed together (e.g., by a user-side device) and, more specifically, generally represent one version of a media presentation that may be selected by an end user for simultaneous decoding and / or presentation. Examples for understanding the concept of "preselection" may be found, for example, as described in MPEG-DASH (ISO / IEC 23009-1). Of course, as will be understood and appreciated by those skilled in the art, in some other possible technical contexts, the term "preselection" may be known (or referred to) by using any other appropriate (equivalent) term, such as, but not limited to, "presentation" as described in, for example, ETSI TS 103190-2, or "preset" as described in, for example, ISO / IEC 23008-3. Thus, a preselection-related box may be a specific box among multiple boxes in a media stream of a specific predefined (or predetermined) type. Such a specific type indicating preselection may be predefined (or predetermined) in advance by using any appropriate means, as will be understood and appreciated by those skilled in the art, which will be described in more detail below.

[0016] If it is determined that the media stream includes a preselection-related box, the method may further include analyzing metadata information corresponding to the preselection-related box, where the metadata information indicates characteristics of the preselection; identifying one or more tracks in the packetized media stream that contribute to the preselection based on the metadata information; and providing the one or more tracks for downstream processing according to the given preselection. As will be understood and appreciated by those skilled in the art, the metadata information presented above may be included (directly) in the media stream (or more specifically, in multiple boxes of the (packetized) media stream) or may be derived (indirectly) from the media stream by using any appropriate means, depending on various implementations. For example, the metadata information may be included or contained in a header box (or a sub-box of another box) that may be associated with or linked to the preselection-related box (e.g., as a sub-box thereof). As mentioned above, preselection generally refers to when a set of media content components are intended to be consumed together, for example, by one or more appropriate downstream devices (e.g., media decoders, media players, etc.). The downstream devices may, in some cases, simply be referred to as "sinks." As a result, depending on various implementations and / or requirements, downstream processing may include, but is not limited to, multiplexing (or re-multiplexing, in some cases), ordering, merging, decoding, or rendering of those contributing tracks, as described in more detail below.

[0017] When configured as described above, the proposed method generally provides an efficient yet flexible way to determine / identify and subsequently signal tracks within a media stream that are configured to contribute to a particular preselection, thereby enabling further appropriate downstream processing of such contributing tracks (e.g., by one or more appropriate downstream devices). Thus, broadly speaking, the proposed method may be considered to provide the possibility and capability of signaling information indicating preselection (and potentially its processing) in a transport layer file (e.g., ISOBMFF) in a uniform manner, which may be considered beneficial in various use cases or scenarios. For example, such a uniform (or, in other words, format / type-agnostic) representation of preselection may be used to implement a uniform data structure for a media playback application programming interface (API) (e.g., used by an application or as a plug-in for a web browser) so that format-specific implementation in the media player is not required, thereby requiring less implementation and / or testing effort and, at the same time, increasing reliability. As another example, such a unified representation of preselection also enables a format-agnostic implementation of preselection data processing in manifest (e.g., MPEG Dynamic Adaptive Streaming over HTTP (DASH) format file, or HTTP Live Stream (HLS) format file) generators, thereby avoiding the need for computationally more expensive operations on binary data, again reducing implementation effort and increasing reliability. As used herein, the term format / type agnostic can generally mean that the proposed representation is generic for all data types (formats).

[0018] In some example implementations, the media stream may further include processing information that indicates how the tracks contributing to the preselection should be processed (e.g., by a downstream device). Similar to the metadata information described above, the processing information may be included (directly) in the media stream (or, more specifically, in the boxes of the (packetized) media stream) or may be derivable (indirectly) from the media stream by using any suitable means, depending on various implementations. For example, the processing information may be included in one particular box (e.g., of a particular (predefined) type), which may be associated with or linked to (e.g., as a sub-box of) a preselection or preselection-related box.

[0019] In some example implementations, the processing information may include ordering information that indicates a track order for processing (e.g., decoding, merging, etc.) the one or more tracks. For example, in some possible cases, the track order may indicate in which order the tracks should be provided to a downstream device (e.g., a decoding device). Similar to the above, such ordering information may be implemented as being included in a (sub)box that is associated with or linked to the processing information.

[0020] In some example implementations, the processing information may include merge information indicating whether one or more tracks should be merged with one or more other tracks for joint (downstream) processing. That is, depending on the implementation of such merge information, in some cases, some tracks may be merged with some other tracks for downstream processing, and in some other cases, some tracks may be treated separately (e.g., routed to individual decode instances). Notably, in some possible cases, such merging may be referred to as multiplexing or any other appropriate term. As will be understood and appreciated by those skilled in the art, track merging (multiplexing) may be achieved by using any appropriate means (e.g., by appending a subsequent track to the end of a preceding track).

[0021] In some example implementations, the method may further include merging the one or more tracks according to the merge information and the ordering information.

[0022] In some example implementations, the ordering information may include, for each track contributing to the preselection, a respective track order value for defining the track order of the tracks. As will be understood and appreciated by those skilled in the art, various appropriate rules may be determined for defining the track order of the tracks by using the respective track order values. As described above, in some possible cases, the track order may indicate in what order the tracks should be presented to a downstream device (e.g., a decoding device). In such a case, one possible example implementation (but not by way of limitation of any kind) may be that a track with a smaller track order value (e.g., 1) is presented to the decoding device earlier than another track with a larger track order value (e.g., 3). In some cases, when multiple tracks have the same track order value, the ordering of those tracks may no longer be significant or important. Additionally, the merge information may similarly include a respective merge flag for each track contributing to the preselection. In particular, a first setting of the merge flag (e.g., “1”) may indicate that the respective track should be merged (or multiplexed) with an adjacent track in track order (e.g., a preceding or following track, depending on various implementations of such merge flag), and a second setting of the merge flag (e.g., “0”) may correspondingly indicate that the respective track should be processed separately (e.g., provided or routed to separate downstream decoding devices). Thus, merging the one or more tracks according to the merge information and ordering information may include successively (or sequentially) scanning the tracks according to the track order and merging the tracks according to their respective merge flags.For example, in some possible cases, if the merge flag flag[i] of track i is set to '1', each sample of this track i may be appended to the sample(s) of the track with the next lower (or higher) track order value (e.g., track i-1 or i+1), while if the merge flag of a track is set to '0', this track i may be fed to a separate decoder instance. As another possible example, in the extreme case where all merge flags for tracks are set to '0', applying the above concept, all tracks will be distributed to several (separate) downstream devices (sinks).

[0023] In some example implementations, the method may further include decoding the one or more tracks for playback of the media stream according to a media presentation indicated by the preselection.

[0024] In some example implementations, the one or more tracks may be decoded by a downstream device (e.g., a media player, a TV, a sound bar, a plug-in, etc.).

[0025] In some example implementations, the merging of the one or more tracks and the decoding of the one or more tracks may be performed by one single device. In other words, there may be use cases in which merging and decoding work in tandem. More specifically, an example (without limitation) of such a use case may be that a JavaScript API expects to work only with a single merged stream, not multiple streams. In such cases, one entity generally accepts multiple incoming streams, merges / multiplexes them as shown above, and then sends them as one merged stream through a single-stream API for decoding on the other side of the API. In some possible cases, the single-stream format may be the Common Media Application Format (CMAF) byte stream format. Of course, as will be understood and appreciated by those skilled in the art, in some other cases, the merging and decoding of tracks may be performed by different (separate) devices. As an illustrative example (and not by way of limitation), a TV may implement stream merging as shown above, but may send the merged stream to a separate downstream device, such as a soundbar, for subsequent decoding.

[0026] In some example implementations, the media stream may include multiple (not just one) preselection-related boxes of a predefined type. Thus, the method may further include selecting (or determining) a preselection-related box from among the multiple preselection-related boxes. As will be understood and appreciated by those skilled in the art, such selection (or determination) of one particular preselection-related box from among the multiple preselection-related boxes may be performed by any suitable means.

[0027] In some example implementations, the preselection-related box may be selected (or determined) by an application (e.g., an application controlling a media player / decoder). For example, in some possible cases, said application may be configured (e.g., based on a predefined algorithm) to (automatically) select (or determine) a preselection-related box that corresponds to a particular setup (e.g., of a decoding or presentation environment).

[0028] In some example implementations, a media stream may include one or more label boxes (which may be associated or linked in some way to a respective preselection or a respective preselection-related box) that each contain descriptive information for a respective media presentation to a user corresponding to a respective preselection. Thus, in such cases, the selection (or determination) of a preselection-related box may be performed based on user input. As an example (and not by way of limitation), label boxes may contain descriptive information indicating selectable subtitles in various languages ​​(e.g., English, German, Chinese, etc.), each of which may be considered a respective preselection (presentation), whereby a user (e.g., controlling an application) may select the corresponding language setting accordingly (e.g., by clicking the keyboard or mouse).

[0029] In some example implementations, the preselection-related boxes may be considered agnostic to the media codec used to encode the media stream before it was packetized. That is, generally speaking, the preselection-related boxes may contain only the information necessary for the respective preselection and may not contain any information that may be linked to a media codec (i.e., codec-specific information). In other words, there is generally no information in the preselection-related box(es) about how the media stream was encoded (e.g., by a particular media encoder) and / or how such media stream should be decoded (e.g., by a particular media decoder).

[0030] In some example implementations, the metadata information corresponding to the preselection-related box may include track identification information indicating one or more track identifiers respectively associated with respective tracks, where the tracks associated with said one or more track identifiers in the metadata information may be relevant to the media presentation. As will be understood and appreciated by those skilled in the art, such track identification information indicating one or more track identifiers may be implemented in any suitable manner, for example, as simply as an array, each element of which (uniquely) indicates a respective track identifier (which itself may be represented by using an integer value or any other suitable form). In that case, in some possible cases, the metadata information corresponding to the preselection-related box may optionally further include a counter (e.g., an integer value) indicating the number of tracks required by (or contributing to) the preselection.

[0031] In some example implementations, the metadata information corresponding to the preselection-related box may include preselection identification information indicating a preselection identifier for identifying the preselection. That is, the metadata information corresponding to the preselection-related box may include necessary information (e.g., represented by using an integer) that enables the preselection to be (uniquely) identifiable to external (e.g., downstream) applications and / or devices, for example, to appropriately assist in the selection / decision of the respective preselection.

[0032] In some example implementations, the metadata information corresponding to a preselection-related box may include unique preselection-specific data for configuring a downstream device (e.g., a downstream media player / decoder) to decode the track according to the preselection. Depending on various implementations and / or requirements, such preselection-specific data may include any suitable information (e.g., codec-specific information in some cases) and may be implemented (represented) in any suitable manner (e.g., integers, arrays, strings, etc.).

[0033] According to a second aspect of the present disclosure, there is provided a method for processing a media stream. The media stream may be an audio stream, a video stream, or a combination thereof. The method may be performed at a user end or, in some cases, in a user end (decode end) environment, which may include, but is not limited to, a TV, a sound bar, a web browser, a media player, a plug-in, etc., depending on various implementations.

[0034] In particular, the method may include receiving a packetized media stream according to a predefined transport format. Similar to the first aspect discussed above, the predefined transport format may be the Base Media File Format (ISOBMFF) as defined by ISO / IEC 14496-12 MPEG-4 Part 12, or any other suitable transport format. The packetized media stream may include multiple hierarchical boxes, each associated with a respective box type identifier. As illustrated above, and as will be understood and appreciated by those skilled in the art, the multiple boxes (or may be referred to using any other appropriate term) may be at the same or different levels (or positions), nested (child / sub-boxes and parent box), etc., depending on various implementations and / or requirements. More specifically, the multiple boxes may include, among other possibilities, one or more track boxes that reference (or in other words, indicate) respective tracks that represent media (content) components of the media stream. In addition, the boxes may also include one or more Track Group boxes (or may be referred to using any other suitable term / name), each associated with a respective pair of Track Group Identifier and Track Group Type that jointly identify a respective Track Group within the media stream. That is, tracks having (e.g., identified by or associated with) the same Track Group Identifier and the same Track Group Type may be considered to belong to the same Track Group. Each such Track Group may generally determine a respective preselection corresponding to a respective media presentation to a user.As already indicated above, the term / phrase preselection (or referred to by using any other appropriate term / name) is generally used to refer to a set of media content components (of a media stream) that are intended to be consumed together (e.g., by a user-side device) and, more specifically, that generally represent one version of a media presentation that can be selected by an end-user for simultaneous decoding / presentation.

[0035] The method may further include checking (e.g., visiting, cycling through, etc.) track boxes in the media stream to determine the full (or complete / entire) set of preselections present in the media stream. In particular, determining the full set of preselections may include: determining a set of unique pairs of track group identifiers and track group types, and addressing the preselections by their respective track group identifiers. As described above, each preselection is associated with a respective track group, which itself is identified by a respective pair of corresponding track group identifier and corresponding track group type. Thus, a preselection may be addressed (or identified) by its associated / linked respective track group identifier.

[0036] The method may further include selecting a preselection from the complete set of preselections. In particular, the preselection may be selected based on attributes (e.g., expressed as metadata or any other suitable format) of each preselection included in a track group box having the same track group identifier.

[0037] The method may also include determining (e.g., identifying) a set of one or more track boxes that contribute to the selected preselection. In particular, a set of one or more track boxes that contribute to the (same) preselection may be determined (identified) by the presence of (respective) track group boxes having the same track group identifier.

[0038] Additionally, the method may further include determining the tracks referenced in each member of the set of one or more track boxes determined above as one or more tracks contributing to the preselection.

[0039] Finally, the method may include providing the one or more tracks for downstream processing according to a preselection. As discussed above, a preselection generally refers to a set of media content components intended to be consumed together by one or more appropriate downstream devices (or, in some possible cases, referred to as sinks), such as, for example, a media decoder, a media player, etc. As a result, depending on various implementations and / or requirements, the downstream processing may include, but is not limited to, multiplexing (or, in some possible cases, remultiplexing), ordering, merging, decoding, or rendering of the contributing tracks, as described in more detail below.

[0040] When configured as described above, the proposed method generally provides an efficient yet flexible way to determine / identify and subsequently signal tracks within a media stream configured to contribute to a particular preselection, thereby enabling further appropriate downstream processing of such contributing tracks (e.g., by one or more downstream devices). In particular, it should be noted that the method proposed above in the first aspect generally seeks to provide relevant information for (all) tracks configured to contribute to a particular preselection in a preselection-related box, thereby enabling indexing (or identification) of all contributing tracks. In this sense, track indexing as described in the first aspect can be considered a type of forward (direct) indexing. In contrast, in the method proposed in this second aspect, the tracks contributing to a particular preselection can be jointly determined by a pair of track group type and track group identifier. More specifically, tracks having (e.g., including) a track group box with a particular (e.g., predefined or predetermined) track group type may generally indicate that the track contributes to preselection. Furthermore, tracks having the same track group identifier may generally indicate that the tracks belong to (contribute to) the same preselection. In this sense, such indexing of tracks described here, as opposed to that proposed in the first aspect, may be considered a kind of reverse indexing. In any case, similar to the first aspect, the method proposed in the second aspect may also provide the possibility and ability to signal information indicating preselection (and potentially its processing) in a transport layer file (e.g., ISOBMFF) in a unified way. This may be considered beneficial in various use cases or scenarios.For example, such a unified (or, in other words, format-agnostic) representation of preselection may be used to implement a unified data structure for a media playback API (e.g., used by an application or as a plug-in for a web browser) so that format-specific implementations in media players are not required, thereby requiring less implementation and / or testing effort and at the same time increasing reliability. As another example, such a unified representation of preselection also enables a format-agnostic implementation of preselection data processing in manifest (e.g., MPEG Dynamic Adaptive Streaming over HTTP, DASH, format file, or HTTP Live Stream, HLS, format file) generators, thereby avoiding the need for computationally more expensive operations on binary data, again reducing implementation effort and increasing reliability.

[0041] In some example implementations, each preselection may be associated with a respective preselection-related box of a predefined type. In particular, a preselection-related box may instantiate (e.g., implicitly, extend, etc.) a track group box having a predefined track group type associated with the preselection. Such a predefined track group type for a preselection may be implemented by using any appropriate means, such as a specific (predefined) string (e.g., "preselection" or "pres"), a specific (predefined) value (e.g., "3"), etc., as will be understood and appreciated by those skilled in the art. Generally speaking, in the method proposed in the second aspect, there is typically one such preselection-related box per preselection and per track corresponding to (contributing to) that preselection. By comparison, in the method proposed in the aforementioned first aspect, there is typically one such preselection-related box per preselection.

[0042] In some example implementations, a preselection-related box may be associated with a preselection processing box that contains processing information indicating how the tracks contributing to the preselection should be processed. The processing information may be included (directly) in the media stream (or more specifically, in multiple boxes of the (packetized) media stream) or may be derived (indirectly) from the media stream by using any suitable means, depending on various implementations. For example, the processing information may be included in a specific box (e.g., of a specific (predefined) type) that may be associated with or linked to the preselection or preselection-related box (e.g., as a sub-box thereof).

[0043] In some example implementations, the preselection association box may be associated with a preselection information box that contains semantic (or descriptive) information that indicates the preselection (eg, attributes, characteristics, etc.).

[0044] In some example implementations, the processing information may include unique preselection-specific data for configuring a downstream device (e.g., a downstream media player / decoder) to decode the track according to the preselection. Depending on various implementations and / or requirements, such preselection-specific data may include any suitable information (e.g., codec-specific information in some cases) and may be implemented (represented) in any suitable manner (e.g., integers, arrays, strings, etc.).

[0045] In some example implementations, the processing information may include ordering information that indicates a track order for ordering the tracks for further downstream processing (e.g., decoding, merging, etc.). For example, in some possible cases, the track order may indicate in what order the tracks should be provided to a downstream device (e.g., a decoding device). Similar to the above, such ordering information may be implemented as being included in a (sub)box associated with or linked to the processing information.

[0046] In some example implementations, the processing information may include merging information indicating whether one or more tracks should be merged with one or more other tracks, e.g., for joint (downstream) processing. That is, depending on the implementation of such merging information, in some cases, some tracks may be merged with some other tracks for downstream processing, and in some other cases, some tracks may be treated separately (e.g., routed to individual decode instances). Notably, in some possible cases, such merging may also be referred to as multiplexing, or any other appropriate term. As will be understood and appreciated by those skilled in the art, track merging (multiplexing) may be achieved by using any appropriate means (e.g., by appending a subsequent track to the end of a preceding track).

[0047] In some example implementations, the method may further include merging the one or more tracks according to the merge information and the ordering information.

[0048] In some example implementations, the ordering information may include, for each track contributing to the preselection, a respective track order value for defining the track order of the tracks. As will be understood and appreciated by those skilled in the art, various appropriate rules may be determined for defining the track order of the tracks by using the respective track order values. As described above, in some possible cases, the track order may indicate in what order the tracks should be presented to a downstream device (e.g., a decoding device). In such a case, one possible example implementation (but not by way of limitation of any kind) may be that a track with a smaller track order value (e.g., 1) is presented to the decoding device earlier than another track with a larger track order value (e.g., 3). In some cases, when multiple tracks have the same track order value, the ordering of those tracks may no longer be significant or important. Additionally, the merge information may similarly include a respective merge flag for each track contributing to the preselection. In particular, a first setting of the merge flag (e.g., “1”) may indicate that the respective track should be merged (or multiplexed) with an adjacent track in track order (e.g., a preceding or following track, depending on various implementations of such merge flag), and a second setting of the merge flag (e.g., “0”) may correspondingly indicate that the respective track should be processed separately (e.g., provided or routed to separate downstream decoding devices). Thus, merging the one or more tracks according to the merge information and ordering information may include successively (or sequentially) scanning the tracks according to the track order and merging the tracks according to their respective merge flags.For example, in some possible cases, if the merge flag flag[i] of track i is set to '1', each sample of this track i may be appended to the sample(s) of the track with the next lower (or higher) track order value (e.g., track i-1 or i+1), while if the merge flag of a track is set to '0', this track i may be fed to a separate decoder instance. As another possible example, in the extreme case where all merge flags for tracks are set to '0', applying the above concept, all tracks will be distributed to several (separate) downstream devices (sinks).

[0049] In some example implementations, the method may further include decoding the one or more tracks for playback of the media stream according to a media presentation indicated by the preselection.

[0050] In some example implementations, the one or more tracks may be decoded by a downstream device (e.g., a media player, a TV, a sound bar, a plug-in, etc.).

[0051] In some example implementations, the merging of the one or more tracks and the decoding of the one or more tracks may be performed by one single device. In other words, there may be use cases in which merging and decoding work in tandem. More specifically, an example (without limitation) of such a use case may be that a JavaScript API only expects to work with a single merged stream, not multiple streams. In such a case, one entity generally accepts multiple incoming streams, merges / multiplexes them as shown above, and then sends them as one merged stream through a single-stream API for decoding on the other side of the API. In some possible cases, the single-stream format may be the Common Media Application Format (CMAF) byte stream format. Of course, as will be understood and appreciated by those skilled in the art, in some other cases, the merging and decoding of tracks may be performed by different (separate) devices. As an illustrative example (and not by way of limitation), a TV may implement stream merging as shown above, but may send the merged stream to a separate downstream device, such as a soundbar, for subsequent decoding.

[0052] In some example implementations, the preselection may be selected (or determined) by an application (e.g., an application controlling a media player / decoder). For example, in some possible cases, the application may be configured (e.g., based on a predefined algorithm) to (automatically) select (or determine) a preselection that corresponds to, for example, a particular setup (e.g., of a decoding or presentation environment).

[0053] In some example implementations, a media stream may include one or more label boxes (which may be associated or linked to a respective preselection or preselection-related box) each containing descriptive information for a respective media presentation to a user corresponding to a respective preselection. Thus, in such cases, preselection selection (or determination) may be performed based on user input. As an example (and not by way of limitation), a label box may contain descriptive information indicating selectable subtitles in various languages ​​(e.g., English, German, Chinese, etc.), each of which may be considered a respective preselection (presentation), thereby allowing a user (e.g., controlling an application) to select the corresponding language setting accordingly (e.g., by clicking a keyboard or mouse). Of course, as will be understood and appreciated by those skilled in the art, preselection selection (or determination) may also be performed by any other suitable means.

[0054] In some example implementations, the media stream may include at least one of an audio stream or a video stream (or a combination thereof), as indicated above. In particular, some possible scenarios (or use cases) to which the methods proposed in this disclosure may be applied may be, for example, a multi-person video conference (where the viewer may have the ability to select one or more video streams), or a TV with picture-in-picture capabilities (e.g., with one picture having a higher bitrate / resolution and another picture having a lower bitrate / resolution).

[0055] According to a third aspect of the present disclosure, there is provided a method for processing a media stream. The media stream may be an audio stream, a video stream, or a combination thereof. The method may be performed on an encoder-side environment (e.g., a media encoder). In some scenarios (or use cases), such an encoder may also be referred to as a (media) packager (i.e., configured to pack / packetize media input).

[0056] In particular, the method may include encapsulating one or more elementary streams according to a predefined transport format to generate a packetized media stream, the packetized media stream including a plurality of hierarchical boxes, each associated with a respective box type identifier. As above, the predefined transport format may be the ISO Base Media File Format (ISOBMFF), as defined by ISO / IEC 14496-12 MPEG-4 Part 12, or any other suitable transport format. Broadly speaking, as will be understood and appreciated by those skilled in the art, elementary streams may be considered to form a compressed binary representation of media data flowing from a single media encoder to a media decoder (either audio or video). When multiplexed (or packetized) into a predefined transport format (e.g., ISOBMFF), these elementary streams may be referred to as "tracks," and "track boxes" describe the properties (or attributes) of each track in the file header. Additionally, as also noted above, multiple boxes (or referred to by using any other appropriate term) may be at the same or different levels (or positions), nested (child / sub-boxes and parent box), etc., depending on various implementations and / or requirements, as will be understood and appreciated by those skilled in the art.

[0057] More specifically, encapsulating the one or more elementary streams may include packetizing media data of the one or more elementary streams in accordance with a transport format to generate one or more track boxes referencing (or indicating) respective tracks of the one or more elementary streams; and generating one or more preselection-related boxes of a predefined type based on header information of the one or more elementary streams, each of the one or more preselection-related boxes indicating a respective preselection corresponding to a media presentation to a user.

[0058] When configured as described above, the proposed method may generally provide an efficient yet flexible way to packetize a media input (e.g., an elementary stream) according to a predefined transport format (e.g., ISOBMFF). More specifically, by generating and including one or more preselection-related boxes (each of which indicates a respective preselection) along with the packetized media stream, the proposed method may also enable representing the preselection in a codec-agnostic and uniform manner, thereby further enabling appropriate downstream processing of the tracks contributing to the corresponding preselection (e.g., according to the methods proposed in the first and second aspects above). Furthermore, as shown above, such a unified representation of preselection also enables a format-agnostic implementation of preselection data processing in manifest (e.g., MPEG Dynamic Adaptive Streaming over HTTP, DASH, format file, or HTTP Live Stream, HLS, format file) generators, thereby avoiding the need for computationally more expensive operations on binary data and at the same time reducing implementation effort and increasing reliability.

[0059] In some example implementations, each of the one or more preselection-related boxes may include metadata information indicating characteristics of the respective preselection. More specifically, the metadata information may include information indicating one or more tracks within the media stream that contribute to the respective preselection. As will be understood and appreciated by those skilled in the art, the metadata information presented above may be included (directly) in the media stream (or, more specifically, in boxes of the (packetized) media stream) or may be derived (indirectly) from the media stream by using any appropriate means, depending on various implementations. For example, the metadata information may be included / contained in a header (or other) box that may be associated with or linked in some way to the preselection-related box (e.g., as a sub-box thereof).

[0060] In some example implementations, the metadata information corresponding to each preselection-related box may further include at least one of preselection identification information indicating a preselection identifier for identifying the respective preselection, or unique preselection-specific data for decoding the track according to the preselection (e.g., for configuring a downstream device such as a downstream media decoder).

[0061] In some example implementations, encapsulating the elementary media stream may further include generating one or more track group boxes, each associated with a respective track group identifier and a respective track group type (pair) that jointly identify a respective track group in the packetized media stream. In particular, tracks having the same track group identifier and the same track group type may be considered to belong to the same track group. Furthermore, generating the one or more preselection association boxes may include: assigning a first unique identifier to each preselection; and, for each track contributing to the respective preselection, generating a respective preselection association box associated with the respective preselection, the preselection association box having a predefined track group type associated with the preselection, instantiating a track group box that sets the track group identifier to the first unique identifier. Such predefined track group types associated with a preselection may be implemented by any suitable means, as will be understood and appreciated by those skilled in the art, such as by using a specific (predefined) string (e.g., "preselection" or "pres"), a specific (predefined) value (e.g., 3), etc. Generally speaking, in the method proposed here, there will generally be one such preselection-related box per preselection and per track that corresponds (contributes) to that preselection.

[0062] In some example implementations, track group boxes may be created by grouping tracks that contribute to a single preselection based on the tracks' respective media types.

[0063] In some example implementations, the media type may include at least one of audio, video, and subtitles, although, of course, any other suitable media type may be used as well, as will be understood and appreciated by those skilled in the art.

[0064] In some example implementations, generating the one or more preselection-related boxes may further include generating one or more preselection processing boxes that include processing information indicating how the tracks contributing to the respective preselection should be processed. The processing information may be included (directly) in the media stream (or more specifically, in boxes of the (packetized) media stream) or may be derived (indirectly) from the media stream by using any suitable means, depending on various implementations. For example, the processing information may be included in a specific box (e.g., of a specific (predefined) type) that may be somehow associated or linked to the preselection or preselection-related box (e.g., as a sub-box thereof).

[0065] In some example implementations, the processing information may include at least one of ordering information indicating a track order for processing the tracks, or merge information indicating whether one or more tracks should be merged with one or more other tracks (e.g., for joint (downstream) processing). For example, in some possible cases, the track order may indicate in which order the tracks should be provided to a downstream device (e.g., a decoding device). Furthermore, depending on the implementation of the merge information, in some cases, some tracks may be merged with some other tracks for downstream processing; while in some other cases, some tracks may be treated separately (e.g., routed to individual decoding instances). Similar to the above, the ordering information as well as the merge information may be implemented as being contained in a (sub)box (or more) associated or linked to the processing information.

[0066] In some example implementations, the method may further include receiving at least one input media; and processing (e.g., encoding) the input media to generate one or more elementary streams, the one or more elementary streams including media data of the input media and corresponding header information. For example, the input media may be processed (e.g., encoded) by using a suitable media encoder to generate the corresponding elementary streams accordingly.

[0067] In some example implementations, the method may further include generating a manifest file based on the one or more preselection-related boxes. Generally speaking, a manifest file typically includes various information, such as information about the media stream (e.g., media type, codec attributes, media-specific properties, etc.). Apart from such information (i.e., information about the media stream itself), the proposed method may further include preselection-related information. The preselection-related information may include, but is not limited to, metadata information, processing information, etc. Compared to conventional techniques, the proposed method for generating a manifest file generally provides a format-agnostic implementation for preselection-related data processing, thereby avoiding computationally more expensive operations (e.g., operations that must be performed on binary data in conventional techniques), reducing implementation and / or testing effort, and increasing reliability.

[0068] In some example implementations, the manifest file may be an MPEG Dynamic Adaptive Streaming over HTTP (DASH) format file, an HTTP Live Stream (HLS) format file, or any other suitable manifest format file, as will be understood and appreciated by those skilled in the art.

[0069] According to a fourth aspect of the present disclosure, there is provided a method for processing a media stream, the media stream may be an audio stream, a video stream, or a combination thereof, the method may be performed by, for example, a manifest generator.

[0070] Specifically, the method can include receiving a packetized media stream according to a predefined transport format. Specifically, the packetized media stream can include a plurality of hierarchical boxes, each associated with a respective box type identifier, the plurality of boxes can include one or more track boxes that reference (e.g., indicate) respective tracks that indicate media components of the media stream, and one or more preselection-related boxes of a predefined type, each preselection-related box indicating a respective preselection corresponding to a media presentation to a user.

[0071] Additionally, the method can include generating a manifest file based on the one or more preselection association boxes.

[0072] Configured as described above, the proposed method can generally provide an efficient yet flexible method for generating a manifest file by also taking into account preselection-related information (e.g., description information and / or processing-related information). More specifically, apart from information about the media stream, the proposed method can further include preselection-related information. The preselection-related information can include, but is not limited to, metadata information, processing information, etc. Compared to conventional (manifest generation) techniques, the proposed method for generating a manifest file generally provides a format-agnostic implementation for preselection-related data processing, thereby avoiding computationally more expensive operations (e.g., operations that must be performed on binary data in conventional techniques), reducing implementation and / or testing effort, and increasing reliability.

[0073] In some example implementations, the manifest file may be an MPEG Dynamic Adaptive Streaming over HTTP (DASH) format file, an HTTP Live Stream (HLS) format file, or any other suitable manifest format file, as will be understood and appreciated by those skilled in the art.

[0074] In some example implementations, media presentation to a user may be characterized by a respective configuration of the language, type, and / or one or more media-specific attributes of the media stream. Of course, any other suitable configuration may also be used, depending on various implementations.

[0075] In some example implementations, the predefined transport format may be the ISO Base Media File Format (ISOBMFF), or any other suitable transport format.

[0076] According to a fifth aspect of the present invention, there is provided a media stream processing device including a processor and a memory coupled to the processor, the processor being adapted to cause the media stream processing device to perform all steps according to any of the exemplary methods described in the previous aspects.

[0077] According to a sixth aspect of the present invention, there is provided a computer program, which may include instructions that, when executed by a processor, cause the processor to perform all of the steps of the methods described throughout the present disclosure.

[0078] According to a seventh aspect of the present invention, there is provided a computer-readable storage medium, which may store the computer program described above.

[0079] It will be understood that apparatus features and method steps can be interchanged in many ways. In particular, details of the disclosed methods can be implemented by a corresponding apparatus (or system), and vice versa, as will be understood by those skilled in the art. Furthermore, it will be understood that any statements made above with respect to a method(s) apply equally to a corresponding apparatus (or system), and vice versa. [Brief explanation of the drawings]

[0080] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which like reference numbers indicate like or similar elements. [Figure 1] 1A and 1B are schematic diagrams illustrating an example implementation of a media playback application programming interface (API), according to an embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram illustrating an exemplary implementation of a system for processing media streams, according to an embodiment of the present disclosure. [Figure 3]1A and 1B are schematic diagrams illustrating an exemplary implementation of a packager according to an embodiment of the present disclosure; [Figure 4] 1A and 1B are schematic diagrams illustrating an exemplary implementation of a manifest generator according to an embodiment of the present disclosure; [Figure 5] 1 is a schematic flowchart illustrating an example of a method for processing a media stream, according to an embodiment of the present disclosure. [Figure 6] 10 is a schematic flowchart illustrating another example of a method for processing a media stream, according to an embodiment of the present disclosure. [Figure 7] 10 is a schematic flowchart illustrating a further example of a method for processing a media stream, according to an embodiment of the present disclosure. [Figure 8] 10 is a schematic flowchart illustrating yet another example of a method for processing a media stream, according to an embodiment of the present disclosure. [Figure 9] 1 is a schematic block diagram of an exemplary apparatus for performing a method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0081] As mentioned above, the same or similar reference numerals in this disclosure, unless otherwise indicated, indicate the same or similar elements, and repeated descriptions thereof may be omitted for reasons of brevity.

[0082] In particular, the drawings and the following description relate to preferred embodiments by way of example only. It should be noted from the following discussion that alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of what is claimed.

[0083] Reference will now be made in detail to several embodiments, examples of which are illustrated in the accompanying drawings. It should be noted that wherever practical, like or similar reference numerals may be used in the figures and may indicate like or similar functionality. The drawings depict embodiments of the disclosed system (or method) for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be used without departing from the principles described herein.

[0084] Furthermore, when connecting elements such as solid or dashed lines or arrows are used in the figures to indicate a connection, relationship, or association between two or more other schematic elements, the absence of such connecting elements is not intended to imply that the connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements may not be shown in the drawings so as not to obscure the invention. Additionally, for ease of explanation, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, when a connecting element represents communication of signals, data, or instructions, those skilled in the art will understand that such element represents one or more signal paths for affecting the communication, as appropriate.

[0085] As indicated above, in a broad sense, this paper seeks to propose techniques that enable the concept of preselection generally within the context of ISOBMFF files, thereby enabling full usability in MPEG-DASH and various other media compression formats.

[0086] Thus, the exemplary embodiments described herein may relate to methods, apparatus, and processes directed to various "preselection" use cases.

[0087] Before going into details for describing exemplary embodiments, for ease of understanding, it may still be worthwhile to explain some of the possible terms that may be used throughout this disclosure, some of which may already be explained (briefly) above.

[0088] In particular, for the purposes of this document, with respect to "preselection," such term generally refers to the simultaneous decoding / presentation of a set of one or more media (content) components, as described, for example, in MPEG-DASH (ISO / IEC 23009-1). As will be understood and recognized by those skilled in the art, in some other possible technical contexts, "preselection" may also be known (or referred to) using other appropriate (equivalent) terms, such as, but not limited to, "presentation," as described, for example, in ETSI TS 103190-2, or "preset," as described, for example, in ISO / IEC 23008-3.

[0089] With respect to "Adaptation Set," such term shall generally be defined as equivalently set forth in MPEG-DASH (ISO / IEC 23009-1), and may also be known (referred to) by other terms, such as "Switching Set" as defined in the Common Media Application Format (CMAF).

[0090] With respect to "Representation," such term shall generally be defined as described in MPEG-DASH (ISO / IEC23009-1) and may also be known as "AudioTrack" (e.g., as defined by the W3C API) or "Tracklist" or "Track" (e.g., as defined by ISOBMFF).

[0091] With respect to a "media content component," such phrase shall generally mean a single contiguous component of media content that has an assigned media content component type, for example as described in MPEG-DASH (ISO / IEC 23009-1).

[0092] With respect to a "media content component type," such phrase is generally intended to mean a single type of media content, such as that described in MPEG-DASH (ISO / IEC 23009-1). As will be understood and appreciated by those skilled in the art, examples of media content component types may include, but are not limited to, audio, video, text, or any other suitable type.

[0093] With regard to "box", such term shall generally mean an object-oriented building block defined by a unique type identifier and length, as described, for example, in ISO 14496-12. As above, note that the term "box" may alternatively be referred to as "atom" in some specifications (including the first definition in MP4), or by any other suitable similar term.

[0094] With regard to "container box", such a phrase is generally intended to mean a box whose sole purpose is to house and group a set of related (sub)boxes, as described, for example, in ISO 14496-12. Nevertheless, it should be noted that a container box is not usually derived from a "full box".

[0095] Finally, with regard to "ISO Base Media File," such term shall generally mean a file that conforms to the file format specified in ISO 14496-12 (i.e., ISOBMFF).

[0096] Therefore, unless expressly stated otherwise, the terms set forth above can be used interchangeably as would be understood and appreciated by one of ordinary skill in the art.

[0097] Before going into the technical details of each exemplary embodiment, it may also be worthwhile to first discuss some possible use cases (scenarios) for preselection, especially from an abstract and broad overall perspective.

[0098] Referring to the drawings, Figure 1A is a schematic diagram illustrating an example implementation of a media playback application programming interface (API). In other words, Figure 1A can be considered an example use case of presentation (preselection) selection in a media player. More specifically, the example of Figure 1A is generally considered one possible implementation for supporting preselection by using conventional techniques.

[0099] In a broader sense, some media content may be delivered using means such as MPEG DASH, which generally utilizes manifest files, but some player frameworks may only consume content in the ISOBMFF format. As a result, media players based on such architectures are required to enable presentation selection based solely on the metadata information contained in the ISOBMFF and, as a result, can no longer rely on signaling from the manifest file(s).

[0100] As shown in FIG. 1A, a media player 1200 receives input media (or files) 1000 to generate a corresponding media output / presentation 1400. The input media 1000 may be a (packetized) container format file or stream (e.g., an ISOBMFF file or stream). The media player 1200 is controlled by an appropriate application 1300. Traditional controls for such an application 1300, and correspondingly, the media player 1200, may include, but are not limited to, "play," "pause," "stop," "fast forward," "rewind," etc. However, with next-generation audio (or similar video technologies), the media player 1200 may be required to present a list of available experiences (or presentations) to the application 1300 and, in return, receive a desired selection from the application. This may generally be implemented in an API 1290 between the media player 1200 and the application 1300.

[0101] Generally speaking, the incoming ISOBMFF 1000 provides (includes) metadata 1010 in a container format-specific representation (e.g., implemented as one or more boxes or in any other suitable form) as well as media elementary streams 1020. In particular, such container format metadata 1010 may generally be configured as byte data that may be codec-specific. Furthermore, the elementary streams 1020 themselves may include header information data (or, in some cases, metadata of some kind) 1030 (which is typically in the form of binary data) as well as coded (or compressed) media data 1040 (which, like the binary metadata 1030, is typically in the form of binary data). As can be understood and appreciated by those skilled in the art, generally, the elementary stream data 1040 may typically be represented as binary information (bit data), and the container data 1010 may typically utilize byte data.

[0102] 1A, the media player 1200 itself may include multiple decoder implementations (or instances) 1220 for many different media formats (or in other words, format or codec specific), and the media player 1200 correspondingly would also be required to provide an appropriate selection element / component (or simply called a selector) 1210 to unpack the elementary streams from the container format data and determine which media decoder 1220 to use. This determination is generally based on the container format metadata 1010, as described above.

[0103] 1A, the preselection-aware decoder 1220 provides (e.g., generates) preselection-related data 1230. Such codec-specific data 1230 must be converted by a corresponding format converter 1240 into a codec-agnostic generalized format 1250 in order to be consumed by an application 1300 via a (typically codec-agnostic) respective API 1290.

[0104] In general, the operations involved in unpacking elementary streams, converting codec / format specific data into a generalized codec-agnostic format, may introduce (undesirable) extra work into the overall procedure (e.g., for implementation, for testing, etc.) and thus may be considered to be to some extent inefficient.

[0105] To address some or all of the issues discussed with respect to Figure 1A, Figure 1B schematically illustrates an exemplary implementation of a media playback API according to an embodiment of the present disclosure. In particular, as previously indicated, identical or similar reference numerals in Figure 1B refer to identical or similar elements in Figure 1A unless otherwise indicated, and repeated descriptions thereof may be omitted for reasons of brevity.

[0106] In particular, compared to FIG. 1A, the media player 1201 shown in FIG. 1B no longer needs to read the respective APIs of each decoder 1221, reformat their output, and provide the data format-independently.

[0107] Instead, using a new data structure 1051 (typically in byte format) related to preselection (as described in more detail below) within the incoming ISOBMFF file or stream 1001, the media player 1201 can expose this data 1051 to the application 1301 in its respective API 1291. The main reason is that such preselection related data 1051 is codec agnostic.

[0108] When configured as proposed, this approach generally eliminates the need to implement multiple format-dependent metadata format converters (e.g., format converter 1240 as shown in Figure 1A). Furthermore, the proposed approach implicitly defines a unified data structure that can be used to present preselection-related metadata (e.g., new data structure 1051) in the media player API 1291. In other words, using the new proposed data structure, a media player is enabled to expose the contained information via an appropriate API to request an application to select a desired presentation (preselection). More specifically, this approach provides a lightweight approach because it works codec-agnostic and does not require parsing into encoded media essences.

[0109] In summary, as will be understood and appreciated by those skilled in the art, the above proposed approach for the Media Player API use case generally provides at least the following advantages: A unified representation for top-level preselection in transport data formats such as ISOBMFF; Unified data structures for media player APIs; No format-specific implementation in the media player is required; · Less effort for implementers, less testing effort, greater reliability and faster uptake in industry.

[0110] Figure 2 is a schematic diagram illustrating an exemplary implementation of a system for processing media streams according to an embodiment of the present invention. In particular, Figure 2 may be viewed as illustrating a possible overall workflow for such a media processing system.

[0111] In particular, as shown in Figure 2, similar to the media player use cases described with respect to Figures 1A and 1B, an (automated) manifest generator 2500 may also need access to the experience (presentation) contained in the incoming media data. Generally, since multiple codecs may have to be handled in such a processing system, these devices 2500 will preferably be operated with codec-agnostic metadata. Furthermore, since a data side path is generally not desirable for robust operation, any metadata structure used may need to be comprehensive and provide all details that should appear in the manifest (or manifest file).

[0112] 2, media (e.g., audio and / or video) encoders 2100-A, 2100-B, and 2100-C generally generate media streams 2200-A, 2200-B, and 2200-C encoded in different formats (e.g., MP3, MP4, AC-4, etc.). These elementary streams 2200-A-C are then provided to packagers 2300-A, 2300-B, and 2300-C for encapsulation into a unified transport format such as ISOBMFF.

[0113] Each of these ISOBMFF files or streams 2400-A, 2400-B, and 2400-C may then be delivered to a respective media player (not shown), but also to a manifest generator 2500 for generating as output a corresponding manifest file 2600. Such files (e.g., MPEG DASH format files or HLS format files) may often be represented in a machine- and / or human-readable format, such as XML (Extensible Markup Language).

[0114] Referring now to FIG. 3A, an exemplary diagram illustrating one possible implementation of a packager (eg, packagers 2300-A-C) is provided.

[0115] In particular, as shown with respect to FIG. 3A, an incoming media elementary stream 3200 (e.g., elementary streams 2400-A-C as illustrated in FIG. 2), consisting of (binary) media header information or metadata 3210 and (binary) encoded / compressed (raw) media information / data 3220, is then packaged into a container format output file or stream (e.g., ISOBMFF) 3400.

[0116] More specifically, in this process, the packager 3300 generally passes the incoming stream data 3200, unchanged, to a corresponding output representation 3420. A suitable internal device (or component) 3310 of the packager 3300 can read the binary header information / metadata 3210 and generate corresponding descriptive metadata information, but in byte format 3410, for the media output 3400. Such a device (or component) 3310 of the packager 3300 may, for example, be a decoder specific information (DSI) generator or any other suitable component. This data 3410 may later be used by a respective media player (e.g., media player 1200 as shown in FIG. 1A) to select an appropriate decoder (e.g., decoder 1220 as shown in FIG. 1A). In the exemplary use case of ISOBMFF, this descriptive byte data 3410 may include, among other things, decoder specific information (DST for short), but may also include some general annotations on the elementary stream. As indicated earlier, elementary streams, when encapsulated in some file format such as ISOBMFF, may be called tracks.

[0117] In comparison with the implementation of Figure 3A, Figure 3B schematically illustrates one possible exemplary implementation of a packager according to an embodiment of the present invention. As above, identical or similar reference numerals in Figure 3B indicate identical or similar elements in Figure 3A unless otherwise indicated, and repeated descriptions thereof may be omitted for reasons of brevity.

[0118] More specifically, as shown with respect to Figure 3B, the packager 3301 may now include a further (internal) device / component 3321 (which may in some possible cases be called a preselection data generator), which is also provided with the binary header information 3211 of the incoming media stream 3201. In some possible cases, the functionality of this device 3321 may be implemented as an additional function as part of the existing (internal) device 3311. However, as clearly shown in Figure 3B, in such cases, an additional byte-format specified data structure 3451 is generated and embedded in the output stream 3400. In particular, as indicated above, this data structure 3451 generally carries (codec-agnostic) preselection-related information (e.g., description information and / or processing-related information) and a media-format-independent representation.

[0119] FIG. 4A is a schematic diagram of an exemplary implementation of a manifest generator.

[0120] In particular, as shown with respect to Figure 4A, the manifest generator 4500 generally reads an input file or stream 4400 (e.g., an ISOBMFF file) and generates the needed information for the corresponding manifest file 4600. Such needed data may include information about the media stream 4620, such as the appropriate media type (sometimes called a MIME type), codec attributes, as well as media-specific properties. All of this information may be generated independently for each media file or stream based on data available from the container-level metadata structure 4410.

[0121] Apart from the information related to these individual media files or streams, it may also be necessary to generate preselection-related information 4610. In conventional techniques, the required information would typically have to be provided, for example, through some kind of out-of-band side data path (as shown in FIG. 4A) or from parsing the binary header data of the encapsulated elementary stream.

[0122] In particular, in the latter case, the corresponding device / component 4510 (which may in some cases be called a preselection data generator or using any other suitable terminology) will therefore not only depend on the codec format, but also require computationally expensive parsing of binary data and additional format-dependent knowledge of how to generate the corresponding preselection from the available content component metadata, depending on the media format used. In some cases, this step may additionally require jointly reading the payloads of multiple media files or streams, making this process difficult and error-prone. In some cases, the manifest generator 4500 may have another (internal) device / component 4520 that may be responsible for performing operations such as media-specific signaling (e.g., "adaptation sets"), as will be understood and appreciated by those skilled in the art.

[0123] To address some or all of the issues discussed with respect to Figure 4A, Figure 4B schematically illustrates an example implementation of a manifest generator according to an embodiment of the present disclosure. In particular, as previously indicated, identical or similar reference numerals in Figure 4B indicate identical or similar elements in Figure 4A, unless otherwise indicated, and repeated descriptions thereof may be omitted for reasons of brevity.

[0124] In particular, as shown with respect to Figure 4B, by utilizing the proposed preselection-related data structure 4451, not only can the need for a codec-specific binary parser be avoided, but the manifest generator 4501 can rely on pre-generated information from the proposed data structure 4451 of the incoming stream 4401. As can be understood and appreciated by those skilled in the art, this data structure 4451 is already aligned with the desired output data 4611. Therefore, unlike the (codec-specific) preselection data generator 4510 as shown in Figure 4A, the manifest generator 4501 proposed herein generally only requires a preselection data converter 4531 responsible for converting the byte information into a manifest file representation, e.g., XML.

[0125] In summary, as will be understood and appreciated by those skilled in the art, the above proposed approach for the manifest generator use case generally provides at least the following advantages: A format-agnostic implementation of preselection data processing in the manifest generator; A byte-only implementation avoids the need for computationally more expensive operations on binary data; Low overhead for implementing preselection metadata in packagers due to high overlap with existing functionality; Equal manifest generator behavior for all media formats; · Less effort for implementers, less testing effort, more reliable and faster uptake by industry.

[0126] As indicated above and elsewhere in this disclosure, the techniques proposed in this disclosure may be applied to the processing of audio streams as well as video streams (or a combination thereof). Specifically, as it relates to video processing, it is worth discussing some of the potential use cases in more detail to further understand the invention.

[0127] One possible use case may include frame rate / resolution scalable video processing. More specifically, some possible MPEG codecs (e.g., in the context of Versatile Video Coding (VCC) as defined in ISO / IEC 23090-3 ("MPEG-I, Part 3"), also known as ITU-T H.266) may provide the option to create multiple display frame rates from the same stream. For example, one stream may be decoded to 50 Hz or, at the expense of higher complexity, to 100 Hz. As can be understood and appreciated by those skilled in the art, being able to decode one stream to either frame rate may be covered by the concept of preselection as proposed in this disclosure. In a related use case, one stream carries only bits for 50 Hz preselection, and a second stream provides additional bits for 100 Hz preselection, so that to decode to 100 Hz, both streams are referenced by the 100 Hz preselection. Similarly, instead of scaling the frame rate, the image resolution may be scaled, so the above example also applies to multi-resolution streams.

[0128] Another possible example use case may relate to joint SDR / HDR streams. This particular use case may be related to the concept of high dynamic range (HDR), which offers higher contrast and sometimes even a wider color gamut (WCG) compared to legacy video in standard dynamic range (SDR). Systems such as those providing backward-compatible enhancement layers may be used to create streams that can be decoded for either SDR displays or HDR-capable displays. The possibility of deriving different experiences from either the same stream or multiple streams may be signaled along with preselections associated with the different experiences.

[0129] Yet another video-related use case may concern the picture-in-picture (pip for short) technique. Specifically, in this use case, parts of a video composition may be replaced by different content (e.g., by using the "sub-picture" concept in VCC, or any other suitable means). An example is a news program where parts of the display are replaced with video captures of a sign language interpreter interpreting the content for hearing-impaired viewers. This is another example where the final experience is composed of different parts, which are chosen (collectively) by pre-selection.

[0130] Having discussed the possible use cases above, reference is now made to Figures 5-8, which schematically show flow charts illustrating example methods for processing media streams, according to embodiments of the present invention.

[0131] In particular, from a broad perspective, the stream processing methods of Figures 5 and 6 may be viewed as two possible solutions for handling (signaling) information related to preselection in transport format files (e.g., ISOBMFF) that may be implemented together or alternatively depending on various implementations and / or circumstances, while the stream processing methods of Figures 7 and 8 may be viewed as corresponding to the possible use cases of the packager and manifest generator described above with reference to Figures 3B and 4B, respectively.

[0132] More specifically, method 5000 shown in FIG. 5 may begin in step S5100 by receiving a packetized media stream according to a predefined transport format. As previously indicated, the media stream may be an audio stream, a video stream, or a combination thereof, such as media stream 3401 prepared by packager 3301 described with reference to FIG. 3B. Method 5000 may be performed at a user end, or in other words, in a user-end (decode-end) environment, which may include, but is not limited to, a TV, a sound bar, a web browser, a media player, a plug-in, etc., depending on various implementations. The predefined transport format may be the Base Media File Format (ISOBMFF) as defined by ISO / IEC 14496-12 MPEG-4 Part 12 by ISO, or any other suitable (transport) format. Specifically, the packetized media stream may include multiple hierarchical boxes, each associated with a respective box type identifier. As noted above, the term "box" as used throughout this disclosure should not be understood to be limited to only that specific term. Rather, the term "box" should be understood generally as any suitable data structure that can serve as a placeholder for media data within a packetized media stream and, therefore, can be referred to by using other appropriate terms. Furthermore, as will be understood and appreciated by those skilled in the art, multiple boxes can be at the same or different levels (or positions), nested (child / sub-boxes and parent box), etc., depending on various implementations and / or requirements. Multiple boxes can contain, among other possibilities, one or more track boxes that reference (or, in other words, indicate) respective tracks that represent media (content) components of a media stream.Broadly speaking, a media (content) component can generally refer to a single / individual continuous component of media content (and may typically be associated with a corresponding (assigned) media content component type, such as audio, video, text, etc.). Examples for understanding the concept of "media (content) component" can be found, for example, as defined / explained in MPEG-DASH (ISO / IEC23009-1).

[0133] Then, in step S5200, method 5000 may include determining whether the media stream includes a preselection-related box of a predefined type indicating a preselection, which may correspond to a media presentation to a user. As discussed above, the term "preselection" is generally used to refer to a set of media content components (of a media stream) that are intended to be consumed together (e.g., by a user-side device) and, more specifically, generally represent one version of a media presentation that may be selected by an end user for simultaneous decoding / presentation. Of course, as will be understood and recognized by those skilled in the art, in some other possible technical contexts, the term "preselection" may be known (or referred to) using other appropriate (equivalent) terms, such as, but not limited to, "presentation," as described in ETSI TS103190-2, or "preset," as described in ISO / IEC 23008-3. Thus, a preselection-related box may be a particular box among multiple boxes in a media stream of a particular predefined (or predetermined) type, and such a particular type indicating preselection may be predefined (or predetermined) in advance by using any suitable means, as will be understood and appreciated by those skilled in the art.

[0134] If it is determined in step S5300 that the media stream includes a preselection-related box, method 5000 may further include parsing metadata information corresponding to the preselection-related box in step S5310, where the metadata information indicates preselection characteristics; identifying one or more tracks in the packetized media stream that contribute to the preselection based on the metadata information in step S5320; and finally, providing the one or more tracks for downstream processing according to the given preselection in step S5330. As will be understood and appreciated by those skilled in the art, the metadata information presented above may be included (directly) in the media stream (or, more specifically, in multiple boxes of the (packetized) media stream) or may be derived (indirectly) from the media stream by using any appropriate means, depending on various implementations. For example, the metadata information may be included in a header box that may be somehow associated with or linked to the preselection-related box (e.g., as a sub-box thereof). As mentioned above, a preselection generally refers to a set of media content components that are intended to be consumed together, for example, by one or more appropriate downstream devices (e.g., media decoders, media players, etc.). The downstream devices may also be referred to simply as "sinks" in some possible cases. As a result, depending on various implementations and / or requirements, downstream processing may include, but is not limited to, multiplexing (or re-multiplexing), ordering, merging, decoding, or rendering those contributing tracks, as described above.

[0135] Configured as described above, the proposed method 5000 generally provides an efficient yet flexible manner for determining / identifying and subsequently signaling tracks within a media stream that are configured to contribute to a particular preselection, thereby enabling further downstream processing of such contributing tracks (e.g., by one or more downstream devices). Thus, broadly speaking, the proposed method may be considered to provide the possibility and capability of signaling information indicating a preselection (and potentially its processing) in a transport layer file (e.g., ISOBMFF) in a unified manner. This may be considered beneficial in various use cases or scenarios, such as those described in detail above with reference to FIGS. 1A, 2, 3B, and 4B.

[0136] In particular, as will be understood and appreciated by those skilled in the art, the International Organization for Standardization (ISO), together with the International Electrotechnical Commission (IEC), has jointly published a functional document for the implementation of certain technologies entitled ISO / IEC 14496-12, Information technology—Coding of audiovisual objects—Part 12: ISO base media file format (the latest version was published in 2020 and is available at https: / / www.iso.org / standard / 74428.html). Originally drafted by the Moving Picture Experts Group (MPEG), this document specifies the ISO base media file format (or ISOBMFF for short), a general format that forms the basis for other, more specific file formats.

[0137] As already indicated above, media information in ISOBMFF is typically structured in a hierarchical, object-oriented manner by utilizing building blocks called "boxes." These data structures are defined by a unique type identifier and length and can be nested or concatenated to form the overall file structure.

[0138] Individual media data within a file format, such as one single video or audio component, are sometimes called "tracks", represented as elementary streams previously generated by a media encoder.

[0139] This file format contains timing, structure, and media information for timed sequences of media data, such as audiovisual presentations. ISO / IEC 14496-12 already includes several means for describing the characteristics of media components. Such descriptive elements or dedicated boxes are assigned to individual tracks and subtracks within the Movie Header ("moov") box. Such available boxes include the "Extended Language Tag" ("elng") box or the Kind ("kind") box.

[0140] As explained in detail above with reference to Figure 5, for the purpose of signaling the properties of a track or sub-track combination, the present disclosure generally proposes to provide a new data structure to be used in parallel with the definition of the track in, for example, the Movie Header (or any other suitable box) as some possible implementations. Nevertheless, it should be noted that, although linking each of these so-called preselections to their constituent media tracks (contributing media tracks) is part of the present proposal, the existing means for signaling their respective properties should be reused as much as possible.

[0141] For example, in some possible implementations, two new boxes may be introduced: one container box that forms the counterpart of the DASH preselection element and may be considered to take all descriptive and structural metadata as child boxes, and another box that may be a (mandatory) header box that takes the structural metadata to create the preselection. Of course, any further appropriate boxes (e.g., processing-related boxes) may also be included if needed, as explained in detail above. In such cases, linking the preselection to tracks can be achieved, for example, by using a unique track identifier (ID) assignment in the track header of each track. Since preselections generally need to be referenced by external applications, a unique identifier shall be assigned accordingly. In particular, the preselection identifier shall be unique within the entire bundle, i.e., across all ISOBMFF files available for this media presentation. Furthermore, since the encoded media format may utilize different assignments, further element(s) may need to be defined and / or configured accordingly (e.g., using corresponding tags to map the respective identifiers used in the media format, etc.).

[0142] In view of that, and in particular for the purpose of implementing the proposed method as described above with respect to Figure 5 in order to support preselection within the context of ISOBMFF, it may be generally suggested in some possible implementations to amend (create or replace) ISO 14496-12, Section 8 with the following text:

[0143] 3.1 Terms and Definitions 3.1.x Preselection A set of one or more media components that represent one version of a media presentation that can be selected by the user for simultaneous decoding / presentation. 3.1.y Bundle A set of multiple tracks, not necessarily contained in one single ISOBMFF file, that gives the entire usable experience of a media presentation.

[0144] 8.18 Preselection Structure 8.18.1 Introduction 8.18.2 Backward Compatibility 8.18.3 Preselection Box definition Box type: 'pres' Containers: Movie Box ('moov'), Movie Fragment Box ('moof') Required: No Amount: Zero or more This is a container box for a single preselection. A preselection provides the description and technical composition of one specific end-user experience. There should be one preselection box for each available preselection in the media presentation.

[0145] The mandatory preselection header box contains the necessary references to all tracks that make up the preselection, and also provides identifiers that can be used for the selection process.

[0146] The properties of a preselection are indicated using additional boxes such as an extended language box or a kind box. When multiple preselections share the same common property, content authors should provide an appropriate text description through a label box. Syntax aligned(8) class PreselectionBox extends Box('pres'){ }

[0147] 8.18.3 Preselection Header Box definition Box type: 'prhd'. Container:Preselection Box ('pres') Required: Yes Quantity: Just 1 This box specifies the characteristics of a single preselection. A preselection contains exactly one preselection header box. The preselection_id and preselection_tag provide the identification of the preselection for an application or media codec, respectively, which can be used to select the presentation. selection_priority should be used to guide any automated selection process when no other distinction is provided. Evaluating the order of preselection elements to determine their priority should be avoided. Syntax [Table 1] Semantics version is an integer (0 in this specification) that specifies the version of this box. flags is a 24-bit integer containing flags, the values ​​of which are not yet defined. The preselection_id is an integer that declares a unique identifier to external applications. The preselection_tag is an integer that declares an identifier for the media codec to be used. selection_priority is an integer that declares the priority of the preselection when no other distinction is possible, such as through the media language. Lower numbers indicate higher priority. n_tracks is an integer that declares the number of tracks needed to form the preselection. track_ID is an array of integers that uniquely identify each track. The list of all tracks in this array is referenced via the track_ID value required for preselection.

[0148] 8.18.4 Label and Group Label Boxes definition Box type: 'labl'. Container: A preselection box ('pres') and optionally others Required: No Amount: Zero or more This box can be used to annotate the container-containing structure with a textual description. This description is intended to be presented to the user in some form of textual display, and is not intended for automated selection processes. Multiple boxes can be used to provide text descriptions in different languages. Syntax aligned(8) class LabelBox extends Box('labl'){ template unsigned int(16) label_id=0; string language String label } Semantics label_id is an integer containing an identifier that can be used by external applications. language is a null-terminated C string containing an RFC4646 (BCP47) compliant language tag string, such as "en-US", "fr-FR", or "zh-CN". label is a null-terminated C string containing a text description.

[0149] As a result, it may be further proposed to update Table 1 (in the same document ISO 14496-12) with the underlined lines as shown below: [Table 2-1] [Table 2-2] [Table 2-3]

[0150] Additionally, it may be proposed to introduce new, container-like preselection header boxes for some of the existing boxes. For example, it may be proposed to modify the existing box definition (in the same document ISO14496-12) as follows:

[0151] 8.4.6 Extended Language Tags 8.4.6.1 Definition Box type: 'elng' Containers: Media Box ('mdia'), Preselection Header Box ('prhd') Required: No Amount: Zero or 1 […]

[0152] As noted above, the modifications suggested above are merely one possible implementation example, and certainly not the only one. Even though specific names for the suggested boxes related to preselection are given / suggested above, these boxes may be named differently. Similarly, even though the preselection-related boxes suggested above appear to be nested under the 'moov' box, these boxes may, of course, be implemented elsewhere. This will be understood and appreciated by those skilled in the art. In one possible example (and, of course, without any limitation), these preselection-related boxes may be associated with (e.g., nested under) a so-called 'EntityToGroupBox' or any other suitable location in the ISOBMFF. Similarly, as will be understood and appreciated by those skilled in the art, as the standard progresses and / or develops, the example table (or elements therein) and example sections described above may have different names and / or (hierarchical) locations. In some possible cases, the exemplary tables (or parts thereof) and / or sections described above may even become partially or completely obsolete (outdated), and the respective preselection-related information may need to be defined elsewhere as it becomes appropriate by then.

[0153] Referring now to Figure 6, there is shown a flowchart of another possible example of a method 6000 for processing a media stream, according to an embodiment of the present invention. As noted above, the method of Figure 6 may generally be viewed as an alternative (or in some possible cases an addition) to the method of Figure 5 described above.

[0154] Specifically, method 6000 begins in step S6100 by receiving a media stream packetized according to a predefined transport format. As above, the media stream may be an audio stream, a video stream, or a combination thereof, such as media stream 3401 prepared by packager 3301 described with reference to FIG. 3B. Method 6000 may be performed at a user end, or in other words, in a user-end (decode-end) environment, which may include, but is not limited to, a TV, a sound bar, a web browser, a media player, a plug-in, etc., depending on various implementations. The predefined transport format may be the Base Media File Format (or ISOBMFF for short), as defined by ISO / IEC 14496-12 MPEG-4 Part 12 by ISO, or any other suitable (transport) format. Specifically, the packetized media stream may include multiple hierarchical boxes, each associated with a respective box type identifier. The multiple boxes may contain, among other possibilities, one or more track boxes that reference (or in other words indicate) respective tracks that represent the media (content) components of the media stream.

[0155] Thereafter, in step S6200, method 6000 may include checking (e.g., visiting, cycling through) track boxes within the media stream to determine the full (or complete / total) set of preselections present in the media stream. Specifically, determining the full set of preselections may include determining a set of unique pairs of track group identifiers and track group types and addressing the preselections by their respective track group identifiers. As noted above, each preselection is associated with a respective track group, which itself is identified by a respective pair of corresponding track group identifier and corresponding track group type. Thus, a preselection may be addressed (or identified) by its associated / linked respective track group identifier.

[0156] Method 6000 may further include, at S6300, selecting a preselection from the full set of preselections. Specifically, the preselection may be selected based on attributes (e.g., represented in metadata or other suitable form) of each preselection included in a track group box with the same track group identifier.

[0157] Thereafter, in step S6400, the method 6000 may include determining a set of one or more track boxes that contribute to the selected preselection. In particular, a set of one or more track boxes that contribute to the same preselection may be identified by the presence of (respective) track group boxes with the same track group identifier.

[0158] Further, in step S6500, the method 6000 may include determining the tracks referenced in each member of the set of one or more track boxes determined above as one or more tracks contributing to the preselection.

[0159] Finally, in step S6600, method 6000 may include providing the one or more tracks for downstream processing according to a preselection. As noted above, a preselection generally refers to a set of media content components intended to be consumed together by one or more appropriate downstream devices (or, in some possible cases, referred to as sinks), such as, for example, a media decoder, a media player, etc. As a result, depending on various implementations and / or requirements, downstream processing may include, but is not limited to, multiplexing (or re-multiplexing), ordering, merging, decoding, or rendering of those contributing tracks, as described in more detail below.

[0160] When configured as described above, the proposed method 6000 generally provides an efficient yet flexible way to determine / identify and subsequently signal tracks within a media stream configured to contribute to a particular preselection, thereby enabling further appropriate downstream processing of such contributing tracks (e.g., by one or more downstream devices). It should be noted, in particular, that the method proposed above in the first aspect generally seeks to provide relevant information for (all) tracks configured to contribute to a particular preselection in a preselection-related box, thereby enabling indexing (or identification) of all contributing tracks. In this sense, track indexing as described in the above method 5000 of FIG. 5 can be considered a type of forward (direct) indexing. In contrast, in the method 6000 shown in FIG. 6, the tracks contributing to a particular preselection can be jointly determined by a pair of track group type and track group identifier. More specifically, tracks having (e.g., including) a track group box with a particular (e.g., predefined or predetermined) track group type may generally indicate that the track contributes to preselection. Furthermore, tracks having the same track group identifier may generally indicate that the tracks belong to (contribute to) the same preselection. In this sense, such indexing of tracks described with reference to FIG. 6 may be considered a kind of reverse indexing, compared to that proposed in FIG. 5. In any case, similar to FIG. 5, the proposed method 6000 of FIG. 6 may also provide the possibility and capability to signal information indicating preselection (and potentially its processing) in a transport layer file (e.g., ISOBMFF) in a unified manner. This may be considered beneficial in various use cases or scenarios detailed above with reference to FIGS. 1A, 2, 3B, and 4B.

[0161] In particular, for the purpose of implementing the proposed method 6000 described above with respect to FIG. 6 as well as FIG. 5, and in particular to support preselection within the context of ISOBMFF, in some possible implementations it may generally be proposed to amend (create or replace) ISO 14496-12 with the following text (instead of or in addition to what is proposed with reference to FIG. 5):

[0162] 3.1 Terms and Definitions 3.1.x Preselection A set of one or more media components that represent one version of a media presentation that can be selected by the user for simultaneous decoding / presentation. 3.1.y Media Components

[0163] It may be generally suggested to amend ISO 14496-12, section 8.3.4.3, with the following underlined lines: track_group_type indicates the grouping_type and shall be set to one of the following values, or a registered value, or a value from a derived specification or registration: 'msrc' indicates that this track belongs to a multi-source presentation, as specified in 8.3.4.4.1. 'ster' indicates that this track is either the left or right view of a stereo pair suitable for playback on a stereoscopic display. 'pres' indicates that this track contributes to preselection, as specified in 8.3.4.4.3. The track_group_id and track_group_type pair identifies a track group within the file. Tracks containing a particular TrackGroupTypeBox with the same values ​​of track_group_id and track_group_type belong to the same track group.

[0164] In general, it can be proposed to amend (or create) ISO14496-12, section 8.3.4.4.3 with the following lines: 8.3.4.1.3 Preselection Box 8.3.4.1.3.1 Definition A TrackGroupTypeBox with track_group_type equal to 'pres' indicates that this track contributes to preselection. Tracks with the same value of track_group_id within a PreselectionGroupBox are part of the same preselection. The preselection may be qualified by media-specific attributes such as language, kind, or audio rendering instructions or channel layout. Attributes signaled in the preselection box shall take precedence over attributes signaled in the contribution tracks. All attributes that uniquely qualify a preselection shall be present in at least one preselection box of the preselection. If they are present in more than one preselection box of a preselection, those boxes shall be identical. Note: The preselection only groups tracks of the same media type. A track that does not contain all the required media components for at least one preselection shall have the track_in_movie flag set to '0' in its track header box. This prevents players that do not understand preselection boxes from playing the track and resulting in an incomplete experience. Note: It is good practice to have one track with the track_in_movie flag set to 1. This means that this track provides at least one complete experience.

[0165] 8.3.4.1.3.2 Syntax aligned(8) class PreselectionGroupBox extends TrackGroupTypeBox('pres') { if (flags & 1) { unsigned int(8) selectionPriority=1 } if (flags & 2) { unsigned int(8) order=0 } PreselectionInformationBox() PreselectionProcessingBox() }

[0166] 8.3.4.1.3.3 Semantics selection_priority is an integer that declares the priority of the preselection when no other distinction is possible, such as through the media language. Lower numbers indicate higher priority. The order specifies the compliance rules for representation within the adaptation set in the preselection according to [MPEG-DASH] from the set listed below. 0: undefined 1: time-ordered 2: Fully-ordered

[0167] 8.3.4.1.3.4 Preselection Information Box 8.3.4.1.3.4.1 Definition Box type: 'prsi' Container: Preselection Box Required: Yes Quantity: Just 1 This box synthesizes all the semantic information about the preselection.

[0168] 8.3.4.1.3.4.2 Syntax aligned(8) class PreselectionInformationBox extends FullBox('prsi', version=0, 0 ){ / / Boxes that describe the preselection } 8.3.4.1.3.4.3 Semantics 8.3.4.1.3.4.5 Preselection Processing Box 8.3.4.1.3.4.5.1 Definition Box type: 'prsp' Container: Preselection Box Required: Yes Quantity: Just 1 This box contains information about how tracks contributing to the preselection are processed. Media type specific boxes MAY be used to describe further processing.

[0169] 8.3.4.1.3.4.5.2 Syntax aligned(8) class PreselectionProcessingBox extends FullBox('prsp', version=0,0){ string preselection_tag; unsigned int(8) track_order unsigned int(1) sample_merge_flag unsigned int(7) reserved / / of tracks contributing to preselection / / Boxes that define additional processing } 8.3.4.1.3.4.5.3 Semantics The preselection_tag is an integer containing an identifier for the label. Labels with the same value belong to a label group. The default value of zero indicates that the label does not belong to any label group. track_order defines the order in which tracks are presented to the decoder. Tracks with smaller track_order are given to the decoder first. If multiple tracks have the same value of track_order, the order does not matter. sample_merge_flag: If this flag is set to '1', each sample of this track is appended to the sample of the track with the next smaller track_order value. If set to '0', this track is fed to a separate decoder instance.

[0170] Furthermore, it may be proposed to make (or amend or replace) ISO 14496-12, section 8.18.4 as follows:

[0171] 8.18.4 Label and Group Label Boxes definition Box Type: 'labl' Container: Container user data box ('udta', preselection box ('pres')) in the truck Required: No Amount: Zero or more Labels provide the ability to annotate data structures within ISOBMFF files to provide a description of the context of the element to which the label is assigned. Such labels may be used, for example, by a playback client to provide choices to the user. Labels may also be used for simple annotation in other contexts. Additionally, a GroupLabel element may be added at a higher level to provide a summary or title for the labels collected in the group. An example would be this being used in a menu to give context for the menu of labels.

[0172] Multiple labels can be used to provide a text description. To annotate the preselection for a multilingual audience, the annotations may be provided in a language different from the language of the preselection.

[0173] Syntax aligned(8) class LabelBox extends FullBox('labl', version=0, 0 ){ unsigned int(8) is_group_label = 0; unsigned int(16) label_id = 0; utf8string language; utf8string label; } Semantics is_group_label specifies whether the label contains a summary label for a group of labels. label_id is an integer containing an identifier for the label. Labels with the same value belong to a label group. The default value of zero indicates that the label does not belong to any label group. language is a null-terminated C string containing an RFC4646 (BCP47) compliant language tag string such as "en-US", "fr-FR", or "zh-CN". Language is the language targeted by the label. label is a null-terminated C string containing a text description.

[0174] It may further be proposed that ISO 14496-12, section 8.18.4 be formulated (or amended or replaced) as follows:

[0175] 8.18.5 Audio Rendering Instructions Box definition Box type: 'ardi' Container:Preselection Box ('pres') Required: No Amount: Zero or 1 The Audio Rendering Instructions box contains hints for the preferred playback channel layout. Syntax aligned(8) class AudioRenderingIndicationBox extends FullBox('ardi', version=0, 0){ unsigned int(8) audio_rendering_indication = 0; } Semantics The audio_rendering_indication contains a hint for the preferred playback channel layout, coded according to Table 2. [Table 3]

[0176] In particular, based on the newly proposed preselection related boxes as explained above, further proposals can be made to the above Table 1 as explained with reference to Figure 5. This will be understood and appreciated by those skilled in the art, and therefore the details thereof will not be repeated for the sake of brevity.

[0177] It may also be proposed to introduce new, container-like preselection header boxes for some of the existing boxes, for example by modifying the existing box definitions with underlined or strikethrough text. 8.4.6 Extended Language Tags 8.4.6.1 Definition Box type: 'elng' Container: Media Box ('mdia'), Preselection Box ('pres') Required: No Amount: 0 or 1 […] 8.10.1 User Data Box 8.10.4.1 Definition Box type: 'UDTA' Containers: Movie Box, Track Box, Movie Fragment Box, [Outside 1] TIFF0007798301000006.tif7170 Track fragment box, or Preselection Box Required: No Amount: 0 or 1 […] 8.10.4 Track Type 8.10.4.1 Definition Box type: 'kind' Container: Container User Data Box ('UDTA') inside the truck or preselection box ('pres') Required: No Amount: 0 or 1 12.2.4 Channel Layout 8.10.4.1 Definition Box type: 'chnl' Container:Audio Sample Entry or Preselection Box Required: No Amount: 0 or 1 […]

[0178] Again, as mentioned above, the above suggested modifications described with reference to the method of Figure 6 should be understood as merely one possible implementation example, and certainly not the only one. Even though specific names for some of the suggested boxes related to preselection are given / suggested above, these boxes may be named differently. Similarly, even though the above suggested preselection-related boxes appear to be associated with some specific box, these boxes may, of course, be implemented elsewhere. This will be understood and appreciated by those skilled in the art.

[0179] Referring now to FIG. 7, an exemplary flowchart of a method 7000 for processing a media stream, according to an embodiment of the present invention, is generally shown. As above, the media stream may be an audio stream, a video stream, or a combination thereof. In particular, the method 7000 may be performed on an encoding environment (e.g., a media encoder). In some scenarios (or use cases), such an encoder may also be referred to as a (media) packager (i.e., configured to pack / packetize media input), as exemplified by 2300-A-C in FIG. 2 or 3301 in FIG. 3B.

[0180] In particular, method 7000 may include, at step S7100, encapsulating one or more elementary streams according to a predefined transport format to generate a packetized media stream. The packetized media stream includes a plurality of hierarchical boxes, each associated with a respective box type identifier. As above, the predefined transport format may be the Base Media File Format (or ISOBMFF for short), as defined by ISO / IEC 14496-12 MPEG-4 Part 12 by ISO, or any other suitable (transport) format.

[0181] More specifically, step S7100 of encapsulating one or more elementary streams includes, in step S7110, packetizing media data of the one or more elementary streams in accordance with a transport format to generate one or more track boxes that reference (or indicate) respective tracks of the one or more elementary streams; and, in step S7120, generating one or more preselection-related boxes of a predefined type based on header information of the one or more elementary streams, each of the one or more preselection-related boxes indicating a respective preselection corresponding to media presentation to a user.

[0182] Configured as described above, the proposed method may generally provide an efficient yet flexible way to packetize media input (e.g., elementary streams) according to a predefined transport format (e.g., ISOBMFF). More specifically, by generating and including one or more preselection-related boxes (each indicating a respective preselection) along with the packetized media stream, the proposed method may also enable representing preselections in a codec-agnostic and uniform manner, thereby further enabling appropriate downstream processing of tracks contributing to the corresponding preselection (e.g., according to the methods proposed in the first and second aspects above). Furthermore, as shown above, such a unified representation of preselection also enables a format-agnostic implementation of preselection data processing in manifest (e.g., MPEG Dynamic Adaptive Streaming over HTTP, DASH, format file, or HTTP Live Stream, HLS, format file) generators, thereby avoiding the need for computationally more expensive operations on binary data, while simultaneously reducing implementation effort and improving reliability.

[0183] 8 is a schematic flowchart illustrating yet another example of a method 8000 for processing a media stream according to an embodiment of the present invention. The media stream may be an audio stream, a video stream, or a combination thereof. The method 8000 may be performed by a manifest generator, such as the manifest generator 2500 of FIG. 2 or 4501 of FIG. 4B.

[0184] In particular, method 8000 may include, in step S8100, receiving a packetized media stream according to a predefined transport format. More specifically, the packetized media stream may include a plurality of hierarchical boxes, each associated with a respective box type identifier, and the plurality of boxes may include one or more track boxes that reference (e.g., indicate) respective tracks that indicate media components of the media stream, and one or more preselection-related boxes of a predefined type, each preselection-related box indicating a respective preselection that corresponds to a media presentation to a user.

[0185] Further, the method 8000 may include, in step S8200, generating a manifest file based on the one or more preselection association boxes.

[0186] Configured as described above, the proposed method can generally provide an efficient yet flexible way for generating a manifest file by also taking into account preselection-related information (e.g., description information and / or processing-related information). More specifically, apart from information about the media stream, the proposed method can further include preselection-related information. The preselection-related information can include, but is not limited to, metadata information, processing information, etc. Compared to conventional (manifest generation) techniques, the proposed method for generating a manifest file generally provides a format-agnostic implementation for handling preselection-related data, thereby avoiding computationally more expensive operations (e.g., operations that must be performed on binary data in conventional techniques), reducing implementation and / or testing effort, and increasing reliability.

[0187] In summary, the methods proposed and described above with reference to the drawings generally relate to techniques for processing media streams by taking "preselection" into account. Such "preselection"-aware techniques may enable a variety of potential use cases. One possible use case may involve selecting among several languages ​​(subtitles), as described above.

[0188] Another possible exemplary use case may relate to the importance of narrative. In particular, techniques such as dialogue enhancement may be required, possibly to address the needs of the hearing impaired or possibly to adapt an audio composition to various listening conditions. This is generally used to improve the ratio of dialogue to background levels. As an extension of such dialogue enhancement techniques, additional measures have been proposed to selectively and progressively drop entire audio elements from a composition in order to improve dialogue intelligibility. In such cases, it is possible to map a continuum from "complete material" to "dialogue only" onto several preselections and use the methods of the present disclosure.

[0189] A further possible use case may relate to audience targeting. More specifically, it may sometimes be beneficial to target a specific audience beyond language. As an example (but not as a limitation of any kind), two sports match commentators may be provided, each biased towards one of the teams. In this case, different preselections may include or refer to different sports match commentators.

[0190] Yet another preselection-related use case may be related to playback environment adaptation: a content creator may generate a specialized version targeting a different playback environment, such as a home cinema setup, compared to a version for the TV's built-in speakers or headphones. In this case, different preselections may relate to different audio for the different playback environments.

[0191] Finally, the present invention also relates to apparatus(es) for performing the methods and techniques described throughout this disclosure. FIG. 9 generally illustrates an example of such an apparatus 9000. In particular, the apparatus 9000 includes a processor 9100 and a memory 9200 coupled to the processor 9100. The memory 9200 may store instructions for the processor 9100. The processor 9100 may also receive, among other things, input data (e.g., media input, packetized media streams, etc.), depending on various use cases and / or implementations. The processor 9100 may be adapted to perform the methods / techniques described throughout this disclosure (e.g., methods 5000, 6000, 7000, and 8000 shown above) and correspondingly generate output data 9400 (e.g., packetized media streams, manifest files, etc.), depending on various use cases and / or implementations. For example, the apparatus 9000 may implement a packager configured to perform the method 7000 for processing a media stream as shown above with respect to FIG. 7, or may implement a manifest generator configured to perform the method 8000 for processing a media stream as shown above with respect to FIG. 8, depending on the circumstances, in accordance with an embodiment of the present invention.

[0192] interpretation A computing device implementing the above techniques may have the following exemplary architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the exemplary architecture includes one or more processors (e.g., a dual-core Intel® Xeon® processor), one or more output devices (e.g., an LCD), one or more network interfaces, one or more input devices (e.g., a mouse, a keyboard, a touch-sensitive display), and one or more computer-readable media (e.g., RAM, ROM, SDRAM, a hard disk, an optical disk, flash memory, etc.). These components may communicate and exchange data over one or more communication channels (e.g., a bus), which may utilize various hardware and software to facilitate the transfer of data and control signals between the components.

[0193] The term "computer-readable medium" refers to any medium that participates in providing instructions to a processor for execution, including, but not limited to, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory), and transmission media, including, but not limited to, coaxial cables, copper wire, and fiber optics.

[0194] The computer-readable medium may further include an operating system (e.g., a Linux operating system), a network communications module, an audio interface manager, an audio processing manager, and a live content distributor. The operating system may be multi-user, multi-processing, multi-tasking, multi-threaded, real-time, etc. The operating system performs basic tasks including, but not limited to, recognizing input from and providing output to network interfaces and / or devices, tracking and managing files and directories on the computer-readable medium (e.g., memory or storage devices), controlling peripheral devices, and managing traffic on one or more communications channels. The network communications module includes various components for establishing and maintaining network connections (e.g., software for implementing communications protocols such as TCP / IP, HTTP, etc.).

[0195] The architecture may be implemented in a parallel processing or peer-to-peer infrastructure, or on a single device having one or more processors. The software may include multiple software components or may be a single body of code.

[0196] The described features may be advantageously implemented in one or more computer programs executable on a programmable system including at least one programmable processor coupled to receive data and instructions from and transmit data and instructions to a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform an activity or bring about a result. Computer programs can be written in any form of programming language, including compiled or interpreted languages ​​(e.g., Objective-C, Java), and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, browser-based web application, or other unit suitable for use in a computing environment.

[0197] Suitable processors for executing a program of instructions include, by way of example, both general-purpose and special-purpose microprocessors, and the sole processor or one of multiple processors or cores of any kind of computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer also includes, or is operatively coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks, magneto-optical disks, and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, for example, semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, application-specific integrated circuits (ASICs).

[0198] To provide for user interaction, features may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or retinal display device, for displaying information to the user. The computer may have a touch surface input device (e.g., a touch screen) or keyboard, and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. The computer may have a voice input device for receiving voice commands from the user.

[0199] Features can be implemented in a computer system that includes back-end components such as data servers, or includes middleware components such as application servers or Internet servers, or includes front-end components such as client computers with graphical user interfaces or Internet browsers, or includes any combination thereof. The components of the system can be connected by any form or medium of digital data communication, such as a communications network. Examples of communications networks include, for example, LANs, WANs, and the computers and networks forming the Internet.

[0200] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server sends data (e.g., HTML pages) to client devices (e.g., to display the data to and receive user input from users interacting with the client devices). Data generated at the client devices (e.g., the results of user interaction) may be received by the server from the client devices.

[0201] One or more computer systems may be configured to perform particular actions by having software, firmware, hardware, or a combination thereof installed on the system that, when running, causes the system to perform the actions. One or more computer programs may be configured to perform particular actions by including instructions that, when executed by a data processing device, cause the device to perform the actions.

[0202] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a certain combination and may initially be claimed as such, one or more features from a claimed combination can, in some cases, be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0203] Similarly, while operations are shown in the figures in a particular order, this should not be understood as requiring such operations to be performed in the particular order shown, or in sequential order, or that all of the shown operations be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged in multiple software products.

[0204] Unless otherwise stated, and as will be apparent from the following description, throughout the description of the present invention, descriptions utilizing terms such as "processing," "computing," "calculating," "determining," "analyzing," and the like will be understood to refer to the actions and / or processes of a computer or computing system or similar electronic computing device that manipulates and / or transforms data represented as physical quantities, such as electronic quantities, into other data also represented as physical quantities.

[0205] Throughout this application, a reference to "one exemplary embodiment," "some exemplary embodiments," or "an exemplary embodiment" means that a particular feature, structure, or characteristic described in connection with that exemplary embodiment is included in at least one exemplary embodiment of the invention. Thus, the appearances of the phrases "in one exemplary embodiment," "in some exemplary embodiments," or "in an exemplary embodiment" in various places throughout this application do not necessarily all refer to the same exemplary embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this application.

[0206] As used herein, unless otherwise specified, the use of ordinal adjectives "first," "second," "third," etc. to describe a common object indicates merely that different instances of a similar object are being referred to and is not intended to imply that the objects so described must be in a given order, temporally, spatially, in ranking, or in any other way.

[0207] It is also to be understood that the phraseology and terminology used herein is for purposes of description and should not be regarded as limiting. The use of "including," "having," or "having," and variations thereof, is intended to encompass the listed items and equivalents thereof, as well as additional items. Unless otherwise specified or limited, the terms "mounted," "connected," "supported," and "coupled," and variations thereof, are used broadly and encompass both direct and indirect mounting, connecting, supporting, and coupling.

[0208] In the claims below and in the description of this specification, the terms "having," "consisting of," or "including" are all open terms meaning the inclusion of at least the recited elements / features but not the exclusion of others. Thus, when used in the claims, the terms "having" / "including" should not be interpreted as being limited to the recited means or elements or steps. For example, the scope of the expression "device having A and B" should not be limited to a device consisting only of elements A and B. The terms "comprise," "include," or "comprising" used in this specification are also all open terms meaning the inclusion of at least the recited elements / features but not the exclusion of others. Thus, "comprise" is synonymous with "have" and means "to have."

[0209] In the foregoing description of exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single exemplary embodiment, figure, or description thereof for the purpose of streamlining the invention and aiding in understanding one or more of the various inventive aspects. However, this method of the invention should not be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single, foregoing disclosed exemplary embodiment. Accordingly, the claims following this specification are expressly incorporated herein, with each claim standing on its own as a separate exemplary embodiment of the present invention.

[0210] Furthermore, it will be understood by those skilled in the art that some exemplary embodiments described herein include some features included in other exemplary embodiments but not other features, and that combinations of features of different exemplary embodiments are intended to form different exemplary embodiments within the scope of the present invention. For example, in the following claims, any of the claimed exemplary embodiments may be used in any combination.

[0211] In the description provided herein, numerous specific details are set forth. However, it will be understood that example embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0212] Thus, while what is believed to be the best mode of the invention has been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended to claim all such changes and modifications as fall within the scope of the invention. For example, any formulas given above are merely representative of procedures that may be used. Functions may be added to or deleted from block diagrams, and operations may be interchanged between functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure.

[0213] Enumerated example embodiments ("EEE") of the present disclosure are described above with respect to methods and systems for determining an indication of audio quality of an audio input. Accordingly, certain embodiments of the present invention may relate to one or more of the following enumerated embodiments.

[0214] EEE1. A method of decoding an encoded media bitstream in a multiplexed format, the multiplexed format including a preselection box for a list of versions having version properties, the preselection box being agnostic to the encoded media format, comprising: determining a list of versions with properties for a selection method from the preselection box; and decoding the encoded media bitstream based on the selected method to output a playable audio version. method. EEE2. The method of EEE1, wherein the selection method is an end-user exposed UI that visualizes the versions and properties to perform instant selection. EEE3. The method of EEE1, wherein the selection method is an automated process based on user preference settings. EEE4. The method of EEE1, wherein the selection method is an automated process based on information about the end device, the geographic region of playback, or other data characteristics. EEE5. The method of any one of EEE1 to EEE4, wherein the multiplexed format is ISOBMFF, Transport Stream, or MXF. EEE6. The method of EEE5, wherein the media bitstream includes an encrypted media payload and the encoding format is clear text. EEE7. The method of any one of EEE1 to EEE4, wherein the decoding method is used for specialized playback architectures for audio, video and virtual reality, specialized types of media or specialized preselection. EEE8. The method of any one of EEE1 to EEE4, wherein the multiplexed format includes information regarding pre-downloaded user preferences. EEE9. A method for packaging the encoded media assets for transmission, wherein the transmission format includes a manifest file listing the assets, the method comprising: packaging the manifest file, the manifest file containing a list of available versions and their properties to be derived from either a single-stream or multi-stream asset, the list being derivable from the asset without access to codec-specific information; method. EEE10. The method of EEE9, wherein the manifest file is in MPEG DASH, HLS, or SDP format. EEE11. A method for processing a media stream, comprising: receiving the media stream packetized according to a predefined transport format, the packetized media stream including a plurality of hierarchical boxes, each associated with a respective box type identifier, the plurality of boxes including one or more track boxes referencing tracks indicating media components of the media stream and one or more track group boxes, each of the one or more track group boxes being associated with a respective track group identifier and a respective track group type that jointly identify a respective track group in the media stream, tracks having the same track group identifier and the same track group type belong to the same track group, and each such track group determining a preselection; visiting all track boxes in the media stream to determine a complete set of preselections; determining the set of all unique pairs of track group identifiers and track group types, and addressing the preselection by track group identifiers; selecting a preselection based on attributes of the preselection found in all track group boxes having the same track group identifier; determining a set of one or more track boxes that contribute to said preselection as identified by the presence of track group boxes having the same track group identifier; determining the contributing tracks referenced in each member of the set of track boxes; method.

Claims

1. 1. A method for processing a media stream, comprising: receiving a packetized media stream according to a predefined transport format, the packetized media stream including a plurality of hierarchical boxes each associated with a respective box type identifier, the plurality of hierarchical boxes including one or more track boxes referencing respective tracks representing media components of the media stream; determining whether the media stream includes a preselection-related box of a predefined type indicating a preselection, the preselection corresponding to a media presentation to a user; If it is determined that the media stream includes the preselection associated box: analyzing metadata information corresponding to the preselection-related box, the metadata information indicating characteristics of the preselection; identifying one or more tracks in the packetized media stream that contribute to the preselection based on the metadata information; and providing the one or more tracks for downstream processing in accordance with the preselection; the media stream further comprises processing information indicating how tracks contributing to the preselection should be processed; the processing information includes sequencing information indicating a track order for processing the one or more tracks; the processing information includes merge information indicating whether one or more tracks should be merged with one or more other tracks for joint processing; The method comprises: merging the one or more tracks according to the merge information and the ordering information; the ordering information includes, for each track contributing to the preselection, a respective track order value for defining a track order for the tracks; the merge information includes a respective merge flag for each track contributing to the preselection, a first setting of the merge flag indicating that the respective track should be merged with an adjacent track in track order, and a second setting of the merge flag indicating that the respective track should be processed separately; Merging the one or more tracks in accordance with the merge information and the ordering information includes: Scan the tracks successively according to track order; including merging tracks according to their respective merge flags, method.

2. 2. The method of claim 1, further comprising decoding the one or more tracks for playback of a media stream according to the media presentation indicated by the preselection.

3. The method of claim 2 , wherein the one or more tracks are decoded by a downstream device.

4. The method of claim 2 , wherein the merging of the one or more tracks and the decoding of the one or more tracks are performed by one single device.

5. 2. The method of claim 1, wherein the media stream includes a plurality of preselection-related boxes of the predefined type, the method further comprising selecting the preselection-related box from among the plurality of preselection-related boxes.

6. The method of claim 5 , wherein the preselection association box is selected by an application.

7. 6. The method of claim 5, wherein the media stream includes one or more label boxes each containing descriptive information for a respective media presentation to the user corresponding to a respective preselection.

8. 10. The method of claim 1, wherein the preselection association box is agnostic to the media codec used to encode the media stream before it is packetized.

9. 2. The method of claim 1, wherein the metadata information corresponding to the preselection associated box includes track identification information indicating one or more track identifiers respectively associated with respective tracks, the tracks associated with the one or more track identifiers in the metadata information being relevant to the media presentation.

10. The method of claim 1 , wherein the metadata information corresponding to the preselection-related box includes preselection identification information indicating a preselection identifier for identifying the preselection.

11. The method of claim 1 , wherein the metadata information corresponding to the preselection-related box includes unique preselection-specific data for configuring a downstream device to decode tracks according to the preselection.

12. 1. A method for processing a media stream, comprising: receiving the media stream packetized according to a predefined transport format, the packetized media stream including a plurality of hierarchical boxes each associated with a respective box type identifier, the plurality of hierarchical boxes including one or more track boxes referencing respective tracks indicating media components of the media stream, and one or more track group boxes each associated with a respective pair of track group identifier and track group type that jointly identify a respective track group in the media stream, tracks having the same track group identifier and the same track group type belong to the same track group, and each such track group determines a preselection corresponding to a media presentation to a user; checking the track boxes in the media stream to determine a full set of preselections present in the media stream, wherein determining the full set of preselections includes: determining a set of unique pairs of track group identifiers and track group types, and addressing preselections by their respective track group identifiers; selecting a preselection from the full set of preselections, the preselection being selected based on attributes of each preselection contained in a track group box having the same track group identifier; determining a set of one or more track boxes that contribute to a selected preselection, the set of one or more track boxes that contribute to the preselection being identified by the presence of track group boxes having the same track group identifier; determining the tracks referenced in each member of said set of one or more track boxes as one or more tracks contributing to said preselection; providing the one or more tracks for downstream processing in accordance with the preselection; each preselection being associated with a respective preselection association box of a predefined type, said preselection association box instantiating a track group box having the predefined track group type associated with the preselection; the preselection association box is associated with a preselection processing box containing processing information indicating how the tracks contributing to the preselection should be processed; the processing information includes ordering information indicating a track order for ordering the tracks; the processing information includes merge information indicating whether the track should be merged with one or more other tracks; The method comprises: further comprising merging the tracks according to the merge information and the ordering information; the ordering information includes a track order value for defining a track order for the tracks; the merge information includes a merge flag, a first setting of the merge flag indicating that the track should be merged with an adjacent track in track order, and a second setting of the merge flag indicating that the track should be processed separately; Merging the one or more tracks in accordance with the merge information and the ordering information includes: scanning the tracks successively according to said track order; including merging tracks according to their respective merge flags, method.

13. The method of claim 12 , wherein the preselection association box is associated with a preselection information box that includes semantic information indicating the preselection.

14. 13. The method of claim 12, wherein the processing information includes unique preselection-specific data for configuring downstream devices to decode tracks according to the preselection.

15. further comprising decoding the tracks for playback of the media stream according to the media presentation indicated by the preselection. The method of claim 12.

16. The method of claim 15 , wherein the one or more tracks are decoded by a downstream device.

17. The method of claim 15 , wherein the merging and decoding of the tracks is performed by one single device.

18. The method of claim 12 , wherein the preselection is determined by an application.

19. the media stream includes one or more label boxes each linked to a respective preselection information box for each preselection, each label box including descriptive information for a respective media presentation to a user; the determination of the preselection is based on user input; The method of claim 12.

20. The method of claim 1 , wherein the media stream comprises at least one of an audio stream or a video stream.

21. 1. A method for processing a media stream, comprising: encapsulating one or more elementary streams according to a predefined transport format to generate a packetized media stream; the packetized media stream includes a plurality of hierarchical boxes, each box associated with a respective box type identifier; Encapsulating the one or more elementary streams includes: packetizing media data of the one or more elementary streams in accordance with the transport format to generate one or more track boxes referencing respective tracks of the one or more elementary streams; generating one or more preselection association boxes of a predefined type based on header information of the one or more elementary streams, each of the one or more preselection association boxes indicating a respective preselection corresponding to a media presentation to a user; generating one or more preselection processing boxes containing processing information indicating how the tracks contributing to each preselection should be processed; the processing information includes sequencing information indicating a track order for processing the one or more tracks; the processing information includes merge information indicating whether one or more tracks should be merged with one or more other tracks for joint processing; the ordering information includes, for each track contributing to the preselection, a respective track order value for defining a track order for the tracks; the merge information includes a respective merge flag for each track contributing to the preselection, a first setting of the merge flag indicating that the respective track should be merged with an adjacent track in track order, and a second setting of the merge flag indicating that the respective track should be processed separately; method.

22. each of the one or more preselection-related boxes includes metadata information indicating characteristics of the respective preselection; the metadata information includes information indicating one or more tracks in the media stream that contribute to each preselection; 22. The method of claim 21.

23. The metadata information corresponding to each preselection related box further comprises: Preselection identification information indicating a preselection identifier for identifying each preselection, or at least one of unique preselection specific data for decoding tracks according to said preselection; 23. The method of claim 22.

24. encapsulating the one or more elementary streams further includes generating one or more track group boxes, each associated with a respective track group identifier and a respective track group type that jointly identify a respective track group in the packetized media stream, wherein tracks having the same track group identifier and the same track group type belong to the same track group; Generating the one or more preselection associated boxes includes: assigning a first unique identifier to each preselection; generating, for each track contributing to each preselection, a respective preselection association box associated with the respective preselection, the preselection association box instantiating a track group box having a predefined track group type associated with the preselection, and setting a track group identifier to the first unique identifier; 22. The method of claim 21.

25. 25. The method of claim 24, wherein the track group boxes are generated by grouping tracks that contribute to a preselection based on the tracks' respective media types.

26. 26. The method of claim 25, wherein the media type includes at least one of audio, video, and subtitles.

27. receiving at least one input media; processing the input media to generate the one or more elementary streams; the one or more elementary streams include the media data of the input media and the corresponding header information; 22. The method of claim 21.

28. The method comprises: generating a manifest file based on the one or more preselection association boxes; 22. The method of claim 21.

29. 29. The method of claim 28, wherein the manifest file is an MPEG Dynamic Adaptive Streaming over HTTP (DASH) format file or an HTTP Live Stream (HLS) format file.

30. 1. A method for processing a media stream, comprising: receiving the media stream packetized according to a predefined transport format, the packetized media stream including a plurality of hierarchical boxes each associated with a respective box type identifier, the plurality of hierarchical boxes including one or more track boxes referencing respective tracks indicating media components of the media stream, and one or more preselection-related boxes of a predefined type, each preselection-related box indicating a respective preselection corresponding to a media presentation to a user; generating a manifest file based on the one or more preselection association boxes; the media stream further comprises processing information indicating how tracks contributing to the preselection should be processed; the processing information includes sequencing information indicating a track order for processing the one or more tracks; the processing information includes merge information indicating whether one or more tracks should be merged with one or more other tracks for joint processing; the ordering information includes, for each track contributing to the preselection, a respective track order value for defining a track order for the tracks; the merge information includes a respective merge flag for each track contributing to the preselection, a first setting of the merge flag indicating that the respective track should be merged with an adjacent track in track order, and a second setting of the merge flag indicating that the respective track should be processed separately; method.

31. 31. The method of claim 30, wherein the manifest file is an MPEG Dynamic Adaptive Streaming over HTTP (DASH) format file or an HTTP Live Stream (HLS) format file.

32. 31. The method of claim 30, wherein the media presentations to the user are characterized by respective configurations of language, type, and / or one or more media-specific attributes of the media stream.

33. 31. The method of claim 30, wherein the predefined transport format is the ISO Base Media File Format (ISOBMFF).

34. 34. A media stream processing device comprising a processor and a memory coupled to the processor, the processor adapted to cause the media stream processing device to perform a method according to any one of claims 1 to 33.

35. A program comprising instructions which, when executed by a processor, cause the processor to carry out a method according to any one of claims 1 to 33.

36. A computer-readable storage medium storing the program according to claim 35.

Citation Information

Patent Citations

  • Method, apparatus and computer program for optimizing transmission of portions of encapsulated media content

    JP2022522388A

  • An apparatus, a method and a computer program for omnidirectional video

    US20200260063A1

  • Signaling of Preselection Information in Media Files Based on a Movie-level Track Group Information Box

    US20230362415A1