CONTENT CONSUMPTION DEVICE SYSTEM
Patent Information
- Application Number
- ARP20220100074
- Authority / Receiving Office
- AR · AR
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-10
- Filing Date
- 2022-01-14
- Publication Date
- 2026-08-26
- Estimated Expiration
- 2038-10-12
Abstract
Description
METHOD AND APPARATUS FOR EFFICIENT DELIVERY AND USE OF AUDIO MESSAGES FOR HIGH QUALITY EXPERIENCE Description 1. Introduction In many applications, delivering audible messages can enhance the user experience during media consumption. One of the most relevant applications of such messages is in Virtual Reality (VR) content. In a VR environment, or similarly in Augmented Reality (AR), Mixed Reality (MR), or 360-degree video environments, the user can typically view the full 360-degree content using, for example, a Head-Mounted Display (HMD) and listen to it through headphones (or speakers, with correct rendering depending on their position). The user can usually move around within the VR / AR space, or at least change the viewing direction—the so-called "viewing window" for video.In 360-degree video environments that use traditional systems (flat-screen displays) instead of HMDs, remote control devices can be employed to emulate the user's movement within the scene, and similar principles apply. It should be noted that 360-degree content can refer to any type of content that comprises more than one viewing angle simultaneously, which the user can select (for example, by tilting their head or using a remote control device). In comparison to traditional content consumption, with VR content creators can no longer control what the user sees in various ways. 238315 1640418 of 78 moments - the current display window. The user is free to choose different display windows at any given time, from among the allowed or available display windows. A common problem with VR content consumption is the risk of users missing important events in the video scene due to selecting the wrong viewing window. To address this issue, the concept of Region of Interest (ROI) was introduced, and several approaches to ROI signaling are considered. While ROI is typically used to indicate to the user the region containing the recommended viewing window, it can also be used for other purposes, such as: indicating the presence of a new character / object in the scene, indicating accessibility features associated with objects in the scene, or essentially any feature that might be associated with an element within the scene. For example, visual messages (e.g., "Turn your head to the left") can be overlaid on the current viewing window.On the other hand, audible sounds, whether natural or synthetic, can be used and played at the ROI location. These audio messages are known as Earcons (short, distinctive sounds used to represent an event). In the context of this application, the Earcon concept is used to characterize the audio messages transmitted to signal ROIs, although the proposed signaling and processing can also be used for generic audio messages for purposes other than ROI signaling. An example of such an audio message is one that conveys information / indications about various options available to the user in an interactive AR / VR / MR environment (e.g., skip the checkout). 238315 1640418 of 78 left to enter room X). In addition, the VR example is used, although the mechanisms described in this document apply to any media consumption environment. 2. Terminology and Definitions The following technical terminology is used: • Audio elements: Audio signals that can be represented, for example, as audio objects, audio channels, scene-based audio (Higher Order Ambisonics, HOA), or a combination of all of them. • Region of Interest (ROI): A region of the video content (or the displayed or simulated environment) that is of interest to the user at a given time. This can commonly be a region of a sphere, for example, or a polygonal selection on a 2D map. The ROI identifies a specific region for a given purpose, defining the boundaries of an object under consideration. • Information about the user's position: location information (e.g., x, y, z coordinates), orientation information (howl, tone, drum roll), direction and speed of movement, etc. • Viewing window: The portion of the spherical video that is currently being displayed and viewed by the user. • Viewpoint: the central point of the Viewport. • 360-degree video (also known as immersive video or spherical video): represents, in the context of this document, video content that contains more than one view (i.e., viewing window) in one direction at the same time. 238315 1640418 of 78 content can be generated, for example, using an omnidirectional camera or a collection of cameras. During playback, the viewer has control of the viewing direction. • Adaptation streams contain one or more media streams. In the simplest case, an Adaptation Stream contains all the audio and video corresponding to the content, but to reduce bandwidth, each stream can be split into a different Adaptation Stream. A common scenario involves one video Adaptation Stream and multiple audio Adaptation Streams (one for each supported language). Adaptation Streams can also contain subtitles or arbitrary metadata. • Representations allow an Adaptation Series to contain the same content encoded in different ways. In most cases, representations are presented at multiple bitrates. This enables clients to request the highest quality content they can play without waiting for buffering. Representations can also be encoded with different codecs, resulting in support for clients with different supported codecs. • Media Presentation Description (MPD) is an XML syntax that contains information about media segments, their relationships, and the information needed to choose between them. In the context of this application, the concepts of Adaptation Series are used more generically, sometimes actually referring to Representations. Furthermore, media streams (audio / video streams) are generally encapsulated first into Media Segments, which are the 238315 1640418 of 78 actual media files played by the client (e.g., DASH client). Various formats can be used for Media Segments, such as the ISO Base Media File Format (ISOBMFF), which is similar to the MPEG (Moving Picture Experts Group)-4 content format, and MPEG-TS. Encapsulation into Media Segments and different Adaptation Representations / Series is independent of the methods described here, as the methods apply to all the various options. Furthermore, the description of the methods in this document can be focused around DASH Server-Client communication, although the methods are generic enough to work with other playback environments, such as MMT, MPEG-2 Transport Stream, DASH ROUTING, File Format for file playback, etc. 3. Current Solutions The current solutions are: [1] . ISO / IEC 23008-3:2015, Information technology - High-efficiency media encoding and transmission in heterogeneous environments -- Part 3: 3D audio [2] . N16950, Study of the ISO / IEC DIS Omnidirectional Media Format 23000-20 [3] . M41184, Use of Earcones for ROI identification in 360 degree video. A transmission mechanism for 360-degree content is defined by the ISO / IEC DIS 23000-20 Omnidirectional Media Format [2]. This standard specifies the media format for encoding, storing, transmitting, and rendering omnidirectional images, video, and associated audio. It provides information on the media codecs to be used. 238315 Section 1640418 of 78 addresses audio and video compression and provides additional metadata information for the proper consumption of 360-degree A / V content. It also specifies restrictions and requirements for transmission channels, such as streaming (real-time transmission) via DASH / MMT or file-based playback. The Earcon concept was first introduced in M41184, Use of Earcons for ROI Identification in 360-degree Video [3], which presents a mechanism for signaling Earcon Audio data to the user. However, some users have reported disappointing feedback about these systems. A large number of ear cones has often proven irritating. When designers reduced the number of ear cones, some users lost important information. Notably, each user has their own knowledge and experience level and would prefer a system tailored to their specific needs. For example, each user would prefer the ear cones to play at a preferred volume (independent, for instance, of the volume used for other audio signals). It has proven difficult for the system designer to create a system that satisfies all potential users. Therefore, a solution has been sought to increase satisfaction for nearly all users. Furthermore, reconfiguring the systems has proven difficult, even for the designers. For example, they have experienced difficulties preparing new Audio Stream releases and updating the Earcons. Furthermore, a restricted system imposes certain limitations on functionality, such as the inability to accurately identify earcones within an audio stream. Moreover, earcones must always be active and 238315 1640418 of 78 can become irritating to the user if played when not needed. Furthermore, the spatial information of the Earcones cannot be signaled or modified, for example, by a DASH Client. Easy access to this information at the Systems level could enable an additional feature for an improved user experience. Furthermore, there is no flexibility to address different types of Earcones (e.g., natural sound, synthetic sound, sound generated in the DASH Client, etc.). All these drawbacks lead to a poor user experience. Therefore, a more flexible architecture would be preferable. 4. The present invention According to the examples, a system for a virtual reality, VR, augmented reality, AR, mixed reality, MR, or 360-degree video environment is presented, configured to: to receive at least one Video Stream associated with an audio and video scene to be played and to receive at least one first Audio Stream associated with the audio and video scene to be played, wherein the System comprises: at least one media video decoder configured to decode at least one video signal from said at least one Video Stream for rendering the audio and video scene to a user and at least one media audio decoder configured to decode at least one Audio signal from said at least one first Audio Stream for the representation of the audio and video scene to the user; a region of interest (ROI) processor configured to: 238315 1640418 of 78 decide, based at least on the user's current display window data, and / or head orientation and / or movement and / or display window metadata and / or audio information message metadata, whether an audio information message associated with at least one ROI should be played or not, wherein the audio information message is independent of said at least one video signal and said at least one audio signal and cause, upon the decision that the information message should be played, the playback of the audio information message. In accordance with the examples, a system for a virtual reality, VR, augmented reality, AR, mixed reality, MR or 360-degree video environment is presented configured to: to receive at least one Video Stream and to receive at least one first Audio Stream, where the system comprises: at least one media video decoder configured to decode at least one video signal from said at least one Video Stream for the representation of a scene in a VR, AR, MR or 360-degree Video environment to a user and at least one media audio decoder configured to decode at least one Audio signal from said at least one first Audio Stream for the representation of an Audio scene to the user; a region of interest (ROI) processor configured to: decide, based on the user's current display window and / or their head orientation and / or motion data and / or display window metadata and / or audio information message metadata, whether an audio information message associated with 238315 1640418 of 78 the at least one ROI must be played or not, where the audio information message is an Earcon and cause, upon the decision that the information message should be played, the playback of the Audio Information Message. The system may consist of: a metadata processor configured to receive and / or process and / or manipulate audio information message metadata in order to cause, upon the decision that the information message should be played, the playback of the Audio Information Message in accordance with the audio information message metadata. The ROI processor can be configured to: to receive data from a user's current viewing window and / or their position and / or head orientation and / or movement and / or other related data and to receive metadata from the current viewing window associated with at least one video signal from said at least one Video Stream, wherein the viewing window metadata defines at least one ROI and to decide, based on at least one of the data from the user's current viewing window and / or their position and / or head orientation and / or movement and the viewing window metadata and / or other criteria, whether an audio information message associated with the at least one ROI should be played or not. The system may comprise: a metadata processor configured to receive and / or process and / or manipulate audio information message metadata describing the audio information message and / or audio metadata describing at least one audio signal encoded in at least one stream of 238315 1640418 of 78 Audio and / or display window metadata, in order to generate the playback of the Audio information message in accordance with the audio information message metadata and / or Audio metadata that describe at least one Audio signal encoded in at least one Audio Stream and / or display window metadata. The ROI processor can be configured to: If at least one ROI is outside the user's current viewing window data and / or their head position and / or orientation and / or movement, generate the playback of an audio information message associated with at least one ROI, in addition to the playback of said at least one audio signal; and if at least one ROI is within the user's current viewing window and / or the head position and / or orientation and / or movement data, disable and / or deactivate the playback of the audio information message associated with at least one ROI. The system can be configured to: to receive said at least one additional Audio Stream in which said at least one additional Audio information message is encoded, wherein the system further comprises: at least one muxer or multiplexer to merge, under the control of the metadata processor and / or the ROI processor and / or another processor, packets of said at least one additional Audio Stream with packets of said at least one first Audio Stream to form a Stream, based on the decision provided by the ROI processor that said at least one additional Audio information message should be played, to generate the playback of the Audio information message in addition to the Audio scene. 238315 1640418 of 78 The system can be configured to: receive at least one Audio metadata describing said at least one Audio signal encoded in said at least one Audio Stream; receive Audio Information Message Metadata associated with at least one additional Audio Information Message from at least one Audio Stream; Given the decision that the information message should be played, modify the Audio Information Message Metadata to enable playback of the Audio Information Message, in addition to playing at least one Audio signal. The system can be configured to: receive at least one Audio metadata describing said at least one Audio signal encoded in said at least one Audio Stream; receive Audio Information Message Metadata associated with at least one additional Audio Information Message from said at least one Audio Stream; Given the decision that the audio information message should be played, modify the Audio Information Message Metadata to enable the playback of an Audio Information Message in connection with said at least one ROI, in addition to the playback of said at least one Audio signal, and modify the Audio Metadata describing said at least one Audio signal to allow a merging of said at least one first Audio Stream and said at least one additional Audio Stream. The system can be configured to: receive at least one Audio metadata describing said at least one Audio signal encoded in said at least one Audio Stream; 238315 1640418 of 78 receive Audio Information Message Metadata associated with at least one additional Audio Information Message from at least one Audio Stream; Upon the decision that the audio information message should be played, send the Audio Information Message Metadata to a Synthetic Audio Generator to create a Synthetic Audio Stream, in order to associate the Audio Information Message Metadata with the Synthetic Audio Stream, and to send the Synthetic Audio Stream and the Audio Information Message Metadata to a multiplexer or muxer to result in the merging of said at least one Audio Stream and the Synthetic Audio Stream. The system can be configured to: Obtain the audio information message metadata from at least one additional audio stream in which the audio information message is encoded. The system can include: An Audio Information Message Metadata Generator configured to generate Audio Information Message Metadata based on the decision that the Audio Information Message associated with at least one ROI should be played. The system can be configured to: Store, for future use, the Audio Information Message Metadata and / or the Audio Information Message Flow. The system can include: A synthetic audio generator configured to synthesize an audio information message based on audio information message metadata associated with at least one ROI. 238315 1640418 of 78 The metadata processor may be configured to control a muxer or multiplexer to merge, based on Audio metadata and / or audio information message metadata, packets from the Audio Information Message Stream with packets from said at least one first Audio Stream to form a Stream in order to obtain the addition of the Audio Information Message to said at least one Audio Stream. Audio information message metadata can be encoded in a configuration frame and / or a data frame that includes at least one of: an identification tag, an integer that uniquely identifies the playback of the Audio Metadata information message, the message type, a status, an indication of scene dependence / independence, position data, gain data, an indication of the presence of an associated text tag, number of available languages, language of the Audio information message, length of data text, data text of the associated text tag and / or description of the Audio information message. The metadata processor and / or the ROI processor can be configured to perform at least one of the following operations: Extract metadata from audio information messages in a stream; 238315 1640418 of 78 Modify Audio Information Message Metadata to activate the Audio Information Message and / or set / change its position; re-include the metadata in a Flow; feed the Stream to an additional media decoder; extract Audio Metadata from at least the first Audio Stream; Extract metadata from audio information messages from an additional stream; Modify audio information message metadata to activate the audio information message and / or set / change its position; modify Audio metadata of said at least one first Audio Stream in order to take into account the existence of the Audio information message and allow merging; Feed a flow to the multiplexer or muxer to multiplex it based on the information received from the ROI processor. The ROI processor can be configured to run a local search for an additional Audio Stream in which the Audio Information message and / or audio information message metadata is encoded and, if it is not obtained, request the additional Audio Stream and / or audio information message metadata from a remote entity. The ROI processor can be configured to run a local search for an additional Audio Stream and / or Audio Information Message Metadata and, if it does not obtain them, have a synthetic Audio generator generate the Audio Information Message Stream and / or Audio Information Message Metadata. The system can be configured to: 238315 1640418 of 78 receive said at least one additional Audio Stream in which is included at least one additional Audio Information message associated with said at least one ROI and decode said at least one additional Audio Stream if the ROI processor decides that an Audio Information message associated with the at least one ROI should be played. The system can include: at least one first Audio decoder to decode said at least one Audio signal from at least one first Audio Stream; at least one additional Audio decoder to decode said at least one additional Audio information message from an additional Audio Stream and at least one mixer and / or renderer to mix and / or overlay the Audio information message from said at least one additional Audio Stream with said at least one Audio signal from said at least one first Audio Stream. The system may be configured to keep track of the metric associated with the historical and / or statistical data associated with the playback of the Audio Information message, in order to disable playback of the Audio Information message if the metric exceeds a predetermined threshold. The ROI processor's decision can be based on the prediction of the user's current viewing window and / or head position and / or orientation data and / or movement relative to the ROI position. The system may be configured to receive at least one first Audio Stream and, upon deciding that the information message should be 238315 1640418 of 78 played, request an Audio Message Information Flow to a remote entity. The system can be configured to determine whether two Audio Information Messages should be played simultaneously or whether a higher priority Audio Information Message should be selected to be played with priority over a lower priority Audio Information Message. The system can be configured to identify an Audio Information Message among a plurality of Audio Information Messages encoded in an additional Audio Stream based on the direction and / or position of the Audio Information Messages in an Audio Stream. Audio Streams can be formatted in the MPEG-H 3D Audio Stream format. The system can be configured to: receiving data about the availability of a plurality of adaptation series, wherein the available adaptation series include at least one Audio adaptation series for said at least one first Audio Stream and at least one Audio message adaptation series corresponding to said at least one additional Audio Stream containing at least one additional Audio information message; create, based on the ROI processor's decision, selection data that identifies which of the adaptation series should be obtained, where the available adaptation series include at least one Audio adaptation series and / or at least one Audio message adaptation series, and request and / or obtain the data corresponding to the adaptation series identified by the selection data, 238315 1640418 of 78 where each adaptation series groups different encodings corresponding to different bit rates. The system may be such that at least one of its elements comprises a Real-Time Adaptive Dynamic Transmission through an HTTP client, DASH and / or is configured to obtain the data corresponding to each of the adaptation series using the ISO Base Media File Format, ISO BMFF, or the MPEG-2 Transport Stream, MPEG-2 TS. The ROI processor can be configured to check for correspondences between the ROI and the current display window and / or head position and / or orientation and / or movement data in order to verify whether the ROI is represented in the current display window, and, if the ROI is outside the current display window and / or head position and / or orientation and / or movement data, to audibly signal the presence of the ROI to the user. The ROI processor may be configured to check for correspondences between the ROI and the current display window and / or head position and / or orientation and / or movement data in order to verify whether the ROI is represented in the current display window, and, if the ROI is within the current display window and / or head position and / or orientation and / or movement data, refrain from audibly indicating the presence of the ROI to the user. The system may be configured to receive, from a remote entity, said at least one Video Stream associated with the video environment scene and said at least one Audio Stream associated with the audio scene, where the audio scene is associated with the video environment scene. The ROI processor can be configured to choose, from a plurality of Audio information messages to be played, the playback 238315 1640418 of 78 of a first audio information message before a second audio information message. The system may include a cache or buffer memory to store an audio information message received from a remote or synthetically generated entity, to reuse the audio information message at different times. The Audio information message may be an Earcon. At least one Video Stream and / or at least one first Audio Stream may be part of the current video environment scene and / or the video audio scene, respectively, and independent of the user's current viewing window and / or the orientation and / or head movement data of the current video environment scene and / or the video audio scene. The system may be configured to request said at least one first Audio Stream and / or at least one Video Stream from a remote entity in connection with the Audio Stream and / or the video environment stream, respectively, and to play said at least one additional Audio information message on the basis of the user's current viewing window and / or head orientation and / or movement data. The system may be configured to request at least one first Audio Stream and / or at least one Video Stream from a remote entity in connection with the Audio Stream and / or the video environment stream, respectively, and to request from the remote entity at least one additional Audio information message based on the user's current viewing window and / or head orientation and / or movement data. The system may be configured to request at least one first Audio Stream and / or at least one Video Stream from a remote entity in 238315 1640418 of 78 connection with the Audio Stream and / or the video environment stream, respectively, and to synthesize said at least one additional Audio information message based on the user's current display window and / or head orientation and / or movement data. The system may be configured to check at least one of the additional criteria for playback of the Audio Information message, where the criteria include a user selection and a user configuration. The system may be configured to check at least one additional criterion for playback of the Audio Information message, where the criteria include the system status. The system may be configured to check at least one additional criterion for playback of the Audio Information message, where the criteria also include the number of playbacks of Audio Information messages that have already been performed. The system may be configured to check at least one additional criterion for playback of the Audio information message, where the criteria also include a flag in a data stream obtained from a remote entity. In accordance with one aspect, a system is presented comprising a client configured as a system of any of the above and / or below examples, and a remote entity configured as a server to transmit said at least one Video Stream and said at least one Audio Stream. The remote entity may be configured to search, in a database, intranet, internet, and / or a geographic network, for at least one additional Audio Stream and / or audio information message metadata and, in the case of 238315 1640418 of 78 obtaining, the transmission of said at least one additional Audio Stream and / or the Audio Information Message Metadata. The remote entity may be configured to synthesize at least one additional Audio Stream and / or generate Audio Information Message Metadata. According to one aspect, a method for a virtual reality, VR, augmented reality, AR, mixed reality, MR, or 360-degree video environment may be presented comprising: decode at least one video signal from said at least one video and audio scene to be reproduced to a user; decode at least one audio signal from the video and audio scene to be reproduced; to decide, based on the user's current viewing window and / or orientation and / or head movement data and / or metadata, whether an audio information message associated with at least one ROI should be played or not, wherein the audio information message is independent of said at least one video signal and said at least one audio signal and cause, upon the decision that the information message should be played, the playback of the audio information message. According to one aspect, a method for a virtual reality, VR, augmented reality, AR, mixed reality, MR, or 360-degree video environment may be presented comprising: decode at least one video signal from said at least one Video Stream for the representation of a scene in a VR, AR, MR or 360-degree Video environment to a user; 238315 1640418 of 78 decode at least one Audio signal from said at least one first Audio Stream for the representation of an Audio scene to the user; decide, based on the user's current viewing window and / or orientation and / or head movement data and / or metadata, whether an audio information message associated with at least one ROI should be played or not, where the audio information message is an Earcon and cause, upon the decision that the information message should be played, the playback of the Audio Information Message. The preceding and / or following methods may include: to receive and / or process and / or manipulate metadata in order to cause, upon the decision that the information message should be played, the playback of the Audio information message in accordance with the metadata in such a way that the Audio information message is part of the Audio scene. The preceding and / or following methods may include: Play the audio and video scene and decide to also play the Audio Information message based on the user's current viewing window and / or orientation and / or head movement data and / or metadata. The preceding and / or following methods may include: Play the audio and video scene and, if at least one ROI is outside the user's current viewing window and / or the head position and / or orientation and / or movement data, generate the playback of an audio information message associated with at least one ROI, in addition to the playback of said at least one audio signal and / or 238315 1640418 of 78 in case that at least one ROI is within the user's current viewing window and / or head position and / or orientation and / or movement data, disable and / or deactivate the playback of the Audio information message associated with at least one ROI. In accordance with the examples, a system for a virtual reality, VR, augmented reality, AR, mixed reality, MR or 360-degree video environment is presented configured to: to receive at least one Video Stream and to receive at least one first Audio Stream, where the system comprises: at least one media video decoder configured to decode at least one video signal from said at least one Video Stream for the representation of a scene in a VR, AR, MR or 360-degree Video environment to a user and at least one media audio decoder configured to decode at least one Audio signal from said at least one first Audio Stream for the representation of an Audio scene to the user; a region of interest (ROI) processor configured to: to decide, based on the user's current viewing window and / or orientation and / or head movement data and / or metadata, whether an audio information message associated with at least one ROI should be played or not, and to cause, upon the decision that the information message should be played, the playback of the audio information message. The examples present a system for a virtual reality, VR, augmented reality, AR, mixed reality, MR or 360-degree video environment configured to: 238315 1640418 of 78 receive at least one Video Stream and receive at least one first Audio Stream, where the system comprises: at least one media video decoder configured to decode at least one video signal from said at least one Video Stream for the representation of a scene in a VR, AR, MR or 360-degree Video environment to a user and at least one media audio decoder configured to decode at least one Audio signal from said at least one first Audio Stream for the representation of an Audio scene to a user; A region of interest (ROI) processor configured to decide, based on the user's current viewing window and / or head position and / or orientation and / or motion data and / or metadata and / or other criteria, whether an audio information message associated with at least one ROI should be played or not, and a metadata processor configured to receive and / or process and / or manipulate metadata in order to cause, upon the decision that the information message should be played, the playback of the audio information message in accordance with the metadata in such a way that the audio information message is part of the audio scene. According to one aspect, a non-transient storage unit is presented that comprises instructions which, when executed by a processor, cause the processor to execute a method according to what has already been stated and / or what is stated below. 5. Description of the drawings Figures 1-5, 5a, and 6 illustrate examples of implementations; Fig. 7 illustrates one method according to an example; 238315 1640418 of 78 Fig. 8 illustrates an example of implementation. 6. Examples 6.1 General Examples Fig. 1 illustrates an example of a 100 system for a virtual reality, VR, augmented reality, AR, mixed reality, MR or 360-degree video environment. System 100 can be associated, for example, with a content consumption device (e.g., Head-Mounted Display or similar), which reproduces visual data on a spherical or semi-spherical screen closely associated with the user's head. System 100 may comprise at least one media video decoder 102 and at least one media audio decoder 112. System 100 may receive at least one Video Stream 106 in which a Video signal is encoded for the representation of a scene in a VR, AR, MR, or 360-degree Video environment 118a to a user. System 100 may receive at least one first Audio Stream 116 in which an Audio signal is encoded for the representation of an Audio scene 118b to a user. System 100 may further comprise a Region of Interest (ROI) processor, 120. The ROI processor 120 can process data associated with an ROI. Generally speaking, the presence of the ROI can be signaled in the display window metadata 131. The display window metadata 131 can be encoded in Video Stream 106 (in other cases, the display window metadata 131 can be encoded in other Streams). The display window metadata 131 can comprise, for example, position information (e.g., coordinate information) associated with the ROI. For example, the ROI in the examples can be understood as a rectangle (identified by coordinates, such as the position of one of the four vertices of the rectangles in the spherical video and the length of 238315 1640418 of 78 sides of the rectangle). The ROI is normally projected onto the spherical video. The ROI is usually associated with a visible element that is believed (according to a specific configuration) to be of interest to the user. For example, the ROI may be associated with a rectangular surface displayed by the content consumption device (or otherwise visible to the user). The ROI 120 processor can control, among other things, the operations of the Media Audio Decoder 112. The ROI 120 processor can obtain data 122 associated with the user's current viewing window and / or head position and / or orientation and / or movement data (in some examples, virtual data associated with virtual position can also be considered part of data 122). This data 122 can be supplied, at least in part, by the content consumption device or by positioning / detection units. The ROI processor 120 can verify the correspondence between the ROI and the user's current display window and / or position data (real or virtual) and / or orientation and / or head movement 122 (other criteria may be used in the examples). For example, the ROI processor can verify whether the ROI is represented in the current display window. If an ROI is only partially represented in the display window (e.g., based on the user's head movements), it can be determined, for example, if only a minimal percentage of the ROI appears on the screen. In any case, the ROI processor 120 is capable of recognizing whether the ROI is not represented or is visible to the user. If the ROI is deemed to be outside the user's current viewing window and / or the head position and / or orientation and / or movement data 122, the ROI processor 120 may audibly signal the 238315 1640418 of 78 presence of the ROI to the user. For example, the ROI processor 120 may request the playback of an Audio Information (Earcon) message in addition to the decoded audio signal of said at least a first Audio Stream 116. If the ROI is deemed to be within the user's current viewing window and / or head position and / or orientation and / or movement data 122, the ROI processor may decide to avoid playing the Audio information message. The Audio Information message can be encoded in an Audio Stream 140 (Audio Information Message Stream), which can be the same as Audio Stream 116 or a different stream. Audio Stream 140 can be generated by the system or obtained from an external entity (e.g., a server). Audio metadata, such as Audio Information Message Metadata 141, can be defined to describe the properties of Audio Information Stream 140. The Audio Information message may be overlaid (or mixed or multiplexed or fused or combined or composed) with the encoded signal in Audio Stream 116 or may not be selected, e.g., simply based on a decision by the ROI processor 120. The ROI processor 120 may base its decision on display window data and / or head position and / or orientation and / or motion data 122, metadata (such as display window metadata 131 or other metadata) and / or other criteria (e.g., selections, system status, number of Audio Information message playbacks already performed, specific functions and / or operations, user preferred settings that may disable the use of Earcons, and so forth). 238315 1640418 of 78 A metadata processor 132 can be implemented. The metadata processor 132 can be interposed, for example, between the ROI processor 120 (by which it can be controlled) and the media audio decoder 112 (which can be controlled from the metadata processor). In the examples, the metadata processor is a section of the ROI processor 120. The metadata processor 132 can receive, generate, process, and / or manipulate metadata from audio information messages 141. The metadata processor 132 can also process and / or manipulate metadata from Audio Stream 116, for example, to multiplex Audio Stream 116 with Audio Information Message Stream 140. Furthermore, the metadata processor 132 can receive metadata from Audio Stream 116, for example, from a server (e.g., a remote entity). Therefore, the metadata processor 132 can change the playback of Audio scenes and adapt the Audio information message to specific situations and / or selections and / or states. Some advantages of certain implementations are discussed here. Audio information messages can be accurately identified, for example, using audio information message metadata 141. Audio information messages can be easily enabled / disabled, for example, by modifying the metadata (e.g., using metadata processor 132). Audio information messages can be enabled / disabled, for example, based on the current display window and ROI information (in addition to desired functions or special effects). The audio information message (containing, for example, status, type, spatial information, and so on) can be easily signaled and modified using common equipment, such as Dynamic Adaptive Transmission in 238315 1640418 of 78 Real-time (Dynamic Adaptive Streaming) through an HTTP Client (DASH), for example. Therefore, easy access to the Audio information message (containing, for example, status, type, spatial information, and more) at the system level can enable other features for an improved user experience. Consequently, System 100 can be easily adapted to user needs and allow for other implementations (e.g., specific applications) that can be carried out by personnel independent of the System 100 designers. Furthermore, flexibility is achieved to address various types of audio information messages (e.g., natural sound, synthetic sound, sound generated in the DASH Client, etc.). Other advantages (which are also evident in the following examples): • Use of text tags in metadata (as a basis for displaying something or generating the Earcon) • Adaptation of the Earcon position based on the device (if it is an HMD I want a precise location, if it is a speaker perhaps it is best to use a different location - directly to a speaker). • Different types of devices: Earcon metadata can be generated in such a way as to signal that the Earcon is active. Some devices know how to parse the metadata and reproduce the Earcon. 238315 1640418 of 78 or Some newer devices that also have a better ROI processor may decide to disable it if it is not needed • More information and an additional figure about adaptation series. Therefore, in a VR / AR environment, the user can typically view all 360-degree content using, for example, a Head-Mounted Display (HMD) and listen to it with headphones. The user can usually move around the VR / AR space, or at least change the viewing direction—the so-called “viewing window” for the video. Compared to traditional content consumption, VR content creators can no longer control what the user views at different points in time—the current viewing window. The user is free to choose different viewing windows at each point in time, from among the permitted or available viewing windows. To indicate the region of interest (ROI) to the user, audible sounds, either natural or synthetic, can be played at the ROI position. These audio cues are known as “earcones.”This invention proposes a solution for the efficient playback of these messages and proposes optimized receiver behavior to utilize Earcones without impacting the user experience or content consumption. This leads to an enhanced Quality of Experience. This can be achieved by using specialized metadata and metadata manipulation mechanisms at the system level to enable or disable Earcones in the final scene. The metadata processor 132 may be configured to receive and / or process and / or manipulate metadata 141 in order to cause, upon the decision that 238315 1640418 of 78 The information message must be played, the playback of the Audio information message according to metadata 141. It can be understood that the audio signals (e.g., those used to represent the scene) are parts of the audio scene (e.g., an Audio scene downloaded from a remote server). The audio signal can be semantically significant, in general, to the audio scene, and all the audio signals present together construct the audio scene. The audio signals can be encoded together in an audio bitstream. The audio signal can be generated by the content creator and / or can be associated with a specific scene and / or can be independent of the ROI. It can be understood that the audio information message (e.g., the Earcon) may not be semantically significant to the audio scene. It can be considered an independent sound that can be artificially generated, such as a recorded sound, a recorded voice, etc. It can also be device-dependent (a system sound generated when a remote control button is pressed, for example). The audio information message (e.g., Earcon) can be understood as intended to guide the user through the scene, without being part of it. The audio information message can be independent of the audio signals as described. According to different examples, it can be included in the same bitstream, transmitted in a separate bitstream, or generated by the system. An example of an audio scene composed of multiple audio signals could be: -- Audio scene, a concert hall containing 5 audio signals: — Audio signal 1: The sound of a piano — Audio signal 2: The singer's voice 238315 1640418 of 78 — Audio signal 3: The voice of Person 1, part of the audience — Audio signal 4: The voice of Person 2, part of the audience — Audio signal 5: The sound created by the clock on the wall The audio information message can be, for example, a recorded sound such as “look at the pianist” (where the piano is the ROI). If the user is already looking at the pianist, the audio message is not played. Another example: a door (e.g., a virtual door) opens behind the user, and a new person enters the room; the user isn't looking in that direction. The Earcon can be triggered, based on this information (regarding the VR environment, such as the virtual position), to alert the user that something is happening behind them. In the examples, each scene (e.g., with the related audio and video streams) is transmitted from the server to the client when the user changes the environment. The audio information message can be flexible. In particular: - The audio information message can be located in the same audio stream associated with the scene to be played; - The audio information message may be located in an additional audio stream; - the Audio Information message may be completely absent, and only the metadata describing the Earcon may be present in the stream and the Audio Information message may be generated in the system; - The Audio Information message may be completely absent, as well as the metadata describing the Audio Information message, in which case the system generates both (the Earcon and the metadata) based on other information about the ROI in the flow. 238315 1640418 of 78 The Audio Information message is generally independent of any audio signal that makes up the audio scene and is not used for the representation of the audio scene. The following are examples of systems that incorporate or include parts that constitute system 100. 6.2 The example in Fig. 2 Figure 2 illustrates a system 200 (which may contain at least one part incorporating system 100) that in this case is represented as subdivided into a server side 202, a media transmission or delivery side 203, a client side 204, and / or a media consumption device side 206. Each of the sides 202, 203, 204, and 206 is a system in itself and can be combined with any other system to obtain another system. In this case, audio information messages are referred to as Earcons, although it is possible to generalize this to any type of audio information message. The client side 204 may receive said at least one Video Stream 106 and / or said at least one Audio Stream 116 from the server side 202 through a media delivery side 203. The delivery side 203 can be based, for example, on a communications system such as a cloud system, a networked system, a geographic communications network, or well-known media transport formats (MPEG-2 Transport Stream TS, DASH, MMT, DASH ROUTE, etc.), or even file-based storage. The delivery side 203 can be capable of executing communications in the form of electronic signals (e.g., wired, wireless, etc.) and / or by distributing data packets (e.g., according to a specific communications protocol) with bitstreams in which audio and video signals are encoded. 238315 1640418 of 78 However, the delivery side 203 may consist of a point-to-point link, a serial or parallel connection, and so on. The delivery side 203 may implement a wireless connection, e.g., according to protocols such as WiFi, Bluetooth, and so on. Client-side port 204 can be associated with a media consumption device, e.g., a head-mounted display (HMD), into which the user's head can be inserted (although other devices can be used). Therefore, the user can experience a video and audio scene (e.g., a VR scene) prepared by client-side port 204 based on video and audio data provided by server-side port 202. However, other implementations are possible. The server side 202 is represented, in this case, by a media encoder 240 (which can encompass video encoders, audio encoders, subtitle encoders, etc.). This encoder 240 can be associated, for example, with an audio and video scene to be rendered. The audio scene might, for instance, recreate an environment and is associated with at least one audio and video data stream 106, 116, which can be encoded based on the user's position (or virtual position) within the VR, AR, or MR environment. Generally speaking, the video stream 106 encodes spherical images, only a portion of which (display windows) is perceived by the user based on their position and movements. The audio stream 116 contains audio data that contributes to the rendering of the audio scene and is intended to be heard by the user.According to the examples, Audio Stream 116 may comprise audio metadata 236 (referring to said at least one Audio signal that is intended to participate for the representation of the. 238315 1640418 of 78 audio scene) and / or Earcon 141 metadata (which may describe the Earcones to be played only in some cases). System 100 is shown here located on the customer side 204. For simplicity, the media video decoder 112 is not illustrated in Fig. 2. To prepare the playback of an Earcon (or other audio information messages), Earcon 141 metadata can be used. Earcon 141 metadata can be considered metadata (which can be encoded in an audio stream) that describes and assigns attributes to the Earcon. Therefore, the Earcon (if it is to be played) can be based on the attributes of the Earcon 141 metadata. Advantageously, the metadata processor 132 can be specifically implemented to process Earcon metadata 141. For example, the metadata processor 132 can control the reception, processing, manipulation, and / or generation of Earcon metadata 141. Once processed, the Earcon metadata can be represented as modified Earcon metadata 234. For example, it is possible to manipulate the Earcon metadata to achieve a specific effect and / or to perform audio processing operations, such as multiplexing, to add the Earcon to the audio signal to be represented in the audio scene. The metadata processor 132 can control the reception, processing, and manipulation of the audio metadata 236 associated with at least one Stream 116. Once processed, the audio metadata 236 can be represented as modified audio metadata 238. The modified metadata 234 and 238 can be sent to the Media Audio Decoder 112 (or to a plurality of decoders in some examples) for playback of the audio scene 118b to the user. 238315 1640418 of 78 In the examples, a synthetic audio generator and / or a storage device 246 may be presented as an optional component. The generator can synthesize an audio stream (e.g., to generate an Earcon that is not encoded in a stream). The storage device allows Earcon streams generated by the generator and / or obtained from a received audio stream to be stored (e.g., in a cache). Therefore, the ROI 120 processor can decide the representation of an Earcon based on the user's current viewing window and / or head position and / or orientation and / or movement data 122. However, the ROI 120 processor can also base its decision on criteria involving other aspects. For example, the ROI processor can enable / disable Earcone playback based on other conditions, such as user selections or selections from higher layers, for example, based on the specific application being used. In the case of a video game application, for example, Earcones or other audio information messages for high levels of gameplay can be avoided. This can be achieved simply by the metadata processor disabling Earcones in the Earcon metadata. It's also possible to disable Earcones based on the system state: for example, if an Earcon has already played, its repetition can be prevented. A timer can be used, for instance, to prevent excessively rapid repetition. The ROI 120 processor can also request controlled playback of a sequence of Earcones associated with all ROIs in the scene. 238315 1640418 of 78 ex., to instruct the user about the elements that the person can see. The metadata processor 132 can control this operation. The ROI 120 processor can also modify the Earcon's position (i.e., its spatial location within the scene) or the Earcon type. For example, some users may prefer to have a specific sound played back at the exact location / position of the ROI, while other users may prefer the Earcon to always be played back in a fixed location (e.g., a central or upper position, like a "voice of God") as a vocal indication of the ROI's position. It is possible to modify the Earcon's playback gain (e.g., to obtain a different volume). This decision can be based on a user selection, for example. Notably, based on the ROI processor's decision, the metadata processor 132 executes the gain modification by modifying the specific attribute associated with the gain within the Earcon's metadata. The original designer of the VR, AR, or MR environment may also be unaware of how the Earcones will actually be rendered. For example, user selections can modify the final rendering of the Earcones. This type of operation can be controlled, for example, by the metadata processor 132, which can modify the metadata of Earcon 141 based on the ROI processor's decision. Therefore, the operations performed on the audio data associated with the Earcon are, in principle, independent of the at least one Audio Stream 116 used to represent the audio scene and can be controlled differently. Earcones can even be generated independently of the audio and video streams 106 and 116 that constitute the 238315 1640418 of 78 audio and video scenes and can be produced by different and independent business groups. Therefore, the examples help increase user satisfaction. For example, a user can make their own selections, such as adjusting the volume of audio information messages, disabling audio information messages altogether, and so on. Thus, each user can have an experience tailored to their preferences. Furthermore, the resulting architecture is more flexible. Audio information messages can be easily updated, for example, by modifying the metadata independently of the audio streams, and / or by modifying the audio information message stream independently of the metadata and the main audio streams. The resulting architecture is also backward compatible with legacy systems: Audio Information Message Streams from the previous technique can be associated with new Audio Information Message Metadata, for example. In the absence of a suitable Audio Information Message Stream, the latter can be easily synthesized (and, for example, stored for later use). The ROI processor can keep track of the metric associated with the historical and / or statistical data associated with the playback of the Audio Information message, in order to disable the playback of the Audio Information message if the metric exceeds a predetermined threshold (this can be used as a criterion). The ROI processor's decision can be based, as a criterion, on the prediction of the user's current viewing window and / or head position and / or orientation data and / or movement 122 in relation to the ROI position. 238315 1640418 of 78 The ROI processor may also be configured to receive at least one first Audio Stream 116 and, upon deciding that the information message should be played, request an Audio Message Information Stream from a remote entity. The ROI processor and / or metadata generator can be further configured to determine whether two Audio Information Messages should be played simultaneously or whether a higher-priority Audio Information Message should be selected for playback over a lower-priority one. Audio information metadata can be used to make this decision. A priority can be obtained, for example, from metadata processor 132 based on the values provided in the audio information message metadata. In some examples, the 240 media encoder can be configured to search a database, intranet, internet, and / or geographic network for additional audio streams and / or audio information message metadata, and, if found, output additional audio streams and / or audio information message metadata. For example, the search can be performed on request from the client side. As explained earlier, this paper proposes a solution for the efficient delivery of Earcon messages alongside audio content. This results in optimized receiver behavior, allowing the use of audio information messages (e.g., Earcones) without impacting the user experience or content consumption. This leads to an improved Quality of Experience. This can be achieved by employing specialized metadata and metadata manipulation mechanisms at the system level to enable or disable Audio information messages in audio scenes 238315 1640418 of 78 final. The Metadata can be used in conjunction with any audio codec and favorably complements the metadata of Next Generation audio codecs (e.g., MPEG-H Audio Metadata). Delivery mechanisms can vary (e.g., DASH / HLS streaming, transmission via DASHROUTE / MMT / MPEG-2 TS, file playback, etc.). This application focuses on DASH delivery, although all concepts apply to the other delivery options. In most cases, audio information messages do not overlap in the time domain; that is, only one ROI is defined at any given time. However, considering more advanced use cases, such as in an interactive environment where the user can change content based on their selections or movements, there might also be use cases that require multiple ROIs. For this purpose, more than one audio information message may be needed at a given point in time. Therefore, a generic solution is described to support all different use cases. The delivery and processing of Audio information messages should complement existing delivery methods for Next Generation Audio. One way to transfer multiple Audio Information Messages corresponding to several time-independent ROIs is to merge all the Audio Information Messages together to obtain an Audio Element (e.g., Audio object) with associated metadata describing the spatial position of each Audio Information Message at different points in time. Since the Audio Information Messages do not overlap in time, they can be 238315 1640418 of 78 independently addressed in a shared audio element. This audio element could contain silence (or missing audio data) interspersed between audio information messages, i.e., whenever there is no audio information message. In this case, the following mechanisms can be applied: • The common Audio Information Message Element can be delivered in the same Elementary Stream (ES) as the audio scene it relates to, or it can be delivered in an auxiliary Stream (dependent or not on the main Stream). • If the Earcon Audio Element is delivered in a Flow dependent on the main Flow, the customer can request the additional Flow whenever a new ROI is present in the visual scene. • The Client (e.g., system 100) can request the Flow, in the examples, in advance of the scene that requires the Earcon. • The Customer can request the Flow, in the examples, based on the current display window, meaning that if the current display window matches the ROI the Customer can decide not to request the additional Earcon flow. • If the Earcon audio element can be delivered in a separate auxiliary stream from the main stream, the client can request the additional stream as before, provided there is a new ROI present in the visual scene. Furthermore, both streams (or more) can be processed using two media decoders and a common rendering / mixing step to blend the decoded Earcon audio data into the audio scene. 238315 1640418 of 78 final. On the other hand, a metadata processor can be used to modify the metadata of the two streams, and a stream merge can be used to merge the two streams. A possible implementation of this metadata processor and stream merge is described below. In alternative examples, multiple Earcones can be transmitted for various ROIs, independent in the time domain or overlapping in the time domain, in multiple Audio Elements (e.g., audio objects) and included either in an elemental stream along with the main audio scene or in multiple auxiliary streams, e.g., each Earcon in an Earcon Stream or a group of Earcones in an Earcon Stream based on a shared property (e.g., all Earcones located on the left share a Stream). • If all Earcon audio elements are transmitted in several auxiliary streams dependent on the main stream (e.g., one Earcon per stream or a group of Earcons per stream), the Client can request, in the examples, an additional stream containing the desired Earcon, provided that the ROI associated with that Earcon is present in the visual scene. • The Client can request, in the examples, the stream containing the Earcon in advance of the scene that requires that Earcon (e.g., based on the user's movements, the ROI processor 120 can execute the decision even if the ROI is not yet part of the scene). • In the examples, the Client can request the Flow based on the current display window, if the display window 238315 1640418 of 78 current coincides with the ROI. The Client may decide not to request the additional Earcon Stream. • If an Earcon audio Element (or a group of Earcons) is transmitted in an auxiliary Stream independent of the main Stream, the Client may request, in the examples as before, the additional Stream provided there is a new ROI present in the visual scene. Furthermore, the two (or more) Streams can be processed using two media decoders and a common Render / Mix step to mix the decoded Earcon audio data to obtain the final audio scene. On the other hand, a metadata processor can be used to modify the metadata of the two Streams, and a Stream Merge to merge the two Streams. A possible implementation of the aforementioned Metadata Processor and Stream Merge is described below. On the other hand, a common (generic) Earcon can be used to signal all ROIs in an audio scene. This can be achieved using the same audio content with different spatial information associated with it at different times. In this case, the ROI processor 120 can request the metadata processor 132 to gather the Earcons associated with the ROIs in the scene and control the sequential playback of the Earcons (e.g., upon user selection or a request to apply a higher layer). On the other hand, an Earcon can be transmitted only once and stored in the client's cache. The client can reuse it for all ROIs of an audio scene with different spatial information associated with the audio content at different times. 238315 1640418 of 78 On the other hand, Earcon audio content can be generated synthetically on the client side. In addition, a metadata generator can be used to create the metadata necessary to signal the spatial information of the Earcon. For example, the Earcon audio content can be compressed and fed to a media decoder along with the main audio content and the new metadata, or it can be mixed with the final audio scene after using the media decoder, or possibly multiple media decoders. On the other hand, the Earcon audio content can be generated synthetically in the Client (e.g., under the control of metadata processor 132), as shown in the examples, while the metadata describing the Earcon is already included in the Stream. Using Earcon-type-specific signaling in the encoder, the metadata can contain the Earcon's spatial information, the specific signaling corresponding to a "Decoder-generated Earcon," but not the audio data corresponding to the Earcon. On the other hand, Earcon audio content can be generated synthetically on the client side, and a Metadata Generator can be used to create the metadata necessary to signal the Earcon's spatial information. For example, Earcon audio content can be compressed and fed to a media decoder along with the main audio content and the new metadata; • or it can be mixed into the final audio scene after the media decoder; • or multiple media decoders can be used. 238315 1640418 of 78 6.3 Examples of metadata corresponding to Audio information messages (e.g., Earcones) An example of Audio Information Message Metadata (Earcons) 141 is presented here, as described above. A structure to describe the properties of Earcones and offer the possibility of easily adjusting these values: Sintaxis No. de bits Mnemónica EarconInfo() { numEarcons 7 uimsbf for ( i=0; i< numEarcons; i++ ) { Earcon_isIndependent[i]; / * independent of la escena de audio * / 1 uimsbf Earcon_id[i]; / * map to group_id * / 7 uimsbf EarconType[i]; / * natural vs sythetic sound; generic vs individual * / 4 uimsbf EarconActive[i]; / * default disabled * / 1 bslbf EarconPosition[i]; / * position change * / 1 bslbf if ( EarconPosition[i] ) { Earcon_azi muth [i]; 8 uimsbf Earcon_elevation[i]; 6 uimsbf Earcon_radius[i]; 4 uimsbf} EarconHasGain; / * gain change * / 1 bslbf if ( EarconHasGain ) { Earcon_gain[i]; 7 uimsbf 238315 1640418 de 78 } EarconHasTextLabel; / *Etiqueta de texto * / 1 bslbf if (EarconHasTextLabel) { Earcon_numLanguages[i]; 4 uimsbf for ( n=0; n< Earcon_numLanguages[i]; n++ ) { Earcon_Language[i][n]; 24 uimsbf Earcon_TextDataLength[i][n]; 8 uimsbf for ( c=0; c< Earcon_TextDataLength[i][n]; c++ ) { Earcon_TextData[i][n][c]; 8 uimsbf}}}}} It is possible that each identifier in the table is intended to be associated with an attribute of the Earcon 132 metadata. The semantics are now described. numEarcons - This field specifies the number of Earcones audio elements available in the Stream. Earcon_isIndependent - This flag defines whether the Earcon audio element is independent of any audio scene. If Earcon_isIndependent == 1, the Earcon audio element is independent of the audio scene. If Earcon_isIndependent == 0, the Earcon audio element is part of the audio scene, and the Earcon_id must have the same value as the mae_groupID associated with the audio element. 238315 1640418 of 78 EarconType - This field defines the Earcon type. The following table specifies the allowed values. EarconType description 0 Not defined 1 Natural sound 2 Synthetic sound 3 Text-to-speech 4 Generic Earcon 5 / * reserved * / 6 / * reserved * / 7 / * reserved * / 8 / * reserved * / 9 / * reserved * / 10 / * reserved * / 11 / * reserved * / 12 / * reserved * / 13 / * reserved * / 14 / * reserved * / 15 Other EarconActive This flag defines whether the Earcon is active. If EarconActive == 1 The Earcon Audio element in the audio scene must be decoded and rendered. EarconPosition This flag defines whether the Earcon has available position information. If Earcon_isIndependent == 0, this position information is used instead of the audio object metadata specified in the 238315 1640418 of 78 structures dynamic_object_metadata() or intracoded_object_metadata_efficient() . Earcon_azimuth the absolute value of the azimuthal angle. Earcon_elevation the absolute value of the elevation angle. Earcon_radius the absolute value of the radius. EarconHasGain This flag defines whether the Earcon has a different Gain value. Earcon_gain This field defines the absolute value corresponding to the Earcon gain. EarconHasTextLabel This flag defines whether the Earcon has an associated text label. Earcon_numLanguages This field specifies the number of languages available for the descriptive text label. Earcon_Language This 24-bit field identifies the language of an Earcon's descriptive text. It contains a 3-character code as defined by ISO 639-2. Either ISO 639-2 / B or ISO 639-2 / T can be used. Each character is encoded in 8 bits according to ISO / IEC 8859-1 and inserted in order into the 24-bit field. EXAMPLE: French has a 3-character code “fre”, which is encoded as follows: “0110 0110 0111 0010 0110 0101”. Earcon_TextDataLength This field defines the length of the description of the next group in the bitstream. Earcon_TextData This field contains a description of an Earcon, that is, a string that describes the content using a high-level description. The format must follow UTF-8 according to ISO / IEC 10646. A structure for identifying Earcones at the system level and associating them with existing display windows. The following two tables offer two ways to implement this structure that can be used in different implementations: 238315 1640418 of 78 aligned(8) class EarconSample() extends SphereRegionSample { for (i = 0; i < num regions; i++){ unsigned int(7) reserved; unsigned int(1) hasEarcon; if (hasEarcon == 1){ unsigned int(8) numRegionEarcons; for (n=0; n <numRegionEarcons; n++){ unsigned int(8) Earcon id; unsigned int(32) Earcon track id;}} }} or on the other hand: aligned(8) class EarconSample() extends SphereRegionSample { for (i = 0; i < num regions; i++) { unsigned int(32) Earcon_track_id; unsigned int(8) Earcon_id; }} Semantics: hasEarcon specifies whether Earcon data is available for a region. numRegionEarcons specifies the number of Earcones available for a region. `Earcon_id` uniquely defines an ID for an Earcon element associated with the spherical region. If the Earcon is part of the Audio scene (i.e., if the Earcon is part of a group of elements identified by a `mae_groupID`), the `Earcon_id` MUST have the same value as the `mae_groupID`. The `Earcon_id` can be used to identify the audio file / track; for example, in the case of DASH delivery, the `AdaptationSet` with the `EarconComponent@tag` element in the MPD is equal to `Earcon_id`. Earcon_track_id - is an integer that uniquely identifies an Earcon track with the spherical region for the entire lifetime of a presentation; that is, if the Earcon track(s) are delivered in the same ISO BMFF file, the Earcon_track_id represents the corresponding 238315 1640418 of 78 track_id of the Earcon track(s). If the Earcon is not delivered within the same BMFF ISO file, this value MUST be set to zero. For easy identification of the Earcon track(s) at the MPD level, the following Attribute / Element EarconComponent@tag can be used: Summary of relevant MPD elements and attributes for MPEG-H Audio Element or Attribute Name Description ContentComponent@tag This field indicates the mae_groupID as stipulated in ISO / IEC 23008-3 [3DA] that is contained in the Media Content Component. EarconComponent@tag This field indicates the Earcon_id as stipulated in ISO / IEC 23008-3 [3DA] that is contained in the Media Content Component. In the case of MPEG-H Audio, this can be implemented, in the examples, by using the MHAS packages: • A new MHAS package can be defined to carry information about Earcons: PACTYP_EARCON which carries the EarconInfo() structure; • a new identification field in a generic MHAS METADATA MHAS package, to carry the EarconInfo() structure. With regard to Metadata, the 132 metadata processor may have at least some of the following capabilities: Extract metadata from audio information messages in a stream; Modify audio information message metadata to activate the audio information message and / or set / change its position and / or write / modify audio information message text labels; re-include the metadata in a Flow; feed the Stream to an additional media decoder; 238315 1640418 of 78 extract Audio Metadata from said at least a first Audio Stream (116); Extract metadata from audio information messages from an additional stream; Modify Audio Information Message Metadata to activate the Audio Information message and / or set / change its position and / or write / modify an Audio Information Message text label; modify Audio metadata of said at least a first Audio Stream (116) in order to take into account the existence of the Audio information message and allow merging; Feed a flow to the multiplexer or muxer to multiplex it based on the information received from the ROI processor. 6.4 Example from Fig. 3 Fig. 3 illustrates a system 300 comprising, on the customer side 204, a system 302 (customer system) which may incorporate, for example, system 100 or 200. System 302 may comprise the ROI processor 120, the metadata processor 132, a decoder group 313 consisting of a plurality of decoders 112. In this example, different Audio Streams are encoded (each in a respective Media Audio Decoder 112) and then mixed and / or rendered together to produce the final Audio scene. At least one Audio Stream is represented here as consisting of two Streams, 116 and 316 (other examples may have a single Stream, as in Fig. 2, or more than two Streams). These are the Audio Streams intended to reproduce the audio scene that the user is expected to experience. 238315 1640418 of 78 In this case, reference is made to Earcones, although it is possible to extend the concept to any Audio Information Message. In addition, the media encoder 240 can produce an Earcon 140 stream. Based on user movements and ROIs indicated in display window metadata 131 and / or other criteria, the ROI processor generates playback of an Earcon from Earcon Stream 140 (also referred to as Additional Audio Stream, as it is added to Audio Streams 116 and 316). Notably, the actual representation of Earcon is based on the metadata of Earcon 141 and the modifications made by the metadata processor 132. In the examples, the Flow can be requested by system 302 (client) from media encoder 240 (server) as needed. For example, the ROI processor might decide that, based on user behavior, a specific Earcon will soon be required and therefore request an appropriate Earcon Flow 140 from media encoder 240. The following aspects of this example should be noted: • Use case: Audio data is delivered in one or more Audio Streams 116, 316 (e.g., a Main Stream and an Auxiliary Stream) while Earcons are delivered in one or more Additional Streams 140 (dependent or independent of the Main Audio Stream) • In a client-side implementation 204, the ROI processor 120 and the metadata processor 132 are used to efficiently process Earcon information • The ROI processor 120 can receive information 122 about the current display window (information from 238315 1640418 of 78 user orientation) on the side of the media consumption device 206 used for content consumption (e.g., based on an HMD). The ROI processor can also receive information about the ROI and signal it in the Metadata (Video Display Windows are signaled according to OMAF). • Based on this information, the ROI 120 processor can decide to activate one (or more) Earcones contained in the Earcon 140 Audio Stream. In addition, the ROI 120 processor can decide to adopt a different location for the Earcones and different gain values (e.g., for a more accurate representation of the Earcon in the current space where the content is consumed). • The ROI processor 120 sends this information to the metadata processor 132. • The metadata processor 132 can analyze the Metadata contained in the Earcon Audio Stream and • enable the Earcon (in order to allow playback) • and, if requested by the ROI 120 processor, modify the spatial position and gain information contained in the Earcon 141 metadata accordingly. • Next, each Audio Stream 116, 316, 140 (based on user position information) is decoded and rendered independently, and mixer or renderer 314 mixes the 238315 1640418 of 78 output of all media decoders as a final step. A different implementation may decode only the compressed audio and send the decoded audio data and metadata to a common general renderer for final rendering of all audio elements (including earcones). • In addition, in a Streaming or Real-Time Transmission environment, based on the same information, the ROI 120 processor can decide to request the Earcon or Earcons 140 Flows in advance (e.g., when the user is looking in the wrong direction) a few seconds before the ROI is enabled. 6.5 Example from Fig. 4 Figure 4 illustrates a 400 system comprising, on the client side 204, a 402 system (client system) which may incorporate, for example, system 100 or 200. In this case, reference is made to the Earcones, although it is possible to extend the concept to any Audio Information Message. The 402 system may comprise the ROI processor 120, the metadata processor 132, and a Stream multiplexer (or muxer) 412. In examples where the multiplexer or muxer 412 is included, the number of operations that the hardware has to perform is advantageously reduced compared to the number of operations that must be performed when using multiple decoders and a mixer or renderer. In this example, different Audio Streams are processed based on their metadata and multiplexed into element 412. At least one Audio Stream is represented here, comprising two Streams, 116 and 316 (other examples may present a single Stream, as in Fig. 2, or more than two Streams). These are the Audio Streams intended for 238315 1640418 of 78 the playback of the audio scene that the user is expected to experience. In addition, an Earcon 140 stream can be provided by the media encoder 240. Based on user movements and the ROIs indicated in the display window metadata 131 and / or other criteria, the ROI processor 120 generates the playback of an Earcon from the Earcon Stream 140 (which is also referred to as an additional Audio Stream since it is added to Audio Streams 116 and 316). Each Audio Stream 116, 316, 140 can include metadata 236, 416, 141, respectively. At least some of this metadata can be manipulated and / or processed before being sent to the Stream 412 multiplexer, where the Audio Stream packets are merged. Consequently, the Earcon can be represented as part of the audio scene. Therefore, the Stream 412 muxer or multiplexer can produce an Audio Stream 414 comprising Modified Audio Metadata 238 and Modified Earcon Metadata 234, which can be sent to an Audio Decoder 112 and decoded and played back for the user. The following aspects of this example should be noted: • Use case: Audio data is delivered in one or more Audio Streams 116, 316 (e.g., a main Stream 116 and an auxiliary Stream 316, although a single Audio Stream can also be produced) while the Earcon (or Earcons) are distributed in one or more additional Streams 140 (dependent or independent of the main Audio Stream 116) 238315 1640418 of 78 • In a client-side implementation 204, the ROI processor 120 and the metadata processor 132 are used to efficiently process Earcon information. • The ROI processor 120 can receive information 122 about the current display window (user orientation information) from the media consumption device used for content consumption (e.g., an HMD). The ROI processor 120 can also receive information about the ROI signaled in the Earcon metadata 141 (Video Display Windows can be signaled in an Omnidirectional Media Application Format, OMAF). • Based on this information, the ROI 120 processor can decide to activate one (or more) Earcones contained in the additional Audio Stream 140. In addition, the ROI 120 processor can decide on a different location of the Earcones and different gain values (e.g., for a more accurate representation of the Earcon in the current space where the content is consumed). • The ROI 120 processor can send this information to the metadata processor 132. • The 132 metadata processor can analyze the metadata contained in the Earcon audio stream and enable Earcon 238315 1640418 of 78 • and, if requested by the ROI processor, modifies the spatial position and / or gain information and / or text labels contained in the Earcon metadata accordingly. • The metadata processor 132 can also analyze the audio metadata 236, 416 of all audio streams 116, 316 and manipulate the specific information of Audio in such a way that the Earcon can be used as part of the audio scene (e.g., if the audio scene has a 5.1 channel bed and 4 objects, the Earcon Audio Element is added to the scene as a fifth object. All metadata fields are updated accordingly). • The audio data of each stream 116, 316 and the modified audio metadata and Earcon metadata are sent to a Stream Muxer or Multiplexer which can generate, based on this, an Audio Stream 414 with a series of Metadata (modified Audio metadata 238 and modified Earcon metadata 234). • This Stream 414 can be decoded by a single Media Audio Decoder 112 based on user position information 122. • Furthermore, in a Streaming environment, based on the same information, the ROI 120 processor can decide to request the Earcon(s) 140 Flow(s) in advance (e.g., when the 238315 1640418 of 78 users look in the wrong direction a few seconds before the ROI is enabled). 6.6 Example from Fig. 5 Figure 5 illustrates a 500 system comprising, on the client side 204, a 502 system (client system) which may incorporate, for example, system 100 or 200. In this case, reference is made to the Earcones, although it is possible to extend the concept to any Audio Information Message. The 502 system may comprise the ROI processor 120, the metadata processor 132, a flow multiplexer or muxer 412. In this example, an Earcon stream is not produced by a remote (client-side) entity, but rather generated by the Synthetic Audio Generator 236 (which may also have the ability to store a Stream for later reuse, or to use a compressed / uncompressed version of a natural sound). However, the metadata for Earcon 141 is provided by the remote entity, e.g., in an Audio Stream 316 (which is not an Earcon stream). Therefore, the Synthetic Audio Generator 236 can be activated to create an Audio Stream 140 based on the attributes of the Earcon 141 metadata. For example, the attributes might refer to a type of synthesized voice (natural sound, synthetic sound, spoken text, etc.) and / or text tags (the Earcon stream can be generated by creating synthetic sound based on the text in the metadata).In the examples, once the Earcon Stream is created, it can be stored for future reuse. On the other hand, the synthetic sound can be a generic sound permanently stored on the device. A Stream 412 muxer or multiplexer can be used to merge packets from Audio Stream 116 (and also in the case of other Streams such as Auxiliary Audio Stream 316) with the Earcon Stream packets generated by the 238315 1640418 of 78 generator 236. Subsequently, an Audio Stream 414 can be obtained which is associated with the modified Audio metadata 238 and modified Earcon metadata 234. The Audio Stream 414 can be decoded by the decoder 112 and played back to the user from the media consumption device side 206. The following aspects of this example should be noted: • Use case: • Audio data is distributed across one or more Audio Streams (e.g., a Main Stream and an Auxiliary Stream). • Earcons are not distributed from the remote device; instead, Earcon 141 metadata is delivered as part of the Main Audio Stream (specific signaling can be used to indicate that the Earcon has no associated audio data). • In a client-side implementation, the ROI processor 120 and the metadata processor 132 are used to efficiently process Earcon information. • The ROI processor 120 can receive information about the current display window (user orientation information) from the device used on the content consumption device side 206 (e.g., an HMD). The ROI processor 120 can also receive information about the ROI signaled in the metadata (Video Display Windows are signaled according to OMAF). 238315 1640418 of 78 • Based on this information, the ROI 120 processor may decide to activate one (or more) Earcones NOT present in Stream 116. In addition, the ROI 120 processor may decide on a different location of the Earcones and different gain values (e.g., for a more accurate representation of the Earcon in the current space where the content is consumed). • The ROI 120 processor can send this information to the metadata processor 132. • The metadata processor 120 can analyze the metadata contained in the audio stream 116 and can • enable an Earcon • and, if requested by the ROI processor 120, modify the spatial position and gain information contained in the Earcon metadata 141 accordingly. • The metadata processor 132 can also analyze the Audio Metadata (e.g., 236, 417) of all audio streams (116, 316) and manipulate the Audio Specific information so that the Earcon can be used as part of the audio scene (e.g., if the audio scene has a 5.1 channel bed and 4 objects, the Earcon Audio Element is added to the scene as a fifth object. All metadata fields are updated accordingly). 238315 1640418 of 78 The modified Earcon metadata and the information obtained from the ROI 120 processor are sent to the Synthetic Audio Generator 246. Based on the received information, the Synthetic Audio Generator 246 can create a synthetic sound (e.g., based on the Earcon's spatial position, a voice signal spelling out the location is generated). Additionally, the Earcon 141 metadata is associated with the generated audio data, forming a new Stream 414. • Similarly, as before, the audio data of each stream (116, 316) and the modified audio metadata and Earcon metadata are then sent to a Stream Multiplexer which can generate, based on this, an individual Audio Stream with a series of Metadata (Audio and Earcon). • This Stream 414 is decoded by a single Media Audio Decoder 112 based on the user's position information. • On the other hand, or in addition, the Earcon's audio data can be retrieved in the Client (e.g., from previous Earcon uses). • On the other hand, the output of the Synthesizer Audio Generator 246 can be uncompressed audio and can be mixed into the final rendered scene. • Furthermore, in a Streaming environment, based on the same information, the ROI processor 120 can decide to request the 238315 1640418 of 78 the Earcon Flows in advance (e.g., when the user looks in the wrong direction a few seconds before the ROI is enabled). 6.7 Example from Fig. 6 Figure 6 illustrates a 600 system comprising, on the client side 204, a 602 system (client system) which may incorporate, for example, system 100 or 200. In this case, reference is made to the Earcones, although it is possible to extend the concept to any Audio Information Message. The 602 system may comprise the ROI processor 120, the metadata processor 132, a flow muxer or multiplexer 412. In this example, an Earcon flow is not provided by a remote (client-side) entity, but is generated by the synthetic Audio generator 236 (which may also have the ability to store a flow for later reuse). In this example, the Earcon 141 metadata is not provided by the remote entity. The Earcon metadata is generated by a metadata generator 432, which can generate Earcon metadata for use (e.g., processing, manipulation, modification) by the metadata processor 132. The Earcon 141 metadata generated by the Earcons metadata generator 432 may have the same structure, format, and / or attributes as the Earcon metadata described in the previous examples. The metadata processor 132 can operate as in the example in Fig. 5. A synthetic audio generator 246 can be activated to create an audio stream 140 based on the attributes in the Earcon metadata 141. For example, the attributes can refer to a type of synthesized voice (natural sound, synthetic sound, spoken text, etc.), and / or the gain, and / or the activation / disactivation state, and so on. In the examples, once 238315 The Earcon Stream 140, created in 1640418 of 78, can be stored (e.g., cached) for future reuse. It is also possible to store (e.g., cached) the Earcon metadata generated by the Earcons metadata generator 432. A Stream 412 muxer or multiplexer can be used to merge packets from Audio Stream 116 (and, also in the case of other Streams, such as Auxiliary Audio Stream 316) with the Earcon Stream packets generated by generator 246. Subsequently, an Audio Stream 414 can be obtained that is associated with modified Audio metadata 238 and modified Earcon metadata 234. Audio Stream 414 can be decoded by decoder 112 and played back to the user from the media consumption device side 206. The following aspects of this example should be noted: • Use case: • Audio data is distributed across one or more Audio Streams (e.g., a main Stream 116 and an auxiliary Stream 316) • Earcon(s) are not distributed from the client side 202, • Earcon metadata is not distributed from the client side 202 • This use case can represent a solution for enabling Earcons for content created using the previous technique that was created without Earcons • In a client-side implementation, the ROI processor 120 and the metadata processor 232 are used to efficiently process Earcon information 238315 1640418 of 78 • The ROI processor 120 can receive information 122 about the current viewing window (user orientation information) from the device used on the content consumption device side 206 (e.g., an HMD). The ROI processor 210 can also receive information about the ROI signaled in the Metadata (Video Viewing Windows are signaled according to OMAF). • Based on this information, the ROI 120 processor can decide to activate one (or more) Earcones that are NOT present in Flow (116, 316). • In addition, the ROI 120 processor can send information about the location of the Earcones and gain values to the Earcone 432 metadata generator. • The ROI 120 processor can send this information to the metadata processor 232. • The 232 metadata processor can analyze the metadata contained in an Earcon audio stream (if present) and can: • Enable Earcon • and, if requested by the ROI 120 processor, modify the spatial position and gain information contained in the Earcon metadata accordingly. 238315 1640418 of 78 • The metadata processor can also analyze the Audio Metadata 236, 417 of all audio streams 116, 316 and manipulate the Audio Specific information in such a way that the Earcon can be used as part of the audio scene (e.g., if the audio scene has a 5.1 channel bed and 4 objects, the Earcon Audio Element is added to the scene as a fifth object. All metadata fields are updated accordingly). • The modified Earcon metadata 234 and the information obtained from the ROI processor 120 are sent to the Synthetic Audio Generator 246. The Synthetic Audio Generator 246 can create, based on the received information, a synthetic sound (e.g., based on the spatial position of the Earcon, a speech signal is generated that spells out the location). Furthermore, the Earcon metadata is associated with the generated Audio data, forming a new Stream. • Similarly, as before, the audio data from each stream, along with the modified audio and Earcon metadata, are then sent to a Stream multiplexer 412, which can generate, based on this, an Audio Stream 414 with a series of Metadata (Audio and Earcon). • This Stream 414 is decoded by a single Media Audio Decoder based on the user's position information 238315 1640418 of 78 • On the other hand, the Earcon audio data may be stored in the Client's cache memory (e.g., from previous Earcon uses) • On the other hand, the output of the Synthesizing Audio Generator may be uncompressed audio and may be mixed with the final rendered scene • In addition, in a Streaming environment, based on the same information, the ROI 120 processor may decide to request the Earcon Stream(s) in advance (e.g., when the user looks in the wrong direction a few seconds before the ROI is enabled) 6.8 Example based on user position It is possible to implement a function that allows an Earcon to be played only when a user does not see the ROI. The ROI processor 120 can periodically check, for example, the user's current viewing window and / or head position and / or orientation and / or movement data 122. If the ROI is visible to the user, Earcon playback is not performed. If, taking into account the user's current viewing window and / or head position and / or orientation and / or movement data, the ROI processor determines that the ROI is not visible to the user, ROI processor 120 can request a playback of the Earcon. In this case, ROI processor 120 can instruct metadata processor 132 to prepare the Earcon playback. Metadata processor 132 can use one of the techniques described for the previous examples. For example, the metadata can be acquired in a stream delivered from the server. 238315 1640418 of 78 202, can be generated by the Earcons 432 metadata generator, and others. Earcon metadata attributes can be easily modified based on the processor's ROI requests and / or various conditions. For example, if a user selection has already disabled the Earcon, the Earcon will not play, even if the user does not see the ROI. For example, if a (previously configured) timer has not yet expired, the Earcon will not play, even if the user does not see the ROI. Furthermore, if from the user's current viewing window data and / or head position and / or orientation and / or movement, the ROI processor determines that the ROI is visible to the user, the ROI processor 120 may request that the Earcon not be played, especially if the Earcon metadata already contains signaling for an active Earcon. In this case, the ROI processor 120 can instruct the metadata processor 132 to disable Earcon playback. The metadata processor 132 can use one of the techniques described for the previous examples. For example, the metadata can be acquired in a stream transmitted by the server side 202, generated by the Earcon metadata generator 432, and so on. Earcon metadata attributes can be easily modified based on requests from the ROI processor and / or various conditions. If the metadata already contains an indication that an Earcon should be played, the metadata is modified, in this case, to indicate that the Earcon is inactive and should not be played. The following aspects of this example should be noted: • Use case: • Audio data is distributed across one or more Audio Streams 116, 316 (e.g., a Main Stream and an Auxiliary Stream) while the Earcone(s) are distributed across the same 238315 1640418 of 78 one or more Audio Streams 116, 316 or in one or more Additional Streams 140 (dependent or independent of the main Audio Stream) • Earcon metadata is configured in such a way as to indicate that Earcon must always be active at specific times. • A first generation of devices that do not include the ROI processor would read the Earcon metadata and generate Earcon playback regardless of whether the user's current viewing window data and / or head position and / or orientation and / or movement data indicate that the ROI is visible to the user. • A newer generation of devices that includes an ROI processor, as described in either system, would utilize the ROI processor's determination. If, based on the user's current viewing window data and / or head position and / or orientation and / or movement data, the ROI processor determines that the ROI is visible to the user, the ROI processor 120 can request that Earcon playback not be performed, especially if the Earcon metadata already contains signaling for an active Earcon.In this case, the ROI 120 processor can cause the metadata processor 132 to disable Earcon playback. The metadata processor 132 can use one of the... 238315 1640418 of 78 techniques described for the previous examples. For example, metadata can be acquired in a stream transmitted by the server side (202), generated by the Earcon metadata generator (432), and so on. Earcon metadata attributes can be easily modified based on ROI processor requests and / or various conditions. If the metadata already contains an indication that an Earcon should be played, the metadata is modified, in this case, to indicate that the Earcon is inactive and should not be played. • Additionally, depending on the playback device, the ROI processor may request modifications to the Earcon metadata. For example, the Earcon's special information may be modified differently depending on whether the sound is played through headphones or speakers. Therefore, the final audio scene experienced by the user is obtained based on the modifications to the metadata executed by the metadata processor. 6.9 Example based on Server-Client communication (Figure 5a) Figure 5a illustrates a 550 system comprising, on the client side 204, a 552 system (client system) which may incorporate, for example, system 100 or 200 or 300 or 400 or 500. In this case, reference is made to the Earcones, although it is possible to extend the concept to any Audio Information Message. The 552 system may comprise the ROI processor 120, the metadata processor 132, and a stream muxer or multiplexer 412. (In the examples, it is 238315 1640418 of 78 decode different Audio Streams (each using a respective Media Audio Decoder 112) and then mix them together and / or render them together to produce the final Audio scene). At least one Audio Stream is represented here, comprising two Streams, 116 and 316 (other examples may produce a single Stream, as in Fig. 2, or more than two Streams). These are the Audio Streams intended to reproduce the audio scene the user is expected to experience. Additionally, an Earcon 140 stream can be provided by the 240 media encoder. Audio streams can be encoded with different bit rates, allowing efficient bit rate adaptation depending on the network connection (i.e., for users using a high-speed connection, the version encoded with a high bit rate is transmitted, while for users with slower network connections, a version at a lower bit rate is transmitted). Audio streams can be stored on a Media Server 554, where for each audio stream, the different encodings with different bit rates are grouped into an Adaptation Series 556 with the appropriate data signaling the availability of all generated Adaptation Series. Audio Adaptation Series 556 and video Adaptation Series 557 can be produced. Based on user movements and the ROIs indicated in the display window metadata 131 and / or other criteria, the ROI processor 120 generates the playback of an Earcon from the Earcon Stream 140 (which is also referred to as Additional Audio Stream since it is added to Audio Streams 116 and 316). In this example: 238315 1640418 of 78 • Client 552 is configured to receive, from the server, data on the availability of all Adaptation Series, where the available Adaptation Series include: or at least one Audio Adaptation series for said at least one Audio Stream or at least one Audio Message Adaptation series corresponding to said at least one additional Audio Stream containing at least one additional Audio Information message. • Similar to the other implementation examples, the ROI processor 120 can receive information 122 about the current display window (user orientation information) from the media consumption device 206 used for content consumption (e.g., based on an HMD). The ROI processor 120 can also receive information about the ROI signaled in the Metadata (Video Display Windows are signaled in OMAF format). Based on this information, the ROI 120 processor can decide to activate one (or more) Earcons contained in the Earcon 140 Audio Stream. Or, in addition, the ROI 120 processor can decide on a different location of the Earcons and different gain values (e.g., for a more accurate representation of the Earcon in the actual space where the content is consumed). 238315 1640418 of 78 or The ROI 120 processor can send this information to a Selection 558 data generator. • A Selection Data Generator 558 may be configured to create, based on the ROI processor's decision, Selection Data 559 that identifies which Adaptation Series are to be received; Adaptation Series include Audio Scene Adaptation Series and Audio Message Adaptation Series. • The Media Server 554 may be configured to send instruction data to the Client 552 to cause the Streaming (real-time playback) client to retrieve the data corresponding to Adaptation Series 556, 557 identified by the selection data that identifies which Adaptation Series are to be received.Where Adaptation Series include Audio Scene Adaptation Series and Audio Message Adaptation Series; a Offload and Switch module 560 is configured to receive the requested Audio Streams from Media Server 554 based on selection data that identifies which Adaptation Series should be received, where Adaptation Series include Audio Scene Adaptation Series and Audio Message Adaptation Series. Offload and Switch module 560 may be further configured to send Audio Metadata and Earcon Metadata 141 to the Metadata Processor 132. • The ROI 120 processor can send this information to the metadata processor 132. 238315 1640418 of 78 • The metadata processor 132 can analyze the Metadata contained in the Audio Stream of Earcon 140 and enable the Earcon (in order to allow its playback) and, if requested by the ROI processor 120, modify the spatial position and gain information contained in the metadata of Earcon 141 accordingly. • The metadata processor 132 can also analyze the audio metadata of all audio streams 116, 316 and manipulate specific audio information so that the Earcon can be used as part of the audio scene (e.g., if the audio scene has a 5.1 channel bed and 4 objects, the Earcon audio element is added to the scene as a fifth object. All metadata fields can be updated accordingly). • The audio data from each stream 116, 316 and the modified audio metadata and Earcon metadata can then be sent to a Stream muxer or multiplexer, which can generate, based on this, an Audio Stream 414 with a series of Metadata (modified Audio metadata 238 and modified Earcon metadata 234). • This Stream can be decoded by a single Media Audio Decoder 112 based on user position information 122. An Adaptation Series may consist of a series of Representations that contain interchangeable versions of the content 238315 1640418 of 78 respectively, e.g., different audio bitrates (e.g., different streams with different bitrates). While a single Representation might theoretically be sufficient to produce a playable Stream, multiple Representations can give the Client the ability to tailor the Media Stream to their current network conditions and bandwidth requirements and thus ensure smooth playback. 6.10 Method All the examples presented can be implemented by steps of the method. In this case, a 700 method (which can be executed by any of the examples presented) is described to complete the description. The method may include: In step 702, receive at least one Video Stream (106) and at least one first Audio Stream (116, 316), In step 704, decode at least one video signal from at least one Video Stream (106) for the representation of a scene in a VR, AR, MR or 360-degree Video environment (118a) to a user and In step 706, decode at least one Audio signal from at least one first Audio Stream (116, 316) for the representation of an Audio scene (118b) to a user; receive a current user display window and / or head position and / or orientation and / or movement data (122) and In step 708, receive display window metadata (131) associated with at least one video signal from said at least one Video Stream (106), wherein the display window metadata defines at least one ROI and In step 710, decide, based on the user's current viewing window data and / or head position and / or orientation and / or 238315 1640418 of 78 of movement (122) and the display window metadata and / or other criteria, whether an audio information message associated with at least one ROI should be played or not and In step 712, receive, process, and / or manipulate Audio Information Message Metadata (141) that describes the Audio Information Message in order to generate playback of the Audio Information Message according to the attributes of the Audio Information Message, such that the Audio Information Message is part of the Audio Scene. Notably, the sequence can also vary. For example, the receive steps 702, 706, and 708 may have a different order, according to the actual order in which the information is delivered. Line 714 refers to the fact that the method can be repeated. Step 712 can be omitted if the ROI processor decides not to play the Audio Information message. 6.11 Other implementations Figure 8 illustrates an 800 system that can implement one of the systems (or a component thereof) or execute the 700 method. The 800 system can comprise an 802 processor and an 806 non-transient memory unit for storing instructions that, when executed by the 802 processor, can cause the processor to perform at least the Stream processing operations described above and / or the metadata processing operations described above. The 800 system can comprise an 804 input / output unit for connection to external devices. The 800 system can implement at least some (or all) of the functions of the ROI processor 120, the metadata processor 232, the generator 246, the muxer or multiplexer 412, the decoder 112, the Earcons metadata generator 432, and others. 238315 1640418 of 78 Depending on certain implementation requirements, the examples can be implemented in hardware. Implementation can be carried out using a digital storage medium, such as a floppy disk, a Digital Versatile Disc (DVD), a Blu-ray Disc, a Compact Disc (CD), a Read-Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electronically Erasable Programmable Read-Only Memory (EEPROM), or a FLASH memory, which has electronically readable control signals stored within it. These signals cooperate (or have the capacity to cooperate) with a programmable computer system in such a way that the respective method is executed. Therefore, the digital storage medium can be read by a computer. In general, examples can be implemented as a computer program product with program instructions. These instructions execute one of the methods when the computer program runs on a computer. The program instructions can be stored, for example, on machine-readable media. Other examples include the computer program for executing one of the methods described here, stored on a machine-readable medium. In other words, an example of the method, therefore, consists of a computer program comprising program instructions for executing one of the methods described here when the computer program is run on a computer. An additional example of the methods, therefore, consists of a data-carrying medium (or digital storage medium, or computer-readable medium). 238315 1640418 of 78 comprising, recorded thereon, the computer program for executing one of the methods described herein. The data carrier medium, the digital storage medium, or the recorded medium are generally tangible and / or non-transient rather than signals, which are intangible and transient. An additional example comprises a processing medium, for example, a computer, a programmable logic device to execute one of the methods described herein. Another additional example involves a computer on which the computer program has been installed to run one of the methods described here. Another example comprises an apparatus or system for transferring (for example, electronically or optically) a computer program for executing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may comprise, for example, a file server for transferring a computer program to the receiver. In some examples, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described here. In some examples, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described here. Generally, the methods are preferably performed by any suitable hardware device. The preceding examples are illustrative of the principles described. It is understood that modifications and variations to the provisions and details described herein should be obvious. Therefore, it is intended to be limited only by the scope of the following patent claims and 238315 1640418 of 78 not because of the specific details presented as a description and explanation of the examples presented here. 238315 1640418 of 78 20225952036 CRISTIAN DANIEL BITTEL - 20225952036 Digitally signed by PORTALTRAMITES - INPI Date: 2022.01.14 16:11:32 -03:00 Reason: Digitally Signed by the INPI Location: Buenos Aires, Argentina 1640418
Claims
1. A content consumption device system for receiving at least one Audio Stream, configured to receive at least one first Audio Stream associated with an Audio Scene, the at least one first Audio Stream having encoded therein at least one Audio signal to be decoded by at least one media Audio decoder, the content consumption device system being characterized in that it comprises at least one processor configured to: decide, based at least on current user motion data and / or selection and / or configuration, whether an Audio Information message should be played, wherein the Audio Information message is independent of the at least one Audio signal;and causing, upon the decision that the Audio Information message should be played, the playback of the Audio Information message, wherein at least one processor is configured to control a muxer or multiplexer to merge packets of the Audio Information message with packets of at least one first Audio Stream into one Stream, wherein the content consumption device system is configured to send, to at least one media Audio decoder, the Stream obtained from the at least one muxer or multiplexer and a modified version of the Audio Information message metadata of the Audio Information message. Eighteen claims follow;