Methods and apparatus for efficient delivery and use of audio messages for a high-quality experience

The system addresses user experience issues in VR, AR, and 360-degree video by dynamically managing audio notifications based on user interaction, ensuring accurate ROI identification and customizable message delivery.

JP2026062890APending Publication Date: 2026-04-10FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
Filing Date
2025-12-27
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing systems for delivering audio messages in VR, AR, and 360-degree video environments struggle with user experience issues due to the inability to accurately identify and manage Regions of Interest (ROIs), leading to missed important events and user inconvenience, lack of flexibility, and difficulty in customizing audio notifications for diverse user preferences.

Method used

A system that determines whether to play audio information messages based on user viewport, head orientation, and motion data, using metadata processors to manage and manipulate audio metadata to enable or disable notifications, allowing for flexible and user-specific audio message delivery.

Benefits of technology

Enhances user experience by accurately identifying and managing ROIs, allowing for customizable and efficient delivery of audio messages, improving user satisfaction and reducing distractions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062890000001_ABST
    Figure 2026062890000001_ABST
Patent Text Reader

Abstract

This invention provides methods and systems for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments. [Solution] A system 100 for a VR, AR, MR, or 360-degree video environment includes a video decoder that receives a video stream associated with the audio and video scene to be played and receives at least one first audio stream associated with the audio and video scene to be played, an audio decoder, and a region of interest (ROI) processor, the ROI processor playing an audio information message associated with at least one ROI, independent of the video signal and audio signal, based on at least the user's current viewport, head orientation, motion data, viewport metadata and / or audio information message metadata.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

Background Art

[0002] 1. Introduction In many applications, the delivery of audible messages can improve the user experience during media consumption. One of the most relevant applications for such messages is provided by virtual reality (VR) content. In a VR environment, or equally an augmented reality (AR) or mixed reality (MR) or 360-degree video environment, the user can typically visualize the entire 360-degree content, for example using a head-mounted display (HMD), and listen to it with headphones (or equally with speakers including correct rendering according to the speaker positions). The user can usually move in the VR / AR space or at least change the viewing direction, which is the so-called "viewport" of the video. In a 360-degree video environment using a conventional playback system (wide display screen) instead of an HMD, a remote control device can be used to emulate the movement of the user within the scene, and the same principle applies. Note that 360-degree content can refer to any type of content composed of multiple viewing angles simultaneously that the user can select (e.g., by the orientation of the user's head or using a remote control device).

[0003] Compared to conventional content consumption, in the case of VR, the content creator cannot control what the user visualizes at various points in time with the current viewport. The user can freely select different viewports from the permitted viewports or available viewports at each instance of time.

[0004] A common problem with consuming VR content is the risk of users missing important events in a video scene due to an incorrect viewport selection. To address this issue, the concept of Region of Interest (ROI) has been introduced, and several concepts for informing users of ROIs have been explored. While ROIs are typically used to indicate to the user the region containing the recommended viewport, they can also be used for other purposes, such as indicating the presence of new characters / objects in the scene, or indicating accessibility features associated with objects in the scene—essentially features that can be associated with elements that make up a video scene. For example, visual messages (e.g., "Turn your head to the left") can be used and overlaid on the current viewport. Alternatively, audible sounds, whether natural or synthesized, can be used by playing them at the ROI location. These audio messages are known as "earcones."

[0005] In this application, the concept of earphones is used to characterize audio messages transmitted to notify ROIs, but the proposed notification and processing can also be used for general audio messages for purposes other than notifying ROIs. An example of such an audio message is provided by an audio message that conveys information / displays of various options the user has in an interactive AR / VR / MR environment (e.g., "To enter room X, jump over the left side of the box"). Furthermore, although a VR example is used, the mechanism described in this document is applicable to any media consumption environment.

[0006] 2. Terms and Definitions The following terms are used in this technical field.

[0007] • Audio elements: Audio signals that can be represented, for example, as audio objects, audio channels, scene-based audio (higher-order ambisonics - HOA), or any combination of all of these.

[0008] • Region of Interest (ROI): A single area of ​​video content (or displayed or simulated environment) that a user is interested in at a given point in time. This is typically, for example, an area on a sphere or a selection of polygons from a 2D map. An ROI identifies a specific area for a particular purpose and defines the boundaries of the object under consideration.

[0009] • User location information: Location information (e.g., x, y, z coordinates), orientation information (yaw, pitch, roll), direction of movement, speed of movement, etc.

[0010] • Viewport: The portion of the 360-degree video currently displayed and being viewed by the user.

[0011] • Viewpoint: The center point of a viewport.

[0012] • 360-degree video (also known as immersive video or spherical video): In the context of this document, this refers to video content that simultaneously includes multiple views (viewports) in one direction. Such content can be created, for example, using an omnidirectional camera or a set of cameras. During playback, the viewer can control the viewing direction.

[0013] An adaptation set contains a media stream or a set of media streams. In the simplest case, it is a single adaptation set containing all the audio and video of the content, but to reduce bandwidth, each stream can be split into different adaptation sets. A common example is having one video adaptation set and multiple audio adaptation sets (one for each supported language). Adaptation sets can also contain subtitles or any metadata.

[0014] • Representations allow adaptation sets to include the same content encoded in different ways. In most cases, representations are provided at multiple bitrates. This allows clients to request the highest quality content that can be played back without waiting for buffering. Representations can also be encoded with various codecs, thus supporting clients with a variety of supported codecs.

[0015] A Media Presentation Description (MPD) is an XML syntax that contains information about media segments, their relationships, and the information needed to select them.

[0016] In the context of this application, the concept of adaptation sets is used more commonly, and sometimes actually refers to representations. Furthermore, media streams (audio / video streams) are typically encapsulated in media segments, which are the actual media files initially played by the client (e.g., a DASH client). Various formats can be used for media segments, such as ISO-based media file formats (ISOBMFF) and MPEG-TS, which are similar to the MPEG-4 container format. Encapsulation into media segments and encapsulation in various representations / adaptation sets are independent of the method described here, and this method applies to all various options.

[0017] Furthermore, while the description of the method in this document can focus on communication between a DASH server and a client, this method is sufficiently general to work in other distribution environments such as MMT, MPEG-2 transport streams, DASH-ROUTE, and file formats for file playback.

[0018] 3. Current Solutions The current solution is as follows:

[0019] [1].ISO / IEC 23008-3:2015,Information technology--High efficiency coding and media delivery in heterogeneous environments--Part 3:3D Audi

[0020] [2].N16950,Study of ISO / IEC DIS 23000-20 Omnidirectional Media Forma

[0021] [3].M41184, Use of Earcons for ROI Identification in 360-degree Video.

[0022] The distribution mechanism for 360-degree content is provided by ISO / IEC 23000-20, Omnidirectional Media Format[2]. This standard specifies a media format for coding, storage, distribution, and rendering of omnidirectional images, videos, and associated audio. It provides information about the media codecs used for audio and video compression, as well as additional metadata information for the proper use of 360-degree A / V content. It also specifies constraints and requirements for distribution channels, such as streaming via DASH / MMT and file-based playback.

[0023] The concept of earcons was first introduced in M41184, “Use of Earcons for ROI Identification in 360-degree Video”[3], which provides a mechanism for notifying the user of earcon audio data.

[0024] However, some users have reported unexpected comments from these systems. In many cases, a large number of icons become bothersome. When designers reduce the number of icons, some users lose important information. In particular, since each user has their own knowledge and experience level, they prefer a system that suits them. For example, each user prefers to play icons at a preferred volume (e.g., independent of the volume used for other audio signals). It has been proven difficult for system designers to obtain a system that provides a satisfactory level for all possible users. Therefore, a solution that can enhance the satisfaction of almost all users has been sought.

[0025] Furthermore, it has been proven difficult for designers to reconfigure the system. For example, it was difficult to prepare a new release of an audio stream or update the icons.

[0026] Furthermore, in a limited system, specific restrictions are imposed on functions, such as not being able to accurately identify an icon for one audio stream. Additionally, the icon always needs to be active, and playing it when not needed may inconvenience the user.

[0027] Furthermore, the icon spatial information cannot be sent or changed, for example, by a DASH client. Since this information can be easily accessed at the system level, additional functions that improve the user experience can be enabled.

[0028] Furthermore, there is no flexibility to handle various types of icons (e.g., natural sounds, synthetic sounds, sounds generated by a DASH client, etc.).

[0029] All these problems lead to a decline in the quality of the user experience. Therefore, a more flexible architecture is desired.

Prior Art Documents

Non-Patent Literature

[0030]

Non-Patent Literature 1

Non-Patent Literature 2

Non-Patent Literature 3

Summary of the Invention

[0031] 4. The present invention According to an example, a system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment is provided, the system comprising: receiving at least one video stream associated with an audio and video scene, being configured to receive at least one first audio stream associated with the reproduced audio and video scene, the system comprising: at least one media video decoder configured to decode at least one video signal from the at least one video stream for presentation of the audio and video scene to a user; at least one media audio decoder configured to decode at least one audio signal from the at least one first audio stream for presentation of the audio and video scene to a user; A region of interest ROI processor is included, and the region of interest ROI processor is Based on at least the user's current viewport and / or head orientation and / or motion data and / or viewport metadata and / or audio information message metadata, it is determined whether to play an audio information message associated with at least one ROI, and the audio information message is independent of at least one video signal and at least one audio signal. When it is decided to play an informational message, the system is configured to play an audio informational message.

[0032] For example, a system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments is provided, and the system is Receive at least one video stream, Configured to receive at least one first audio stream, The system is A media video decoder configured to decode at least one video signal from at least one video stream in order to present a VR, AR, MR, or 360-degree video environment scene to the user, For the representation of an audio scene to the user, at least one media audio decoder configured to decode at least one audio signal from at least one first audio stream, A region of interest ROI processor is included, and the region of interest ROI processor is Based on the user's current viewport and / or head orientation and / or motion data and / or viewport metadata and / or audio info message metadata, it is determined whether to play an audio info message associated with at least one ROI, where the audio info message is an earphone. When it is decided to play an informational message, the system is configured to play an audio informational message.

[0033] The system is The metadata processor may further include a metadata processor configured to receive and / or process and / or manipulate audio information message metadata and, when it decides to play back the information message, to play back the audio information message according to the audio information message metadata.

[0034] ROI processors are Receives data on the user's current viewport and / or position and / or head orientation and / or movement and / or other user-related data. Receive viewport metadata associated with at least one video signal from at least one video stream, and the viewport metadata defines at least one ROI, The system may be configured to determine whether to play an audio informational message associated with at least one ROI based on data from the user's current viewport and / or position and / or head orientation and / or motion, and viewport metadata.

[0035] The system is The metadata processor may further include one configured to receive and / or process and / or manipulate audio metadata and / or viewport metadata describing audio information messages and / or at least one audio signal encoded in at least one audio stream, and to play back audio information messages according to the audio information message metadata and / or audio metadata and / or viewport metadata describing at least one audio signal encoded in at least one audio stream.

[0036] ROI processors are If at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or motion data, then in addition to playing at least one audio signal, play at least one audio informational message associated with the ROI. The system may be configured to disallow and / or deactivate playback of audio informational messages associated with at least one ROI if at least one ROI is located within the user's current viewport and / or position and / or head orientation and / or motion data.

[0037] The system is It may be further configured to receive at least one additional audio stream in which at least one audio information message is encoded. The system is The system further includes at least one maxer or multiplexer that, under the control of a metadata processor and / or an ROI processor and / or another processor, merges packets of at least one additional audio stream with packets of at least one first audio stream in one stream, and plays an audio information message in addition to the audio scene, based on a decision provided by the ROI processor to play at least one audio information message.

[0038] The system is It receives at least one audio metadata describing at least one audio signal encoded into at least one audio stream, Receive audio information message metadata associated with at least one audio information message from at least one audio stream, When it is decided to play an informational message, the system may be configured to play at least one audio signal, as well as modify the audio informational message metadata to enable playback of the audio informational message.

[0039] The system is It receives at least one audio metadata describing at least one audio signal encoded into at least one audio stream, Receive audio information message metadata associated with at least one audio information message from at least one audio stream, When it is decided to play an audio information message, in addition to playing at least one audio signal, the audio information message metadata is modified to enable the playback of an audio information message associated with at least one ROI. The audio metadata describing at least one audio signal may be modified to allow merging of at least one first audio stream with at least one additional audio stream.

[0040] The system is It receives at least one audio metadata describing at least one audio signal encoded into at least one audio stream, Receive audio information message metadata associated with at least one audio information message from at least one audio stream, When it is decided to play an audio information message, the system may be configured to provide the audio information message metadata to a synthetic audio generator to create a synthetic audio stream, associate the audio information message metadata with the synthetic audio stream, and provide the synthetic audio stream and the audio information message metadata to a multiplexer or maxer to enable merging of at least one audio stream with the synthetic audio stream.

[0041] The system is The system may be configured to retrieve audio information message metadata from at least one additional audio stream in which the audio information message is encoded.

[0042] The system is The system may include an audio information message metadata generator configured to generate audio information message metadata based on a decision to play an audio information message associated with at least one ROI.

[0043] The system is It may be configured to store audio information message metadata and / or audio information message streams for future use.

[0044] The system is The system may include a synthesized audio generator configured to synthesize an audio information message based on audio information message metadata associated with at least one ROI.

[0045] The metadata processor may be configured to control a maxer or multiplexer to merge packets of an audio information message stream with packets of at least one first audio stream in one stream, in order to obtain the addition of audio information messages to at least one audio stream based on audio metadata and / or audio information message metadata.

[0046] Audio information message metadata may be encoded in a configuration frame and / or a data frame, and the data frame is Identification tag, An integer that uniquely identifies the playback of audio information message metadata. Message type, status Displaying dependencies / independence from the scene. Location data, Gain data, Displaying the existence of associated text labels. Number of available languages, Language of audio information messages, Length of data text, The data text of the associated text label, and / or It includes at least one of the descriptions of an audio information message.

[0047] The metadata processor and / or ROI processor are Extract audio information message metadata from the stream, Modify the audio information message metadata to activate the audio information message and / or set / change its position. Embed metadata into the stream The stream is fed to an additional media decoder. Extract audio metadata from at least one first audio stream, Extract audio information message metadata from the additional stream, Modify the audio information message metadata to activate the audio information message and / or set / change its position. Modify the audio metadata of at least one first audio stream so that it can be merged while taking into account the presence of audio information messages. It may be configured to perform at least one of the following operations: supplying a stream to a multiplexer or maxer for multiplexing or multiplexing information received from the ROI processor.

[0048] The ROI processor may be configured to perform a local lookup for additional audio streams and / or audio information message metadata in which the audio information message is encoded, and, if it cannot be found, to request the additional audio streams and / or audio information message metadata from a remote entity.

[0049] The ROI processor may be configured to perform a local search for additional audio streams and / or audio information message metadata, and if it is not possible to find them, to cause a synthesized audio generator to generate audio information message streams and / or audio information message metadata.

[0050] The system is Receive at least one additional audio stream containing at least one audio informational message associated with at least one ROI, The ROI processor may be configured to decode at least one additional audio stream if it decides to play an audio information message associated with at least one ROI.

[0051] The system is A first audio decoder for decoding at least one audio signal from at least one first audio stream, At least one additional audio decoder for decoding at least one audio information message from an additional audio stream, The system may include at least one mixer and / or renderer for mixing and / or superimposing audio information messages from at least one additional audio stream with at least one audio signal from at least one first audio stream.

[0052] The system may be configured to maintain a track of metrics associated with historical and / or statistical data related to the playback of audio informational messages, and to disable the playback of audio informational messages if the metrics exceed a predetermined threshold.

[0053] The ROI processor's decision may also be based on predictions of the user's current viewport and / or position and / or head orientation and / or motion data in relation to the ROI's location.

[0054] The system may be configured to receive at least one first audio stream and, once it is determined to play an informational message, request an audio message informational stream from a remote entity.

[0055] The system may be configured to determine whether to play two audio information messages simultaneously, or to select a higher-priority audio information message to be played preferentially over a lower-priority audio information message.

[0056] The system may be configured to identify an audio information message from among multiple audio information messages encoded in one additional audio stream, based on the address and / or position of the audio information message in the audio stream.

[0057] The audio stream may also be formatted in the MPEG-H 3D audio stream format.

[0058] The system is The system receives data regarding the availability of multiple adaptation sets, and the available adaptation sets include an adaptation set of at least one audio scene from at least one first audio stream, and an adaptation set of at least one audio message from at least one additional audio stream containing at least one audio information message, and the system Based on the ROI processor's decision, selection data is created to identify which of the available adaptation sets to search, and the available adaptation sets include at least one audio scene adaptation set and / or at least one audio message adaptation set. Request and / or retrieve data for the adaptation set identified by the selected data. Each adaptation set may be configured to group different encodings with different bitrates.

[0059] The system may be configured such that at least one of its elements includes HTTP, DASH, dynamic adaptive streaming via a client, and / or uses the ISO-based media file format ISO BMFF, or the MPEG-2 transport stream MPEG-2 TS to retrieve data for each of the adaptation sets.

[0060] The ROI processor may be configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or motion data to determine if the ROI is represented in the current viewport, and to notify the user audibly of the ROI's presence if the ROI is outside the current viewport and / or position and / or head orientation and / or motion data.

[0061] The ROI processor may be configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or motion data to determine if the ROI is represented in the current viewport, and to suppress audible notification to the user of the ROI's presence if the ROI is present in the current viewport and / or position and / or head orientation and / or motion data.

[0062] The system may be configured to receive from a remote entity at least one video stream associated with a video environment scene and at least one audio stream associated with an audio scene, the audio scene being associated with the video environment scene.

[0063] The ROI processor may be configured to select the playback of one first audio information message that precedes a second audio information message from among several audio information messages to be played.

[0064] The system may include cache memory for storing audio informational messages received from or synthetically generated from remote entities, and for reusing audio informational messages in different time instances.

[0065] Audio information messages may also be transmitted via earphones.

[0066] At least one video stream and / or at least one first audio stream may each be part of the current video environment scene and / or video audio scene, and may be independent of the user's current viewport and / or head orientation and / or motion data in the current video environment scene and / or video audio scene.

[0067] The system may be configured to request at least one first audio stream and / or at least one video stream from remote entities associated with the audio stream and / or video environment stream, respectively, and to play at least one audio informational message based on the user's current viewport and / or head orientation and / or motion data.

[0068] The system may be configured to request at least one first audio stream and / or at least one video stream from remote entities associated with the audio stream and / or video environment stream, respectively, and to request at least one audio informational message from the remote entities based on the user's current viewport and / or head orientation and / or motion data.

[0069] The system may be configured to request at least one first audio stream and / or at least one video stream from remote entities associated with the audio stream and / or video environment stream, respectively, and to synthesize at least one audio informational message based on the user's current viewport and / or head orientation and / or motion data.

[0070] The system may be configured to check at least one of additional criteria for playback of audio information messages, the criteria may further include user selection and / or user settings.

[0071] The system may be configured to check at least one of additional criteria for playback of audio information messages, the criteria further including the state of the system.

[0072] The system may be configured to check at least one of additional criteria for playback of an audio information message, the criteria further including the number of times an audio information message has already been played.

[0073] The system may also be configured to check at least one of additional criteria for playback of audio information messages, the criteria further including flags in a data stream retrieved from a remote entity.

[0074] In one embodiment, a system is provided that includes a client configured as one of the above and / or below examples, and a remote entity configured as a server for delivering at least one video stream and at least one audio stream.

[0075] A remote entity may be configured to search for at least one additional audio stream and / or audio information message metadata in a database, intranet, internet, and / or geographic network, and, if found, deliver at least one additional audio stream and / or audio information message metadata.

[0076] The remote entity may be configured to synthesize at least one additional audio stream and / or generate audio information message metadata.

[0077] According to one embodiment, a method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment may be provided, and the method is A step of decoding at least one video signal from at least one video and audio scene to be played back to the user, A step of decoding at least one audio signal from the video and audio scenes to be played back, A step of determining whether to play an audio information message associated with at least one ROI based on the user's current viewport and / or head orientation and / or motion data and / or metadata, wherein the audio information message is independent of at least one video signal and at least one audio signal, The process includes the step of playing an audio informational message once it is decided to play an informational message.

[0078] According to one embodiment, a method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment may be provided, and the method is A step of decoding at least one video signal from at least one video stream in order to represent a VR, AR, MR, or 360-degree video environment scene to the user, To represent the audio scene to the user, the process involves decoding at least one audio signal from at least one first audio stream, A step of determining whether to play an audio informational message associated with at least one ROI, based on the user's current viewport and / or head orientation and / or motion data and / or metadata, wherein the audio informational message is an earphone, and When it is decided to play an informational message, the steps include playing an audio informational message, Includes.

[0079] The above and / or below methods are, Once it is decided to play an informational message, the process may include receiving and / or processing and / or manipulating metadata in order to play the audio informational message according to the metadata, such that the audio informational message is part of an audio scene.

[0080] The above and / or below methods are, Steps to play audio and video scenes, The step may include deciding to play further audio informational messages based on the user's current viewport and / or head orientation and / or motion data and / or metadata.

[0081] The above and / or below methods are, Steps to play audio and video scenes, If at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or motion data, then play at least one audio informational message associated with at least one ROI, in addition to playing at least one audio signal, and / or The steps may include: if at least one ROI is in the user's current viewport and / or position and / or head orientation and / or motion data, disallowing and / or deactivating the playback of audio informational messages associated with at least one ROI.

[0082] For example, a system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments is provided, and the system is Receive at least one video stream, Configured to receive at least one first audio stream, The system is A media video decoder configured to decode at least one video signal from at least one video stream in order to present a VR, AR, MR, or 360-degree video environment scene to the user, For the representation of an audio scene to the user, at least one media audio decoder configured to decode at least one audio signal from at least one first audio stream, A region of interest ROI processor is included, and the region of interest ROI processor is Based on the user's current viewport and / or head orientation and / or motion data and / or metadata, determine whether to play an audio informational message associated with at least one ROI. When it is decided to play an informational message, the system is configured to play an audio informational message.

[0083] In the example, a system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments is provided, and the system is Receive at least one video stream, Configured to receive at least one first audio stream, The system is A media video decoder configured to decode at least one video signal from at least one video stream in order to present a VR, AR, MR, or 360-degree video environment scene to the user, For the representation of an audio scene to the user, at least one media audio decoder configured to decode at least one audio signal from at least one first audio stream, A region of interest ROI processor configured to determine whether to play an audio informational message associated with at least one ROI based on the user's current viewport and / or position and / or head orientation and / or motion data and / or metadata and / or other criteria, A metadata processor is configured to receive and / or process and / or manipulate metadata and, when it decides to play back an informational message, to play back an audio informational message according to the metadata such that the audio informational message is part of an audio scene.

[0084] According to one embodiment, a non-transient storage unit is provided, which, when executed by a processor, includes instructions that cause the processor to perform the methods described above and / or below.

[0085] 5. Description of the drawings [Brief explanation of the drawing]

[0086] [Figure 1] This figure shows an example of an embodiment. [Figure 2] This figure shows an example of an embodiment. [Figure 3] This figure shows an example of an embodiment. [Figure 4] This figure shows an example of an embodiment. [Figure 5] This figure shows an example of an embodiment. [Figure 5a] This figure shows an example of an embodiment. [Figure 6] This figure shows an example of an embodiment. [Figure 7] This figure shows an example of the method. [Figure 8] This figure shows an example of an embodiment. [Modes for carrying out the invention]

[0087] 6. Example 6.1 General Examples Figure 1 shows an example of system 100 for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment. System 100 may be associated with, for example, a content consumption device (e.g., a head-mounted display) that plays visual data on a spherical or hemispherical display closely associated with the user's head.

[0088] System 100 may include at least one media video decoder 102 and at least one media audio decoder 112. System 100 may receive at least one video stream 106 in which a video signal is encoded to represent a VR, AR, MR, or 360-degree video environment scene 118a to the user. System 100 may receive at least one first audio stream 116 in which an audio signal is encoded for representing an audio scene 118b to the user.

[0089] System 100 may also include a Region of Interest (ROI) processor 120. The ROI processor 120 can process data associated with the ROI. Generally speaking, the presence of an ROI may be indicated by viewport metadata 131. The viewport metadata 131 may be encoded in the video stream 106 (in other examples, the viewport metadata 131 may be encoded in other streams). The viewport metadata 131 may include, for example, location information (e.g., coordinate information) associated with the ROI. For example, the ROI may be understood as a rectangle in this example (identified by coordinates such as the position of one of the four vertices of the rectangle in the spherical video and the length of the sides of the rectangle). The ROI is typically projected onto the spherical video. The ROI is typically associated with a visible element that is considered to be of interest to the user (according to a particular configuration). For example, the ROI may be associated with a rectangular area that is displayed by a content-consuming device (or is visible to the user in some way).

[0090] The ROI processor 120 can, in particular, control the operation of the media audio decoder 112.

[0091] The ROI processor 120 can obtain data 122 associated with the user's current viewport and / or position and / or head orientation and / or movement (virtual data associated with a virtual position can also be understood as part of the data 122 in some examples). This data 122 may be provided, for example, by a content consuming device or by a positioning / detection unit, at least in part.

[0092] The ROI processor 120 can check the correspondence between the ROI and the user's current viewport and / or position (actual or virtual) and / or head orientation and / or movement data 122 (for example, other criteria may be used). For example, the ROI processor can check whether the ROI is represented in the current viewport. If the ROI is only partially represented in the viewport (for example, based on the user's head movement), it can determine, for example, whether a minimum percentage of the ROI is visible on the screen. In any case, the ROI processor 120 can recognize whether the ROI is not represented or not visible to the user.

[0093] If the ROI is considered to be outside the user's current viewport and / or position and / or head orientation and / or motion data 122, the ROI processor 120 may audibly notify the user of the presence of the ROI. For example, the ROI processor 120 may request the playback of an audio information message (earcon) in addition to an audio signal decoded from at least one first audio stream 116.

[0094] If the ROI is considered to be within the user's current viewport and / or position and / or head orientation and / or motion data122, the ROI processor may decide to avoid playing the audio information message.

[0095] The audio information message may be encoded into an audio stream 140 (audio information message stream), which may be the same as or different from the audio stream 116. The audio stream 140 may be generated by the system 100 or obtained from an external entity (e.g., a server). Audio metadata, such as audio information message metadata 141, may be defined to describe the properties of the audio information stream 140.

[0096] Audio information messages may be superimposed on (or mixed, multiplexed, merged, combined, or configured with) the signal encoded in the audio stream 116, or they may not be selected, for example, simply based on a decision by the ROI processor 120. The ROI processor 120 may make that decision based on viewport and / or position and / or head orientation and / or motion data 122, metadata (such as viewport metadata 131 or other metadata), and / or other criteria (e.g., selection, system state, the number of audio information messages already played, specific functions and / or operations, user preference settings which may disable the use of earphones, etc.).

[0097] A metadata processor 132 may be implemented. The metadata processor 132 can be inserted, for example, between the ROI processor 120 (which may control the metadata processor 132) and the media audio decoder 112 (which may be controlled by the metadata processor). In the example, the metadata processor is part of the ROI processor 120. The metadata processor 132 can receive, generate, process, and / or manipulate audio information message metadata 141. The metadata processor 132 can also process and / or manipulate the metadata of the audio stream 116, for example, to multiplex the audio stream 116 with the audio information message stream 140. Furthermore or alternatively, the metadata processor 132 can receive metadata for the audio stream 116 from, for example, a server (e.g., a remote entity).

[0098] Therefore, the metadata processor 132 can modify the playback of the audio scene and adapt the audio information message to a specific situation and / or selection and / or state.

[0099] Here, we will describe some of the advantages of several embodiments.

[0100] Audio information messages can be accurately identified, for example, using audio information message metadata 141.

[0101] Audio information messages can be easily activated / deactivated, for example, by modifying the metadata (e.g., by the metadata processor 132). Audio information messages can also be enabled / disabled, for example, based on the current viewport and ROI information (and any special functions or effects achieved).

[0102] Audio informational messages (including status, type, spatial information, etc.) can be easily notified and modified by common devices such as dynamic adaptive streaming via an HTTP (DASH) client.

[0103] Therefore, since audio information messages (including status, type, spatial information, etc.) can be easily accessed at the system level, additional features can be enabled to improve the user experience. Thus, system 100 can be easily customized, allowing for further embodiments (e.g., specific applications) that can be implemented by personnel independent of the system 100's designers.

[0104] Furthermore, flexibility is achieved in handling various types of audio information messages (e.g., natural sounds, synthesized sounds, sounds generated by the DASH client, etc.).

[0105] Other advantages (as will become clear in the following examples): • Use of text labels within metadata (as a basis for displaying something or generating earphones) • Adjusting the earphone position based on the device (precise positioning is required for HMDs, and for speakers, it might be better to use a different position - directly to one speaker).

[0106] • Different device classes: • Earcon metadata can be created in a way that indicates that the earcon is active.

[0107] • Some devices only recognize how to play earphones by parsing metadata. • Some newer devices with better ROI processors can decide to deactivate them when not needed. • Further information and additional diagrams regarding the adaptation set.

[0108] Therefore, in a VR / AR environment, users can typically visualize the entire 360-degree content using, for example, a head-mounted display (HMD) and listen through headphones. Users can typically move around in the VRJAR space, or at least change their viewing direction, which is the so-called "viewport" of the video. Compared to conventional content consumption, in the case of VR, content creators lose control of what the user visualizes at various points in time in the current viewport. Users are free to select different viewports for each instance of time from the permitted or available viewports. To indicate a region of interest (ROI) to the user, audible sounds (natural or synthesized sounds) can be used by playing them at the location of the ROI. These audio messages are known as "earcones." This invention proposes a solution for the efficient delivery of such messages and proposes optimized receiver operation for utilizing earcones without affecting the user experience and content consumption. This improves the quality of the experience. This can be achieved by using dedicated metadata and metadata manipulation mechanisms at the system level to enable or disable earcones in the final scene.

[0109] The metadata processor 132 can be configured to receive and / or process and / or manipulate metadata 141 and, in a decision to play an informational message, to play an audio informational message according to the metadata 141. An audio signal (e.g., one that represents a scene) can be understood as part of an audio scene (e.g., an audio scene downloaded from a remote server). Audio signals generally have semantic meaning in relation to an audio scene, and all audio signals that exist together constitute an audio scene. Audio signals can be encoded together into a single audio bitstream. Audio signals may be created by a content creator and / or associated with a particular scene and / or independent of an ROI.

[0110] Audio informational messages (e.g., earphones) may be understood as having no semantic meaning within the audio scene. They can be understood as independent sounds that can be artificially generated, such as recorded sounds or the voice of a person on a recorder. They may also be device-dependent (e.g., system sounds generated when a button is pressed on a remote control). Audio informational messages (e.g., earphones) may be understood not as part of the scene, but as something that guides the user within the scene.

[0111] Audio information messages may be independent of audio signals as described above. According to a different example, they may be included in the same bitstream, transmitted in a separate bitstream, or generated by system 100.

[0112] An example of an audio scene composed of multiple audio signals is as follows:

[0113] - Audio Scene: Concert room containing 5 audio signals: ---Audio signal 1: Piano sound ---Audio signal 2: Singer's voice ---Audio signal 3: Voice of person 1, who is part of the audience. ---Audio signal 4: Voices of two people who are part of the audience ---Audio signal 5: Sound generated by a wall clock The audio information message may be a recorded voice message, for example, "Look at the pianist" (the piano is the ROI). If the user is already looking at the pianist, the audio message will not play.

[0114] Another example: A door behind the user (e.g., a virtual door) opens, and a new person enters the room. The user is not looking. The earphones can be triggered based on this (information about the VR environment, such as virtual position) to notify the user that something has happened behind them.

[0115] In the example, when the user changes the environment, each scene (for example, the associated audio and video streams) is sent from the server to the client.

[0116] Audio information messages may be flexible. In particular: - Audio information messages can be placed in the same audio stream associated with the scene being played.

[0117] - Audio information messages can be placed in additional audio streams.

[0118] - Audio information messages may be completely missing, but only metadata describing the earphones may be present in the stream, and audio information messages can be generated by the system.

[0119] - Audio information messages and the metadata describing them may be completely missing, in which case the system will generate both (earcon and metadata) based on other information about ROIs in the stream.

[0120] Audio information messages are generally independent of the audio signal portion of an audio scene and are not used to represent the audio scene.

[0121] Examples of systems that embody or include part of System 100 are shown below.

[0122] 6.2 Example in Figure 2 Figure 2 shows a system 200 (which may include at least some implementing systems 100) that is here subdivided into a server side 202, a media distribution side 203, a client side 204, and / or a media consuming device side 206. Each of sides 202, 203, 204, and 206 is a system itself and can be combined with other systems to form another system. Here, we refer to an audio information message as an earcon, although it can be generalized to any type of audio information message.

[0123] The client-side 204 can receive at least one video stream 106 and / or at least one audio stream 116 from the server-side 202 via the media distribution side 203.

[0124] The distributor 203 may be based on a communication system or file-based storage, such as a cloud system, network system, geographical communication network, or well-known media transport formats (e.g., MPEG-2 TS transport stream, DASH, MMT, DASH ROUTE). The distributor 203 may perform communication by delivering data packets in the form of electrical signals (e.g., via cable, wirelessly, etc.) and / or in bitstreams in which audio and video signals are encoded (e.g., according to a specific communication protocol). However, the distributor 203 may be embodied by point-to-point links, serial or parallel connections, etc. The distributor 203 may perform wireless connections according to protocols such as WiFi, Bluetooth, etc.

[0125] The client-side 204 can be associated with a media-consuming device, such as an HND (Human-Wing Device) into which the user can insert their head (however, other devices may be used). Thus, the user can experience a video and audio scene (e.g., a VR scene) prepared by the client-side 204 based on the video and audio data provided by the server-side 202. However, other embodiments are also possible.

[0126] The server-side 202 is represented here as having a media encoder 240 (which may cover video encoders, audio encoders, subtitle encoders, etc.). This encoder 240 may be associated with, for example, the represented audio and video scenes. The audio scene may be, for example, for reproducing an environment and may be associated with at least one audio and video data stream 106, 116, which may be encoded based on the position (or virtual position) reached by the user in the VR, AR, or MR environment. Typically, the video stream 106 encodes a spherical image, and only a portion of it (viewport) is displayed to the user according to its position and movement. The audio stream 116 contains audio data that participates in the audio scene representation and is intended to be heard by the user. For example, the audio stream 116 may include audio metadata 236 (which refers to at least one audio signal intended to participate in the audio scene representation) and / or earphone metadata 141 (which may, in some cases, describe only the earphones to be played).

[0127] System 100 is represented here as being located on the client side 204. For simplicity, the media video decoder 112 is not shown in Figure 2.

[0128] The earcon metadata 141 can be used to prepare the earcon (or other audio information message) for playback. The earcon metadata 141 can be understood as metadata (which may be encoded in an audio stream) that describes and provides attributes associated with the earcon. Thus, the earcon (if played) can be based on the attributes of the earcon metadata 141.

[0129] Advantageously, the metadata processor 132 may be specifically implemented for processing the earphone metadata 141. For example, the metadata processor 132 can control the reception, processing, manipulation, and / or generation of the earphone metadata 141. Once processed, the earphone metadata is represented as modified earphone metadata 234. For example, the earphone metadata can be manipulated to obtain specific effects, as well as to perform audio processing operations such as multiplexing or multiplexing to add earphones to the audio signals represented in the audio scene.

[0130] The metadata processor 132 can control the reception, processing, and manipulation of audio metadata 236 associated with at least one stream 116. Once processed, the audio metadata 236 can be represented as modified audio metadata 238.

[0131] The modified metadata 234 and 238 can be provided to the media audio decoder 112 (or multiple decoders in some examples) for playback of the audio scene 118b to the user.

[0132] In the example, a synthesis audio generator and / or a storage device 246 may be provided as optional components. The generator can synthesize an audio stream (for example, to generate earcons that are not encoded into a stream). The storage device allows the earcon stream obtained from the audio stream generated and / or received by the generator to be stored (for example, for future use) (for example, in cache memory).

[0133] Therefore, the ROI processor 120 can determine the representation of the earcon based on the user's current viewport and / or position and / or head orientation and / or motion data 122. However, the ROI processor 120 may also make that determination based on criteria including other aspects.

[0134] For example, the ROI processor can enable / disable earphone playback based on other conditions, such as user selection or higher-level selection, or based on the specific application it is intended to be consumed by. For instance, in a video game application, earphone and other audio information messages can be avoided if the video game level is high. This can be easily achieved by the metadata processor by disabling earphone metadata.

[0135] Furthermore, the earphones can be disabled based on the system state. For example, if the earphones are already playing, their repetition will be prohibited. A timer may also be used to avoid repetition that is too fast.

[0136] The ROI processor 120 can also request controlled playback of a set of earphones (e.g., earphones associated with all ROIs in the scene) to instruct the user about elements that the user can see. The metadata processor 132 can control this operation.

[0137] The ROI processor 120 can also change the earphone position (i.e., spatial position within the scene) or earphone type. For example, some users may prefer to play a specific sound at the exact location / position of the ROI as the earphone, while others may prefer to play the earphone at one fixed position (e.g., a "voice of God" in the center or at the top) to audibly indicate where the ROI is located.

[0138] The gain of the earphone playback can be changed (for example, to obtain a different volume). This decision may follow, for example, a user selection. In particular, based on the ROI processor's decision, the metadata processor 132 performs the gain change by modifying a specific attribute associated with the gain among the earphone metadata associated with the earphone.

[0139] Even the original designers of VR, AR, and MR environments may not be aware of how the ear cons will actually be reproduced. For example, the final rendering of the ear cons may change depending on the user's choices. Such behavior can be controlled, for example, by a metadata processor 132 that can modify the ear con metadata 141 based on a decision by an ROI processor.

[0140] Therefore, operations performed on the audio data associated with the earcon are, in principle, independent of and can be managed in a different way from at least one audio stream 116 used to represent the audio scene. The earcon can also be generated separately from the audio and video streams 106, 116 that constitute the audio and video scene, and can be generated by different independent groups of entrepreneurs.

[0141] Therefore, this example makes it possible to increase user satisfaction. For example, users can make their own choices, such as changing the volume of audio information messages or disabling audio information messages. Thus, each user can get an experience that is more suited to their preferences. Furthermore, the acquired architecture is more flexible. Audio information messages can be easily updated, for example, by changing the metadata independently of the audio stream, and / or by changing the audio information message stream independently of the metadata and the main audio stream.

[0142] The resulting architecture is compatible with legacy systems; for example, legacy audio information message streams can be associated with the new audio information message metadata. If a suitable audio information message stream does not exist, the latter can be easily synthesized (and, for example, stored for subsequent use).

[0143] The ROI processor maintains a track of metrics associated with historical and / or statistical data related to the playback of audio informational messages, and can disable the playback of audio informational messages if the metrics exceed a predetermined threshold (which can be used as a criterion).

[0144] The ROI processor's determination may, as a criterion, be based on predictions of the user's current viewport and / or position and / or head orientation and / or motion data in relation to the ROI location.

[0145] The ROI processor may be further configured to receive at least one first audio stream 116 and, once it is determined to play an informational message, request an audio message informational stream from a remote entity.

[0146] The ROI processor and / or metadata generator may be further configured to determine whether to play two audio information messages simultaneously or to select a higher-priority audio information message to be played preferentially over a lower-priority audio information message. Audio information metadata can be used to perform this decision. Priority can be obtained by the metadata processor 132 based, for example, on a value in the audio information message metadata.

[0147] In some examples, the media encoder 240 may be configured to allow remote entities to search for additional audio streams and / or audio information message metadata in databases, intranets, the Internet, and / or geographical networks, and to deliver additional audio streams and / or audio information message metadata if found. For example, the search may be performed based on a client-side request.

[0148] As explained above, a solution is proposed here for efficiently delivering earphone messages along with audio content. Optimized receiver operation is achieved to utilize audio information messages (e.g., earphones) without impacting user experience and content consumption. This improves the quality of the experience.

[0149] This can be achieved by using dedicated metadata and metadata manipulation mechanisms at the system level to enable or disable audio information messages in the final audio scene. The metadata can be used with any audio codec and appropriately complements next-generation audio codec metadata (e.g., MPEG-H audio metadata).

[0150] There can be various distribution mechanisms (e.g., streaming via DASH / HLS, broadcasting via DASH-ROUTE / MMT / MPEG-2 TS, file playback, etc.). While this application considers DASH distribution, all concepts are valid for other distribution options as well.

[0151] In most cases, audio informational messages do not overlap in the time domain; that is, only one ROI is defined at a given point in time. However, considering more advanced use cases, such as interactive environments where content can change based on user selection / movement, there may be use cases that require multiple ROIs. For this purpose, multiple audio informational messages may be needed at once. Therefore, we will discuss general solutions to support all these different use cases.

[0152] The delivery and processing of audio information messages must complement existing methods of next-generation audio delivery.

[0153] One way to transmit multiple audio information messages from multiple time-domain independent ROIs is to mix all audio information messages into a single audio element (e.g., an audio object) using associated metadata that describes the spatial location of each audio information message in different time instances. Since the audio information messages do not overlap in time, they can be individually addressed by a single shared audio element. This audio element can contain silence (or no audio data) between audio information messages, i.e., whenever there are no audio information messages. In this case, the following mechanism applies:

[0154] Audio elements, which are common audio information messages, can be delivered on the same primary stream (ES) as the associated audio scene, or on a single auxiliary stream (dependent on or independent of the main stream).

[0155] If the earphone audio elements are delivered via an auxiliary stream that depends on the main stream, the client can request additional streams whenever a new ROI exists in the visual scene.

[0156] A client (e.g., system 100) can request a stream before a scene that requires earphones, for example.

[0157] • In this example, the client can request a stream based on the current viewport. That is, if the current viewport matches the ROI, the client can decide not to request an additional earphone stream.

[0158] If the earphone audio elements are delivered on an auxiliary stream independent of the main stream, the client can, as before, request additional streams whenever a new ROI exists in the visual scene. Furthermore, two (or more) streams can be processed using two media decoders and a common rendering / mixing step to mix the decoded earphone audio data into the final audio scene. Alternatively, a metadata processor can be used to modify the metadata of the two streams, and a "stream merger" can be used to merge the two streams. Possible embodiments of such metadata processors and stream mergers are described below.

[0159] Alternatively, in another example, multiple ear cones from several ROIs, which are either time-independent or overlapping in the time domain, could be delivered by multiple audio elements (such as audio objects) and embedded in a single base stream alongside the main audio scene, or embedded in multiple auxiliary streams, for example, each ear cone in a single ES or in a group of ear cones based on a shared property (for example, all ear cones on the left share a single stream).

[0160] If all earphone audio elements are delivered via several auxiliary streams that depend on the main stream (for example, one earphone per stream or a group of earphones per stream), the client can request an additional stream containing the desired earphone whenever the ROI associated with that earphone exists in the visual scene.

[0161] The client can, for example, request a stream with the earphone before a scene that requires it (for example, based on user movement, the ROI processor 120 can make a decision even if the ROI is not yet part of the scene).

[0162] • In the example, the client can request a stream based on the current viewport, and if the current viewport matches the ROI, the client can decide not to request an additional earphone stream.

[0163] If one earphone audio element (or group of earphones) is delivered on an auxiliary stream independent of the main stream, the client can request additional streams whenever a new ROI exists in the visual scene, for example, as before. Furthermore, two (or more) streams can be processed using two media decoders and a common rendering / mixing step to mix the decoded earphone audio data into the final audio scene. Alternatively, a metadata processor can be used to modify the metadata of the two streams, and a "stream merger" can be used to merge the two streams. Possible embodiments of such metadata processors and stream mergers are described below.

[0164] Alternatively, a single common (general-purpose) earphone can be used to notify all ROIs within a single audio scene. This can be achieved by using the same audio content with different spatial information associated with audio content in different time instances. In this case, the ROI processor 120 can collect earphones associated with ROIs in the scene and request the metadata processor 132 to sequentially control the playback of the earphones (for example, upon user selection or a request from a higher-level application).

[0165] Alternatively, a single earphone can be sent only once and cached to the client. The client can then reuse it for all ROIs within a single audio scene, using different spatial information associated with audio content from different time instances.

[0166] Alternatively, the earphone audio content can be synthesized and generated on the client side. In conjunction with this, a metadata generator can be used to create the metadata necessary to communicate the spatial information of the earphones. For example, the earphone audio content can be compressed and fed into a single media decoder along with the main audio content and new metadata, or it can be mixed into the final audio scene after the media decoder, or multiple media decoders can be used.

[0167] Alternatively, the earphone audio content can be generated synthetically on the client side (for example, under the control of the metadata processor 132) while metadata describing the earphone is already embedded in the stream. The metadata may include spatial information of the earphone, a specific unification of the "earphone generated by the decoder," using specific notifications of the earphone type in the encoder, but may not include the earphone audio data.

[0168] Alternatively, the client can synthesize and generate the earphone audio content, and use a metadata generator to create the metadata necessary to communicate the earphone's spatial information. For example, the earphone audio content is • The main audio content and new metadata are compressed and fed into a single media decoder.

[0169] Alternatively, it can be mixed into the final audio scene after the media decoder.

[0170] Alternatively, multiple media decoders can be used.

[0171] 6.3 Example of metadata for audio information messages (e.g., earphones) As mentioned above, an example of audio information message (earcon) metadata 141 is presented here.

[0172] It provides a single structure for describing ear controller properties and the possibility of easily adjusting these values.

[0173] [Table 1] Each identifier in the table is intended to be associated with an attribute in IACON metadata 132.

[0174] Here, we will explain semantics.

[0175] numEarcons - This field specifies the number of earphone audio elements available in the stream.

[0176] Earcon_isIndependent - This flag defines whether an earcon audio element is independent of any audio scene. If Earcon_isIndependent==1, the earcon audio element is independent of the audio scene. If Earcon_isIndependent==0, the earcon audio element is part of the audio scene, and Earcon_id must have the same value as the mae_groupID associated with the audio element.

[0177] EarconType - This field defines the type of earphone. The following table shows the acceptable values.

[0178] [Table 2] EarconActive This flag defines whether the earphones are active. If EarconActive==1, the earphone audio elements are decoded and rendered into the audio scene.

[0179] EarconPosition This flag defines whether the earcon has available position information. If Earcon_isIndependent==0, this position information is used instead of the audio object metadata specified in the dynamic_object_metadata() or intracoded_object_metadata_efficient() structure.

[0180] Earcon_azimuth: The absolute value of the azimuth angle.

[0181] Earcon_elevation: The absolute value of the elevation angle.

[0182] Earcon_radius: The absolute value of the radius.

[0183] The EarconHasGain flag defines whether the earcon gain values ​​are different.

[0184] The `Earcon_gain` field defines the absolute value of the earcon's gain.

[0185] The `EarconHasTextLabel` flag defines whether an earcon has a text label associated with it.

[0186] The `Earcon_numLanguages` field specifies the number of available languages ​​for the descriptive text label.

[0187] The 24-bit field Earcon_Language identifies the language of the earcon's description text. It contains a 3-character code specified in ISO 639-2. Both ISO 639-2 / B and ISO 639-2 / T can be used. Each character is coded to 8 bits according to ISO / IEC 8859-1 and inserted sequentially into the 24-bit field. For example, French has a 3-character code "fre" which is coded as "0110 0110 0111 0010 0110 0101".

[0188] The Earcon_TextDataLength field defines the length of the next group description in the bitstream.

[0189] Earcon_TextData: This field contains a description of the earcon, a string that describes the content at a higher level. The format must conform to UTF-8 according to ISO / IEC 10646.

[0190] A single structure for identifying ear controllers at the system level and associating them with existing viewports. The following two tables show two ways of implementing such a structure, which can be used in various embodiments. aligned(8)class EarconSample()extends SphereRegionSample{ for(i=0;i <num_regions;i++){ unsigned int(7)reserved; unsigned int(1) has Earcon; if (hasEarcon == 1) { unsigned int(8)numRegionEarcons; for(n=0;n <numRegionEarcons;n++){ unsigned int(8)Earcon_id; unsigned int(32)Earcon_track_id; } } } } Or instead: aligned(8)class EarconSample()extends SphereRegionSample{ for(i=0;i <num_regions;i++){ unsigned int(32)Earcon_track_id; unsigned int(8)Earcon_id; } } Semantics: hasEarcon specifies whether earcon data is available in a given region.

[0191] numRegionEarcons specifies the number of earcons available in a single region.

[0192] Earcon_id uniquely defines the ID of a single earcon element associated with a spherical region. If the earcon is part of an audio scene (i.e., if the earcon is part of a group of elements identified by a single mae_groupID), then Earcon_id must have the same value as mae_groupID. Earcon_id can be used for identification in audio files / tracks, for example, in the case of DASH delivery, in the EarconComponent of MPD.

[0193] The AdaptationSet containing the tag element is equal to the Earcon_id.

[0194] The Earcon_track_id is an integer that uniquely identifies a single earcon track associated with a spherical region throughout the lifetime of a single presentation. In other words, if earcon tracks are delivered within the same ISO BMFF file, Earcon_track_id represents the corresponding track_id of the earcon track. If the earcon tracks are not delivered within the same ISO BMFF file, this value should be set to zero.

[0195] To easily identify earcon tracks at the MPD level, add the following attribute / element to EarconComponent

[0196] It can be used as a tag.

[0197] Overview of MPD elements and attributes associated with MPEG-H audio

[0198] [Table 3] In the case of MPEG-H audio, this can be done using MHAS packets, as shown in the example.

[0199] A new MHAS packet can be defined to carry information about earcon: PACTYP_EARCON, which carries the EarconInfo() structure; • A new identification field in the common MHAS METADATA MHAS packet for carrying the EarconInfo() structure.

[0200] With regard to metadata, the metadata processor 132 may have at least some of the following functions: Extract audio information message metadata from the stream, Modify the audio information message metadata to activate the audio information message, and / or set / change its position, and / or write / change the text label of the audio information message. Embed metadata into the stream The stream is fed to an additional media decoder. Extract audio metadata from at least one first audio stream (116), Extract audio information message metadata from the additional stream, Modify the audio information message metadata to activate the audio information message, and / or set / change its position, and / or write / change the text label of the audio information message. Modify the audio metadata of at least one first audio stream (116) so that the presence of audio information messages can be taken into account during merging. Based on the information received from the ROI processor, a stream is supplied to a multiplexer or maxer for multiplexing or multiplexing.

[0201] 6.4 Example in Figure 3 Figure 3 shows a system 300 on the client side 204, which includes a system 302 (client system) that can implement, for example, system 100 or 200.

[0202] The system 302 may include an ROI processor 120, a metadata processor 132, and a decoder group 313 formed by multiple decoders 112.

[0203] In this example, different audio streams are decoded (each by the media audio decoder 112), then mixed together and / or rendered to provide the final audio scene.

[0204] Here, at least one audio stream is represented as containing two streams 116 and 316 (other examples may provide a single stream or three or more streams, as shown in Figure 2). These are audio streams for playing the audio scene that the user is expected to experience. While we are referring to earphones here, it is also possible to generalize the concept of audio information messages.

[0205] Furthermore, the earphone stream 140 may be provided by the media encoder 240. Based on the user's movement and the ROI indicated in the viewport metadata 131 and / or other criteria, the ROI processor plays the earphones from the earphone stream 140 (also indicated as additional audio streams, as they are added to audio streams 116, 316).

[0206] In particular, the actual representation of the ear controller is based on the changes made by the ear controller metadata 141 and the metadata processor 132.

[0207] In this example, a stream can be requested from the media encoder 240 (server) by the system 302 (client) when needed. For example, the ROI processor can determine, based on user movement, that a specific earphone is needed immediately and therefore request the appropriate earphone stream 140 from the media encoder 240.

[0208] The following aspects of this example can be noted.

[0209] • Use case: Audio data is delivered via one or more audio streams 116, 316 (e.g., one main stream and an auxiliary stream), while the earphones are delivered via one or more additional streams 140 (dependent on or independent of the main audio stream).

[0210] In one embodiment of the client-side 204, the ROI processor 120 and the metadata processor 132 are used to efficiently process earcon information.

[0211] The ROI processor 120 can receive information 122 (user orientation information) about the current viewport from the media consumption device side 206 used for content consumption (for example, based on an HMD). The ROI processor can also receive ROIs notified in metadata (video viewports are notified as in OMAF).

[0212] Based on this information, the ROI processor 120 can decide to activate one (or more) earphones included in the earphone audio stream 140. Furthermore, the ROI processor 120 can determine different locations and different gain values ​​for the earphones (for example, for a more accurate representation of the earphones in the current space where the content is consumed).

[0213] The ROI processor 120 provides this information to the metadata processor 132.

[0214] The metadata processor 132 analyzes the metadata contained in the earphone audio stream. • Enable the earphones (to allow playback). Furthermore, if requested by the ROI processor 120, the spatial position and gain information contained in the earphone metadata 141 can be modified accordingly.

[0215] Each audio stream 116, 316, and 140 is decoded and rendered independently (based on the user's location information), and the outputs of all media decoders are mixed together as a final step by a mixer or renderer 314. In another embodiment, only compressed audio can be decoded, and the decoded audio data and metadata can be provided to a general common renderer for the final rendering of all audio elements (including earphones).

[0216] Furthermore, in a streaming environment, the ROI processor 120 can decide in advance to request the earphone stream 140 based on the same information (for example, if the user looked in the wrong direction a few seconds before the ROI becomes active).

[0217] 6.5 Example in Figure 4 Figure 4 shows a system 400 on the client side 204, which includes a system 402 (client system) that can embody, for example, system 100 or 200. Here, earphones are used as a reference, but the concept of audio information messages can also be generalized.

[0218] System 402 may include an ROI processor 120, a metadata processor 132, and a stream multiplexer or maxer 412. In the example where the multiplexer or maxer 412 is present, the number of operations performed by the hardware is favorably reduced compared to the number of operations performed when multiple decoders and one mixer or renderer are used.

[0219] In this example, different audio streams are processed based on metadata in element 412 and multiplexing or multiplexing.

[0220] Here, at least one audio stream is represented as containing two streams 116 and 316 (other examples may provide a single stream or three or more streams, as shown in Figure 2). These are the audio streams for playing the audio scene that the user is expected to experience.

[0221] Furthermore, the earphone stream 140 may be provided by the media encoder 240. Based on the user's movements and the ROI indicated in the viewport metadata 131 and / or other criteria, the ROI processor 120 plays the earphones from the earphone stream 140 (also indicated as additional audio streams, as they are added to audio streams 116, 316).

[0222] Each audio stream 116, 316, and 140 may contain metadata 236, 416, and 141, respectively. At least a portion of this metadata is manipulated and / or processed so that it is provided to a stream maxer or multiplexer 412 to which the audio stream packets are merged together. Thus, the earphones can be represented as part of an audio scene.

[0223] Therefore, the stream maxer or multiplexer 412 can provide an audio stream 414 containing modified audio metadata 238 and modified earphone metadata 234, which can then be provided to the audio decoder 112 for decoding and playback to the user.

[0224] The following aspects of this example can be noted.

[0225] • Use case: Audio data is delivered in one or more audio streams 116, 316 (for example, one main stream 116 and an auxiliary stream 316 are provided, but a single audio stream may also be provided), while the earphones are delivered in one or more additional streams 140 (dependent on or independent of the main audio stream 116).

[0226] In one embodiment of the client-side 204, the ROI processor 120 and the metadata processor 132 are used to efficiently process earcon information.

[0227] The ROI processor 120 can receive information 122 (user orientation information) about the current viewport from a media consumption device (e.g., HMD) used for content consumption. The ROI processor 120 can also receive information about the ROI notified in the earphone metadata 141 (the video viewport can be notified in Omnidirectional Media Application Format, OMAF).

[0228] Based on this information, the ROI processor 120 may decide to activate one (or more) earbuds included in the additional audio stream 140. Furthermore, the ROI processor 120 may determine different locations and different gain values ​​for the earbuds (for example, for a more accurate representation of the earbuds in the current space where the content is consumed).

[0229] The ROI processor 120 can provide this information to the metadata processor 132.

[0230] The metadata processor 132 analyzes the metadata contained in the earphone audio stream. • Enable earphones · Also, when requested by the ROI processor, the spatial position and / or gain information and / or text label included in the iAC metadata can be appropriately changed.

[0231] · The metadata processor 132 also analyzes the audio metadata 236, 416 of all the audio streams 116, 316 and can manipulate the audio-specific information so that the iAC can be used as part of the audio scene (for example, there is an audio scene 5.1 channel bed and four objects, and the iAC audio element is added to the scene as a fifth object. All metadata fields are updated accordingly).

[0232] · The audio data, the modified audio metadata, and the iAC metadata of each stream 116, 316 are provided to a stream muxer or multiplexer that can generate one audio stream 414 having a set of metadata (modified audio metadata 238 and modified iAC metadata 234) based on this.

[0233] · This stream 414 may be decoded by a single media audio decoder 112 based on the user position information 122.

[0234] · Further, in a streaming environment, the ROI processor 120 can determine to request the iAC stream 140 in advance based on the same information (for example, when the user looks in the wrong direction a few seconds before the ROI becomes effective).

[0235] 6.6 Example of FIG. 5 FIG. 5 shows a system 500 including a system 502 (client system) that can implement, for example, the system 100 or 200 on the client side 204. Here, an iAC is referred to, but it is also possible to generalize the concept of the audio information message.

[0236] System 502 may include an ROI processor 120, a metadata processor 132, and a stream multiplexer or maxar 412.

[0237] In this example, the earcon stream is not provided by the remote entity (on the client side), but is generated by the synthetic audio generator 236 (for later reuse or using stored compressed / uncompressed versions of natural sounds). The earcon metadata 141 is provided by the remote entity, for example, in the audio stream 316 (not the earcon stream). Thus, the synthetic audio generator 236 can be activated to create the audio stream 140 based on the attributes of the earcon metadata 141. For example, the attributes could refer to the type of synthesized speech (natural sound, synthesized speech, speech text, etc.) and / or a text label (the earcon can be generated by creating a synthesized speech based on the text of the metadata). In the example, after the earcon stream is created, the same thing is stored for future reuse. Alternatively, the synthesized speech may be a generic sound persistently stored on the device.

[0238] Using a stream maxer or multiplexer 412, packets of audio stream 116 (and other streams such as auxiliary audio stream 316) can be merged with packets of earphone stream generated by generator 236. The audio stream 414 associated with the modified audio metadata 238 and modified earphone metadata 234 can then be obtained. The audio stream 414 may be decoded by decoder 112 and played back to the user on media consumption device 206.

[0239] The following aspects of this example can be noted.

[0240] ·Use case: • Audio data is delivered via one or more audio streams (for example, one main stream and an auxiliary stream).

[0241] • The earphones are not delivered from the remote device, but the earphone metadata 141 is delivered as part of the main audio stream (a specific notification may be used to indicate that no audio data is associated with the earphones).

[0242] In one client-side embodiment, the ROI processor 120 and the metadata processor 132 are used to efficiently process earcon information.

[0243] The ROI processor 120 can receive information about the current viewport (user orientation information) from the device used by the content consuming device 206 (e.g., HMD). The ROI processor 120 can also receive ROIs notified by metadata (video viewports are notified as in OMAF).

[0244] Based on this information, the ROI processor 120 may decide to activate one (or more) earbuds that are not present in stream 116. Furthermore, the ROI processor 120 may determine different locations and different gain values ​​for earbuds (for example, for a more accurate representation of earbuds in the current space where content is consumed).

[0245] The ROI processor 120 can provide this information to the metadata processor 132.

[0246] The metadata processor 120 analyzes the metadata contained in the audio stream 116, • Enable earphones Furthermore, if requested by the ROI processor 120, the spatial position and gain information contained in the earphone metadata 141 can be modified accordingly.

[0247] · The metadata processor 132 also analyzes the audio metadata (such as 236, 417) of all audio streams (116, 316), and can manipulate the audio-specific information so that the iacon can be used as part of the audio scene (for example, there is an audio scene 5.1 channel bed and four objects, and the iacon audio element is added to the scene as the fifth object. All metadata fields are updated accordingly).

[0248] · The modified iacon metadata and the information from the ROI processor 120 are provided to the synthetic audio generator 246. The synthetic audio generator 246 can create a synthetic sound based on the received information (for example, based on the spatial position of the iacon, an audio signal is generated to spell out the position). Also, the iacon metadata 141 is associated with the generated audio data to become a new stream 414.

[0249] · Similarly, as before, the audio data (116, 316) of each stream and the modified audio metadata and iacon metadata are provided to the stream muxer, and the stream muxer can generate based on this one audio stream having a set of metadata (audio and iacon).

[0250] · This stream 414 is decoded by a single media audio decoder 112 based on the user's position information.

[0251] · Instead or additionally, the audio data of the iacon can be monetized at the client (for example, from previous iacon usage).

[0252] · Alternatively, the output of the synthetic audio generator 246 can be uncompressed audio and can be mixed into the final rendered scene.

[0253] Furthermore, in a streaming environment, based on the same information, the ROI processor 120 can decide in advance to request an earphone stream (for example, if the user looked in the wrong direction a few seconds before the ROI becomes active).

[0254] 6.7 Example in Figure 6 Figure 6 shows a system 600 on the client side 204, which includes a system 602 (client system) that can embody, for example, system 100 or 200. Here, ear controllers are referred to, but the concept of audio information messages can also be generalized.

[0255] System 602 may include an ROI processor 120, a metadata processor 132, and a stream multiplexer or maxar 412.

[0256] In this example, the earphone stream is not provided by the remote entity (on the client side), but is generated by the synthesized audio generator 236 (which can store the stream for later reuse).

[0257] In this example, the earcon metadata 141 is not provided by a remote entity. The earcon metadata is generated by a metadata generator 432, which can generate earcon metadata that is used (e.g., processed, manipulated, or modified) by the metadata processor 132. The earcon metadata 141 generated by the earcon metadata generator 432 may have the same structure and / or format and / or attributes as the earcon metadata described in the previous example.

[0258] The metadata processor 132 can operate as shown in the example in Figure 5. Based on the attributes of the earphone metadata 141, the synthesized audio generator 246 can be activated to create the audio stream 140. For example, the attributes may refer to the type of synthesized speech (natural sound, synthesized voice, speech text, etc.), and / or gain, and / or activation / deactivation status, etc. In the example, after the earphone stream 140 is created, it may be stored (e.g., cached) for future reuse. Earphone metadata generated by the earphone metadata generator 432 can also be stored (e.g., cached).

[0259] Using a stream maxer or multiplexer 412, packets of audio stream 116 (and other streams such as auxiliary audio stream 316) can be merged with packets of earphone stream generated by generator 246. The audio stream 414 associated with the modified audio metadata 238 and modified earphone metadata 234 can then be obtained. The audio stream 414 may be decoded by decoder 112 and played back to the user on media consumption device 206.

[0260] The following aspects of this example can be noted.

[0261] ·Use case: Audio data is delivered via one or more audio streams (for example, one main stream 116 and an auxiliary stream 316).

[0262] • The earphones will not be streamed from client-side 202. • Client-side 202 does not deliver earphone metadata.

[0263] This use case can represent a solution for enabling earphones in legacy content created without them.

[0264] In one client-side embodiment, the ROI processor 120 and the metadata processor 232 are used to efficiently process earcon information.

[0265] The ROI processor 120 can receive information 122 (user orientation information) about the current viewport from the device used by the content consuming device 206 (e.g., HMD). The ROI processor 210 can also receive ROIs and ROIs notified by metadata (video viewports are notified as in OMAF).

[0266] Based on this information, the ROI processor 120 can decide to activate one (or more) earphones that are not present in the streams (116, 316).

[0267] Furthermore, the ROI processor 120 can provide the earcon metadata generator 432 with information regarding the earcon's position and gain value.

[0268] The ROI processor 120 can provide this information to the metadata processor 232.

[0269] The metadata processor 232 analyzes the metadata contained in the earphone audio stream (if any), • Enable earphones • If requested by the ROI processor 120, the spatial position and gain information included in the earcon metadata can be modified accordingly.

[0270] The metadata processor also parses the audio metadata 236 and 417 of all audio streams 116 and 316, and can manipulate audio-specific information so that the earcon can be used as part of the audio scene (for example, if there is an audio scene 5.1 channel bed and 4 objects, and the earcon audio element is added to the scene as a fifth object, all metadata fields are updated accordingly).

[0271] The modified earphone metadata 234 and information from the ROI processor 120 are provided to the synthesized audio generator 246. The synthesized audio generator 246 can create synthesized sound based on the received information (for example, based on the spatial position of the earphones, an audio signal is generated to spell out the position). The earphone metadata is also associated with the generated audio data to form a new stream.

[0272] Similarly, as before, the audio data and modified audio metadata and earphone metadata for each stream are provided to a stream maxer or multiplexer 412, which can generate a set of metadata (audio and earphone) based on this one audio stream 414.

[0273] This stream 414 is decoded by a single media audio decoder based on user location information.

[0274] Alternatively, the client can monetize the audio data from the earphones (for example, from previous earphone usage).

[0275] Alternatively, the output of the synthesized audio generator can be uncompressed audio, which can then be mixed into the final rendered scene. Furthermore, in a streaming environment, the ROI processor 120 can decide in advance to request an earphone stream based on the same information (for example, if the user looked in the wrong direction a few seconds before the ROI becomes active).

[0276] 6.8 Example based on user location A feature can be implemented that allows the earphones to play only if the user does not display the ROI.

[0277] The ROI processor 120 can periodically check, for example, the user's current viewport and / or position and / or head orientation and / or motion data 122. If the ROI is displayed to the user, earphone playback is not performed.

[0278] If the ROI processor determines, based on the user's current viewport and / or position and / or head orientation and / or motion data, that the ROI is not visible to the user, the ROI processor 120 can request earcon playback. In this case, the ROI processor 120 can cause the metadata processor 132 to prepare earcon playback. The metadata processor 132 can use one of the techniques described in the above example. For example, metadata can be obtained in a stream delivered by the server-side 202 and generated by the earcon metadata generator 432. The attributes of the earcon metadata can be easily changed based on the ROI processor's request and / or various conditions. For example, if the earcon was previously disabled by the user's choice, the earcon will not be played even if the user is not looking at the ROI. For example, if a timer (previously set) has not yet expired, the earcon will not be played even if the user is not looking at the ROI.

[0279] Furthermore, if the ROI processor determines from the user's current viewport and / or position and / or head orientation and / or motion data that the ROI is visible to the user, the ROI processor 120 can request that the earphones not be played, especially if the earphone metadata already contains notifications for active earphones.

[0280] In this case, the ROI processor 120 can cause the metadata processor 132 to disable earcon playback. The metadata processor 132 can use one of the techniques described for the above example. For example, metadata can be obtained in a stream delivered by the server-side 202 and can be generated by the earcon metadata generator 432. The attributes of the earcon metadata can be easily modified based on the ROI processor's request and / or various conditions. If the metadata already contains an instruction that the earcon needs to be played back, in this case the metadata is modified to indicate that the earcon is inactive and cannot be played back.

[0281] The following aspects of this example can be noted.

[0282] ·Use case: Audio data is delivered via one or more audio streams 116, 316 (for example, one main stream and an auxiliary stream), but the earphones are delivered via either the same one or more audio streams 116, 316, or one or more additional streams 140 (dependent on or independent of the main audio stream).

[0283] • The earphone metadata is configured to indicate that the earphones are always active at a specific moment.

[0284] • First-generation devices without an ROI processor read earcon metadata, and the user's current viewport and / or position and / or head orientation and / or motion data trigger the earcon to play, regardless of the fact that the ROI is visible to the user.

[0285] • A new generation of devices, including an ROI processor as described in either system, utilizes the ROI processor's decisions. If the ROI processor determines, based on the user's current viewport and / or position and / or head orientation and / or motion data, that the ROI is visible to the user, the ROI processor 120 can request that the earcon playback not occur, especially if the earcon metadata already contains a notification that the earcon is active. In this case, the ROI processor 120 can cause the metadata processor 132 to disable earcon playback. The metadata processor 132 can use one of the techniques described in the above example. For example, the metadata can be obtained in a stream delivered by the server-side 202 and can be generated by the earcon metadata generator 432. The attributes of the earcon metadata can be easily modified based on the ROI processor's request and / or various conditions. If the metadata already contains an instruction that the earcon needs to be played, in this case the metadata is modified to indicate that the earcon is inactive and cannot be played.

[0286] Furthermore, depending on the playback device, the ROI processor may request changes to the earphone metadata. For example, the spatial information of the earphones may be modified differently depending on whether the sound is played through headphones or speakers.

[0287] Therefore, the final audio scene experienced by the user is obtained based on the metadata changes performed by the metadata processor.

[0288] 6.9 Example based on server-client communication (Figure 5a) Figure 5a shows a system 550 that includes a system 552 (client system) which can embody, for example, system 100, 200, 300, 400, or 500 on the client side 204. Here, ear controllers are referred to, but the concept of audio information messages can also be generalized.

[0289] System 552 may include an ROI processor 120, a metadata processor 132, and a stream multiplexer or maxer 412. (In the example, different audio streams are decoded (each by a media audio decoder 112), then mixed together and / or rendered to provide the final audio scene.)

[0290] Here, at least one audio stream is represented as containing two streams 116 and 316 (other examples may provide a single stream or three or more streams, as shown in Figure 2). These are the audio streams for playing the audio scene that the user is expected to experience.

[0291] Furthermore, the earphone stream 140 may be provided by a media encoder 240.

[0292] Audio streams can be encoded at various bitrates, enabling efficient bitrate adaptation depending on the network connection (i.e., users with high-speed connections receive a higher bitrate encoded version, while users with slower network connections receive a lower bitrate version).

[0293] The audio streams may be stored in the media server 554, and for each audio stream, different encodings at different bitrates are grouped into one adaptation set 556 along with appropriate data indicating the availability of all created adaptation sets. Audio adaptation set 556 and video adaptation set 557 may be provided.

[0294] Based on user movement and the ROI indicated in viewport metadata 131 and / or other criteria, the ROI processor 120 plays the earphones from earphone stream 140 (also indicated as additional audio streams, as they are added to audio streams 116 and 316).

[0295] In this example: Client 552 is configured to receive data from the server regarding the availability of all adaptation sets.

[0296] • At least one audio scene adaptation set for at least one audio stream. And • At least one audio message adaptation set for at least one additional audio stream containing at least one audio informational message. Similar to other exemplary embodiments, the ROI processor 120 can receive information 122 (user orientation information) about the current viewport from the media consuming device side 206 used for content consumption (e.g., based on an HMD). The ROI processor 120 can also receive ROIs and ROIs notified by metadata (video viewports are notified as in OMAF).

[0297] Based on this information, the ROI processor 120 can decide to activate one (or more) earphones included in the earphone audio stream 140.

[0298] Furthermore, the ROI processor 120 can determine different locations and different gain values ​​for the earphones (for example, for a more accurate representation of the earphones in the current space where the content is consumed).

[0299] The ROI processor 120 can provide this information to the selection data generator 558.

[0300] The selection data generator 558 may be configured to create selection data 559 that, based on the ROI processor's decision, identifies which adaptation sets to receive. The adaptation sets include audio scene adaptation sets and audio message adaptation sets.

[0301] The media server 554 may be configured to provide command data to the client 552, causing the streaming client to retrieve data for adaptation sets 556, 557, which are identified by selection data that specifies which adaptation set to receive. The adaptation sets include audio scene adaptation sets and audio message adaptation sets.

[0302] The download and switching module 560 is configured to receive requested audio streams from the media server 554 based on selection data that identifies which adaptation sets to receive. The adaptation sets include audio scene adaptation sets and audio message adaptation sets. The download and switching module 560 may be further configured to provide audio metadata and earphone metadata 141 to the metadata processor 132.

[0303] The ROI processor 120 can provide this information to the metadata processor 132.

[0304] The metadata processor 132 analyzes the metadata contained in the earphone audio stream 140. • Enable the earphones (to allow playback). Furthermore, if requested by the ROI processor 120, the spatial position and gain information contained in the earphone metadata 141 can be modified accordingly.

[0305] The metadata processor 132 also parses the audio metadata of all audio streams 116,316 and can manipulate audio-specific information so that the earcon can be used as part of the audio scene (for example, if there is an audio scene 5.1 channel bed and four objects, and the earcon audio element is added to the scene as a fifth object, all metadata fields may be updated accordingly).

[0306] The audio data and modified audio metadata and earphone metadata of each stream 116, 316 may be provided to a stream maxer or multiplexer that can generate a single audio stream 414 having a set of metadata (modified audio metadata 238 and modified earphone metadata 234) based on this.

[0307] This stream may be decoded by a single media audio decoder 112 based on user location information 122.

[0308] An adaptation set may consist of a set of representations containing interchangeable versions of each piece of content, for example, different audio bitrates (e.g., different streams with different bitrates). While theoretically one representation may suffice to provide a playable stream, using multiple representations allows the client to adapt the media stream to current network conditions and bandwidth requirements, ensuring smooth playback.

[0309] 6.10 Method All of the above examples can be carried out by method steps. Here, method 700 (which can be carried out by any of the above examples) is fully described. This method includes the following:

[0310] In step 702, at least one video stream (106) and at least one first audio stream (116, 316) are received.

[0311] In step 704, decode at least one video signal from at least one video stream (106) to represent a VR, AR, MR, or 360-degree video environment scene (118a) to the user.

[0312] In step 706, in order to represent the audio scene (118b) to the user, decode at least one audio signal from at least one first audio stream (116, 316), Receives data (122) of the user's current viewport and / or position and / or head orientation and / or motion.

[0313] In step 708, viewport metadata (131) associated with at least one video signal from at least one video stream (106) is received, and the viewport metadata defines at least one ROI.

[0314] In step 710, determine whether to play an audio informational message associated with at least one ROI based on the user's current viewport and / or position and / or head orientation and / or motion data (122) and viewport metadata and / or other criteria.

[0315] In step 712, receive, process, and / or manipulate audio information message metadata (141) that describes the audio information message in order to play the audio information message according to the audio information message attributes in such a way that the audio information message is part of the audio scene.

[0316] In particular, the sequence may also be different. For example, receiving steps 702, 706, and 708 may have a different order depending on the actual order in which the information is delivered.

[0317] Line 714 refers to the fact that the method may be repeated. If the ROI processor decides not to play the audio information message, step 712 is skipped.

[0318] 6.11 Other Embodiments Figure 8 shows a system 800 that can implement one of the systems (or its components) or perform method 700. System 800 may include a processor 802 and a non-temporary memory unit 806 that stores instructions that, when executed by the processor 802, can cause the processor to perform at least the above-described stream processing operations and / or metadata processing operations. System 800 may include an input / output unit 804 for connection to external devices.

[0319] System 800 can implement at least some (or all) of the following functions: ROI processor 120, metadata processor 232, generator 246, maxer or multiplexer 412, decoder 112m, earcon metadata generator 432, etc.

[0320] Depending on the specific embodiment, the embodiments can be implemented in hardware. The embodiments can be implemented using digital storage media that store electronically readable control signals and which cooperate (or can cooperate) with a computer system that is programmable to perform each method, such as floppy disks, digital multipurpose disks (DVDs), Blu-ray discs, compact discs (CDs), read-only memory (ROMs), programmable read-only memory (PROMs), erasable and programmable read-only memory (EPROMs), electrically erasable programmable read-only memory (EEPROMs), or flash memory. Thus, the digital storage media may be computer-readable.

[0321] Generally, the embodiments may be implemented as a computer program product including program instructions, which operate to perform one of the methods when the computer program product is executed on a computer. The program instructions may be stored, for example, in a machine-readable medium.

[0322] Other embodiments include a computer program for performing one of the methods described herein, stored in a machine-readable carrier. In other words, an example of the method is a computer program having program instructions for performing one of the methods described herein when the computer program is executed on a computer.

[0323] Therefore, a further example of the present method includes a computer program for performing one of the methods described herein, which is recorded on a data carrier medium (or digital storage medium, or computer-readable medium). The data carrier medium, digital storage medium, or recorded medium is tangible and / or non-temporary, rather than intangible and transient signals.

[0324] Further examples include processing units, such as computers or programmable logical devices, that perform one of the methods described herein.

[0325] Further examples include a computer on which a computer program for performing one of the methods described herein is installed.

[0326] Further examples include devices or systems for transferring (e.g., electronically or optically) a computer program to perform one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The device or system may include, for example, a file server for transferring the computer program to the receiver.

[0327] In some examples, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the method described herein. In some examples, the field-programmable gate array may work with a microprocessor to perform one of the methods described herein. In general, the method may be performed by any suitable hardware device.

[0328] The above examples illustrate the principles described above. It should be understood that modifications and changes to the arrangements and details described herein are obvious. Therefore, it is intended to be limited by the imminent claims rather than by the specific details presented as part of the description and explanation of the embodiments herein.

Claims

1. A system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, wherein the system is The system receives at least one video stream (106) associated with the audio and video scene to be played back. It is configured to receive at least one first audio stream (116, 316) associated with the audio and video scenes being played back, The aforementioned system, For the representation of the audio and video scenes to the user, at least one media video decoder (102) configured to decode at least one video signal from the at least one video stream (106), For the representation of the audio and video scenes to the user, at least one media audio decoder (112) configured to decode at least one audio signal from the at least one first audio stream (116, 316), The system includes a region of interest ROI processor (120), wherein the region of interest ROI processor (120) is Based on at least the user's current viewport and / or head orientation and / or motion data (122) and / or viewport metadata (131) and / or audio information message metadata (141), a decision is made as to whether to play the audio information message associated with the at least one ROI, wherein the audio information message is independent of the at least one video signal and the at least one audio signal. When it is decided to play the aforementioned informational message, the audio informational message is played. A system configured in such a way.

2. A system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, wherein the system is Receive at least one video stream (106), It is configured to receive at least one first audio stream (116, 316), The aforementioned system, To represent a VR, AR, MR, or 360-degree video environment scene (118a) to the user, at least one media video decoder (102) configured to decode at least one video signal from the at least one video stream (106), For the representation of the audio scene (118b) to the user, at least one media audio decoder (112) configured to decode at least one audio signal from the at least one first audio stream (116, 316), The system includes a region of interest ROI processor (120), wherein the region of interest ROI processor (120) is Based on the user's current viewport and / or head orientation and / or motion data (122) and / or viewport metadata (131) and / or audio information message metadata (141), it is determined whether to play the audio information message associated with the at least one ROI, wherein the audio information message is an earphone. When it is decided to play the aforementioned informational message, the audio informational message is played. A system configured in such a way.

3. The system according to claim 1 or 2, further comprising a metadata processor (132) configured to receive and / or process and / or manipulate audio information message metadata (141) and, when it decides to play the information message, to play the audio information message according to the audio information message metadata (141).

4. The ROI processor (120) is The system receives data on the user's current viewport and / or position and / or head orientation and / or motion and / or other user-related data (122), The viewport metadata (131) associated with at least one video signal is received from the at least one video stream (106), and the viewport metadata (131) defines at least one ROI. Based on the user's current viewport and / or position and / or head orientation and / or motion data (122) and at least one of the viewport metadata, it is determined whether to play the audio information message associated with the at least one ROI. The system according to any one of claims 1 to 3, configured as described above.

5. The system according to any one of claims 1 to 4, further comprising a metadata processor (132) configured to receive and / or process and / or manipulate audio metadata (236) and / or viewport metadata (131) describing the audio information message metadata (141) and / or at least one audio signal encoded in at least one audio stream (116) to reproduce the audio information message according to the audio metadata (236) and / or viewport metadata (131).

6. The ROI processor (120) is If the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or motion data (122), in addition to playing the at least one audio signal, the audio information message associated with the at least one ROI is played. If the at least one ROI is in the user's current viewport and / or position and / or head orientation and / or motion data (122), then the playback of the audio information message associated with the at least one ROI is disabled and / or deactivated. The system according to any one of claims 1 to 5, configured as described above.

7. Further configured to receive the at least one additional audio stream (140) in which the at least one audio information message is encoded, The aforementioned system, The system according to any one of claims 1 to 6, further comprising at least one maxer or multiplexer (412) which, under the control of the metadata processor (132) and / or the ROI processor (120) and / or another processor, merges packets of the at least one additional audio stream (140) with packets of the at least one first audio stream (116, 316) in one stream (414), and plays the audio information message in addition to the audio scene, based on the decision provided by the ROI processor (120) to play the at least one audio information message.

8. The system receives at least one audio metadata (236) describing the at least one audio signal encoded in the at least one audio stream (116), The system receives audio information message metadata (141) associated with at least one audio information message from at least one audio stream (116), When it is decided to play the information message, in addition to playing the at least one audio signal, the audio information message metadata (141) is modified to enable the playback of the audio information message. The system according to any one of claims 1 to 7, further configured as follows.

9. The system receives at least one audio metadata (141) describing the at least one audio signal encoded in the at least one audio stream (116), The audio information message metadata (141) associated with at least one audio information message is received from the at least one audio stream (116). When it is decided to play the audio information message, in addition to playing the at least one audio signal, the audio information message metadata (141) is modified to enable the playback of the audio information message associated with the at least one ROI. Modify the audio metadata (236) describing the at least one audio signal to enable merging of the at least one first audio stream (116) and the at least one additional audio stream (140). The system according to any one of claims 1 to 8, further configured as follows.

10. The system receives at least one audio metadata (236) describing the at least one audio signal encoded in the at least one audio stream (116), The system receives audio information message metadata (141) associated with at least one audio information message from at least one audio stream (116), When it is decided to play the audio information message, the audio information message metadata (141) is provided to a synthesis audio generator (246) to create a synthesis audio stream (140), the audio information message metadata (141) is associated with the synthesis audio stream (140), and the synthesis audio stream (140) and the audio information message metadata (141) are provided to a multiplexer or maxer (412) to enable merging of the at least one audio stream (116) and the synthesis audio stream (140). The system according to any one of claims 1 to 9, further configured as follows.

11. The system according to any one of claims 1 to 10, further configured to retrieve the audio information message metadata (141) from the at least one additional audio stream (140) in which the audio information message is encoded.

12. The system according to any one of claims 1 to 11, further comprising an audio information message metadata generator (432) configured to generate audio information message metadata (141) based on the decision to play an audio information message associated with the at least one ROI.

13. The system according to any one of claims 1 to 12, further configured to store the audio information message metadata (141) and / or the audio information message stream (140) for future use.

14. The system further includes a synthesis audio generator (432) configured to synthesize an audio information message based on the audio information message metadata (141) associated with the at least one ROI, The system according to any one of claims 1 to 13.

15. The system according to any one of claims 1 to 14, wherein the metadata processor (132) is configured to control a maxer or multiplexer (412) to merge packets of the audio information message stream (140) with packets of the at least one first audio stream (116) in a single stream (414) in order to obtain the addition of the audio information message to the at least one audio stream (116) based on the audio metadata and / or audio information message metadata.

16. The audio information message metadata (141) is encoded into a configuration frame and / or a data frame, and the data frame is Identification tag, An integer that uniquely identifies the playback of the aforementioned audio information message metadata, The type of the aforementioned message, status Display of dependency / independence from the aforementioned scene, Location data, Gain data, Displaying the existence of associated text labels. Number of available languages, The language of the aforementioned audio information message, Length of data text, The associated text label's data text, and / or The system according to any one of claims 1 to 15, comprising at least one of the descriptions of the audio information message.

17. The metadata processor (132) and / or the ROI processor (120) Extract audio information message metadata from the stream, Modify the audio information message metadata to activate the audio information message and / or set / change its position. Embed metadata into the stream The aforementioned stream is supplied to an additional media decoder, Audio metadata is extracted from the at least one first audio stream (116) mentioned above. Extract audio information message metadata from the additional stream, Modify the audio information message metadata to activate the audio information message and / or set / change its position. Modify the audio metadata of at least one first audio stream (116) so that it can be merged taking into account the presence of the aforementioned audio information message, The system according to any one of claims 1 to 16, configured to perform at least one of the operations of supplying a stream to a multiplexer or maxer for multiplexing or multiplexing the information received from the ROI processor.

18. The system according to any one of claims 1 to 17, wherein the ROI processor (120) is configured to perform a local lookup for an additional audio stream (140) and / or audio information message metadata in which the audio information message is encoded, and if it cannot be found, to request the additional audio stream (140) and / or audio information message metadata from a remote entity.

19. The system according to any one of claims 1 to 18, wherein the ROI processor (120) is configured to perform a local search for additional audio streams (140) and / or audio information message metadata, and if it is not possible to find them, to cause a synthesized audio generator (432) to generate the audio information message streams and / or audio information message metadata.

20. Receiving the at least one additional audio stream (140) which includes at least one audio information message associated with the at least one ROI, If the ROI processor decides to play back the audio information message associated with the at least one ROI, it decodes the at least one additional audio stream (140). The system according to any one of claims 1 to 19, further configured as follows.

21. A first audio decoder (112) for decoding the at least one audio signal from at least one first audio stream (116), At least one additional audio decoder (112) for decoding the at least one audio information message from an additional audio stream (140), A mixer and / or renderer (314) for mixing and / or superimposing the audio information messages from the at least one additional audio stream (140) with the at least one audio signal from the at least one first audio stream (116), The system according to claim 20, further comprising:

22. The system according to any one of claims 1 to 21, further configured to maintain a track of metrics associated with historical data and / or statistical data associated with the playback of the audio information message, and to disable the playback of the audio information message if the metrics exceed a predetermined threshold.

23. The system according to any one of claims 1 to 22, wherein the determination of the ROI processor is based on a prediction of the user's current viewport and / or position and / or head orientation and / or motion data (122) in relation to the position of the ROI.

24. The system according to any one of claims 1 to 23, further configured to receive at least one first audio stream (116) and, when it is determined to play the informational message, to request an audio message informational stream from a remote entity.

25. The system according to any one of claims 1 to 24, further configured to determine whether to play two audio information messages simultaneously or to select a higher-priority audio information message to be played preferentially over a lower-priority audio information message.

26. The system according to any one of claims 1 to 25, further configured to identify an audio information message from among a plurality of audio information messages encoded in one additional audio stream (140) based on the address and / or position of the audio information message in the audio stream.

27. The system according to any one of claims 1 to 26, wherein the audio stream is formatted in the MPEG-H 3D audio stream format.

28. The system receives data relating to the availability of multiple adaptation sets (556, 557), wherein the available adaptation sets include an adaptation set of at least one audio scene of the at least one first audio stream (116, 316) and an adaptation set of at least one audio message of the at least one additional audio stream (140) which includes at least one audio information message. Based on the decision of the ROI processor, selection data (559) is created to identify which of the adaptation sets to search, wherein the available adaptation sets include at least one audio scene adaptation set and / or at least one audio message adaptation set. Request and / or retrieve the data of the adaptation set identified by the selected data. Each adaptation set groups different encodings with different bitrates. The system according to any one of claims 1 to 27, further configured as follows.

29. The system according to claim 28, wherein at least one of its elements is configured to include, and / or to retrieve the data for each of the adaptation sets using HTTP, DASH, dynamic adaptive streaming via client, and / or an ISO-based media file format ISO BMFF, or an MPEG-2 transport stream MPEG-2 TS.

30. The system according to any one of claims 1 to 29, wherein the ROI processor (120) is configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or motion data (122) in order to check whether the ROI is represented in the current viewport, and to notify the user of the presence of the ROI by voice if the ROI is outside the current viewport and / or position and / or head orientation and / or motion data (122).

31. The system according to any one of claims 1 to 30, wherein the ROI processor (120) checks the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or motion data (122) to check whether the ROI is represented in the current viewport, and is configured to suppress audible notification of the presence of the ROI to the user if the ROI is present in the current viewport and / or position and / or head orientation and / or motion data (122).

32. The system according to any one of claims 1 to 31, configured to receive from a remote entity (202) the at least one video stream (116) associated with the video environment scene and the at least one audio stream (106) associated with the audio scene, wherein the audio scene is associated with the video environment scene.

33. The system according to any one of claims 1 to 32, wherein the ROI processor (120) is configured to select the playback of one first audio information message preceding a second audio information message from among a plurality of audio information messages to be played back.

34. The system according to any one of claims 1 to 33, further comprising a cache memory (246) for storing audio information messages received from or synthesized from a remote entity (204) and for reusing the audio information messages in different time instances.

35. The system according to any one of claims 1 and 3 to 34, wherein the audio information message is an earphone.

36. The system according to any one of claims 1 to 35, wherein the at least one video stream and / or the at least one first audio stream are each part of the current video environment scene and / or video audio scene and are independent of the user's current viewport and / or head orientation and / or motion data (122) in the current video environment scene and / or video audio scene.

37. The system according to any one of claims 1 to 36, configured to request the at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environment stream, respectively, and to play the at least one audio informational message based on the user's current viewport and / or head orientation and / or motion data (122).

38. The system according to any one of claims 1 to 37, configured to request the at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environment stream, respectively, and to request the at least one audio informational message from the remote entity based on the user's current viewport and / or head orientation and / or motion data (122).

39. The system according to any one of claims 1 to 38, configured to request the at least one first audio stream and / or at least one video stream from remote entities associated with the audio stream and / or video environment stream, respectively, and to synthesize the at least one audio informational message based on the user's current viewport and / or head orientation and / or motion data (122).

40. The system according to any one of claims 1 to 39, configured to check at least one of the additional criteria for playback of the audio information message, wherein the criteria further include user selection and / or user settings.

41. The system according to any one of claims 1 to 40, configured to check at least one of the additional criteria for playback of the audio information message, wherein the criteria further include the state of the system.

42. The system according to any one of claims 1 to 41, configured to check at least one of additional criteria for the playback of the audio information message, the criteria further comprising the number of times the audio information message has been played back.

43. The system according to any one of claims 1 to 42, configured to check at least one of the additional criteria for playback of the audio information message, wherein the criteria further include a flag in a data stream obtained from a remote entity.

44. A system comprising a client configured as a system according to any one of claims 1 to 43, and remote entities (202, 240) configured as servers for distributing the at least one video stream (106) and the at least one audio stream (116).

45. The system according to claim 44, wherein the remote entities (202, 240) are configured to retrieve the at least one additional audio stream (140) and / or audio information message metadata in a database, intranet, internet, and / or geographic network, and to deliver the at least one additional audio stream (140) and / or audio information message metadata if found.

46. The system according to claim 45, wherein the remote entities (202, 240) are configured to synthesize the at least one additional audio stream (140) and / or generate the audio information message metadata.

47. A method for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, The steps include decoding at least one video signal from at least one video and audio scene to be played back to the user, The steps include decoding at least one audio signal from the video and audio scenes to be played back, A step of determining whether to play an audio information message associated with the at least one ROI based on the user's current viewport and / or head orientation and / or motion data (122) and / or metadata, wherein the audio information message is independent of the at least one video signal and the at least one audio signal, When it is decided to play the aforementioned informational message, the steps include playing the aforementioned audio informational message, A method that includes this.

48. A method for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, The steps include decoding at least one video signal from the at least one video stream (106) in order to present a VR, AR, MR, or 360-degree video environment scene (118a) to the user, The steps include: for representing the audio scene (118b) to the user, decoding at least one audio signal from the at least one first audio stream (116, 316); A step of determining whether to play an audio information message associated with the at least one ROI based on the user's current viewport and / or head orientation and / or motion data (122) and / or metadata, wherein the audio information message is an earphone, When it is decided to play the aforementioned informational message, the steps include playing the aforementioned audio informational message, A method that includes this.

49. The method according to claim 47 or 48, further comprising the step of receiving and / or processing and / or manipulating the metadata (141) in order to play the audio information message according to the metadata (141) such that the audio information message is part of the audio scene, once it is decided to play the information message.

50. The steps include playing the aforementioned audio and video scenes, The steps include determining to further play the audio information message based on the user's current viewport and / or head orientation and / or movement data (122) and / or metadata, The method according to any one of claims 47 to 49, further comprising:

51. The steps include playing the aforementioned audio and video scenes, If the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or motion data (122), in addition to playing the at least one audio signal, play an audio information message associated with the at least one ROI, and / or If the at least one ROI is in the user's current viewport and / or position and / or head orientation and / or motion data (122), the steps include disallowing and / or deactivating the playback of the audio information message associated with the at least one ROI, The method according to any one of claims 47 to 50, further comprising:

52. A non-transient storage unit, which, when executed by a processor, includes an instruction causing the processor to perform the method according to any one of claims 47 to 51.

Citation Information

Patent Citations

  • IEC23008-3