Method and apparatus for efficient delivery and use of audio messages for a high quality experience - Patents.com
The system addresses the limitations of existing audio message delivery in VR/AR by dynamically controlling earcons based on user interaction, enhancing user experience through flexible and customizable audio message delivery.
Patent Information
- Application Number
- JP2024003075
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-10-12
- Filing Date
- 2024-01-12
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2038-10-10
AI Technical Summary
Existing systems for delivering audio messages in VR, AR, and 360-degree video environments face issues such as user irritation from excessive earcons, inability to pinpoint earcons to a single audio stream, and lack of flexibility in accommodating different types of earcons, leading to a poor user experience.
A system that determines whether to play audio information messages based on user viewport, head orientation, and movement data, using metadata processors to enable or disable earcons and merge them with audio streams, allowing for flexible and user-specific delivery of audio messages.
Enhances user experience by providing customizable and efficient delivery of audio messages, reducing user annoyance and improving content consumption quality.
Smart Images

Figure 0007801052000004 
Figure 0007801052000005 
Figure 0007801052000006
Abstract
Description
[Technical Field]
[0001] [Background technology]
[0002] 1. Introduction In many applications, the delivery of audible messages can improve the user experience during media consumption. One of the most relevant applications of such messages is provided by virtual reality (VR) content. In a VR environment, or similarly in augmented reality (AR) or mixed reality (MR) or 360-degree video environments, the user typically visualizes the entire 360-degree content using, for example, a head-mounted display (HMD) and listens to it through headphones (or, equivalently, through speakers with the correct rendering depending on the speaker's position). The user can usually move around in the VR / AR space or at least change the viewing direction, which is the so-called "viewport" of the video. In a 360-degree video environment that uses a traditional playback system (widescreen) instead of an HMD, a remote control device can be used to emulate the user's movement within the scene, and similar principles apply. Note that 360-degree content can refer to any type of content consisting of multiple simultaneous viewing angles that the user can select (e.g., by the orientation of the user's head or using a remote control device).
[0003] Compared to traditional content consumption, in VR, content creators no longer have control over what users visualize in the current viewport at various times: users are free to choose different viewports at different instances of time from the allowed or available viewports.
[0004] A common issue with VR content consumption is the risk that users miss important events in a video scene due to incorrect viewport selection. To address this issue, the concept of region of interest (ROI) has been introduced, and several concepts for ROI notification have been explored. ROIs are typically used to indicate to users the region containing the recommended viewport, but they can also be used for other purposes, such as indicating the presence of a new character / object in the scene and indicating accessibility features associated with objects in the scene—essentially, features that can be associated with elements that make up the video scene. For example, a visual message (e.g., "Turn your head left") can be used to overlay the current viewport. Alternatively, audible sounds, either natural or synthesized, can be used by playing them at the ROI location. These audio messages are known as "earcons."
[0005] In this application scenario, we use the concept of earcons to characterize audio messages conveyed to notify a ROI, but the proposed notification and processing can also be used for general audio messages for purposes other than notifying a ROI. An example of such an audio message is provided by an audio message (e.g., "To enter room X, jump over the left side of the box") to convey information / indications of various options the user has in an interactive AR / VR / MR environment. Furthermore, although we use a VR example, the mechanisms described in this document apply to any media consumption environment.
[0006] 2. Terms and Definitions The following terms are used in the art:
[0007] Audio Elements: Audio signals that can be represented, for example, as audio objects, audio channels, scene-based audio (Higher Order Ambisonics - HOA), or any combination of all.
[0008] Region of Interest (ROI): A single area of video content (or displayed or simulated environment) that is of interest to the user at a given time. This is typically, for example, a region on a sphere, or a polygonal selection from a 2D map. The ROI identifies a specific area for a specific purpose and defines the boundaries of the object under consideration.
[0009] User location information: location information (e.g., x, y, z coordinates), orientation information (yaw, pitch, roll), movement direction, movement speed, etc.
[0010] Viewport: The portion of the spherical video that is currently displayed and viewed by the user.
[0011] Viewpoint: The center point of the viewport.
[0012] 360-degree video (also known as immersive video or spherical video): In the context of this document, this refers to video content that includes multiple views (viewports) in one direction at the same time. Such content can be created, for example, using an omnidirectional camera or collection of cameras. During playback, the viewer can control the viewing direction.
[0013] An adaptation set contains a media stream or a set of media streams. In the simplest case there is one adaptation set that contains all the audio and video of the content, but to reduce bandwidth each stream can be split into different adaptation sets. A common case is to have one video adaptation set and several audio adaptation sets (one for each supported language). Adaptation sets can also contain subtitles or any metadata.
[0014] Representations allow an adaptation set to contain the same content encoded in different ways. In most cases, representations are offered at multiple bitrates, allowing clients to request the highest quality content possible without waiting for buffering. Representations can also be encoded with different codecs, allowing clients with a variety of supported codecs to be supported.
[0015] · A Media Presentation Description (MPD) is an XML syntax that contains information about media segments, their relationships, and the information needed to select them.
[0016] In this application context, the notion of adaptation set is more commonly used and may actually refer to a representation. Also, media streams (audio / video streams) are typically encapsulated into media segments, which are the actual media files that are first played by a client (e.g., a DASH client). Media segments can use a variety of formats, such as the ISO Base Media File Format (ISOBMFF), which is similar to the MPEG-4 container format, and MPEG-TS. The encapsulation into media segments and in the various representations / adaptation sets is independent of the method described here, which applies to all the various options.
[0017] Additionally, while the method description in this document may focus on DASH server and client communication, the method is general enough to work in other delivery environments such as MMT, MPEG-2 Transport Stream, DASH-ROUTE, and file formats for file playback.
[0018] 3. Current Solution The current solution is as follows:
[0019] [1].ISO / IEC 23008-3:2015,Information technology--High efficiency coding and media delivery in heterogeneous environments--Part 3:3D Audi
[0020] [2].N16950,Study of ISO / IEC DIS 23000-20 Omnidirectional Media Forma
[0021] [3].M41184, Use of Earcons for ROI Identification in 360-degree Video.
[0022] The delivery mechanism for 360-degree content is provided by ISO / IEC 23000-20, Omnidirectional Media Format[2]. This standard specifies a media format for the coding, storage, distribution, and rendering of omnidirectional images, video, and associated audio. It provides information about the media codecs used for audio and video compression, as well as additional metadata information for the correct use of 360-degree A / V content. It also specifies constraints and requirements for delivery channels, such as streaming via DASH / MMT or file-based playback.
[0023] The concept of earcons was first introduced in M41184, "Use of Earcons for ROI Identification in 360-degree Video" [3], which provides a mechanism to communicate earcon audio data to the user.
[0024] However, some users have reported disappointing results from these systems. The large number of earcons often becomes irritating. When designers reduce the number of earcons, some users lose important information. In particular, each user has their own knowledge and experience level, and therefore prefers a system that suits them. For example, each user prefers to play earcons at a preferred volume (e.g., independent of the volumes used for other audio signals). It has proven difficult for system designers to obtain a system that provides a satisfactory level for all possible users. Therefore, a solution that can increase the satisfaction of almost all users has been sought.
[0025] Furthermore, even for the designers, reconfiguring the system proved difficult, for example to prepare a new release of the audio stream or update the earcons.
[0026] Furthermore, limited systems impose certain limitations on functionality, such as the inability to pinpoint earcons to a single audio stream, and earcons must be active at all times, potentially annoying users if they play when not needed.
[0027] Furthermore, earcon spatial information cannot be signaled or modified by, for example, a DASH client. Having easy access to this information at a system level can enable additional features that improve the user experience.
[0028] Furthermore, it lacks the flexibility to accommodate different types of earcons (e.g., natural sounds, synthetic sounds, sounds generated by the DASH client, etc.).
[0029] All of these issues lead to a poor quality of experience for users, so a more flexible architecture is desirable. [Prior art documents] [Non-patent literature]
[0030] [Non-Patent Document 1] ISO / IEC 23008-3:2015, Information technology--High efficiency coding and media delivery in heterogeneous environments--Part 3:3D audio [Non-patent document 2] N16950,Study of ISO / IEC DIS 23000-20 Omnidirectional Media Format [Non-patent document 3] M41184,Use of Earcons for ROI Identification in 360-degree Video Summary of the Invention
[0031] 4. The present invention According to an example, a system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment is provided, the system comprising: receiving at least one video stream associated with the audio and video scenes; configured to receive at least one first audio stream associated with the audio and video scene to be played; The system is at least one media video decoder configured to decode at least one video signal from the at least one video stream for presentation of audio and video scenes to a user; at least one media audio decoder configured to decode at least one audio signal from the at least one first audio stream for presentation of an audio and video scene to a user; a region of interest ROI processor, the region of interest ROI processor comprising: determining whether to play an audio information message associated with the at least one ROI based on at least the user's current viewport and / or head orientation and / or movement data and / or viewport metadata and / or audio information message metadata, wherein the audio information message is independent of the at least one video signal and the at least one audio signal; If it is determined that an information message should be played, an audio information message is played.
[0032] According to an example, a system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment is provided, the system comprising: Receive at least one video stream; configured to receive at least one first audio stream; The system is at least one media video decoder configured to decode at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; at least one media audio decoder configured to decode at least one audio signal from the at least one first audio stream for representation of an audio scene to a user; a region of interest ROI processor, the region of interest ROI processor comprising: determining whether to play an audio information message associated with at least one ROI based on the user's current viewport and / or head orientation and / or movement data and / or viewport metadata and / or audio information message metadata, the audio information message being an earcon; If it is determined that an information message should be played, an audio information message is played.
[0033] The system is The device may further include a metadata processor configured to receive and / or process and / or manipulate the audio information message metadata and, when determining to play the information message, play the audio information message in accordance with the audio information message metadata.
[0034] The ROI processor receiving a user's current viewport and / or position and / or head orientation and / or movement data and / or other user-related data; receiving viewport metadata associated with at least one video signal from at least one video stream, the viewport metadata defining at least one ROI; The device may be configured to determine whether to play an audio information message associated with at least one ROI based on at least one of the user's current viewport and / or position and / or head orientation and / or movement data and viewport metadata.
[0035] The system is The audio information message may further include a metadata processor configured to receive and / or process and / or manipulate the audio information message metadata describing the audio information message and / or the audio metadata and / or viewport metadata describing the at least one audio signal encoded in the at least one audio stream to play the audio information message in accordance with the audio information message metadata and / or the audio metadata and / or viewport metadata describing the at least one audio signal encoded in the at least one audio stream.
[0036] The ROI processor If the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or movement data, in addition to playing the at least one audio signal, playing an audio information message associated with the at least one ROI; The device may be configured to disallow and / or deactivate playback of an audio information message associated with at least one ROI if the at least one ROI is within the user's current viewport and / or position and / or head orientation and / or movement data.
[0037] The system is may be further configured to receive at least one additional audio stream having at least one audio information message encoded therein; The system is The audio processing device further includes at least one muxer or multiplexer that, under control of the metadata processor and / or the ROI processor and / or another processor, merges packets of the at least one additional audio stream with packets of the at least one first audio stream in one stream and plays the audio information message in addition to the audio scene based on a decision provided by the ROI processor to play the at least one audio information message.
[0038] The system is receiving at least one audio metadata describing at least one audio signal encoded into at least one audio stream; receiving audio information message metadata associated with at least one audio information message from at least one audio stream; If it is decided to play an information message, in addition to playing the at least one audio signal, the audio information message metadata may be modified to enable the playback of the audio information message.
[0039] The system is receiving at least one audio metadata describing at least one audio signal encoded into at least one audio stream; receiving audio information message metadata associated with at least one audio information message from at least one audio stream; when it is determined to play an audio information message, modifying the audio information message metadata to enable playback of the audio information message associated with the at least one ROI in addition to playing the at least one audio signal; The method may be configured to modify audio metadata describing the at least one audio signal to enable merging of the at least one first audio stream with the at least one additional audio stream.
[0040] The system is receiving at least one audio metadata describing at least one audio signal encoded into at least one audio stream; receiving audio information message metadata associated with at least one audio information message from at least one audio stream; When it is determined to play an audio information message, the audio information message metadata may be provided to a synthesized audio generator to create a synthesized audio stream, the audio information message metadata may be associated with the synthesized audio stream, and the synthesized audio stream and the audio information message metadata may be provided to a multiplexer or muxer to enable merging of at least one audio stream with the synthesized audio stream.
[0041] The system is It may be configured to obtain audio information message metadata from at least one additional audio stream in which the audio information message is encoded.
[0042] The system is The system may include an audio information message metadata generator configured to generate audio information message metadata based on a decision to play an audio information message associated with the at least one ROI.
[0043] The system is It may be configured to store the audio information message metadata and / or the audio information message stream for future use.
[0044] The system is The audio processing system may include a synthetic audio generator configured to synthesize an audio information message based on audio information message metadata associated with the at least one ROI.
[0045] The metadata processor may be configured to control a muxer or multiplexer to merge packets of the audio information message stream with packets of at least one first audio stream in one stream to obtain the addition of an audio information message to the at least one audio stream based on the audio metadata and / or the audio information message metadata.
[0046] The audio information message metadata may be encoded into a configuration frame and / or a data frame, the data frame comprising: Identification tags, an integer that uniquely identifies the playback of the Audio Information Message metadata; The type of message, status Dependent / independent display from the scene, location data, Gain data, Indication of the presence of an associated text label, the number of languages available, the language of the audio information message, The length of the data text, the data text of the associated text label, and / or The audio information message includes at least one of the following:
[0047] The Metadata Processor and / or ROI Processor: Extracting Audio Information Message metadata from the stream; Modifying Audio Info Message metadata to activate and / or set / change the position of Audio Info Messages; Embed metadata into the stream, feeding the stream to an additional media decoder; extracting audio metadata from at least one first audio stream; Extracting audio information message metadata from the additional stream; Modifying Audio Info Message metadata to activate and / or set / change the position of Audio Info Messages; modifying audio metadata of at least one first audio stream so that the audio metadata can be merged taking into account the presence of the audio information message; It may be configured to perform at least one of the operations of: feeding the streams to a multiplexer or muxer in order to multiplex or multiplex them based on information received from the ROI processor.
[0048] The ROI processor may be configured to perform a local search for additional audio streams and / or audio information message metadata in which the audio information message is encoded, and if unable to do so, to request the additional audio streams and / or audio information message metadata from a remote entity.
[0049] The ROI processor may be configured to perform a local search for additional audio streams and / or audio information message metadata and, if unable to do so, cause a synthetic audio generator to generate the audio information message streams and / or audio information message metadata.
[0050] The system is receiving at least one additional audio stream including at least one audio information message associated with at least one ROI; The ROI processor may be configured to decode at least one additional audio stream if it determines to play an audio information message associated with at least one ROI.
[0051] The system is at least one first audio decoder for decoding at least one audio signal from at least one first audio stream; at least one additional audio decoder for decoding at least one audio information message from the additional audio stream; and at least one mixer and / or renderer for mixing and / or superimposing audio information messages from the at least one additional audio stream with at least one audio signal from the at least one first audio stream.
[0052] The system may be configured to keep track of metrics associated with historical and / or statistical data associated with the playback of audio information messages and to disable the playback of audio information messages if the metrics exceed a predetermined threshold.
[0053] The ROI processor's determination may be based on predictions of the user's current viewport and / or position and / or head orientation and / or movement data relative to the location of the ROI.
[0054] The system may be configured to receive at least one first audio stream and, upon determining to play an information message, request an audio message information stream from a remote entity.
[0055] The system may be configured to establish whether to play two audio information messages simultaneously or to select a higher priority audio information message to be played in preference to a lower priority audio information message.
[0056] The system may be configured to identify an audio information message from among a plurality of audio information messages encoded in one additional audio stream based on the address and / or position of the audio information message in the audio stream.
[0057] The audio stream may be formatted in the MPEG-H 3D audio stream format.
[0058] The system is receiving data regarding availability of a plurality of adaptation sets, the available adaptation sets including an adaptation set for at least one audio scene of at least one first audio stream and an adaptation set for at least one audio message of at least one additional audio stream including at least one audio information message, the system comprising: generating selection data identifying which of the adaptation sets to retrieve based on the determination of the ROI processor, the available adaptation sets including an adaptation set for at least one audio scene and / or an adaptation set for at least one audio message; Requesting and / or retrieving data in the adaptation set identified by the selection data; Each adaptation set may be configured to group different encodings at different bit rates.
[0059] The system may be configured to retrieve data for each of the adaptation sets, at least one of whose elements includes HTTP, DASH, dynamic adaptive streaming via a client, and / or using the ISO Base Media File Format ISO BMFF, or an MPEG-2 Transport Stream MPEG-2 TS.
[0060] The ROI processor may be configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or movement data to check whether the ROI is represented in the current viewport, and to audibly notify the user of the presence of the ROI if the ROI is outside the current viewport and / or position and / or head orientation and / or movement data.
[0061] The ROI processor may be configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or movement data to check whether the ROI is represented in the current viewport, and to suppress audio notification to the user of the presence of the ROI if the ROI is within the current viewport and / or position and / or head orientation and / or movement data.
[0062] The system may be configured to receive, from a remote entity, at least one video stream associated with a video environment scene and at least one audio stream associated with an audio scene, the audio scene being associated with the video environment scene.
[0063] The ROI processor may be configured to select, from among the plurality of audio information messages to be played, one first audio information message to be played before a second audio information message.
[0064] The system may include a cache memory for storing audio information messages received from remote entities or synthetically generated, and for reusing audio information messages at different time instances.
[0065] The audio information message may be an earcon.
[0066] The at least one video stream and / or the at least one first audio stream may be part of the current video environment scene and / or video audio scene, respectively, and may be independent of data of the user's current viewport and / or head orientation and / or movement in the current video environment scene and / or video audio scene.
[0067] The system may be configured to request at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environmental stream, respectively, and play at least one audio information message based on data of the user's current viewport and / or head orientation and / or movement.
[0068] The system may be configured to request at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environmental stream, respectively, and to request at least one audio information message from the remote entity based on data of the user's current viewport and / or head orientation and / or movement.
[0069] The system may be configured to request at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environmental stream, respectively, and synthesize at least one audio information message based on data of the user's current viewport and / or head orientation and / or movement.
[0070] The system may be configured to check at least one of additional criteria for playing the audio information message, which criteria may further include user selection and / or user settings.
[0071] The system may be configured to check at least one of additional criteria for the playing of the audio information message, the criteria further including the state of the system.
[0072] The system may be configured to check at least one of additional criteria for the playback of an audio information message, the criteria further including the number of playbacks of the audio information message already performed.
[0073] The system may be configured to check at least one of additional criteria for the playback of the audio information message, the criteria further including a flag in the data stream obtained from the remote entity.
[0074] According to one aspect, there is provided a system including a client configured as a system of any of the above and / or below examples, and a remote entity configured as a server for delivering at least one video stream and at least one audio stream.
[0075] The remote entity may be configured to search a database, an intranet, the Internet, and / or a geographic network for the at least one additional audio stream and / or audio information message metadata and, if found, deliver the at least one additional audio stream and / or audio information message metadata.
[0076] The remote entity may be configured to synthesize at least one additional audio stream and / or generate audio information message metadata.
[0077] According to one aspect, there may be provided a method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment, the method comprising: decoding at least one video signal from at least one video and audio scene to be played to a user; decoding at least one audio signal from the video and audio scenes being played; determining whether to play an audio information message associated with at least one ROI based on data and / or metadata of a user's current viewport and / or head orientation and / or movement, wherein the audio information message is independent of the at least one video signal and the at least one audio signal; If it is determined that an informational message should be played, playing an audio informational message.
[0078] According to one aspect, there may be provided a method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment, the method comprising: decoding at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; decoding at least one audio signal from the at least one first audio stream for representation of an audio scene to a user; determining whether to play an audio information message associated with at least one ROI based on data and / or metadata of the user's current viewport and / or head orientation and / or movement, the audio information message being an earcon; if it is determined that an information message should be played, playing an audio information message; Includes.
[0079] The above and / or the following methods: When it is decided to play the information message, the step may include receiving and / or processing and / or manipulating the metadata to play the audio information message in accordance with the metadata so that the audio information message is part of the audio scene.
[0080] The above and / or the following methods: playing the audio and video scenes; and determining to play further audio information messages based on the user's current viewport and / or head orientation and / or movement data and / or metadata.
[0081] The above and / or the following methods: playing the audio and video scenes; and / or if the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or movement data, in addition to playing the at least one audio signal, playing an audio information message associated with the at least one ROI; and disallowing and / or deactivating playback of an audio information message associated with at least one ROI if the at least one ROI is within the user's current viewport and / or position and / or head orientation and / or movement data.
[0082] According to an example, a system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment is provided, the system comprising: Receive at least one video stream; configured to receive at least one first audio stream; The system is at least one media video decoder configured to decode at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; at least one media audio decoder configured to decode at least one audio signal from the at least one first audio stream for representation of an audio scene to a user; a region of interest ROI processor, the region of interest ROI processor comprising: determining whether to play an audio information message associated with at least one ROI based on the user's current viewport and / or head orientation and / or movement data and / or metadata; If it is determined that an information message should be played, an audio information message is played.
[0083] In an example, a system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment is provided, the system comprising: Receive at least one video stream; configured to receive at least one first audio stream; The system is at least one media video decoder configured to decode at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; at least one media audio decoder configured to decode at least one audio signal from the at least one first audio stream for representation of an audio scene to a user; a region of interest ROI processor configured to determine whether to play an audio information message associated with at least one ROI based on a user's current viewport and / or position and / or head orientation and / or movement data and / or metadata and / or other criteria; and a metadata processor configured to receive and / or process and / or manipulate the metadata and, when determining to play an information message, play the audio information message in accordance with the metadata such that the audio information message is part of an audio scene.
[0084] According to one aspect, there is provided a non-transitory storage unit containing instructions that, when executed by a processor, cause the processor to perform the above and / or below described methods.
[0085] 5. Description of the drawings [Brief explanation of the drawings]
[0086] [Figure 1] FIG. 1 illustrates an example of an embodiment. [Figure 2] FIG. 1 illustrates an example of an embodiment. [Figure 3] FIG. 1 illustrates an example of an embodiment. [Figure 4] FIG. 1 illustrates an example of an embodiment. [Figure 5] FIG. 1 illustrates an example of an embodiment. [Figure 5a] FIG. 1 illustrates an example of an embodiment. [Figure 6] FIG. 1 illustrates an example of an embodiment. [Figure 7] FIG. 1 illustrates a method according to an example. [Figure 8] FIG. 1 illustrates an example of an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0087] 6. Example 6.1 General Examples 1 illustrates an example of a system 100 for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment. System 100 may be associated with, for example, a content consumption device (e.g., a head-mounted display, etc.), which reproduces visual data on a spherical or hemispherical display closely associated with a user's head.
[0088] The system 100 can include at least one media video decoder 102 and at least one media audio decoder 112. The system 100 can receive at least one video stream 106 in which a video signal is encoded for presenting a VR, AR, MR, or 360-degree video environment scene 118a to a user. The system 100 can receive at least one first audio stream 116 in which an audio signal is encoded for presenting an audio scene 118b to a user.
[0089] The system 100 may also include a region of interest (ROI) processor 120. The ROI processor 120 may process data associated with the ROI. Generally speaking, the presence of the ROI may be signaled in viewport metadata 131. The viewport metadata 131 may be encoded in the video stream 106 (or, in other examples, the viewport metadata 131 may be encoded in another stream). The viewport metadata 131 may include, for example, location information (e.g., coordinate information) associated with the ROI. For example, the ROI may be understood as a rectangle (identified by coordinates such as the location of one of the rectangle's four vertices in the spherical video and the length of the rectangle's sides). The ROI is typically projected onto the spherical video. The ROI is typically associated with a visible element that is considered to be of user interest (according to a particular configuration). For example, the ROI may be associated with a rectangular region displayed by the content consumption device (or otherwise visible to the user).
[0090] The ROI processor 120 may, among other things, control the operation of the media audio decoder 112 .
[0091] The ROI processor 120 may obtain data 122 associated with the user's current viewport and / or position and / or head orientation and / or movement (virtual data associated with a virtual position may also be understood as part of the data 122 in some examples). These data 122 may be provided, at least in part, by the content consumption device or by a positioning / detection unit, for example.
[0092] The ROI processor 120 can check for correspondence between the ROI and the user's current viewport and / or position (real or virtual) and / or head orientation and / or movement data 122 (e.g., other criteria may be used). For example, the ROI processor can check whether the ROI is represented in the current viewport. If the ROI is only partially represented in the viewport (e.g., based on the user's head movement), it can determine, for example, whether a minimum percentage of the ROI is displayed on the screen. In either case, the ROI processor 120 can recognize whether the ROI is not represented or is not visible to the user.
[0093] If the ROI is deemed to be outside the user's current viewport and / or position and / or head orientation and / or movement data 122, the ROI processor 120 may audibly notify the user of the presence of the ROI. For example, the ROI processor 120 may request the playing of an audio information message (ear-con) in addition to the audio signal decoded from the at least one first audio stream 116.
[0094] If the ROI is deemed to be within the user's current viewport and / or position and / or head orientation and / or movement data 122, the ROI processor may decide to avoid playing an audio information message.
[0095] The audio information messages may be encoded into audio stream 140 (audio information message stream), which may be the same as or a different stream from audio stream 116. Audio stream 140 may be generated by system 100 or obtained from an external entity (e.g., a server). Audio metadata, such as audio information message metadata 141, may be defined to describe properties of audio information stream 140.
[0096] The audio information messages may be superimposed (or mixed, multiplexed, merged, combined, composed, etc.) on the signal encoded in the audio stream 116, or may be selected, for example, solely based on a decision of the ROI processor 120. The ROI processor 120 may make its decision based on viewport and / or position and / or head orientation and / or movement data 122, metadata (such as viewport metadata 131 or other metadata), and / or other criteria (e.g., selections, system state, number of audio information message playbacks already performed, specific features and / or operations, user preference settings that may disable the use of earcons, etc.).
[0097] A metadata processor 132 may be implemented. The metadata processor 132 may be inserted, for example, between the ROI processor 120 (which may control the metadata processor 132) and the media audio decoder 112 (which may be controlled from the metadata processor). In an example, the metadata processor is part of the ROI processor 120. The metadata processor 132 may receive, generate, process, and / or manipulate audio information message metadata 141. The metadata processor 132 may also process and / or manipulate metadata for the audio stream 116, for example, to multiplex the audio stream 116 with the audio information message stream 140. Additionally or alternatively, the metadata processor 132 may receive metadata for the audio stream 116, for example, from a server (e.g., a remote entity).
[0098] Thus, the metadata processor 132 can modify the playback of audio scenes and adapt audio information messages to particular situations and / or selections and / or conditions.
[0099] Some advantages of some implementations are described here.
[0100] The audio information message may be precisely identified using, for example, the audio information message metadata 141 .
[0101] Audio information messages can be easily activated / deactivated, for example, by changing metadata (e.g., by metadata processor 132). Audio information messages can be enabled / disabled, for example, based on the current viewport and ROI information (and special features or effects to be achieved).
[0102] Audio information messages (including, for example, status, type, spatial information, etc.) can be easily signaled and modified by common devices, for example, Dynamic Adaptive Streaming over HTTP (DASH) clients.
[0103] Thus, easy system-level access to audio information messages (including status, type, spatial information, etc.) can enable additional functionality to enhance the user experience. Thus, system 100 can be easily customized, allowing for further implementations (e.g., specific applications) that can be performed by personnel independent of the designers of system 100.
[0104] Furthermore, flexibility is achieved in dealing with different types of audio information messages (eg, natural sounds, synthetic sounds, sounds generated by the DASH client, etc.).
[0105] Other advantages (as will become clear in the examples below): Using text labels in metadata (as a basis for displaying something or generating earcons) Adjusting earcon position based on device (in case of HMD you will need exact position, in case of speakers it may be better to use different position - directly on one speaker).
[0106] Different device classes: Earcon metadata can be created in a way that signals that the earcon is active.
[0107] Some devices only know how to play earcons by parsing the metadata Some newer devices with better ROI processors can decide to deactivate it when not needed Further information and additional illustrations for the adaptation set.
[0108] Therefore, in VR / AR environments, users typically visualize the entire content in 360 degrees using, for example, a head-mounted display (HMD) and listen to it through headphones. Users can typically move around in the VR / AR space or at least change their viewing direction, which is the so-called "viewport" of the video. Compared to traditional content consumption, in VR, content creators no longer control what the user visualizes at various times through the current viewport. Users are free to select different viewports at different times from the allowed or available viewports. To indicate a region of interest (ROI) to the user, audible sounds (natural or synthesized) can be used by playing them at the ROI location. These audio messages are known as "earcons." This invention proposes a solution for the efficient delivery of such messages and an optimized receiver operation to utilize earcons without affecting the user experience and content consumption. This improves the quality of the experience. This can be achieved by using dedicated metadata and metadata manipulation mechanisms at the system level to enable or disable earcons in the final scene.
[0109] The metadata processor 132 can be configured to receive and / or process and / or manipulate the metadata 141 and, upon deciding to play an information message, play an audio information message in accordance with the metadata 141. An audio signal (e.g., intended to represent a scene) can be understood to be part of an audio scene (e.g., an audio scene downloaded from a remote server). Audio signals are generally semantically meaningful to an audio scene, and all audio signals present together constitute an audio scene. Audio signals can be encoded together into one audio bitstream. Audio signals may be created by a content creator and / or associated with a particular scene and / or independent of the ROI.
[0110] Audio information messages (e.g., earcons) may be understood as having no semantic meaning to the audio scene. They may be understood as independent sounds that can be artificially generated, such as recorded sounds or the voice of a human recorder. They may also be device dependent (e.g., a system sound generated by pressing a button on a remote control). Audio information messages (e.g., earcons) may be understood as not being part of the scene, but meant to guide the user through the scene.
[0111] The audio information message may be separate from the audio signal as described above, and according to different examples, it may be included in the same bitstream, transmitted in a separate bitstream, or generated by the system 100.
[0112] An example of an audio scene made up of multiple audio signals is as follows:
[0113] -Audio Scene Concert room with 5 audio signals: ---Audio signal 1: Piano sound ---Audio signal 2: Singer's voice ---Audio signal 3: Voice of person 1, who is part of the audience ---Audio signal 4: Voice of person 2, part of the audience ---Audio signal 5: The sound produced by a wall clock The audio information message may be, for example, a recorded voice saying "Look at the piano player" (the piano is the ROI). If the user is already looking at the piano player, the audio message will not be played.
[0114] Another example: A door behind the user (e.g. a virtual door) opens and a new person enters the room, without the user looking. Earcons can be triggered based on this (information about the VR environment, such as virtual location) to notify the user that something has happened behind them.
[0115] In the example, as the user changes the environment, each scene (e.g., associated audio and video streams) is sent from the server to the client.
[0116] Audio information messages may be flexible, in particular: Audio information messages can be placed in the same audio stream associated with the scene being played.
[0117] Audio information messages can be placed in additional audio streams.
[0118] -Audio Info messages may be missing entirely, only metadata describing earcons may be present in the stream, and Audio Info messages may be generated by the system.
[0119] -It is possible that the Audio Info message and the metadata describing the Audio Info message are missing entirely, in which case the system will generate both (earcons and metadata) based on other information about the ROI in the stream.
[0120] Audio information messages are generally independent of the audio signal parts of an audio scene and are not used in the representation of the audio scene.
[0121] Examples of systems that embody or include portions of system 100 are provided below.
[0122] 6.2 Example of Figure 2 2 illustrates a system 200 (which may include at least a portion of an implementing system 100), depicted here as being subdivided into a server side 202, a media distribution side 203, a client side 204, and / or a media consumption device side 206. Each of sides 202, 203, 204, and 206 is a system in itself and can be combined with other systems to obtain another system. Audio information messages are referred to herein as earcons, even though they can be generalized to any kind of audio information message.
[0123] The client side 204 can receive at least one video stream 106 and / or at least one audio stream 116 from the server side 202 via the media delivery side 203 .
[0124] The distribution side 203 can be based on, for example, a cloud system, a network system, a geographical communication network, or a communication system such as a known media transport format (MPEG-2 TS Transport Stream, DASH, MMT, DASH ROUTE, etc.) or a file-based storage. The distribution side 203 can perform communication by delivering data packets in the form of electrical signals (e.g., via cable, wireless, etc.) and / or bitstreams in which audio and video signals are encoded (e.g., according to a specific communication protocol). However, the distribution side 203 may also be embodied by a point-to-point link, a serial or parallel connection, etc. The distribution side 203 can perform a wireless connection, for example, according to protocols such as WiFi, Bluetooth, etc.
[0125] The client side 204 may be associated with a media consumption device, such as a headset into which the user can insert their head (although other devices may be used). Thus, the user can experience video and audio scenes (e.g., VR scenes) prepared by the client side 204 based on video and audio data provided by the server side 202. However, other implementations are possible.
[0126] The server side 202 is represented here as having a media encoder 240 (which may include a video encoder, an audio encoder, a subtitle encoder, etc.). This encoder 240 may be associated with, for example, an audio and video scene to be rendered. The audio scene may, for example, be intended to reproduce an environment and associated with at least one audio and video data stream 106, 116, which may be encoded based on the position (or virtual position) reached by the user in a VR, AR, or MR environment. Typically, the video stream 106 encodes a spherical image, only a portion of which (the viewport) is displayed to the user according to its position and movement. The audio stream 116 contains audio data that participates in the audio scene representation and is intended to be heard by the user. According to an example, the audio stream 116 may include audio metadata 236 (which refers to at least one audio signal intended to participate in the audio scene representation) and / or earcon metadata 141 (which may, in some cases, describe only the earcons to be rendered).
[0127] The system 100 is depicted here as residing on the client side 204. For simplicity, the media video decoder 112 is not depicted in FIG.
[0128] To prepare the playback of earcons (or other audio information messages), earcon metadata 141 can be used. Earcon metadata 141 can be understood as metadata (which may be encoded into the audio stream) that describes and provides attributes associated with earcons. Thus, earcons (if played) can be based on the attributes of earcon metadata 141.
[0129] Advantageously, the metadata processor 132 may be specifically implemented for processing earcon metadata 141. For example, the metadata processor 132 may control the reception, processing, manipulation, and / or generation of the earcon metadata 141. Once processed, the earcon metadata is represented as modified earcon metadata 234. For example, the earcon metadata may be manipulated to obtain certain effects and / or perform audio processing operations such as multiplexing or muxing to add earcons to audio signals represented in an audio scene.
[0130] The metadata processor 132 may control the receipt, processing, and manipulation of audio metadata 236 associated with at least one stream 116. Once processed, the audio metadata 236 may be represented as modified audio metadata 238.
[0131] The modified metadata 234, 238 may be provided to the media audio decoder 112 (or multiple decoders in some examples) for playback of the audio scene 118b to the user.
[0132] In an example, a synthesized audio generator and / or storage device 246 may be provided as an optional component. The generator may synthesize an audio stream (e.g., to generate earcons that are not encoded in the stream). The storage device may store (e.g., in a cache memory) earcon streams generated by the generator and / or obtained in a received audio stream (e.g., for future use).
[0133] Thus, the ROI processor 120 may determine the representation of earcons based on the user's current viewport and / or position and / or head orientation and / or movement data 122. However, the ROI processor 120 may also make its determination based on criteria that include other aspects.
[0134] For example, the ROI processor may enable / disable the playback of earcons based on the particular application they are intended to be consumed in, for example, based on user selection, higher layer selection, or other conditions. For example, in the case of a video game application, earcons and other audio information messages may be avoided when the video game level is high. This can be easily obtained by the metadata processor by disabling earcons in the earcon metadata.
[0135] Additionally, earcons can be disabled based on system state, e.g., if an earcon is already playing, its repetition is prohibited, e.g., a timer may be used to avoid repetition too quickly.
[0136] The ROI processor 120 can also request controlled playback of a series of earcons (e.g., earcons associated with all ROIs in a scene), for example, to instruct the user about the elements they can see. The metadata processor 132 can control this operation.
[0137] The ROI processor 120 can also change the earcon position (i.e., spatial location within the scene) or earcon type. For example, some users prefer to play a specific sound as an earcon at the exact location / position of the ROI, while other users prefer to always play the earcons in one fixed position (e.g., a "voice of God" in the center or at the top) to provide an audio indication of where the ROI is located.
[0138] The gain of the playback of the earcons can be changed (e.g., to obtain a different volume). This decision may be based on, for example, a user selection. In particular, based on the decision of the ROI processor, the metadata processor 132 performs the gain change by modifying certain attributes related to gain in the earcon metadata associated with the earcons.
[0139] The original designer of the VR, AR, or MR environment may not have been aware of how the earcons would actually be rendered. For example, user selections may alter the final rendering of the earcons. Such behavior can be controlled, for example, by the metadata processor 132, which can modify the earcon metadata 141 based on decisions of the ROI processor.
[0140] Therefore, operations performed on the audio data associated with earcons are, in principle, independent of and can be managed differently from the at least one audio stream 116 used to represent the audio scene. Earcons can also be generated separately from the audio and video streams 106, 116 that make up the audio and video scene, and can even be generated by different, independent entrepreneurial groups.
[0141] This example therefore allows for increased user satisfaction. For example, users can make their own choices, for example by changing the volume of the audio information messages, disabling the audio information messages, etc. Thus, each user can get an experience that is better suited to their preferences. Furthermore, the obtained architecture is more flexible. Audio information messages can be easily updated, for example by changing the metadata independently of the audio stream and / or by changing the audio information message stream independently of the metadata and the main audio stream.
[0142] The resulting architecture is also compatible with legacy systems, e.g., legacy audio information message streams can be associated with new audio information message metadata, and in cases where no suitable audio information message stream exists, the latter can be easily synthesized (and, e.g., stored for subsequent use).
[0143] The ROI processor may keep track of metrics associated with historical and / or statistical data associated with the playback of audio information messages and disable the playback of the audio information messages if the metrics exceed a predetermined threshold (which may be used as a criterion).
[0144] The ROI processor's determination may be based on a prediction of the user's current viewport and / or position and / or head orientation and / or movement data 122 relative to the location of the ROI as a criterion.
[0145] The ROI processor may be further configured to receive the at least one first audio stream 116 and, upon determining to play an information message, request an audio message information stream from a remote entity.
[0146] The ROI processor and / or metadata generator may be further configured to determine whether to play two audio information messages simultaneously or to select a higher priority audio information message to be played in preference to a lower priority audio information message. Audio information metadata may be used to perform this determination. The priority may be obtained by the metadata processor 132, for example, based on values in the audio information message metadata.
[0147] In some examples, media encoder 240 may be configured to allow a remote entity to search a database, an intranet, the Internet, and / or a geographic network for additional audio streams and / or audio information message metadata and, if found, deliver the additional audio streams and / or audio information message metadata. For example, the search may be performed based on a client-side request.
[0148] As explained above, a solution is proposed here for efficiently delivering earcon messages together with audio content. Optimized receiver operation is obtained to utilize audio information messages (e.g., earcons) without affecting the user experience and content consumption, thereby improving the quality of experience.
[0149] This can be achieved by using dedicated metadata and metadata manipulation mechanisms at the system level to enable or disable audio information messages in the final audio scene. The metadata can be used with any audio codec and nicely complements next generation audio codec metadata (e.g. MPEG-H audio metadata).
[0150] The delivery mechanism can be different (e.g. streaming via DASH / HLS, broadcast via DASH-ROUTE / MMT / MPEG-2 TS, file playback, etc.). In this application, DASH delivery is considered, but all concepts are valid for other delivery options.
[0151] In most cases, audio information messages do not overlap in the time domain, i.e., at a given time point, only one ROI is defined. However, when considering more advanced use cases, for example interactive environments where the user can change content based on selection / movement, there may be use cases that require multiple ROIs. For this purpose, several audio information messages may be needed at once. Therefore, we describe a general solution to support all the different use cases.
[0152] The delivery and processing of audio information messages must complement existing delivery methods for next generation audio.
[0153] One way to convey multiple Audio Information Messages for multiple ROIs that are independent in the time domain is to mix all Audio Information Messages into one Audio Element (e.g., an Audio Object) with associated metadata that describes the spatial location of each Audio Information Message at different time instances. Because the Audio Information Messages do not overlap in time, they can be addressed individually with one shared Audio Element. This Audio Element may contain silence (or no audio data) between Audio Information Messages, i.e., whenever there are no Audio Information Messages. In this case, the following mechanism applies:
[0154] · Audio elements, which are common audio information messages, can be delivered in the same elementary stream (ES) as the associated audio scene, or in one auxiliary stream (dependent or independent of the main stream).
[0155] If earcon audio elements are delivered in auxiliary streams that depend on the main stream, the client can request the additional streams whenever a new ROI is present in the visual scene.
[0156] A client (e.g., system 100) can request a stream before a scene that requires earcons, for example.
[0157] The client can, for example, request streams based on the current viewport, i.e., if the current viewport matches the ROI, the client can decide not to request additional earcon streams.
[0158] If earcon audio elements are delivered in an auxiliary stream independent of the main stream, the client can still request the additional stream whenever a new ROI is present in the visual scene. Furthermore, the two (or more) streams can be processed using two media decoders and a common rendering / mixing step to mix the decoded earcon audio data into the final audio scene. Alternatively, a metadata processor can be used to modify the metadata of the two streams and a "stream merger" can be used to merge the two streams. Possible implementations of such a metadata processor and stream merger are described below.
[0159] In alternative examples, multiple earcons for several ROIs, independent in the time domain or overlapping in the time domain, can be delivered in multiple audio elements (e.g., audio objects) and embedded in one elementary stream together with the main audio scene, or in multiple auxiliary streams, e.g., each earcon in one ES or a group of earcons in one ES based on a shared property (e.g., all earcons on the left share one stream).
[0160] If all earcon audio elements are delivered in several auxiliary streams that depend on the main stream (e.g., one earcon per stream or a group of earcons per stream), the client may request, for example, one additional stream containing the earcon of interest whenever the ROI associated with that earcon is present in the visual scene.
[0161] A client can request a stream with earcons, for example, before a scene that requires them (e.g., based on user movement, the ROI processor 120 can make a decision even if the ROI is not yet part of the scene).
[0162] The client can, for example, request streams based on the current viewport, and if the current viewport matches the ROI, the client can decide not to request additional earcon streams.
[0163] If one earcon audio element (or a group of earcons) is delivered in an auxiliary stream independent of the main stream, the client can request the additional stream, for example, whenever a new ROI is present in the visual scene, as before. Furthermore, the two (or more) streams can be processed using two media decoders and a common rendering / mixing step to mix the decoded earcon audio data into the final audio scene. Alternatively, a metadata processor can be used to modify the metadata of the two streams and a "stream merger" can be used to merge the two streams. Possible implementations of such a metadata processor and stream merger are described below.
[0164] Alternatively, one common (generic) earcon can be used to signal all ROIs within an audio scene. This can be achieved by using the same audio content with different spatial information associated with the audio content at different time instances. In this case, the ROI processor 120 can collect earcons related to ROIs within a scene and request the metadata processor 132 to control the playback of the earcons in order (e.g., upon user selection or higher-layer application request).
[0165] Alternatively, one earcon can be sent only once and cached on the client, which can reuse it for all ROIs within an audio scene, using different spatial information associated with the audio content at different time instances.
[0166] Alternatively, the earcon audio content can be synthesized and generated on the client, along with a metadata generator to create the metadata needed to signal the spatial information of the earcons. For example, the earcon audio content can be compressed and fed to one media decoder along with the main audio content and new metadata, or mixed into the final audio scene after the media decoder, or multiple media decoders can be used.
[0167] Alternatively, the earcon audio content can be synthetically generated at the client (e.g., under the control of the metadata processor 132), for example, while metadata describing the earcons is already embedded in the stream. The metadata can include spatial information for the earcons, a specific unification of "decoder-generated earcons" using specific notification of the earcon type at the encoder, but cannot include the audio data of the earcons.
[0168] Alternatively, the earcon audio content can be synthesized and generated on the client, and a metadata generator can be used to create the metadata necessary to communicate the spatial information of the earcons. For example, the earcon audio content can be The main audio content is compressed together with the new metadata and fed to a single media decoder.
[0169] Or it can be mixed into the final audio scene after the media decoder.
[0170] Or multiple media decoders can be used.
[0171] 6.3 Example of metadata for audio information messages (e.g., earcons) As mentioned above, an example of audio information message (earcon) metadata 141 is presented here.
[0172] It provides a single structure for describing earcon properties and the possibility to easily adjust these values.
[0173] [Table 1] Each identifier in the table is intended to be associated with an attribute in the earcon metadata 132.
[0174] Here we explain the semantics.
[0175] numEarcons - This field specifies the number of earcon audio elements available in the stream.
[0176] Earcon_isIndependent - This flag defines if the earcon audio element is independent of any audio scene. If Earcon_isIndependent==1, the earcon audio element is independent of the audio scene. If Earcon_isIndependent==0, the earcon audio element is part of an audio scene and Earcon_id must have the same value as the mae_groupID associated with the audio element.
[0177] EarconType - This field defines the type of earcon. The following table shows the allowed values.
[0178] [Table 2] EarconActive This flag defines if the earcon is active. If EarconActive==1, the earcon audio elements will be decoded and rendered into the audio scene.
[0179] EarconPosition This flag defines if the earcon has position information available. If Earcon_isIndependent==0, this position information will be used instead of the audio object metadata specified in the dynamic_object_metadata() or intracoded_object_metadata_efficient() structures.
[0180] Earcon_azimuth Absolute value of the azimuth angle.
[0181] Earcon_elevation Absolute value of the elevation angle.
[0182] Earcon_radius Absolute value of the radius.
[0183] EarconHasGain This flag defines if the earcons have different gain values.
[0184] Earcon_gain This field defines the absolute value of the earcon gain.
[0185] EarconHasTextLabel This flag defines whether the earcon has a text label associated with it.
[0186] Earcon_numLanguages This field specifies the number of available languages for the descriptive text label.
[0187] Earcon_Language This 24-bit field identifies the language of the earcon description text. It contains the 3-letter code specified in ISO 639-2. Both ISO 639-2 / B and ISO 639-2 / T can be used. Each character is coded into 8 bits according to ISO / IEC 8859-1 and inserted sequentially into the 24-bit field. Example: French has the 3-letter code "fre", coded as "0110 0110 0111 0010 0110 0101".
[0188] Earcon_TextDataLength This field defines the length of the next group description in the bitstream.
[0189] Earcon_TextData This field contains the earcon description, a string that describes the content at a high level. The format must be UTF-8 according to ISO / IEC 10646.
[0190] A structure for identifying earcons at a system level and associating them with existing viewports. The following two tables show two ways to achieve such a structure that can be used in various implementations. aligned(8)class EarconSample()extends SphereRegionSample{ for(i=0;i <num_regions;i++){ unsigned int(7)reserved; unsigned int(1)hasEarcon; if(hasEarcon==1){ unsigned int(8)numRegionEarcons; for(n=0;n <numRegionEarcons;n++){ unsigned int(8)Earcon_id; unsigned int(32)Earcon_track_id; } } } } Or alternatively: aligned(8)class EarconSample()extends SphereRegionSample{ for(i=0;i <num_regions;i++){ unsigned int(32)Earcon_track_id; unsigned int(8)Earcon_id; } } Semantics: hasEarcon specifies whether earcon data is available for a region.
[0191] numRegionEarcons specifies the number of earcons available in a region.
[0192] Earcon_id uniquely defines the ID of one earcon element associated with a spherical region. If the earcon is part of an audio scene (i.e., if the earcon is part of one group of elements identified by one mae_groupID), Earcon_id must have the same value as mae_groupID. Earcon_id can be used to identify in audio files / tracks, e.g., in case of DASH distribution, it is specified in the EarconComponent of the MPD.
[0193] The AdaptationSet in which the tag element is included is equal to Earcon_id.
[0194] Earcon_track_id is an integer that uniquely identifies one earcon track associated with a sphere region throughout the lifetime of one presentation. That is, if the earcon tracks are delivered in the same ISO BMFF file, Earcon_track_id represents the corresponding track_id of the earcon track. If the earcons are not delivered in the same ISO BMFF file, this value should be set to zero.
[0195] To easily identify earcon tracks at the MPD level, add the following attributes / elements to EarconComponent
[0196] Can be used as a tag.
[0197] Overview of MPD elements and attributes associated with MPEG-H audio
[0198] [Table 3] For MPEG-H audio, this can be implemented using MHAS packets, for example.
[0199] · A new MHAS packet can be defined to carry information about earcons: PACTYP_EARCON carrying an EarconInfo() structure; ·New identification field in the general MHAS METADATA MHAS packet to carry the EarconInfo() structure.
[0200] With respect to metadata, the metadata processor 132 may have at least some of the following functions: Extracting Audio Information Message metadata from the stream; Modifying Audio Info message metadata to activate and / or set / change the position of an Audio Info message and / or write / change the text label of an Audio Info message; Embed metadata into the stream, feeding the stream to an additional media decoder; extracting audio metadata from at least one first audio stream (116); Extracting audio information message metadata from the additional stream; Modifying Audio Info message metadata to activate and / or set / change the position of an Audio Info message and / or write / change the text label of an Audio Info message; modifying audio metadata of at least one first audio stream (116) so that it can be merged taking into account the presence of the audio information message; The streams are fed to a multiplexer or muxer to multiplex or multiplex them based on the information received from the ROI processor.
[0201] 6.4 Example of Figure 3 FIG. 3 illustrates a system 300 that includes a system 302 (client system) on the client side 204 that may, for example, embody systems 100 or 200 .
[0202] The system 302 may include an ROI processor 120 , a metadata processor 132 , and a decoder group 313 formed by multiple decoders 112 .
[0203] In this example, the different audio streams are decoded (respectively by the media audio decoder 112) and then mixed and / or rendered together to provide the final audio scene.
[0204] Here, the at least one audio stream is represented as including two streams 116, 316 (other examples could provide one single stream, as in FIG. 2, or three or more streams). These are the audio streams for playing the audio scene that the user is expected to experience. Here, reference is made to earcons, but the concept of audio information messages can also be generalized.
[0205] Additionally, the earcon stream 140 may be provided by the media encoder 240. Based on the user's movements and the ROI indicated in the viewport metadata 131 and / or other criteria, the ROI processor plays the earcons from the earcon stream 140 (also shown as an additional audio stream because they are added to the audio stream 116, 316).
[0206] In particular, the actual representation of the earcons is based on the earcon metadata 141 and the modifications performed by the metadata processor 132.
[0207] In an example, streams can be requested by the system 302 (client) from the media encoder 240 (server) when needed. For example, the ROI processor can determine, based on user movement, that certain earcons will be needed soon and therefore request the appropriate earcon stream 140 from the media encoder 240.
[0208] The following aspects of this example can be noted.
[0209] Use case: Audio data is delivered in one or more audio streams 116, 316 (e.g. one main stream and a supplementary stream), but earcons are delivered in one or more additional streams 140 (dependent on or independent of the main audio stream).
[0210] In one implementation of the client side 204, the ROI processor 120 and the metadata processor 132 are used to efficiently process the earcon information.
[0211] The ROI processor 120 can receive information 122 about the current viewport (user orientation information) from the media consumption device 206 used for content consumption (e.g., based on an HMD). The ROI processor can also receive ROI and ROI information signaled in metadata (video viewport signaled like OMAF).
[0212] Based on this information, the ROI processor 120 can decide to activate one (or more) earcon(s) included in the earcon audio stream 140. Additionally, the ROI processor 120 can determine different locations and different gain values for the earcon(s) (e.g., for a more accurate representation of the earcon(s) in the current space where the content is consumed).
[0213] The ROI processor 120 provides this information to the metadata processor 132.
[0214] The metadata processor 132 analyzes the metadata contained in the earcon audio stream, Enable earcons (to allow them to play) and, if required by the ROI processor 120, modify the spatial position and gain information contained in the earcon metadata 141 accordingly.
[0215] Each audio stream 116, 316, 140 is decoded and rendered independently (based on user location information) and the output of all media decoders is mixed together as a final step by a mixer or renderer 314. In another implementation, only the compressed audio can be decoded and the decoded audio data and metadata can be provided to a general common renderer for final rendering of all audio elements (including earcons).
[0216] Furthermore, in a streaming environment, the ROI processor 120 can decide to request the earcon stream 140 in advance based on the same information (e.g., if the user looks in the wrong direction a few seconds before the ROI becomes active).
[0217] 6.5 Example of Figure 4 4 illustrates a system 400 including a system 402 (client system) that may embody, for example, systems 100 or 200 on the client side 204. While reference is made here to earcons, the concept can be generalized to audio information messages.
[0218] The system 402 may include an ROI processor 120, a metadata processor 132, and a stream multiplexer or muxer 412. In instances where a multiplexer or muxer 412 is present, the number of operations performed by the hardware is advantageously reduced relative to the number of operations performed when multiple decoders and one mixer or renderer are used.
[0219] In this example, different audio streams are processed based on the metadata and multiplexing in element 412 .
[0220] Here, the at least one audio stream is represented as including two streams 116, 316 (other examples could provide one single stream, as in FIG. 2, or three or more streams), which are the audio streams for playing the audio scene that the user is expected to experience.
[0221] Additionally, the earcon stream 140 may be provided by the media encoder 240. Based on the user's movements and the ROI indicated in the viewport metadata 131 and / or other criteria, the ROI processor 120 plays the earcons from the earcon stream 140 (also shown as an additional audio stream because they are added to the audio stream 116, 316).
[0222] Each audio stream 116, 316, 140 may contain metadata 236, 416, 141, respectively. At least some of this metadata is manipulated and / or processed to be provided to a stream muxer or multiplexer 412 where packets of the audio streams are merged together. Thus, earcons can be represented as part of an audio scene.
[0223] Thus, the stream muxer or multiplexer 412 can provide an audio stream 414 including the modified audio metadata 238 and the modified earcon metadata 234, which can be provided to the audio decoder 112 to be decoded and played to the user.
[0224] The following aspects of this example can be noted.
[0225] Use case: Audio data is delivered in one or more audio streams 116, 316 (e.g. one main stream 116 and one auxiliary stream 316 are provided, but a single audio stream could also be provided), while earcons are delivered in one or more additional streams 140 (dependent on or independent of the main audio stream 116).
[0226] In one implementation of the client side 204, the ROI processor 120 and the metadata processor 132 are used to efficiently process the earcon information.
[0227] The ROI processor 120 can receive information about the current viewport 122 (user orientation information) from the media consumption device (e.g., HMD) used for content consumption. The ROI processor 120 can also receive information about the ROI signaled in earcon metadata 141 (video viewports can be signaled in Omnidirectional Media Application Format, OMAF).
[0228] Based on this information, the ROI processor 120 can decide to activate one (or more) earcon(s) included in the additional audio stream 140. Additionally, the ROI processor 120 can determine different locations and different gain values for the earcon(s) (e.g., for a more accurate representation of the earcon(s) in the current space where the content is consumed).
[0229] The ROI processor 120 can provide this information to the metadata processor 132.
[0230] The metadata processor 132 analyzes the metadata contained in the earcon audio stream, Enable earcons Additionally, if requested by the ROI Processor, the spatial position and / or gain information and / or text labels contained in the earcon metadata may be modified accordingly.
[0231] The metadata processor 132 also analyzes the audio metadata 236, 416 of all audio streams 116, 316 and is able to manipulate the audio specific information so that earcons can be used as part of the audio scene (e.g. an audio scene has a 5.1 channel bed and four objects, an earcon audio element is added to the scene as a fifth object, all metadata fields are updated accordingly).
[0232] The audio data, modified audio metadata and earcon metadata of each stream 116, 316 are provided to a stream muxer or multiplexer that can generate an audio stream 414 with a set of metadata (modified audio metadata 238 and modified earcon metadata 234) based on this.
[0233] This stream 414 may be decoded by a single media audio decoder 112 based on user location information 122.
[0234] Furthermore, in a streaming environment, the ROI processor 120 can decide to request the earcon stream 140 in advance based on the same information (e.g., if the user looks in the wrong direction a few seconds before the ROI becomes active).
[0235] 6.6 Example of Figure 5 5 illustrates a system 500 including a system 502 (client system) on the client side 204 that may embody, for example, systems 100 or 200. While reference is made here to earcons, the concept can be generalized to audio information messages.
[0236] The system 502 may include an ROI processor 120 , a metadata processor 132 , and a stream multiplexer or muxer 412 .
[0237] In this example, the earcon stream is not provided by a remote entity (on the client side), but is generated by the synthesized audio generator 236 (which uses compressed / uncompressed versions of natural sounds for later reuse or saved). The earcon metadata 141 is provided by the remote entity, for example in the audio stream 316 (not the earcon stream). Thus, the synthesized audio generator 236 can be activated to create the audio stream 140 based on the attributes of the earcon metadata 141. For example, the attributes can refer to the type of synthesized sound (natural sound, synthetic sound, speech-to-text, etc.) and / or a text label (earcons can be generated by creating synthesized sounds based on the text in the metadata). In the example, after the earcon stream is created, the same is stored for future reuse. Alternatively, the synthesized sound can be a generic sound permanently stored on the device.
[0238] A stream muxer or multiplexer 412 can be used to merge packets of the audio stream 116 (and also other streams, such as the auxiliary audio stream 316) with packets of the earcon stream generated by the generator 236. An audio stream 414 can then be obtained that is associated with the modified audio metadata 238 and the modified earcon metadata 234. The audio stream 414 can be decoded by the decoder 112 and played to the user on the media consumption device side 206.
[0239] The following aspects of this example can be noted.
[0240] ·Use case: · Audio data is delivered in one or more audio streams (e.g. one main stream and one auxiliary stream).
[0241] No earcons are delivered from the remote device, but the earcon metadata 141 is delivered as part of the main audio stream (a specific notification may be used to indicate that no audio data is associated with the earcons).
[0242] In one client-side implementation, the ROI processor 120 and metadata processor 132 are used to efficiently process the earcon information.
[0243] The ROI processor 120 can receive information about the current viewport (user orientation information) from the device used on the content consumption device side 206 (e.g., HMD). The ROI processor 120 can also receive ROI and ROI notified in metadata (video viewport is notified like in OMAF).
[0244] Based on this information, the ROI processor 120 can decide to activate one (or more) earcon(s) that are not present in the stream 116. Additionally, the ROI processor 120 can determine different locations and different gain values for the earcon(s) (e.g., for a more accurate representation of the earcon(s) in the current space where the content is consumed).
[0245] The ROI processor 120 can provide this information to the metadata processor 132.
[0246] A metadata processor 120 analyzes the metadata contained in the audio stream 116, Enable earcons and, if required by the ROI processor 120, modify the spatial position and gain information contained in the earcon metadata 141 accordingly.
[0247] The metadata processor 132 also analyzes the audio metadata (e.g. 236, 417) of all audio streams (116, 316) and is able to manipulate the audio specific information so that earcons can be used as part of the audio scene (e.g. an audio scene has a 5.1 channel bed and four objects, an earcon audio element is added to the scene as a fifth object, all metadata fields are updated accordingly).
[0248] The modified earcon metadata and information from the ROI processor 120 are provided to a synthetic audio generator 246. The synthetic audio generator 246 can create synthetic sounds based on the received information (e.g., based on the spatial location of the earcons, an audio signal is generated to spell out the location). The earcon metadata 141 is also associated with the generated audio data into a new stream 414.
[0249] Similarly, as before, the audio data (116, 316) of each stream and the modified audio and earcon metadata are provided to the stream muxer, which can generate based on this one audio stream with one set of metadata (audio and earcons).
[0250] This stream 414 is decoded by a single media audio decoder 112 based on the user's location information.
[0251] Alternatively or additionally, the audio data of the earcon can be monetized by the client (e.g., from a previous use of the earcon).
[0252] Alternatively, the output of the synthesized audio generator 246 can be uncompressed audio, which can be mixed into the final rendered scene.
[0253] Furthermore, in a streaming environment, based on the same information, the ROI processor 120 can decide to request an earcon stream in advance (e.g., if the user looks in the wrong direction a few seconds before the ROI becomes active).
[0254] 6.7 Example of Figure 6 6 illustrates a system 600 including a system 602 (client system) that may embody, for example, systems 100 or 200 on the client side 204. While reference is made here to earcons, the concept can be generalized to audio information messages.
[0255] The system 602 may include a ROI processor 120 , a metadata processor 132 , and a stream multiplexer or muxer 412 .
[0256] In this example, the earcon stream is not provided by a remote entity (on the client side), but is generated by the synthesized audio generator 236 (which can store the stream for later reuse).
[0257] In this example, the earcon metadata 141 is not provided by a remote entity. The earcon metadata is generated by a metadata generator 432 that can generate earcon metadata that is used (e.g., processed, manipulated, modified) by the metadata processor 132. The earcon metadata 141 generated by the earcon metadata generator 432 may have the same structure and / or format and / or attributes as the earcon metadata described in the previous example.
[0258] The metadata processor 132 may operate as in the example of Figure 5. Based on the attributes of the earcon metadata 141, the synthesized audio generator 246 may be activated to create the audio stream 140. For example, the attributes may refer to the type of synthesized audio (natural sound, synthetic sound, speech-to-text, etc.), and / or gain, and / or activation / deactivation state, etc. In an example, after the earcon stream 140 is created, the same may be stored (e.g., cached) for future reuse. The earcon metadata generated by the earcon metadata generator 432 may also be stored (e.g., cached).
[0259] A stream muxer or multiplexer 412 can be used to merge packets of the audio stream 116 (and also other streams such as the auxiliary audio stream 316) with packets of the earcon stream generated by the generator 246. An audio stream 414 can then be obtained that is associated with the modified audio metadata 238 and the modified earcon metadata 234. The audio stream 414 can be decoded by the decoder 112 and played to the user on the media consumption device side 206.
[0260] The following aspects of this example can be noted.
[0261] ·Use case: Audio data is delivered in one or more audio streams (e.g., one main stream 116 and one auxiliary stream 316).
[0262] -Earcons are not delivered from the client side 202. No earcon metadata is delivered from the client side 202.
[0263] This use case can represent a solution to enable earcons for legacy content that was created without them.
[0264] In one client-side implementation, the ROI processor 120 and metadata processor 232 are used to efficiently process the earcon information.
[0265] The ROI processor 120 can receive information about the current viewport 122 (user orientation information) from the device used on the content consumption device side 206 (e.g., HMD). The ROI processor 210 can also receive ROI and ROIs signaled in metadata (video viewport signaled like OMAF).
[0266] Based on this information, the ROI processor 120 can decide to activate one (or more) earcon(s) that are not present in the stream (116, 316).
[0267] Additionally, the ROI processor 120 can provide information about the earcon positions and gain values to the earcon metadata generator 432.
[0268] The ROI processor 120 can provide this information to the metadata processor 232.
[0269] The metadata processor 232 analyzes the metadata contained in the earcon audio stream (if present), Enable earcons If required by the ROI processor 120, the spatial position and gain information contained in the earcon metadata can be modified accordingly.
[0270] The metadata processor also analyzes the audio metadata 236, 417 of all audio streams 116, 316 and is able to manipulate the audio specific information so that earcons can be used as part of the audio scene (e.g. an audio scene has a 5.1 channel bed and four objects, an earcon audio element is added to the scene as a fifth object, all metadata fields are updated accordingly).
[0271] The modified earcon metadata 234 and information from the ROI processor 120 are provided to a synthetic audio generator 246. The synthetic audio generator 246 can create synthetic sounds based on the received information (e.g., based on the spatial location of the earcons, an audio signal is generated to spell out the location), and the earcon metadata is associated with the generated audio data into a new stream.
[0272] Similarly, as before, the audio data and modified audio and earcon metadata for each stream are provided to a stream muxer or multiplexer 412 which can generate a set of metadata (audio and earcons) based on this one audio stream 414.
[0273] This stream 414 is decoded by a single media audio decoder based on user location information.
[0274] Alternatively, the earcon audio data can be cashed out by the client (e.g. from a previous earcon usage).
[0275] Alternatively, the output of the synthetic audio generator is uncompressed audio that can be mixed into the final rendered scene. Additionally, in a streaming environment, the ROI processor 120 can decide to request an earcon stream in advance based on the same information (e.g., if the user looks in the wrong direction a few seconds before the ROI becomes active).
[0276] 6.8 Example based on user location A feature can be implemented that allows earcons to be played only when the user is not viewing the ROI.
[0277] The ROI processor 120 may, for example, periodically check the user's current viewport and / or position and / or head orientation and / or movement data 122. If the ROI is displayed to the user, earcons will not be played.
[0278] If the ROI processor determines that the ROI is not visible to the user based on the user's current viewport and / or position and / or head orientation and / or movement data, the ROI processor 120 can request the playback of earcons. In this case, the ROI processor 120 can have the metadata processor 132 prepare the earcons for playback. The metadata processor 132 can use one of the techniques described for the above example. For example, the metadata can be obtained in a stream delivered by the server side 202 and generated by the earcon metadata generator 432. Attributes of the earcon metadata can be easily changed based on the request of the ROI processor and / or various conditions. For example, if earcons were previously disabled by the user's selection, earcons will not be played even if the user is not looking at the ROI. For example, if a (previously set) timer has not yet expired, earcons will not be played even if the user is not looking at the ROI.
[0279] Additionally, if the ROI processor determines that the ROI is visible to the user based on the user's current viewport and / or position and / or head orientation and / or movement data, the ROI processor 120 may request that earcons not be played, especially if the earcon metadata contains notification of already active earcons.
[0280] In this case, the ROI processor 120 can have the metadata processor 132 disable the playback of the earcons. The metadata processor 132 can use one of the techniques described for the examples above. For example, the metadata can be obtained in a stream delivered by the server side 202 and generated by the earcon metadata generator 432. The attributes of the earcon metadata can be easily modified based on the requests of the ROI processor and / or various conditions. If the metadata already contains an indication that the earcons should be played, in this case the metadata is modified to indicate that the earcons are inactive and cannot be played.
[0281] The following aspects of this example can be noted.
[0282] ·Use case: · Audio data is delivered in one or more audio streams 116, 316 (e.g. one main stream and a supplementary stream), while earcons are delivered either in the same one or more audio streams 116, 316 or in one or more additional streams 140 (dependent on or independent of the main audio stream).
[0283] The earcon metadata is set to indicate that the earcon will always be active at a specific moment.
[0284] First generation devices that do not include an ROI processor will read the earcon metadata and play the earcons regardless of the fact that the user's current viewport and / or position and / or head orientation and / or movement data indicates that the ROI is visible to the user.
[0285] New generation devices that include the ROI processor described in either system utilize the ROI processor's determination. If the ROI processor determines that the ROI is visible to the user based on the user's current viewport and / or position and / or head orientation and / or movement data, the ROI processor 120 can request that earcons not be played, especially if the earcon metadata contains a notification of an already active earcon. In this case, the ROI processor 120 can have the metadata processor 132 disable earcon playback. The metadata processor 132 can use one of the techniques described for the example above. For example, the metadata can be obtained in a stream delivered by the server side 202 or generated by the earcon metadata generator 432. Attributes of the earcon metadata can be easily modified based on the ROI processor's request and / or various conditions. If the metadata already contains an indication that an earcon should be played, the metadata is modified to indicate that the earcon is inactive and cannot be played.
[0286] Additionally, depending on the playback device, the ROI processor may require earcon metadata to be modified. For example, the spatial information of earcons may be modified differently if the sound is played through headphones or speakers.
[0287] Thus, the final audio scene experienced by the user is obtained based on the metadata modifications performed by the metadata processor.
[0288] 6.9 Example based on server-client communication (Figure 5a) 5a shows a system 550 including, on the client side 204, a system 552 (client system) that may embody, for example, systems 100 or 200 or 300 or 400 or 500. Here, reference is made to earcons, but the concept can also be generalized to audio information messages.
[0289] The system 552 may include an ROI processor 120, a metadata processor 132, and a stream multiplexer or muxer 412. (In an example, different audio streams are decoded (respectively by the media audio decoder 112) and then mixed and / or rendered together to provide a final audio scene.)
[0290] Here, the at least one audio stream is represented as including two streams 116, 316 (other examples could provide one single stream, as in FIG. 2, or three or more streams), which are the audio streams for playing the audio scene that the user is expected to experience.
[0291] Additionally, the earcon stream 140 may be provided by the media encoder 240 .
[0292] Audio streams can be encoded at different bitrates, allowing for efficient bitrate adaptation depending on the network connection (i.e., a higher bitrate coded version is delivered to users with fast connections, and a lower bitrate version is delivered to users with slower network connections).
[0293] The audio streams may be stored on a media server 554, and for each audio stream, different encodings at different bit rates are grouped into one adaptation set 556 along with appropriate data signaling the availability of all created adaptation sets. Audio adaptation sets 556 and video adaptation sets 557 may be provided.
[0294] Based on the user's movements and the ROI indicated in the viewport metadata 131 and / or other criteria, the ROI processor 120 plays earcons from the earcon stream 140 (also shown as an additional audio stream because it is added to the audio stream 116, 316).
[0295] In this example: The client 552 is configured to receive data about the availability of all adaptation sets from the server.
[0296] At least one audio scene adaptation set for at least one audio stream; and At least one audio message adaptation set for at least one additional audio stream containing at least one audio information message As with other exemplary implementations, the ROI processor 120 can receive information 122 about the current viewport (user orientation information) from the media consumption device 206 used for content consumption (e.g., based on an HMD). The ROI processor 120 can also receive ROI and ROI information signaled in metadata (video viewport signaled, such as in OMAF).
[0297] Based on this information, the ROI processor 120 can decide to activate one (or more) earcon(s) included in the earcon audio stream 140.
[0298] Additionally, the ROI processor 120 may determine different locations and different gain values for the earcons (e.g., for a more accurate representation of the earcons in the current space where the content is consumed).
[0299] The ROI processor 120 can provide this information to the selection data generator 558.
[0300] The selection data generator 558 may be configured to generate selection data 559 specifying which adaptation sets to receive based on the determination of the ROI processor. The adaptation sets include an audio scene adaptation set and an audio message adaptation set.
[0301] The media server 554 may be configured to provide instruction data to the client 552 to cause the streaming client to retrieve data in the adaptation sets 556, 557 identified by selection data specifying which adaptation sets to receive. The adaptation sets include audio scene adaptation sets and audio message adaptation sets.
[0302] The downloading and switching module 560 is configured to receive the requested audio streams from the media server 554 based on selection data that identifies which adaptation sets to receive. The adaptation sets include audio scene adaptation sets and audio message adaptation sets. The downloading and switching module 560 may be further configured to provide the audio metadata and earcon metadata 141 to the metadata processor 132.
[0303] The ROI processor 120 can provide this information to the metadata processor 132.
[0304] The metadata processor 132 analyzes the metadata contained in the earcon audio stream 140, Enable earcons (to allow them to play) and, if required by the ROI processor 120, modify the spatial position and gain information contained in the earcon metadata 141 accordingly.
[0305] The metadata processor 132 also analyzes the audio metadata of all audio streams 116, 316 and can manipulate the audio specific information so that earcons can be used as part of the audio scene (e.g. an audio scene has a 5.1 channel bed and four objects, an earcon audio element is added to the scene as a fifth object, all metadata fields may be updated accordingly).
[0306] The audio data, modified audio metadata and earcon metadata of each stream 116, 316 may be provided to a stream muxer or multiplexer that can generate a single audio stream 414 with a set of metadata (modified audio metadata 238 and modified earcon metadata 234) based on this.
[0307] This stream may be decoded by a single media audio decoder 112 based on user location information 122.
[0308] An adaptation set may be formed by a set of representations containing interchangeable versions of the respective content, e.g., different audio bitrates (e.g., different streams at different bitrates). While theoretically one representation is sufficient to provide a playable stream, using multiple representations allows the client to adapt the media stream to the current network conditions and bandwidth requirements, ensuring smooth playback.
[0309] 6.10 Method All the above examples can be implemented by method steps. Here, the method 700 (which can be performed by any of the above examples) will be fully described. The method includes:
[0310] In step 702, at least one video stream (106) and at least one primary audio stream (116, 316) are received.
[0311] At step 704, at least one video signal from at least one video stream (106) is decoded to present a VR, AR, MR, or 360-degree video environment scene (118a) to a user.
[0312] In step 706, decoding at least one audio signal from at least one first audio stream (116, 316) for presentation of an audio scene (118b) to a user; Receive data (122) about the user's current viewport and / or position and / or head orientation and / or movement.
[0313] Step 708 receives viewport metadata (131) associated with at least one video signal from at least one video stream (106), the viewport metadata defining at least one ROI.
[0314] In step 710, a determination is made as to whether to play an audio information message associated with at least one ROI based on the user's current viewport and / or position and / or head orientation and / or movement data (122) and viewport metadata and / or other criteria.
[0315] In step 712, audio information message metadata (141) describing the audio information message is received, processed, and / or manipulated to play the audio information message according to the audio information message attributes in a manner such that the audio information message is part of an audio scene.
[0316] In particular, the sequence may also be different: for example, the receiving steps 702, 706, 708 may have a different order according to the actual order in which the information is delivered.
[0317] Line 714 refers to the fact that the method may be repeated. In the event of the ROI processor's decision not to play an audio information message, step 712 is skipped.
[0318] 6.11 Other Implementations 8 shows a system 800 that may implement one of the systems (or components thereof) or perform method 700. System 800 may include a processor 802 and a non-transitory memory unit 806 that stores instructions that, when executed by processor 802, may cause the processor to perform at least the above-described stream processing operations and / or the above-described metadata processing operations. System 800 may include an input / output unit 804 for connection with external devices.
[0319] The system 800 may implement at least some (or all) of the functionality of the ROI processor 120, the metadata processor 232, the generator 246, the muxer or multiplexer 412, the decoder 112m, the earcon metadata generator 432, and the like.
[0320] Depending on the particular implementation, the embodiments may be implemented in hardware. The embodiments may be implemented using a digital storage medium, such as a floppy disk, a digital versatile disk (DVD), a Blu-ray disk, a compact disk (CD), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory, on which electronically readable control signals are stored that cooperate (or can cooperate) with a programmable computer system to perform the respective methods. Thus, the digital storage medium may be computer-readable.
[0321] In general, embodiments may be implemented as a computer program product including program instructions that operate to perform one of the methods when the computer program product is run on a computer. The program instructions may be stored on, for example, a machine-readable medium.
[0322] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine-readable carrier. In other words, an example of a method is therefore a computer program having program instructions for performing one of the methods described herein, when the computer program runs on a computer.
[0323] Thus, a further example of the method is a data carrier medium (or digital storage medium, or computer readable medium) comprising and having recorded thereon a computer program for performing one of the methods described herein. The data carrier medium, digital storage medium, or recorded medium is tangible and / or non-transitory, rather than an intangible, transitory signal.
[0324] Further examples include a processing unit, such as a computer, or a programmable logic device, that performs one of the methods described herein.
[0325] A further example comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0326] Further examples include an apparatus or system that transfers (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0327] In some examples, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some examples, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods may be performed by any suitable hardware apparatus.
[0328] Further examples include: [1] 1. A system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment, the system comprising: receiving at least one video stream (106) associated with the audio and video scenes to be played; configured to receive at least one first audio stream (116, 316) associated with the audio and video scene to be played; The system comprises: at least one media video decoder (102) configured to decode at least one video signal from said at least one video stream (106) for presentation of said audio and video scenes to a user; at least one media audio decoder (112) configured to decode at least one audio signal from said at least one first audio stream (116, 316) for presentation of said audio and video scene to said user; a region of interest ROI processor (120), wherein the region of interest ROI processor (120) determining whether to play an audio information message associated with the at least one ROI based on at least the user's current viewport and / or head orientation and / or movement data (122) and / or viewport metadata (131) and / or audio information message metadata (141), wherein the audio information message is independent of the at least one video signal and the at least one audio signal; if it is determined that the information message should be played, playing the audio information message; It is a system that is configured as follows. Further examples include: [2] 1. A system for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment, the system comprising: receiving at least one video stream (106); configured to receive at least one first audio stream (116, 316); The system comprises: at least one media video decoder (102) configured to decode at least one video signal from the at least one video stream (106) to present a VR, AR, MR, or 360-degree video environment scene (118a) to a user; at least one media audio decoder (112) configured to decode at least one audio signal from the at least one first audio stream (116, 316) for presentation of an audio scene (118b) to the user; a region of interest ROI processor (120), wherein the region of interest ROI processor (120) determining whether to play an audio information message associated with the at least one ROI based on the user's current viewport and / or head orientation and / or movement data (122) and / or viewport metadata (131) and / or audio information message metadata (141), wherein the audio information message is an earcon; if it is determined that the information message should be played, playing the audio information message; It is a system that is configured as follows. Further examples include: [3] The system according to any one of claims 1 to 2, further comprising a metadata processor (132) configured to receive and / or process and / or manipulate audio information message metadata (141) and, when it is determined to play the information message, play the audio information message in accordance with the audio information message metadata (141). Further examples include: [4] The ROI processor (120) receiving a user's current viewport and / or position and / or head orientation and / or movement data and / or other user-related data (122); receiving viewport metadata (131) associated with at least one video signal from the at least one video stream (106), the viewport metadata (131) defining at least one ROI; determining whether to play an audio information message associated with the at least one ROI based on at least one of the user's current viewport and / or position and / or head orientation and / or movement data (122) and viewport metadata; The system according to any one of [1] to [3] above is configured as follows. Further examples include: [5] The system according to any one of [1] to [4], further comprising a metadata processor (132) configured to receive and / or process and / or manipulate audio information message metadata (141) describing the audio information message and / or audio metadata (236) describing at least one audio signal encoded in at least one audio stream (116) and / or viewport metadata (131) to play the audio information message in accordance with the audio information message metadata (141) and / or the audio metadata (236) describing the at least one audio signal encoded in at least one audio stream (116) and / or the viewport metadata (131). Further examples include: [6] The ROI processor (120) If the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or movement data (122), in addition to playing the at least one audio signal, play an audio information message associated with the at least one ROI; disallowing and / or deactivating the playback of the audio information message associated with the at least one ROI if the at least one ROI is within the user's current viewport and / or position and / or head orientation and / or movement data (122); The system according to any one of [1] to [5] above is configured as follows. Further examples include: [7] further configured to receive the at least one additional audio stream (140) in which the at least one audio information message is encoded; The system comprises: The system of any one of [1] to [6] further comprises at least one muxer or multiplexer (412) that, under the control of the metadata processor (132) and / or the ROI processor (120) and / or another processor, merges packets of the at least one additional audio stream (140) with packets of the at least one first audio stream (116, 316) in one stream (414) and plays the audio information message in addition to the audio scene based on the decision to play the at least one audio information message provided by the ROI processor (120). Further examples include: [8] receiving at least one audio metadata (236) describing the at least one audio signal encoded into the at least one audio stream (116); receiving audio information message metadata (141) associated with at least one audio information message from at least one audio stream (116); when it is decided to play the information message, in addition to playing the at least one audio signal, modifying the audio information message metadata (141) to enable the playback of the audio information message. The system according to any one of [1] to [7] above, further configured as follows: Further examples include: [9] receiving at least one audio metadata (141) describing the at least one audio signal encoded into the at least one audio stream (116); receiving audio information message metadata (141) associated with at least one audio information message from said at least one audio stream (116); when it is determined to play the audio information message, modifying the audio information message metadata (141) to enable playback of an audio information message associated with the at least one ROI in addition to playing the at least one audio signal; modifying the audio metadata (236) describing the at least one audio signal to enable merging of the at least one first audio stream (116) with the at least one additional audio stream (140); The system according to any one of [1] to [8] above, further configured as follows: Further examples include:
[10] receiving at least one audio metadata (236) describing the at least one audio signal encoded into the at least one audio stream (116); receiving audio information message metadata (141) associated with at least one audio information message from at least one audio stream (116); when it is determined to play the audio information message, providing the audio information message metadata (141) to a synthesized audio generator (246) to create a synthesized audio stream (140), associating the audio information message metadata (141) with the synthesized audio stream (140), and providing the synthesized audio stream (140) and the audio information message metadata (141) to a multiplexer or muxer (412) to enable merging of the at least one audio stream (116) with the synthesized audio stream (140); The system according to any one of [1] to [9] above, further configured as follows: Further examples include:
[11] The system of any one of [1] to
[10] , further configured to obtain the audio information message metadata (141) from the at least one additional audio stream (140) in which the audio information message is encoded. Further examples include:
[12] The system of any one of [1] to
[11] further includes an audio information message metadata generator (432) configured to generate audio information message metadata (141) based on the decision to play an audio information message associated with the at least one ROI. Further examples include:
[13] The system according to any one of [1] to
[12] , further configured to store the audio information message metadata (141) and / or the audio information message stream (140) for future use. Further examples include:
[14] a synthesized audio generator (432) configured to synthesize an audio information message based on audio information message metadata (141) associated with the at least one ROI; The system according to any one of [1] to
[13] above. Further examples include:
[15] The system according to any one of [1] to
[14] , wherein the metadata processor (132) is configured to control a muxer or multiplexer (412) to merge packets of the audio information message stream (140) with packets of the at least one first audio stream (116) in one stream (414) based on the audio metadata and / or audio information message metadata to obtain the addition of the audio information message to the at least one audio stream (116). Further examples include:
[16] The audio information message metadata (141) is encoded into configuration frames and / or data frames, the data frames comprising: Identification tags, an integer that uniquely identifies said playback of audio information message metadata; the type of the message; status Dependent / independent display from said scene, location data, Gain data, Indication of the presence of an associated text label, the number of languages available, the language of said audio information message; The length of the data text, the data text of said associated text label; and / or The system according to any one of [1] to
[15] , including at least one of the descriptions of the audio information message. Further examples include:
[17] The metadata processor (132) and / or the ROI processor (120) Extracting Audio Information Message metadata from the stream; Modifying audio information message metadata to activate and / or set / change the position of said audio information message; Embed metadata into the stream, providing said stream to an additional media decoder; extracting audio metadata from the at least one first audio stream (116); Extracting audio information message metadata from the additional stream; Modifying audio information message metadata to activate and / or set / change the position of said audio information message; modifying audio metadata of the at least one first audio stream (116) so that it can be merged taking into account the presence of the audio information message; The system of any one of [1] to
[16] , configured to perform at least one of the operations of: supplying streams to the multiplexer or muxer to multiplex or multiplex them based on the information received from the ROI processor. Further examples include:
[18] The system of any one of [1] to
[17] , wherein the ROI processor (120) is configured to perform a local search for additional audio streams (140) in which the audio information message is encoded and / or audio information message metadata, and if unable to do so, to request the additional audio streams (140) and / or audio information message metadata from a remote entity. Further examples include:
[19] The system of any one of [1] to
[18] , wherein the ROI processor (120) is configured to perform a local search for an additional audio stream (140) and / or audio information message metadata, and if unable to do so, to cause a synthetic audio generator (432) to generate the audio information message stream and / or audio information message metadata. Further examples include:
[20] receiving the at least one additional audio stream (140) including at least one audio information message associated with the at least one ROI; decoding the at least one additional audio stream (140) if the ROI processor determines to play an audio information message associated with the at least one ROI; The system according to any one of [1] to
[19] above, further configured as follows: Further examples include: 〔twenty one〕 at least one first audio decoder (112) for decoding said at least one audio signal from at least one first audio stream (116); at least one additional audio decoder (112) for decoding said at least one audio information message from an additional audio stream (140); at least one mixer and / or renderer (314) for mixing and / or superimposing the audio information messages from the at least one additional audio stream (140) with the at least one audio signal from the at least one first audio stream (116); The system according to
[20] , further comprising: Further examples include: 〔twenty two〕 The system of any one of [1] to
[21] , further configured to keep track of metrics associated with historical and / or statistical data associated with the playback of the audio information message, and to disable playback of the audio information message if the metrics exceed a predetermined threshold. Further examples include: 〔twenty three〕 The system of any one of [1] to
[22] , wherein the ROI processor's determination is based on a prediction of the user's current viewport and / or position and / or head orientation and / or movement data (122) relative to the position of the ROI. Further examples include: 〔twenty four〕 The system described in any one of [1] to
[23] is further configured to, upon receiving the at least one first audio stream (116) and determining to play the information message, request an audio message information stream from a remote entity. Further examples include: 〔twenty five〕 The system of any one of [1] to
[24] , further configured to establish whether to play two audio information messages simultaneously or to select a higher priority audio information message to be played in preference to a lower priority audio information message. Further examples include:
[26] The system of any one of [1] to
[25] , further configured to identify an audio information message from among a plurality of audio information messages encoded in one additional audio stream (140) based on the address and / or position of the audio information message in the audio stream. Further examples include:
[27] The system according to any one of [1] to
[26] , wherein the audio stream is formatted in the MPEG-H 3D audio stream format. Further examples include:
[28] receiving data on the availability of a plurality of adaptation sets (556, 557), the available adaptation sets including an adaptation set for at least one audio scene of the at least one first audio stream (116, 316) and an adaptation set for at least one audio message of the at least one additional audio stream (140), the adaptation set including at least one audio information message; generating selection data (559) specifying which of the adaptation sets to retrieve based on the determination of the ROI processor, the available adaptation sets including an adaptation set for at least one audio scene and / or an adaptation set for at least one audio message; requesting and / or retrieving the data in the adaptation set identified by the selection data; Each adaptation set groups different encodings at different bitrates. The system according to any one of [1] to
[27] , further configured as follows: Further examples include:
[29] The system of
[28] , wherein at least one of the elements includes HTTP, DASH, dynamic adaptive streaming via a client, and / or is configured to retrieve the data for each of the adaptation sets using the ISO Base Media File Format ISO BMFF, or an MPEG-2 Transport Stream MPEG-2 TS. Further examples include:
[30] The system of any one of [1] to
[29] , wherein the ROI processor (120) is configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or movement data (122) to check whether the ROI is represented in the current viewport, and to notify the user of the presence of the ROI by audio if the ROI is outside the current viewport and / or position and / or head orientation and / or movement data (122). Further examples include:
[31] The system of any one of [1] to
[30] , wherein the ROI processor (120) is configured to check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or movement data (122) to check whether the ROI is represented in the current viewport, and if the ROI is within the current viewport and / or position and / or head orientation and / or movement data (122), to suppress audio notification of the presence of the ROI to the user. Further examples include:
[32] The system described in any one of [1] to
[31] is configured to receive from a remote entity (202) at least one video stream (116) associated with the video environment scene and at least one audio stream (106) associated with the audio scene, wherein the audio scene is associated with the video environment scene. Further examples include:
[33] The system of any one of [1] to
[32] , wherein the ROI processor (120) is configured to select, from among a plurality of audio information messages to be played, one first audio information message to be played before a second audio information message. Further examples include:
[34] The system according to any one of [1] to
[33] , further comprising a cache memory (246) for storing audio information messages received from a remote entity (204) or synthetically generated, and for reusing the audio information messages at different time instances. Further examples include:
[35] The system according to any one of [1] and [3] to
[34] , wherein the audio information message is an earcon. Further examples include:
[36] The system described in any one of [1] to
[35] , wherein the at least one video stream and / or the at least one first audio stream are part of the current video environment scene and / or video audio scene, respectively, and are independent of the user's current viewport and / or head orientation and / or movement data (122) in the current video environment scene and / or video audio scene. Further examples include:
[37] The system of any one of [1] to
[36] 36 is configured to request the at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environmental stream, respectively, and play the at least one audio information message based on the user's current viewport and / or head orientation and / or movement data (122). Further examples include:
[38] The system according to any one of [1] to
[37] , configured to request the at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environmental stream, respectively, and to request the at least one audio information message from the remote entity based on data (122) of the user's current viewport and / or head orientation and / or movement. Further examples include:
[39] The system according to any one of [1] to
[38] , configured to request the at least one first audio stream and / or at least one video stream from a remote entity associated with the audio stream and / or video environmental stream, respectively, and synthesize the at least one audio information message based on data (122) of the user's current viewport and / or head orientation and / or movement. Further examples include:
[40] The system of any one of [1] to
[39] , configured to check at least one additional criterion for the playback of the audio information message, the criterion further including user selection and / or user settings. Further examples include:
[41] The system of any one of [1] to
[40] , configured to check at least one additional criterion for the playback of the audio information message, the criterion further including the state of the system. Further examples include:
[42] The system of any one of [1] to
[41] is configured to check at least one of additional criteria for the playback of the audio information message, the criteria further including the number of playbacks of the audio information message already performed. Further examples include:
[43] The system of any one of [1] to
[42] is configured to check at least one of additional criteria for the playback of the audio information message, the criteria further including a flag in a data stream obtained from a remote entity. Further examples include:
[44] A system including a client configured as a system described in any one of [1] to
[43] , and a remote entity (202, 240) configured as a server for delivering the at least one video stream (106) and the at least one audio stream (116). Further examples include:
[45] The system of
[44] , wherein the remote entity (202, 240) is configured to search a database, an intranet, the Internet, and / or a geographical network for the at least one additional audio stream (140) and / or audio information message metadata, and, if found, deliver the at least one additional audio stream (140) and / or audio information message metadata. Further examples include:
[46] The system according to
[45] , wherein the remote entity (202, 240) is configured to synthesize the at least one additional audio stream (140) and / or generate the audio information message metadata. Further examples include:
[47] 1. A method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment, comprising: decoding at least one video signal from said at least one video and audio scene to be played to a user; decoding at least one audio signal from said video and audio scenes being played; determining whether to play an audio information message associated with the at least one ROI based on the user's current viewport and / or head orientation and / or movement data (122) and / or metadata, wherein the audio information message is independent of the at least one video signal and the at least one audio signal; if it is determined that the information message should be played, playing the audio information message; The method includes: Further examples include:
[48] 1. A method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment, comprising: decoding at least one video signal from the at least one video stream (106) to present a VR, AR, MR, or 360-degree video environment scene (118a) to a user; decoding at least one audio signal from said at least one first audio stream (116, 316) for presentation of an audio scene (118b) to said user; determining whether to play an audio information message associated with the at least one ROI based on the user's current viewport and / or head orientation and / or movement data (122) and / or metadata, wherein the audio information message is an earcon; if it is determined that the information message should be played, playing the audio information message; The method includes: Further examples include:
[49] The method of any one of
[47] to
[48] , further comprising the step of receiving and / or processing and / or manipulating the metadata (141) to play the audio information message in accordance with the metadata (141) when it is decided to play the information message, so that the audio information message is part of the audio scene. Further examples include:
[50] playing the audio and video scenes; determining to further play said audio information message based on said user's current viewport and / or head orientation and / or movement data (122) and / or metadata; The method according to any one of
[47] to
[49] , further comprising: Further examples include:
[51] playing the audio and video scenes; If the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or movement data (122), in addition to playing the at least one audio signal, play an audio information message associated with the at least one ROI; and / or disallowing and / or deactivating the playback of the audio information message associated with said at least one ROI if said at least one ROI is within said user's current viewport and / or position and / or head orientation and / or movement data (122); The method according to any one of
[47] to
[50] , further comprising: Further examples include:
[52] A non-transitory storage unit containing instructions that, when executed by a processor, cause the processor to perform the method of any one of
[47] to
[51] . The above examples are illustrative of the principles described above. It is understood that modifications and variations of the arrangements and details described herein will be apparent. It is therefore the intention to be limited by the scope of the appended claims and not by the specific details presented by way of description and illustration of the embodiments herein.
Claims
1. 1. A system comprising: configured to receive at least one first audio stream (116, 316) associated with the audio scene to be played; at least one audio decoder (112) configured to decode at least one audio signal from said at least one first audio stream (116, 316) for presentation of said audio scene to a user; at least one processor (120, 132); the processor (120, 132) determines whether to play uncompressed earcons based on at least the user's movement data (122) and / or earcon metadata (141) describing properties of earcons and / or a user selection; When it is determined to play the earcons, the earcons are played; the system further comprises a mixer (314) for mixing the at least one audio signal decoded from the at least one first audio stream (116, 316) with the earcons; system.
2. The system of claim 1 , wherein the earcons are obtained from an external entity.
3. The at least one processor (120, 132) modifying the earcon metadata (141) to activate the earcon; Embed metadata into the stream, extracting audio metadata from the at least one first audio stream (116); Extracting the earcon metadata (141) from the audio stream; 2. The system of claim 1, further configured to perform at least one of the following operations: modifying audio metadata of the at least one first audio stream (116) so that mixing can take into account the presence of the earcons.
4. The system of claim 1 , wherein the earcons are associated with accessibility features associated with objects in the audio scene.
5. The system of claim 1 , further configured to store the earcon metadata (141) for future use.
6. The system of claim 1 , further configured to store the earcons (140) for future use.
7. The system described in claim 1, further comprising a synthetic audio generator (432) configured to generate the earcons based on the earcon metadata (141).
8. The system of claim 1 , wherein the earcon metadata (141) is written in a configuration frame or in a data frame that includes at least gain data.
9. The system of claim 1 , wherein the earcons are encoded and / or configured to perform a local search of the earcon metadata (141).
10. 2. The system of claim 1, configured to perform a local search for the earcons and / or the earcon metadata and, if unavailable, have a synthetic audio generator generate the earcon metadata.
11. The system of claim 1 , configured to perform a local search for the earcon metadata (141) and, if not available, to have a metadata generator generate the earcon metadata (141).
12. 2. The system of claim 1, further configured to keep track of metrics related to historical and / or statistical data related to the playback of the earcons in order to modify the configuration of the earcons or the earcon metadata (141).
13. The system of claim 1 , further configured to establish whether to play two earcons simultaneously or to select a higher priority earcon to be played preferentially over a lower priority earcon.
14. The system of claim 1 , wherein the audio stream is formatted in an MPEG-H 3D audio stream format.
15. The system of claim 1, including dynamic adaptive streaming via HTTP, DASH, a client, and / or configured to retrieve the data for each of the adaptation sets using the ISO Base Media File Format ISO BMFF, or an MPEG-2 Transport Stream MPEG-2 TS.
16. Generating the earcons based on the earcon metadata (141). The system of claim 1 configured to:
17. The system of claim 1 , further comprising a media viewing device (206) for playing the decoded audio signal.
Citation Information
Patent Citations
IEC23008-3
JP1974016547A
Transmission device, transmission method, reception device, and reception method
JP2016201643A
Connectionless transport for user input control for wireless display devices
JP2016511965A
Adaptive data streaming method by push message control
JP2016531466A