Method and apparatus for efficient delivery and use of high quality of experience audio messages
Patent Information
- Application Number
- CN202311468058.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-10-12
- Filing Date
- 2018-10-10
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2038-10-10
Smart Images

Figure CN117692673B_ABST
Abstract
Description
[0001] This application is a divisional application of the application filed on October 10, 2018, with international application number PCT / EP2018 / 077556, Chinese application number “201880080159.5”, and entitled “Method and apparatus for efficient delivery and use of high-quality audio messages”. Technical Field
[0002] This invention relates to scene reproduction, and more specifically, to an audio and / or video reproduction method and apparatus that can improve the quality of user experience. Background Technology
[0003] 1. Introduction
[0004] In many applications, the delivery of audible messages can improve the user experience during media consumption. Virtual reality (VR) content provides one of the most relevant applications for such messages. In VR environments, or similarly in augmented reality (AR) or mixed reality (MR) or 360-degree video environments, users can typically use, for example, a head-mounted display (HMD) to visualize full 360-degree content and listen to it through headphones (or similarly through speakers, including proper rendering depending on their position). Users can typically move within the VR / AR space, or at least change their viewing orientation—the so-called “viewport” of the video. In 360-degree video environments using classic reproduction systems (wide displays) instead of HMDs, remote control devices can be used to simulate user movement within the scene, and similar principles apply. It should be noted that 360-degree content can refer to any type of content from which the user can select (e.g., by the user’s head orientation or using a remote control device), which includes more than one perspective at a time.
[0005] Compared to traditional content consumption, with VR, content creators no longer have control over what the user sees at each moment—the current viewport. Users can freely choose different viewports from those that are allowed or available at each time instance.
[0006] A common risk in VR content consumption is that users may miss important events in a video scene due to incorrect viewport selection. To address this, the concept of Region of Interest (ROI) has been introduced, and several concepts for signaling ROIs have been considered. While ROIs are typically used to indicate to the user the area containing the recommended viewport, they can also be used for other purposes, such as: indicating the presence of new characters / objects in the scene; indicating accessibility features associated with objects in the scene; and essentially any feature that can be associated with the elements that make up the video scene. For example, a visual message (e.g., "Turn your head to the left") can be used and overlaid on the current viewport. Alternatively, audible sounds (natural or synthetic) can be used at the location of the ROI. These audio messages are called "Earcons".
[0007] In the context of this application, the concept of "Earcon" will be used to characterize audio messages conveyed to signal an ROI; however, the proposed signaling and processing can also be used for generic audio messages not intended to signal an ROI. An example of such an audio message is given by an audio message used to convey information / instructions about various options available to the user / user in an interactive AR / VR / MR environment (e.g., "Skip the box on your left to enter room X"). Furthermore, a VR example will be used, but the mechanisms described in this document are applicable to any media consumption environment.
[0008] 2. Terms and Definitions
[0009] The following terms are used in the technical field:
[0010] • Element: can be represented as, for example, an audio object, an audio channel, scene-based audio (higher-order Ambisonics (HOA)), or an audio signal of all these combinations.
[0011] • Region of Interest (ROI): An area of video content (or a displayed or simulated environment) that a user is interested in at any given moment. For example, this could typically be an area on a sphere or a selection of polygons in a 2D map. ROIs identify specific areas for a particular purpose, defining the boundaries of the object under consideration.
[0012] • User location information: location information (e.g., x, y, z coordinates), orientation information (yaw, pitch, roll), direction of motion, and speed, etc.
[0013] • Viewport: The portion of the spherical video currently displayed and viewed by the user.
[0014] • Viewpoint: The center point of the viewport.
[0015] • 360-degree video (also known as immersive video or spherical video): In the context of this document, it refers to "video content" that contains more than one view (i.e., viewport) in one direction at any given time. Such content can be created, for example, using an omnidirectional camera or a set of cameras. During playback, the viewer can control the viewing direction.
[0016] An adaptive set contains a media stream or a collection of media streams. In its simplest case, an adaptive set contains all the audio and video of the content, but to reduce bandwidth, each stream can be split into a different adaptive set. A common scenario is having one video adaptive set and multiple audio adaptive sets (one audio adaptive set for each supported language). Adaptive sets can also contain subtitles or arbitrary metadata.
[0017] • This indicates that the adaptation set can contain the same content encoded in different ways. In most cases, the representation will be provided at multiple bitrates. This allows clients to request the highest quality content they can play without having to wait for buffering. The representation can also be encoded using different codecs, thus supporting clients with different supported codecs.
[0018] Media representation description (MPD) is an XML syntax that contains information about media segments, the relationships between media segments, and the information necessary to select between media segments.
[0019] In the context of this application, the concept of an adaptive set is used more generally, sometimes actually referring to a representation. Furthermore, media streams (audio / video streams) are typically first encapsulated into media segments, which are the actual media files played by a client (e.g., a DASH client). Various formats can be used for media segments, such as the ISO Basic Media File Format (ISOBMFF) and MPEG-TS, which are similar to MPEG-4 container formats. The encapsulation of media segments, and in different representation / adaptive sets, is independent of the methods described herein, which are applicable to all the various options.
[0020] Furthermore, while the methods described in this document may focus on DASH server-client communication, these methods are general enough to be used with other transport environments, such as MMT, MPEG-2 transport streams, DASH-ROUTE, and file formats used for file playback.
[0021] 3. Current Solution
[0022] The current solution is:
[0023] [1]. ISO / IEC 23008-3:015, Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D Audio
[0024] [2]. N16950, Study of ISO / IEC DIS23000-20 Omnidirectional Media Format
[0025] [3].M41184, Use of Earcons for ROI Identification in 360-degree Video.
[0026] ISO / IEC 23000-20 Omnidirectional Media Format [2] provides a mechanism for delivering 360-degree content. This standard specifies the media format used for encoding, storing, delivering, and rendering omnidirectional images, videos, and associated audio. It provides information related to media encoding and decoding for audio and video compression, as well as additional metadata information for the proper consumption of 360-degree A / V content. It also specifies constraints and requirements for delivery channels, such as streaming over DASH / MMT or file-based playback.
[0027] The concept of Earcon was first introduced in M41184 “ROI identification using Earcons in 360-degree video”[3], which provides a mechanism for signaling Earcon audio data signals to users.
[0028] However, some users have reported disappointing reviews of these systems. Often, a large number of earcons is annoying. When designers reduce the number of earcons, some users lose important information. It's worth noting that each user has their own level of knowledge and experience, and preferences for systems suited to their individual needs. To give just one example, each user prefers to reproduce earcons at a preferred volume (e.g., regardless of the volume used for other audio signals). It has proven difficult for system designers to achieve a system that provides a good level of satisfaction for all possible users. Therefore, a solution has been sought to allow for increased satisfaction for virtually all users.
[0029] Furthermore, it has been demonstrated that reconfiguring a system is difficult, even for designers. For example, they encounter difficulties when preparing new versions of audio streams and updating Earcons.
[0030] Furthermore, the restricted system imposes certain limitations on functionality, such as the inability to accurately identify Earcons within an audio stream. Additionally, Earcons must always be active, and playing them back when they are not needed can be frustrating for the user.
[0031] Furthermore, Earth spatial information cannot be signaled or modified by clients such as DASH. Easy access to this information at the system level enables additional features to provide a better user experience.
[0032] Furthermore, it lacks flexibility in handling various types of Earcons (e.g., natural sounds, synthesized sounds, sounds generated in the DASH Client, etc.).
[0033] All these issues result in a poor user experience. Therefore, a more flexible system architecture would be preferred. Summary of the Invention
[0034] 4. This invention
[0035] According to the example, a system is provided for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, the system being configured as follows:
[0036] Receive at least one video stream associated with the audio-visual scene to be reproduced; and
[0037] Receive at least one first audio stream associated with the audio-video scene to be reproduced.
[0038] The system includes:
[0039] At least one media video decoder is configured to decode at least one video signal from at least one video stream to represent an audio-video scene to a user; and
[0040] At least one media audio decoder is configured to decode at least one audio signal from at least one first audio stream to represent an audio-video scene to a user;
[0041] The Region of Interest (ROI) processor is configured as follows:
[0042] Based at least on the user's current viewport and / or head orientation and / or motion data and / or viewport metadata and / or audio information message metadata, determine whether to reproduce an audio information message associated with at least one ROI, wherein the audio information message is independent of the at least one video signal and the at least one audio signal; and
[0043] If it is decided to reproduce the information message, then the audio information message is reproduced.
[0044] According to the example, a system is provided for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, the system being configured to:
[0045] Receive at least one video stream; and
[0046] Receive at least one first audio stream.
[0047] The system includes:
[0048] At least one media video decoder is configured to decode at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; and
[0049] At least one media audio decoder is configured to decode at least one audio signal from the at least one first audio stream to represent an audio scene to a user;
[0050] The Region of Interest (ROI) processor is configured as follows:
[0051] Based on the user's current viewport and / or head orientation and / or motion data and / or viewport metadata and / or audio information message metadata, a decision is made as to whether to reproduce an audio information message associated with at least one ROI, wherein the audio information message is an earcon; and
[0052] If it is decided to reproduce the information message, then the audio information message is reproduced.
[0053] The system may include:
[0054] A metadata processor is configured to receive and / or process and / or manipulate audio information message metadata so that, when it is determined that the information message should be reproduced, the audio information message is reproduced based on the audio information message metadata.
[0055] The ROI processor can be configured as follows:
[0056] Receive the user's current viewport and / or position and / or head orientation and / or motion data and / or other user-related data; and
[0057] Receive viewport metadata associated with at least one video signal from the at least one video stream, the viewport metadata defining at least one ROI; and
[0058] Based on at least one of the user's current viewport and / or position and / or head orientation and / or motion data, as well as the viewport metadata and / or other criteria, a decision is made as to whether to reproduce the audio information message associated with the at least one ROI.
[0059] The system may include:
[0060] A metadata processor is configured to receive and / or process and / or manipulate audio information message metadata describing the audio information message and / or audio metadata describing the at least one audio signal encoded in the at least one audio stream and / or the viewport metadata, so as to reproduce the audio information message based on the audio information message metadata and / or the audio metadata describing the at least one audio signal encoded in the at least one audio stream and / or the viewport metadata.
[0061] The ROI processor can be configured as follows:
[0062] In cases where the at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or motion data, in addition to reproducing the at least one audio signal, the audio information message associated with the at least one ROI is also reproduced; and
[0063] If at least one ROI is within the user's current viewport and / or position and / or head orientation and / or motion data, the reproduction of audio information messages associated with the at least one ROI is disabled and / or deactivated.
[0064] The system can be configured as follows:
[0065] Receive at least one additional audio stream, wherein the at least one audio information message is encoded in the at least one additional audio stream.
[0066] The system also includes:
[0067] At least one multiplexer or multiplexer is used in the metadata processor and
[0068] / or under the control of the ROI processor and / or another processor, based on the ROI
[0069] The processor provides a decision to reproduce the at least one audio information message, which will...
[0070] A group of at least one additional audio stream and a group of the at least one first audio stream.
[0071] The audio is merged into a single stream so that the audio is reproduced in addition to the audio scene.
[0072] Information message.
[0073] The system can be configured as follows:
[0074] Receive at least one audio metadata describing the at least one audio signal encoded in the at least one audio stream;
[0075] Receive audio information message metadata associated with at least one audio information message from at least one audio stream;
[0076] If it is decided to reproduce the information message, the audio information message metadata is modified so that, in addition to reproducing the at least one audio signal, the audio information message can also be reproduced.
[0077] The system can be configured as follows:
[0078] Receive at least one audio metadata describing the at least one audio signal encoded in the at least one audio stream;
[0079] Receive audio information message metadata associated with at least one audio information message from the at least one audio stream;
[0080] If it is decided to reproduce the audio information message, the audio information message metadata is modified so that, in addition to reproducing the at least one audio signal, the audio information message associated with the at least one ROI can also be reproduced; and
[0081] Modify the audio metadata describing the at least one audio signal to allow merging the at least one first audio stream and the at least one additional audio stream.
[0082] The system can be configured as follows:
[0083] Receive at least one audio metadata describing the at least one audio signal encoded in the at least one audio stream;
[0084] Receive audio information message metadata associated with at least one audio information message from at least one audio stream;
[0085] If it is decided to reproduce the audio information message, the audio information message metadata is provided to the synthesized audio generator to create a synthesized audio stream, so that the audio information message metadata is associated with the synthesized audio stream, and the synthesized audio stream and the audio information message metadata are provided to a multiplexer or multiplexer to allow merging of the at least one audio stream and the synthesized audio stream.
[0086] The system can be configured as follows:
[0087] At least one additional audio stream, wherein the audio information message is encoded, is obtained from the at least one additional audio stream.
[0088] The system may include:
[0089] An audio information message metadata generator is configured to generate audio information message metadata based on a decision to reproduce an audio information message associated with the at least one ROI.
[0090] The system can be configured as follows:
[0091] Store the audio information message metadata and / or the audio information message stream for future use.
[0092] The system may include:
[0093] A synthesized audio generator is configured to synthesize audio information messages based on audio information message metadata associated with the at least one ROI.
[0094] The metadata processor is configured to: based on the audio metadata and / or audio information message metadata, merge the packets of the audio information message stream with the packets of the at least one first audio stream into one stream, so as to obtain the addition of the audio information message to the at least one audio stream.
[0095] The audio information message metadata can be encoded in a configuration frame and / or data frame that includes at least one of the following:
[0096] Identify labels,
[0097] An integer that uniquely identifies the reproduction of the audio information message metadata.
[0098] Message type
[0099] state,
[0100] Indicators of scene dependency / independence
[0101] Location data,
[0102] Gain data,
[0103] Indication of the presence of associated text tags.
[0104] The number of available languages
[0105] The language of audio information messages
[0106] Data text length,
[0107] The associated text label data text, and / or
[0108] Description of the audio message.
[0109] The metadata processor and / or the ROI processor can be configured to perform at least one of the following operations:
[0110] Extract audio information message metadata from the stream;
[0111] Modify the audio information message metadata to activate the audio information message and / or set / change the position of the audio information message;
[0112] Embed metadata back into the stream;
[0113] Feed the stream to the additional media decoder;
[0114] Extract audio metadata from at least one first audio stream;
[0115] Extract audio information message metadata from the attached stream;
[0116] Modify the audio information message metadata to activate the audio information message and / or set / change the position of the audio information message;
[0117] Modify the audio metadata of at least one first audio stream to account for the presence of audio information messages and allow merging;
[0118] Based on the information received from the ROI processor, the stream is fed to a multiplexer or multiplexer for multiplexing or multiplexing.
[0119] The ROI processor can be configured to: perform a local search for the additional audio stream and / or audio information message metadata in which the audio information message is encoded, and, if not found, request the additional audio stream and / or audio information message metadata from a remote entity.
[0120] The ROI processor is configured to perform a local search for the additional audio stream and / or audio information message metadata, and, if not found, to cause the synthesized audio generator to generate the audio information message stream and / or audio information message metadata.
[0121] The system can be configured as follows:
[0122] Receive the at least one additional audio stream, the at least one additional audio stream including at least one audio information message associated with the at least one ROI; and
[0123] If the ROI processor decides to reproduce the audio information message associated with the at least one ROI, then the at least one additional audio stream is decoded.
[0124] The system may include:
[0125] At least one first audio decoder for decoding the at least one audio signal from at least one first audio stream;
[0126] At least one additional audio decoder for decoding the at least one audio information message from an additional audio stream; and
[0127] At least one mixer and / or renderer for mixing and / or overlaying audio information messages from the at least one additional audio stream with at least one audio signal from the at least one first audio stream.
[0128] The system can be configured to track metrics associated with historical data and / or statistics related to the reproduction of the audio information message, so as to disable the reproduction of the audio information message if the metrics exceed a predetermined threshold.
[0129] The ROI processor's decision can be based on predictions of the user's current viewport and / or position and / or head orientation and / or motion data 122 relative to the ROI.
[0130] The system can be configured to receive at least one first audio stream and, when deciding to reproduce an information message, request an audio message information stream from a remote entity.
[0131] The system can be configured to determine whether to reproduce two audio information messages simultaneously, or whether to prioritize the reproduction of a higher-priority audio information message relative to a lower-priority audio information message.
[0132] The system can be configured to identify the audio information message among multiple audio information messages encoded in an additional audio stream, based on the address and / or position of the audio information message in the audio stream.
[0133] The audio stream can be formatted as MPEG-H 3D audio stream format.
[0134] The system can be configured as follows:
[0135] Receive data on the availability of multiple adaptive sets, the available adaptive sets including at least one audio scene adaptive set for the at least one first audio stream and at least one audio message adaptive set for at least one additional audio stream, the at least one additional audio stream containing at least one audio information message;
[0136] Based on the decision of the ROI processor, selection data is created, which identifies which adaptive set to retrieve, and the available adaptive sets include at least one audio scene adaptive set and / or at least one audio message adaptive set; and
[0137] Request and / or retrieve data from the adaptive set identified by the selected data.
[0138] Each adaptive set groups different codes for different bit rates.
[0139] The system can enable at least one of its elements to include HTTP-based, DASH-based, client-based dynamic adaptive streaming, and / or to retrieve data for each adaptive set using the ISO-based media file format ISO BMFF or the MPEG-2 transport stream MPEG-2TS.
[0140] The ROI processor can be configured to: check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or motion data, so as to check whether the ROI is represented in the current viewport, and if the ROI is not represented in the current viewport and / or position and / or head orientation and / or motion data, signal the presence of the ROI to the user in the form of sound.
[0141] The ROI processor can be configured to: check the correspondence between the ROI and the current viewport and / or position and / or head orientation and / or motion data, so as to check whether the ROI is represented in the current viewport, and if the ROI is within the current viewport and / or position and / or head orientation and / or motion data, not to notify the user of the presence of the ROI in the form of sound.
[0142] The system can be configured to receive from a remote entity at least one video stream associated with the video environment scene and at least one audio stream associated with the audio scene, wherein the audio scene is associated with the video environment scene.
[0143] The ROI processor can be configured to select a first audio message to reproduce before a second audio message from among a plurality of audio messages to be reproduced.
[0144] The system may include: a high-speed buffer memory for storing audio information messages received from or synthesized from remote entities, so as to reuse the audio information messages at different time instances.
[0145] The audio information message can be an earcon.
[0146] The at least one video stream and / or the at least one first audio stream may be part of the current video environment scene and / or video audio scene, respectively, and are independent of the user's current viewport and / or head orientation and / or motion data in the current video environment scene and / or video audio scene.
[0147] The system can be configured to request the at least one first audio stream and / or at least one video stream from a remote entity in association with the audio stream and / or video environment stream, respectively, and to reproduce the at least one audio information message based on the user's current viewport and / or head orientation and / or motion data.
[0148] The system can be configured to request the at least one first audio stream and / or at least one video stream from a remote entity in association with the audio stream and / or video environment stream, respectively, and to request the at least one audio information message from the remote entity based on the user's current viewport and / or head orientation and / or motion data.
[0149] The system can be configured to: request the at least one first audio stream and / or at least one video stream from a remote entity in association with the audio stream and / or video environment stream, respectively, and synthesize the at least one audio information message based on the user's current viewport and / or head orientation and / or motion data.
[0150] The system can be configured to: check at least one of the additional criteria for reproducing the audio information message, the criteria further including user selection and / or user settings.
[0151] The system is also configured to check at least one of the additional standards for reproducing the audio information message, the standard further including the state of the system.
[0152] The system can be configured to: check at least one of the additional criteria for reproducing the audio information message, the criteria further including the number of audio information messages reproduced that have been performed.
[0153] The system can be configured to: examine at least one of the additional criteria used to reproduce the audio information message, the criteria further including a flag in a data stream obtained from a remote entity.
[0154] According to one aspect, a system is provided, the system comprising: a client configured according to any of the examples above and / or below; and a remote entity configured as a server for transmitting at least one video stream and at least one audio stream.
[0155] The remote entity can be configured to: search for at least one additional audio stream and / or audio information message metadata in a database, intranet, Internet and / or geographic network, and, if found, transmit the at least one additional audio stream and / or the audio information message metadata.
[0156] The remote entity can be configured to synthesize at least one additional audio stream and / or generate audio information message metadata.
[0157] According to one aspect, a method for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments can be provided, the method comprising:
[0158] Decode at least one video signal from at least one video audio scene to reproduce it to the user;
[0159] Decode at least one audio signal from the video audio scene for reproduction;
[0160] Based on the user's current viewport and / or head orientation and / or motion data and / or metadata, determine whether to reproduce an audio information message associated with at least one ROI, wherein the audio information message is independent of the at least one video signal and the at least one audio signal; and
[0161] If it is decided to reproduce the information message, then the audio information message is reproduced.
[0162] According to one aspect, a method for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments can be provided, the method comprising:
[0163] Decode at least one video signal from at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to the user;
[0164] Decode at least one audio signal from at least one first audio stream to represent an audio scene to the user;
[0165] Based on the user's current viewport and / or head orientation and / or motion data and / or metadata, determine whether to reproduce an audio information message associated with at least one ROI, wherein the audio information message is an earcon; and
[0166] If it is decided to reproduce the information message, then the audio information message is reproduced.
[0167] The above and / or the following methods may include:
[0168] Receive and / or process and / or manipulate metadata so that, if it is determined that an information message should be reproduced, the audio information message is reproduced according to the metadata so that the audio information message is part of an audio scene.
[0169] The above and / or the following methods may include:
[0170] Reproducing audio and video scenes; and
[0171] Based on the user's current viewport and / or head orientation and / or motion data and / or metadata, it is determined whether to reproduce the audio information message.
[0172] The above and / or the following methods may include:
[0173] Reproducing audio and video scenes; and
[0174] In cases where at least one ROI is outside the user's current viewport and / or position and / or head orientation and / or motion data, in addition to reproducing the at least one audio signal, the audio information message associated with the at least one ROI is also reproduced; and / or
[0175] If at least one ROI is within the user's current viewport and / or position and / or head orientation and / or motion data, the reproduction of audio information messages associated with the at least one ROI is disabled and / or deactivated.
[0176] According to the example, a system is provided for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments, the system being configured to:
[0177] Receive at least one video stream; and
[0178] Receive at least one first audio stream.
[0179] The system includes:
[0180] At least one media video decoder is configured to decode at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; and
[0181] At least one media audio decoder is configured to decode at least one audio signal from the at least one first audio stream to represent an audio scene to a user;
[0182] The Region of Interest (ROI) processor is configured as follows:
[0183] Based on the user's current viewport and / or head orientation and / or motion data and / or metadata, determine whether to reproduce audio information messages associated with at least one ROI;
[0184] and
[0185] If it is decided to reproduce the information message, then the audio information message is reproduced.
[0186] In the example, a system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments is provided, the system being configured to:
[0187] Receive at least one video stream; and
[0188] Receive at least one first audio stream.
[0189] The system includes:
[0190] At least one media video decoder is configured to decode at least one video signal from the at least one video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user; and
[0191] At least one media audio decoder is configured to decode at least one audio signal from the at least one first audio stream to represent an audio scene to a user;
[0192] The Region of Interest (ROI) processor is configured to: determine, based on the user's current viewport and / or position and / or head orientation and / or motion data, as well as viewport metadata and / or other criteria, whether to reproduce audio information messages associated with at least one ROI; and
[0193] A metadata processor is configured to receive and / or process and / or manipulate metadata so that, if it is determined that an information message should be reproduced, the audio information message is reproduced based on the metadata, such that the audio information message is part of an audio scene.
[0194] According to one aspect, a non-transitory storage unit is provided for storing instructions that, when executed by a processor, cause the processor to perform the methods described above and / or below. Attached Figure Description
[0195] Figure 1 Figures 5a, 5b and 6 show examples of the implementation;
[0196] Figure 7 The method shown is based on the example;
[0197] Figure 8 An example of the implementation is shown. Detailed Implementation
[0198] 6.1 General Example
[0199] Figure 1 An example of a system 100 for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments is shown. System 100 may be associated, for example, with a content consumption device (e.g., a head-mounted display, etc.) that reproduces visual data in a spherical or hemispherical display closely associated with the user's head.
[0200] System 100 may include at least one media video decoder 102 and at least one media audio decoder 112. System 100 may receive at least one video stream 106, in which video signals are encoded to represent VR, AR, MR, or 360-degree video environment scene 118a to a user. System 100 may receive at least one first audio stream 116, in which audio signals are encoded to represent audio scene 118b to a user.
[0201] System 100 may also include a Region of Interest (ROI) processor 120. The ROI processor 120 can process data associated with the ROI. Generally, the presence of the ROI can be signaled in viewport metadata 131. Viewport metadata 131 can be encoded in video stream 106 (in other examples, viewport metadata 131 can be encoded in other streams). Viewport metadata 131 may include, for example, location information associated with the ROI (e.g., coordinate information). For example, in this example, the ROI can be understood as a rectangle (identified by coordinates such as the position of one of the four vertices of a rectangle in a spherical video and the length of the rectangle's sides). The ROI is typically projected into the spherical video. The ROI is typically associated with a visible element that is considered (depending on a specific configuration) to be of interest to the user. For example, the ROI may be associated with a rectangular area displayed by a content consuming device (or be visible to the user in some way).
[0202] The ROI processor 120 can specifically control the operation of the media audio decoder 112.
[0203] The ROI processor 120 can obtain user's current viewport and / or position and / or head orientation and / or motion data 122 associated with the user's current viewport and / or position and / or head orientation and / or motion (in some examples, virtual data associated with a virtual position may also be understood as part of data 122). This user's current viewport and / or position and / or head orientation and / or motion data 122 may be provided at least in part, for example, by a content consuming device or by a positioning / detection unit.
[0204] The ROI processor 120 can examine the correspondence between the ROI and the user's current viewport and / or position (actual or virtual) and / or head orientation and / or motion data 122 (other criteria may be used in this example). For example, the ROI processor can check whether the ROI is represented in the current viewport. If the ROI is only partially represented in the viewport (e.g., based on the user's head movement), it can determine, for example, whether a minimum percentage of the ROI is displayed on the screen. In any case, the ROI processor 120 is able to identify whether the ROI is not represented or is not visible to the user.
[0205] If the ROI is deemed to be outside the user's current viewport and / or position and / or head orientation and / or motion data 122, the ROI processor 120 may auditorily signal to the user the presence of the ROI. For example, in addition to the audio signal decoded from at least one first audio stream 116, the ROI processor 120 may also request the reproduction of an audio information message (Earcon).
[0206] If the ROI is considered to be within the user's current viewport and / or position and / or head orientation and / or motion data 122, the ROI processor may decide not to reproduce the audio information message.
[0207] Audio information messages can be encoded in audio stream 140 (audio information message stream), which may be the same as or different from audio stream 116. Audio stream 140 may be generated by system 100 or obtained from an external entity (e.g., a server). Audio metadata, such as audio information message metadata 141, can be defined to describe the attributes of audio information stream 140.
[0208] Audio information messages may be superimposed (or mixed, multiplexed, merged, combined, or composed) onto the signal encoded in audio stream 116, or may not be selected, for example, based solely on the decision of ROI processor 120. The decision of ROI processor 120 may be based on viewport and / or position and / or head orientation and / or motion data 122, metadata (such as viewport metadata 131 or other metadata), and / or other criteria (e.g., selection, system status, number of audio information messages reproduced, specific functions and / or operations, user preference settings that can be disabled using Earcons, etc.).
[0209] A metadata processor 132 can be implemented. The metadata processor 132 can be, for example, inserted between an ROI processor 120 (which can be controlled via the metadata processor 132) and a media audio decoder 112 (which can be controlled via the metadata processor). In this example, the metadata processor is part of the ROI processor 120. The metadata processor 132 can receive, generate, process, and / or manipulate audio information message metadata 141. The metadata processor 132 can also process and / or manipulate the metadata of the audio stream 116, for example, by mixing the audio stream 116 with the audio information message stream 140. Alternatively or additionally, the metadata processor 132 can receive the metadata of the audio stream 116, for example, from a server (e.g., a remote entity).
[0210] Therefore, the metadata processor 132 can alter the audio scene reproduction and adapt the audio information message to specific situations and / or choices and / or states.
[0211] This section discusses some of the advantages of this implementation.
[0212] Audio information messages can be accurately identified, for example, by using audio information message metadata 141.
[0213] Audio information messages can be easily activated / deactivated, for example, by modifying metadata (e.g., via metadata processor 132). Audio information messages can be enabled / disabled based on the current viewport and ROI information (and any specific functionality or effect to be implemented).
[0214] For example, audio information messages (containing information such as status, type, spatial information, etc.) can be easily signaled and modified via common devices—such as HTTP Dynamic Adaptive Streaming (DASH) clients.
[0215] Therefore, easy access to audio information messages (including, for example, status, type, spatial information, etc.) at the system level can enable additional functionality to provide a better user experience. Consequently, System 100 can be easily customized, allowing for further implementations (e.g., application-specific implementations) that can be performed by personnel independent of the designer of System 100.
[0216] Furthermore, it offers flexibility in handling various types of audio information messages (e.g., natural sounds, synthesized sounds, sounds generated in the DASH client, etc.).
[0217] Other advantages (which will also be apparent in the following examples):
[0218] • Use text tags in metadata (as a basis for displaying content or generating Earcon).
[0219] • Adaptive use of the device's earcon position (if the device is an HMD, accurate positioning is desirable, while if the device is a speaker, a better approach might be to use different positions within a single speaker).
[0220] Different equipment categories:
[0221] Earthcon metadata can be created by signaling Earthcon to indicate that it is active.
[0222] Some devices will only know how to parse metadata and reproduce Earcon.
[0223] Some newer devices with additional, better ROI processors can decide to activate the ROI processor when it is not needed.
[0224] More information and additional graphics about adaptive sets.
[0225] Therefore, in VR / AR environments, users can typically visualize 360-degree content using, for example, a head-mounted display (HMD) and listen to it through headphones. Users can typically move within the VR / AR space, or at least change their viewing orientation—the so-called "viewport" of the video. Compared to traditional content consumption, for VR, content creators no longer have control over what the user visualizes at each moment—the current viewport. Users can freely choose different viewports from allowed or available viewports at each time instance. To indicate a region of interest (ROI) to the user, audible sound (natural or synthetic) can be used at the location of the ROI. These audio messages are called "Earcons." This invention proposes a solution for efficiently delivering such messages and presents an optimized receiver behavior to utilize Earcons without impacting user experience and content consumption. This results in improved experience quality. This can be achieved at the system level using dedicated metadata and metadata manipulation mechanisms to enable or disable Earcons in the final scene.
[0226] Metadata processor 132 can be configured to receive and / or process and / or manipulate audio information message metadata 141 so that, when it is determined whether to reproduce the information message, the audio information message is reproduced according to the audio information message metadata 141. Audio signals (e.g., signals used to represent a scene) can be understood as part of an audio scene (e.g., an audio scene downloaded from a remote server). Audio signals are generally semantically meaningful to an audio scene, and all audio signals appearing together constitute the audio scene. Audio signals can be encoded together in an audio bitstream. Audio signals can be created by a content creator and / or can be associated with a specific scene and / or can be independent of the ROI.
[0227] Audio information messages (e.g., earcon) can be understood as having no semantic meaning to the audio context. This audio information message can be understood as an independent sound that can be artificially generated, such as a recorded sound, a human voice on a tape recorder, etc. It can also be device-dependent (e.g., a system sound generated when a button on a remote control is pressed). Audio information messages (e.g., earcon) can be understood as intended to guide the user within a context, rather than being part of the context itself.
[0228] The audio information message can be independent of the audio signal as described above. Depending on the example, it can be included in the same bitstream, or it can be sent in a separate bitstream, or it can be generated by system 100.
[0229] An example of an audio scene consisting of multiple audio signals could be:
[0230] --Audio scene, a concert hall containing the following 5 audio signals:
[0231] ---Audio signal 1: The sound of a piano
[0232] ---Audio signal 2: The singer's voice
[0233] ---Audio signal 3: The voices of the audience members (part 1)
[0234] ---Audio signal 4: The sound of the audience members (part 2)
[0235] ---Audio signal 5: The sound produced by the clock on the wall.
[0236] Audio messages can be recorded sounds such as "Look at the pianist" (the piano is the ROI). If the user is already looking at the pianist, the audio message will not be played back.
[0237] Another example: A door opens behind the user (e.g., a virtual door), and a new person enters the room; the user is not looking in that direction. Based on this (information about the VR environment, such as virtual location), an earcon can be triggered to notify the user that something is happening behind them.
[0238] In the example, each scene (e.g., with associated audio and video streams) is sent from the server to the client as the user changes the environment.
[0239] Audio messages can be flexible. Specifically:
[0240] - Audio information messages can reside in the same audio stream associated with the scene to be reproduced;
[0241] - Audio information messages can be located in an additional audio stream;
[0242] - Audio information messages may be completely lost, but only the metadata describing Earcons can appear in the stream and audio information messages can be generated in the system;
[0243] - The audio message and the metadata describing the audio message may be completely lost. In this case, the system will generate both (earcon and metadata) based on other information related to the ROI in the stream.
[0244] Audio information messages are typically independent of any audio signal portion of an audio scene and are not used to represent an audio scene.
[0245] Below are examples of systems that embody or include parts that embody System 100.
[0246] 6.2 Example from Figure 2
[0247] Figure 2 illustrates system 200 (which may include at least a portion of system 100), which is represented here as a subdivision of server side 202, media delivery side 203, client side 204, and / or media consumption device side 206. Each of server side 202, media delivery side 203, client side 204, and media consumption device side 206 is a system in itself and can be combined with any other system to obtain another system. Here, even though audio information messages can be generalized as any kind of audio information message, audio information messages are referred to as Earcons.
[0248] The client side 204 can receive at least one video stream 106 and / or at least one audio stream 116 from the server side 202 via the media transmission side 203.
[0249] The media delivery side 203 can be based on communication systems such as cloud systems, network systems, geographic communication networks, or well-known media transmission formats (MPEG-2 TS transport streams, DASH, MMT, DASH ROUTE, etc.), or even file-based storage. The media delivery side 203 can perform communication in the form of electrical signals (e.g., via cable, wireless, etc.) and / or by distributing packets (e.g., according to a specific communication protocol) together with bitstreams in which audio and video signals are encoded. However, the media delivery side 203 can be embodied through point-to-point links, serial connections, or parallel connections. The media delivery side 203 can perform wireless connections, for example, according to protocols such as WiFi, Bluetooth, etc.
[0250] Client-side 204 can be associated with a media consumption device (e.g., an HND), into which the user's head can be inserted (however, other devices can be used). Therefore, the user can experience a video and audio scene (e.g., a VR scene) prepared by client-side 204 based on video and audio data provided by server-side 202. However, other implementations are also possible.
[0251] Server-side 202 is represented here as having a media encoder 240 (which may cover a video encoder, audio encoder, subtitle encoder, etc.). This media encoder 240 may be associated, for example, with an audio-visual scene to be represented. The audio scene may be used, for example, to reconstruct an environment and is associated with at least one audio stream 116 and at least one video stream 106, which may be encoded based on the user's position (or virtual position) in a VR, AR, or MR environment. Generally, at least one video stream 106 encodes a spherical image, only a portion of which (the viewport) is seen by the user based on its position and motion. The audio stream 116 contains audio data that participates in the audio scene representation and is intended to be heard by the user. According to an example, the audio stream 116 may include audio metadata 236 (which refers to at least one audio signal intended to participate in the audio scene representation) and / or Earcon metadata (audio information message metadata) 141 (which may describe Earcons that will only be reproduced in certain situations).
[0252] System 100 is represented here as client-side 204. For simplicity, media video decoder 102 is not shown in Figure 2.
[0253] To prepare for the reproduction of an Earcon (or other audio information message), Earcon metadata 141 can be used. Earcon metadata 141 can be understood as metadata that describes and provides attributes associated with the Earcon (which can be encoded in the audio stream). Therefore, the Earcon (if to be reproduced) can be based on the attributes of Earcon metadata 141.
[0254] Advantageously, the metadata processor 132 can be specifically implemented for processing Earcon metadata 141. For example, the metadata processor 132 can control the reception, processing, manipulation, and / or generation of Earcon metadata 141. When processed, the Earcon metadata can be represented as modified Earcon metadata 234. For example, it is possible to manipulate the Earcon metadata to obtain specific effects and / or to perform audio processing operations, such as multiplexing or remultiplexing, for adding the Earcon to the audio signal to be represented in the audio scene.
[0255] The metadata processor 132 can control the reception, processing, and manipulation of audio metadata 236 associated with at least one audio stream 116. When processed, the audio metadata 236 can be represented as modified audio metadata 238.
[0256] The modified Earcon metadata 234 and the modified audio metadata 238 can be provided to the media audio decoder 112 (or, in some examples, to multiple decoders) to reproduce the audio scene 118b to the user.
[0257] In the example, a synthesized audio generator 246 and / or a storage device may be provided as optional components. The generator can synthesize audio streams (e.g., to generate Earcon streams not encoded in the stream). The storage device allows (e.g., in a cache) the storage of the Earcon streams generated by the generator and / or obtained from the received audio streams (e.g., for future use).
[0258] Therefore, the ROI processor 120 can determine the representation of the earcon based on the user's current viewport and / or position and / or head orientation and / or motion data 122. However, the ROI processor 120 can also make its decisions based on criteria involving other aspects.
[0259] For example, an ROI processor can enable / disable Earcon reproduction based on other conditions, such as user choice or higher-level choices, like based on a specific application intended for consumption. For video game applications, for example, for high-level video games, Earcon or other audio information messages can be avoided. This can be easily achieved by a metadata processor by disabling Earcons in the Earcon metadata.
[0260] Furthermore, Earcon can be disabled based on the following system status: for example, if Earcon has been reproduced, its repetition can be prevented. Timers can be used, for example, to prevent repetition from happening too quickly.
[0261] The ROI processor 120 can also request a controlled reproduction of a set of Earcons (e.g., Earcons associated with all ROIs in the scene), for example, to guide the user about the elements he / she can see. The metadata processor 132 can control this operation.
[0262] The ROI processor 120 can also modify the Earcon location (i.e., spatial location in the scene) or the Earcon type. For example, some users may prefer to play a specific sound as an Earcon at the exact location / position of the ROI, while other users may prefer to always play an Earcon as a sound indication of the ROI's location from a fixed position (e.g., the center or top position).
[0263] The gain of the Earcon's reproduction can be modified (e.g., to obtain a different volume). This decision can, for example, follow the user's choice. It is worth noting that, based on the ROI processor's decision, the metadata processor 132 will perform the gain modification by modifying specific attributes associated with the gain in the Earcon metadata associated with the Earcon.
[0264] The original designers of VR, AR, and MR environments may not yet know how to actually reproduce the Earcons. For example, user choices can modify the final rendering of the Earcons. This operation can be controlled, for example, by a metadata processor 132, which can modify the Earcon metadata 141 based on the decisions of the ROI processor.
[0265] Therefore, the operations performed on the audio data associated with the Earcon are, in principle, independent of at least one audio stream 116 used to represent the audio scene, and can be managed differently. The Earcon can even be generated independently of the audio stream 116 and video stream 106 constituting the audio-video scene, and the Earcon can be generated by different and independent entrepreneurial groups.
[0266] Therefore, this example allows for increased user satisfaction. For instance, a user can make their own choice by disabling audio message messages (e.g., by modifying the volume of the audio message). Thus, each user can have an experience more suited to their preferences. Furthermore, the resulting architecture is more flexible. Audio message messages can be easily updated, for example, by modifying metadata independently of the audio stream, and / or by modifying the audio message stream independently of both the metadata and the main audio stream.
[0267] The resulting architecture is also compatible with legacy systems: for example, a legacy audio message stream can be associated with new audio message metadata. In the absence of a suitable audio message stream, in this example, the latter can be easily synthesized (and, for example, stored for later use).
[0268] The ROI processor can track metrics associated with historical data and / or statistics related to the reproduction of audio information messages so that the reproduction of audio information messages can be disabled if the metrics exceed a predetermined threshold (this can be used as a criterion).
[0269] As a standard, the ROI processor's decision can be based on a prediction of the user's current viewport and / or position and / or head orientation and / or motion data 122 relative to the ROI.
[0270] The ROI processor can also be configured to: receive at least one first audio stream 116; and, when deciding to reproduce the information message, request the audio message information stream from the remote entity.
[0271] The ROI processor and / or metadata generator can also be configured to determine whether to reproduce two audio information messages simultaneously, or whether to prioritize the reproduction of a higher-priority audio information message relative to a lower-priority audio information message. To perform this decision, audio information metadata can be used. Priority can be obtained, for example, by the metadata processor 132 based on values in the audio information message metadata.
[0272] In some examples, the media encoder 240 can be configured to search for additional audio stream and / or audio information message metadata in a database, intranet, internet, and / or geographic network, and, if retrieved, deliver the additional audio stream and / or audio information message metadata. For example, a search can be performed on a request from the client side.
[0273] As described above, a solution is proposed here for efficiently delivering Earcon messages along with audio content. Optimized receiver behavior is achieved to utilize audio information messages (e.g., Earcons) without impacting user experience and content consumption. This will result in improved experience quality.
[0274] This can be achieved by using dedicated metadata and metadata manipulation mechanisms at the system level to enable or disable audio information messages in the final audio scene. The metadata can be used with any audio codec and can complement next-generation audio codec metadata in a well-structured way (e.g., MPEG-H audio metadata).
[0275] The transmission mechanism can be varied (e.g., streaming via DASH / HLS, broadcasting via DASH-ROUTE / MMT / MPEG-2TS, file playback, etc.). In this application, DASH transmission is considered, but all concepts are valid for other transmission options.
[0276] In most cases, audio information messages do not overlap in the temporal domain; that is, only one ROI is defined at a specific point in time. However, considering more advanced use cases, such as interactive environments where users can change content based on their selections / movements, there may be use cases requiring multiple ROIs. Therefore, more than one audio information message may be needed at any given moment. Thus, a general solution to support all these different use cases is described.
[0277] The transmission and processing of audio information messages should complement existing methods of next-generation audio transmission.
[0278] One approach to transmitting multiple audio information messages for several temporally independent Regions of Interest (ROIs) is to combine all audio information messages into a single audio element (e.g., an audio object), where associated metadata describes the spatial location of each audio information message at different time instances. Since the audio information messages do not overlap temporally, they can be addressed independently within a shared audio element. This audio element can contain silence (or no audio data) between audio information messages, i.e., whenever no audio information message is present. In this case, the following mechanism can be applied:
[0279] • A generic audio message audio element can be transmitted in the same base stream (ES) as the audio scene it is associated with, or a generic audio message audio element can be transmitted in an auxiliary stream (dependent on or independent of the main stream).
[0280] • If Earcon audio elements are delivered in an auxiliary stream that relies on the main stream, the client can request an additional stream whenever a new ROI appears in the visual scene.
[0281] • In the example, the client (e.g., system 100) can request the stream before the scenario requires Earcon.
[0282] • In the example, the client can request a stream based on the current viewport; that is, if the current viewport matches the ROI, the client can decide not to request an additional Earcon stream.
[0283] If Earcon audio elements can be delivered in an auxiliary stream independent of the main stream, then clients can request additional streams as before whenever a new ROI appears in the visual scene. Furthermore, two media decoders and a common rendering / mixing step can be used to process two (or more) streams to blend the decoded Earcon audio data into the final audio scene. Alternatively, a metadata processor can be used to modify the metadata of the two streams, and a "stream merger" can be used to merge the two streams. Possible implementations of such a metadata processor and stream merger are described below.
[0284] In alternative examples, multiple Earcones of several ROIs that are temporally independent or temporally overlapping can be delivered in multiple audio elements (e.g., audio objects) and embedded in a base stream along with the main audio scene, or embedded in multiple auxiliary streams, such as each Earcon in an ES, or a group of Earcones in an ES based on shared properties (e.g., all Earcones on the left share a stream).
[0285] • If all Earcon audio elements rely on several mainstream auxiliary streams
[0286] Transmitted in (e.g., one Earcon per stream or a group of Earcones per stream),
[0287] For example, whenever the ROI associated with this Earcon exists in the visual scene
[0288] At that time, the client can request an additional stream containing the required Earcon.
[0289] In the example, the client can request the necessary equipment before any scenario requires Earcon.
[0290] Earcon's flow (e.g., even if the ROI is not yet part of the scene, it is based on...)
[0291] The ROI processor 120 can also make decisions regarding the user's movements.
[0292] In the example, the client can request a stream based on the current viewport, if the current viewport...
[0293] If the port matches the ROI, the client can decide not to request an additional Earcon stream.
[0294] • If an Earcon audio element (or a group of Earcon) is independent of the mainstream
[0295] If transmitted in the auxiliary stream, then in the example, whenever a new ROI appears in the visual field...
[0296] In this scenario, the client can request additional streams as before. Furthermore, it can enable...
[0297] Use two media decoders and a common rendering / blending step to process two (or) media decoders.
[0298] (More) streams to mix the decoded Earcon audio data into the final audio.
[0299] In high-frequency scenarios, alternatively, a metadata processor can be used to modify the two streams.
[0300] Metadata can be used to merge two streams using a "stream merger". The following describes...
[0301] Possible implementations of such metadata processors and stream mergers.
[0302] Alternatively, a single, generic Earcons can be used to signal all ROIs within an audio scene. This can be achieved by using the same video content, where different spatial information is associated with audio content at different temporal instances. In this case, the ROI processor 120 can: request the metadata processor 132 to collect the Earcons associated with the ROIs in the scene and sequentially control the reproduction of the Earcons (e.g., at the user's selection or at a higher-level application request).
[0303] Alternatively, an Earcon can only be transmitted once and cached on the client. The client can reuse an Earcon for all ROIs within an audio scene, where different spatial information is associated with audio content at different temporal instances.
[0304] Alternatively, the Earcon audio content can be generated in-house on the client side. In addition, a metadata generator can be used to create the necessary metadata to signal the spatial information of the Earcon. For example, the Earcon audio content can be compressed along with the main audio content and new metadata and fed into a media decoder, or the Earcon audio content can be mixed into the final audio scene after the media decoder, or several media decoders can be used.
[0305] Alternatively, in the example, the Earcon audio content can be synthesized in the client (e.g., under the control of metadata processor 132), with metadata describing the Earcon already embedded in the stream. The encoder uses specific signaling of the Earcon type, and the metadata can contain spatial information about the Earcon, specific signaling for the "decoder-generated Earcon," but without the Earcon's audio data.
[0306] Alternatively, the Earcon audio content can be generated in the client, and a metadata generator can be used to create the necessary metadata to signal the spatial information of the Earcon. For example, the Earcon audio content could:
[0307] • Compressed along with the main audio content and new metadata and fed into a media decoder;
[0308] Alternatively, the Earcon audio content can be mixed into the final audio scene after the media decoder;
[0309] Alternatively, multiple media decoders can be used.
[0310] 6.3 Examples of metadata for audio information messages (e.g., Earcons)
[0311] As mentioned above, an example of audio information message (Earcons) metadata 141 is provided here.
[0312] A structure for describing the Earcon attribute and providing the possibility of easily adjusting these values:
[0313]
[0314]
[0315]
[0316] Each identifier in this table can be intended to be associated with an attribute of the Earcon metadata.
[0317] This section discusses semantics.
[0318] numEarcons - This field specifies the number of Earcons audio elements available in the stream.
[0319] Earcon_isIndependent - This flag defines whether an Earcon audio element is independent of any audio scene. If Earcon_isIndependent == 1, the Earcon audio element is independent of the audio scene. If Earcon_isIndependent == 0, the Earcon audio element is part of the audio scene, and Earcon_id should have the same value as the mae_groupID associated with the same audio element.
[0320] EarconType - This field defines the type of Earcon. The following table specifies the allowed values.
[0321] EarconType describe 0 Undefined 1 Natural Sounds 2 Synthesized sound 3 Spoken text 4 General Earcon 5 / *Reserved* / 6 / *Reserved* / 7 / *Reserved* / 8 / *Reserved* / 9 / *Reserved* / 10 / *Reserved* / 11 / *Reserved* / 12 / *Reserved* / 13 / *Reserved* / 14 / *Reserved* / 15 other
[0322] The EarconActive flag defines whether the Earcon is active. If EarconActive == 1, the EarconAudio element should be decoded and rendered into the Audio scene.
[0323] EarconPosition This flag defines whether the Earcon has available location information. If Earcon_isIndependent == 0, this location information should be used instead of the audio object metadata specified in the dynamic_object_metadata() or intracoded_object_metadata_efficient() structure.
[0324] Earcon_azimuth: The absolute value of the azimuth angle.
[0325] Earcon_elevation: The absolute value of the elevation angle. The absolute value of the radius.
[0326] Earcon_radius
[0327] The EarconHasGain flag defines whether Earcon has different gain values.
[0328] The Earcon_gain field defines the absolute value of the Earcon gain.
[0329] The EarconHasTextLabel flag defines whether an Earcon has an associated text label.
[0330] The Earcon_numLanguages field specifies the number of languages available for the description text tag.
[0331] The Earcon_Language 24-bit field identifies the language of the Earcon's description text. It contains a 3-character code specified by ISO 639-2. Both ISO 639-2 / B and ISO 639-2 / T can be used. According to ISO / IEC 8859-1, each character is encoded as 8 bits and inserted sequentially into the 24-bit field. Example: The French code "fre" with 3 characters is encoded as: "011001100111001001100101".
[0332] The Earcon_TextDataLength field defines the length of the following group of descriptions in the bitstream.
[0333] The Earcon_TextData field contains a description of the Earcon, i.e., a string used to describe the content using a high-level description. The format should conform to UTF-8 according to ISO / IEC 10646.
[0334] A structure for identifying Earcons at the system level and associating them with existing viewports. The following two tables provide two ways to implement this structure in different implementations:
[0335]
[0336] Or alternatively:
[0337]
[0338] Semantics:
[0339] hasEarcon specifies whether Earcon data is available for a region.
[0340] numRegionEarcons specifies the number of Earcons available for a region.
[0341] Earcon_id uniquely defines the ID of an Earcon element associated with a spherical region. If the Earcon is part of an audio scene (i.e., the Earcon is part of a group of elements identified by a mae_groupID), then the Earcon_id should have the same value as the mae_groupID. Earcon_id can be used to identify audio files / tracks; for example, in the case of DASH delivery, the AdaptationSet in the MPD with the EarearComponent@tag element has the same Earcon_id.
[0342] Earcon_track_id is an integer that uniquely identifies an Earcon track associated with the spherical region throughout its entire lifetime. Specifically, if one or more Earcon tracks are transmitted within the same ISO BMFF file, Earcon_track_id represents the corresponding track_id for that Earcon track(s). If Earcon tracks are not transmitted within the same ISO BMFF file, this value should be set to zero.
[0343] To easily identify one or more Earcon tracks at the MPD level, the following attributes / elements can be used with the EarconComponent@tag:
[0344] Summary of MPD elements and attributes related to MPEG-H audio
[0345]
[0346]
[0347] For example, for MPEG-H audio, this can be achieved by using MHAS grouping:
[0348] • New MHAS packets can be defined to carry information about Earcons:
[0349] The PACTYP EARCON structure carries the EarconInfo() structure;
[0350] • New identifier field in the generic MHAS METADATA MHAS group.
[0351] Used to hold the EarconInfo() structure.
[0352] For metadata, metadata processor 132 may have at least some of the following functions:
[0353] Extract audio information message metadata from the stream;
[0354] Modify audio message metadata to activate the audio message, and / or set /
[0355] Change the position of the audio message and / or write / modify the audio message text label;
[0356] Embed metadata back into the stream;
[0357] Feed the stream to the additional media decoder;
[0358] Extract audio metadata from at least one first audio stream (116);
[0359] Extract audio information message metadata from the attached stream;
[0360] Modify audio message metadata to activate the audio message, and / or set /
[0361] Change the position of the audio message and / or write / modify the audio message text label;
[0362] Modify the audio metadata of at least one first audio stream (116) to take into account
[0363] The existence of audio information messages allows for merging;
[0364] Based on the information received from the ROI processor, the stream is fed to the multiplexer.
[0365] Use a device or multiplexer to multiplex or reuse it.
[0366] 6.4 Example from Figure 3
[0367] Figure 3 illustrates system 300, which on the client side 204 includes system 302 (client system) that may embody, for example, system 100 or 200.
[0368] System 302 may include ROI processor 120, metadata processor 132, and decoder group 313 formed by multiple media audio decoders 112.
[0369] In this example, different audio streams are decoded (each audio stream is decoded separately by the corresponding media audio decoder 112), and then the different audio streams are mixed together and / or rendered together to provide the final audio scene.
[0370] Here, at least one audio stream is represented as comprising two audio streams, 116 and 316 (other examples may provide a single stream as shown in Figure 2, or two or more streams). These are audio streams designed to reproduce the audio scene that the user expects to experience. This concept can be generalized to any audio message, also referencing Earcons.
[0371] Additionally, the media encoder 240 can provide an Earcon stream 140. Based on the ROI and / or other criteria indicated in the user's motion and viewport metadata 131, the ROI processor will enable the reproduction of the Earcon from the Earcon stream 140 (also indicated as an additional audio stream in addition to audio streams 116 and 316).
[0372] It is worth noting that the actual representation of the Earcon will be based on the Earcon metadata 141 and on the modifications performed by the metadata processor 132.
[0373] In the example, system 302 (client) can request a stream from media encoder 240 (server) when necessary. For example, the ROI processor can determine that a specific earcon will be needed soon based on the user's movement, and therefore can request the appropriate earcon stream 140 from media encoder 240.
[0374] The following aspects of this example are worth noting:
[0375] • Use case: Transmit audio data in one or more audio streams 116, 316 (e.g., a main stream and a secondary stream), and transmit (one or more) Earcon in one or more additional streams 140 (depending on or independent of the main audio stream).
[0376] In one implementation of client-side 204, ROI processor 120 and metadata processor 132 are used to efficiently process Earcon information.
[0377] The ROI processor 120 can receive information about the current viewport (user orientation information) from the media consumption device side 206 (e.g., HMD-based) used for content consumption. The ROI processor can also receive information about metadata and ROIs signaled in the metadata (e.g., signaling the video viewport via OMAF).
[0378] Based on this information, the ROI processor 120 can decide to activate one (or more) of the Earcones contained in the Earcon audio stream 140. Additionally, the ROI processor 120 can determine different locations of the Earcones and different gain values (e.g., to more accurately represent the Earcon in the current space where the content is consumed).
[0379] The ROI processor 120 provides this information to the metadata processor 132.
[0380] • Metadata processor 132 can parse the metadata contained in the Earcon audio stream, and
[0381] • Enable Earcon (to allow its reproduction)
[0382] Furthermore, if requested by the ROI processor 120, the metadata processor 132 accordingly modifies the spatial location and gain information contained in the Earcon metadata 141.
[0383] Then, each audio stream 116, 316, 140 is independently decoded and rendered (based on user location information), and the mixer or renderer 314 mixes the outputs of all media decoders together as the final step. Different implementations can only decode the compressed audio and provide the decoded audio data and metadata to the general renderer for the final rendering of all audio elements (including Earcons).
[0384] Additionally, in a streaming environment, based on the same information, the ROI processor 120 can decide to pre-request the Earcon stream 140 (e.g., when a user is looking in the wrong direction a few seconds before the ROI is enabled).
[0385] Example from Figure 4, No. 6.5
[0386] Figure 4 illustrates system 400, which on the client side 204 includes system 402 (client system) that can embody, for example, system 100 or 200. Here, it is also possible to generalize the concept to any audio information message, and also refer to Earcons.
[0387] System 402 may include ROI processor 120, metadata processor 132, and stream multiplexer or multiplexer 412. In the example of multiplexer or multiplexer 412, the number of operations to be performed by hardware is advantageously reduced compared to the number of operations to be performed when using multiple decoders and a mixer or renderer.
[0388] In this example, different audio streams are processed based on their metadata and then multiplexed or reused by multiplexer 412.
[0389] Here, at least one audio stream is represented as including two audio streams 116 and 316 (other examples may provide a single stream as shown in Figure 2, or two or more streams). These are audio streams designed to reproduce the audio scene that the user expects to experience.
[0390] Additionally, media encoder 240 can provide Earcon stream 140. Based on the ROI and / or other criteria indicated in the user's motion and viewport metadata 131, ROI processor 120 will enable the reproduction of the Earcon from Earcon stream 140 (also indicated as an additional audio stream in addition to audio streams 116 and 316).
[0391] Each audio stream 116, 316, 140 may include metadata 236, 416, 141, respectively. At least some of this metadata may be manipulated and / or processed to provide to a stream multiplexer or multiplexer 412, in which packets of audio streams are combined together. Thus, Earcon can be represented as part of the audio scene.
[0392] The stream multiplexer or multiplexer 412 can therefore provide an audio stream 414 including modified audio metadata 238 and modified Earcon metadata 234, which can be provided to the media audio decoder 112 and decoded and reproduced to the user.
[0393] The following aspects of this example are worth noting:
[0394] • Use case: Transmit audio data in one or more audio streams 116, 316 (e.g., a main audio stream 116 and an auxiliary audio stream 316, but a single audio stream may also be provided), and transmit (one or more) Earcon in one or more additional streams 140 (depending on or independent of the main audio stream).
[0395] In one implementation of client-side 204, ROI processor 120 and metadata processor 132 are used to efficiently process Earcon information.
[0396] The ROI processor 120 can receive information about the current viewport (user orientation information) from the media consumption device side 206 (e.g., HMD) used for content consumption. The ROI processor 120 can also receive information about the Earcon metadata 141 and the ROIs that are signaled in the Earcon metadata 141 (which can be signaled to the video viewport in the omnidirectional media application format OMAF).
[0397] Based on this information, the ROI processor 120 can decide to activate one (or more) earcons contained in the additional audio stream 140. Additionally, the ROI processor 120 can determine different positions of the earcons and different gain values (e.g., to more accurately represent the earcons in the current space where the content is consumed).
[0398] • The ROI processor 120 can provide this information to the metadata processor 132.
[0399] • Metadata processor 132 can parse the metadata contained in the Earcon audio stream, and
[0400] • Enable Earcon
[0401] Furthermore, if requested by the ROI processor, the metadata processor 132 accordingly modifies the spatial location and / or gain information and / or text labels contained in the Earcon metadata.
[0402] The metadata processor 132 can also parse the audio metadata 236 and 416 of all audio streams 116 and 316, and can manipulate audio-specific information so that the Earcon can be used as part of an audio scene (for example, if the audio scene has a 5.1 channel bed and 4 objects, the Earcon audio element is added to the scene as a fifth object). All metadata fields are updated accordingly.
[0403] Then, the audio data of each audio stream 116, 316, along with the modified audio metadata and Earcon metadata, are provided to a stream multiplexer or multiplexer, which can then generate an audio stream 414 with a set of metadata (modified audio metadata 238 and modified Earcon metadata 234).
[0404] • This stream 414 can be decoded by a single media audio decoder 112 based on user location information.
[0405] Furthermore, in a streaming environment, based on the same information, the ROI processor 120 can decide to pre-request the Earcon stream 140 (e.g., when a user is looking in the wrong direction a few seconds before the ROI is enabled).
[0406] 6.6 Example from Figure 5a
[0407] Figure 5a illustrates system 500, which on the client side 204 includes system 502 (client system) that can embody, for example, system 100 or 200. Here, it is also possible to generalize the concept to any audio information message, and also refer to Earcons.
[0408] System 502 may include ROI processor 120, metadata processor 132, stream multiplexer or multiplexer 412.
[0409] In this example, the Earcon stream is not provided by a remote entity (on the client side), but generated by a synthesized audio generator 246 (which may also have the ability to store the stream for later reuse or for use with compressed / uncompressed versions of the stored natural sound). However, the Earcon metadata 141 is provided by a remote entity, for example, in audio stream 116 (not the Earcon stream). Therefore, the synthesized audio generator 246 can be activated to create audio stream 140 based on the attributes of the Earcon metadata 141. For example, attributes may refer to the type of synthesized speech (natural sound, synthesized sound, spoken text, etc.) and / or text tags (Earcon can be generated by creating synthesized sound based on text in the metadata). In this example, after the Earcon stream is created, it can be stored for future reuse. Alternatively, the synthesized sound could be a generic sound permanently stored in the device.
[0410] A stream multiplexer or multiplexer 412 can be used to merge packets of audio stream 116 (and, in the case of other streams, packets of auxiliary audio stream 316) with packets of the Earcon stream generated by synthesized audio generator 246. Then, an audio stream 414 associated with modified audio metadata 238 and modified Earcon metadata 234 can be obtained. Audio stream 414 can be decoded by media audio decoder 112 and reproduced to the user on media consumption device side 206.
[0411] The following aspects of this example are worth noting:
[0412] • Use cases:
[0413] • Audio data in one or more audio streams (e.g., a mainstream and secondary)
[0414] Transmitted in the stream.
[0415] • The remote device did not transmit the Earcon, but the Earcon metadata 141 is used as the primary...
[0416] A portion of the audio stream is transmitted (specific signaling can be used to instruct the Earcon).
[0417] (No associated audio data).
[0418] • In one client-side implementation, there is an ROI processor 120 and a metadata processor.
[0419] 132 is used to efficiently process Earcon information.
[0420] The ROI processor 120 can receive information about the current viewport (user orientation information) from the device used by the media consumption device side 206 (e.g., HMD). The ROI processor 120 can also receive information about metadata and ROIs signaled in the metadata (e.g., signaling the video viewport via OMAF).
[0421] Based on this information, the ROI processor 120 can decide to activate one (or more) earcons that are not present in the audio stream 116. Additionally, the ROI processor 120 can determine different locations of the earcons and different gain values (e.g., to more accurately represent the earcon in the current space where the content is consumed).
[0422] • The ROI processor 120 can provide this information to the metadata processor 132.
[0423] • The metadata processor can parse the metadata contained in audio stream 116, and can
[0424] • Enable Earcon
[0425] Furthermore, if requested by the ROI processor 120, the metadata processor 132 accordingly modifies the spatial location and gain information contained in the Earcon metadata 141.
[0426] The metadata processor 132 can also parse the audio metadata (e.g., 236, 417) of all audio streams (116, 316) and can manipulate audio-specific information so that the Earcon can be used as part of an audio scene (e.g., if the audio scene has a 5.1 channel bed and 4 objects, the Earcon audio element is added to the scene as a fifth object). All metadata fields are updated accordingly.
[0427] The modified Earthcon metadata and information from the ROI processor 120 are provided to the synthesized audio generator 246. The synthesized audio generator 246 can create synthesized sounds based on the received information (e.g., generating a speech signal spelled out based on the spatial location of the Earthcon). Furthermore, the Earthcon metadata 141 is associated with the generated audio data in a new stream 414.
[0428] Similarly, as before, the audio data for each audio stream (116, 316), along with the modified audio metadata and Earcon metadata, is then provided to the stream multiplexer, which can then generate an audio stream with a set of metadata (audio and Earcon).
[0429] • This stream 414 is decoded by a single media audio decoder 112 based on user location information.
[0430] Alternatively or additionally, the Earcon audio data can be cached on the client (e.g., based on previous Earcon usage).
[0431] Alternatively, the output of the synthesized audio generator 246 can be uncompressed audio, which can be mixed into the final rendered scene.
[0432] Furthermore, in a streaming environment, based on the same information, the ROI processor 120 can decide to request an Earcon stream in advance (e.g., when a user is looking in the wrong direction a few seconds before the ROI is enabled).
[0433] 6.7 Example from Figure 6
[0434] Figure 6 illustrates system 600, which on the client side 204 includes system 602 (client system) that can embody, for example, system 100 or 200. Here, it is also possible to generalize the concept to any audio information message, and also refer to Earcons.
[0435] System 602 may include ROI processor 120, metadata processor 132, stream multiplexer or multiplexer 412.
[0436] In this example, the Earcon stream is not provided by a remote entity (on the client side), but is generated by a synthesized audio generator 246 (which may also have the ability to store the stream for later reuse).
[0437] In this example, the remote entity does not provide Earcon metadata 141. Earcon metadata is generated by metadata generator 432, which can generate Earcon metadata that metadata processor 132 will use (e.g., process, manipulate, modify). The Earcon metadata 141 generated by Earcon metadata generator 432 may have the same structure and / or format and / or attributes as the Earcon metadata discussed for the previous example.
[0438] Metadata processor 132 can operate as shown in the example of Figure 5a. Synthetic audio generator 246 can be activated to create audio stream 140 based on attributes of Earthcon metadata 141. For example, attributes could refer to the type of synthesized speech (natural voice, synthesized voice, speech text, etc.), and / or gain, and / or active / inactive state, etc. In the example, after creating Earthcon stream 140, it can be stored (e.g., cached) for future reuse. Earthcon metadata generated by Earthcon metadata generator 432 can also be stored (e.g., cached).
[0439] A stream multiplexer or multiplexer 412 can be used to merge packets of audio stream 116 (and, in the case of other streams, packets of auxiliary audio stream 316) with packets of the Earcon stream generated by synthesized audio generator 246. Then, an audio stream 414 associated with modified audio metadata 238 and modified Earcon metadata 234 can be obtained. Audio stream 414 can be decoded by media audio decoder 112 and reproduced to the user on media consumption device side 206.
[0440] The following aspects of this example are worth noting:
[0441] • Use cases:
[0442] • Audio data in one or more audio streams (e.g., a main audio stream 116)
[0443] It is transmitted in the auxiliary audio stream 316.
[0444] • Server-side 202 Not Transmitted (one or more) Earcon.
[0445] • Server-side 202 error: Earcon metadata not transmitted.
[0446] This use case can be represented for creating [something] without Earcons.
[0447] A solution for enabling Earcons for legacy content.
[0448] • In one client-side implementation, there is an ROI processor 120 and a metadata processor.
[0449] 232 is used to efficiently process Earcon information.
[0450] • The ROI processor 120 can be sourced from the media consumption device side 206 (e.g., HMD).
[0451] The device used receives information about the current viewport (user orientation information).
[0452] (Information). The ROI processor 210 can also receive information about metadata and
[0453] ROIs signaled in metadata (e.g., signaled via OMAF)
[0454] (Video viewport).
[0455] Based on this information, the ROI processor 120 can decide to activate the audio stream.
[0456] One (or more) Earcons that do not exist in (116, 316).
[0457] In addition, the ROI processor 120 can determine the location of Earcons.
[0458] Information about the gain value is provided to the Earcon metadata generator 432.
[0459] • The ROI processor 120 can provide this information to the metadata processor 232.
[0460] • Metadata processor 232 can parse metadata contained in the Earcon audio stream (if it exists), and can:
[0461] • Enable Earcon
[0462] Furthermore, if requested by the ROI processor 120, the metadata processor 132 accordingly modifies the spatial location and gain information contained in the Earcon metadata.
[0463] The metadata processor can also parse the audio metadata 236 and 417 of all audio streams 116 and 316, and can manipulate audio-specific information so that Earcon can be used as part of an audio scene (for example, if an audio scene has a 5.1 channel bed and 4 objects, the Earcon audio element is added to the scene as a fifth object). All metadata fields are updated accordingly.
[0464] The modified Earcon metadata 234 and information from the ROI processor 120 are provided to the synthesized audio generator 246. The synthesized audio generator 246 can create synthesized sounds based on the received information (e.g., generating a speech signal spelled out based on the spatial location of the Earcon). Furthermore, the Earcon metadata is associated with the generated audio data in a new stream.
[0465] Similarly, as before, the audio data of each stream, along with the modified audio metadata and Earcon metadata, is then provided to the stream multiplexer or multiplexer 412, which can then generate an audio stream 414 with a set of metadata (audio and Earcon).
[0466] • This stream 414 is decoded by a single media audio decoder based on user location information.
[0467] Alternatively, the Earcon audio data can be cached on the client (e.g., based on previous Earcon usage).
[0468] Alternatively, the output of the synthesized audio generator can be uncompressed audio, which can be mixed into the final rendered scene.
[0469] Furthermore, in a streaming environment, based on the same information, the ROI processor 120 can decide to request an Earcon stream in advance (e.g., when a user is looking in the wrong direction a few seconds before the ROI is enabled).
[0470] 6.8 Examples Based on User Location
[0471] It can enable the reproduction of the Earcon only when the user cannot see the ROI.
[0472] For example, the ROI processor 120 can periodically check the user's current viewport and / or position and / or head orientation and / or motion data 122. If the ROI is visible to the user, it will not trigger the reproduction of the earcon.
[0473] If the ROI processor determines from the user's current viewport and / or position and / or head orientation and / or motion data that the ROI is not visible to the user, the ROI processor 120 may request the reproduction of the Earth. In this case, the ROI processor 120 may cause the metadata processor 132 to prepare for the reproduction of the Earth. The metadata processor 132 may use one of the techniques described for the above examples. For example, the metadata may be retrieved in a stream passed by the server side 202, may be generated by the Earth metadata generator 432, and so on. The attributes of the Earth metadata may be easily modified based on the ROI processor's request and / or various conditions. For example, if the user's selection has previously disabled the Earth, the Earth will not be reproduced even if the user does not see the ROI. For example, if a (previously set) timer has not yet expired, the Earth will not be reproduced even if the user does not see the ROI.
[0474] Additionally, if the ROI processor determines that the ROI is visible to the user based on the user's current viewport and / or position and / or head orientation and / or motion data, the ROI processor 120 may request that the Earcon not be reproduced, especially if the Earcon metadata already contains signaling for the active Earcon.
[0475] In this scenario, ROI processor 120 can disable the reproduction of the Earcon by metadata processor 132. Metadata processor 132 can use one of the techniques described in the examples above. For example, the metadata can be retrieved from a stream passed by server-side 202, can be generated by Earcon metadata generator 432, and so on. The attributes of the Earcon metadata can be easily modified based on requests from the ROI processor and / or various conditions. If the metadata already contains an indication corresponding to the reproduction of the Earcon, then in this case, the metadata is modified to indicate that the Earcon is not active and should not be reproduced.
[0476] The following aspects of this example are worth noting:
[0477] • Use cases:
[0478] • In one or more audio streams 116, 316 (e.g., a mainstream and secondary)
[0479] Audio data is transmitted in one or more audio streams (116, ...).
[0480] 316 or in one or more additional streams 140 (depending on or independent of the main stream)
[0481] (Audio stream) is transmitted via Earcon.
[0482] • The Earcon metadata is set to indicate the Earcon
[0483] The clock is active at a specific time.
[0484] First-generation devices, excluding the ROI processor, will read Earcon metadata and
[0485] This allows for the reproduction of the Earcon, regardless of the user's current viewport.
[0486] And / or location and / or head orientation and / or motion data indicate the ROI for the user
[0487] visible.
[0488] • Next-generation devices, including any system-specific ROI processors, will utilize
[0489] The ROI processor determines this. This is based on the user's current viewport and / or...
[0490] The ROI processor determines the location and / or head orientation and / or motion data.
[0491] If the ROI is visible to the user, then the ROI processor 120 can request not to reproduce it.
[0492] Earcon, especially if Earcon metadata already contains information about the activity.
[0493] Earcon signaling. In this case, ROI processor 120 can...
[0494] Metadata processor 132 disables Earcon reproduction. Metadata processor 132
[0495] One of the techniques described in the examples above can be used. For example, metadata.
[0496] It can be retrieved in the stream transmitted by the server-side 202, and can be...
[0497] Earcon metadata generator 432 generates, etc. Earcon metadata...
[0498] Attributes can be customized based on ROI processor requests and / or various conditions.
[0499] Modify in a different location. If the metadata already contains instructions for reproducing the Earth,
[0500] In this case, modify the metadata to indicate that Earcon is not activated and
[0501] It should not be reproduced.
[0502] Furthermore, depending on the playback device, the ROI processor can determine the modification request.
[0503] Earcon metadata. For example, if reproduced through headphones or speakers.
[0504] Sound, on the other hand, can modify the Earcon spatial information differently.
[0505] Therefore, the final audio scene of the user experience will be obtained based on the metadata modifications performed by the metadata processor.
[0506] 6.9 Example of server-client communication (Figure 5b)
[0507] Figure 5b illustrates system 550, which on the client side 204 includes system 552 (client system) that can embody, for example, system 100, 200, 300, 400, or 500. Here, it is also possible to generalize the concept to any audio information message, also referring to Earcons.
[0508] System 552 may include ROI processor 120, metadata processor 132, and stream multiplexer or multiplexer 412. (In the example, different audio streams are decoded (each audio stream is decoded separately by the corresponding media audio decoder 112), and then the different audio streams are mixed together and / or rendered together to provide the final audio scene.)
[0509] Here, at least one audio stream is represented as including two audio streams 116 and 316 (other examples may provide a single stream as shown in Figure 2, or two or more streams). These are audio streams designed to reproduce the audio scene that the user expects to experience.
[0510] Additionally, the media encoder 240 can provide an Earcon stream 140.
[0511] Audio streams can be encoded at different bitrates, which allows for effective bitrate adaptation depending on the network connection (i.e., a high-bitrate encoded version is delivered to users with high-speed connections, while a low-bitrate version is delivered to users with lower-speed network connections).
[0512] The audio stream can be stored on a media server 554, where for each audio stream, different encodings at different bit rates are grouped into an adaptive set 556, and appropriate data is signaled to all created adaptive sets regarding their availability. Audio adaptive set 556 and video adaptive set 557 can be provided.
[0513] Based on the ROI and / or other criteria indicated in the user's motion and viewport metadata 131, the ROI processor 120 will enable the reproduction of the Earcon from the Earcon stream 140 (also indicated as an additional audio stream in addition to audio streams 116 and 316).
[0514] In this example:
[0515] Client 552 is configured to receive data from the server regarding the availability of all adaptive sets, including:
[0516] o at least one audio scene adaptive set of at least one audio stream; and
[0517] o An adaptive set of at least one audio message, including at least one additional audio stream and at least one audio information message.
[0518] Similar to other example implementations, the ROI processor 120 can receive information about the current viewport (user orientation information) from the media consumption device side 206 (e.g., HMD-based) used for content consumption. The ROI processor 120 can also receive information about metadata and ROIs signaled in the metadata (e.g., signaling the video viewport via OMAF).
[0519] Based on this information, the ROI processor 120 can decide to activate one (or more) of the Earcon contained in the Earcon audio stream 140.
[0520] Additionally, the ROI processor 120 can determine different locations of the Earcon and different gain values (e.g., to more accurately represent the Earcon in the current space where the content is consumed).
[0521] oROI processor 120 can provide this information to selection data generator 558.
[0522] • The selection data generator 558 can be configured to: based on the ROI processor's decision, create selection data 559 that identifies which adaptive sets to receive; the adaptive sets include audio scene adaptive sets and audio message adaptive sets;
[0523] Media server 554 can be configured to provide instruction data to client 552, enabling the streaming client to retrieve data from adaptive sets 556 and 557 identified by selection data, which identifies which adaptive sets to receive; the adaptive sets include audio scene adaptive sets and audio message adaptive sets.
[0524] The download and switching module 560 is configured to receive the requested audio stream from the media server 554 based on selection data identifying which adaptive sets to receive; the adaptive sets include audio scene adaptive sets and audio message adaptive sets. The download and switching module 560 can also be further configured to provide audio metadata and Earcon metadata 141 to the metadata processor 132.
[0525] • The ROI processor 120 can provide this information to the metadata processor 132.
[0526] • Metadata processor 132 can parse the metadata contained in the Earcon audio stream 140, and
[0527] Enable Earcon (so that its reproduction can be allowed).
[0528] Furthermore, if requested by the ROI processor 120, the metadata processor 132 accordingly modifies the spatial location and gain information contained in the Earcon metadata 141.
[0529] The metadata processor 132 can also parse the audio metadata of all audio streams 116 and 316, and can manipulate audio-specific information so that the Earcon can be used as part of an audio scene (for example, if the audio scene has a 5.1 channel bed and 4 objects, the Earcon audio element is added to the scene as a fifth object. All metadata fields can be updated accordingly).
[0530] Then, the audio data of each audio stream 116, 316, along with the modified audio metadata and Earcon metadata, can be provided to a stream multiplexer or multiplexer, which can then generate an audio stream 414 with a set of metadata (modified audio metadata 238 and modified Earcon metadata 234).
[0531] This stream can be decoded by a single media audio decoder 112 based on user location information.
[0532] An adaptive set can be formed by a set of representations containing interchangeable versions of various contents, such as different audio bitrates (e.g., different streams at different bitrates). While a single representation is theoretically sufficient to provide a playable stream, multiple representations can give clients the possibility to adapt the media stream to their current network conditions and bandwidth requirements, thus ensuring smoother playback.
[0533] 6.10 Method
[0534] All the examples above can be implemented using method steps. Here, for completeness, method 700 is described (method 700 can be executed by any of the examples above). This method may include:
[0535] In step 702, at least one video stream (106) and at least one first audio stream (116, 316) are received;
[0536] In step 704, at least one video signal is decoded from at least one video stream (106) to represent a VR, AR, MR, or 360-degree video environment scene to the user (118a); and
[0537] In step 706, at least one audio signal is decoded from at least one first audio stream (116, 316) to represent an audio scene to the user (118b);
[0538] Receive the user's current viewport and / or position and / or head orientation and / or motion data (122); and
[0539] In step 708, viewport metadata (131) associated with at least one video signal is received from at least one video stream (106), the viewport metadata defining at least one ROI; and
[0540] In step 710, based on the user's current viewport and / or position and / or head orientation and / or motion data (122) and viewport metadata and / or other criteria, it is determined whether to reproduce the audio information message associated with at least one ROI; and
[0541] At step 712, audio information message metadata (141) describing the audio information message is received, processed, and / or manipulated, such that the audio information message is reproduced according to the audio information message attributes, so that the audio information message is part of the audio scene.
[0542] Obviously, the order can also be changed. For example, depending on the actual order in which the information is transmitted, the receiving steps 702, 706, and 708 can have different orders.
[0543] Line 714 refers to the fact that this method can be repeated. If the ROI processor decides not to reproduce the audio information message, step 712 can be skipped.
[0544] 6.11 Other Implementations
[0545] Figure 8 System 800 is illustrated, which may implement one of the systems (or components thereof) or execution method 700. System 800 may include processor 802 and a non-transitory storage unit 806 for storing instructions, which, when executed by processor 802, enable the processor to perform at least the stream processing operations and / or the metadata processing operations discussed above. System 800 may include an input / output unit 804 for connection to external devices.
[0546] System 800 can implement at least some (or all) of the functions of ROI processor 120, metadata processor 232, synthesized audio generator 246, multiplexer or multiplexer 412, decoder 112m, Earcon metadata generator 432, etc.
[0547] Depending on certain implementation requirements, an example can be implemented in hardware. This implementation can be performed using digital storage media, such as floppy disks, digital versatile disks (DVDs), Blu-ray discs, compact discs (CDs), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory, on which electronically readable control signals are stored (or are capable of) cooperating with a programmable computer system to perform the corresponding methods. Therefore, the digital storage medium can be computer-readable.
[0548] Typically, an example can be implemented as a computer program product with program code operable to perform one of the methods when the computer program product is run on a computer. Program instructions may, for example, be stored on a machine-readable medium.
[0549] Other examples include a computer program stored on a machine-readable medium that performs one of the methods described herein. In other words, method examples are therefore computer programs having program instructions for performing one of the methods described herein when the computer program is run on a computer.
[0550] Therefore, another example of the method is a data carrier (or digital storage medium or computer-readable medium) on which a computer program is recorded, which is used to perform one of the methods described herein. The data carrier medium, digital storage medium, or recording medium is a tangible and / or non-transitory signal, rather than an intangible and transient one.
[0551] Another example includes a processing unit, such as a computer or programmable logic device, which performs one of the methods described herein.
[0552] Another example includes a computer on which a computer program is installed to perform one of the methods described herein.
[0553] Another example includes an apparatus or system for transmitting a computer program to a receiver (e.g., electronically or optically) for performing one of the methods described herein. The receiver may be, for example, a computer, mobile device, storage device, etc. The apparatus or system may include, for example, a file server for transmitting the computer program to the receiver.
[0554] In some examples, programmable logic devices (e.g., field-programmable gate arrays) can be used to perform some or all of the functions described herein. In some examples, field-programmable gate arrays can cooperate with microprocessors to perform one of the methods described herein. Generally, the methods can be implemented by any suitable hardware device.
[0555] The examples above are illustrative of the principles disclosed herein. It should be understood that modifications and variations to the arrangements and details described herein will be readily apparent. Therefore, the scope is intended to be limited by the appended patent claims rather than by the specific details given by way of description and explanation of the examples herein.
Claims
1. A content consumption device system configured to receive at least one first audio stream associated with a scene to be reproduced, at least one audio signal being encoded in the audio stream, the content consumption device system including or configured to be connected to at least one media audio decoder, the at least one media audio decoder being used to decode the at least one audio signal from at least one first audio stream to represent the scene to a user. in, The content consumption device system includes a first processor, which is configured to: Based at least on the user's motion data and / or the user's selection and / or metadata, a decision is made as to whether to reproduce an audio information message, wherein the audio information message is uncompressed and independent of the at least one audio signal; as well as When deciding to reproduce the audio information message, the audio information message is reproduced.
2. The content consumption device system according to claim 1, further comprising: A metadata processor configured to receive and / or generate and / or process and / or modify audio information message metadata so as to reproduce the audio information message based on the audio information message metadata when it is determined that the audio information message should be reproduced.
3. The content consumption device system according to claim 2 is further configured to: generate the audio information message based on the audio information message metadata.
4. The content consumption device system according to claim 1, further comprising: An audio generator is used to generate the audio information message.
5. The content consumption device system according to claim 1, further comprising: A mixer for mixing the audio information message with the at least one audio signal.
6. The content consumption device system according to claim 1, wherein, The audio information message is associated with an accessibility feature, which is associated with an object in the scene.
7. The content consumption device system according to claim 1, further comprising: A multiplexer or multiplexer is used to combine the audio information message with the at least one audio signal.
8. The content consumption device system according to claim 1, wherein, The audio information message is Earcon.
9. The content consumption device system according to claim 2, wherein, The metadata processor is configured to perform at least one of the following operations: Embed metadata in the stream; Extract audio metadata from the at least one first audio stream; Modify the audio information message metadata to activate the audio information message and / or set / change the position of the audio information message; as well as Modify the audio metadata of the at least one first audio stream to take into account the presence of the audio information message and allow merging.
10. The content consumption device system of claim 2 is further configured to store the audio information message metadata for future use.
11. The content consumption device system according to claim 2, wherein, The audio information message metadata includes at least one of the following: Identify labels, The type of the audio information message, state, Indicators of scene dependency / independence Location data, Indication of the presence of associated text tags. The number of available languages Data text length, and Description of the audio information message.
12. The content consumption device system according to claim 2, wherein, The audio information message metadata includes at least one gain data.
13. The content consumption device system according to claim 2, wherein, The audio information message metadata includes at least one of the languages of the audio information message.
14. The content consumption device system of claim 1 is further configured to perform at least one of the following operations: Receive at least one audio metadata describing at least one audio signal encoded in at least one first audio stream; Process the at least one audio information message associated with the at least one accessibility feature; Generate audio information message metadata; The audio information message is generated based on the audio information message metadata; as well as The at least one first audio stream is multiplexed, mixed, or combined with the audio information message into a single audio signal.
15. The content consumption device system according to claim 1, further configured as follows: Processing audio streams and audio information message metadata, the audio stream including at least one encoded audio signal and at least one uncompressed audio information message; and The at least one audio signal and the at least one audio information message are decoded and reproduced based on the audio information message metadata.
16. The content consumption device system of claim 1, configured to: perform a local search on the audio information message, and, if not found, request audio information message metadata from a remote entity.
17. The content consumption device system of claim 3, configured to: perform a local search on the audio information message, and, if not found, cause an audio generator to generate the audio information message.
18. The content consumption device system according to claim 1, wherein, The audio stream is formatted as MPEG-H 3D audio stream format.
19. The content consumption device system according to claim 1, further configured as follows: Receive data on the availability of multiple adaptive sets, the available adaptive sets including at least one audio scene adaptive set for the at least one first audio stream and at least one audio information message adaptive set for the audio information message; Based on the decision of the first processor, selection data is created, which identifies which adaptive sets to retrieve, and the available adaptive sets include at least one audio scene adaptive set and / or at least one audio information message adaptive set. as well as Request and / or retrieve data from the adaptive set identified by the selected data. Each adaptive set groups different codes for different bit rates.
20. The content consumption device system according to claim 1, configured to: select a first audio information message to be reproduced before a second audio information message from a plurality of audio information messages to be reproduced.
21. The content consumption device system of claim 1, wherein the cache memory stores audio information messages synthesized or received from a remote entity for reuse at different times.
22. The content consumption device system of claim 1, configured to: check at least one criterion for reproducing the audio information message, the criterion including user selection and / or user settings.
23. The content consumption device system according to claim 22 is configured to: check whether the user has deactivated the setting for reproducing audio information messages or system sounds.
24. The content consumption device system according to claim 1, further comprising: A metadata processor is configured to receive audio information message metadata, receive a request from the first processor to modify the audio information message metadata, and modify the audio information message metadata to the modified audio information message metadata according to the request from the first processor. as well as The first processor is further configured to reproduce the audio information message based on the modified audio information message metadata.
25. The content consumption device system of claim 1, further configured to: receive at least one video stream associated with the scene to be reproduced; and Also includes: At least one media video decoder is configured to decode at least one video signal from the at least one video stream to represent a scene to a user.
26. The content consumption device system according to claim 1, wherein, The audio information message is associated with a general audio message.
27. The content consumption device system according to claim 1, wherein, The at least one first processor is configured to determine whether to reproduce the audio information message based on an indication of accessibility features associated with an object in the scene or accessibility feature indication metadata.
28. The content consumption device system according to claim 1, configured to: generate uncompressed audio information messages.
29. A method for receiving at least one first audio stream associated with a scene to be reproduced, wherein at least one audio signal is encoded in the audio stream, the method comprising: Based at least on the user's motion data or the user's selection or metadata, a decision is made as to whether to reproduce an audio information message, wherein the audio information message is uncompressed and independent of the at least one audio signal; as well as When deciding to reproduce the audio information message, the audio information message is reproduced, wherein reproducing the audio information message includes: decoding the at least one audio signal from the at least one first audio stream or decoding the at least one audio signal from the at least one first audio stream to represent the scene to the user.
30. The method of claim 29, comprising: Processing or modifying the audio information message metadata so that reproduction includes: reproducing the audio information message based on the audio information message metadata.
31. A non-transitory storage unit, comprising instructions that, when executed by a processor, cause the processor to perform a method of receiving at least one first audio stream associated with a scene to be reproduced, wherein at least one audio signal is encoded in the audio stream, the method comprising: Based at least on the user's motion data or the user's selection or metadata, a decision is made as to whether to reproduce an audio information message, wherein the audio information message is uncompressed and independent of the at least one audio signal; as well as When deciding to reproduce the audio information message, the audio information message is reproduced, wherein reproducing the audio information message includes: decoding the at least one audio signal from the at least one first audio stream or decoding the at least one audio signal from the at least one first audio stream to represent the scene to the user.
Citation Information
Patent Citations
A method and an apparatus for processing an audio signal
KR1020160136716A
Generating and transmitting metadata for virtual reality
US20160381398A1