Optimizing audio delivery for virtual reality applications
The system addresses the challenge of managing audio in VR, AR, and MR by dynamically adjusting audio streams based on user-specific data, resulting in a more efficient and realistic audio experience.
Patent Information
- Application Number
- JP2025046112
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-10-12
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
AI Technical Summary
Current systems for virtual reality (VR), augmented reality (AR), and mixed reality (MR) struggle to efficiently manage audio content as users move within or transition between environments, leading to complex communication requirements and high bitrate demands.
A system that requests and receives audio streams based on user-specific data such as viewport, head orientation, movement, and interaction metadata, allowing for adaptive bitrate and quality adjustments to optimize audio delivery.
This approach enables a more realistic and efficient audio experience in VR, AR, and MR environments by dynamically adjusting audio streams based on user position and movement, reducing payload and improving user experience.
Smart Images

Figure 2025090843000001_ABST
Abstract
Description
Technical Field
[0001]
Background Art
[0002] In a virtual reality (VR) environment, or similarly an augmented reality (AR) or mixed reality (MR) or 360-degree video environment, a user can usually visualize the entire 360-degree content, for example, using a head-mounted display (HMD), and can listen through headphones (or similarly, speakers including correct rendering according to the position).
[0003] In a simple use case, the content is created such that only one audio / video scene (for example, a 360-degree video) is played at a certain moment. The audio / video scene has a fixed position (for example, a sphere with the user at the center), and the user cannot move within the scene and can only rotate the head in various directions (yaw, pitch, roll). In this case, different videos and audio are played (different viewports are displayed) based on the orientation of the user's head.
[0004] In the case of video, the video content is distributed for the entire 360-degree scene together with metadata (for example, stitch information, projection mapping, etc.) for describing the rendering process, and is selected based on the current user's viewport. However, in the case of audio, the content is the same throughout the scene. Based on the metadata, the audio content is adapted to the current user's viewport (for example, audio objects are rendered differently based on the viewport / user orientation information). Note that 360-degree content refers to any type of content composed of multiple viewing angles that a user can select (for example, by the orientation of the user's head or a remote control device).
[0005] In more complex scenarios, when the user moves within the VR scene or "jumps" from one scene to the next, the audio content may also change (for example, an audio source that was inaudible in one scene becomes audible in the next - "the door opens"). In existing systems, a complete audio scene can be encoded into one stream and, if necessary (depending on the main stream), into additional streams. Such systems are known as next - generation audio systems (e.g., MPEG - H 3D audio, etc.). Such use cases can include the following.
[0006] · Example 1: The user selects to enter a new room and the entire audio / video scene changes · Example 2: When the user moves within the VR scene and opens and passes through a door, it means that an audio transition from one scene to the next is required For the purpose of explaining this scenario, the concept of discrete viewpoints in space is introduced as discrete positions in the space (or VR environment) where various audio / video content is available.
[0007] A "straight - forward" solution is to provide a real - time encoder that changes the encoding (number of audio elements, spatial information, etc.) based on feedback from the playback device regarding the user's position / orientation. This solution implies very complex communication between the client and the server in a streaming environment, for example.
[0008] · The client (usually assumed to use only simple logic) requires an advanced mechanism not only to request various streams but also to convey complex information about the encoding details that enable appropriate content processing based on the user's position.
[0009] · Media servers typically have various streams pre - input (formatted in a specific format that enables "per - segment" delivery), and the main function of the server is to provide information about the available streams and perform delivery when requested. To enable scenarios that allow encoding based on feedback from playback devices, the media server requires a high - level communication link with multiple live media encoders and the ability to create on - the - fly all signaling information (e.g., media presentation descriptions) that can change in real - time.
[0010] Such a system can be imagined, but its complexity and computational requirements exceed the capabilities and features of currently available devices and systems, or those that will be developed in the next few decades.
[0011] Alternatively, it is also possible to always deliver content that represents a complete VR environment ("complete world"). This solves the problem, but it requires a huge bitrate that exceeds the capacity of the available communication links.
[0012] This is complex in a real - time environment, and alternative solutions have been proposed to enable this use case with low complexity using available systems.
[0013] 2. Terms and Definitions The following terms are used in this technical field.
[0014] · Audio elements: For example, audio objects, audio channels, scene - based audio (higher - order ambisonics - HOA), or audio signals that can be represented as any combination thereof.
[0015] ·Region of Interest (ROI): One area of video content (or a displayed or simulated environment) that a user is interested in at a given point in time. This is typically, for example, an area on a sphere, or a polygon selection from a 2D map. The ROI identifies a specific area for a particular purpose and defines the boundaries of the object under consideration.
[0016] ·User position information: Position information (e.g., x, y, z coordinates), orientation information (yaw, pitch, roll), direction of movement, speed of movement, etc.
[0017] ·Viewport: A part of the omnidirectional video that is currently being displayed and viewed by the user.
[0018] ·Viewpoint: The center point of the viewport.
[0019] ·360-degree video (also known as immersive video or omnidirectional video): In the context of this document, represents video content that includes multiple views (viewports) in multiple directions simultaneously. Such content can be created, for example, using an omnidirectional camera or a set of cameras. During playback, the viewer can control the viewing direction.
[0020] ·Media Presentation Description (MPD) is a syntax such as XML, for example, that contains information about media segments, their relationships, and the information necessary to select them.
[0021] ·An adaptation set contains a media stream or a set of media streams. In the simplest case, there is one adaptation set that includes all the audio and video of the content, but for bandwidth reduction, each stream can be split into different adaptation sets. A common case is to have one video adaptation set and multiple audio adaptation sets (one for each supported language). An adaptation set can also include subtitles or any metadata.
[0022] · Depending on the representation, the adaptation set can include the same content encoded in different ways. In most cases, the representation is provided at multiple bitrates. This allows the client to request the highest quality content that can be played without waiting for buffering. Since the representation can also be encoded with various codecs, it is possible to support clients with various supported codecs.
[0023] In the context of this application, the concept of an adaptation set is more commonly used and may actually refer to a representation. Also, a media stream (audio / video stream) is usually encapsulated in media segments, which are the actual media files that are initially played by a client (e.g., a DASH client). Various formats can be used for the media segments, such as the ISO Base Media File Format (ISOBMFF) similar to the MPEG-4 container format or the MPEG-2 Transport Stream (TS). Encapsulation into media segments and encapsulation in various representations / adaptation sets are irrelevant to the methods described here, and the present method applies to all such various options.
[0024] Furthermore, although the description of the method in this document focuses on the communication between the DASH server and the client, the present method is general enough to function in other delivery environments such as MMT, MPEG-2 TS, DASH-ROUTE, and file formats for file playback.
[0025] Generally, the adaptation set is at a higher layer with respect to the stream and can include metadata (e.g., associated with a location). The stream can include multiple audio elements. An audio scene can be associated with multiple streams distributed as part of multiple adaptation sets.
[0026] 3. Current Solutions The current solutions are as follows.
[0027] [1].ISO / IEC 23008-3:2015, Information technology--High efficiency coding and media delivery in heterogeneous environments--Part 3:3D audio
[0028] [2].N16950, Study of ISO / IEC DIS 23000-20 Omnidirectional Media Format The current solutions are limited and can provide an independent VR experience at one fixed location. Therefore, the user can change the orientation but cannot move within the VR environment.
Prior Art Documents
Non-Patent Documents
[0029]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
[0030] According to one embodiment, a system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments may be configured to receive video and audio streams for playback on a media consumption device. The system may include at least one media video decoder configured to decode a video signal from the video stream to represent a VR, AR, MR, or 360-degree video environment scene to a user, and at least one audio decoder configured to decode an audio signal from at least one audio stream. The system may be configured to request from a server at least one audio stream and / or one audio element of the audio stream and / or one adaptation set based at least on data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data.
[0031] According to one aspect, the system may be configured to provide to the server data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data to obtain from the server at least one audio stream and / or one audio element of the audio stream and / or one adaptation set.
[0032] In one embodiment, at least one scene is associated with at least one audio element, and each audio element is associated with a position and / or region within a visual environment where the audio element is audible. Different audio streams may be configured to be provided based on different user positions and / or viewports and / or head orientations and / or movement and / or interaction metadata and / or virtual location data within the scene. According to another aspect, the system may be configured to determine whether to play at least one audio element of an audio stream and / or one adaptation set with respect to the current user's viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position in a scene, and the system may be configured to request and / or receive at least one audio element at the current user's virtual position.
[0033] According to one aspect, the system may be configured to predictively determine whether at least one audio element of an audio stream and / or one adaptation set will be relevant and / or audible based on at least the current user's viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data, and the system may be configured to request and / or receive at least one audio element and / or an audio stream and / or an adaptation set at a specific user's virtual position prior to the predicted movement and / or interaction of the user in the scene, and the system may be configured to play at least one audio element and / or an audio stream at the specific user's virtual position after the movement and / or interaction of the user in the scene upon reception.
[0034] One embodiment of the system may be configured to request and / or receive at least one audio element at a lower bitrate and / or quality level at the user's virtual position before the movement and / or interaction of the user in the scene, and the system may be configured to request and / or receive at least one audio element at a higher bitrate and / or quality level at the user's virtual position after the movement and / or interaction of the user in the scene.
[0035] According to one aspect, the system may be configured such that at least one audio element is associated with at least one scene, and each audio element is associated with a position and / or region within the visual environment associated with the scene. The system may be configured to request and / or receive a stream at a higher bitrate and / or quality for audio elements closer to the user than for audio elements farther from the user.
[0036] According to one aspect of the system, at least one audio element is associated with at least one scene, and at least one audio element may be associated with a position and / or region within the visual environment associated with the scene. The system may be configured to request different streams at different bitrates and / or quality levels of the audio elements based on the relevance and / or auditing ability level at the virtual position of each user in the scene. The system may be configured to request an audio stream at a higher bitrate and / or quality level for an audio element that is more relevant and / or has higher audibility at the virtual position of the current user, and / or may be configured to request an audio stream at a lower bitrate and / or quality level for an audio element that is less relevant and / or has lower audibility at the virtual position of the current user.
[0037] In one embodiment of the system, at least one audio element may be associated with a scene, each audio element is associated with a position and / or region within the visual environment associated with the scene, and the system may be configured to periodically transmit data on the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual position data to the server, whereby at a first position, a higher bitrate and / or quality stream is provided by the server, and at a second position, a lower bitrate and / or quality stream is provided by the server, and the first position is closer to at least one audio element than the second position.
[0038] In one embodiment, the system may be defined for a plurality of visual environments such as an environment where a plurality of scenes are adjacent and / or proximate, and a first stream associated with a first current scene is provided, and when the user transitions to a second further scene, both the stream associated with the first scene and a second stream associated with the second scene are provided.
[0039] In one embodiment, the system may be defined for a plurality of scenes with respect to first and second visual environments, the first and second environments being adjacent and / or proximate environments, and a first stream associated with the first scene is provided by the server for playback of the first scene when the user's position or virtual position is in the first environment associated with the first scene, and a second stream associated with the second scene is provided by the server for playback of the second scene when the user's position or virtual position is in the second environment associated with the second scene, and when the user's position or virtual position is at a transition position between the first scene and the second scene, both the first stream associated with the first scene and the second stream associated with the second scene are provided.
[0040] In one embodiment, the system may be defined for first and second visual environments where a plurality of scenes are adjacent and / or proximate environments, and the system is configured to request and / or receive a first stream associated with a first scene associated with the first environment for playback of the first scene when the user's virtual position is in the first environment, and the system is configured to request and / or receive a second stream associated with a second scene associated with the second environment for playback of the second scene when the user's virtual position is in the second environment, and the system may be configured to request and / or receive both the first stream associated with the first scene and the second stream associated with the second scene when the user's virtual position is at a transition position between the first environment and the second environment.
[0041] According to one aspect, the system may be configured such that the first stream associated with the first scene is obtained at a higher bitrate and / or quality when the user is in the first environment associated with the first scene, while the second stream associated with the second scene associated with the second environment is obtained at a lower bitrate and / or quality when the user is at the start of the transition position from the first scene to the second scene, and the first stream associated with the first scene is obtained at a lower bitrate and / or quality when the user is at the end of the transition position from the first scene to the second scene, and the second stream associated with the second scene is obtained at a higher bitrate and / or quality, and the lower bitrate and / or quality is lower than the higher bitrate and / or quality.
[0042] According to one aspect, the system may be configured such that a plurality of scenes are defined for a plurality of environments such as adjacent and / or neighboring environments. The system may obtain a stream associated with a first current scene associated with a first current environment. When the distance of the user's position or virtual position from the scene boundary is less than a predetermined threshold, the system may further obtain an audio stream associated with a second adjacent and / or neighboring environment associated with a second scene.
[0043] According to one aspect, the system may be configured such that a plurality of scenes may be defined for a plurality of visual environments. The system requests and / or obtains a stream associated with the current scene at a higher bitrate and / or quality, and a stream associated with a second scene at a lower bitrate and / or quality, where the lower bitrate and / or quality is lower than the higher bitrate and / or quality.
[0044] According to one aspect, the system may be configured such that a plurality of N audio elements may be defined. When the distance of the user from the position or region of these audio elements is greater than a predetermined threshold, the N audio elements are processed to obtain a smaller number M (M < N) of audio elements associated with a position or region close to the position or region of the N audio elements. Thereby, when the distance of the user from the position or region of the N audio elements is less than a predetermined threshold, at least one audio stream associated with the N audio elements is provided to the system, or when the distance of the user from the position or region of the N audio elements is greater than a predetermined threshold, at least one audio stream associated with the M audio elements is provided to the system.
[0045] According to one aspect, the system may be configured such that at least one visual environment scene is associated with at least one plurality of N audio elements (N >= 2), each audio element being associated with a position and / or region within the visual environment. At least one plurality of N audio elements may be provided in at least one representation at a high bitrate and / or quality level, and at least one plurality of N audio elements may be provided in at least one representation at a low bitrate and / or quality level. At least one representation is obtained by processing the N audio elements to obtain a smaller number M (M < N) of audio elements associated with a position or region close to the position or region of the N audio elements. The system may be configured to request a representation at a higher bitrate and / or quality level for an audio element when the audio element is more relevant and / or more audible at the current virtual position of the user in the scene, and the system may be configured to request a representation at a lower bitrate and / or quality level for an audio element when the audio element is less relevant and / or less audible at the current virtual position of the user in the scene.
[0046] According to one aspect, the system may be configured such that different streams are obtained for different audio elements when the distance and / or relevance and / or audible level and / or angular orientation of the user is below a predetermined threshold.
[0047] In one embodiment, the system may be configured to request and / or obtain a stream based on the orientation of the user in the scene and / or the direction of movement of the user and / or the interaction of the user.
[0048] In one embodiment, the viewport of the system may be associated with position and / or virtual position and / or movement data and / or the head.
[0049] According to one aspect, the system may be configured such that different audio elements are provided in different viewports, and the system may be configured to request and / or receive a first audio element with a higher bitrate than a second audio element not within the viewport when the first audio element is within the viewport.
[0050] According to one aspect, the system may be configured to request and / or receive a first audio stream and a second audio stream, where the first audio element of the first audio stream is more relevant and / or more audible than the second audio element of the second audio stream, and the first audio stream is requested and / or received at a higher bitrate and / or quality than the bitrate and / or quality of the second audio stream.
[0051] According to one aspect, the system may be configured such that at least two visual environment scenes are defined, at least one first and second audio element is associated with a first scene associated with a first visual environment, at least one third audio element is associated with a second scene associated with a second visual environment, the system may be configured to obtain metadata describing that at least one second audio element is further associated with the second visual environment scene, the system may be configured to request and / or receive at least the first and second audio elements when the user's virtual position is in the first visual environment, the system may be configured to request and / or receive at least the second and third audio elements when the user's virtual position is in the second visual environment scene, and the system may be configured to request and / or receive at least the first, second, and third audio elements when the user's virtual position is transitioning between the first visual environment scene and the second visual environment scene.
[0052] One embodiment of the system may be configured such that at least one first audio element is provided with at least one audio stream and / or an adaptation set, at least one second audio element is provided with at least one second audio stream and / or an adaptation set, at least one third audio element is provided with at least one third audio stream and / or an adaptation set, at least a first visual environment scene is described by metadata as a complete scene that requires at least the first and second audio streams and / or an adaptation set, a second visual environment scene is described by metadata as an incomplete scene that requires at least the third audio stream and / or an adaptation set, and at least the second audio stream associated with the first visual environment scene, and the system includes a metadata processor configured to be able to merge, when the user's virtual position is in the second visual environment, the second audio stream belonging to the first visual environment and the third audio stream associated with the second visual environment into a new single stream by manipulating the metadata.
[0053] According to one aspect, the system includes a metadata processor configured to manipulate metadata within at least one audio stream in front of at least one audio decoder based on data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual position data.
[0054] According to one aspect, the metadata processor may be configured to enable and / or disable at least one audio element in at least one audio stream before at least one audio decoder based on data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data, and when the system determines that the audio element will no longer be played as a result of the current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual location data, the metadata processor may be configured to disable at least one audio element in at least one audio stream before at least one audio decoder, and when the system determines that the audio element will be played as a result of the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual location data, the metadata processor may be configured to enable at least one audio element in at least one audio stream before at least one audio decoder.
[0055] According to one aspect, the system may be configured to disable the decoding of audio elements selected based on data of the user's current viewport and / or head orientation and / or movement and / or metadata and / or virtual location.
[0056] According to one aspect, the system may be configured to merge at least one first audio stream associated with the current audio scene into at least one stream associated with an adjacent, proximate, and / or future audio scene.
[0057] According to one aspect, the system may be configured to obtain and / or collect statistical data or aggregated data regarding the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual location data, and send a request to a server associated with the statistical data or aggregated data.
[0058] According to one aspect, the system may be configured to deactivate the decoding and / or playback of at least one stream based on metadata associated with the at least one stream and based on the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual location data.
[0059] According to one aspect, the system may be configured to operate metadata associated with a selected group of audio streams based on at least the user's current or estimated viewport and / or head orientation and / or movement data and / or metadata and / or virtual location data, to select and / or enable and / or activate audio elements that constitute the audio scene to be played, and / or to enable the merging of all selected audio streams into a single audio stream.
[0060] According to one aspect, the system may be configured to control at least one stream request to the server based on the distance of the user's position from the boundaries of adjacent and / or proximate environments associated with different scenes, or other metrics associated with the user's position in the current environment or predictions in future environments.
[0061] According to one aspect of the system, information may be provided from the server system for each audio element or audio object, and the information includes descriptive information about the location where the sound scene or audio element is active.
[0062] According to one aspect, the system may be configured to select between playing one scene and synthesizing, mixing, multiplexing, overlaying, or combining at least two scenes based on current or future or data and / or metadata and / or virtual positions and / or user selections of the viewport and / or head orientation and / or movement, where the two scenes are associated with different adjacent and / or proximate environments.
[0063] According to one aspect, the system may be configured to create or use at least adaptation sets, where several adaptation sets are associated with one audio scene, and / or additional information is provided associating each adaptation set with one view point or one audio scene, and / or information regarding the boundaries of one audio scene, and / or information regarding the relationship between one adaptation set and one audio scene (e.g., the audio scene is encoded in three streams encapsulated in three adaptation sets), and / or additional information may be provided including information regarding the connection between the boundaries of the audio scene and the plurality of adaptation sets.
[0064] According to one aspect, the system may be configured to receive streams of scenes associated with adjacent or proximate environments and to start decoding and / or playing the streams of the adjacent or proximate environments upon detection of a transition at the boundary between the two environments.
[0065] According to one aspect, the system may be configured to operate as a client and as a server configured to deliver video and / or audio streams to be played on a media consumption device.
[0066] According to one aspect, the system requests and / or receives at least one first adaptation set including at least one audio stream associated with at least one first audio scene, requests and / or receives at least one second adaptation set including at least one second audio stream associated with at least two audio scenes including at least one first audio scene, and based on metadata available regarding the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual location data, and / or information describing the association of at least one first adaptation set to at least one first audio scene and / or the association of at least one second adaptation set to at least one first audio scene, it may be configured to merge at least one first audio stream and at least one second audio stream into a new decoded audio stream.
[0067] According to one aspect, the system may be configured to receive information regarding the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual location data, and / or information characterizing a change triggered by the user's action, and to receive information regarding the availability of an adaptation set and information describing the association of at least one adaptation set to at least one scene and / or viewport and / or viewport and / or location and / or virtual location and / or movement data and / or orientation.
[0068] According to one aspect, the system determines whether to play at least one audio element from at least one audio scene embedded in at least one stream and at least one additional audio element from at least one additional audio scene embedded in at least one additional stream, and in the case of an affirmative determination, is configured to perform an operation of merging or synthesizing or multiplexing or superimposing or combining at least one additional stream of the additional audio scene with at least one stream of at least one audio scene.
[0069] According to one aspect, the system manipulates audio metadata associated with a selected audio stream based on at least the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual location data to select and / or enable and / or activate the audio elements that constitute the audio scene determined to be played, and may be configured to enable merging of all selected audio streams into a single audio stream.
[0070] According to one aspect, a server may be provided for delivering audio and video streams for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment to a client, the video and audio streams being played on a media consumption device, the server may include an encoder for encoding and / or a storage device for storing a video stream that describes a visual environment, the visual environment being associated with an audio scene, the server may further include an encoder for encoding and / or a storage device for storing a plurality of streams and / or audio elements and / or adaptation sets to be delivered to the client, the streams and / or audio elements and / or adaptation sets being associated with at least one audio scene, the server selects and delivers a video stream based on a request from the client, the video stream being associated with an environment, and based on a request from the client, selects an audio stream and / or audio elements and / or an adaptation set, the request being based on at least data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data, and being associated with an audio scene associated with the environment, and is configured to deliver the audio stream to the client.
[0071] According to one aspect, a stream may be encapsulated in an adaptation set, each adaptation set including a plurality of streams associated with different representations at different bitrates and / or qualities of the same audio content, and the selected adaptation set is selected based on a request from the client.
[0072] According to one aspect, the system may operate as both a client and a server.
[0073] According to one aspect, the system may include a server.
[0074] According to one aspect, a method for a virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environment configured to receive video and audio streams to be played on a media consumption device (e.g., a playback device) may be provided, the method including decoding a video signal from a video stream for presentation to a user of a VR, AR, MR, or 360-degree video environment scene, decoding an audio signal from an audio stream, and requesting and / or obtaining at least one audio stream from a server based on the user's current viewport and / or position data and / or head orientation and / or movement data and / or metadata and / or virtual location data and / or metadata.
[0075] According to one aspect, a computer program including instructions that, when executed by a processor, cause the processor to perform the above method may be provided. BRIEF DESCRIPTION OF THE DRAWINGS
[0076]
Figure 1.1
Figure 1.2
Figure 1.3
Figure 1.4
Figure 1.5
Figure 1.6
Figure 1.7
Figure 1.8
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 8A
Figure 8B
Embodiments for Carrying out the Invention
[0077] Examples of systems according to aspects of the present invention are disclosed below in this specification (for example, after Figure 1.1).
[0078] Examples of the system of the present invention (which may be embodied by different examples disclosed below) are collectively denoted by reference numeral 102. The system 102 can obtain, for example, from an audio and / or video stream of a server system (for example, 120) for the representation of an audio scene and / or visual environment to a user, and thus may be a client system. The client system 102 may also receive, for example, metadata that provides side and / or auxiliary information regarding the audio and / or video stream from the server system 120.
[0079] The system 102 may be associated with (or in some examples may include) a media consumption device (MCD) that actually plays audio and / or video signals to the user. In some examples, the user may wear the MCD.
[0080] System 102 can execute requests to server system 120, and these requests are associated with data of the current viewport and / or head orientation (e.g., angular orientation) and / or movement and / or interaction metadata and / or virtual location data 110 of at least one user. (Some metrics may be provided). The data of the viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data 110 may be provided in the feedback from the MCD to client system 102, and based on this feedback, client system 102 may provide requests to server system 120.
[0081] In some cases, the request (indicated by reference numeral 112) may include the data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data 110 (or its displayed or processed version). Based on the data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data 110, server system 120 provides the required audio and / or video streams and / or metadata. In this case, server system 120 can have knowledge of the user's position (e.g., in a virtual environment) and can associate the correct stream with the user's position.
[0082] In other cases, the request 112 from the client system 102 can include an explicit request for a specific audio and / or video stream. In this case, the request 112 can be based on the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual location data 110. The client system 102 has knowledge of the audio and video signals that need to be rendered to the user even if the client system 102 does not store the required stream therein. The client system 102 can, for example, target a specific stream within the server system 120.
[0083] The client system 102 may be a system for virtual reality (VR), augmented reality (AR), mixed reality (MR), or 360-degree video environments configured to receive video and audio streams for playback on a media consumption device. The system 102 includes at least one media video decoder configured to decode video signals from a video stream to represent a VR, AR, MR, or 360-degree video environment scene to the user, and at least one audio decoder 104 configured to decode an audio signal (108) from at least one audio stream 106. The system 102 is configured to request 112 from the server 120 at least one audio stream 106 and / or one audio element of the audio stream and / or one adaptation set based at least on the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual location data 110.
[0084] In VR, AR, and MR environments, it should be noted that the user 140 may mean being in a specific environment (e.g., a specific room). The environment is described, for example, by a video signal encoded on the server side (not necessarily including the server system 120, but including another encoder that previously encoded the video stream stored in the storage of the server 120 on the side of the server system 120). At each instant, in some examples, the user can only enjoy some of the video signals (e.g., the viewport).
[0085] Generally, each environment may be associated with a specific audio scene. The audio scene can be understood as the collection of all sounds played to the user over a specific period in a specific environment.
[0086] Conventionally, environments have been understood in discrete numbers. Therefore, the number of environments has been understood to be finite. For the same reason, the number of audio scenes has been understood to be finite. Therefore, in the prior art, VR, AR, and MR systems are designed as follows. - The user always aims to be in one environment. Therefore, for each environment: o The client system 102 requests from the server system 120 only the video stream associated with a single environment.
[0087] o The client system 102 requests from the server system 120 only the audio stream associated with a single scene.
[0088] This approach has become inconvenient.
[0089] For example, all audio streams are collectively delivered to the client system 102 for each scene / environment. When the user moves to another environment, it is necessary to deliver a completely new audio stream (e.g., when the user passes through a door, it means the transfer of the environment / scene).
[0090] Furthermore, in some cases, an unnatural experience may occur. For example, when the user is close to a wall (such as a virtual wall in a virtual room), sounds should be heard from the other side of the wall. However, this experience is impossible in a conventional environment. The set of audio streams associated with the current scene clearly does not include the streams associated with adjacent environments / scenes.
[0091] On the other hand, increasing the bitrate of the audio stream usually improves the user experience. However, this may cause further problems. The higher the bitrate, the higher the payload that the server system needs to deliver to the client system 102. For example, if an audio scene contains multiple audio sources (transmitted as audio elements), some of them are close to the user's position and others are far from the user's position, the distant sound sources will be less audible. Therefore, if all audio elements are delivered at the same bitrate or quality level, the bitrate may become very high. This means inefficient audio stream delivery. If the server system 120 delivers the audio stream at the highest possible bitrate, inefficient delivery occurs because it requires as high a bitrate for related sounds generated near the user as for those with a low audible level or low relevance to the overall audio scene. Therefore, if all the audio streams of one scene are delivered at the highest bitrate, the communication between the server system 120 and the client system 102 will unnecessarily increase the payload. If all the audio streams of one scene are delivered at a lower bitrate, the user experience will not be satisfactory.
[0092] Communication problems exacerbate the inconveniences described above. When a user passes through a door, the environment / scene changes instantaneously, and the server system 120 needs to instantaneously provide all streams to the client system 102.
[0093] Consequently, the above problems could not be solved conventionally.
[0094] However, with the present invention, these problems can be solved. The client system 102 requests from the server system 120, which may be based on data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data (and not only based on the environment / scene). Thus, the server system 120 can provide, at each instant, for example, an audio stream rendered for each user location.
[0095] For example, when the user is not approaching a wall, the client system 102 does not need to request a stream of the adjacent environment (e.g., the client system 102 may request only when the user approaches a wall). Further, the stream coming from outside the wall may be audible at a low volume, so the bitrate may be reduced. In particular, more relevant streams (e.g., streams from audio objects within the current environment) are delivered from the server system 120 to the client system 102 at the highest bitrate and / or highest quality level (as a result, less relevant streams, due to their low bitrate and quality level, leave bandwidth available for more relevant streams).
[0096] Lower quality levels can be obtained, for example, by reducing the bitrate or processing the audio elements so that less data needs to be transmitted, while keeping the bitrate per audio signal constant. For example, if ten audio objects are all at various positions far from the user, these objects can be mixed into a smaller number of signals based on the user's position.
[0097] - At positions very far from the user's position (e.g., positions higher than a first threshold), the object is mixed into two signals (other numbers are possible based on spatial position and semantics) and delivered as two "virtual objects".
[0098] - At positions close to the user's position (e.g., lower than the first threshold but higher than a second threshold smaller than the first threshold), the object is mixed into five signals (based on their spatial position and semantics) and delivered as five (other numbers are possible) "virtual objects".
[0099] - At positions very close to the user's position (positions lower than the first and second thresholds), the ten objects are delivered as ten audio signals providing the highest quality.
[0100] Although all the highest-quality audio signals may be considered very important and audible, the user may still be able to identify each object individually. If the quality level at a far position is lower, some audio objects may become less relevant or inaudible, and thus the user may not be able to localize the audio signals in space individually. Therefore, even if the quality level for delivering these audio signals is reduced, the quality of the user experience will not degrade.
[0101] Another example is when the user crosses a door. At a transition location (e.g., the boundary between two different environments / scenes), the server system 120 provides streams for both scenes / environments, but at a lower bitrate. This is because the user experiences sound from two different environments (where sound may be merged from different audio streams originally associated with different scenes / environments), and the highest quality levels of each sound source (or audio element) are not required.
[0102] In view of the above, the present invention enables going beyond the conventional approach of discrete numbers of visual environments and audio scenes, but enables a gradual representation of different environments / scenes, giving the user a more realistic experience.
[0103] Hereinafter, each visual environment (e.g., virtual environment) is considered to be associated with an audio scene (the attributes of the environment may also be the attributes of the scene). Each environment / scene may be associated with, for example, a geometric coordinate system (which may be a virtual geometric coordinate system). Since there may be a boundary between environments / scenes, when the user's position (e.g., virtual position) crosses the boundary, another environment / scene is reached. The boundary may be based on the coordinate system used. The environment may include audio objects (audio elements, sound sources) that can be placed at some specific coordinates of the environment / scene. For example, with respect to the relative position and / or orientation of the user with respect to the audio object (audio element, sound source), the client system 102 can request different streams, and / or the server system 120 can provide different streams (e.g., at a higher / lower bitrate and / or quality level depending on the distance and / or direction).
[0104] More generally, client system 102 can request from, and / or obtain from, server system 120 different streams (e.g., different representations of the same sound at different bitrates and / or quality levels) based on audibility and / or relevance. Audibility and / or relevance may be determined based on, for example, at least data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data.
[0105] In some examples, it may be possible to merge different streams. In some cases, it may be possible to synthesize, mix, multiplex, overlay, or combine at least two scenes. For example, use a mixer and / or renderer (e.g., used downstream of multiple decoders, each decoding at least one audio stream), or perform a stream multiplexing operation, such as upstream of stream decoding. In other cases, it may be possible to decode various streams and render them with various speaker settings.
[0106] Note that the present invention does not necessarily reject the concepts of visual environments and audio scenes. In particular, in the present invention, audio and video streams associated with a particular scene / environment may be delivered from server system 120 to client system 102 when the user enters the environment / scene. Nevertheless, within the same environment / scene, different audio streams and / or audio objects and / or adaptation sets may be requested, addressed, and / or delivered. In particular, the following may be possible.
[0107] - At least a portion of the video data associated with the visual environment is delivered from server 120 to client 102 at the user's entrance to the scene, and / or - At least some of the audio data (stream, object, adaptation set, etc.) is delivered to the client system 102 only based on current (or future) viewport and / or head orientation and / or movement data and / or metadata and / or virtual position and / or user selection / interaction, and / or - Optionally, (regardless of current or future position, viewport and / or head orientation and / or movement data and / or metadata and / or virtual position and / or user selection), based on the current scene, some audio data is delivered to the client system 102, while the remaining audio data is delivered based on current or future viewport and / or head orientation and / or movement data and / or metadata and / or virtual position and / or user selection.
[0108] Note that various elements (server system, client system, MCD, etc.) can represent different hardware devices or elements of the same one (for example, the client and MCD can be implemented as part of the same mobile phone, or, similarly, the client can be placed on a PC connected to a secondary screen that constitutes the MCD).
[0109] Examples One embodiment of the system 102 (client) shown in FIG. 1.1 is configured to receive an (audio) stream 106 based on a defined position within an environment (e.g., a virtual environment) that can be understood to be associated with a video and audio scene (hereinafter referred to as scene 150). Different positions within the same scene 150 generally mean different streams 106 provided to the audio decoder 104 of the system 102 (e.g., from the media server 120) or different metadata associated with the stream 106. The system 102 is connected to a media consumer device (MCD) and receives feedback associated with the position and / or virtual position of the user in the same environment therefrom. Hereinafter, the position of the user within the environment may be associated with a particular viewport that the user enjoys (e.g., the viewport is intended to be a surface assumed as a rectangular surface projected onto a sphere and displayed to the user).
[0110] In an exemplary scenario, when a user moves within a VR, AR, and / or MR scene 150, it can be imagined that the audio content is virtually generated by one or more audio sources 152 that may change. The audio source 152 can be understood as a virtual audio source in the sense that it can refer to a position within the virtual environment. The rendering of each audio source is adapted to the user's position (e.g., in a simplified example, the level of the audio source is higher the closer the user is to the position of the audio source and lower the farther the user is from the audio source). Nevertheless, each audio element (audio source) is encoded into an audio stream provided to the decoder. The audio stream can be associated with various positions and / or regions within the scene. For example, an audio source 152 that is inaudible in one scene may become audible in the next scene when, for example, a door in the VR, AR, and / or MR scene 150 is opened. Next, the user can choose to enter a new scene / environment 150 (e.g., a room), and the entire audio scene changes. For the purpose of explaining this scenario, the term discrete view points of the space can be used as discrete positions in the space (or VR environment) where different audio content is available.
[0111] Generally speaking, media server 120 can provide stream 106 associated with a particular scene 150 based on the position of the user within scene 150. Stream 106 can be encoded by at least one encoder 154 and provided to media server 120. Media server 120 can transmit stream 113 using communication 113 (e.g., via a communication network). The provision of stream 113 may be based on a request 112 set by system 102 based on the position 110 of the user (e.g., in a virtual environment). The position 110 of the user can also be understood as being associated with a viewport (there is one single rectangle represented for each position) and a viewpoint (the viewpoint is the center of the viewport) that the user enjoys. Thus, in some examples, the provision of the viewport may be the same as the provision of the position.
[0112] The system 102 shown in FIG. 1.2 is configured to receive an (audio) stream 113 based on another configuration on the client side. In this exemplary embodiment, on the encoding side, a plurality of media encoders 154 are provided and one or more streams 106 can be created using them for each available scene 150 associated with one sound scene portion of one viewpoint.
[0113] Media server 120 can store a plurality of audio and video adaptation sets (not shown) including different encodings of the same audio and video streams at different bitrates. Further, the media server may include description information of all the adaptation sets, including the availability of all the created adaptation sets. The adaptation set may also include information describing the association of one adaptation set to one particular audio scene and / or viewpoint. In this way, each adaptation set can be associated with one of the available audio scenes.
[0114] The adaptation set may further include information describing the boundaries of each audio scene and / or viewpoint, which may include, for example, a complete audio scene or simply individual audio objects. The boundaries of one audio scene may be defined, for example, as the geometric coordinates of a sphere (e.g., center and radius).
[0115] The client-side system 102 can receive information regarding the current viewport and / or head orientation and / or movement data and / or interaction metadata and / or any information characterizing changes caused by the user's virtual position or the user's actions. Further, the system 102 can also receive information regarding the availability of all adaptation sets, as well as information describing the association of one adaptation set to one audio scene and / or viewpoint, and / or information describing the "boundaries" of each audio scene and / or viewpoint (which can include, for example, a complete audio scene or only individual objects). For example, such information can be provided as part of the Media Presentation Description (MPD) XML syntax in a DASH delivery environment.
[0116] The system 102 can provide an audio signal to a Media Consumption Device (MCD) used for content consumption. Also, the media consumption device serves to collect collection information regarding the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions) as position and transition data 110.
[0117] The viewport processor 1232 may be configured to receive position and transition data 110 from the media consumption device side. The viewport processor 1232 may also be able to receive information regarding the ROI signaled with metadata and all the information available at the receiving end (system 102). Next, the viewport processor 1232 can determine which audio viewport should be played at a particular instant based on all the information received and / or derived from the received and / or available metadata. For example, the viewport processor 1232 can determine to play one complete audio scene, and one new audio scene 108 has to be created from all the available audio scenes, for example, only some audio elements of a plurality of audio scenes are played, while the other remaining audio elements of these audio scenes are not played. The viewport processor 1232 can also determine whether it is necessary to play the transition between two or more audio scenes.
[0118] The selection part 1230 can be provided to select one or more adaptation sets from the available adaptation sets signaled with the information received by the receiving end based on the information received from the viewport processor 1232, and the selected adaptation set completely describes the audio scene to be played at the user's current location. This audio scene may be one complete audio scene defined on the encoding side, or it may be necessary to create a new audio scene from all the available audio scenes.
[0119] Furthermore, when a transition between two or more audio scenes is about to occur based on an instruction from the viewport processor 1232, the selection part can be configured to select one or more adaptation sets from the available adaptation sets signaled with the information received by the receiving end, and the selected adaptation set fully describes the audio scene that needs to be reproduced in the near future (for example, when the user walks at a specific speed in the direction of the next audio scene, the next audio scene is predicted to be needed and is selected prior to playback).
[0120] Furthermore, several adaptation sets corresponding to adjacent locations are first selected at a lower bitrate and / or a lower quality level. For example, a representation encoded at a lower bitrate is selected from the representations available in one adaptation set, and based on the change in position, the quality is improved by selecting a higher bitrate for those specific adaptation sets. For example, a representation encoded at a higher bitrate is selected from the representations available in one adaptation set.
[0121] Based on an instruction received from the selection part, a download and switching part 1234 may be provided to request one or more adaptation sets from the available adaptation sets from the media server, receive one or more adaptation sets from the available adaptation sets from the media server, and extract metadata information from all the received audio streams.
[0122] The metadata processor 1236 may be provided to receive information that can include audio metadata corresponding to each received audio stream from download and switching information about the received audio stream. The metadata processor 1236 may also process and manipulate the audio metadata associated with each audio stream 113 based on information received from the viewport processor 1232 that can include information regarding the user's position and / or orientation and / or direction of movement 110 to select / enable the audio elements 152 necessary to compose a new audio scene, as indicated by the viewport processor 1232, such that all the audio streams 113 can be merged into a single audio stream 106.
[0123] The stream muxer / merger 1238 may be configured to merge all the selected audio streams into one audio stream 106 based on information received from the metadata processor 1236 that can include the modified and processed audio metadata corresponding to all the received audio streams 113.
[0124] The media decoder 104 is configured to receive and decode at least one audio stream for playback of a new audio scene, as indicated by the viewport processor 1232, based on information regarding the user's position and / or orientation and / or direction of movement.
[0125] In another embodiment, the system 102 shown in FIG. 1.7 may be configured to receive audio streams 106 at different audio bitrates and / or quality levels. The hardware configuration of this embodiment is the same as that of FIG. 1.2. At least one visual environment scene 152 can be associated with at least one plurality of N audio elements (N>=2), and each audio element is associated with a position and / or region within the visual environment. At least one plurality of N audio elements 152 are provided in at least one representation at a high bitrate and / or quality level, at least one plurality of N audio elements 152 are provided in at least one representation at a low bitrate and / or quality level, and at least one representation is obtained by processing the N audio elements 152 to obtain a smaller number M (M<N) of audio elements 152 associated with a position or region close to the position or region of the N audio elements 152.
[0126] The processing of the N audio elements 152 may be, for example, a simple addition of audio signals, or an active downmix based on their spatial positions 110, or a rendering of the audio signals to a new virtual position located between the audio signals using their spatial positions. The system may be configured to request a representation at a higher bitrate and / or quality level for an audio element when the audio element is more relevant and / or more audible at the current virtual position of the user in the scene, and the system may be configured to request a representation at a lower bitrate and / or quality level for an audio element when the audio element is less relevant and / or less audible at the current virtual position of the user in the scene.
[0127] FIG. 1.8 shows an example of a system (which may be system 102), and shows system 102 for a virtual reality VR, augmented reality AR, mixed reality MR, or 360-degree video environment configured to receive video stream 1800 and audio stream 106 reproduced by a media consumption device, System 102 is, at least one media video decoder 1804 configured to decode video signal 1808 from video stream 1800 to represent a VR, AR, MR, or 360-degree video environment to a user, and at least one audio decoder 104 configured to decode audio signal 108 from at least one audio stream 106, may be included.
[0128] System 102 may be configured to request (112) at least one audio stream 106 and / or one audio element of the audio stream and / or one adaptation set from a server (eg 120) based on at least data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual position data 110 (eg provided as feedback from media consumption device 180).
[0129] System 102 may be the same as system 102 of FIGS. 1.1 to 1.7 and / or may obtain the scenarios from FIG. 2a onwards.
[0130] This example also refers to a method for a virtual reality VR, augmented reality AR, mixed reality MR, or 360-degree video environment configured to receive a video stream and an audio stream reproduced by a media consumption device [eg, a playback device], and this method includes decoding a video signal from a video stream for presentation of a VR, AR, MR, or 360-degree video environment scene to a user, and Decoding an audio signal from an audio stream; Requesting and / or obtaining at least one audio stream from a server based on the user's current viewport and / or position data and / or head orientation and / or movement data and / or metadata and / or virtual position data and / or metadata;
[0131] Case 1 Different scenes / environments 150 generally mean receiving different streams 106 from server 120. However, the stream 106 received by audio decoder 104 may also be conditioned by the user's position within the same scene 150.
[0132] At a first (starting) time point (t = t1) shown in FIG. 2a, the user is, for example, located within scene 150 and has a first defined position within a VR environment (or AR environment, or MR environment). In a Cartesian XYZ coordinate system (e.g., horizontal, etc.), the user's first viewport (position) 110’ is associated with coordinates x’ u and y’ u (the Z axis is oriented out of the paper here). In this first scene 150, two audio elements 152-1 and 152-1 are arranged, having coordinates x’1 and y’1 of audio element 1 (152-1), and x’2 and y’2 of audio element 2 (152-2), respectively. The distance d’1 from the user to audio element 1 (152-1) is smaller than the distance d’2 (152-1) from the user to audio element 2. All user position (viewport) data is transmitted from the MCD to system 102.
[0133] At a second exemplary time point (t = t2) shown in FIG. 2b, the user is, for example, within the same scene 150 but located at a second different position. In a Cartesian XY coordinate system, the user's second viewport (position) 110” is associated with new coordinates x” u and y”u is associated with (axis Z is directed out of the paper here). Here, the user's distance d”1 from the audio element 1 (152-1) is greater than the user's distance d”2 from the audio element 2 (152-2). All user position (viewport) data is sent back from the MCD to the system 102.
[0134] The user equipped with the MCD for visualizing a specific viewport within the 360-degree environment may, for example, be listening via headphones. The user can enjoy the reproduction of different sounds for different positions shown in FIGS. 2a and 2b of the same scene 150.
[0135] For example, data on any position and / or transition and / or viewport and / or virtual position and / or head orientation and / or movement within the scene from FIG. 2a to FIG. 2b can be periodically (e.g., via feedback) sent as signal 110 from the MCD to the system 102 (client). The client can resend the position and transition data 110’ or 110” (e.g., viewport data) to the server 120. The client 102 or the server 120 can determine the audio stream 106 necessary to reproduce the correct audio scene at the current user position based on the position and transition data 110’ or 110” (e.g., viewport data). The client can determine and send a request 112 for the corresponding audio stream 106, and the server 120 can be configured to appropriately distribute the stream 106 in response to the position information provided by the client (system 102). Alternatively, the server 120 may determine and distribute the stream 106 accordingly in response to the position information provided by the client (system 102).
[0136] The client (System 102) can request the transmission of a stream to be decoded to represent the scene 150. In some examples, System 102 can transmit information regarding the highest quality level reproduced by the MCD (in other examples, it is the server 120 that determines the quality level to be reproduced by the MCD based on the user's position within the scene). In response, the server 120 can select one of a number of representations associated with the represented audio scene and deliver at least one stream 106 according to the user's position 110' or 110". Thus, the client (System 102) may be configured to deliver the audio signal 108 to the user, for example, via the audio decoder 104, and reproduce the sound associated with the user's actual (valid) position 110' or 110" (the adaptation set 113 may be used. For example, different variations of the same stream at different bitrates may be used for different positions of the user).
[0137] The stream 106 (which may be pre-processed or generated on-the-fly) can be sent to the client (System 102) and configured for a number of viewpoints associated with a particular sound scene.
[0138] Note that different qualities (e.g., different bitrates) may be provided for different streams 106 according to a specific position of a user (e.g., 110’ or 110”) in a (for example, virtual) environment. For example, in the case of multiple audio sources 152-1 and 152-2, each audio source 152-1 and 152-2 may be associated with a specific position within the scene 150. The closer the user's position 110’ or 110’ is to the first audio source 152-1, the higher the required resolution and / or quality of the stream associated with the first audio source 152-2 becomes. This exemplary case can be applied to the audio element 1 (152-1) in FIG. 2a and the audio element 2 (152-2) in FIG. 2b. The farther the user's position 110 is from the second audio source 152-2, the lower the required resolution of the stream 106 associated with the second audio source 152-2 becomes. This exemplary case can be applied to the audio element 2 (152-2) in FIG. 2a and the audio element 1 (152-1) in FIG. 2b.
[0139] In fact, first, the closer audio source sounds at a higher level (and thus is provided at a higher bitrate), and second, the farther audio source can be made to sound at a lower level (enabling it to require a lower resolution).
[0140] Therefore, based on the position 110’ or 110” in the environment provided by the client 102, the server 120 can provide different streams 106 at different bitrates (or other qualities). Based on the fact that distant audio elements do not require a high quality level, the overall user quality experience is maintained even when delivered at a lower bitrate or quality level.
[0141] Therefore, different quality levels can be used for several audio elements at different user positions while maintaining the quality of the experience.
[0142] Without this solution, all streams 106 would be provided from the server 120 to the client at the highest bitrate, thereby increasing the payload of the communication channel from the server 120 to the client.
[0143] Case 2 Figure 3 (Case 2) shows an embodiment of another exemplary scenario (represented in the vertical plane XZ of the space XYZ, with the axis Y represented as going into the paper), where the user moves in the first VR, AR, and / or MR scene A (150A), opens a door, and walks through the door (transition 150AB), which means an audio transition from the first scene 150A at time t1 through a temporary position (150AB) at time t2 to the next (second) scene B (150B) at time t3.
[0144] At time t1, the user may be at the x-direction position x1 in the first VR, AR, and / or MR scene. At time t3, the user may be in a different second VR, AR, and / or MR scene B (150B) at position x3. At instant t2, the user may be at the transition position 150AB while opening and passing through the door (e.g., a virtual door). Thus, the transition means a transition of audio information from the first scene 150A to the second scene 150B.
[0145] In this situation, the user changes their position 110 from, for example, the first VR environment (characterized by the first view point (A) as shown in FIG. 1.1) to the second VR environment (characterized by the second view point (B) as shown in FIG. 1.1). In certain cases, for example, during the transition through a door at the x-direction position x2, some audio elements 152A and 152B may be present at both view points (positions A and B).
[0146] The user (equipped with an MCD) is changing the position 110 (x1 - x3) towards the door, which means that at the transition position x2, the audio element belongs to both the first scene 150A and the second scene 150B. The MCD sends the new position and transition data 110 to the client, and the client resends it to the media server 120. The user may be able to listen to the appropriate audio source defined by the intermediate position x2 between the first position x1 and the second position x3.
[0147] Any position and transition from the first position (x1) to the second position (x3) is sent periodically (e.g., continuously) from the MCD to the client. The client 102 can resend the position and transition data 110 (x1~x3) to the media server 120, and the media server 120 is configured to deliver one dedicated item such as a new set of the pre-processed stream 106 in the form of the actual adaptation set 113’ according to the received position and transition data 110 (x1~x3).
[0148] The media server 120 can select one of a number of representations associated with the aforementioned information, not only regarding the function of the MCD that displays the highest bitrate, but also regarding the position and transition data 110 (x1 - x3) of the user during movement from one position to another. (In this situation, an adaptation set can be used. The media server 120 can determine which adaptation set 113’ optimally represents the virtual transition of the user without interfering with the rendering ability of the MCD.) Therefore, the media server 120 can deliver the dedicated stream 106 according to the position transition (e.g., as the new adaptation set 113’). The client 102 may be configured to deliver the audio signal 108 to the user 140 accordingly, for example, via the media audio decoder 104.
[0149] Stream 106 (generated on-the-fly and / or pre-processed) can be transmitted to client 102 at an adaptation set 113’ realized periodically (e.g., continuously).
[0150] When the user walks through the door, server 120 can transmit both the stream 106 of the first scene 150A and the stream 106 of the second scene 150B. This is to give the user a realistic impression by mixing or multiplexing or composing or playing these streams 106 simultaneously. Thus, based on the user's position 110 (e.g., "the position corresponding to the door"), server 120 transmits different streams 106 to the client.
[0151] Even in this case, since different streams 106 are heard simultaneously, they can have different resolutions and may be transmitted from server 120 to the client at different resolutions. When the user completes the transition and is in the second (position) scene 150A (and when the door behind the user is closed), server 120 can reduce or refrain from transmitting the stream 106 of the first scene 150 (if server 120 has already provided the streams to client 102, client 102 can decide not to use them).
[0152] Case 3 FIG. 4 (Case 3) shows an embodiment with another exemplary scenario (represented by the vertical plane XZ of the space XYZ, where the axis Y is represented as going into the paper), which means the transition of audio from one first position at time t1 to a second position within the first scene 150A at time t2 as the user moves within the VR, AR, and / or MR scene 150A. The user at the first position may be far from the wall at a distance d1 from the wall at time t1 and may be close to the wall at a distance d2 from the wall at time t2. Here, d1 > d2. At distance d1, the user hears only the source 152A of scene 150A, but can also hear the source 152B of scene 150B beyond the wall.
[0153] When the user is at the second position (d2), the client 102 sends data regarding the user's position 110 (d2) to the server 120 and receives from the server 120 not only the audio stream 106 of the first scene 150A but also the audio stream 106 of the second scene 150B. For example, based on the metadata provided by the server 120, the client 102 plays the stream 106 of the second scene 150B (through the wall) at a low volume, for example, via the decoder 104.
[0154] Even in this case, the bitrate (quality) of the stream 106 of the second scene 150B may be low, and thus it is necessary to reduce the transmission payload from the server 120 to the client. In particular, the position 110 (d1, d2) of the client (and / or the viewport) defines the audio stream 106 provided by the server 120.
[0155] For example, the system 102 may be configured to obtain a stream associated with a first current scene (150A) associated with a first current environment, and the distance of the user's position or virtual position from the boundary of the scene (for example, corresponding to a wall) is less than a predetermined threshold (for example, d2 < d しきい値 ) In this case, the system 102 further obtains a second audio stream associated with a second scene (150B) associated with an adjacent and / or proximate environment.
[0156] Case 4 Figures 5a and 5b show an embodiment with another exemplary scenario (represented by the horizontal plane XY of the space XYZ, and the axis Z is represented as coming out of the paper), where the user is located in the same VR, AR, and / or MR scene 150 but is arranged at different instants at different distances from, for example, two audio elements.
[0157] At the first instant t = t1 shown in FIG. 5a, the user is, for example, placed at the first position. At this first position, the first audio element 1 (152-1) and the second audio element 2 (152-2) are respectively (e.g., substantially) placed at distances d1 and d2 from the user equipped with the MCD. In this case, both distances d1 and d2 may be greater than the defined threshold distance d しきい値 and thus the system 102 is configured to group both audio elements into a single virtual source 152-3. The position and properties (such as spatial extent) of the single virtual source can be calculated based on the positions of the original two sources in such a way as to best mimic the original sound field generated by the two sources (e.g., two well-localized point sources can be reproduced as a single source at the center of the distance between them). The user position data 110 (d1, d2) can be transmitted from the MCD to the system 102 (client) and subsequently to the server 120, and the server 120 can determine to transmit an appropriate audio stream 106 to be rendered by the server system 120 (in other embodiments, it is the client 102 that determines the stream to be transmitted from the server 120). By grouping both audio elements into a single virtual source 152-3, the server 120 can select one of a number of representations associated with the aforementioned information. (For example, it is possible to deliver a dedicated stream 106 accordingly, and an adaptation set 113' associated with, for example, one single channel accordingly.) Thus, the user can receive, via the MCD, the audio signal transmitted from a single virtual audio element 152-3 placed between the actual audio elements 1 (152-1) and 2 (152-2).
[0158] At the second instant t = t2 shown in FIG. 5b, the user is, for example, located within the same scene 150 and has a second defined position in the same VR environment as in FIG. 5a. At this second position, two audio elements 152-1 and 152-2 are arranged (e.g., substantially) at distances d3 and d4 from the user, respectively. Both distances d3 and d4 may be shorter than the threshold distance d しきい値 and thus the grouping of the audio elements 152-1 and 152-2 into a single virtual source 152-3 is no longer used. The user position data is transmitted from the MCD to the system 102 and subsequently to the server 120, which can determine to send another appropriate audio stream 106 to be rendered by the system server 120 (in other embodiments, this determination is made by the client 102). By avoiding grouping the audio elements, the server 120 can select different representations associated with the aforementioned information and accordingly distribute a dedicated stream 106 with an adaptation set 113' associated with different channels for each audio element. As a result, the user can receive the audio signals 108 transmitted from two different audio elements 1 (152-1) and 2 (152-2) via the MCD. Therefore, the closer the user's position 110 is to the audio sources 1 (152-1) and 2 (152-2), the higher the required quality level of the streams associated with the audio sources needs to be selected.
[0159] In fact, as shown in FIG. 5b, the closer the audio sources 1 (152-1) and 2 (152-2) are to the user, the higher the level needs to be adjusted, so the audio signals 108 are rendered at a higher quality level. In contrast, the remotely located audio sources 1 and 2 represented in FIG. 5b need to be heard at a lower level when reproduced by a single virtual source and are thus rendered, for example, at a lower quality level.
[0160] In a similar configuration, a number of audio elements are placed in front of the user, and all of them are placed at a distance greater than the threshold distance from the user. In one embodiment, two groups of five audio elements may each be coupled to two virtual sources. The user position data is transmitted from the MCD to the system 102 and subsequently to the server 120, and the server 120 can determine to transmit an appropriate audio stream 106 to be rendered by the system server 120. By grouping all ten audio elements into only two single virtual sources, the server 120 can select one of a number of representations associated with the aforementioned information and accordingly distribute a dedicated stream 106 with, for example, an adaptation set 113' associated with two single audio elements. As a result, the user can receive, via the MCD, an audio signal transmitted from two different virtual audio elements placed in the same placement area as the actual audio elements.
[0161] At a subsequent instant, the user is approaching a number (ten) of audio elements. In this subsequent scene, all the audio elements are at a threshold distance d しきい値Since it is arranged at a smaller distance, the system 102 is configured to end the grouping of audio elements. New user position data is sent from the MCD to the system 102 and subsequently to the server 120, which can decide to send another appropriate audio stream 106 to be rendered by the server system 120. By not grouping the audio elements, the server 120 can select different representations associated with the aforementioned information and accordingly distribute a dedicated stream 106 with an adaptation set 113' associated with different channels for each audio element. As a result, the user can receive an audio signal transmitted from 10 different audio elements via the MCD. Therefore, the closer the user's position 110 is to the audio source, the higher the required resolution of the stream associated with the audio source needs to be selected.
[0162] Case 5 Figure 6 (Case 5) shows a user 140 at one position in a single scene 150 wearing a media consumer device (MCD) that can be directed in three exemplary different directions, each associated with a different viewport 160-1, 160-2, 160-3. These directions shown in Figure 6 may have directions (e.g., angular directions) in a polar coordinate system and / or a Cartesian XY coordinate system that point to a first viewport 801 at, for example, 180° at the bottom of Figure 6, a second viewport 802 located at, for example, 90° on the right side of Figure 6, and a third viewport 803 located at, for example, 0° at the top of Figure 6. Each of these viewports is associated with the orientation of the user 140 wearing the media consumer device (MCD), and the user located in the center is provided with a specific viewport displayed by the MCD that renders the corresponding audio signal 108 according to the orientation of the MCD.
[0163] In this particular VR environment, the first audio element s1(152) is located in the first viewport 160-1, which is near the viewpoint located, for example, at 180°, and the second audio element s2(152) is located in the third viewport 160-3, which is near the viewpoint located, for example, at 180°. Before changing the user's orientation, the user 140 experiences that the sound associated with the user's actual (effective) position is louder from the audio element s1 than from the audio element s2 in the first orientation towards the viewpoint 801 (viewport 160-1).
[0164] By changing the user's orientation, the user 140 experiences that the sound associated with the user's actual position 110 comes from the side at approximately the same volume from both audio elements s1 and s2 in the second orientation towards the viewpoint 802.
[0165] Finally, by changing the user's orientation, the user 140 can experience the sound associated with the audio element 2 as louder than the sound associated with the audio element s1 in the third orientation towards the viewpoint 801 (viewport 160-3) (in fact, the sound from the audio element 2 arrives from the front and the sound from the audio element 1 arrives from the back).
[0166] Therefore, different viewports and / or orientations and / or virtual position data can be associated with different bitrates and / or qualities.
[0167] Other Cases and Examples FIG. 7A shows an embodiment of a method for receiving an audio stream by a system in the form of a series of operation steps in the figure. At any given moment, the user of system 102 is associated with data and / or interaction metadata and / or virtual position of the user's current viewport and / or head orientation and / or movement. At a particular moment, the system can determine, in step 701 of FIG. 7A, the audio elements to be reproduced based on the data and / or interaction metadata and / or virtual position of the current viewport and / or head orientation and / or movement. Thus, in the next step 703, the relevance and audibility levels of each audio element can be determined. As described above with reference to FIG. 6, the VR environment can have different audio elements placed near or further away from the user within a particular scene 150, but may have a particular orientation within the 360-degree surroundings. All of these factors determine the relevance and audibility levels of each audio element.
[0168] In the next step 705, system 102 can request an audio stream according to the relevance and audible levels determined for each of the audio elements from media server 120.
[0169] In the next step 707, system 102 can receive an audio stream 113 appropriately prepared by media server 120, and the streams of different bitrates can reflect the relevance and audible levels determined in the previous steps.
[0170] In the next step 709, system 102 (e.g., an audio decoder) can decode the received audio stream 113, whereby, in step 711, a particular scene 150 is reproduced (e.g., by MCD) according to the data and / or interaction metadata and / or virtual position of the current viewport and / or head orientation and / or movement.
[0171] Figure 7B shows the interaction between the media server 120 and the system 102 according to the aforementioned series of operation diagrams. At a specific moment, the media server can transmit the audio stream 750 at a lower bitrate according to the lower relevance and audible level of the relevant audio elements of the aforementioned scene 150. The system can determine at a subsequent moment 752 that a change in interaction or position data has occurred. Such an interaction can result, for example, from a change in position data in the same scene 150 or, for example, from activating the door handle while the user attempts to enter a second scene separated from the first scene by a door provided by the door handle.
[0172] The current viewport and / or head orientation and / or movement data and / or interaction metadata and / or change in virtual position can result in a request 754 being sent by the system 102 to the media server 120. This request can reflect the higher relevance and audibility level of the relevant audio elements determined for the subsequent scene 150. In response to request 754, the media server transmits stream 756 at a higher bitrate, enabling a more plausible realistic playback of scene 150 at the current user's virtual position by the system 102.
[0173] Figure 8A also shows another embodiment of a method for receiving an audio stream by the system in the form of a series of operation steps in the figure. At a specific moment 801, a determination of the first current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position can be performed. By subtracting the affirmative cases, a request for a stream associated with a first position defined by a low bitrate can be prepared and transmitted by the system 102 at step 803.
[0174] A decision step 805 having three different results can be executed at a subsequent instant. One or two defined thresholds may be associated at this step to make a predictive determination regarding, for example, subsequent viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position. Thus, regarding the probability of change to a second position, a comparison with the first and / or second threshold can be performed, and as a result, for example, three different subsequent steps are executed.
[0175] For example, in a result reflecting a very low probability (e.g., associated with a comparison with the first predetermined threshold above), a new comparison step 801 is executed.
[0176] In a result reflecting a low probability (e.g., higher than the first predetermined threshold but in the example, higher than the first threshold and lower than the second predetermined threshold), a request for a low bitrate audio stream 113 can occur at step 809.
[0177] In a result reflecting a high probability (e.g., higher than the second predetermined threshold), at step 807, a request for a high bitrate audio stream 113 can be executed. Thus, a subsequent step executed after executing step 807 or 809 can also be the decision step 801.
[0178] Figure 8B shows the interaction between media server 120 and system 102 according to only one of the sequences of the aforementioned operation diagrams. At a particular moment, the media server can transmit an audio stream 850 at a low bitrate according to the aforementioned determined low relevance and audible level of the audio elements of the aforementioned scene 150. The system can determine at a subsequent moment 852 that an interaction will occur predictably. Current viewport and / or head orientation and / or movement data and / or interaction metadata and / or predicted changes in virtual position can result in an appropriate request 854 being sent from system 102 to media server 120. This request can reflect one of the above cases where it is likely to reach a second position associated with a high bitrate according to the audible level of the audio elements required for each subsequent scene 150. In response, the media server transmits a stream 856 at a higher bitrate, enabling a plausible and realistic playback of scene 150 at the current user's virtual position by system 102.
[0179] The system 102 shown in FIG. 1.3 is configured to receive an audio stream 113 based on another configuration on the client side, and the system architecture can use discrete view points based on a solution using a plurality of audio decoders 1320, 1322. On the client side, the system 102 can embody, for example, in addition or alternatively, the part of the system described in FIG. 1.2 that includes a plurality of audio decoders 1320, 1322, which can be configured to decode individual audio streams, as indicated by metadata processor 1236, for example, with some audio elements deactivated.
[0180] A mixer / renderer 1238 configured to play the final audio scene based on information regarding the position and / or orientation and / or direction of movement of the user may be provided in the system 102, i.e., for example, some audio elements that are not audible at that particular location are disabled or not rendered.
[0181] The following embodiments shown in FIGS. 1.4, 1.5 and 1.6 are based on an independent adaptation set for discrete viewpoints with a flexible adaptation set. When the user moves within the VR environment, the audio scene may change continuously. In order to ensure an excellent audio experience, it is necessary to make all audio elements constituting the audio scene at a particular point in time available to the media decoder, which can utilize the position information to create the final audio scene.
[0182] If the content is pre-encoded, the system can accurately play the audio scenes of these specific locations on the premise that at several predefined locations, these audio scenes do not overlap and the user can "jump / switch" from one location to the next.
[0183] However, when the user "walks" from one location to the next, the user can hear the audio elements of two (or more) audio scenes simultaneously. The solution to this use case does not rely on the mechanisms provided for decoding multiple audio streams (using either a maxer with a single media decoder or multiple media decoders with additional mixers / renderers), as provided in previous system examples, and it is necessary to provide the client with an audio stream that describes the complete audio scene.
[0184] Optimization is provided below by introducing the concept of common audio elements among multiple audio streams.
[0185] Description of Aspects and Examples Solution 1: Independent adaptation sets for discrete positions (viewpoints).
[0186] One way to solve the above problem is to use completely independent adaptation sets for each location. To better understand this solution, Figure 1.1 is used as a scenario example. In this example, three different individual viewpoints (composed of three different audio scenes) are used to create a complete VR environment in which the user can move. Therefore, · Some independent or overlapping audio scenes are encoded into several audio streams. For each audio scene, either one main stream can be used, or one main stream and additional auxiliary streams can be used depending on the use case (for example, some audio objects including different languages can be encoded into independent streams for efficient delivery). In the provided example, audio scene A is encoded into two streams (A1 and A2), audio scene B is encoded into three streams (B1, B2, and B3), and audio scene C is encoded into three streams (C1, C2, and C3). Note that audio scene A and audio scene B share some common elements (two audio objects in this example). Since all scenes need to be complete and independent (for example, for independent playback on non-VR playback devices), the common elements need to be encoded twice in each scene.
[0187] · Since all audio streams are encoded at different bitrates (i.e., different representations), efficient bitrate adaptation is possible according to the network connection (i.e., users using a high-speed connection are provided with a high-bitrate coded version, and users with a low-speed network connection are delivered a lower-bitrate version).
[0188] · The audio stream is stored in the media server, and for each audio stream, different encodings with different bitrates (i.e., different representations) are grouped into one adaptation set, and the availability of all adaptation sets for which appropriate data has been created is notified.
[0189] · Further, in addition to the adaptation set, the media server receives information regarding the position of the "boundary" of each audio scene and its relationship to each adaptation set (e.g., including only a complete audio scene or individual objects). In this way, each adaptation set can be associated with one of the available audio scenes. The boundary of one audio scene may be defined, for example, as the geometric coordinates of a sphere (e.g., center and radius).
[0190] o Each adaptation set also includes descriptive information regarding where the sound scene or audio element is active. For example, if one or more objects are included in one auxiliary stream, the adaptation set can include information such as the location where the object can be heard (e.g., the coordinates and radius of the center of the sphere).
[0191] · The media server provides the client (such as a DASH client) with information regarding the location of the "boundary" associated with each adaptation set. For example, in the case of a DASH delivery environment, this may be embedded in the Media Presentation Description (MPD) XML syntax.
[0192] · The client receives information regarding the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions).
[0193] · The client receives information regarding each adaptation set, and based on this and the user's location and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions, such as x, y, z coordinates or yaw, pitch, roll values), the client selects one or more adaptation sets that fully describe the audio scene being played at the user's current location.
[0194] · The client requests one or more adaptation sets o Further, the client selects more adaptation sets that fully describe multiple audio scenes, and uses the audio streams corresponding to the multiple audio scenes to create a new audio scene that needs to be played at the user's current location. For example, the user is walking within a VR environment and at a certain point is in between (or in a place where the effect of hearing two audio scenes is present).
[0195] o When the audio streams become available, multiple media decoders are used to decode the individual audio streams, and an additional mixer / renderer 1238 is used to play the final audio scene based on information regarding the user's location and / or orientation and / or direction of movement (i.e., for example, some audio elements that cannot be heard at that specific location are disabled or not rendered).
[0196] o Alternatively, using a metadata processor 1236, by operating on the audio metadata associated with all audio streams based on information regarding the user's location and / or orientation and / or direction of movement, · Select / activate the audio elements 152 necessary to compose the new audio scene.
[0197] · Also, enable the merging of all audio streams into a single audio stream.
[0198] · The media server distributes the necessary adaptation set.
[0199] · Alternatively, the client provides information regarding the user's positioning to the media server, and the media server provides an instruction regarding the necessary adaptation set.
[0200] Figure 1.2 shows another implementation example of such a system.
[0201] · Encoding side o Multiple media encoders that can be used to create one or more audio streams for each available audio scene associated with one sound scene portion of one viewpoint o Multiple media encoders that can be used to create one or more video streams for each available video scene associated with one video scene part of one viewpoint. For simplicity, the video encoder is not shown in the figure.
[0202] o A media server that stores multiple audio and video adaptation sets including different encodings of the same audio and video streams at different bitrates (i.e., different representations). Further, the media server stores the description information of all adaptation sets, which can include the following.
[0203] · The availability of all created adaptation sets.
[0204] · Information describing the association between one adaptation set and one audio scene and / or viewpoint. In this way, each adaptation set can be associated with one of the available audio scenes.
[0205] · Information describing the "boundaries" of each audio scene and / or viewpoint (which may include only a complete audio scene or an individual object). The boundaries of one audio scene may be defined, for example, as the geometric coordinates of a sphere (e.g., center and radius).
[0206] · On the client side, a system (client system) including any of the following.
[0207] o A receiving side capable of receiving the following, · Information regarding the user's position and / or orientation and / or movement direction (or information characterizing changes triggered by the user's actions) · Information regarding the availability of all adaptation sets, information describing the association between one adaptation set and one audio scene and / or viewpoint, and / or information describing the "boundaries" of each audio scene and / or viewpoint (which may include only a complete audio scene or an individual object). For example, such information may be provided as part of the Media Presentation Description (MPD) XML syntax in the case of a DASH delivery environment.
[0208] o On the side of the media consumption device used for content consumption (e.g., based on an HMD). Also, the media consumption device serves to collect information regarding the user's position and / or orientation and / or movement direction (or information characterizing changes triggered by the user's actions).
[0209] o The viewport processor 1232 can be configured as follows.
[0210] · Receive information regarding the current viewport including the user's position and / or orientation and / or movement direction (or information characterizing changes triggered by the user's actions) from the side of the media consumption device.
[0211] · Receive information regarding the ROI notified by the metadata (video viewport notified in the OMAF specification).
[0212] · Receive all information available on the receiving side.
[0213] · Based on all the information received from the received and / or available metadata and received and / or derived, determine which audio / video viewport to play at a specific moment. For example, the viewport processor 1232 determines as follows.
[0214] · Play one complete audio scene.
[0215] · It is necessary to create one new audio scene from all available audio scenes (for example, only some audio elements of multiple audio scenes are played, and the other remaining audio elements of these audio scenes are not played).
[0216] · It is necessary to reproduce the transition between two or more audio scenes.
[0217] o A selection part 1230 configured to select one or more adaptation sets from the available adaptation sets notified by the information received at the receiving end based on the information received from the viewport processor 1232. The selected adaptation set completely describes the audio scene to be played at the user's current location. This audio scene is either one complete audio scene defined on the encoding side or it is necessary to create a new audio scene from all available audio scenes.
[0218] Furthermore, when a transition between two or more audio scenes is about to occur based on an instruction from the viewport processor 1232, the selection portion 1230 can be configured to select one or more adaptation sets from the available adaptation sets signaled by the information received at the receiving end, and the selected adaptation set(s) fully describe(s) the audio scene(s) that need to be reproduced in the near future (e.g., if the user is walking at a specific speed in the direction of the next audio scene, the next audio scene is predicted to be needed and is selected prior to playback).
[0219] · Additionally, some adaptation sets corresponding to adjacent locations may be initially selected at a low bitrate (i.e., a representation encoded at a low bitrate is selected from the representations available in one adaptation set), and based on the change in position, the quality is improved by selecting a higher bitrate for these specific adaptation sets (i.e., a representation encoded at a higher bitrate is selected from the representations available in one adaptation set).
[0220] o A download and switching portion that can be configured as follows · Based on the instruction received from the selection portion 1230, request one or more adaptation sets from the available adaptation sets at the media server 120.
[0221] · Receive one or more adaptation sets (i.e., one representation out of all the representations available within each adaptation set) from the available adaptation sets at the media server 120.
[0222] · Extract metadata information from all the received audio streams A metadata processor 1236 that can be configured as follows · Receive information that can include audio metadata corresponding to each received audio stream from the download and switching information about the received audio stream.
[0223] · Based on the information received from the viewport processor 1232 that can include information about the user's position and / or orientation and / or direction of movement, by processing and manipulating the audio metadata associated with each audio stream, · Select / enable the audio elements 152 necessary to compose a new audio scene as indicated by the viewport processor 1232.
[0224] · All audio streams can be merged into a single audio stream.
[0225] o The stream muxer / merger 1238 may be configured to merge all selected audio streams into one audio stream based on the information received from the metadata processor 1236 and that can include the modified and processed audio metadata corresponding to all received audio streams.
[0226] o The media decoder is configured to receive and decode at least one audio stream for the playback of the new audio scene as indicated by the viewport processor 1232 based on the information about the user's position and / or orientation and / or direction of movement.
[0227] Figure 1.3 shows a system including a system (client system) that can embody a part of the system described in, for example, Figure 1.2 on the client side, which further or alternatively includes the following.
[0228] · The plurality of media decoders can be configured to decode individual audio streams as indicated by the metadata processor 1236 (e.g., some audio elements are deactivated).
[0229] · The mixer / renderer 1238 can be configured to play the final audio scene based on information regarding the user's position and / or orientation and / or direction of movement (i.e., for example, some audio elements that cannot be heard at that particular location are disabled or not rendered).
[0230] Solution 2 Figures 1.4, 1.5, and 1.6 show examples based on Solution 2 of the present invention (which may be an embodiment of the examples of Figures 1.1 and / or 1.2 and / or 1.3): independent adaptation sets for discrete positions (viewpoints) with a flexible adaptation set.
[0231] When the user moves within the VR environment, the audio scene 150 may change continuously. To ensure an excellent audio experience, all audio elements 152 that make up the audio scene 150 at a particular point in time need to be made available to the media decoder, and the media decoder can utilize the position information to create the final audio scene.
[0232] If the content is pre-encoded, under the premise that at some predefined locations these audio scenes do not overlap and the user can "jump / switch" from one location to the next, the system can accurately play the audio scenes of these specific locations.
[0233] However, when the user "walks" from one location to the next, the user can hear the audio elements 152 of two (or more) audio scenes 150 simultaneously. The solution to this use case is provided in examples of previous systems that do not rely on the mechanisms provided to decode multiple audio streams (using either a maxer with a single media decoder or multiple media decoders with additional mixers / renderers 1238), and it is necessary to provide the client / system 102 with an audio stream that describes the complete audio scene 150.
[0234] By introducing the concept of common audio elements 152 among multiple audio streams, optimizations are provided below.
[0235] Figure 1.4 shows an example where different scenes share at least one audio element (such as an audio object, a sound source, etc.). Thus, the client 102 can receive, for example, one main stream 106A associated only with one scene A (e.g., associated with the environment where the user is currently located) and associated with the object 152A, and one auxiliary stream 106B shared by a different scene B (e.g., a stream within the boundary between the current scene A of the user and an adjacent or neighboring stream B that shares the object 152B) and associated with the object 152B.
[0236] Thus, as shown in Figure 1.4, · Several independent or overlapping audio scenes are encoded into several audio streams. The audio streams 106 are created in the following ways.
[0237] o For each audio scene 150, one main stream can be created that includes only the audio elements 152 that are part of that respective audio scene and not part of other audio scenes. And / or For all audio scenes 150 that share the audio element 152, the common audio element 152 is encoded only in an auxiliary audio stream that is associated with only one of the audio scenes, and appropriate metadata information indicating its association with other audio scenes is created. Or, put another way, the additional metadata indicates that some audio streams may be used with multiple audio scenes. And / or o Depending on the use case, additional auxiliary streams may be created (e.g., some audio objects containing different languages may be encoded into independent streams for efficient delivery).
[0238] o In the provided embodiments, · Audio scene A is encoded as follows: · Main audio stream (A1, 106A), · Auxiliary audio stream (A2, 106B) · Metadata information indicating that some audio elements 152B of audio scene A are encoded not in these audio streams A but in an auxiliary stream A2 (106B) belonging to a different audio scene (audio scene B). · Audio scene B is encoded as follows: · Main audio stream (B1, 106C), · Auxiliary audio stream (B2), · Auxiliary audio stream (B3), · Metadata information indicating that the audio element 152B from audio stream B2 is a common audio element 152B that also belongs to audio scene A.
[0239] · Audio scene C is encoded into three streams (C1, C2, and C3).
[0240] · The audio streams 106 (106A, 106B, 106C …) are encoded at different bitrates (i.e., different representations), and for example, depending on the network connection, efficient bitrate adaptation becomes possible (i.e., a high-bitrate encoded version is delivered to users using a high-speed connection, and a low-bitrate version is delivered to users using a low-speed network connection).
[0241] · The audio stream 106 is stored in the media server 120. For each audio stream, different encodings at different bitrates (i.e., different representations) are grouped into one adaptation set, and the availability of all adaptation sets for which appropriate data has been created is notified. (Multiple representations of a stream associated with the same audio signal may exist in the same adaptation set at different bitrates and / or qualities and / or resolutions.) · Further, in addition to the adaptation set, the media server 120 can receive information regarding the position of the “boundary” of each audio scene and the relationship with each adaptation set (e.g., including only a complete audio scene or an individual object). In this way, each adaptation set can be associated with one or more of the available audio scenes 150. The boundary of one audio scene may be defined, for example, as the geometric coordinates of a sphere (e.g., the center and radius).
[0242] o Each adaptation set may also include descriptive information regarding where the sound scene or audio element 152 is active. For example, if one or more objects are included in one auxiliary stream (e.g., A2, 106B), the adaptation set can include information such as the location where the object can be heard (e.g., the coordinates of the center of the sphere and the radius).
[0243] Optionally or alternatively, each adaptation set (e.g., the adaptation set associated with scene B) can include descriptive information (e.g., metadata) that can indicate that the audio elements (e.g., 152B) of one audio scene (e.g., B) are (also or further) encoded in an audio stream (e.g., 106B) belonging to another audio scene (e.g., A).
[0244] · The media server 120 can provide information regarding the position of the "boundaries" associated with each adaptation set to the system 102 (client), e.g., a DASH client. For example, in the case of a DASH delivery environment, this may be embedded in a Media Presentation Description (MPD) XML syntax.
[0245] · The system 102 (client) can receive information regarding the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions).
[0246] · The system 102 (client) can receive information regarding each adaptation set, and based on this and / or the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions, such as x, y, z coordinates and yaw, pitch, roll values), the system 102 (client) can select one or more adaptation sets that fully or partially describe the audio scene 150 being played at the user 140's current location.
[0247] · The system 102 (client) can request one or more adaptation sets.
[0248] o Further, the system 102 (client) can select one or more adaptation sets that fully or partially describe a plurality of audio scenes 150, and use the audio streams 106 corresponding to the plurality of audio scenes 150 to create a new audio scene 150 that is played at the current location of the user 140.
[0249] o Based on metadata indicating that the audio element 152 is part of a plurality of audio scenes 150, the common audio element 152 can be requested only once per complete audio scene instead of twice to create a new audio scene.
[0250] o When the audio stream becomes available at the client system 102, in an example, one or more media decoders (104) are used to decode the individual audio streams and / or an additional mixer / renderer is used to play the final audio scene based on information regarding the user's position and / or orientation and / or direction of movement (i.e., for example, parts of the audio elements that cannot be heard at that particular location are muted or not rendered).
[0251] o Alternatively or additionally, a metadata processor is used to manipulate the audio metadata associated with all audio streams based on information regarding the user's position and / or orientation and / or direction of movement, · to select / enable the audio elements 152 (152A - 152c) necessary to construct a new audio scene. And / or · to enable all audio streams to be merged into a single audio stream.
[0252] · The media server 120 can deliver the required adaptation sets.
[0253] ·Alternatively, the system 102 (client) provides information regarding the positioning of the user 140 to the media server 120, and the media server provides instructions regarding the necessary adaptation set.
[0254] Figure 1.5 shows another exemplary embodiment of such a system.
[0255] ·Encoding side A plurality of media encoders 154 that can be used to create one or more audio streams 106 that embed audio elements 152 from one or more available audio scenes 150 associated with one sound scene portion of one viewpoint.
[0256] ·For each audio scene 150, one main stream can be created that includes only the audio elements 152 that are part of that respective audio scene 150 and not part of other audio scenes.
[0257] ·Additional auxiliary streams can be created for the same audio scene (e.g., some audio objects including different languages can be encoded into independent streams for efficient delivery).
[0258] ·Additional auxiliary streams can be created that include the following.
[0259] ·Audio elements 152 common to a plurality of audio scenes 150.
[0260] ·Metadata information indicating the association of this auxiliary stream with all other audio scenes 150 that share the common audio element 152. Or put another way, the metadata indicates the possibility that some audio streams can be used together with multiple audio scenes.
[0261] A plurality of media encoders that can be used to create one or more video streams of each available video scene associated with one video scene portion of one view point. For simplicity, the video encoder is not shown in the figure.
[0262] o A media server 120 that stores a plurality of audio and video adaptation sets including different encodings of the same audio and video streams at different bitrates (i.e., different representations). Further, the media server 120 stores the description information of all the adaptation sets, which can include the following.
[0263] · The availability of all created adaptation sets.
[0264] · Information describing the association between one adaptation set and one audio scene and / or view point. In this way, each adaptation set can be associated with one of the available audio scenes.
[0265] · Information describing the "boundaries" of each audio scene and / or view point (e.g., may include only the complete audio scene or individual objects). The boundaries of one audio scene may be defined, for example, as the geometric coordinates of a sphere (e.g., center and radius).
[0266] · Information indicating the association between one adaptation set and a plurality of audio scenes that share at least one common audio element.
[0267] · On the client side, a system (client system) including any of the following.
[0268] o A receiving side that can receive the following, · Information regarding the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions) · Information regarding the availability of all adaptation sets, as well as information describing the association of one adaptation set with one audio scene and / or viewpoint, and / or information describing the "boundaries" of each audio scene and / or viewpoint (e.g., may include only a complete audio scene or individual objects). For example, such information may be provided as part of the Media Presentation Description (MPD) XML syntax in the case of a DASH delivery environment.
[0269] · Information indicating the association of one adaptation set with multiple audio scenes that share at least one common audio element.
[0270] o On the media consumption device side used for content consumption (e.g., based on an HMD). Also, the media consumption device serves to collect information regarding the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions).
[0271] o The viewport processor 1232 can be configured as follows.
[0272] · Receive information regarding the current viewport including the user's position and / or orientation and / or direction of movement (or information characterizing changes triggered by the user's actions) from the media consumption device side.
[0273] · Receive information regarding the ROI notified by the metadata (video viewport notified in the OMAF specification).
[0274] · Receive all information available on the receiving side.
[0275] ·Based on all the information received from and / or derived from the received and / or available metadata, determine which audio / video viewport to play at a specific moment. For example, the viewport processor 1232 determines as follows.
[0276] ·One complete audio scene is reproduced ·It is necessary to create one new audio scene from all the available audio scenes (for example, only some audio elements of multiple audio scenes are reproduced, and the other remaining audio elements of these audio scenes are not reproduced).
[0277] ·It is necessary to reproduce the transition between two or more audio scenes o A selection part 1230 configured to select one or more adaptation sets from the available adaptation sets notified by the information received by the receiving end based on the information received from the viewport processor 1232. The selected adaptation set completely or partially describes the audio scene to be played at the user's current location. This audio scene is one or some complete audio scenes defined on the encoding side, or it is necessary to create a new audio scene from all the available audio scenes.
[0278] ·Furthermore, when the audio element 152 belonging to multiple audio scenes is selected based on the information indicating the association between at least one adaptation set and multiple audio scenes including the same audio element 152.
[0279] Furthermore, when a transition between two or more audio scenes is about to occur based on an instruction from the viewport processor 1232, the selection part 1230 can be configured to select one or more adaptation sets from the available adaptation sets signaled by the information received at the receiving end, and the selected adaptation set(s) fully describe the audio scene(s) that need to be reproduced in the near future (e.g., when the user walks at a specific speed in the direction of the next audio scene, the next audio scene is predicted to be needed and is selected prior to playback).
[0280] · Additionally, some adaptation sets corresponding to adjacent locations may be initially selected at a low bitrate (i.e., a representation encoded at a low bitrate is selected from the representations available in one adaptation set), and based on the change in position, the quality is improved by selecting a higher bitrate for these specific adaptation sets (i.e., a representation encoded at a higher bitrate is selected from the representations available in one adaptation set).
[0281] o A download and switching part that can be configured as follows · Based on the instruction received from the selection part 1230, request one or more adaptation sets from the available adaptation sets at the media server 120.
[0282] · Receive one or more adaptation sets (i.e., one representation out of all the representations available within each adaptation set) from the available adaptation sets at the media server 120.
[0283] · Extract metadata information from all the received audio streams A metadata processor 1236 that can be configured as follows · Receive information that can include audio metadata corresponding to each received audio stream from the download and switching information about the received audio stream.
[0284] · Based on information received from the viewport processor 1232 that can include information about the user's position and / or orientation and / or direction of movement, by processing and manipulating the audio metadata associated with each audio stream, · Select / enable the audio elements 152 necessary to compose a new audio scene as indicated by the viewport processor 1232.
[0285] · All audio streams can be merged into a single audio stream.
[0286] o The stream muxer / merger 1238 may be configured to merge all selected audio streams into one audio stream based on information received from the metadata processor 1236 that can include the modified and processed audio metadata corresponding to all received audio streams.
[0287] o The media decoder is configured to receive and decode at least one audio stream for playback of the new audio scene as indicated by the viewport processor 1232 based on information about the user's position and / or orientation and / or direction of movement.
[0288] Figure 1.6 shows a system including a system (client system) that can embody a part of the system described, for example, in Figure 5 on the client side, which further or alternatively includes the following.
[0289] · Multiple media decoders can be configured to decode individual audio streams as indicated by the metadata processor 1236 (e.g., some audio elements are deactivated).
[0290] · The mixer / renderer 1238 can be configured to play the final audio scene based on information regarding the user's position and / or orientation and / or direction of movement (i.e., for example, some audio elements that cannot be heard at that particular location are disabled or not rendered).
[0291] Update of File Format for File Playback In the case of the usage scenario of the file format, multiple main streams and auxiliary streams can be encapsulated as individual tracks in a single ISOBMFF file. A single track of such a file represents a single audio element as described above. Since the MPD containing the information necessary for the correct layout is not available, it is necessary to provide the information at the file format level, for example, by providing / introducing a specific file format box or track and a movie-level specific file format box. Depending on the usage scenario, there are various information necessary to enable the correct rendering of the encapsulated audio scene, but the following set of information is fundamental and must always be present.
[0292] · Information regarding the contained audio scene, such as "location boundaries" · Information regarding all available audio elements, especially which audio elements are encapsulated in which tracks · Information regarding the location of the encapsulated audio elements · A list of all audio elements belonging to one audio scene, and one audio element may belong to multiple audio scenes.
[0293] With this information, all of the use cases mentioned should function in a file-based environment, including cases where additional metadata processors or shared encodings are used.
[0294] Further Considerations Regarding the Above Examples In an example (e.g., at least one of FIGS. 1.1 to 6), at least one scene can be associated with at least one audio element (audio source 152), and each audio element is associated with a position and / or region in the visual environment where the audio element is audible, such that different audio streams are provided from the server system 120 to the client system 102 for different user positions and / or viewports and / or head orientations and / or movement data and / or interaction metadata and / or virtual position data within the scene.
[0295] In an example, the client system 102 may be configured to determine whether to play at least one audio element 152 of an audio stream (e.g., A1, A2) and / or one adaptation set in the presence of data on the current user's viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual position in the scene, and the system 102 is configured to request and / or receive at least one audio element at the current user's virtual position.
[0296] In an example, a client system (e.g., 102) may be configured to predictively determine whether at least one audio element (152) of an audio stream and / or one adaptation set becomes relevant and / or audible based on at least data of a user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data (110). The system may be configured to request and / or receive at least one audio element and / or an audio stream and / or an adaptation set at a virtual location of a particular user before a predicted movement and / or interaction of the user in a scene, and the system, when received, may be configured to play at least one audio element and / or an audio stream at a virtual location of the particular user after the movement and / or interaction of the user in the scene. For example, see FIGS. 8A and 8B above. In some examples, at least one of the operations of system 102 or 120 may be performed based on prediction data and / or statistical data and / or aggregated data.
[0297] In an example, a client system (e.g., 102) may be configured to request and / or receive at least one audio element (e.g., 152) at a lower bitrate and / or quality level at a virtual location of a user before a movement and / or interaction of the user in a scene, and the system may be configured to request and / or receive at least one audio element at a higher bitrate and / or quality level at a virtual location of the user after the movement and / or interaction of the user in the scene. For example, see FIG. 7B.
[0298] In an example, at least one audio element is associated with at least one scene, at least one audio element may be associated with a position and / or region within the visual environment associated with the scene, and the system is configured to request different streams at different bitrates and / or quality levels of the audio element based on the relevance and / or auditability level at the virtual position of each user in the scene, and the system is configured to request an audio stream at a higher bitrate and / or quality level for an audio element that is more relevant and / or more audible at the virtual position of the current user, and / or request an audio stream at a lower bitrate and / or quality level for an audio element that is less relevant and / or less audible at the virtual position of the current user. Generally, refer to FIG. 7A. Also refer to FIGS. 2a and 2b (where a more relevant and / or more audible source may be closer to the user), FIG. 3 (where a more relevant and / or more audible source is the source of scene 150a when the user is at position x1, and a more relevant and / or more audible source is the source of scene 150b when the user is at position x3), FIG. 4 (at time t2, a more relevant and / or more audible source may be from the first scene), FIG. 6 (a more audible source may be what the user sees from the front).
[0299] In an example, at least one audio element (152) is associated with a scene, each audio element is associated with a position and / or region within a visual environment associated with the scene, and the client system 102 is configured to periodically send data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual position data (110) to the server system 120, whereby, at a position close to at least one audio element (152), a higher bitrate and / or quality is provided from the server, and at a position farther from at least one audio element (152), a lower bitrate and / or quality stream is provided from the server. For example, refer to FIGS. 2a and 2b.
[0300] In an example, a plurality of scenes (e.g., 150A, 150B) may be defined for a plurality of visual environments such as adjacent and / or proximate environments, a first stream associated with a first current scene (e.g., 150A) is provided, and when the user transitions (150AB) to a second further scene (e.g., 150B), both the stream associated with the first scene and a second stream associated with the second scene are provided. For example, refer to FIG. 3.
[0301] In an example, a plurality of scenes are defined for first and second visual environments, the first and second environments are adjacent and / or proximate environments, a first stream associated with the first scene is provided from the server for playback of the first scene when the user's virtual position is in the first environment associated with the first scene, a second stream associated with the second scene is provided from the server for playback of the second scene when the user's virtual position is in the second environment associated with the second scene, and when the user's virtual position is at a transition position between the first scene and the second scene, both the first stream associated with the first scene and the second stream associated with the second scene are provided. For example, refer to FIG. 3.
[0302] In an example, the first stream associated with the first scene is obtained at a higher bitrate and / or quality when the user is in the first environment associated with the first scene, while the second stream associated with the second scene environment associated with the second environment is obtained at a lower bitrate and / or quality when the user is at the start of the transition position from the first scene to the second scene, and when the user is at the end of the transition position from the first scene to the second scene, the first stream associated with the first scene is obtained at a lower bitrate and / or quality, and the second stream associated with the second scene is obtained at a higher bitrate and / or quality. This is the case, for example, in FIG. 3.
[0303] In an example, a plurality of scenes (e.g., 150A, 150B) are defined for a plurality of visual environments (e.g., adjacent environments), and the system 102 can request and / or obtain a stream associated with the current scene at a higher bitrate and / or quality and a stream associated with the second scene at a lower bitrate and / or quality. See, for example, FIG. 4.
[0304] In an example, a plurality of N audio elements are defined, and when the distance of the user to the position or region of these audio elements is greater than a predetermined threshold, the N audio elements are processed to obtain a smaller number M (M < N) of audio elements associated with a position or region closer to the position or region of the N audio elements, whereby, when the distance of the user to the position or region of the N audio elements is less than a predetermined threshold, at least one audio stream associated with the N audio elements is provided to the system, or when the distance of the user to the position or region of the N audio elements is greater than a predetermined threshold, at least one audio stream associated with the M audio elements is provided to the system. See, for example, FIG. 1.7.
[0305] In an example, at least one visual environment scene is associated with at least one plurality of N audio elements (N >= 2), each audio element is associated with a position and / or region within the visual environment, at least one plurality of N audio elements may be provided in at least one representation at a high bitrate and / or quality level, at least one plurality of N audio elements are provided in at least one representation at a low bitrate and / or quality level, at least one representation is obtained by processing the N audio elements to obtain a smaller number M (M < N) of audio elements associated with a position or region close to the position or region of the N audio elements, the system is configured to request a representation at a higher bitrate and / or quality level for an audio element if the audio element is more relevant and / or more audible at the current virtual position of the user in the scene, and the system is configured to request a representation at a lower bitrate and / or quality level for an audio element if the audio element is less relevant and / or less audible at the current virtual position of the user in the scene. For example, see Figure 1.7.
[0306] In an example, different streams are obtained for different audio elements when the user's distance and / or relevance and / or audible level and / or angular orientation is below a predetermined threshold. For example, see Figure 1.7.
[0307] In an example, since different audio elements are provided in different viewports, if a first audio element is within the current viewport, the first audio element is obtained at a higher bitrate than a second audio element that is not within the viewport. For example, see Figure 6.
[0308] In the example, at least two visual environment scenes are defined, at least one first and second audio element is associated with a first scene associated with the first visual environment, at least one third audio element is associated with a second scene associated with the second visual environment, system 102 is configured to obtain metadata describing that at least one second audio element is further associated with the second visual environment scene, the system is configured to request and / or receive at least the first and second audio elements when the user's virtual position is in the first visual environment, the system is configured to request and / or receive at least the second and third audio elements when the user's virtual position is in the second visual environment scene, and the system is configured to request and / or receive at least the first, second, and third audio elements when the user's virtual position is transitioning between the first visual environment scene and the second visual environment scene. See, for example, FIG. 1.4. This also applies to FIG. 3.
[0309] In an example, at least one first audio element may be provided with at least one audio stream and / or an adaptation set, at least one second audio element is provided with at least one second audio stream and / or an adaptation set, at least one third audio element is provided with at least one third audio stream and / or an adaptation set, at least the first visual environment scene is described by metadata as a complete scene that requires at least the first and second audio streams and / or an adaptation set, the second visual environment scene is described by metadata as an incomplete scene that requires at least the third audio stream and / or an adaptation set, and at least the second audio stream and / or an adaptation set associated with at least the first visual environment scene, and the system is configured to operate the metadata to merge the second audio stream belonging to the first visual environment and the third audio stream associated with the second visual environment into a new single stream when the user's virtual position is in the second visual environment. For example, see FIGS. 1.2 to 1.3, FIG. 1.5, and FIG. 1.6.
[0310] In an example, system 102 may include a metadata processor (e.g., 1236) configured to manipulate metadata within at least one audio stream prior to at least one audio decoder based on data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual position data.
[0311] In an example, a metadata processor (e.g., 1236) may be configured to enable and / or disable at least one audio element within at least one audio stream in front of at least one audio decoder based on data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data. When the system determines as a result of the data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data that the audio element will no longer be played, the metadata processor may be configured to disable at least one audio element within at least one audio stream in front of at least one audio decoder. When the system determines as a result of the data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data that the audio element will be played, the metadata processor may be configured to enable at least one audio element within at least one audio stream in front of at least one audio decoder.
[0312] Server Side Also referred to herein is a server (120) for delivering audio and video streams for virtual reality VR, augmented reality AR, mixed reality MR, or 360-degree video environments to a client, the video and audio streams being played back on a media consumption device, the server (120) including an encoder for encoding and / or a storage device for storing a video stream that describes a visual environment, the visual environment being associated with an audio scene, the server further including an encoder for encoding and / or a storage device for storing a plurality of streams and / or audio elements and / or adaptation sets to be delivered to the client, the streams and / or audio elements and / or adaptation sets being associated with at least one audio scene, the server selecting and delivering a video stream based on a request from the client, the video stream being associated with an environment, and based on a request from the client, selecting an audio stream and / or audio elements and / or an adaptation set, the request being associated with at least data of the user's current viewport and / or head orientation and / or movement and / or interaction metadata and / or virtual location data, and associated with an audio scene associated with the environment, configured to deliver the audio stream to the client.
[0313] Further Embodiments and Variations Depending on a particular embodiment, an example can be implemented in hardware. The embodiment can be implemented, for example, using a digital storage medium storing electronically readable control signals that cooperate (or can cooperate) with a programmable computer system so that respective methods are executed, such as a floppy disk, a digital versatile disk (DVD), a Blu-ray disk, a compact disk (CD), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable and programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory. Thus, the digital storage medium may be computer-readable.
[0314] In general, an example can be implemented as a computer program product including program instructions, which, when the computer program product is executed on a computer, operate to execute one of the methods. The program instructions may be stored, for example, on a machine-readable medium.
[0315] Other examples include a computer program stored on a machine-readable carrier for executing one of the methods described herein. In other words, thus, an example of a method is a computer program having program instructions for executing one of the methods described herein when the computer program is executed on a computer.
[0316] Thus, a further example of a method includes a computer program for executing one of the methods described herein, which is a data carrier medium (or a digital storage medium, or a computer-readable medium) on which it is recorded. The data carrier medium, digital storage medium, or recorded medium is tangible and / or non-transitory, not an intangible and transitory signal.
[0317] Further examples include a processing unit, such as a computer or a programmable logic device, that executes one of the methods described herein.
[0318] Further examples include a computer having installed thereon a computer program for executing one of the methods described herein.
[0319] Further examples include an apparatus or system that transfers (e.g., electronically or optically) a computer program for executing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0320] In some examples, a programmable logic device (e.g., a field programmable gate array) may be used to execute some or all of the functions of the methods described herein. In some examples, the field programmable gate array may cooperate with a microprocessor to execute one of the methods described herein. Generally, the methods may be executed by any suitable hardware device.
[0321] The above examples illustrate the principles described above. It should be understood that modifications and changes to the arrangements and details described herein will be apparent. Accordingly, it is intended to be limited not by the specific details presented as descriptions and explanations of the embodiments herein, but by the claims that are imminently pending.
Claims
1. A system (102) for a virtual reality (VR), augmented reality (AR), mixed reality (MR) or 360-degree video environment configured to receive video and audio streams to be played on a media consumption device, comprising: The system (102) comprises: at least one media video decoder configured to decode a video signal from the video stream to present a VR, AR, MR, or 360-degree video environment scene to a user; at least one audio decoder (104) configured to decode an audio signal (108) from at least one audio stream (106); The system (102) is configured to request (112) at least one audio stream (106) and / or one audio element of an audio stream and / or one adaptation set from a server (120) based on at least the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data (110).
2. The system of claim 1, configured to provide the server (120) with the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data (110) in order to retrieve the at least one audio stream (106) and / or an audio element of an audio stream and / or an adaptation set from the server (120).
3. 3. The system of claim 1 or 2, wherein at least one scene is associated with at least one audio element (152), each audio element being associated with a position and / or area in the visual environment where the audio element is audible, and wherein different audio streams are provided for different user positions and / or viewports and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data within the scene.
4. configured to determine whether to play at least one audio element and / or one adaptation set of an audio stream relative to a current user viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position in said scene; The system according to claim 1 , wherein the system is configured to request and / or receive the at least one audio element at the current user's virtual location.
5. the system is configured to predictively determine whether at least one audio element (152) of an audio stream and / or an adaptation set will be relevant and / or audible based on at least a current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data (110) of the user, the system is configured to request and / or receive the at least one audio element and / or audio stream and / or adaptation set at a specific user virtual position prior to the predicted user movement and / or interaction in the scene, The system of claim 1 , wherein the system is configured, upon receipt, to play the at least one audio element and / or audio stream at a virtual position of the particular user after the user's movement and / or interaction in the scene.
6. configured to request and / or receive the at least one audio element (152) at a lower bitrate and / or quality level at a virtual position of the user prior to the user's movement and / or interaction in the scene; The system of claim 1 , wherein the system is configured to request and / or receive the at least one audio element at a higher bitrate and / or quality level at a virtual position of the user after the user's movement and / or interaction in the scene.
7. At least one audio element (152) is associated with at least one scene, each audio element being associated with a location and / or area within the visual environment associated with the scene; 7. The system of claim 1, wherein the system is configured to request and / or receive streams at a higher bit rate and / or quality for audio elements closer to the user than for audio elements further away from the user.
8. at least one audio element (152) is associated with at least one scene, said at least one audio element being associated with a location and / or area within said visual environment associated with said scene; the system is configured to request different streams at different bit rates and / or quality levels of audio elements based on a relevance and / or auditability level at each user's virtual location in the scene; the system is configured to request audio streams at higher bit rates and / or quality levels for audio elements that are more relevant and / or more audible at the current user's virtual location; and / or The system of claim 1 , configured to request audio streams at lower bit rates and / or quality levels for audio elements that are less relevant and / or less audible at the current user's virtual location.
9. At least one audio element (152) is associated with a scene, each audio element being associated with a location and / or area within the visual environment associated with the scene; The system is configured to periodically transmit the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data (110) to the server, whereby In a first location, a higher bit rate and / or quality stream is provided by the server; In a second location, a lower bit rate and / or quality stream is provided by the server; The system of claim 1 , wherein the first location is closer to the at least one audio element (152) than the second location.
10. A plurality of scenes (150A, 150B) are defined for a plurality of visual environments, such as adjacent and / or proximate environments; 10. The system of claim 1, wherein a first stream associated with a first current scene is provided, and when a user transitions to a second further scene, both the stream associated with the first scene and the second stream associated with the second scene are provided.
11. A plurality of scenes (150A, 150B) are defined for a first and a second visual environment, said first and second environments being adjacent and / or proximate environments; a first stream associated with the first scene is provided from the server for playback of the first scene when the user's location or virtual location is in a first environment associated with the first scene; a second stream associated with the second scene is provided from the server for playback of the second scene when the user's location or virtual location is in a second environment associated with the second scene; 11. The system of claim 1, wherein when the user's position or virtual position is at a transition position between the first scene and the second scene, both a first stream associated with the first scene and a second stream associated with the second scene are provided.
12. A plurality of scenes (150A, 150B) are defined for first and second visual environments, the first and second visual environments being adjacent and / or proximate environments; the system is configured to request and / or receive a first stream associated with a first scene associated with the first environment for playback of the first scene when the user's virtual location is in the first environment; the system is configured to request and / or receive a second stream associated with the second scene associated with the second environment for playback of the second scene when the user's virtual location is in the second environment; 12. The system of claim 1, wherein the system is configured to request and / or receive both a first stream associated with the first scene and a second stream associated with the second scene when the user's virtual location is at a transition location (150AB) between the first environment and the second environment.
13. the first stream associated with the first scene is acquired at a higher bitrate and / or quality when the user is in the first environment associated with the first scene; Meanwhile, the second stream associated with the second scene associated with the second environment is acquired at a lower bit rate and / or quality when the user is at the beginning of a transition position from the first scene to the second scene, when the user is at the end of a transition position from the first scene to the second scene, the first stream associated with the first scene is acquired at a lower bitrate and / or quality and the second stream associated with the second scene is acquired at a higher bitrate and / or quality; 13. The system of claim 10, wherein the lower bitrate and / or quality is lower than the higher bitrate and / or quality.
14. A plurality of scenes (150A, 150B) are defined for a plurality of environments, such as adjacent and / or nearby environments; The system is configured to obtain the stream associated with a first current scene associated with a first current environment; 14. The system of claim 1, wherein if the distance of the user's position or virtual position from a boundary of the scene is less than a predefined threshold, the system further acquires an audio stream associated with a second adjacent and / or nearby environment associated with the second scene.
15. A plurality of scenes (150A, 150B) are defined for a plurality of visual environments; the system requests and / or obtains the stream associated with the current scene at a higher bitrate and / or quality and the stream associated with the second scene at a lower bitrate and / or quality; 15. The system of claim 1, wherein the lower bitrate and / or quality is lower than the higher bitrate and / or quality.
16. A plurality of N audio elements are defined, and if the distance of the user to a position or region of these audio elements is greater than a predefined threshold, the N audio elements are processed to obtain a smaller number M (M<N) of audio elements associated with positions or regions close to the position or region of the N audio elements, whereby providing at least one audio stream associated with the N audio elements to the system if the user's distance to the location or area of the N audio elements is less than a predetermined threshold; or 16. The system of claim 1 , further comprising: providing at least one audio stream associated with the M audio elements to the system when the user's distance to the location or area of the N audio elements is greater than a predetermined threshold.
17. At least one visual environment scene is associated with at least one plurality of N audio elements, N>=2, each audio element being associated with a location and / or region within the visual environment; said at least one plurality of N audio elements being provided in at least one representation at a high bit rate and / or quality level; the at least one plurality of N audio elements is provided in at least one representation at a lower bit rate and / or quality level, the at least one representation being obtained by processing the N audio elements to obtain a smaller number M, M<N, of audio elements associated with positions or regions close to the position or region of the N audio elements; the system is configured to request the representation at a higher bitrate and / or quality level for an audio element if the audio element is more relevant and / or more audible at the user's current virtual position in the scene; 17. The system of claim 1, wherein the system is configured to request the representation at a lower bitrate and / or quality level for an audio element if the audio element is less relevant and / or less audible at the user's current virtual position in the scene.
18. The system according to claims 16 and 17, wherein different streams are obtained for the different audio elements if the user's distance and / or relevance and / or audibility level and / or angular orientation is below a predefined threshold.
19. 19. The system of claim 1 , wherein the system is configured to request and / or obtain the stream based on an orientation of the user in the scene and / or a direction of the user's movement and / or a user's interaction.
20. The system of claim 1 , wherein the viewport is associated with the position and / or virtual position and / or movement data and / or head.
21. 21. The system of claim 1, wherein different audio elements are provided in different viewports, the system being configured such that when one first audio element (S1) is within a viewport (160-1), the system requests and / or receives a first audio element of a higher bitrate than a second audio element (S2) that is not within the viewport.
22. configured to request and / or receive a first audio stream and a second audio stream, the first audio element of the first audio stream being more relevant and / or more audible than the second audio element of the second audio stream; 22. The system of claim 1, wherein the first audio stream is requested and / or received at a higher bitrate and / or quality than the bitrate and / or quality of the second audio stream.
23. At least two visual environment scenes are defined, at least one first and second audio element is associated with a first scene associated with the first visual environment, and at least one third audio element is associated with a second scene associated with the second visual environment; the system is configured to obtain metadata describing at least one second audio element further associated with a second visual environment scene; the system is configured to request and / or receive the at least first and second audio elements when the user's virtual location is in the first visual environment; the system is configured to request and / or receive the at least second and third audio elements when the user's virtual location is in the second visual environment scene; 23. The system of claim 1, wherein the system is configured to request and / or receive the at least first, second and third audio elements when the user's virtual position is transitioning between the first visual environment scene and the second visual environment scene.
24. the at least one first audio element is provided in at least one audio stream and / or adaptation set, the at least one second audio element is provided in at least one second audio stream and / or adaptation set, the at least one third audio element is provided in at least one third audio stream and / or adaptation set, the at least one first visual environment scene is described by metadata as a complete scene requiring the at least one first and second audio streams and / or adaptation sets, the second visual environment scene is described by metadata as an incomplete scene requiring the at least one third audio stream and / or adaptation set, and the at least one second audio stream and / or adaptation set associated with the at least one first visual environment scene; 24. The system of claim 23, further comprising a metadata processor configured to manipulate the metadata to enable merging of the second audio stream belonging to the first visual environment and the third audio stream associated with the second visual environment into a new single stream when the user's virtual location is in the second visual environment.
25. 25. The system of claim 1, further comprising a metadata processor configured to manipulate the metadata in at least one audio stream prior to the at least one audio decoder based on the user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data.
26. the metadata processor is configured to enable and / or disable at least one audio element in at least one audio stream prior to the at least one audio decoder based on a user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data; wherein the metadata processor is configured to disable at least one audio element in the at least one audio stream prior to the at least one audio decoder if the system determines that the audio element is not to be played anymore as a result of a current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data; 26. The system of claim 25, wherein the metadata processor is configured to enable at least one audio element in at least one audio stream before the at least one audio decoder if the system determines that the audio element is played as a result of a user's current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data.
27. 27. The system of claim 1 , configured to disable decoding of selected audio elements based on the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual position.
28. 28. The system of claim 1 , configured to merge at least one first audio stream associated with the current audio scene with at least one stream associated with an adjacent, nearby and / or future audio scene.
29. 29. The system of claim 1 , configured to obtain and / or collect statistical or aggregated data regarding the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual position data, and to send a request to the server associated with said statistical or aggregated data.
30. 30. The system of claim 1, configured to deactivate decoding and / or playback of at least one stream based on metadata associated with the at least one stream and based on the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual position data.
31. manipulating metadata associated with the selected group of audio streams based on at least said user's current or estimated viewport and / or head orientation and / or movement data and / or metadata and / or virtual position data; Selecting and / or validating and / or activating the audio elements that make up the audio scene to be played, and / or 31. The system of any one of claims 1 to 30, further configured to enable merging of all selected audio streams into a single audio stream.
32. 32. The system of claim 1 , configured to control the request for the at least one stream to the server based on the distance of the user's position from the boundaries of adjacent and / or nearby environments associated with different scenes or other metrics associated with the user's position in the current environment or a prediction in the future environment.
33. 33. The system of claim 1, wherein for each audio element or audio object, information is provided from the server system (120), said information including descriptive information about the sound scene or where the audio element is active.
34. 34. The system of claim 1, configured to select between playing one scene and compositing, mixing, multiplexing, superimposing or combining at least two scenes, the two scenes being associated with different adjacent and / or nearby environments, based on the current or future or viewport and / or head orientation and / or movement data and / or metadata and / or virtual position and / or user selection.
35. configured to create or use at least said adaptation set; Several adaptation sets are associated with one audio scene, and / or Additional information is provided that associates each adaptation set with one viewpoint or one audio scene; and / or information about the boundaries of an audio scene, and / or Information about the relationship between an adaptation set and an audio scene (e.g., an audio scene is encoded into three streams encapsulated in three adaptation sets), and / or information regarding connections between the boundaries of the audio scene and the plurality of adaptation sets; 35. The system of claim 1 , wherein additional information is provided which may include:
36. receiving a stream of scenes associated with an adjacent or proximate environment; upon detection of the transition of the boundary between two environments, starting the decoding and / or playing of the stream of the adjacent or proximate environment.
36. The system of claim 1 , configured to:
37. 37. A system comprising: the system of any one of claims 1 to 36 configured to operate as a client; and a server configured to deliver video and / or audio streams to be played on a media consumption device.
38. The system comprises: Requesting and / or receiving at least one first adaptation set including at least one audio stream associated with at least one first audio scene; requesting and / or receiving at least one second adaptation set including at least one second audio stream associated with at least two audio scenes including the at least one first audio scene; enabling merging of the at least one first audio stream and the at least one second audio stream into a new audio stream to be decoded based on available metadata regarding a user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual position data and / or information describing an association of the at least one first adaptation set to the at least one first audio scene and / or an association of the at least one second adaptation set to the at least one first audio scene, 38. The system of claim 1 , further configured to:
39. receiving information about a user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual position data and / or information characterizing changes triggered by actions of said user; receiving information regarding the availability of an adaptation set and information describing an association of the at least one adaptation set to at least one scene and / or viewpoint and / or viewport and / or position and / or virtual position and / or motion data and / or orientation; 39. The system of any one of claims 1 to 38, configured to:
40. determining whether to play at least one audio element from the at least one audio scene embedded in the at least one stream and at least one additional audio element from the at least one additional audio scene embedded in the at least one additional stream; - in case of a positive determination, performing an operation of merging or synthesizing or multiplexing or superimposing or combining said at least one additional stream of said additional audio scene with said at least one stream of said at least one audio scene.
40. The system of any one of claims 1 to 39, configured to:
41. Manipulating audio metadata associated with the selected audio stream based on at least the user's current viewport and / or head orientation and / or movement data and / or metadata and / or virtual position data; selecting and / or validating and / or activating the audio elements that compose the audio scene determined to be played, Allows merging all selected audio streams into a single audio stream, 41. The system of any one of claims 1 to 40, configured to:
42. A server (120) for delivering audio and video streams for a virtual reality (VR), an augmented reality (AR), a mixed reality (MR) or a 360-degree video environment to a client, said video and audio streams being played on a media consumption device; The server (120) includes an encoder for encoding and / or a storage device for storing a video stream describing a visual environment, the visual environment being associated with an audio scene; the server further comprises an encoder for encoding and / or a storage device for storing a plurality of streams and / or audio elements and / or adaptation sets to be delivered to the client, the streams and / or audio elements and / or adaptation sets being associated with at least one audio scene, The server, Selecting and delivering a video stream based on a request from the client, the video stream being associated with an environment; selecting audio streams and / or audio elements and / or adaptation sets based on a request from the client, the request being associated with at least a current viewport and / or head orientation and / or movement data and / or interaction metadata and / or virtual position data of the user and an audio scene associated with the environment; Delivering the audio stream to the client; The server (120) is configured to:
43. said streams are encapsulated into adaptation sets, each adaptation set comprising multiple streams associated with different representations, at different bit rates and / or qualities, of the same audio content; 43. The server of claim 42, wherein the selected adaptation set is selected based on the request from the client.
44. 42. A system comprising the system of any one of claims 1 to 41 acting as a client and as the server.
45. 45. A system according to claim 44, comprising the server according to claim 42 or 43.
46. 1. A method for a virtual reality (VR), augmented reality (AR), mixed reality (MR) or 360° video environment configured to receive video and audio streams played on a media consumption device, comprising: decoding a video signal from the video stream for presentation to a user of a VR, AR, MR, or 360 degree video environment scene; decoding an audio signal from the audio stream; - requesting and / or obtaining from a server at least one audio stream based on a current viewport and / or position data and / or head orientation and / or movement data and / or metadata and / or virtual position data and / or metadata of said user; The method includes:
47. 47. A computer program comprising instructions which, when executed by a processor, cause the processor to perform the method of claim 46.
Citation Information
Patent Citations
Game device, sound data creating method, and program
JP2007029506A
Stereoscopic audio visual system for interactive type
JP2009043274A
Interactive audio system
US20020103554A1
IEC23008-3