Method and system for handling global transitions between listening positions in a virtual reality environment
The described method and system address the limitation of existing audio rendering systems by enabling efficient 6DoF audio transitions through fade-out/fade-in gains and pre-rendered objects, ensuring smooth and computationally light transitions in virtual reality environments.
Patent Information
- Application Number
- JP2023151837
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-18
- Filing Date
- 2023-09-20
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2038-12-18
AI Technical Summary
Existing audio rendering systems are limited to handling rotational movements of the listener's head, failing to efficiently manage translational changes in listening position and associated degrees of freedom in virtual reality environments.
A method and system for rendering audio in virtual reality environments that includes determining listener movements between audio scenes, applying fade-out and fade-in gains to audio signals, and using pre-rendered virtual objects to facilitate smooth transitions while reducing computational load.
Enables efficient and acoustically consistent 6DoF audio experiences by minimizing computational requirements during transitions between audio scenes, maintaining audio source positions, and enhancing audio quality through environmental and directional considerations.
Smart Images

Figure 0007715775000001 
Figure 0007715775000002 
Figure 0007715775000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to the following priority applications: U.S. Provisional Application No. 62 / 599,841, filed on Dec. 18, 2017 (Docket No. D17085USP1); and European Application No. 17208088.9, filed on Dec. 18, 2017 (Docket No. D17085EP). The contents of these applications are hereby incorporated by reference.
[0002] Technical Field This document relates to efficiently and consistently handling transitions between auditory viewports and / or listening positions in a virtual reality (VR) rendering environment.
Background Art
[0003] Virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications are rapidly evolving to include increasingly sophisticated acoustic models of sound sources and scenes that can be enjoyed from different viewpoints / perspectives or listening positions. Two different classes of flexible audio representation may be used, for example, for VR applications: sound field representation and object - based representation. Sound field representation is a physics - based approach that encodes the wavefronts incident at a listening position. For example, techniques such as B - format or higher - order ambisonics (HOA) use spherical harmonic decomposition to represent spatial wavefronts. Object - based approaches represent a complex auditory scene as a collection of individual elements that include audio waveforms and possibly time - varying associated parameters or metadata.
[0004] Enjoying VR, AR, and MR applications can involve users experiencing different auditory perspectives or viewpoints. For example, room-based virtual reality may be provided based on a mechanism using six degrees of freedom (DoF). FIG. 1 shows an example of a 6 DoF interaction demonstrating translational movement (forward / backward, up / down, and left / right) and rotational movement (pitch, yaw, roll). Different from the 3 DoF spherical video experience restricted to head rotation, content created for 6 DoF interaction allows navigation within the virtual environment (such as physically walking indoors) in addition to head rotation. This can be achieved based on a position tracker (such as a camera-based one) and an orientation tracker (such as a gyroscope and / or accelerometer). 6 DoF tracking technology can be available on high-end mobile VR platforms (such as Google Tango) as well as high-end mobile VR platforms (such as PlayStation®VR, Oculus Rift, HTC Vive). The user's experience of the directionality and spatial spread of a sound source or audio source is critically important for the realism of the 6 DoF experience, especially the experience of navigation around virtual audio sources within a scene.
[0005] Available audio rendering systems (such as MPEG-H 3D renderers) are typically limited to rendering 3 DoF (i.e., the rotational movement of the audio scene caused by the movement of the listener's head). The translational changes in the listener's listening position and the associated DoF typically cannot be handled by such renderers. SUMMARY OF THE INVENTION PROBLEM TO BE SOLVED BY THE INVENTION
[0006] This document is directed to the technical problem of providing a resource-efficient method and system for handling translational movement in the context of audio rendering. MEANS FOR SOLVING THE PROBLEM
[0007] According to one aspect, a method of rendering audio in a virtual reality rendering environment is described. The method includes rendering an audio signal of an audio source of a starting audio scene from a starting source position on a sphere around a listener's listening position. Further, the method includes determining that the listener moves from the listening position in the starting audio scene to a listening position in a different ending audio scene. Additionally, the method includes applying a fade-out gain to the starting audio signal to determine a modified starting audio signal. The method further includes rendering the modified starting audio signal of the starting audio source from the starting source position on the sphere around the listening position.
[0008] According to a further aspect, a virtual reality audio renderer for rendering audio in a virtual reality rendering environment is described. The virtual reality audio renderer is configured to render an audio signal of an audio source of a starting audio scene from a starting source position on a sphere around a listener's listening position. Additionally, the virtual reality audio renderer is configured to determine that the listener moves from the listening position in the starting audio scene to a listening position in a different ending audio scene. Further, the virtual reality audio renderer is configured to apply a fade-out gain to the starting audio signal to determine a modified starting audio signal and to render the modified starting audio signal of the starting audio source from the starting source position on the sphere around the listening position.
[0009] According to a further aspect, a method for generating a bitstream indicative of an audio signal to be rendered within a virtual reality rendering environment is described. The method includes: determining an initial audio signal of an initial audio source of an initial audio scene; determining initial position data regarding an initial source position of the initial audio source; generating a bitstream including the initial audio signal and the initial position data; receiving an indication that a listener moves from the initial audio scene to a final audio scene within the virtual reality rendering environment; determining a final audio signal of a final audio source of the final audio scene; determining final position data regarding a final source position of the final audio source; and generating a bitstream including the final audio signal and the final position data.
[0010] According to another aspect, an encoder configured to generate a bitstream indicative of an audio signal to be rendered within a virtual reality rendering environment is described. The encoder is configured to: determine an initial audio signal of an initial audio source of an initial audio scene; determine initial position data regarding an initial source position of the initial audio source; generate a bitstream including the initial audio signal and the initial position data; receive an indication that a listener moves from the initial audio scene to a final audio scene within the virtual reality rendering environment; determine a final audio signal of a final audio source of the final audio scene; determine final position data regarding a final source position of the final audio source; and generate a bitstream including the final audio signal and the final position data.
[0011] According to a further aspect, a virtual reality audio renderer for rendering an audio signal in a virtual reality rendering environment is described. The audio renderer has a 3D audio renderer configured to render an audio signal of an audio source from a source position on a sphere around a listening position of a listener within the virtual reality rendering environment. Further, the virtual reality audio renderer has a preprocessing unit configured to determine a new listening position of the listener within the virtual reality rendering environment. Further, the preprocessing unit is configured to update the audio signal and the source position of the audio source with respect to the sphere around the new listening position. The 3D audio renderer is configured to render the updated audio signal of the audio source from the updated source position on the sphere around the new listening position.
[0012] According to a further aspect, a software program is described.
[0013] The software program may be adapted for execution on a processor and may be adapted to execute the method steps outlined herein when executed on the processor.
[0014] According to another aspect, a storage medium is described. The storage medium may be adapted for execution on a processor and may have a software program adapted to execute the method steps outlined herein when executed on the processor.
[0015] According to a further aspect, a computer program product is described. The computer program may include executable instructions for executing the method steps outlined herein when executed on a computer.
[0016] The methods and systems, including the preferred embodiments outlined in this patent application, may be used alone or in combination with other methods and systems disclosed in this document. Furthermore, all aspects of the methods and systems outlined in this patent application may be arbitrarily combined. In particular, the features of the claims may be combined with each other in any way.
Brief Description of the Drawings
[0017] The present invention will be described below in an exemplary manner with reference to the accompanying drawings.
Figure 1a
Figure 1b
Figure 1c
Figure 2
Figure 3
Figure 4a
Figure 4b
Figure 5a
Figure 5b
Figure 6
Figure 7
Figure 8
Figure 9a
Figure 9b
Figure 9c
Figure 9d
DETAILED DESCRIPTION OF THE INVENTION
[0018] As outlined above, this document relates to the efficient provision of 6DoF in a 3D (three-dimensional) audio environment. FIG. 1a shows a block diagram of an exemplary audio processing system 100. An acoustic environment 110, such as a stadium, includes various different audio sources 113. Exemplary audio sources 113 within the stadium are individual spectators, stadium speakers, players on the field, etc. The acoustic environment 110 may be subdivided into different audio scenes 111, 112. By way of example, the first audio scene 111 may correspond to a home team cheering block, and the second audio scene 111 may correspond to a guest team cheering block. Depending on where a listener is located within the audio environment, the listener perceives audio sources from either the first audio scene 111 or the second audio scene 112.
[0019] The different audio sources 113 of the audio environment 110 may be captured using an audio sensor 120, in particular using a microphone array. In particular, the one or more audio scenes 111, 112 of the audio environment 110 may be described using a multi-channel audio signal, one or more audio objects and / or a higher order ambisonics (HOA) signal. In the following, it is assumed that the audio source 113 is associated with audio data captured by the audio sensor 120. Here, the audio data indicates the audio signal and the position of the audio source 113 as a function of time (at a specific sampling rate, for example 20 ms).
[0020] A 3D audio renderer, such as an MPEG-H 3D audio renderer, typically assumes that the listener is located at a specific listening position within the audio scenes 111, 112. The audio data for the various audio sources 113 of the audio scenes 111, 112 is typically provided under the assumption that the listener is located at this specific listening position. The audio encoder 130 may have a 3D audio encoder 131 configured to encode the audio data of the audio sources 113 of the one or more audio scenes 111, 112.
[0021] Furthermore, VR (virtual reality) metadata may be provided. This enables the listener to change the listening position within the audio scenes 111, 112 and / or move between different audio scenes 111, 112. The encoder 130 may have a metadata encoder 132 configured to encode the VR metadata. The encoded VR metadata and the encoded audio data of the audio source 113 may be combined in a combination unit 133 to provide a bitstream 140 indicating the audio data and the VR metadata. The VR metadata may include, for example, environmental data describing the acoustic characteristics of the audio environment 110.
[0022] The bitstream 140 may be decoded using a decoder 150 to provide (decoded) audio data and (decoded) VR metadata. An audio renderer 160 for rendering audio within a rendering environment 180 that permits 6DoF may have a preprocessing unit 161 and a (normal) 3D audio renderer 162 (such as MPEG-H 3D audio). The preprocessing unit 161 may be configured to determine the listening position 182 of the listener 181 within the listening environment 180. The listening position 182 may indicate the audio scene 111 in which the listener 181 is located. Further, the listening position 182 may indicate the exact position within the audio scene 111. The preprocessing unit 161 may further be configured to determine a 3D audio signal for the current listening position 182, possibly based on the (decoded) VR metadata, based on the (decoded) audio data. The 3D audio signal may then be rendered using the 3D audio renderer 162.
[0023] It should be noted that the concepts and methods described herein may be specified in a frequency-varying manner, defined globally or in an object / media-dependent manner, applied directly in the spectral or time domain, and / or hard-coded into the VR renderer 160, or specified via a corresponding input interface.
[0024] Figure 1b shows an exemplary rendering environment 180. The listener 181 may be located within the starting audio scene 111. For rendering purposes, the audio sources 113, 194 may be assumed to be arranged at various rendering positions on a (unit) sphere 114 around the listener 181. The rendering positions of the various audio sources 113, 194 may change over time (according to a given sampling rate). Various situations may occur within the VR rendering environment 180: The listener 181 may perform a global transition 191 from the starting audio scene 111 to the ending audio scene 112. Alternatively or additionally, the listener 181 may perform a local transition 192 to a different listening position 182 within the same audio scene 111. Alternatively or additionally, the audio scene 111 may exhibit acoustically significant environmental characteristics (such as walls), which may be described using environmental data 193 and should be taken into account when a change in the listening position 182 occurs. Alternatively or additionally, the audio scene 111 may include one or more ambient audio sources 194 (such as for background noise), which should be taken into account when a change in the listening position 182 occurs.
[0025] Figure 1c shows an exemplary global transition 191 from a starting audio scene 111 having audio sources 113A1 through A n to an ending audio scene 112 having audio sources 113B1 through B m In particular, each audio source 113 may be included in only one of the starting audio scene 111 and the ending audio scene 112. For example, the audio sources 113A1 through A n are included in the starting audio scene 111 but not in the ending audio scene 112, and the audio sources 113B1 through B mis included in the end audio scene 112 but not in the start audio scene 111. The audio source 113 may be characterized by corresponding between-object characteristics (coordinates, directivity, distance attenuation function, etc.). The global transition 191 may be executed within a certain transition time interval (for example, within a range of 5 seconds, 1 second, or less than 1 second). The listening position 182 in the start scene 111 at the beginning of the global transition 191 is marked with "A". Further, the listening position 182 in the end scene 112 at the end of the global transition 191 is marked with "B". Further, FIG. 1c shows a local transition 192 in the end scene 112 between the listening position "B" and the listening position "C".
[0026] FIG. 2 shows a global transition 191 from the start scene 111 (or start viewport) to the end scene 112 (or end viewport) during the transition time interval t. Such a transition 191 may occur when the listener 181 switches between different scenes or viewports 111, 112, for example, within a stadium. Therefore, the global transition 191 from the start scene 111 to the end scene 112 does not necessarily need to correspond to the actual physical movement of the listener 181, and may simply be initiated by a listener's command to switch or transition to another viewport 111, 112. Nevertheless, the present disclosure refers to the position of the listener. This is understood to be the position of the listener in a VR / AR / MR environment. At the intermediate time point 213, the listener 181 may be located at an intermediate position between the start scene 111 and the end scene 112. The 3D audio signal 203 rendered at the intermediate position and / or at the intermediate time point 213 takes into account the sound propagation of each audio source 113, and is determined by determining the contribution of each of the audio sources 113A1 to A n of the start scene 111 and each of the audio sources 113B1 to B m of the end scene 112. However, this will result in a relatively high computational load (especially in the case of a relatively large number of audio sources 113).
[0027] At the beginning of the global transition 191, the listener 181 may be positioned at the starting listening position 201. During the entire transition 191, a 3D starting audio signal A G may be generated with respect to the starting listening position 201. Here, the starting audio signal depends only on the audio source 113 of the starting scene 111 (and does not depend on the audio source 113 of the ending scene 112). The global transition 191 does not affect the apparent source position of the audio source 113 of the starting scene 111. Thus, assuming a static audio source 113 in the starting scene 111, the rendering position of the audio source 113 during the global transition 191 with respect to the listening position 201 does not change even if the listening position (for the listener) transitions from the starting scene to the ending scene. Furthermore, at the beginning of the global transition 191, it may be fixed that the listener 181 arrives at the ending listening position 202 within the ending scene 112 at the end of the global transition 191. During the entire transition 191, a 3D ending audio signal B G may be generated with respect to the ending listening position 202. Here, the ending audio signal depends only on the audio source 113 of the ending scene 112 (and does not depend on the audio source 113 of the source scene 111). The global transition 191 does not affect the apparent source position of the audio source 113 of the ending scene 112 (for the listener).
[0028] To determine the intermediate audio signal 203 at an intermediate position and / or intermediate time point 213 during the global transition 191, the start audio signal at the intermediate time point 213 may be combined with the end audio signal at the intermediate time point 213. In particular, a fade-out factor or gain derived from the fade-out function 211 may be applied to the start audio signal. The fade-out function 211 may be such that the fade-out factor or gain "a" decreases within an increasing distance from the start scene 111 to the intermediate position. Further, a fade-in factor or gain derived from the fade-in function 212 may be applied to the end audio signal. The fade-in function 212 may be such that the fade-in factor or gain "b" increases with a decreasing distance from the end scene 112 to the intermediate position. Exemplary fade-out function 211 and exemplary fade-in function 212 are shown in FIG. 2. The intermediate audio signal may then be given by the weighted sum of the start audio signal and the end audio signal, where the weights correspond to the fade-out gain and the fade-in gain, respectively.
[0029] Thus, a fade-in function or curve 212 and a fade-out function or curve 211 can be defined for the global transition 191 between different 3DoF viewports 201, 202. The functions 211, 212 may be applied to pre-rendered virtual objects or 3D audio signals representing the start audio scene 111 and the end audio scene 112. By doing so, a consistent audio experience can be provided with reduced VR audio rendering calculations during the global transition 191 between different audio scenes 111, 112.
[0030] Intermediate position x i The intermediate audio signal 203 at may be determined using linear interpolation of the start audio signal and the end audio signal. The intensity F of the audio signal is F(x i ) = a * F(A G )+(1 - a) * F(B G) may be provided by. The factors "a" and "b = 1 - a" may be provided by a norm function a = a() that depends on the starting listening position 201, the ending listening position 202, and the intermediate position.
[0031] As an alternative to the function, a lookup table a = [1,…,0] may be provided for various intermediate positions. In the above, in order to allow a smooth transition from the starting scene 111 to the ending scene 112, for a plurality of intermediate positions x i it is understood that for, an intermediate audio signal 203 can be determined and rendered.
[0032] During the global transition 191, additional effects (such as the Doppler effect and / or reverberation) may be taken into account. The functions 211, 212 may be adapted by the content provider, for example, to reflect artistic intentions. Information about the functions 211, 212 may be included in the bitstream 140 as metadata. Thus, the encoder 130 may be configured to provide information about the fade-in function 212 and / or the fade-out function 211 as metadata within the bitstream 140. Alternatively or additionally, the audio renderer 160 may apply the functions 211, 212 stored in the audio renderer 160.
[0033] To indicate to the renderer 160 that the global transition 191 is executed from the starting scene 111 to the ending scene 112, a flag may be transmitted from the listener to the renderer 160, particularly to the VR preprocessing unit 161. The flag may trigger the audio processing described in this document for generating the intermediate audio signal during the transition phase. The flag may be signaled explicitly or implicitly through related information (for example, via the coordinates of a new viewport or the listening position 202). The flag may be sent from any data interface side (for example, server / content, user / scene, auxiliary). Along with the flag, the starting audio signal A Gand an end audio signal B G Information about may be provided. As an example, the ID of one or more audio objects or audio sources may be provided. Alternatively, a request to calculate the start audio signal and / or the end audio signal may be provided to the renderer 160.
[0034] Thus, a VR renderer 160 having a preprocessing unit 161 for a 3DoF renderer 162 is described, which enables a 6DoF function in a resource-efficient manner. The preprocessing unit 161 allows the use of a standard 3DoF renderer 162 such as an MPEG-H 3D audio renderer. The VR preprocessing unit 161 includes pre-rendered virtual audio objects A G and B G configured to efficiently perform calculations for the global transition 191 using. During the global transition 191, the computational load is reduced by utilizing only two pre-rendered virtual objects. Each virtual object may include multiple audio signals for multiple audio sources. Further, during the transition 191, only the pre-rendered virtual audio objects A G and B G may be provided within the bitstream 140, so the bitrate requirement can be reduced. Further, the processing delay can be reduced.
[0035] A 3DoF function may be provided for all intermediate positions along the global transition trajectory. This may be achieved by overlapping the start audio object and the end audio object using fade-out / fade-in functions 211, 212. Further, additional audio objects may be rendered and / or additional audio effects may be included.
[0036] Figure 3 shows an exemplary local transition 192 from a starting listening position B 301 to an ending listening position C 302 within the same audio scene 111. The audio scene 111 includes different audio sources or objects 311, 312, 313. The different audio sources or objects 311, 312, 313 may have different directivity profiles 332. Further, the audio scene 111 may have environmental characteristics, particularly one or more obstacles, that affect the propagation of audio within the audio scene 111. The environmental characteristics may be described using environmental data 193. Further, the relative distances 321, 322 of the audio object 311 to the listening positions 301, 302 may be known.
[0037] Figures 4a and 4b show a method for dealing with the effect of the local transition 192 on the intensities of different audio sources or objects 311, 312, 313. As outlined above, the audio sources 311, 312, 313 of the audio scene 111 are typically assumed to be located on a sphere 114 around the listening position 301 by the 3D audio renderer 162. Thus, at the beginning of the local transition 192, the audio sources 311, 312, 313 may be arranged on a starting sphere 114 around the starting listening position 301, and at the end of the local transition 192, the audio sources 311, 312, 313 may be arranged on an ending sphere 114 around the ending listening position 302.
[0038] The audio sources 311, 312, 313 may be remapped from the starting sphere 114 to the ending sphere 114. For this purpose, rays going from the ending listening position 302 to the source positions of the audio sources 311, 312, 313 on the starting sphere 114 may be considered. The audio sources 311, 312, 313 may be arranged at the intersections of those rays with the ending sphere 114.
[0039] The intensities F of the audio sources 311, 312, 313 on the end ball 114 typically differ from the intensities on the start ball 114. The intensity F may be modified using an intensity gain function or distance function 415 that gives a distance gain 410 as a function of the distances 420 of the audio sources 311, 312, 313 from the listening positions 301, 302. The distance function 415 typically indicates a cut-off distance 421 at which a zero distance gain 410 is applied beyond that. The start distance 321 from the start listening position 301 of the audio source 311 gives a start gain 411. Further, the end distance 322 to the end listening position 302 of the audio source 311 gives an end gain 412. The intensity F of the audio source 311 may be re-scaled using the start gain 411 and the end gain 412, thereby giving the intensity F of the audio source 311 on the end ball 114. In particular, the intensity F of the start audio signal of the audio source 311 on the start ball 114 may be divided by the start gain 411 and multiplied by the end gain 412 to give the intensity F of the end audio signal of the audio source 311 on the end ball 114.
[0040] Thus, the position of the audio source 311 after the local transition 192 may be determined as C (e.g., using a geometric transformation) i =source_remap_function(B i , C). Further, the intensity of the audio source 311 after the local transition 192 may be determined as F(C i ) = F(B i ) * distance_function(B i , C i , C). Thus, the distance attenuation may be modeled by the corresponding intensity gain given by the distance function 415.
[0041] Figures 5a and 5b show an audio source 312 with a non-uniform directivity profile 332. The directivity profile can be defined using a directivity gain 510 that indicates gain values for various directions or directivity angles 520. In particular, the directivity profile of the audio source 312 may be defined using a directivity gain function 515 that indicates the directivity gain 510 as a function of the directivity angle 520 (where the angle 520 can range from 0° to 360°). For a 3D audio source 312, the directivity angle 520 is typically a two-dimensional angle including an azimuth angle and an elevation angle. Thus, the directivity gain function 515 is typically a two-dimensional function of the two-dimensional directivity angle 520.
[0042] The directivity profile 332 of the audio source 312 may be taken into account in the context of the local transition 192 by determining a starting directivity angle 521 of a starting ray between the audio source 312 and the starting listening position 301 (the audio source is disposed on a starting sphere 114 around the starting listening position 301) and an ending directivity angle 522 of an ending ray between the audio source 312 and the ending listening position 302 (the audio source is disposed on an ending sphere 114 around the ending listening position 302). Using the directivity gain function 515 of the audio source 312, a starting directivity gain 511 and an ending directivity gain 512 can be determined as function values of the directivity gain function 515 for the starting directivity angle 521 and the ending directivity angle 522, respectively (see Figure 5b). Then, the intensity F of the audio source 312 at the starting listening position 301 may be divided by the starting directivity gain 511 and multiplied by the ending directivity gain 512 so as to determine the intensity F of the audio source 312 at the ending listening position 302.
[0043] Therefore, the sound source directivity may be parameterized by a directivity factor or gain 510 indicated by a directivity gain function 515. The directivity gain function 515 may represent the intensity of an audio source 312 at some distance as a function of the angle 520 with respect to the listening positions 301, 302. The directivity gain 510 may be defined as the ratio to the gain of an audio source 312 that is at the same distance and has the same total power, where the total power is radiated uniformly in all directions. The directivity profile 332 may be parameterized by a set of gains 510 corresponding to vectors that originate at the center of the audio source 312 and end at points distributed on a unit sphere around the center of the audio source 312. The directivity profile 332 of the audio source 312 may depend on the use case scenario and the available data (e.g., a uniform distribution for 3D flight cases, a flattened distribution for 2D+ use cases, etc.).
[0044] The resulting audio intensity of the audio source 312 at the end listening position 302 may be estimated as F(C i ) = F(B i ) * Distance_function() * Directivity_gain_function(C i , C, Directivity_parametrization). Here, the Directivity_gain_function [directivity gain function] depends on the directivity profile 332 of the audio source 312. The Distance_function() [distance function] takes into account the modified intensity caused by the change in the distances 321, 322 of the audio source 312 due to the transition of the audio source 312.
[0045] FIG. 6 shows an exemplary obstacle 603 that may need to be taken into account in the context of a local transition 192 between different listening positions 301, 302. Specifically, the audio source 313 may be hidden behind the obstacle 603 at the end listening position 302. The obstacle 603 may be described by environmental data 193 that includes a set of parameters. The parameters may be, for example, the spatial dimensions of the obstacle 603 and an obstacle attenuation function indicating the attenuation of sound caused by the obstacle 603.
[0046] The audio source 313 may indicate an obstacle-free distance 602 (OSD) to the end listening position 302. The OFD 602 may indicate the length of the shortest path between the audio source 313 and the end listening position 302 that does not pass through the obstacle 603. Further, the audio source 313 may indicate a going-through distance 601 (GHD) to the end listening position 302. The GHD 601 may indicate the length of the shortest path between the audio source 313 and the end listening position 302 that typically passes through the obstacle 603. The obstacle attenuation function may be a function of the OFD 602 and the GHD 601. Further, the obstacle attenuation function may be a function of the intensity F(B i ) of the audio source 313.
[0047] The intensity of the audio source C at the end listening position 302 i may be a combination of the sound from the audio source 313 passing around the obstacle 603 and the sound from the audio source 313 passing through the obstacle 603.
[0048] Thus, the VR renderer 160 may be provided with parameters for controlling the influence of the environmental geometry and the medium. The obstacle geometry / media data 193 or parameters may be provided by the content provider and / or the encoder 130. The audio intensity of the audio source 313 is: F(C i ) = F(B i) It can be estimated as *Distance_function(OFD)*Directivity_gain_function(OFD)+Obstacle_attenuation_function(F(Bi), OFD, GHD). The first term corresponds to the contribution of the sound that bypasses the obstacle 603. The second term corresponds to the contribution of the sound that passes through the obstacle 603.
[0049] The minimum obstacle-free distance (OFD) 602 may be determined using A* Dijkstra's path discovery algorithm and may be used to control direct sound attenuation. The passing distance (GHD) 601 may be used to control reverberation and distortion. Alternatively or additionally, a raycasting technique may be used to describe the effect of the obstacle 603 on the intensity of the audio source 313.
[0050] FIG. 7 shows an exemplary field of view 701 of a listener 181 located at the end listening position 302. Further, FIG. 7 shows an exemplary focus of interest 702 of a listener located at the end listening position 302. The field of view 701 and / or the focus of interest 702 may be used to enhance (e.g., amplify) audio from an audio source within the field of view 701 and / or the focus of interest 702. The field of view 701 may be considered a user-driven effect and may be used to enable an audio amplifier for an audio source 311 related to the user's field of view 701. In particular, a "cocktail party effect" simulation may be performed by removing frequency tiles from background audio sources to improve the intelligibility of the speech signal related to the audio source 311 within the listener's field of view 701. The focus of attention 702 may be seen as a content-driven effect and may be used to enable an audio amplifier for an audio source 311 related to the content area of interest (e.g., to draw the user's attention to look at and / or move in the direction of the audio source 311).
[0051] The audio intensity of the audio source 311 is: F(Bi ) = Field_of_view_function(C, F(B i ), Field_of_view_data) may be modified. Here, Field_of_view_function [field of view function] describes the modification applied to the audio signal of the audio source 311 within the field of view 701 of the listener 181. Further, the audio intensity of the audio source within the focus of interest 702 of the listener is: F(B i ) = Attention_focus_function(F(B i ), Attention_focus_data) may be modified. Here, attention_focus_function [focus of interest function] describes the modification applied to the audio signal of the audio source 311 within the focus of interest 702.
[0052] The functions described in this document for handling the transition of the listener 181 from the starting listening position 301 to the ending listening position 302 may be applied to the position changes of the audio sources 311, 312, 313 in a similar manner.
[0053] Thus, this document describes an efficient means for calculating the coordinates and / or audio intensity of virtual audio objects or audio sources 311, 312, 313 representing the local VR audio scene 111 at any listening positions 301, 302. The coordinates and / or intensity can be determined taking into account the sound source distance attenuation curve, sound source orientation and directivity, environmental geometry / media influence and / or "field of view" and "focus of interest" data for additional audio signal enhancement. The described methods can significantly reduce the computational amount by performing the calculation only when the listening positions 301, 302 and / or the positions of the audio objects / sources 311, 312, 313 change.
[0054] Furthermore, this document describes concepts for the specification of distances, directivities, geometric functions, processing, and / or signaling mechanisms for the VR renderer 160. Additionally, concepts for a minimum "obstacle-free distance" for controlling direct sound attenuation and a "passage distance" for controlling reverberation and distortion are described. Further, concepts for source directivity parameterization are described.
[0055] Figure 8 shows the handling of ambient sound sources 801, 802, 803 in the context of a local transition 192. Specifically, Figure 8 shows three different ambient sound sources 801, 802, 803. Here, the ambient sound may be attributed to point audio sources. An ambient sound flag may be provided to the preprocessing unit 161 to indicate that the point audio source 311 is the ambient sound source 801. The processing during local and / or global transitions of the listening positions 301, 302 may depend on the value of the ambient sound flag.
[0056] In the context of the global transition 191, the ambient sound source 801 may be treated like a normal audio source 311. Figure 8 shows the local transition 192. The positions of the ambient sound sources 801, 802, 803 may be copied from the start sphere 114 to the end sphere 114, thereby providing the positions of the ambient sound sources 811, 812, 813 at the end listening position 302. Further, if the environmental conditions remain unchanged, the intensity of the ambient sound source 801 may be kept constant. That is, F(C Ai ) = F(B Ai ). On the other hand, in the case of the obstacle 603, the intensities of the ambient sound sources 803, 813 may be determined, for example, as F(C Ai ) = F(B Ai ) * Distance_function Ai (OFD) + Obstacle_attenuation_function(F(B Ai ), OFD, GHD).
[0057] FIG. 9a shows a flowchart of an exemplary method 900 for rendering audio in a virtual reality rendering environment 180. Method 900 may be executed by a VR audio renderer 160. Method 900 includes rendering 901 an origin audio signal of an audio source 113 of an origin audio scene 111 from an origin source position on a sphere 114 around a listening position 201 of a listener 181. Rendering 901 may be performed using a 3D audio renderer 162 that is limited to handling only 3DoF, and in particular may be limited to handling only rotational movements of the listener's 181 head. In particular, 3D audio renderer 162 is not configured to handle translational movements of the listener's head. 3D audio renderer 162 may include, or alternatively be, an MPEG-H audio renderer.
[0058] Note that the expression "rendering an audio signal of audio source 113 from a particular source position" indicates that the listener perceives the audio signal as coming from that particular source position. This expression should not be understood as a limitation on how the audio signal is actually rendered. Various different rendering techniques may be used to "render an audio signal from a particular source position", i.e., to provide the listener 181 with the perception that the audio signal is coming from a particular source position.
[0059] Furthermore, method 900 includes determining 902 that listener 181 moves from a listening position 201 within a starting audio scene 111 to a listening position 202 within a different ending audio scene 112. Thus, a global transition 191 from the starting audio scene 111 to the ending audio scene 112 can be detected. In this context, method 900 may include receiving an indication that listener 181 moves from the starting audio scene 111 to the ending audio scene 112. The indication may include a flag or may be a flag. The indication may be communicated from listener 181 to VR audio renderer 160, for example, via a user interface of VR audio renderer 160.
[0060] Typically, the starting audio scene 111 and the ending audio scene 112 each include one or more audio sources 113 that are different from each other. Specifically, the starting audio signals of the one or more starting audio sources 113 may not be audible within the ending audio scene 112, and / or the ending audio signals of the one or more ending audio sources 113 may not be audible within the starting audio scene 111.
[0061] Method 900 may include applying a fade-out gain to the starting audio signal 903 (in response to determining that a global transition 191 to a new ending audio scene 112 is being performed), to determine a modified starting audio signal. In particular, the starting audio signal is generated such that it would be perceived at the listening position within the starting audio scene, regardless of the movement of the listener 181 from the listening position 201 within the starting audio scene 111 to the listening position 202 within the ending audio scene 112. Further, method 900 may include rendering 904 the modified starting audio signal of the starting audio source 113 from the starting source position on the sphere 114 around the listener positions 201, 202 (in response to determining that a global transition 191 to a new ending audio scene 112 is being performed). These operations may be repeated during the global transition 191, for example, at regular time intervals.
[0062] Thus, by gradually fading out the starting audio signal of the one or more starting audio sources 113 of the starting audio scene 111, a global transition 191 between different audio scenes 111, 112 can be performed. As a result, a computationally efficient and acoustically consistent global transition 191 between different audio scenes 111, 112 is provided.
[0063] It may be determined that the listener 181 moves from the starting audio scene 111 to the ending audio scene 112 during a certain transition time interval. Here, the transition time interval typically has a certain duration (e.g., 2s, 1s, 500 ms or less). The global transition 191 may be gradually performed within the transition time interval. Specifically, during the global transition 191, an intermediate time point 213 within the transition time interval may be determined (e.g., according to a certain sampling rate such as 100 ms, 50 ms, 20 ms or less). Then, the fade-out gain may be determined based on the relative position of the intermediate time point 213 within the transition time interval.
[0064] Specifically, the transition time interval for the global transition 191 may be subdivided into a sequence of intermediate time points 213. For each intermediate time point 213 of the sequence of intermediate time points 213, a fade - out gain for modifying the start - audio signal of the one or more start - audio sources may be determined. Further, at each intermediate time point 213 of the sequence of intermediate time points 213, the modified start - audio signal of the one or more start - audio sources 113 may be rendered from the start - source position on the sphere 114 around the listening positions 201, 202. By doing so, an acoustically consistent global transition 191 can be performed in a computationally efficient manner.
[0065] The method 900 may include providing a fade - out function 211 that indicates fade - out gains at various intermediate time points 213 within the transition time interval. Here, the fade - out function 211 typically is such that the fade - out gain decreases with the progressing intermediate time point 213, thereby providing a smooth global transition 191 to the end - audio scene 112. Specifically, the fade - out function 211 can be such that the start - audio signal remains unmodified at the beginning of the transition time interval, the start - audio signal is increasingly attenuated at the progressing intermediate time point 213, and / or the start - audio signal is completely attenuated at the end of the transition time interval.
[0066] The start - source position of the start - audio source 113 on the sphere 114 around the listening positions 201, 202 may be maintained as the listener 181 moves from the start - audio scene 111 to the end - audio scene 112 (especially, throughout the transition time interval). Alternatively or additionally, it may be assumed that the listener 181 remains at the same listening positions 201, 202 (throughout the transition time interval). By doing so, the computational amount for the global transition 191 between the audio scenes 111, 112 can be further reduced.
[0067] Method 900 may further include determining an end audio signal of an end audio source 113 of an end audio scene 112. Further, method 900 may include determining an end source position on a sphere 114 around listening positions 201, 202. In particular, the end audio signal is generated such that it would be perceived at a listening position within the end audio scene, regardless of the movement of listener 181 from a listening position 201 within start audio scene 111 to a listening position 202 within end audio scene 112. Further, method 900 may include applying a fade-in gain to the end audio signal to determine a modified end audio signal. The modified end audio signal of end audio source 113 may then be rendered from the end source position on sphere 114 around listening positions 201, 202. These operations may be repeated during global transition 191 and may be performed, for example, at regular time intervals.
[0068] Thus, similar to the fade-out of the start audio signal of the one or more start audio sources 113 of start scene 111, the end audio signal of the one or more end audio sources 113 of end scene 112 may be faded in, thereby providing a smooth global transition 191 between audio scenes 111, 112.
[0069] As described above, listener 181 may move from start audio scene 111 to end audio scene 112 during the transition time interval. The fade-in gain may be determined based on the relative position of an intermediate time point 213 within the transition time interval. Specifically, a sequence of fade-in gains may be determined for a corresponding sequence of intermediate time points 213 during global transition 191.
[0070] The fade-in gain may be determined using a fade-in function 212 that indicates the fade-in gain at various intermediate time points 213 within the transition time interval. Here, the fade-in function 212 is typically such that the fade-in gain increases with the progressing intermediate time point 213. Specifically, the fade-in function 212 may be such that the end-point audio signal is fully attenuated at the beginning of the transition time interval, the attenuation of the end-point audio signal decreases at the progressing intermediate time point 213, and / or the end-point audio signal remains unmodified at the end of the transition time interval, thereby providing a smooth global transition 191 between the audio scenes 111, 112 in a computationally efficient manner.
[0071] Similar to the start source position of the start audio source 113, the end source position of the end audio source 113 on the sphere 114 around the listening positions 201, 202 may be maintained, particularly throughout the transition time interval, when the listener 181 moves from the start audio scene 111 to the end audio scene 112. Alternatively or additionally, it may be assumed that the listener 181 remains at the same listening positions 201, 202 (throughout the transition time interval). By doing so, the computational load for the global transition 191 between the audio scenes 111, 112 can be further reduced.
[0072] The fade-out function 211 and / or the fade-in function 212 may be derived from a bit stream indicating a start audio signal and / or an end audio signal. The bit stream 140 may be provided to the VR audio renderer 160 by the encoder 130. Thus, the global transition 191 can be controlled by the content provider. Alternatively or additionally, the fade-out function 211 and / or the fade-in function 212 may be derived from a storage unit of a virtual reality (VR) audio renderer 160 configured to render a start audio signal and / or an end audio signal within the virtual reality rendering environment 180, thereby providing reliable operation during the global transition 191 between the audio scenes 111, 112.
[0073] The method 900 may include sending an indicator (e.g., a flag indicating this) to the encoder 130 that the listener 181 is moving from the start audio scene 111 to the end audio scene 112. Here, the encoder 130 may be configured to generate a bit stream 140 indicating a start audio signal and / or an end audio signal. Based on the indicator, the encoder 130 can selectively provide the audio signal for the one or more audio sources 113 of the start audio scene 111 and / or for the one or more audio sources 113 of the end audio scene 112 within the bit stream 140. Thus, by providing an indicator for the upcoming global transition 191, it is possible to reduce the required bandwidth for the bit stream 140.
[0074] As already shown above, the starting audio scene 111 may include a plurality of starting audio sources 113. Thus, method 900 may include rendering a plurality of starting audio signals of the corresponding plurality of starting audio sources 113 from a plurality of different starting source positions on the sphere 114 around the listening positions 201, 202. Further, method 900 may include applying a fade-out gain to the plurality of starting audio signals to determine a plurality of modified starting audio signals. Further, method 900 may include rendering the plurality of modified starting audio signals of the starting audio sources 113 from the corresponding plurality of different starting source positions on the sphere 114 around the listening positions 201, 202.
[0075] Similarly, method 900 may include determining a plurality of ending audio signals of the corresponding plurality of ending audio sources 113 of the ending audio scene 112. Further, method 900 may include determining a plurality of ending source positions on the sphere 114 around the listening positions 201, 202. Further, method 900 may include applying a fade-in gain to the plurality of ending audio signals to determine the corresponding plurality of modified ending audio signals. Further, method 900 includes rendering the plurality of modified ending audio signals of the plurality of ending audio sources 113 from the corresponding plurality of ending source positions on the sphere 114 around the listening positions 201, 202.
[0076] Alternatively or additionally, the starting audio signal rendered during the global transition 191 may be an overlap of the audio signals of a plurality of starting audio sources 113. Specifically, at the beginning of the transition time interval, the audio signals of (all) the audio sources 113 of the starting audio scene 111 may be combined to give a combined starting audio signal. This starting audio signal may be modified using a fade-out gain. Further, the starting audio signal may be updated at a specific sampling rate (e.g., 20 ms) during the transition time interval. Similarly, the ending audio signal may correspond to a combination of the audio signals of a plurality of ending audio sources 113 (in particular, all ending audio sources 113). The combined ending audio source may then be modified during the transition time interval using a fade-in gain. By combining the audio signals of the starting audio scene 111 and the ending audio scene 112 respectively, the computational amount can be further reduced.
[0077] Furthermore, a virtual reality audio renderer 160 for rendering audio in the virtual reality rendering environment 180 is described. As outlined in this document, the VR audio renderer 160 may have a preprocessing unit 161 and a 3D audio renderer 162. The virtual reality audio renderer 160 may be configured to render the starting audio signal of the starting audio source 113 of the starting audio scene 111 from the starting source position on the sphere 114 around the listening position 201 of the listener 181. Further, the VR audio renderer 160 is configured to determine that the listener 181 moves from the listening position 201 within the starting audio scene 111 to the listening position 202 within a different ending audio scene 112. Further, the VR audio renderer 160 is configured to apply a fade-out gain to the starting audio signal to determine a modified starting audio signal and render the modified starting audio signal of the starting audio source 113 from the starting source position on the sphere 114 around the listening positions 201, 202.
[0078] Furthermore, an encoder 130 is described that is configured to generate a bitstream 140 indicative of an audio signal to be rendered within a virtual reality rendering environment 180. The renderer 130 may be configured to determine an origin audio signal of an origin audio source 113 of an origin audio scene 111. Further, the encoder 130 may be configured to determine origin position data regarding an origin source position of the origin audio source 113. The encoder 130 may then generate a bitstream 140 that includes the origin audio signal and the origin position data.
[0079] The encoder 130 may receive an indication (via a feedback channel from the VR audio renderer 160 to the encoder 130) that the listener 181 is moving within the virtual reality rendering environment 180 from the origin audio scene 111 to the destination audio scene 112.
[0080] The encoder 130 may then determine a destination audio signal of a destination audio source 113 of the destination audio scene 112 and destination position data regarding a destination source position of the destination audio source 113 (in particular, only in response to receiving such an indication). Further, the encoder 130 may generate a bitstream 140 that includes the destination audio signal and the destination position data. Thus, the encoder 130 may be configured to provide the destination audio signal of one or more destination audio sources 113 of the destination audio scene 112 only upon receiving an indication regarding a global transition 191 to the destination audio scene 112. By doing so, the required bandwidth for the bitstream 140 may be reduced.
[0081] Figure 9b shows a flowchart of a corresponding method 930 for generating a bitstream 140 that represents an audio signal to be rendered within a virtual reality rendering environment 180. Method 930 includes determining 931 an initial audio signal of an initial audio source 113 of an initial audio scene 111. Further, method 930 includes determining 932 initial position data regarding an initial source position of the initial audio source 113. Further, method 930 includes generating 933 a bitstream 140 that includes the initial audio signal and the initial position data.
[0082] Method 930 includes receiving 934 an indication that a listener 181 moves from an initial audio scene 111 to a final audio scene 112 within the virtual reality rendering environment 180. In response thereto, method 930 may include determining 935 a final audio signal of a final audio source 113 of the final audio scene 112 and determining 936 final position data regarding a final source position of the final audio source 113. Further, method 930 includes generating 937 a bitstream 140 that includes the final audio signal and the final position data.
[0083] Figure 9c shows a flowchart of an exemplary method 910 for rendering an audio signal within a virtual reality rendering environment 180. Method 910 may be performed by a VR audio renderer 160.
[0084] Method 910 includes rendering 911 the initial audio signals of audio sources 311, 312, 313 from initial source positions on an initial sphere 114 around an initial listening position 301 of the listener 181. Rendering 911 may be performed using a 3D audio renderer 162. In particular, rendering 911 may be performed under the assumption that the initial listening position 301 is fixed. Thus, rendering 911 may be limited to three degrees of freedom (in particular, to rotational movements of the head of the listener 181).
[0085] To account for the additional three degrees of freedom regarding the translational movement of the listener 181, method 910 may include determining 912 that the listener 181 moves from the starting listening position 301 to the ending listening position 302. Here, the ending listening position 302 is typically within the same audio scene 111. Thus, the listener 181 may be determined 912 to perform a local transition 192 within the same audio scene 111.
[0086] In response to determining that the listener 181 performs a local transition 192, method 910 may include determining 913 the ending source positions of the audio sources 311, 312, 313 on the ending sphere 114 around the ending listening position 302 based on the starting source positions. In other words, the source positions of the audio sources 311, 312, 313 may be transferred from the starting sphere 114 around the starting listening position 301 to the ending sphere 114 around the ending position 302. This may be achieved by projecting the starting source positions from the starting sphere 114 to the ending sphere. In particular, the ending source positions may be determined such that the ending source positions correspond to the intersections of the rays between the ending listening position 302 and the starting source positions with the ending sphere 114.
[0087] Furthermore, method 910 may include determining 914 the ending audio signals of the audio sources 311, 312, 313 based on the starting audio signals (in response to determining that the listener 181 performs a local transition 192). In particular, the intensity of the ending audio signals may be determined based on the intensity of the starting audio signals. Alternatively or additionally, the spectral composition of the ending audio signals may be determined based on the spectral composition of the starting audio signals. Thus, how the audio signals of the audio sources 311, 312, 313 are perceived from the ending listening position 302 may be determined (in particular, the intensity and / or spectral composition of the audio signals may be determined).
[0088] The above-described determining steps 913, 914 may be executed by the preprocessing unit 161 of the VR audio renderer 160. The preprocessing unit 161 may handle the translational movement of the listener 181 by transferring the audio signals of one or more audio sources 311, 312, 313 from the starting sphere 114 around the starting listening position 301 to the ending sphere 114 around the ending listening position 302. As a result, the transferred audio signals of the one or more audio sources 311, 312, 313 may also be rendered using a 3D audio renderer 162 (which may be limited to 3DoF). Thus, the method 910 allows for the efficient provision of 6DoF within the VR audio rendering environment 180.
[0089] As a result, the method 910 may include rendering 915 the ending audio signals of the audio sources 311, 312, 313 (using a 3D audio renderer such as, for example, an MPEG-H audio renderer) from the ending source position on the ending sphere 114 around the ending listening position 302.
[0090] Determining 914 the ending audio signal may include determining the ending distance 322 between the starting source position and the ending listening position 302. The ending audio signal (in particular, the intensity of the ending audio signal) may then be determined (in particular, scaled) based on the ending distance 322. In particular, determining 914 the ending audio signal may include applying a distance gain 410 to the starting audio signal. Here, the distance gain 410 depends on the ending distance 322.
[0091] A distance function 415 may be provided that represents the distance gain 410 as a function of the distances 321, 322 between the source positions of the audio signals 311, 312, 313 and the listening positions 301, 302 of the listeners 181. The distance gain 410 applied to the starting audio signal (to determine the ending audio signal) may be determined based on the function value of the distance function 415 for the ending distance 322. By doing so, the ending audio signal may be determined efficiently and precisely.
[0092] Furthermore, determining the ending audio signal 914 may include determining the starting distance 321 between the starting source position and the starting listening position 301. Then, the ending audio signal may be determined based on the starting distance 321 as well. In particular, the distance gain 410 applied to the starting audio signal may be determined based on the function value of the distance function 415 for the starting distance 321. In one preferred example, the function value of the distance function 415 for the starting distance 321 and the function value of the distance function 415 for the ending distance 322 are used to rescale the intensity of the starting audio signal to determine the ending audio signal. Thus, an efficient and precise local transition 191 within the audio scene 111 may be provided.
[0093] Determining the ending audio signal 914 may include determining the directivity profile 332 of the audio sources 311, 312, 313. The directivity profile 332 may indicate the intensity of the starting audio signal in various directions. Then, the ending audio signal may be determined based on the directivity profile 332 as well. By taking the directivity profile 332 into account, the acoustic quality of the local transition 192 may be improved.
[0094] The directivity profile 332 may indicate a directivity gain 510 that is applied to the starting audio signal to determine the ending audio signal. In particular, the directivity profile 332 may indicate a directivity gain function 515. Here, the directivity gain function 515 may indicate the directivity gain 510 as a function of a (possibly two-dimensional) directivity angle 520 between the source positions of the audio sources 311, 312, 313 and the listening positions 301, 302 of the listener 181.
[0095] Thus, determining 914 the ending audio signal may include determining an ending angle 522 between the ending source position and the ending listening position 302. Then, the ending audio signal may be determined based on the ending angle 522. In particular, the ending audio signal may be determined based on the function value of the directivity gain function 515 for the ending angle 522.
[0096] Alternatively or additionally, determining 914 the ending audio signal may include determining a starting angle 521 between the starting source position and the starting listening position 301. Then, the ending audio signal may be determined based on the starting angle 521. In particular, the ending audio signal may be determined based on the function value of the directivity gain function 515 for the starting angle 521. In one preferred example, the ending audio signal may be determined by modifying the intensity of the starting audio signal using the function values of the directivity gain function 515 for the starting angle 521 and for the ending angle 522 to determine the intensity of the ending audio signal.
[0097] Furthermore, method 910 may include end - point environment data 193 indicating the audio propagation characteristics of the medium between the end - point source position and the end - point listening position 302. The end - point environment data 193 may indicate an obstacle 603 located on the direct path between the end - point source position and the end - point listening position 302; indicate information regarding the spatial dimensions of the obstacle 603; and / or indicate the attenuation suffered by the audio signal on the direct path between the end - point source position and the end - point listening position 302. In particular, the end - point environment data 193 may indicate an obstacle attenuation function of the obstacle 603, and the attenuation function may indicate the attenuation suffered by the audio signal passing through the obstacle 603 located on the direct path between the end - point source position and the end - point listening position 302.
[0098] The end - point audio signal may be determined based on the end - point environment data 193, thereby further enhancing the quality of the audio rendered within the VR rendering environment 180.
[0099] As shown above, the end - point environment data 193 may indicate an obstacle 603 on the direct path between the end - point source position and the end - point listening position 302. Method 910 may include determining a passing distance 601 between the end - point source position and the end - point listening position 302 on the direct path. Then, the end - point audio signal may be determined based on the passing distance 601. Alternatively or additionally, a non - obstacle distance 602 between the end - point source position and the end - point listening position 302 on an indirect path that does not pass through the obstacle 603 may be determined. Then, the end - point audio signal may be determined based on the non - obstacle distance 602.
[0100] Specifically, the indirect component of the end - point audio signal may be determined based on the start - point audio signal propagating along the indirect path. Furthermore, the direct component of the end - point audio signal may be determined based on the start - point audio signal propagating along the direct path. Then, the end - point audio signal may be determined by combining the indirect component and the direct component. By doing so, the acoustic effects of the obstacle 603 can be taken into account in a precise and efficient manner.
[0101] Furthermore, method 910 may include determining focus information regarding the field of view 701 and / or area of interest 702 of listener 181. The endpoint audio signal may then be determined based on the focus information. Specifically, the spectral composition of the audio signal may be adapted depending on the focus information. By doing so, the VR experience of listener 181 may be further improved.
[0102] Furthermore, method 910 may include determining that audio sources 311, 312, 313 are ambience audio sources. In this context, an indicator (e.g., a flag) may be received within bitstream 140 from encoder 130. For example, the indicator indicates that audio sources 311, 312, 313 are ambience audio sources. Ambience audio sources typically provide background audio signals. The starting source position of the ambience audio source may be maintained as the endpoint source position. Alternatively or additionally, the intensity of the starting audio signal of the ambience audio source may be maintained as the intensity of the endpoint audio signal. By doing so, ambience audio sources can be handled efficiently and consistently in the context of local transition 192.
[0103] The above aspects are applicable to an audio scene 111 including a plurality of audio sources 311, 312, 313. In particular, the method 910 may include rendering a plurality of starting audio signals of the corresponding plurality of audio sources 311, 312, 313 from a plurality of different starting source positions on a starting sphere 114. Further, the method 910 may include determining, for each of the corresponding plurality of audio sources 311, 312, 313 on an ending sphere 114, a plurality of ending source positions based on the plurality of starting source positions. Further, the method 910 may include determining a plurality of ending audio signals of the corresponding plurality of audio sources 311, 312, 313 based on the plurality of starting audio signals. Then, the plurality of ending audio signals of the corresponding plurality of audio sources 311, 312, 313 may be rendered from corresponding plurality of ending source positions on the ending sphere 114 around an ending listening position 302.
[0104] Furthermore, a virtual reality audio renderer 160 for rendering audio signals in a virtual reality rendering environment 180 is described. The audio renderer 160 is configured to render starting audio signals of audio sources 311, 312, 313 from starting source positions on a starting sphere 114 around a starting listening position 301 of a listener 181 (in particular, using a 3D audio renderer 162 of the VR audio renderer 160).
[0105] Furthermore, the VR audio renderer 160 may be configured to determine that the listener 181 moves from the starting listening position 301 to the ending listening position 302. In response thereto, the VR audio renderer 160 may be configured to determine, for example, within a preprocessing unit 161 of the VR audio renderer 160, ending source positions of the audio sources 311, 312, 313 on an ending sphere 114 around the ending listening position 302 based on the starting source positions, and to determine ending audio signals of the audio sources 311, 312, 313 based on the starting audio signals.
[0106] Furthermore, the VR audio renderer 160 (e.g., the 3D audio renderer 162) may be configured to render the end audio signals of the audio sources 311, 312, 313 from the end source positions on the end sphere 114 around the end listening position 302.
[0107] Thus, the virtual reality audio renderer 160 may have a preprocessing unit 161 configured to determine the end source positions and the end audio signals of the audio sources 311, 312, 313. Furthermore, the VR audio renderer 160 may have a 3D audio renderer 162 configured to render the end audio signals of the audio sources 311, 312, 313. The 3D audio renderer 162 may be configured to adapt the rendering of the audio signals of the audio sources 311, 312, 313 on the (unit) sphere 114 around the listening positions 301, 302 of the listener 181 according to the rotational movement of the head of the listener 181 (to provide 3DoF within the rendering environment 180). On the other hand, the 3D audio renderer 162 may not be configured to adapt the rendering of the audio signals of the audio sources 311, 312, 313 according to the translational movement of the head of the listener 181. Thus, the 3D audio renderer 162 may be limited to 3DoF. Then, the translational DoF may be provided in an efficient manner using the preprocessing unit 161. Thereby, an overall VR audio renderer 160 with 6DoF is provided.
[0108] Furthermore, an audio encoder 130 configured to generate a bitstream 140 is described. The bitstream 140 represents audio signals of at least one audio source 311, 312, 313 and is generated to represent the positions of the at least one audio source 311, 312, 313 within a rendering environment 180. Additionally, the bitstream 140 may represent environment data 193 regarding the audio propagation characteristics of the audio within the rendering environment 180. By communicating the environment data 193 regarding the audio propagation characteristics, local transitions 192 within the rendering environment 180 can be enabled in a precise manner.
[0109] Furthermore, a bitstream 140 is described that represents audio signals of at least one audio source 311, 312, 313, the positions of the at least one audio source 311, 312, 313 within a rendering environment 180, and environment data 193 regarding the audio propagation characteristics of the audio within the rendering environment 180. Alternatively or additionally, the bitstream 140 may indicate whether the audio sources 311, 312, 313 are ambient audio sources 801.
[0110] FIG. 9d shows a flowchart of an exemplary method 920 for generating a bitstream. The method 920 includes determining 921 audio signals of at least one audio source 311, 312, 313. Additionally, the method 920 includes determining 922 position data regarding the positions of the at least one audio source 311, 312, 313 within a rendering environment 180. Further, the method 920 may include determining 923 environment data 193 regarding the audio propagation characteristics of the audio within the rendering environment 180. The method 920 further includes inserting 934 the audio signals, the position data, and the environment data 193 into the bitstream 140. Alternatively or additionally, an indicator of whether the audio sources 311, 312, 313 are ambient audio sources 801 may be inserted into the bitstream 140.
[0111] Therefore, in this paper, a virtual reality audio renderer 160 (corresponding method) for rendering an audio signal in a virtual reality rendering environment 180, and audio sources 311, 312, 313 are described. The audio renderer 160 has a 3D audio renderer 162 configured to render the audio signals of the audio sources 113, 311, 312, 313 from source positions on a sphere 114 around the listening positions 301, 302 of a listener 181 within the virtual reality rendering environment 180. Further, the virtual reality audio renderer 160 has a preprocessing unit 161 configured to determine new listening positions 301, 302 of the listener 181 within the virtual reality rendering environment 180 (within the same or different audio scenes 111, 112). Further, the preprocessing unit 161 is configured to update the audio signals and the source positions of the audio sources 113, 311, 312, 313 with respect to a sphere 114 around the new listening positions 301, 302. The 3D audio renderer 162 is configured to render the updated audio signals of the audio sources 311, 312, 313 from the updated source positions on a sphere 114 around the new listening positions 301, 302.
[0112] The methods and systems described in this paper may be implemented as software, firmware and / or hardware. Certain components may be implemented as software running on a digital signal processor or a microprocessor. Other components may be implemented as hardware or as an application specific integrated circuit. Signals encountered in the methods and systems described may be stored on a medium such as a random access memory or an optical storage medium. The signals may be transferred via a network such as a radio network, a satellite network, a wireless network or a wired network, for example the Internet. A typical apparatus using the methods and systems described in this paper is a portable electronic device or other consumer equipment used for storing and / or rendering audio signals.
[0113] The numbered examples (enumerated example, EE) of this manuscript are as follows: 〔EE1〕 A method (900) for rendering audio in a virtual reality rendering environment (180), the method comprising: · Rendering the start audio signal of the start audio source (113) of the start audio - scene (111) from a start - source position on a sphere (114) around the listening position (201) of the listener (181) (step 901); · Determining that the listener (181) moves from the listening position (201) in the start audio - scene (111) to a listening position (202) in a different end audio - scene (112) (step 902); · Applying a fade - out gain to the start audio signal to determine a modified start audio signal (step 903); · Rendering the modified start audio signal of the start audio source (113) from a start - source position on a sphere (114) around the listening positions (201, 202) (step 904), Method. 〔EE2〕 The method further comprising: · Determining that the listener (181) moves from the start audio - scene (111) to the end audio - scene (112) during a transition time interval; · Determining an intermediate time point (213) within the transition time interval; · Determining the fade - out gain based on the relative position of the intermediate time point (213) within the transition time interval, The method according to EE1. 〔EE3〕 · The method includes providing a fade - out function (211) that indicates the fade - out gain at various intermediate time points (213) within the transition time interval, · The fade - out function (211) is such that the fade - out gain decreases with the progressing intermediate time point (213), The method according to EE2. 〔EE4〕 The fade-out function (211) is such that: · the starting audio signal remains unmodified at the start of the transition time interval; and / or · the starting audio signal is increasingly attenuated at an intermediate time point (213) as it progresses; and / or · the starting audio signal is completely attenuated at the end of the transition time interval, a method according to EE3. 〔EE5〕 The method includes: · maintaining the starting source position of the starting audio source (113) on the sphere (114) around the listening position (201, 202) when the listener (181) moves from the starting audio scene (111) to the ending audio scene (112); and / or · maintaining the listening position (201, 202) unchanged when the listener (181) moves from the starting audio scene (111) to the ending audio scene (112), a method according to any one of EE1 to 4. 〔EE6〕 The method includes: · determining an ending audio signal of the ending audio source (113) of the ending audio scene (112); · determining an ending source position on the sphere (114) around the listening position (201, 202); · applying a fade-in gain to the ending audio signal to determine a modified ending audio signal; · rendering the modified ending audio signal of the ending audio source (113) from the ending source position on the sphere (114) around the listening position (201, 202), a method according to any one of EE1 to 5. 〔EE7〕 The method is: · determining that the listener (181) moves from the starting audio scene (111) to the ending audio scene (112) during the transition time interval; · determining an intermediate time point (213) within the transition time interval; · determining the fade-in gain based on the relative position of the intermediate time point (213) within the transition time interval, the method according to EE6. [EE8] · the method includes providing a fade-in function (212) indicating the fade-in gain at various intermediate time points (213) within the transition time interval, · the fade-in function (212) is such that the fade-in gain increases with the progressing intermediate time point (213), the method according to EE7. [EE9] the fade-in function (211) is · the ending audio signal remains unmodified at the end of the transition time interval; and / or · the ending audio signal becomes less and less attenuated at the progressing intermediate time point (213); and / or · the ending audio signal is completely attenuated at the start of the transition time interval, the method according to EE8. [EE10] the method includes · maintaining the ending source position of the ending audio source (113) on the sphere (114) around the listening position (201, 202) when the listener (181) moves from the starting audio scene (111) to the ending audio scene (112); and / or · maintaining the listening position (201, 202) unchanged when the listener (181) moves from the starting audio scene (111) to the ending audio scene (112), the method according to any one of EE6 to 9. [EE11] The method described in EE8 when EE8 cites EE3, wherein the fade-out function (211) and the fade-in function (212) are combined to provide a constant gain for a plurality of different intermediate time points (213). 〔EE12〕 The fade-out function (211) and / or the fade-in function (212) is · Derived from the bitstream (140) indicating the starting audio signal and / or the ending audio signal; and / or · Derived from the storage unit of the virtual reality audio renderer (160) configured to render the starting audio signal and / or the ending audio signal within the virtual reality rendering environment (180). The method described in EE8 when EE8 cites EE3. 〔EE13〕 The method according to any one of EE1 to 12, wherein the method includes receiving an indication that the listener (181) moves from the starting audio scene (111) to the ending audio scene (112). 〔EE14〕 The method according to EE13, wherein the indication includes a flag. 〔EE15〕 The method according to any one of EE1 to 14, wherein the method includes sending an indication that the listener (181) moves from the starting audio scene (111) to the ending audio scene (112) to the encoder (130); and the encoder (130) is configured to generate a bitstream (140) indicating the starting audio signal. 〔EE16〕 The method according to any one of EE1 to 15, wherein the first audio signal is rendered using a 3D audio renderer (162), particularly an MPEG-H audio renderer. 〔EE17〕 The method includes · Rendering a plurality of start audio signals of a corresponding plurality of start audio sources (113) from a plurality of different start source positions on a sphere (114) around the listening position (201, 202); · Applying the fade-out gain to the plurality of start audio signals to determine a plurality of modified start audio signals; · Rendering the plurality of modified start audio signals of the start audio source (113) from the corresponding plurality of start source positions on a sphere (114) around the listening position (201, 202), including The method according to any one of EE1 to 16. 〔EE18〕 The method is · Determining a plurality of end audio signals of a corresponding plurality of end audio sources (113) of the end audio scene (112); · Determining a plurality of end source positions on a sphere (114) around the listening position (201, 202); · Applying the fade-in gain to the plurality of end audio signals to determine a corresponding plurality of modified end audio signals; · Rendering the plurality of modified end audio signals of the plurality of end audio sources (113) from the corresponding plurality of end source positions on a sphere (114) around the listening position (201, 202), including The method according to any one of EE6 to 17. 〔EE19〕 The method according to any one of EE1 to 18, wherein the start audio signal is an overlap of audio signals of a plurality of start audio sources (113). 〔EE20〕 A virtual reality audio renderer (160) for rendering audio in a virtual reality rendering environment (180), the virtual reality audio renderer (160) being · Rendering a start audio signal of a start audio source (113) of a start audio scene (111) from a start source position on a sphere (114) around a listening position (201) of a listener (181); · Determine that a listener (181) moves from a listening position (201) within a starting audio scene (111) to a listening position (202) within a different ending audio scene (112); · Apply a fade - out gain to the starting audio signal to determine a modified starting audio signal; · Render the modified starting audio signal from the starting source position on a sphere (114) around the listening positions (201, 202) of the starting audio source (113), A virtual reality audio renderer. 〔EE21〕 An encoder (130) configured to generate a bitstream (140) representing an audio signal to be rendered within a virtual reality rendering environment (180), the encoder (130) comprising: · Determine a starting audio signal of a starting audio source (113) of a starting audio scene (111); · Determine starting position data regarding the starting source position of the starting audio source (113); · Generate a bitstream (140) including the starting audio signal and the starting position data; · Receive an indication that a listener (181) moves from the starting audio scene (111) to an ending audio scene (112) within the virtual reality rendering environment (180); · Determine an ending audio signal of an ending audio source (113) of the ending audio scene (112); · Determine ending position data regarding the ending source position of the ending audio source (113); · Generate a bitstream (140) including the ending audio signal and the ending position data. Encoder. 〔EE22〕 A method (930) for generating a bitstream (140) representing an audio signal to be rendered within a virtual reality rendering environment (180), the method comprising: ·Determine the starting audio signal of the starting audio source (113) of the starting audio scene (111) (931); ·Determine the starting position data regarding the starting source position of the starting audio source (113) (932); ·Generate a bitstream (140) including the starting audio signal and the starting position data (933); ·Receive an indicator that a listener (181) moves from the starting audio scene (111) to an ending audio scene (112) within the virtual reality rendering environment (180) (934); ·Determine the ending audio signal of the ending audio source (113) of the ending audio scene (112) (935); ·Determine the ending position data regarding the ending source position of the ending audio source (113) (936); ·Generate a bitstream (140) including the ending audio signal and the ending position data (937), Method. 〔EE23〕 A virtual reality audio renderer (160) for rendering an audio signal in a virtual reality rendering environment (180), the audio renderer includes ·A 3D audio renderer (162) configured to render the audio signal of an audio source (113) from a source position on a sphere (114) around a listening position (201, 202) of a listener (181) within the virtual reality rendering environment (180); ·A preprocessing unit (161), ·Determine a new listening position (201, 202) of a listener (181) within the virtual reality rendering environment (180); ·The preprocessing unit (161) configured to update the audio signal and the source position of the audio source (201, 202) regarding a sphere (114) around the new listening position (201, 202), The 3D audio renderer (162) is configured to render the updated audio signal of the audio source (113) from an updated source position on a sphere (114) around the new listening positions (201, 202). Virtual reality audio renderer.
Claims
1. A method for rendering audio in a virtual reality rendering environment using an audio renderer for rendering three degrees of freedom (3DoF), the method comprising: rendering, by the audio renderer, an audio signal of a starting audio source of a starting audio scene from a starting source position on a sphere around a starting listening position of a listener within the virtual reality rendering environment; determining the presence of movement of the listener, the movement being a movement from the starting listening position within the starting audio scene to an ending listening position within an ending audio scene within the virtual reality rendering environment; determining a modified starting audio signal by applying a fade-out gain to the starting audio signal based on the determination of the movement; determining an audio signal of an ending audio source of the ending audio scene; determining an ending source position on a sphere around the ending listening position; determining a modified ending audio signal by applying a fade-in gain to the ending audio signal; rendering, by the audio renderer, the modified starting audio signal of the starting audio source from the starting source position on the sphere around the starting listening position; rendering, by the audio renderer, the modified ending audio signal of the ending audio source from the ending source position on the sphere around the ending listening position, wherein the method further comprises: determining that the listener moves from the starting audio scene to the ending audio scene during a transition time interval; determining an intermediate time point within the transition time interval; determining the fade-out gain based on a relative position of the intermediate time point within the transition time interval. A method.
2. The method according to claim 1, wherein the modified starting audio signal is rendered from the same position with respect to the listener throughout the movement from the starting listening position within the starting audio scene to the ending listening position within the ending audio scene.
3. The method according to claim 1, wherein the ending audio scene does not include the starting audio source.
4. A non-transitory computer-readable storage medium storing executable instructions for causing a computer to execute the method according to claim 1.
5. A system for rendering audio in a virtual reality rendering environment using an audio renderer for rendering three degrees of freedom (3DoF), the system comprising: A first renderer that renders, by the audio renderer, an audio signal of a starting audio source of a starting audio scene from a starting source position on a sphere around a starting listening position of a listener within the virtual reality rendering environment; A first processor for determining the presence of movement of the listener, the movement being a movement from the starting listening position in the starting audio scene to an ending listening position in an ending audio scene within the virtual reality rendering environment; A second processor for determining a modified starting audio signal by applying a fade-out gain to the starting audio signal based on the determination of the movement; A third processor for determining an audio signal of an ending audio source of the ending audio scene; A fourth processor for determining an ending source position on a sphere around the ending listening position; A third processor for determining a modified ending audio signal by applying a fade-in gain to the ending audio signal; A second renderer that renders, by the audio renderer, the modified starting audio signal of the starting audio source from the starting source position on a sphere around the starting listening position;
Citation Information
Patent Citations
Video game processing device, and video game processing program
JP2014222306A
Method, computer readable storage medium and apparatus for determining target sound scene at target position from two or more source sound scenes
JP2017188873A
Spatialized audio output based on predicted position data
US20170295446A1