Method and system for handling global transitions between listening positions in virtual reality environment

The method addresses the challenge of translational audio rendering in VR by using fade-out and fade-in gains on pre-rendered virtual objects, enabling efficient 6 DoF transitions with consistent audio quality.

JP2025165970APending Publication Date: 2025-11-05DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025120162
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-12-18
Filing Date
2025-07-17
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

Existing audio rendering systems, such as MPEG-H 3D renderers, are limited to handling rotational movements (3 DoF) and struggle to efficiently manage translational changes in listening positions, which are crucial for providing a realistic 6 DoF experience in virtual reality environments.

Method used

A method and system that applies fade-out and fade-in gain functions to origin and destination audio signals during transitions between audio scenes, utilizing pre-rendered virtual objects to maintain a consistent audio experience while reducing computational overhead.

Benefits of technology

Enables efficient 6 DoF audio rendering by smoothly transitioning between audio scenes with reduced computational requirements, ensuring a seamless and realistic audio experience in virtual reality environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025165970000001_ABST
    Figure 2025165970000001_ABST
Patent Text Reader

Abstract

To provide a method and a system to handle a transition between a listening position in a virtual reality rendering environment in an efficient and consistent manner.SOLUTION: A method comprises the steps of: rendering an origin audio signal of an origin audio source of an origin audio scene 111 from an origin source position on a sphere around a listening position 201 of a listener; determining that the listener moves from the listening position in the origin audio scene to a listening position 202 in a different end audio scene 112; applying a fade-out gain derived from a fade-out function 211 to the origin audio signal to determine a modified origin audio signal; applying a fade-in gain derived from a fade-in function 212 to an end point audio signal to determine a modified end point audio signal; rendering the modified origin audio signal of the origin audio source from the origin source position on the sphere around the listening position.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 62 / 599,841 (Docket No. D17085USP1), filed December 18, 2017, and European Application No. 17208088.9 (Docket No. D17085EP), filed December 18, 2017, the contents of which are incorporated herein by reference.

[0002] Technical Field This paper concerns handling transitions between auditory viewports and / or listening positions in a virtual reality (VR) rendering environment in an efficient and consistent manner. [Background technology]

[0003] Virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications are rapidly evolving to include increasingly sophisticated acoustic models of sound sources and scenes that can be enjoyed from different viewpoints / perspectives or listening positions. Two different classes of flexible audio representations are sometimes used, for example, for VR applications: sound field representations and object-based representations. Sound field representations are physically based methods that encode the wavefront incident at the listening position. For example, methods such as B-format or Higher Order Ambisonics (HOA) represent spatial wavefronts using spherical harmonic decomposition. Object-based methods represent complex auditory scenes as a collection of single elements that contain audio waveforms or signals and possibly time-varying associated parameters or metadata.

[0004] Enjoying VR, AR, and MR applications can involve users experiencing different auditory perspectives or viewpoints. For example, room-based virtual reality may be provided based on mechanisms using six degrees of freedom (DoF). Figure 1 shows an example of 6 DoF interaction, illustrating translational movement (forward / backward, up / down, and left / right) and rotational movement (pitch, yaw, and roll). Unlike 3 DoF spherical video experiences, which are limited to head rotation, content created for 6 DoF interaction allows for navigation within the virtual environment (e.g., physically walking around a room) in addition to head rotation. This can be achieved based on position trackers (e.g., camera-based) and orientation trackers (e.g., gyroscopes and / or accelerometers). 6 DoF tracking technology may be available on high-end mobile VR platforms (e.g., Google Tango) as well as on PlayStation®VR, Oculus Rift, and HTC Vive. The user's experience of the directionality and spatial extent of a sound or audio source is crucial to the realism of the 6 DoF experience, especially the experience of navigation within a scene and around a virtual audio source.

[0005] Available audio rendering systems (e.g., MPEG-H 3D renderers) are typically limited to rendering 3 DoF (i.e., rotational movement of the audio scene caused by the listener's head movement). Translational changes in the listener's listening position and the associated DoF typically cannot be handled by such renderers. Summary of the Invention [Problem to be solved by the invention]

[0006] This paper addresses the technical challenge of providing a resource-efficient method and system for handling translation in the context of audio rendering. [Means for solving the problem]

[0007] According to one aspect, a method for rendering audio in a virtual reality rendering environment is described. The method includes rendering an origin audio signal of an origin audio source of an origin audio scene from an origin source position on a sphere around a listening position of a listener. The method further includes determining that the listener moves from the listening position in the origin audio scene to a listening position in a different destination audio scene. In addition, the method includes applying a fade-out gain to the origin audio signal to determine a modified origin audio signal. The method further includes rendering the modified origin audio signal of the origin audio source from the origin source position on the sphere around the listening position.

[0008] According to a further aspect, a virtual reality audio renderer for rendering audio in a virtual reality rendering environment is described. The virtual reality audio renderer is configured to render an origin audio signal of an origin audio source in an origin audio scene from an origin source position on a sphere around a listening position of a listener. Additionally, the virtual reality audio renderer is configured to determine when the listener moves from the listening position in the origin audio scene to a listening position in a different destination audio scene. Furthermore, the virtual reality audio renderer is configured to apply a fade-out gain to the origin audio signal to determine a modified origin audio signal and render the modified origin audio signal of the origin audio source from an origin source position on a sphere around the listening position.

[0009] According to a further aspect, a method for generating a bitstream indicative of an audio signal to be rendered in a virtual reality rendering environment is described, the method including: determining an origin audio signal of an origin audio source of an origin audio scene; determining origin position data related to an origin source position of the origin audio source; generating a bitstream including the origin audio signal and the origin position data; receiving an indication that a listener moves from the origin audio scene to an end audio scene in the virtual reality rendering environment; determining an end audio signal of an end audio source of the end audio scene; determining end position data related to an end source position of the end audio source; and generating a bitstream including the end audio signal and the end position data.

[0010] According to another aspect, an encoder configured to generate a bitstream indicative of an audio signal to be rendered in a virtual reality rendering environment is described, the encoder being configured to: determine an origin audio signal for an origin audio source of an origin audio scene; determine origin position data related to an origin source position of the origin audio source; generate a bitstream including the origin audio signal and the origin position data; receive an indication that a listener moves from the origin audio scene to an end audio scene in the virtual reality rendering environment; determine an end audio signal for an end audio source of the end audio scene; determine end position data related to an end source position of the end audio source; and generate a bitstream including the end audio signal and the end position data.

[0011] According to a further aspect, a virtual reality audio renderer for rendering an audio signal in a virtual reality rendering environment is described. The audio renderer includes a 3D audio renderer configured to render an audio signal of an audio source from a source position on a sphere around a listening position of a listener within the virtual reality rendering environment. The virtual reality audio renderer further includes a pre-processing unit configured to determine a new listening position of the listener within the virtual reality rendering environment. The pre-processing unit is further configured to update the audio signal and the source position of the audio source relative to the sphere around the new listening position. The 3D audio renderer is configured to render the updated audio signal of the audio source from the updated source position on the sphere around the new listening position.

[0012] According to a further aspect, a software program is described.

[0013] The software program may be adapted for execution on a processor and may be adapted to perform the method steps outlined herein when executed on the processor.

[0014] According to another aspect, a storage medium is described, which may have a software program adapted for execution on a processor and adapted to perform the method steps outlined herein when executed on the processor.

[0015] According to a further aspect, a computer program product is described, which may include executable instructions for performing the method steps outlined herein when executed on a computer.

[0016] The methods and systems, including the preferred embodiments, outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief explanation of the drawings]

[0017] The invention is described below, by way of example, with reference to the accompanying drawings, in which: [Figure 1a] 1 illustrates an exemplary audio processing system that provides 6DoF audio. [Figure 1b] 1 illustrates an exemplary situation within a 6DoF audio and / or rendering environment. [Figure 1c] 1 shows an exemplary transition from a source audio scene to a destination audio scene. [Figure 2] 1 illustrates an exemplary scheme for determining a spatial audio signal during a transition between different audio scenes. [Figure 3] An exemplary audio scene is shown. [Figure 4a] Demonstrates the remapping of audio sources in response to changes in listening position within the audio scene. [Figure 4b] 1 shows an exemplary distance function. [Figure 5a] Illustrates an audio source with a non-uniform directional profile. [Figure 5b] 1 illustrates an exemplary directivity function for an audio source. [Figure 6] An example audio scene with acoustically significant obstructions is shown. [Figure 7] Indicates the listener's field of view and focus of attention. [Figure 8] Demonstrates the treatment of ambient audio when the listening position changes within the audio scene. [Figure 9a]1 shows a flowchart of an exemplary method for rendering a 3D audio signal during a transition between different audio scenes. [Figure 9b] 1 shows a flowchart of an exemplary method for generating a bitstream for transitions between different audio scenes. [Figure 9c] 1 shows a flowchart of an exemplary method for rendering a 3D audio signal during a transition in an audio scene. [Figure 9d] 1 shows a flowchart of an exemplary method for generating a bitstream for a local transition. DETAILED DESCRIPTION OF THE INVENTION

[0018] As outlined above, this paper relates to efficiently providing 6DoF in a 3D (three-dimensional) audio environment. FIG. 1a shows a block diagram of an exemplary audio processing system 100. A stadium-like acoustic environment 110 includes a variety of different audio sources 113. Exemplary audio sources 113 in a stadium include individual spectators, stadium speakers, and players on the field. The acoustic environment 110 may be subdivided into different audio scenes 111, 112. For example, a first audio scene 111 may correspond to a home team cheering block, and a second audio scene 111 may correspond to a guest team cheering block. Depending on where the listener is located within the audio environment, the listener perceives audio sources 113 from the first audio scene 111 or audio sources from the second audio scene 112.

[0019] The different audio sources 113 of the audio environment 110 may be captured using audio sensors 120, in particular using a microphone array. In particular, said one or more audio scenes 111, 112 of the audio environment 110 may be described using multi-channel audio signals, one or more audio objects and / or Higher Order Ambisonics (HOA) signals. In the following, it is assumed that the audio sources 113 are associated with audio data captured by the audio sensors 120, where the audio data indicates the audio signal and the position of the audio source 113 as a function of time (at a particular sampling rate, for example 20 ms).

[0020] A 3D audio renderer, such as an MPEG-H 3D audio renderer, typically assumes that the listener is located at a particular listening position within the audio scene 111, 112. Audio data for the various audio sources 113 of the audio scene 111, 112 is typically provided under the assumption that the listener is located at this particular listening position. The audio encoder 130 may include a 3D audio encoder 131 configured to encode audio data for the audio sources 113 of one or more of the audio scenes 111, 112.

[0021] Additionally, virtual reality (VR) metadata may be provided, which allows a listener to change their listening position within the audio scenes 111, 112 and / or move between different audio scenes 111, 112. The encoder 130 may comprise a metadata encoder 132 configured to encode the VR metadata. The encoded VR metadata and the encoded audio data of the audio source 113 may be combined in a combination unit 133 to provide a bitstream 140 indicative of the audio data and the VR metadata. The VR metadata may include, for example, environmental data describing the acoustic characteristics of the audio environment 110.

[0022] The bitstream 140 may be decoded using a decoder 150 to provide (decoded) audio data and (decoded) VR metadata. An audio renderer 160 for rendering audio in a rendering environment 180 that allows 6DoF may include a preprocessing unit 161 and a (conventional) 3D audio renderer 162 (such as MPEG-H 3D audio). The preprocessing unit 161 may be configured to determine a listening position 182 of a listener 181 in the listening environment 180. The listening position 182 may indicate the audio scene 111 in which the listener 181 is located. Furthermore, the listening position 182 may indicate an exact position within the audio scene 111. The preprocessing unit 161 may further be configured to determine a 3D audio signal for the current listening position 182 based on the (decoded) audio data and possibly based on the (decoded) VR metadata. The 3D audio signal may then be rendered using a 3D audio renderer 162 .

[0023] It should be noted that the concepts and methods described herein may be specified in a frequency-varying manner, may be defined globally or in an object / media-dependent manner, may be applied directly in the spectral or time domain, and / or may be hard-coded into the VR renderer 160 or specified via a corresponding input interface.

[0024] FIG. 1b shows an exemplary rendering environment 180. A listener 181 may be located within an origin audio scene 111. For rendering purposes, audio sources 113, 194 may be assumed to be located at various rendering positions on a (unit) sphere 114 around the listener 181. The rendering positions of the various audio sources 113, 194 may change over time (according to a given sampling rate). Various situations may arise within the VR rendering environment 180: the listener 181 may perform a global transition 191 from the origin audio scene 111 to the destination audio scene 112. Alternatively or additionally, the listener 181 may perform a local transition 192 to a different listening position 182 within the same audio scene 111. Alternatively or additionally, the audio scene 111 may exhibit acoustically significant environmental features (e.g., walls), which may be described using environmental data 193 and should be taken into account when a change in listening position 182 occurs. Alternatively or additionally, the audio scene 111 may include one or more ambient audio sources 194 (e.g., for background noise), which should be taken into account when changes in listening position 182 occur.

[0025] FIG. 1c shows audio sources 113A1 to A n The source audio scene 111 has audio sources 113B1 to B m 1 shows an exemplary global transition 191 to a destination audio scene 112 having an audio source 113A1 through an audio source 113A2. In particular, each audio source 113 may be included in only one of the source audio scene 111 and the destination audio scene 112. For example, audio sources 113A1 through A1 n is included in the starting audio scene 111 but not included in the ending audio scene 112, and audio sources 113B1 to B mis included in the destination audio scene 112 but not in the origin audio scene 111. The audio source 113 may be characterized by corresponding inter-location object properties (coordinates, directivity, distance sound attenuation function, etc.). The global transition 191 may be performed within a transition time interval (e.g., within a range of 5 seconds, 1 second, or less). The listening position 182 in the origin scene 111 at the beginning of the global transition 191 is marked with "A". Furthermore, the listening position 182 in the destination scene 112 at the end of the global transition 191 is marked with "B". Furthermore, Figure 1c shows a local transition 192 in the destination scene 112 between listening positions "B" and "C".

[0026] 2 illustrates a global transition 191 from an origin scene 111 (or origin viewport) to an end scene 112 (or end viewport) during a transition time interval t. Such a transition 191 may occur when a listener 181 switches between different scenes or viewports 111, 112, for example, in a stadium. Thus, the global transition 191 from the origin scene 111 to the end scene 112 need not correspond to an actual physical movement of the listener 181, but may simply be initiated by the listener's command to switch or transition to another viewport 111, 112. Nevertheless, this disclosure refers to the listener's position, which is understood to be the listener's position in the VR / AR / MR environment. At the intermediate point 213, the listener 181 may be located at an intermediate position between the origin scene 111 and the destination scene 112. The 3D audio signal 203 rendered at the intermediate position and / or at the intermediate point 213 may be generated by extracting the audio signals from the audio sources 113A1 to A1 of the origin scene 111, taking into account the sound propagation of each audio source 113. n and the audio sources 113B1 to B of the end scene 112. m However, this would lead to a relatively high amount of computation (especially in the case of a relatively large number of audio sources 113).

[0027] At the beginning of the global transition 191, the listener 181 may be positioned at the origin listening position 201. During the entire transition 191, the 3D origin audio signal A G may be generated, where the origin audio signal only depends on the audio source 113 in the origin scene 111 (and not on the audio source 113 in the destination scene 112). The global transition 191 does not affect the apparent source position of the audio source 113 in the origin scene 111. Thus, assuming a static audio source 113 in the origin scene 111, the rendering position of the audio source 113 during the global transition 191 relative to the listening position 201 does not change (relative to the listener) as the listening position transitions from the origin scene to the destination scene. Furthermore, at the beginning of the global transition 191, it may be fixed that the listener 181 will arrive at the destination listening position 202 in the destination scene 112 at the end of the global transition 191. During the entire transition 191, the 3D destination audio signal B G may be generated for the destination listening position 202, where the destination audio signal only depends on the audio sources 113 in the destination scene 112 (and not on the audio sources 113 in the source scene 111). The global transition 191 does not affect the apparent source position (relative to the listener) of the audio sources 113 in the destination scene 112.

[0028] To determine an intermediate audio signal 203 at an intermediate position and / or intermediate time point 213 during the global transition 191, the originating audio signal at the intermediate time point 213 may be combined with the destination audio signal at the intermediate time point 213. In particular, a fade-out factor or gain derived from a fade-out function 211 may be applied to the originating audio signal. The fade-out function 211 may be such that a fade-out factor or gain "a" decreases within increasing distance of the intermediate position from the originating scene 111. Additionally, a fade-in factor or gain derived from a fade-in function 212 may be applied to the destination audio signal. The fade-in function 212 may be such that a fade-in factor or gain "b" increases with decreasing distance of the intermediate position from the destination scene 112. An exemplary fade-out function 211 and an exemplary fade-in function 212 are shown in FIG. 2. The intermediate audio signal may then be given by a weighted sum of the source and destination audio signals, the weights corresponding to the fade-out and fade-in gains, respectively.

[0029] Thus, a fade-in function or curve 212 and a fade-out function or curve 211 may be defined for the global transition 191 between different 3DoF viewports 201, 202. The functions 211, 212 may be applied to pre-rendered virtual objects or 3D audio signals representing the originating audio scene 111 and the destination audio scene 112. In this way, a consistent audio experience may be provided during the global transition 191 between different audio scenes 111, 112 with reduced VR audio rendering calculations.

[0030] intermediate position x i The intermediate audio signal 203 at may be determined using linear interpolation of the source and destination audio signals. The strength F of the audio signal is given by F(x i )=a*F(A G )+(1-a)*F(B GThe factors "a" and "b=1-a" may be given by a norm function a=a() that depends on the starting listening position 201, the ending listening position 202 and the intermediate positions.

[0031] As an alternative to a function, a look-up table a=[1,...,0] may be provided for various intermediate positions. In the above, to allow a smooth transition from the starting scene 111 to the ending scene 112, multiple intermediate positions x i It will be appreciated that for , the intermediate audio signal 203 can be determined and rendered.

[0032] During the global transition 191, additional effects (e.g., Doppler effect and / or reverberation) may be taken into account. The functions 211, 212 may be adapted by the content provider, for example to reflect artistic intent. Information about the functions 211, 212 may be included as metadata in the bitstream 140. Thus, the encoder 130 may be configured to provide information about the fade-in function 212 and / or the fade-out function 211 as metadata in the bitstream 140. Alternatively or additionally, the audio renderer 160 may apply the functions 211, 212 stored in the audio renderer 160.

[0033] A flag may be communicated from the listener to the renderer 160, particularly to the VR pre-processing unit 161, to indicate to the renderer 160 that a global transition 191 is being performed from the origin scene 111 to the destination scene 112. The flag may trigger the audio processing described herein to generate intermediate audio signals during the transition phase. The flag may be signaled explicitly or implicitly through related information (e.g., via the coordinates of a new viewport or listening position 202). The flag may be sent from any data interface side (e.g., server / content, user / scene, auxiliary). Together with the flag, the origin audio signal A Gand destination audio signal B G For example, the IDs of one or more audio objects or audio sources may be provided. Alternatively, a request may be provided to the renderer 160 to compute the source and / or destination audio signals.

[0034] Thus, a VR renderer 160 is described that has a pre-processing unit 161 for a 3DoF renderer 162, enabling 6DoF functionality in a resource-efficient manner. The pre-processing unit 161 allows the use of a standard 3DoF renderer 162, such as an MPEG-H 3D audio renderer. The VR pre-processing unit 161 generates pre-rendered virtual audio objects A, B, C, D, D, E ... G and B G The global transition 191 may be configured to efficiently perform the calculations for the global transition 191 using the pre-rendered virtual objects A and B. The amount of calculations is reduced by utilizing only two pre-rendered virtual objects during the global transition 191. Each virtual object may contain multiple audio signals for multiple audio sources. Furthermore, during the transition 191, the pre-rendered virtual audio object A G and B G The bit rate requirement may be reduced since only the .times. ...

[0035] 3DoF functionality may be provided for all intermediate positions along the global transition trajectory, which may be achieved by overlapping the origin and destination audio objects using fade-out / fade-in functions 211, 212. Furthermore, additional audio objects may be rendered and / or additional audio effects may be included.

[0036] 3 illustrates an exemplary local transition 192 from an origin listening position B 301 to an end listening position C 302 within the same audio scene 111. The audio scene 111 includes different audio sources or objects 311, 312, and 313. The different audio sources or objects 311, 312, and 313 may have different directional profiles 332. Furthermore, the audio scene 111 may have environmental characteristics, in particular one or more obstacles, that have an effect on the propagation of audio within the audio scene 111. The environmental characteristics may be described using environmental data 193. Furthermore, the relative distances 321 and 322 of the audio object 311 to the listening positions 301 and 302 may be known.

[0037] 4a and 4b show a scheme for handling the effect of a local transition 192 on the intensity of different audio sources or objects 311, 312, 313. As outlined above, the audio sources 311, 312, 313 of the audio scene 111 are typically assumed by the 3D audio renderer 162 to be located on a sphere 114 around the listening position 301. Thus, at the beginning of the local transition 192, the audio sources 311, 312, 313 may be located on the origin sphere 114 around the origin listening position 301, and at the end of the local transition 192, the audio sources 311, 312, 313 may be located on the destination sphere 114 around the destination listening position 302.

[0038] The audio sources 311, 312, 313 may be remapped from the origin sphere 114 to the destination sphere 114. For this purpose, a ray may be considered that goes from the destination listening position 302 to the source positions of the audio sources 311, 312, 313 on the origin sphere 114. The audio sources 311, 312, 313 may be positioned at the intersection of the ray with the destination sphere 114.

[0039] The intensity F of the audio sources 311, 312, 313 on the destination sphere 114 is typically different from their intensity on the origin sphere 114. The intensity F may be modified using an intensity gain function or distance function 415 that provides a distance gain 410 as a function of the distance 420 of the audio sources 311, 312, 313 from the listening positions 301, 302. The distance function 415 typically indicates a cutoff distance 421 beyond which a zero distance gain 410 is applied. The origin distance 321 of the audio source 311 to the origin listening position 301 provides the origin gain 411. Furthermore, the destination distance 322 of the audio source 311 to the destination listening position 302 provides the destination gain 412. The intensity F of the audio source 311 may be rescaled using the origin gain 411 and the destination gain 412 to provide the intensity F of the audio source 311 on the destination sphere 114. In particular, the strength F of the origin audio signal of the audio source 311 on the origin sphere 114 may be divided by the origin gain 411 and multiplied by the destination gain 412 to give the strength F of the destination audio signal of the audio source 311 on the destination sphere 114.

[0040] Thus, the position of the audio source 311 after the local transition 192 is given by (e.g., using a geometric transformation) C i =source_remap_function(B i , C). Furthermore, the intensity of the audio source 311 after the local transition 192 may be determined as F(C i )=F(B i )*distance_function(B i ,C i , C). Thus, the distance attenuation can be modeled by a corresponding intensity gain given by the distance function 415.

[0041] 5a and 5b illustrate an audio source 312 with a non-uniform directivity profile 332. The directivity profile may be defined using a directivity gain 510 that indicates gain values ​​for various directions or directivity angles 520. In particular, the directivity profile of the audio source 312 may be defined using a directivity gain function 515 that indicates the directivity gain 510 as a function of the directivity angle 520 (where the angle 520 may range from 0° to 360°). For a 3D audio source 312, the directivity angle 520 is typically a two-dimensional angle that includes an azimuth angle and an elevation angle. Thus, the directivity gain function 515 is typically a two-dimensional function of the two-dimensional directivity angle 520.

[0042] The directivity profile 332 of the audio source 312 may be taken into account in the context of the local transition 192 by determining an origin directivity angle 521 of the origin ray between the audio source 312 and the origin listening position 301 (the audio source is positioned on the origin sphere 114 around the origin listening position 301) and an end directivity angle 522 of the end ray between the audio source 312 and the end listening position 302 (the audio source is positioned on the end sphere 114 around the end listening position 302). Using a directivity gain function 515 of the audio source 312, the origin directivity gain 511 and the end directivity gain 512 may be determined as function values ​​of the directivity gain function 515 for the origin directivity angle 521 and the end directivity angle 522, respectively (see FIG. 5b). The strength F of the audio source 312 at the origin listening position 301 may then be divided by the origin directivity gain 511 and multiplied by the destination directivity gain 512 to determine the strength F of the audio source 312 at the destination listening position 302.

[0043] Thus, sound source directivity may be parameterized by a directional factor or gain 510, indicated by a directional gain function 515. The directional gain function 515 may indicate the strength of an audio source 312 at some distance as a function of the angle 520 relative to the listening positions 301, 302. The directional gain 510 may be defined as the ratio to the gain of an audio source 312 at the same distance and with the same total power, where the total power is radiated uniformly in all directions. The directional profile 332 may be parameterized by a set of gains 510 corresponding to vectors emanating from the center of the audio source 312 and terminating at points distributed on a unit sphere around the center of the audio source 312. The directional profile 332 of the audio source 312 may depend on the use case scenario and available data (e.g., a uniform distribution for a 3D flight case, a flattened distribution for a 2D+ use case, etc.).

[0044] The resulting audio intensity of the audio source 312 at the endpoint listening position 302 is F(C i )=F(B i )*Distance_function()*Directivity_gain_function(C i , C,Directivity_parametrization), where Directivity_gain_function depends on the directional profile 332 of the audio source 312. Distance_function() takes into account the modified intensity caused by the change in distance 321, 322 of the audio source 312 due to the transition of the audio source 312.

[0045] 6 shows an example obstacle 603 that may need to be taken into account in the context of the local transition 192 between different listening positions 301, 302. Specifically, the audio source 313 may be hidden behind the obstacle 603 at the end listening position 302. The obstacle 603 may be described by the environment data 193, which includes a set of parameters, such as the spatial dimensions of the obstacle 603 and an obstacle attenuation function that indicates the attenuation of sound caused by the obstacle 603.

[0046] The audio source 313 may indicate an obstacle-free distance 602 (OSD) to the end listening position 302. The OFD 602 ​​may indicate the length of the shortest path between the audio source 313 and the end listening position 302 that does not pass through an obstacle 603. Additionally, the audio source 313 may indicate a going-through distance 601 (GHD) to the end listening position 302. The GHD 601 may indicate the length of the shortest path between the audio source 313 and the end listening position 302 that typically passes through an obstacle 603. An obstacle attenuation function may be a function of the OFD 602 ​​and the GHD 601. Additionally, the obstacle attenuation function may be a function of the intensity F(B i ) may also be a function of

[0047] Audio source C at endpoint listening position 302 i The intensity of may be a combination of sound from the audio source 313 passing around the obstacle 603 and sound from the audio source 313 passing through the obstacle 603 .

[0048] Thus, the VR renderer 160 may be provided with parameters to control the influence of the environment geometry and media. The obstacle geometry / media data 193 or parameters may be provided by the content provider and / or the encoder 130. The audio intensity of the audio source 313 is: F(C i )=F(B i)*Distance_function(OFD)*Directivity_gain_function(OFD)+Obstacle_attenuation_function(F(Bi),OFD,GHD). The first term corresponds to the contribution of sound that goes around the obstacle 603. The second term corresponds to the contribution of sound that passes through the obstacle 603.

[0049] The minimum Obstacle Free Distance (OFD) 602 may be determined using an A* Dijkstra pathfinding algorithm and may be used to control direct sound attenuation. The Ground Path Distance (GHD) 601 may be used to control reverberation and distortion. Alternatively or additionally, ray casting techniques may be used to describe the effect of obstacles 603 on the intensity of the audio source 313.

[0050] FIG. 7 illustrates an exemplary field of view 701 of a listener 181 located at an end listening position 302. Additionally, FIG. 7 illustrates an exemplary focus of attention 702 of a listener located at an end listening position 302. The field of view 701 and / or focus of interest 702 may be used to enhance (e.g., amplify) audio coming from audio sources within the field of view 701 and / or focus of interest 702. The field of view 701 may be considered a user-driven effect and may be used to enable sound enhancer for audio sources 311 associated with the user's field of view 701. In particular, a "cocktail party effect" simulation may be performed by removing frequency tiles from background audio sources to improve the intelligibility of speech signals associated with audio sources 311 within the listener's field of view 701. The attention focus 702 may be considered a content-driven effect and may be used to enable sound enhancer for audio sources 311 associated with a content region of interest (e.g., to draw the user's attention to look and / or move in the direction of the audio sources 311).

[0051] The audio intensity of audio source 311 is: F(Bi )=Field_of_view_function(C,F(B i ), Field_of_view_data), where Field_of_view_function describes the modification applied to the audio signal of the audio source 311 that is within the field of view 701 of the listener 181. Furthermore, the audio intensity of the audio source that is within the listener's focus of interest 702 may be modified as: F(B i )=Attention_focus_function(F(B i ), Attention_focus_data), where attention_focus_function describes the modification to be applied to the audio signal of the audio source 311 that is within the focus of interest 702.

[0052] The functions described herein for handling the transition of the listener 181 from the origin listening position 301 to the destination listening position 302 may be applied in a similar manner to the change in position of the audio sources 311, 312, 313.

[0053] Thus, this paper describes efficient means for calculating the coordinates and / or audio intensity of virtual audio objects or audio sources 311, 312, 313 representing the local VR audio scene 111 at any listening position 301, 302. The coordinates and / or intensity may be determined taking into account sound source distance attenuation curves, sound source orientation and directionality, environmental geometry / media effects, and / or "field of view" and "focus of interest" data for additional audio signal enhancement. The described methods may significantly reduce the amount of computation by only performing calculations when the listening positions 301, 302 and / or audio objects / sources 311, 312, 313 change position.

[0054] Additionally, this paper describes concepts for specifying distance, directivity, geometric functions, processing, and / or signaling mechanisms for the VR renderer 160. Additionally, concepts are described for minimum "obstruction-free distance" to control direct sound attenuation and "passing distance" to control reverberation and distortion. Additionally, concepts are described for sound source directivity parameterization.

[0055] 8 illustrates the handling of ambient sound sources 801, 802, 803 in the context of local transition 192. Specifically, FIG. 8 illustrates three different ambient sound sources 801, 802, 803, where the ambient sounds may be attributed to point audio sources. An ambient sound flag may be provided to pre-processing unit 161 to indicate that point audio source 311 is ambient sound audio source 801. Processing during local and / or global transitions of listening positions 301, 302 may depend on the value of the ambient sound flag.

[0056] In the context of the global transition 191, the ambient sound source 801 may be treated like a normal audio source 311. Figure 8 shows the local transition 192. The positions of the ambient sound sources 801, 802, 803 may be copied from the origin sphere 114 to the destination sphere 114, thereby giving the positions of the ambient sound sources 811, 812, 813 at the destination listening position 302. Furthermore, if the environmental conditions remain unchanged, the intensity of the ambient sound source 801 may remain unchanged. That is, F(C Ai )=F(B Ai On the other hand, in the case of the obstacle 603, the intensity of the ambient sound sources 803, 813 is calculated using the obstacle attenuation function, e.g., F(C Ai )=F(B Ai )*Distance_function Ai (OFD)+Obstacle_attenuation_function(F(B Ai ), OFD, GHD).

[0057] FIG. 9a shows a flowchart of an exemplary method 900 for rendering audio in a virtual reality rendering environment 180. The method 900 may be performed by a VR audio renderer 160. The method 900 includes rendering 901 an origin audio signal of an audio source 113 of an origin audio scene 111 from an origin source position on a sphere 114 around a listening position 201 of a listener 181. The rendering 901 may be performed using a 3D audio renderer 162 that may be limited to handling only 3DoF, and in particular, limited to handling only rotational movement of the listener's 181 head. In particular, the 3D audio renderer 162 is not configured to handle translational movement of the listener's head. The 3D audio renderer 162 may include or be an MPEG-H audio renderer.

[0058] Note that the phrase "rendering the audio signal of the audio source 113 from a particular source position" indicates that a listener perceives the audio signal as coming from that particular source position. This phrase should not be understood as a limitation on how the audio signal is actually rendered. A variety of different rendering techniques may be used to "render the audio signal from a particular source position," i.e., to provide the listener 181 with the perception that the audio signal is coming from a particular source position.

[0059] Further, the method 900 includes determining 902 that the listener 181 moves from a listening position 201 in the source audio scene 111 to a listening position 202 in a different destination audio scene 112. Thus, a global transition 191 from the source audio scene 111 to the destination audio scene 112 may be detected. In this context, the method 900 may include receiving an indication that the listener 181 moves from the source audio scene 111 to the destination audio scene 112. The indication may include or be a flag. The indication may be communicated from the listener 181 to the VR Audio Renderer 160, for example, via a user interface of the VR Audio Renderer 160.

[0060] Typically, the origin audio scene 111 and the destination audio scene 112 each include one or more different audio sources 113. In particular, the origin audio signals of the one or more origin audio sources 113 may not be audible in the destination audio scene 112 and / or the destination audio signals of the one or more destination audio sources 113 may not be audible in the origin audio scene 111.

[0061] The method 900 may include applying 903 a fade-out gain to the origin audio signal to determine a modified origin audio signal (in response to determining that the global transition 191 to the new destination audio scene 112 is to be performed). In particular, the origin audio signal is generated such that it would be perceived at a listening position within the origin audio scene, regardless of a movement of the listener 181 from the listening position 201 within the origin audio scene 111 to the listening position 202 within the destination audio scene 112. Furthermore, the method 900 may include rendering 904 the modified origin audio signal of the origin audio source 113 from an origin source position on a sphere 114 around the listener positions 201, 202 (in response to determining that the global transition 191 to the new destination audio scene 112 is to be performed). These operations may be performed repeatedly during the global transition 191, for example at regular time intervals.

[0062] Thus, a global transition 191 between different audio scenes 111, 112 may be performed by gradually fading out the origin audio signal of the one or more origin audio sources 113 of the origin audio scene 111. This results in a computationally efficient and acoustically consistent global transition 191 between different audio scenes 111, 112.

[0063] It may be determined that the listener 181 moves from the originating audio scene 111 to the destination audio scene 112 during a transition time interval, where the transition time interval typically has a duration (e.g., 2 s, 1 s, 500 ms, or less). A global transition 191 may be performed gradually within the transition time interval. Specifically, during the global transition 191, an intermediate point 213 within the transition time interval may be determined (e.g., according to a sampling rate of 100 ms, 50 ms, 20 ms, or less). A fade-out gain may then be determined based on the relative position of the intermediate point 213 within the transition time interval.

[0064] Specifically, the transition time interval for the global transition 191 may be subdivided into a sequence of intermediate points in time 213. For each intermediate point in time 213 in the sequence of intermediate points in time 213, a fade-out gain may be determined for modifying the origin audio signal of the one or more origin audio sources. Furthermore, at each intermediate point in time 213 in the sequence of intermediate points in time 213, the modified origin audio signal of the one or more origin audio sources 113 may be rendered from an origin source position on a sphere 114 around the listening positions 201, 202. In this way, an acoustically consistent global transition 191 may be performed in a computationally efficient manner.

[0065] The method 900 may include providing a fade-out function 211 indicating a fade-out gain at various intermediate points 213 within the transition time interval, where the fade-out function 211 is typically such that the fade-out gain decreases with the progressing intermediate points 213, thereby providing a smooth global transition 191 to the destination audio scene 112. Specifically, the fade-out function 211 may be such that the source audio signal remains unmodified at the beginning of the transition time interval, the source audio signal is increasingly attenuated at the progressing intermediate points 213, and / or the source audio signal is completely attenuated at the end of the transition time interval.

[0066] The origin source position of the origin audio source 113 on the sphere 114 around the listening positions 201, 202 may be maintained (in particular for the entire transition time interval) as the listener 181 moves from the origin audio scene 111 to the destination audio scene 112. Alternatively or additionally, it may be assumed that the listener 181 stays at the same listening position 201, 202 (for the entire transition time interval). This may further reduce the amount of computation for the global transition 191 between audio scenes 111, 112.

[0067] The method 900 may further include determining an end audio signal of the end audio source 113 of the end audio scene 112. Additionally, the method 900 may include determining an end source position on a sphere 114 around the listening positions 201, 202. In particular, the end audio signal is generated such that it will be perceived at the listening position in the end audio scene, regardless of a movement of the listener 181 from the listening position 201 in the origin audio scene 111 to the listening position 202 in the end audio scene 112. Additionally, the method 900 may include applying a fade-in gain to the end audio signal to determine a modified end audio signal. The modified end audio signal of the end audio source 113 may then be rendered from the end source position on the sphere 114 around the listening positions 201, 202. These operations may be performed repeatedly during the global transition 191, for example, at regular time intervals.

[0068] Thus, similar to the fading out of the origin audio signals of the one or more origin audio sources 113 of the origin scene 111, the destination audio signals of one or more destination audio sources 113 of the destination scene 112 may be faded in, thereby providing a smooth global transition 191 between the audio scenes 111, 112.

[0069] As described above, the listener 181 may move from the originating audio scene 111 to the destination audio scene 112 during the transition time interval. Fade-in gains may be determined based on the relative positions of the intermediate time points 213 within the transition time interval. Specifically, a sequence of fade-in gains may be determined for a corresponding sequence of intermediate time points 213 during the global transition 191.

[0070] The fade-in gain may be determined using a fade-in function 212 that indicates the fade-in gain at various intermediate points 213 within the transition time interval, where the fade-in function 212 is typically such that the fade-in gain increases with the progressing intermediate point 213. In particular, the fade-in function 212 may be such that the end audio signal is fully attenuated at the beginning of the transition time interval, becomes less attenuated at the progressing intermediate points 213 of the end audio signal, and / or the end audio signal remains unmodified at the end of the transition time interval, thereby providing a smooth global transition 191 between the audio scenes 111, 112 in a computationally efficient manner.

[0071] Similar to the origin source position of the origin audio source 113, the destination source position of the destination audio source 113 on the sphere 114 around the listening positions 201, 202 may be maintained, in particular during the entire transition time interval, when the listener 181 moves from the origin audio scene 111 to the destination audio scene 112. Alternatively or additionally, it may be assumed that the listener 181 stays at the same listening position 201, 202 (during the entire transition time interval). This may further reduce the amount of computation for the global transition 191 between audio scenes 111, 112.

[0072] The fade-out function 211 and / or the fade-in function 212 may be derived from a bitstream representing the originating and / or destination audio signals. The bitstream 140 may be provided to the VR audio renderer 160 by the encoder 130. Thus, the global transition 191 may be controlled by a content provider. Alternatively or additionally, the fade-out function 211 and / or the fade-in function 212 may be derived from a storage unit of the virtual reality (VR) audio renderer 160 configured to render the originating and / or destination audio signals within the virtual reality rendering environment 180, thereby providing reliable operation during the global transition 191 between the audio scenes 111, 112.

[0073] The method 900 may include sending an indication (e.g., a flag indicating this) to the encoder 130 that the listener 181 is moving from the source audio scene 111 to the destination audio scene 112. The encoder 130 may then be configured to generate a bitstream 140 indicating the source audio signal and / or the destination audio signal. The indication enables the encoder 130 to selectively provide in the bitstream 140 the audio signals for the one or more audio sources 113 of the source audio scene 111 and / or for the one or more audio sources 113 of the destination audio scene 112. Thus, providing an indication of the upcoming global transition 191 allows a reduction in the required bandwidth for the bitstream 140.

[0074] As already indicated above, the audio scene of origin 111 may include multiple audio sources of origin 113. Accordingly, the method 900 may include rendering multiple audio signals of the corresponding audio sources of origin 113 from multiple different source positions on the sphere 114 around the listening positions 201, 202. Additionally, the method 900 may include applying a fade-out gain to the multiple audio signals of origin to determine multiple modified audio signals. Additionally, the method 900 may include rendering multiple modified audio signals of the audio source of origin 113 from multiple different source positions on the sphere 114 around the listening positions 201, 202.

[0075] Similarly, the method 900 may include determining a plurality of destination audio signals for a corresponding plurality of destination audio sources 113 of the destination audio scene 112. Further, the method 900 may include determining a plurality of destination source positions on a sphere 114 around the listening positions 201, 202. Further, the method 900 may include applying a fade-in gain to the plurality of destination audio signals to determine a corresponding plurality of modified destination audio signals. Further, the method 900 includes rendering the plurality of modified destination audio signals for the plurality of destination audio sources 113 from a corresponding plurality of destination source positions on the sphere 114 around the listening positions 201, 202.

[0076] Alternatively or additionally, the origin audio signal rendered during the global transition 191 may be a superposition of audio signals of multiple origin audio sources 113. Specifically, at the beginning of the transition time interval, the audio signals of (all) audio sources 113 of the origin audio scene 111 may be combined to provide a combined origin audio signal. This origin audio signal may be modified using a fade-out gain. Furthermore, the origin audio signal may be updated at a certain sampling rate (e.g., 20 ms) during the transition time interval. Similarly, the destination audio signal may correspond to a combination of audio signals of multiple destination audio sources 113 (in particular, all destination audio sources 113). The combined destination audio source may then be modified during the transition time interval using a fade-in gain. By combining the audio signals of the origin audio scene 111 and the destination audio scene 112, respectively, the amount of computation may be further reduced.

[0077] Additionally, a virtual reality audio renderer 160 for rendering audio in a virtual reality rendering environment 180 is described. As outlined herein, the VR audio renderer 160 may include a pre-processing unit 161 and a 3D audio renderer 162. The virtual reality audio renderer 160 may be configured to render an origin audio signal of an origin audio source 113 in an origin audio scene 111 from an origin source position on a sphere 114 around a listening position 201 of a listener 181. Furthermore, the VR audio renderer 160 is configured to determine when the listener 181 moves from the listening position 201 in the origin audio scene 111 to a listening position 202 in a different destination audio scene 112. Further, the VR audio renderer 160 is configured to apply a fade-out gain to the origin audio signal to determine a modified origin audio signal, and to render the modified origin audio signal of the origin audio source 113 from an origin source position on a sphere 114 around the listening positions 201, 202.

[0078] Further described is an encoder 130 configured to generate a bitstream 140 indicative of an audio signal to be rendered within a virtual reality rendering environment 180. The renderer 130 may be configured to determine an origin audio signal of an origin audio source 113 of the origin audio scene 111. Further, the encoder 130 may be configured to determine origin position data relating to an origin position of the origin audio source 113. The encoder 130 may then generate the bitstream 140 including the origin audio signal and the origin position data.

[0079] The encoder 130 may receive an indication (via a feedback channel from the VR audio renderer 160 to the encoder 130) that the listener 181 is moving from the source audio scene 111 to the destination audio scene 112 within the virtual reality rendering environment 180.

[0080] The encoder 130 may then determine (particularly only in response to receiving such an indication) an end audio signal of the end audio source 113 of the end audio scene 112 and end position data relating to the end source position of the end audio source 113. Further, the encoder 130 may generate a bitstream 140 including the end audio signal and the end position data. Thus, the encoder 130 may be configured to provide the end audio signal of one or more end audio sources 113 of the end audio source 112 only in response to receiving an indication for the global transition 191 to the end audio scene 112. By doing so, the required bandwidth for the bitstream 140 may be reduced.

[0081] 9b shows a flowchart of a corresponding method 930 for generating a bitstream 140 indicative of an audio signal to be rendered within a virtual reality rendering environment 180. The method 930 includes determining 931 an origin audio signal of an origin audio source 113 of the origin audio scene 111. Further, the method 930 includes determining 932 origin position data relating to an origin position of the origin audio source 113. Further, the method 930 includes generating 933 a bitstream 140 including the origin audio signal and the origin position data.

[0082] The method 930 includes receiving 934 an indication that the listener 181 is moving from the originating audio scene 111 to the destination audio scene 112 within the virtual reality rendering environment 180. In response, the method 930 may include determining 935 an destination audio signal for the destination audio source 113 of the destination audio scene 112 and determining 936 destination position data related to the destination source position of the destination audio source 113. Further, the method 930 includes generating 937 a bitstream 140 including the destination audio signal and the destination position data.

[0083] 9c shows a flowchart of an example method 910 for rendering an audio signal in a virtual reality rendering environment 180. The method 910 may be performed by the VR audio renderer 160.

[0084] The method 910 includes rendering 911 origin audio signals of audio sources 311, 312, 313 from origin source positions on a sphere of origin 114 around a listening position of origin 301 of a listener 181. The rendering 911 may be performed using a 3D audio renderer 162. In particular, the rendering 911 may be performed under the assumption that the listening position of origin 301 is fixed. Thus, the rendering 911 may be limited to three degrees of freedom (in particular, rotational movement of the head of the listener 181).

[0085] To account for the additional three degrees of freedom (for translational movement of the listener 181), the method 910 may include determining 912 that the listener 181 moves from the origin listening position 301 to the destination listening position 302, where the destination listening position 302 is typically within the same audio scene 111. Thus, the listener 181 may be determined 912 to perform a local transition 192 within the same audio scene 111.

[0086] In response to determining that the listener 181 performs the local transition 192, the method 910 may include determining 913, based on the origin source positions, end source positions of the audio sources 311, 312, 313 on the end sphere 114 around the end listening position 302. In other words, the source positions of the audio sources 311, 312, 313 may be translated from the origin sphere 114 around the origin listening position 301 to the end sphere 114 around the end position 302. This may be achieved by projecting the origin source positions from the origin sphere 114 onto the end sphere. In particular, the end source positions may be determined such that they correspond to the intersection of a ray between the end listening position 302 and the origin source positions with the end sphere 114.

[0087] Further, the method 910 may include determining 914 an end audio signal of the audio source 311, 312, 313 based on the origin audio signal (in response to determining that the listener 181 performs the local transition 192). In particular, the intensity of the end audio signal may be determined based on the intensity of the origin audio signal. Alternatively or additionally, the spectral composition of the end audio signal may be determined based on the spectral composition of the origin audio signal. Thus, how the audio signal of the audio source 311, 312, 313 is perceived from the end listening position 302 may be determined (in particular, the intensity and / or the spectral composition of the audio signal may be determined).

[0088] The above-mentioned determining steps 913, 914 may be performed by a pre-processing unit 161 of the VR audio renderer 160. The pre-processing unit 161 may handle the translational movement of the listener 181 by relocating the audio signals of one or more audio sources 311, 312, 313 from the origin sphere 114 around the origin listening position 301 to the destination sphere 114 around the destination listening position 302. As a result, the relocated audio signals of the one or more audio sources 311, 312, 313 may also be rendered using the 3D audio renderer 162 (which may be limited to 3 DoF). Thus, the method 910 allows for the efficient provision of 6 DoF within the VR audio rendering environment 180.

[0089] As a result, the method 910 may include rendering 915 (e.g., using a 3D audio renderer such as an MPEG-H audio renderer) the destination audio signals of the audio sources 311, 312, 313 from the destination source positions on the destination sphere 114 around the destination listening position 302.

[0090] Determining 914 the end audio signal may include determining an end distance 322 between the origin source position and the end listening position 302. The end audio signal (particularly the intensity of the end audio signal) may then be determined (particularly scaled) based on the end distance 322. In particular, determining 914 the end audio signal may include applying a distance gain 410 to the origin audio signal, where the distance gain 410 is dependent on the end distance 322.

[0091] A distance function 415 may be provided that indicates the distance gain 410 as a function of the distance 321, 322 between the source positions of the audio signals 311, 312, 313 and the listening positions 301, 302 of the listener 181. The distance gain 410 to be applied to the source audio signal (to determine the destination audio signal) may be determined based on the function value of the distance function 415 for the destination distance 322. In this way, the destination audio signal may be determined efficiently and accurately.

[0092] Furthermore, determining 914 the destination audio signal may include determining an origin distance 321 between the origin source position and the origin listening position 301. The destination audio signal may then be determined based on the origin distance 321. In particular, the distance gain 410 applied to the origin audio signal may be determined based on a function value of the distance function 415 for the origin distance 321. In one preferred example, the function value of the distance function 415 for the origin distance 321 and the function value of the distance function 415 for the destination distance 322 are used to rescale the intensity of the origin audio signal to determine the destination audio signal. Thus, an efficient and precise local transition 191 within the audio scene 111 may be provided.

[0093] Determining 914 the destination audio signal may include determining a directional profile 332 of the audio sources 311, 312, 313. The directional profile 332 may indicate the strength of the origin audio signal in various directions. The destination audio signal may then be determined (also) based on the directional profile 332. By taking the directional profile 332 into account, the acoustic quality of the local transition 192 may be improved.

[0094] The directivity profile 332 may indicate a directional gain 510 to be applied to the source audio signal to determine the destination audio signal. In particular, the directivity profile 332 may indicate a directional gain function 515, which may indicate the directional gain 510 as a function of a (possibly two-dimensional) directivity angle 520 between the source positions of the audio sources 311, 312, 313 and the listening positions 301, 302 of the listener 181.

[0095] Thus, determining 914 the endpoint audio signal may include determining an endpoint angle 522 between the endpoint source position and the endpoint listening position 302. The endpoint audio signal may then be determined based on the endpoint angle 522. In particular, the endpoint audio signal may be determined based on a function value of the directivity gain function 515 for the endpoint angle 522.

[0096] Alternatively or additionally, determining 914 the end audio signal may include determining an origin angle 521 between the origin source position and the origin listening position 301. The end audio signal may then be determined based on the origin angle 521. In particular, the end audio signal may be determined based on a function value of the directivity gain function 515 for the origin angle 521. In a preferred example, the end audio signal may be determined by modifying the intensity of the origin audio signal using the function values ​​of the directivity gain function 515 for the origin angle 521 and for the end angle 522 to determine the intensity of the end audio signal.

[0097] Additionally, the method 910 may include endpoint environment data 193 indicative of audio propagation characteristics of the medium between the endpoint source location and the endpoint listening position 302. The endpoint environment data 193 may be indicative of obstacles 603 located on the direct path between the endpoint source location and the endpoint listening position 302; may be indicative of information regarding the spatial dimensions of the obstacles 603; and / or may be indicative of the attenuation experienced by an audio signal on the direct path between the endpoint source location and the endpoint listening position 302. In particular, the endpoint environment data 193 may be indicative of an obstacle attenuation function for the obstacle 603, which may be indicative of the attenuation experienced by an audio signal passing through the obstacle 603 on the direct path between the endpoint source location and the endpoint listening position 302.

[0098] The destination audio signal may be determined based on the destination environment data 193, thereby further enhancing the quality of the audio rendered within the VR rendering environment 180.

[0099] As indicated above, the endpoint environment data 193 may indicate an obstacle 603 on a direct path between the endpoint source location and the endpoint listening position 302. The method 910 may include determining a passing distance 601 between the endpoint source location and the endpoint listening position 302 on the direct path. The endpoint audio signal may then be determined based on the passing distance 601. Alternatively or additionally, an obstacle-free distance 602 between the endpoint source location and the endpoint listening position 302 on an indirect path that does not pass through the obstacle 603 may be determined. The endpoint audio signal may then be determined based on the obstacle-free distance 602.

[0100] Specifically, the indirect component of the destination audio signal may be determined based on the source audio signal propagating along the indirect path. Furthermore, the direct component of the destination audio signal may be determined based on the source audio signal propagating along the direct path. The destination audio signal may then be determined by combining the indirect and direct components. In this way, the acoustic effect of the obstacle 603 may be taken into account in a precise and efficient manner.

[0101] Furthermore, the method 910 may include determining focus information related to the field of view 701 and / or focus of interest 702 of the listener 181. The end audio signal may then be determined based on the focus information. Specifically, the spectral composition of the audio signal may be adapted depending on the focus information. This may further improve the VR experience of the listener 181.

[0102] Further, the method 910 may include determining that the audio sources 311, 312, 313 are ambience audio sources. In this context, an indicator (e.g., a flag) may be received in the bitstream 140 from the encoder 130. For example, the indicator may indicate that the audio sources 311, 312, 313 are ambience audio sources. Ambient audio sources typically provide background audio signals. The origin source positions of the ambience audio sources may be maintained as destination source positions. Alternatively or additionally, the intensity of the origin audio signal of the ambience audio source may be maintained as the intensity of the destination audio signal. In this way, the ambience audio sources can be handled efficiently and consistently in the context of the local transition 192.

[0103] The above-described aspects are applicable to an audio scene 111 including a plurality of audio sources 311, 312, 313. In particular, the method 910 may include rendering a plurality of origin audio signals for the corresponding plurality of audio sources 311, 312, 313 from a plurality of different origin source positions on the origin sphere 114. Further, the method 910 may include determining a plurality of destination source positions for the corresponding plurality of audio sources 311, 312, 313 on the destination sphere 114 based on the plurality of origin source positions, respectively. Further, the method 910 may include determining a plurality of destination audio signals for the corresponding plurality of audio sources 311, 312, 313 based on the plurality of origin audio signals, respectively. The plurality of destination audio signals for the corresponding plurality of audio sources 311, 312, 313 may then be rendered from a corresponding plurality of destination source positions on the destination sphere 114 around the destination listening position 302.

[0104] Additionally, a virtual reality audio renderer 160 is described for rendering audio signals in a virtual reality rendering environment 180. The audio renderer 160 is configured to render (in particular using a 3D audio renderer 162 of the VR audio renderer 160) origin audio signals of audio sources 311, 312, 313 from origin source positions on a sphere of origin 114 around an origin listening position 301 of a listener 181.

[0105] Additionally, VR audio renderer 160 may be configured to determine that listener 181 moves from origin listening position 301 to destination listening position 302. In response, VR audio renderer 160 may be configured to determine (e.g., within preprocessing unit 161 of VR audio renderer 160) destination source positions of audio sources 311, 312, 313 on destination sphere 114 around destination listening position 302 based on the origin source positions, and to determine destination audio signals for audio sources 311, 312, 313 based on the origin audio signals.

[0106] Additionally, the VR audio renderer 160 (e.g., the 3D audio renderer 162 ) may be configured to render the destination audio signals of the audio sources 311 , 312 , 313 from destination source positions on the destination sphere 114 around the destination listening position 302 .

[0107] Thus, the virtual reality audio renderer 160 may comprise a pre-processing unit 161 configured to determine destination source positions and destination audio signals of the audio sources 311, 312, 313. Furthermore, the VR audio renderer 160 may comprise a 3D audio renderer 162 configured to render the destination audio signals of the audio sources 311, 312, 313. The 3D audio renderer 162 may be configured to adapt the rendering of the audio signals of the audio sources 311, 312, 313 on the (unit) sphere 114 around the listening positions 301, 302 of the listener 181 according to rotational movements of the head of the listener 181 (to provide 3DoF within the rendering environment 180). On the other hand, the 3D audio renderer 162 may not be configured to adapt the rendering of the audio signals of the audio sources 311, 312, 313 according to translational movements of the head of the listener 181. In this way, the 3D audio renderer 162 may be limited to 3 DoF, and translational DoF can then be provided in an efficient manner using the pre-processing unit 161, thereby providing an overall VR audio renderer 160 with 6 DoF.

[0108] Further described is an audio encoder 130 configured to generate a bitstream 140. The bitstream 140 is generated to indicate an audio signal of at least one audio source 311, 312, 313 and to indicate a position of said at least one audio source 311, 312, 313 within a rendering environment 180. Furthermore, the bitstream 140 may indicate environment data 193 related to audio propagation characteristics of the audio within the rendering environment 180. By signaling the environment data 193 related to the audio propagation characteristics, local transitions 192 within the rendering environment 180 may be enabled in a precise manner.

[0109] Further described is a bitstream 140 indicating environment data 193 relating to an audio signal of at least one audio source 311, 312, 313; the position of said at least one audio source 311, 312, 313 within the rendering environment 180; and audio propagation characteristics of the audio within the rendering environment 180. Alternatively or additionally, the bitstream 140 may indicate whether the audio source 311, 312, 313 is an ambient sound audio source 801.

[0110] 9d shows a flowchart of an exemplary method 920 for generating a bitstream. The method 920 includes determining 921 an audio signal of at least one audio source 311, 312, 313. Further, the method 920 includes determining 922 position data related to a position of the at least one audio source 311, 312, 313 within the rendering environment 180. Further, the method 920 may include determining 923 environment data 193 related to audio propagation characteristics of audio within the rendering environment 180. The method 920 may further include inserting 934 the audio signal, the position data, and the environment data 193 into the bitstream 140. Alternatively or additionally, an indication of whether the audio source 311, 312, 313 is an ambient audio source 801 may be inserted into the bitstream 140.

[0111] Thus, this document describes a virtual reality audio renderer 160 (and corresponding methods) for rendering audio signals from audio sources 311, 312, 313 in a virtual reality rendering environment 180. The audio renderer 160 comprises a 3D audio renderer 162 configured to render audio signals from audio sources 113, 311, 312, 313 from source positions on a sphere 114 around a listening position 301, 302 of a listener 181 in the virtual reality rendering environment 180. Furthermore, the virtual reality audio renderer 160 comprises a pre-processing unit 161 configured to determine a new listening position 301, 302 for the listener 181 in the virtual reality rendering environment 180 (within the same or a different audio scene 111, 112). Further, the pre-processing unit 161 is configured to update the audio signals and source positions of the audio sources 113, 311, 312, 313 with respect to a sphere 114 around the new listening positions 301, 302. The 3D audio renderer 162 is configured to render the updated audio signals of the audio sources 311, 312, 313 from the updated source positions on the sphere 114 around the new listening positions 301, 302.

[0112] The methods and systems described herein may be implemented as software, firmware, and / or hardware. Certain components may be implemented as software running on a digital signal processor or microprocessor. Other components may be implemented as hardware or as application-specific integrated circuits. The signals encountered in the described methods and systems may be stored on media such as random access memory or optical storage media. The signals may be transmitted over a network, such as an airwave network, a satellite network, a wireless network, or a wired network, e.g., the Internet. Typical devices utilizing the methods and systems described herein are portable electronic devices or other consumer equipment used to store and / or render audio signals.

[0113] The enumerated example (EE) for this article is as follows: [EE1] A method (900) for rendering audio in a virtual reality rendering environment (180), the method comprising: Rendering (901) an origin audio signal of an origin audio source (113) of an origin audio scene (111) from an origin source position on a sphere (114) around a listening position (201) of a listener (181); determining (902) that a listener (181) moves from said listening position (201) in a source audio scene (111) to a listening position (202) in a different destination audio scene (112); applying a fade-out gain to the source audio signal to determine a modified source audio signal (903); Rendering (904) the modified origin audio signal of the origin audio source (113) from the origin source position on a sphere (114) around the listening positions (201, 202), method. [EE2] The method: determining that a listener (181) moves from the starting audio scene (111) to the ending audio scene (112) during a transition time interval; determining an intermediate point (213) within the transition time interval; determining the fade-out gain based on the relative position of the intermediate point (213) within the transition time interval; Method described in EE1. [EE3] the method includes providing a fade-out function (211) indicative of the fade-out gain at various intermediate points (213) within the transition time interval; the fade-out function (211) is such that the fade-out gain decreases with increasing intermediate time point (213); Method described in EE2. [EE4] The fade-out function (211) is the source audio signal remains unmodified at the beginning of the transition time interval; and / or the source audio signal is increasingly attenuated at intermediate points in time (213) as it progresses; and / or the source audio signal is completely attenuated at the end of the transition time interval; The method described in EE3 is as follows. [EE5] The method comprises: maintaining the origin source position of the origin audio source (113) on a sphere (114) around the listening position (201, 202) as the listener (181) moves from the origin audio scene (111) to the destination audio scene (112); and / or maintaining the listening position (201, 202) unchanged as the listener (181) moves from the source audio scene (111) to the destination audio scene (112); A method described in any one of EE1 to EE4. [EE6] The method comprises: determining an end audio signal of an end audio source (113) of said end audio scene (112); determining end point source positions on a sphere (114) around said listening positions (201, 202); applying a fade-in gain to the destination audio signal to determine a modified destination audio signal; rendering the modified destination audio signal of the destination audio source (113) from the destination source position on a sphere (114) around the listening position (201, 202), A method described in any one of EE1 to EE5. [EE7] The method: determining that a listener (181) moves from the starting audio scene (111) to the ending audio scene (112) during a transition time interval; determining an intermediate point (213) within the transition time interval; determining the fade-in gain based on the relative position of the intermediate point (213) within the transition time interval; Method described in EE6. [EE8] the method includes providing a fade-in function (212) indicative of the fade-in gain at various intermediate points (213) within the transition time interval; the fade-in function (212) is such that the fade-in gain increases with the progression of intermediate times (213); Method described in EE7. [EE9] The fade-in function (211) is the end audio signal remains unmodified at the end of the transition time interval; and / or the end audio signal becomes less and less attenuated at intermediate points in time (213) as it progresses; and / or the end audio signal is completely attenuated at the beginning of the transition time interval; The method of claim EE8, [EE10] The method comprises: maintaining a destination source position of the destination audio source (113) on a sphere (114) around the listening position (201, 202) as the listener (181) moves from the source audio scene (111) to the destination audio scene (112); and / or maintaining the listening position (201, 202) unchanged as the listener (181) moves from the source audio scene (111) to the destination audio scene (112); The method of any one of EE6 to EE9. [EE11] The method of claim EE8, wherein the fade-out function (211) and the fade-in function (212) combine to provide a constant gain for a plurality of different intermediate points (213). [EE12] The fade-out function (211) and / or the fade-in function (212) derived from a bitstream (140) representing the source audio signal and / or the destination audio signal; and / or derived from a storage unit of a virtual reality audio renderer (160) configured to render the source audio signal and / or the destination audio signal within a virtual reality rendering environment (180); How to write EE8 when EE8 cites EE3. [EE13] 13. The method of any one of claims EE1 to EE12, comprising receiving an indication that a listener (181) is moving from the starting audio scene (111) to the ending audio scene (112). [EE14] The method of claim EE13, wherein the indicator comprises a flag. [EE15] The method of any one of claims EE1 to EE14, comprising sending an indication to an encoder (130) that a listener (181) is moving from the starting audio scene (111) to the ending audio scene (112); and the encoder (130) is configured to generate a bitstream (140) indicative of the starting audio signal. [EE16] 16. The method of any one of claims EE1 to EE15, wherein the first audio signal is rendered using a 3D audio renderer (162), in particular an MPEG-H audio renderer. [EE17] The method comprises: Rendering a plurality of origin audio signals of a corresponding plurality of origin audio sources (113) from a plurality of different origin source positions on a sphere (114) around said listening position (201, 202); applying the fade-out gains to the plurality of source audio signals to determine a plurality of modified source audio signals; rendering the modified origin audio signals of the origin audio source (113) from the corresponding origin source positions on a sphere (114) around the listening positions (201, 202), The method of any one of EE1 to 16. [EE18] The method comprises: determining a plurality of destination audio signals of a corresponding plurality of destination audio sources (113) of the destination audio scenes (112); determining a plurality of end point source locations on a sphere (114) around the listening positions (201, 202); applying the fade-in gains to the plurality of end audio signals to determine a corresponding plurality of modified end audio signals; rendering the modified end-point audio signals of the end-point audio sources (113) from the corresponding end-point source positions on a sphere (114) around the listening positions (201, 202), The method of any one of EE6 to 17. [EE19] 19. The method of any one of claims EE1 to EE18, wherein the source audio signal is a superposition of audio signals of multiple source audio sources (113). [EE20] A virtual reality audio renderer (160) for rendering audio in a virtual reality rendering environment (180), the virtual reality audio renderer (160) comprising: Rendering an origin audio signal of an origin audio source (113) of an origin audio scene (111) from an origin source position on a sphere (114) around a listening position (201) of a listener (181); determining that a listener (181) moves from said listening position (201) in a source audio scene (111) to a listening position (202) in a different destination audio scene (112); applying a fade-out gain to the source audio signal to determine a modified source audio signal; configured to render the modified origin audio signal of the origin audio source (113) from an origin source position on a sphere (114) around the listening positions (201, 202), Virtual reality audio renderer. [EE21] An encoder (130) configured to generate a bitstream (140) representing an audio signal to be rendered within a virtual reality rendering environment (180), the encoder (130) comprising: determining an origin audio signal of an origin audio source (113) of the origin audio scene (111); determining origin location data relating to an origin location of the origin audio source (113); generating a bitstream (140) including the origin audio signal and the origin position data; receiving an indication that a listener (181) is moving from the source audio scene (111) to the destination audio scene (112) within the virtual reality rendering environment (180); determining an end audio signal of an end audio source (113) of said end audio scene (112); determining destination location data relating to the destination source location of said destination audio source (113); configured to generate a bitstream (140) including the end audio signal and the end position data; Encoder. [EE22] A method (930) for generating a bitstream (140) representing an audio signal to be rendered within a virtual reality rendering environment (180), the method comprising: determining (931) an origin audio signal of an origin audio source (113) of an origin audio scene (111); determining (932) origin location data relating to an origin source location of the origin audio source (113); generating (933) a bitstream (140) including the origin audio signal and the origin position data; receiving (934) an indication that a listener (181) is moving from the source audio scene (111) to the destination audio scene (112) within the virtual reality rendering environment (180); determining (935) an end audio signal of an end audio source (113) of the end audio scene (112); determining (936) destination location data relating to the destination source location of the destination audio source (113); generating (937) a bitstream (140) including the end audio signal and the end position data; method. [EE23] A virtual reality audio renderer (160) for rendering an audio signal in a virtual reality rendering environment (180), the audio renderer comprising: a 3D audio renderer (162) configured to render an audio signal of an audio source (113) from a source position on a sphere (114) around a listening position (201, 202) of a listener (181) within a virtual reality rendering environment (180); A pre-treatment unit (161), determining a new listening position (201, 202) for a listener (181) within a virtual reality rendering environment (180); a pre-processing unit (161) configured to update the audio signal and the source positions of the audio sources (201, 202) with respect to a sphere (114) around the new listening position (201, 202), the 3D audio renderer (162) is configured to render the updated audio signal of the audio source (113) from an updated source position on a sphere (114) around the new listening position (201, 202); Virtual reality audio renderer.

Claims

1. 1. A method of rendering audio in a virtual reality rendering environment using an audio renderer for rendering three degrees of freedom (3DoF), the method comprising: rendering, by the audio renderer, an origin audio signal of an origin audio source of an origin audio scene from an origin source position on a sphere around an origin listening position of a listener within a virtual reality rendering environment; determining the presence of listener movement, said movement being movement within a virtual reality rendered environment from said origin listening position within an origin audio scene to a destination listening position within a destination audio scene; determining a modified origin audio signal by applying a fade-out gain to the origin audio signal based on the determination of the movement; determining an end audio signal of an end audio source of the end audio scene; determining an end source location on a sphere around the end listening location; applying a fade-in gain to the end audio signal to determine a modified end audio signal; rendering, by the audio renderer, the modified origin audio signal of the origin audio source from the origin source position on a sphere around the origin listening position; A method comprising:

2. 2. The method of claim 1, wherein the modified source audio signal is rendered to the listener from the same position throughout the movement from the source listening position in the source audio scene to the destination listening position in the destination audio scene.

3. The method of claim 1 , wherein the destination audio scene does not include the origin audio source.

4. The method comprises: determining that a listener moves from the source audio scene to the destination audio scene during a transition time interval; determining an intermediate point within the transition time interval; determining the fade-out gain based on the relative position of the intermediate point within the transition time interval. The method of claim 1.

5. A non-transitory computer-readable storage medium storing executable instructions for causing a computer to perform the method of claim 1.

6. 1. A system for rendering audio in a virtual reality rendering environment using an audio renderer for rendering three degrees of freedom (3DoF), the system comprising: a first renderer that renders, by the audio renderer, an origin audio signal of an origin audio source of an origin audio scene from an origin source position on a sphere around an origin listening position of a listener within the virtual reality rendering environment; a first processor for determining the presence of listener movement, the movement being movement within a virtual reality rendering environment from the origin listening position within an origin audio scene to a destination listening position within a destination audio scene; a second processor that determines a modified origin audio signal by applying a fade-out gain to the origin audio signal based on the determination of the movement; a third processor for determining an end audio signal of an end audio source of the end audio scene; a fourth processor for determining an end source location on a sphere around the end listening location; a fifth processor that applies a fade-in gain to the destination audio signal to determine a modified destination audio signal; a second renderer that renders the modified origin audio signal of the origin audio source from the origin source position on a sphere around the origin listening position by the audio renderer; A system having: