Encoder for handling global transitions between listener positions in a virtual reality environment.
The method addresses the challenge of handling translational motion in virtual reality audio rendering by applying gradual output gains and interpolation techniques, ensuring a seamless audio experience during position changes.
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- DOLBY INTERNATIONAL AB
- Filing Date
- 2018-12-18
- Publication Date
- 2026-07-07
AI Technical Summary
Existing audio rendering systems are limited to handling rotational movements and fail to efficiently manage translational motion in virtual reality environments, leading to inconsistent audio experiences during transitions between different listener positions.
A method and system for rendering audio in virtual reality environments that involves determining listener movements, applying gradual output gains to source audio signals, and rendering modified signals from pre-defined positions on a sphere, allowing for efficient handling of translational motion through techniques like linear interpolation and gradual gain functions.
This approach provides a computationally efficient and acoustically consistent experience during transitions between audio scenes, reducing computational complexity and processing delays while maintaining audio coherence.
Smart Images

Figure 00000000_0000_ABST
Description
1 / 50 “ENCODER FOR HANDLING GLOBAL TRANSITIONS BETWEEN LISTENER POSITIONS IN A VIRTUAL REALITY ENVIRONMENT” Separated from BR112020012299-8, filed on 12 / 18 / 2018 CROSS-REFERENCE TO RELATED REQUESTS
[001] This application claims priority for the following priority applications: United States Provisional Application No. 62 / 599,841 (reference: D17085USP1), filed December 18, 2017, and EP application 17208088.9 (reference: D17085EP), filed December 18, 2017, the descriptions of which are incorporated herein by reference. FIELD OF TECHNIQUE
[002] This document refers to efficient and consistent handling of transitions between auditory viewing ports and / or listener positions in a virtual reality (VR) rendering environment. BACKGROUND OF THE INVENTION
[003] Virtual reality (VR), augmented reality (AR), and mixed reality (MR) applications are rapidly evolving to include increasingly refined acoustic models of sound sources and scenes that can be appreciated from different viewpoints / perspectives or listener positions. Two different classes of flexible audio representations can, for example, be employed for VR applications: sound field representations and object-based representations. Sound field representations are physically based approaches that encode the incident wavefront at the listener's position. For example, approaches such as higher-order ambisonics (HOA) or B-format represent the spatial wavefront using a spherical harmonic decomposition. Object-based approaches represent a complex auditory scene as a collection of singular elements comprising an audio waveform or audio signal and associated parameters or metadata. Petition 870260038484, dated 04 / 24 / 2026, page 11 / 132 2 / 50 possibly with temporal variation.
[004] Enjoying VR, AR, and MR applications can include experiencing different viewpoints or auditory perspectives by the user. For example, room-based virtual reality can be provided based on a mechanism using 6 degrees of freedom (DoF). FIG. 1 illustrates an example of 6 DoF interaction showing translational movement (forward / backward, up / down, and left / right) and rotational movement (tilt, yaw, and roll). Unlike a 3 DoF spherical video experience, limited to head rotations, content created for 6 DoF interaction also allows navigation within a virtual environment (e.g., physically walking within a room), in addition to head rotations. This can be achieved based on positional trackers (e.g., camera-based) and orientational trackers (e.g., gyroscopes and / or accelerometers).6DoF tracking technology may be available in high-end desktop VR systems (e.g., PlayStation®VR, Oculus Rift, HTC Vive) as well as high-end mobile VR platforms (e.g., Google Tango). The user experience of directionality and spatial extension of sound or audio sources is fundamental to the realism of 6DoF experiences, particularly an experience of navigating through a scene and around virtual audio sources.
[005] Available audio rendering systems (such as the MPEG-H 3D audio renderer) are generally limited to rendering 3 DoFs (i.e., rotational movement of an audio scene caused by a listener's head movement). Transactional changes to a listener's position and associated DoFs typically cannot be handled by these renderers.
[006] This document addresses the technical problem of providing resource-efficient methods and systems for handling translational motion in the context of audio rendering. Petition 870260038484, dated 04 / 24 / 2026, page 12 / 132 3 / 50 SUMMARY
[007] According to one aspect, a method for rendering audio in a virtual reality rendering environment is described. The method comprises rendering an audio signal from a source audio source of a source audio scene from a source source position on a sphere surrounding a listener source position. Furthermore, the method comprises determining that the listener moves from the listener position within the source audio scene to a listener position within a different target audio scene. Additionally, the method comprises applying a gradual output gain to the source audio signal to determine a modified source audio signal. The method further comprises rendering the modified source audio signal from the source audio source from the source source position on the sphere surrounding the listener position.
[008] According to a further aspect, a virtual reality audio renderer is described for rendering audio in a virtual reality rendering environment. The virtual reality audio renderer is configured to render an audio signal from a source audio source of a source audio scene from a source source position on a sphere around a listener position. Furthermore, the virtual reality audio renderer is configured to determine that the listener moves from the listener position within the source audio scene to a listener position within a different target audio scene.Furthermore, the virtual reality audio renderer is configured to apply a gradual output gain to the source audio signal to determine a modified source audio signal, and to render the modified source audio signal from the source audio source's position on the sphere surrounding the listener's position.
[009] According to an additional aspect, a method for generating is described. Petition 870260038484, dated 04 / 24 / 2026, page 13 / 132 4 / 50 a bitstream indicative of an audio signal to be rendered within a virtual reality rendering environment. The method comprises: determining a source audio signal from a source audio source of a source audio scene; determining source position data relative to a source position of the source audio source; generating a bitstream comprising the source audio signal and the source position data; receiving an indication that a listener moves from the source audio scene to a target audio scene within the virtual reality rendering environment; determining a target audio signal from a target audio source of the target audio scene; determining target position data relative to a target source position of the target audio source; and generating a bitstream comprising the target audio signal and the target position data.
[010] According to another aspect, an encoder is described that is configured to generate a bitstream indicative of an audio signal to be rendered within a virtual reality rendering environment.The encoder is configured to: determine a source audio signal from a source audio source of a source audio scene; determine the source position data relative to a source source position of the source audio source; generate a bitstream comprising the source audio signal and the source position data; receive an indication that a listener is moving from the source audio scene to a target audio scene within the virtual reality rendering environment; determine a target audio signal from a target audio source of the target audio scene; determine the target position data relative to a target source position of the target audio source; and generate a bitstream comprising the target audio signal and the target position data.
[011] According to a further aspect, a virtual reality audio renderer is described for rendering an audio signal in a virtual reality environment. Petition 870260038484, dated 04 / 24 / 2026, page 14 / 132 5 / 50 Virtual Reality Rendering. The audio renderer comprises a 3D audio renderer that is configured to render an audio signal from an audio source from a source position on a sphere surrounding a listener position within the virtual reality rendering environment. Additionally, the virtual reality audio renderer comprises a pre-processing unit that is configured to determine a new listener position within the virtual reality rendering environment. Furthermore, the pre-processing unit is configured to update the audio signal and the source position of the audio source relative to a sphere surrounding the new listener position. The 3D audio renderer is configured to render the updated audio signal from the audio source from the updated source position on the sphere surrounding the new listener position.
[012] According to another aspect, a software program is described. The software program can be adapted for execution on a processor and to execute the method steps described in this document when executed on the processor.
[013] According to another aspect, a storage medium is described. The storage medium may comprise a software program adapted for execution on a processor and to execute the method steps described in this document when executed on the processor.
[014] According to another aspect, a computer program product is described. The computer program may comprise executable instructions to perform the method steps described in this document when executed on a computer.
[015] It should be noted that the methods and systems, including their preferred embodiments, as described in this patent application, may be used independently or in combination with other methods and systems. Petition 870260038484, dated 04 / 24 / 2026, page 15 / 132 6 / 50 disclosed in this document. Furthermore, all aspects of the methods and systems described in this patent application may be arbitrarily combined. In particular, the features of the claims may be combined with each other in an arbitrary manner. BRIEF DESCRIPTION OF THE FIGURES
[016] The invention is explained below in an exemplary manner with reference to the accompanying drawings, in which
[017] Fig. 1a shows an example of an audio processing system to provide 6 DoF audio;
[018] Fig. 1 b shows examples of situations within a 6 DoF audio and / or rendering environment;
[019] Fig. 1 c shows an example of transitioning from a source audio scene to a destination audio scene;
[020] Fig. 2 illustrates an example of a scheme for determining spatial audio signals during a transition between different audio scenes;
[021] Fig. 3 shows an example of an audio scene;
[022] Fig. 4a illustrates the remapping of audio sources in reaction to a change in the listener's position within an audio scene;
[023] Fig. 4b shows an example of a distance function;
[024] Fig. 5a illustrates an audio source with a non-uniform directivity profile;
[025] Fig. 5b shows an example of the directivity function of an audio source;
[026] Fig. 6 shows an example of an audio scene with an acoustically relevant obstacle;
[027] Fig. 7 illustrates a listener's field of vision and focus of attention;
[028] Fig. 8 illustrates the handling of ambient audio in case of a change in Petition 870260038484, dated 04 / 24 / 2026, page 16 / 132 7 / 50 listener position within an audio scene;
[029] Fig. 9a shows a flowchart of an example method for rendering a 3D audio signal during a transition between different audio scenes;
[030] Fig. 9b shows a flowchart of an example method for generating a bitstream for the transition between different audio scenes;
[031] Fig. 9c shows a flowchart of an example method for rendering a 3D audio signal during a transition within an audio scene; and
[032] Fig. 9d shows a flowchart of an example method for generating a bitstream for the transition. DETAILED DESCRIPTION
[033] As described above, this document refers to the efficient provision of 6DoF in a 3D (three-dimensional) audio environment. Fig. 1a illustrates a block diagram of an example audio processing system 100. An acoustic environment 110, such as a stadium, may comprise several different audio sources 113. Examples of audio sources 113 within a stadium are individual spectators, a stadium loudspeaker, the players on the field, etc. The acoustic environment 110 can be subdivided into different audio scenes 111, 112. By way of example, a first audio scene 111 may correspond to the home team's support block and a second audio scene 112 may correspond to the visiting team's support block. Depending on where a listener is positioned within the audio environment, the listener will perceive audio sources 113 from the first audio scene 111 or audio sources 113 from the second audio scene 112.
[034] The different audio sources 113 of an audio environment 110 can be captured using audio sensors 120, mainly using arrays of Petition 870260038484, dated 04 / 24 / 2026, page 17 / 132 8 / 50 microphone. In particular, one or more audio scenes 111, 112 of an audio environment 110 can be described using multichannel audio signals, one or more audio objects and / or higher-order ambisonic signals (HOA). Hereinafter, an audio source 113 is assumed to be associated with audio data that is captured by audio sensors 120, wherein the audio data indicates an audio signal and the position of the audio source 113 as a function of time (at a specific sampling rate of, for example, 20 ms).
[035] A 3D audio renderer, such as the MPEG-H 3D audio renderer, typically assumes that a listener is positioned at a specific listener position within an audio scene 111, 112. The audio data for the different audio sources 113 of an audio scene 111, 112 are typically provided under the assumption that the listener is positioned at that specific listener position. An audio encoder 130 may comprise a 3D audio encoder 131 that is configured to encode the audio data from the audio sources 113 of one or more audio scenes 111, 112.
[036] In addition, VR (virtual reality) metadata may be provided, enabling the listener to change listener position within an audio scene 111, 112 and / or move between different audio scenes 111, 112. The encoder 130 may comprise a metadata encoder 132 that is configured to encode VR metadata. The encoded VR metadata and the encoded audio data from the audio sources 113 may be combined in the combination unit 133 to provide a bitstream 140 that is indicative of the audio data and VR metadata. VR metadata may, for example, comprise environment data describing the acoustic properties of an audio environment 110.
[037] The 140 bitstream can be decoded using a 150 decoder to provide the (decoded) audio data and VR metadata. Petition 870260038484, dated 04 / 24 / 2026, page 18 / 132 9 / 50 (decoded). An audio renderer 160 for rendering audio within a rendering environment 180 that allows 6DoFs may comprise a preprocessing unit 161 and a 3D (conventional) audio renderer 162 (such as MPEG-H 3D audio). The preprocessing unit 161 may be configured to determine the listener position 182 of a listener 181 within the listener environment 180. The listener position 182 may indicate the audio scene 111 within which the listener 181 is positioned. Furthermore, the listener position 182 may indicate the exact position within an audio scene 111. The preprocessing unit 161 may also be configured to determine a 3D audio signal for the current listener position 182 based on audio data (decoded) and possibly VR metadata (decoded). The 3D audio signal can then be rendered using the 162 3D audio renderer.
[038] It should be noted that the concepts and schemes described in this document may be specified in a frequency-varying manner, may be defined globally or in an object / media-dependent manner, may be applied directly in the spectral or temporal domain and / or may be encoded in the VR 160 renderer or may be specified via a corresponding input interface.
[039] Fig. 1b shows an example of a 180 rendering environment. The listener 181 can be positioned within a source audio scene 111. For rendering purposes, it can be assumed that the audio sources 113, 194 are placed in different rendering positions on a sphere (unit) 114 around the listener 181. The rendering positions of the different audio sources 113, 194 can change over time (according to a given sampling rate). Different situations can occur within a VR 180 rendering environment: The listener 181 can perform a global transition 191 from the source audio scene 111 to a target audio scene 112. Petition 870260038484, dated 04 / 24 / 2026, page 19 / 132 10 / 50 Alternatively or furthermore, listener 181 may perform a local transition 192 to a different listener position 182 within the same audio scene 111. Alternatively or furthermore, an audio scene 111 may exhibit acoustically relevant environmental properties (such as a wall) that can be described using environment data 193 and that must be taken into account when a change in listener position 182 occurs. Alternatively or furthermore, an audio scene 111 may comprise one or more ambient audio sources 194 (e.g., for background noise) that must be taken into account when a change in listener position 182 occurs.
[040] Fig. 1c shows an example of global transition 191 from a source audio scene 111 with audio sources 113 Ai to An to a destination audio scene 112 with audio sources 113 Bi to Bm. Notably, each audio source 113 can be included only in one source audio scene 111 and in the destination audio scene 112, for example, audio sources 113 Ai to An are included in the source audio scene 111, but not in the destination audio scene 112, while audio sources 113 Bi to Bm are included in the destination audio scene 112, but not in the source audio scene 111.
[041] An audio source 113 can be characterized by the corresponding object properties between locations (coordinates, directivity, distance sound attenuation function, etc.). The global transition 191 can be performed within a certain transition time interval (e.g., within the interval of 5 seconds, 1 second or less). The listener position 182 in the source scene 111, at the beginning of the global transition 191, is marked with A. Furthermore, the listener position 182 within the destination scene 112, at the end of the global transition 191, is marked with B. In addition, Fig. 1c illustrates a local transition 192 within the destination scene 112 between listener position B and listener position C.
[042] Fig. 2 shows the global transition 191 from the source scene 111 Petition 870260038484, dated 04 / 24 / 2026, page 20 / 132 11 / 50 (or source display port) to destination scene 112 (or destination display port) during the transition time interval t. This transition 191 can occur when a listener 181 switches between different scenes or display ports 111, 112, for example, within a stadium. As such, the global transition 191 from source scene 111 to destination scene 112 need not correspond to the listener's actual physical movement 181, but may simply be initiated by the listener's command to switch or transition to another display port 111, 112. Notwithstanding, the present disclosure refers to a listener position, which is understood as a listener position in the VR / AR / MR environment.
[043] At an intermediate time instant 213, listener 181 can be positioned at an intermediate position between source scene 111 and destination scene 112. The 3D audio signal 203, which must be rendered at the intermediate position and / or at the intermediate time instant 213, can be determined by determining the contribution of each of the audio sources 113 Ai to An from source scene 111 and each of the audio sources 113 Bi to Bm from destination scene 112, taking into account the sound propagation of each audio source 113. This, however, would be linked to a relatively high computational complexity (especially in the case of a relatively high number of audio sources 113).
[044] At the start of global transition 191, listener 181 can be positioned at source listener position 201. Throughout transition 191, a 3D Ag source audio signal can be generated relative to source listener position 201, where the source audio signal depends only on audio sources 113 of source scene 111 (and does not depend on audio sources 113 of destination scene 112). Global transition 191 does not affect the source positions a) of the audio sources 113 of source scene 111. Therefore, assuming stationary audio sources 113 of source scene 111, the rendering positions of the sources Petition 870260038484, dated 04 / 24 / 2026, page 21 / 132 12 / 50 audio 113 during the global transition 191 in relation to the listener's position 201 do not change, even if the listener's position may transition from the source scene to the destination scene (in relation to the listener).
[045] Furthermore, it can be fixed at the beginning of global transition 191 that listener 181 will arrive at target listener position 202 within target scene 112 at the end of global transition 191. Throughout transition 191, a target audio signal 3D Ag can be generated relative to target listener position 202, wherein the target audio signal depends only on the audio sources 113 of target scene 112 (and does not depend on the audio sources 113 of source scene 111). Global transition 191 does not affect the apparent source positions of the audio sources 113 of target scene 112 (relative to the listener).
[046] To determine the intermediate 3D audio signal 203 at an intermediate position and / or at an intermediate time instant 213 during the global transition 191, the source audio signal at the intermediate time instant 213 can be combined with the destination audio signal at the intermediate time instant 213. In particular, a gradual output factor or gain derived from a gradual output function 211 can be applied to the source audio signal. The gradual output function 211 can be such that the gradual output factor or gain a decreases within an increasing distance from the intermediate position of the source scene 111. Furthermore, a gradual input factor or gain derived from a gradual input function 212 can be applied to the destination audio signal. The gradual input function 212 can be such that the gradual input factor or gain b increases with decreasing distance from the intermediate position of the destination scene 112.An example of a graduated output function 211 and an example of a graduated input function 212 are shown in Fig. 2. The intermediate audio signal can then be given by the weighted sum of the source audio signal and the destination audio signal, where the weights correspond to the graduated output gain and graduated input gain. Petition 870260038484, dated 04 / 24 / 2026, page 22 / 132 13 / 50 respectively.
[047] Therefore, a gradual input function or curve 212 and a gradual output function or curve 211 can be defined for a global transition 191 between different 3DoF viewports 201, 202. The functions 211, 212 can be applied to pre-rendered virtual objects or 3D audio signals representing the source audio scene 111 and the destination audio scene 112. By doing so, a consistent audio experience can be provided during a global transition 191 between different audio scenes 111, 112, with reduced VR audio rendering calculations.
[048] The intermediate audio signal 203 at an intermediate position xi can be determined using linear interpolation of the source audio signal and the destination audio signal. The intensity F of the audio signals can be given by: F(xi)=a*F(AG)+(1-a)*F(BG). The factor aeb = 1-a can be given by a norm function a = a (), which depends on the source listener position 201, the destination listener position 202, and the intermediate position. Alternatively to a function, a lookup table a = [1,..., 0] can be provided for different intermediate positions.
[049] Above, it is understood that the intermediate audio signal 203 can be determined and rendered for a plurality of intermediate positions xi to allow a smooth transition from the source scene 111 to the destination scene 112.
[050] During a global transition, 191 additional effects (e.g., Doppler effect and / or reverb) may be taken into account. Functions 211, 212 may be adapted by a content provider, for example, to reflect an artistic intention. Information about functions 211, 212 may be included as metadata in bitstream 140. Therefore, an encoder 130 may be configured to provide information about a gradual input function 212 and / or a gradual output function 211 as metadata within a bitstream 140. Petition 870260038484, dated 04 / 24 / 2026, page 23 / 132 14 / 50 Alternatively or in addition, an audio renderer 160 can apply a function 211,212 stored in audio renderer 160.
[051] A flag can be signaled from a listener to the renderer 160, notably to the VR preprocessing unit 161, to indicate to the renderer 160 that a global transition 191 should be performed from a source scene 111 to a destination scene 112. The flag can trigger the audio processing described in this document to generate an intermediate audio signal during the transition phase. The flag can be signaled explicitly or implicitly through related information (e.g., through coordinates of the new display window or listener position 202). The flag can be sent from either side of the data interface (e.g., server / content, user / scene, auxiliary). Along with the flag, information about the source audio signal Ag and the destination audio signal Bg can be provided. By way of example, an ID of one or more audio objects or audio sources can be provided.Alternatively, a request to calculate the source audio signal and / or the destination audio signal can be provided to renderer 160.
[052] Therefore, a VR renderer 160 comprising a preprocessing unit 161 for a 3DoF renderer 162 is described to enable 6DoF functionality in a resource-efficient manner. The preprocessing unit 161 allows the use of a standard 3DoF renderer 162, such as the MPEG-H 3D audio renderer. The VR preprocessing unit 161 can be configured to efficiently perform calculations for a global transition 191 using pre-rendered virtual audio objects Ag and Bg representing the source scene 111 and the destination scene 112, respectively. Computational complexity is reduced by using only two pre-rendered virtual objects during a global transition 191. Each virtual object can comprise a plurality of signals. Petition 870260038484, dated 04 / 24 / 2026, page 24 / 132 15 / 50 audio for a plurality of audio sources. Furthermore, bitrate requirements can be reduced, as during the 191 transition only the pre-rendered virtual audio objects Ag and Bg can be provided in the 140 bitstream. Additionally, processing delays can be reduced.
[053] 3DoF functionality can be provided for all intermediate positions along the global transition path. This can be achieved by overlapping the source audio object and the destination audio object using the gradual exit / gradual entry functions 211, 212. In addition, additional audio objects can be rendered and / or extra audio effects can be included.
[054] Fig. 3 shows an example of local transition 192 from a source listener position B 301 to a destination listener position C 302 within the same audio scene 111. The audio scene 111 comprises different audio sources or objects 311, 312, 313. The different audio sources or objects 311, 312, 313 may have different directivity profiles 332. In addition, the audio scene 111 may have environmental properties, notably one or more obstacles, that influence the propagation of audio within the audio scene 111. The environmental properties can be described using environment data 193. Furthermore, the relative distances 321, 322 of an audio object 311 to listener positions 301, 302 can be known.
[055] Figures 4a and 4b illustrate a scheme for handling the effects of a local transition 192 on the intensity of the different audio sources or objects 311, 312, 313. As described above, the audio source 311, 312, 313 of an audio scene 111 is typically assumed by a 3D audio renderer 162 to be positioned on a sphere 114 around the listener position 301. As such, at the beginning of a local transition 192, the audio sources 311, 312, 313 can be placed on a source sphere 114 around the source listener position 301 and at the end of the local transition 192, the audio sources 311, 312, 313 can be placed on a Petition 870260038484, dated 04 / 24 / 2026, page 25 / 132 16 / 50 target sphere 114 around target listener position 302. An audio source 311, 312, 313 can be remapped from source sphere 114 to target sphere 114. For this purpose, a ray going from target listener position 302 to the source position of the audio source 311, 312, 313 on source sphere 114 can be considered. The audio source 311, 312, 313 can be placed at the intersection of the ray with target sphere 114.
[056] The intensity F of an audio source 311, 312, 313 in the destination sphere 114 typically differs from the intensity in the source sphere 114. The intensity F can be modified using an intensity gain function or distance function 415, which provides a distance gain 410 as a function of the distance 420 of an audio source 311, 312, 313 from the listener position 301, 302. The distance function 415 typically exhibits a cutoff distance 421 above which a distance gain 410 of zero is applied. The source distance 321 from an audio source 311 to the source listener position 301 provides a source gain 411. Furthermore, the destination distance 322 from the audio source 311 to the destination listener position 302 provides a destination gain 412. The intensity F of the audio source 311 can be scaled using the source gain 411 and the destination gain 412, thus providing the intensity F of the audio source 311 at the destination sphere 114.In particular, the intensity F of the source audio signal from audio source 311 in source sphere 114 can be divided by the source gain 411 and multiplied by the destination gain 412 to give the intensity F of the destination audio signal from audio source 311 in destination sphere 114.
[057] Therefore, the position of an audio source 311 subsequent to a local transition 192 can be determined as: Ci = source_remap_function(Bi, C) (e.g., using a geometric transformation). Furthermore, the intensity of an audio source 311 subsequent to a local transition 192 can be determined as: F(Ci) = F(Bi) * distance_function(Bi, Ci, C). Distance attenuation Petition 870260038484, dated 04 / 24 / 2026, page 26 / 132 17 / 50 can therefore be modeled by the corresponding intensity gains provided by the distance function 415.
[058] Figures 5a and 5b illustrate an audio source 312 having a non-uniform directivity profile 332. The directivity profile can be defined using directivity gains 510 which indicate a gain value for different directions or angles of directivity 520. In particular, the directivity profile 332 of an audio source 312 can be defined using a directivity gain function 515 which indicates the directivity gain 510 as a function of the angle of directivity 520 (where the angle 520 can vary from 0° to 360°). It should be noted that, for 3D audio sources 312, the angle of directivity 520 is typically a two-dimensional angle comprising an azimuthal angle and an elevation angle. Therefore, the directivity gain function 515 is typically a two-dimensional function of the two-dimensional angle of directivity 520.
[059] The directivity profile 332 of an audio source 312 can be taken into account in the context of a local transition 192, determining the source directivity angle 521 of the source ray between the audio source 312 and the source listener position 301 (with the audio source 312 being placed on the source sphere 114 around the source listener position 301) and the destination directivity angle 522 of the destination ray between the audio source 312 and the destination listener position 302 (with the audio source 312 being placed on the destination sphere 114 around the destination listener position 302). Using the directivity gain function 515 of the audio source 312, the source directivity gain 511 and the destination directivity gain 512 can be determined as the function values of the directivity gain function 515 for the source directivity angle 521 and the destination directivity angle 522, respectively (see Fig. 5b).The intensity F of the audio source 312 at the listener position of origin 301 can then be divided by the directivity gain of origin 511 and multiplied by the gain of. Petition 870260038484, dated 04 / 24 / 2026, page 27 / 132 18 / 50 directivity of target 512 to determine the intensity F of the audio source 312 at the target listener position 302.
[060] Therefore, the directivity of the sound source can be parameterized by a directivity factor or gain 510 indicated by a directivity gain function 515. The directivity gain function 515 can indicate the intensity of the audio source 312 at some distance as a function of the angle 520 relative to the listener's position 301, 302. The directivity gains 510 can be defined as ratios relative to the gains of an audio source 312 at the same distance, having the same total power that is radiated uniformly in all directions. The directivity profile 332 can be parameterized by a set of gains 510 that correspond to vectors originating at the center of the audio source 312 and ending at points distributed on a unit sphere around the center of the audio source 312.The 332 directivity profile of a 312 audio source may depend on a use case scenario and available data (e.g., a uniform distribution for a 3D flight case, a planar distribution for 2D+ use cases, etc.).
[061] The resulting audio intensity from an audio source 312 at a target listener position 302 can be estimated as: F(Ci) = F(Bi)*distance_function()*directivity_gain_function(Ci,C,directivity_parameterization), where directivity_gain_function is dependent on the directivity profile 332 of the audio source 312. The distance_function() takes into account the modified intensity caused by the change in distance 321, 322 from the audio source 312 due to the audio source 312 transition.
[062] Fig. 6 shows an example of obstacle 603 that may need to be taken into account in the context of a local transition 192 between different listener positions 301, 302. In particular, the audio source 313 may be hidden behind obstacle 603 at the target listener position 302. The obstacle Petition 870260038484, dated 04 / 24 / 2026, page 28 / 132 19 / 50 603 can be described by environmental data 193 comprising a set of parameters, such as spatial dimensions of the obstacle 603 and an obstacle attenuation function, which indicates the attenuation of sound caused by the obstacle 603.
[063] An audio source 313 can display an obstacle-free distance 602 (OFD) to the destination listener position 302. The OFD 602 can indicate the shortest path length between the audio source 313 and the destination listener position 302 that does not cross the obstacle 603. Additionally, the audio source 313 can display a through distance 601 (GHD) to the destination listener position 302. The GHD 601 can indicate the shortest path length between the audio source 313 and the destination listener position 302 that normally passes through the obstacle 603. The obstacle attenuation function can be a function of the OFD 602 and the GHD 601. Furthermore, the obstacle attenuation function can be a function of the intensity F(Bi) of the audio source 313.
[064] The intensity of the audio source Ci at the target listener position 302 may be a combination of the sound coming from the audio source 313, which passes around the obstacle 603 and the sound coming from the audio source 313 that passes through the obstacle 603.
[065] Therefore, the VR renderer 160 can be provided with parameters to control the influence of geometry and environment. Obstacle geometry / media data 193 or parameters can be provided by a content provider and / or encoder 130. The audio intensity of an audio source 313 can be estimated as: F(Ci)=F(Bi)* Distance_function(OFD)*Directivity_gain_function(OFD)+Obstacle_attenuation_function(F(Bi), OFD, GHD). The first term corresponds to the contribution of sound passing around an obstacle 603. The second term corresponds to the contribution of sound passing through an obstacle 603.
[066] The minimum obstacle-free distance (OFD) 602 can be determined Petition 870260038484, dated 04 / 24 / 2026, page 29 / 132 20 / 50 using A*Dijkstra's pathfinding algorithm and can be used to control direct sound attenuation. The path distance (GHD) 601 can be used to control reverberation and distortion. Alternatively, or in addition, a broadcast approach can be used to describe the effects of an obstacle 603 on the intensity of an audio source 313.
[067] Fig. 7 illustrates an example of a field of view 701 of a listener 181 placed in the target listener position 302. Furthermore, Fig. 7 shows an example of a focus of attention 702 of a listener placed in the target listener position 302. The field of view 701 and / or the focus of attention 702 can be used to enhance (e.g., amplify) the audio from an audio source that is within the field of view 701 and / or the focus of attention 702. The field of view 701 can be considered a user-activated effect and can be used to activate a sound enhancer for audio sources 311 associated with the user's field of view 701. In particular, a cocktail party effect simulation can be performed by removing frequency blocks from a background audio source to improve the comprehensibility of a speech signal associated with the audio source 311 that is within the listener's field of view 701.The 702 attention focus can be viewed as a content-triggered effect and can be used to activate a sound enhancer for 311 audio sources associated with a region of content of interest (e.g., drawing the user's attention to look at and / or move in the direction of a 311 audio source).
[068] The audio intensity of an audio source 311 can be modified as: F(Bi)=Field_of_view_function(C,F(Bi), Field_of_view_data), where the field of view function describes the modification that is applied to an audio signal from an audio source 311 that is within the listener's field of view 701. Furthermore, the audio intensity of an audio source located within the listener's focus of attention 702 can be modified as: F(Bi)= Petition 870260038484, dated 04 / 24 / 2026, page 30 / 132 21 / 50 Attention_focus_function(F(Bi), Attention_focus_data), where attention_focus_function describes the modification that is applied to an audio signal from an audio source 311 that is located inside the attention focus 702.
[069] The functions described in this document for handling the transition of listener 181 from a source listener position 301 to a destination listener position 302 can be applied analogously to a change of position of an audio source 311, 312, 313.
[070] Therefore, this document describes efficient means for calculating coordinates and / or audio intensities of virtual audio objects or audio sources 311, 312, 313 representing a local VR audio scene 111 at arbitrary listener positions 301, 302. The coordinates and / or intensities can be determined by taking into account the distance attenuation curves of the sound source, the orientation and directivity of the sound source, the environmental geometry / media influence, and / or field-of-view and focus-of-attention data for further enhancements of the audio signal. The schemes described can significantly reduce computational complexity by performing calculations only if the listener position 301, 302 and / or the position of an audio object / source 311, 312, 313 changes.
[071] Furthermore, this document describes concepts for specifying distances, directivity, geometric functions, processing and / or signaling mechanisms for a VR 160 renderer. Additionally, a concept of minimum “obstacle-free distance” is described for controlling direct sound attenuation and “passage distance” for controlling reverberation and distortion. Furthermore, a concept for parameterizing the directivity of the sound source is described.
[072] Fig. 8 illustrates the handling of ambient sound sources 801, 802, 803 in the context of a local transition 192. In particular, Fig. 8 shows three sources of Petition 870260038484, dated 04 / 24 / 2026, page 31 / 132 22 / 50 different ambient sounds 801, 802, 803, wherein an ambient sound can be assigned to a point audio source. An ambient flag can be provided to the preprocessing unit 161, for the purpose of indicating that a point audio source 311 is an ambient audio source 801. Processing during a local and / or global transition of listener position 301, 302 may depend on the value of the ambient flag.
[073] In the context of a global transition 191, an ambient sound source 801 can be manipulated as a normal audio source 311. Fig. 8 illustrates a local transition 192. The position of an ambient sound source 801, 802, 803 can be copied from the source sphere 114 to the destination sphere 114, thus providing the position of the ambient sound source 811, 812, 813 at the destination listener position 302. Furthermore, the intensity of the ambient sound source 801 can be kept unchanged if the ambient conditions remain unchanged, F(Caí) = F(Baí). On the other hand, in the case of an obstacle 603, the intensity of an ambient sound source 803, 813 can be determined using the obstacle attenuation function, for example, as F(CAi)=F(BAi)*Distance_functionAi(OFD)+Obstacle_attenuation_function(F(BAi), OFD, GHD).
[074] Fig. 9a shows the flowchart of an example of method 900 for rendering audio in a virtual reality rendering environment 180. Method 900 can be executed by a VR audio renderer 160. Method 900 comprises rendering 901, an audio signal from an audio source 113 of an audio scene 111 from a source position on a sphere 114 around a listener position 201 of a listener 181. Rendering 901 can be performed using a 3D audio renderer 162 which may be limited to handling only 3DoF, notably which may be limited to handling rotational head movements of the listener 181. In particular, the 3D audio renderer 162 may not be configured to handle Petition 870260038484, dated 04 / 24 / 2026, page 32 / 132 23 / 50 head translation movements of the listener. The 3D audio renderer 162 may understand or may be an MPEG-H audio renderer.
[075] It should be noted that the expression render an audio signal from an audio source 113 from a specific source position indicates that the listener 181 perceives the audio signal as coming from the specific source position. The expression should not be understood as a limitation of how the audio signal is actually rendered. Several different rendering techniques can be used to render an audio signal from a specific source position, that is, to provide a listener 181 with the perception that an audio signal is coming from a specific source position.
[076] Furthermore, method 900 comprises determining 902 that listener 181 moves from listener position 201 within source audio scene 111 to listener position 202 within a different destination audio scene 112. Therefore, a global transition 191 from source audio scene 111 to destination audio scene 112 can be detected. In this context, method 900 may comprise receiving an indication that listener 181 moves from source audio scene 111 to destination audio scene 112. The indication may comprise or may be a flag. The indication may be signaled from listener 181 to VR audio renderer 160, for example, via a VR audio renderer 160 user interface.
[077] Typically, the source audio scene 111 and the destination audio scene 112 comprise one or more audio sources 113 that are different from each other. In particular, the source audio signals from one or more source audio sources 113 may not be audible in the destination audio scene 112 and / or the destination audio signals from one or more destination audio sources 113 may not be audible within the source audio scene 111.
[078] Method 900 may include (in response to the determination that Petition 870260038484, dated 04 / 24 / 2026, page 33 / 132 24 / 50 a global transition 191 to a new target audio scene 112 is performed) applying 903 a gradual output gain to the source audio signal to determine a modified source audio signal. Notably, the source audio signal is generated as it would be perceived at the listener's position in the source audio scene 111, regardless of the listener's movement 181 from listener position 201 within the source audio scene 111 to listener position 202 within the destination audio scene 112. Furthermore, method 900 can understand (in reaction to the determination that a global transition 191 to a new destination audio scene 112 is performed) rendering 904 the modified source audio signal from the source audio source 113 from the source source's position in the sphere 114 around listener positions 201, 202. These operations can be performed repeatedly, for example, at regular time intervals, during the global transition 191.
[079] Therefore, a global transition 191 between different audio scenes 111, 112 can be achieved by progressively dimming the source audio signals from one or more source audio sources 113 of the source audio scene 111. As a result, a computationally efficient and acoustically consistent global transition 191 is provided between different audio scenes 111, 112.
[080] It can be determined that listener 181 moves from source audio scene 111 to destination audio scene 112 during a transition time interval, wherein the transition time interval typically has a certain duration (e.g., 2s, 1s, 500ms or less). The global transition 191 can be performed progressively within the transition time interval. In particular, during the global transition 191, an intermediate time instant 213 within the transition time interval can be determined (e.g., according to a certain sampling rate of, for example, 100 ms, 50 ms, 20 ms or less). The gradual output gain can be determined based on a Petition 870260038484, dated 04 / 24 / 2026, page 34 / 132 25 / 50 relative location of the intermediate time instant 213 within the transition time interval.
[081] In particular, the transition time interval for the global transition 191 can be subdivided into a sequence of intermediate time instants 213. For each intermediate time instant 213 of the sequence of intermediate time instants 213, a gradual output gain to modify the source audio signals from one or more source audio sources can be determined. Furthermore, at each intermediate time instant 213 of the sequence of intermediate time instants 213, the modified source audio signals from one or more source audio sources 113 can be rendered from the source source position on the sphere 114 around the listener position 201,202. By doing this, an acoustically consistent global transition 191 can be achieved in a computationally efficient manner.
[082] Method 900 may comprise providing a gradual output function 211 indicating the gradual output gain at different intermediate time instants 213 within the transition time interval, wherein the gradual output function 211 is typically such that the gradual output gain decreases with progressing intermediate time instants 213, thus providing a smooth overall transition 191 to the target audio scene 112. In particular, the gradual output function 211 may be such that the source audio signal remains unmodified at the beginning of the transition time interval, that the source audio signal is increasingly attenuated at the ongoing intermediate time instants 213, and / or that the source audio signal is fully attenuated at the end of the transition time interval.
[083] The position of the source audio source 113 in the sphere 114 around the listener position 201,202 can be maintained as listener 181 moves from source audio scene 111 to destination audio scene 112 Petition 870260038484, dated 04 / 24 / 2026, page 35 / 132 26 / 50 (primarily throughout the entire transition time interval). Alternatively, or in addition, one can assume (throughout the entire transition time interval) that listener 181 remains in the same position as listeners 201, 202. By doing so, the computational complexity for a global transition 191 between audio scenes 111, 112 can be further reduced.
[084] Method 900 may further comprise determining a target audio signal from a target audio source 113 of the target audio scene 112. In addition, method 900 may comprise determining a target source position on the sphere 114 around listener positions 201, 202. Notably, the target audio signal is generated as it would be perceived at the listener position in the target audio scene 112, regardless of the listener 181’s movement from listener position 201 within the source audio scene 111 to listener position 202 within the target audio scene 112. Furthermore, method 900 may comprise applying a gradual input gain to the target audio signal to determine a modified target audio signal. The modified target audio signal from target audio source 113 can then be rendered from the target source position on sphere 114 around listener positions 201, 202.These operations can be performed repeatedly, for example, at regular time intervals, during the global transition 191.
[085] Therefore, in a manner analogous to the disappearance of source audio signals from one or more source audio sources 113 of source scene 111, destination audio signals from one or more destination audio sources 113 of destination scene 112 can be gradually entered, thus providing a smooth overall transition 191 between audio scenes 111, 112.
[086] As indicated above, listener 181 can move from source audio scene 111 to destination audio scene 112 during a transition time interval. The gradual input gain can be determined based on a Petition 870260038484, dated 04 / 24 / 2026, page 36 / 132 27 / 50 relative location of the intermediate time instant 213 within the transition time interval. In particular, a gradual input gain sequence can be determined for a corresponding sequence of intermediate time instants 213 during the global transition 191.
[087] Gradual input gains can be determined using a gradual input function 212 that indicates the gradual input gain at different intermediate time instants 213 within the transition time interval, wherein the gradual input function in 212 is typically such that the gradual input gain increases with the progress of intermediate time instants 213. In particular, the gradual input function 212 may be such that the target audio signal is fully attenuated at the beginning of the transition time interval, that the target audio signal is attenuated in a decreasing manner at the intermediate time instants 213, and / or that the target audio signal remains unchanged at the end of the transition time interval, thus providing a smooth overall transition 191 between audio scenes 111, 112 in a computationally efficient manner.
[088] In the same way that the position of the source audio of source audio 113, the position of the destination audio of destination audio 113 on the sphere 114 around the listener position 201,202 can be maintained as listener 181 moves from source audio scene 111 to destination audio scene 112, especially during the entire transition time interval. Alternatively or, in addition, one can assume (during the entire transition time interval) that listener 181 remains in the same position as listener 201, 202. By doing so, the computational complexity for a global transition 191 between audio scenes 111, 112 can be further reduced.
[089] The gradual output function 211 and the gradual input function 212 in combination provide a constant gain for a plurality of different Petition 870260038484, dated 04 / 24 / 2026, page 37 / 132 28 / 50 intermediate time instants 213. In particular, the gradual exit function 211 and the gradual entry function 212 can add up to a constant value (e.g., 1) for a plurality of different intermediate time instants 213. Therefore, the gradual entry function 212 and the gradual exit function 211 can be interdependent, thus providing a consistent audio experience during the overall transition 191.
[090] The gradual exit function 211 and / or the gradual entry function 212 may be derived from a bitstream 140 that is indicative of the source audio signal and / or the destination audio signal. The bitstream 140 may be provided by an encoder 130 to the VR audio renderer 160. Therefore, the global transition 191 may be controlled by a content provider. Alternatively or in addition, the gradual exit function 211 and / or the gradual entry function 212 may be derived from a virtual reality (VR) audio rendering storage unit 160 that is configured to render the source audio signal and / or the destination audio signal within the virtual reality rendering environment 180, thus providing reliable operation during global transitions 191 between audio scenes 111, 112.
[091] Method 900 may comprise sending an indication (e.g., a flag indicating) that listener 181 is moving from source audio scene 111 to destination audio scene 112 to an encoder 130, wherein the encoder 130 may be configured to generate a bitstream 140 that is indicative of the source audio signal and / or the destination audio signal. The indication may enable the encoder 130 to selectively provide the audio signals to one or more audio sources 113 of the source audio scene 111 and / or to one or more audio sources 113 of the destination audio scene 112 within the bitstream 140. Therefore, providing an indication for an upcoming global transition 191 allows a reduction in the bandwidth required for the bitstream 140. Petition 870260038484, dated 04 / 24 / 2026, page 38 / 132 29 / 50
[092] As indicated above, source audio scene 111 may comprise a plurality of source audio sources 113. Therefore, method 900 may comprise rendering a plurality of source audio signals from a corresponding plurality of source audio sources 113 from a plurality of different source source positions in the sphere 114 around listener positions 201, 202. Furthermore, method 900 may comprise applying gradual output gain to the plurality of source audio signals to determine a plurality of modified source audio signals. Additionally, method 900 may comprise rendering the plurality of modified source audio signals from source audio 113 from the corresponding plurality of source source positions in the sphere 114 around listener positions 201, 202.
[093] In an analogous manner, method 900 may comprise determining a plurality of target audio signals from a corresponding plurality of target audio sources 113 from the target audio scene 112. Furthermore, method 900 may comprise determining a plurality of target source positions in the sphere 114 around listener positions 201, 202. Furthermore, method 900 may comprise applying gradual input gain to the plurality of target audio signals to determine a corresponding plurality of modified target audio signals. Method 900 further comprises rendering the plurality of modified target audio signals from the plurality of target audio sources 113 from the corresponding plurality of target source positions in the sphere 114 around listener positions 201, 202.
[094] Alternatively or, in addition, the source audio signal that is rendered during a global transition 191 may be an overlay of audio signals from a plurality of source audio sources 113. In particular, at the beginning of the transition time interval, the audio signals from (all) audio sources Petition 870260038484, dated 04 / 24 / 2026, page 39 / 132 30 / 50 The audio signals from source audio scene 111 and target audio scene 112 can be combined to provide a combined source audio signal. This source audio signal can be modified with gradual output gain. Furthermore, the source audio signal can be updated at a specific sampling rate (e.g., 20 ms) during the transition time interval. Analogously, the target audio signal can correspond to a combination of the audio signals from a plurality of target audio sources 113 (primarily from all target audio sources 113). The combined target audio source can then be modified during the transition time interval using gradual input gain. By combining the audio signal from source audio scene 111 and target audio scene 112, respectively, the computational complexity can be further reduced.
[095] Furthermore, a virtual reality audio renderer 160 is described for rendering audio in a virtual reality rendering environment 180. As described in the present document, the VR audio renderer 160 may comprise a preprocessing unit 161 and a 3D audio renderer 162. The virtual reality audio renderer 160 is configured to render an audio signal from a source audio source 113 of a source audio scene 111 from a source source position on a sphere 114 around a listener position 201 of a listener 181. Furthermore, the VR audio renderer 160 is configured to determine that the listener 181 moves from the listener position 201 within the source audio scene 111 to a listener position 202 within a different target audio scene 112.Furthermore, the VR 160 audio renderer is configured to apply a gradual output gain to the source audio signal to determine a modified source audio signal and to render the modified source audio signal from source audio 113 from the source source position on sphere 114 around listener positions 201, 202. Petition 870260038484, dated 04 / 24 / 2026, page 40 / 132 31 / 50
[096] Furthermore, an encoder 130 that is configured to generate a bitstream 140 indicative of an audio signal to be rendered within a virtual reality rendering environment 180 is described. The encoder 130 can be configured to determine an audio signal from an audio source 113 of an audio scene 111. In addition, the encoder 130 can be configured to determine source position data relative to a source position of the audio source 113. The encoder 130 can then generate a bitstream 140 comprising the source audio signal and the source position data.
[097] Encoder 130 can be configured to receive an indication that a listener 181 is moving from source audio scene 111 to destination audio scene 112 within the virtual reality rendering environment 180 (for example, via a feedback channel from a VR audio renderer 160 towards encoder 130).
[098] The encoder 130 can then determine a target audio signal from a target audio source 113 of the target audio scene 112 and target position data relative to a target source position of the target audio source 113 (primarily only in reaction to receiving such an indication). Furthermore, the encoder 130 can generate a bitstream 140 comprising the target audio signal and the target position data. Therefore, the encoder 130 can be configured to supply the target audio signals from one or more target audio sources 113 of the target audio scene 112 selectively, only subject to receiving an indication for a global transition 191 to the target audio scene 112. By doing so, the bandwidth required for the bitstream 140 can be reduced.
[099] Fig. 9b shows a flowchart of a corresponding method 930 for generating a bitstream 140 indicative of an audio signal to be rendered within Petition 870260038484, dated 04 / 24 / 2026, page 41 / 132 32 / 50 a virtual reality rendering environment 180. Method 930 comprises determining 931 an audio signal from an audio source 113 of an audio scene 111. Furthermore, method 930 comprises determining 932 source position data relative to a source position of the audio source 113. Furthermore, method 930 comprises generating 933 a bitstream 140 comprising the audio signal and the source position data.
[0100] Method 930 comprises receiving 934 an indication that a listener 181 is moving from source audio scene 111 to destination audio scene 112 within the virtual reality rendering environment 180. In response to this, method 930 may comprise determining 935 a destination audio signal from a destination audio source 113 of the destination audio scene 112 and determining 936 destination position data relative to a destination source position of the destination audio source 113. Furthermore, method 930 comprises generating 937 a bitstream 140 comprising the destination audio signal and the destination position data.
[0101] Fig. 9c shows a flowchart of an example of method 910 for rendering an audio signal in a virtual reality rendering environment 180. Method 910 can be executed by a VR audio renderer 160.
[0102] Method 910 comprises rendering 911 an audio signal from an audio source 311, 312, 313 from an origin source position on an origin sphere 114 around an origin listener position 301 of a listener 181. Rendering 911 can be performed using a 3D audio renderer 162. In particular, rendering 911 can be performed under the assumption that the origin listener position 301 is fixed. Therefore, rendering 911 can be limited to three degrees of freedom (notably to a rotational movement of the head of listener 181). Petition 870260038484, dated 04 / 24 / 2026, page 42 / 132 33 / 50
[0103] To account for three additional degrees of freedom (for example, for a translational movement of listener 181), the method 910 can comprise determining 912 that listener 181 moves from source listener position 301 to destination listener position 302, wherein destination listener position 302 normally lies within the same audio scene 111. Therefore, it can be determined 912 that listener 181 performs a local transition 192 within the same audio scene 111.
[0104] In response to the determination that listener 181 performs a local transition 192, method 910 may comprise determining 913 a target source position of the audio source 311, 312, 313 on a target sphere 114 around the target listener position 302 based on the source source position. In other words, the position of the audio source 311, 312, 313 can be transferred from an origin sphere 114 around the origin listener position 301 to a destination sphere 114 around the destination position 302. This can be achieved by projecting the origin position from the origin sphere 114 to the destination sphere 114. In particular, the destination source position can be determined such that the destination source position corresponds to a ray intersection between the destination listener position 302 and the origin source position with the destination sphere 114.
[0105] Furthermore, method 910 can comprise (in reaction to the determination that listener 181 performs a local transition 192) determining 914 a target audio signal from audio source 311, 312, 313 based on the source audio signal. In particular, the intensity of the target audio signal can be determined based on the intensity of the source audio signal. Alternatively or further, the spectral composition of the target audio signal can be determined based on the spectral composition of the source audio signal. Therefore, it can be determined how the audio signal from audio source 311, 312, 313 is Petition 870260038484, dated 04 / 24 / 2026, page 43 / 132 34 / 50 perceived from the target listener position 302 (notably the intensity and / or spectral composition of the audio signal can be determined).
[0106] The aforementioned determination steps 913, 914 can be performed by a preprocessing unit 161 of the VR audio renderer 160. The preprocessing unit 161 can handle a translational movement of the listener 181 by transferring audio signals from one or more audio sources 311, 312, 313 from a source sphere 114 around the source listener position 114 to a destination sphere 301 to a destination sphere 114 around the destination listener position 302. As a result, the audio signals transferred from one or more audio sources 311, 312, 313 can also be rendered using a 3D audio renderer 162 (which may be limited to 3DoFs). Therefore, the 910 method allows for efficient provisioning of 6DoFs within a 180 VR audio rendering environment.
[0107] Consequently, method 910 may comprise rendering 915 the target audio signal from audio source 311, 312, 313 from target source position on target sphere 114 around target listener position 302 (for example, using a 3D audio renderer, such as the MPEG-H audio renderer).
[0108] Determining the target audio signal 914 may involve determining a target distance 322 between the source position and the target listener position 302. The target audio signal (notably the target audio signal intensity) may then be determined (notably scaled) based on the target distance 322. In particular, determining the target audio signal 914 may involve applying a distance gain 410 to the source audio signal, wherein the distance gain 410 is dependent on the target distance 322.
[0109] A distance function 415 can be provided, which is indicative of Petition 870260038484, dated 04 / 24 / 2026, page 44 / 132 35 / 50 distance gain 410 as a function of a distance 321, 322 between a source position of an audio signal 311, 312, 313 and a listener position 301, 302 of a listener 181. The distance gain 410 that is applied to the source audio signal (to determine the destination audio signal) can be determined based on the functional value of the distance function 415 for the destination distance 322. By doing this, the destination audio signal can be determined efficiently and accurately.
[0110] Furthermore, determining the target audio signal 914 may involve determining a source distance 321 between the source position and the source listener position 301. The target audio signal may then be determined (also) based on the source distance 321. In particular, the distance gain 410 that is applied to the source audio signal may be determined based on the functional value of the distance function 415 for the source distance 321. In a preferred example, the functional value of the distance function 415 for the source distance 321 and the functional value of the distance function 415 for the target distance 322 are used to rescale the intensity of the source audio signal to determine the target audio signal. Therefore, an efficient and accurate local transition 191 within an audio scene 111 can be provided.
[0111] Determining the target audio signal 914 may involve determining a directivity profile 332 of the audio source 311, 312, 313. The directivity profile 332 may be indicative of the intensity of the source audio signal in different directions. The target audio signal may then be determined (also) based on the directivity profile 332. By considering the directivity profile 332, the acoustic quality of a local transition 192 may be improved.
[0112] The 332 directivity profile may be indicative of a 510 directivity gain to be applied to the source audio signal to determine the destination audio signal. In particular, the 332 directivity profile may be indicative of a 515 directivity gain function, wherein the 515 directivity gain function Petition 870260038484, dated 04 / 24 / 2026, page 45 / 132 36 / 50 may indicate the directivity gain 510 as a function of a directivity angle (possibly two-dimensional) 520 between a source position of an audio source 311, 312, 313 and a listener position 301, 302 of a listener 181.
[0113] Therefore, determining the target audio signal 914 may involve determining a target angle 522 between the target source position and the target listener position 302. The target audio signal can then be determined based on the target angle 522. In particular, the target audio signal can be determined based on the functional value of the directivity gain function 515 for the target angle 522.
[0114] Alternatively or, in addition, determining the target audio signal may involve determining an origin angle 521 between the source position and the listener position 301. The target audio signal may then be determined based on the origin angle 521. In particular, the target audio signal may be determined based on the functional value of the directivity gain function 515 for the origin angle 521. In a preferred example, the target audio signal may be determined by modifying the intensity of the source audio signal using the functional value of the directivity gain function 515 for the origin angle 521 and for the target angle 522, to determine the intensity of the target audio signal.
[0115] In addition, method 910 may comprise determining target environment data 193 that are indicative of an audio propagation property of the medium between the target source position and the target listener position 302. The target environment data 193 may be indicative of an obstacle 603 that is positioned in a direct path between the target source position and the target listener position 302; indicative of information about the spatial dimensions of the obstacle 603; and / or indicative of an attenuation incurred by an audio signal in the direct path between the target source position and the target listener position 302. Petition 870260038484, dated 04 / 24 / 2026, page 46 / 132 37 / 50 destination listener 302. In particular, the destination environment data 193 may be indicative of an obstacle attenuation function of an obstacle 603, wherein the attenuation function may indicate an attenuation incurred by an audio signal passing through the obstacle 603 on the direct path between the destination source position and the destination listener position 302.
[0116] The target audio signal can then be determined based on the target environment data 193, further increasing the quality of the rendered audio within a VR rendering environment 180.
[0117] As indicated above, the target environment data 193 may be indicative of an obstacle 603 in the direct path between the target source position and the target listener position 302. The method 910 may comprise determining a through distance 601 between the target source position and the target listener position 302 in the direct path. The target audio signal may then be determined based on the through distance 601. Alternatively, or in addition, an obstacle-free distance 602 between the target source position and the target listener position 302 in an indirect path, which does not cross the obstacle 603, may be determined. The target audio signal may then be determined based on the obstacle-free distance 602.
[0118] In particular, an indirect component of the target audio signal can be determined based on the source audio signal propagating along the indication path. Furthermore, a direct component of the target audio signal can be determined based on the source audio signal propagating along the direct path. The target audio signal can then be determined by combining the indirect component and the direct component. By doing so, the acoustic effects of an obstacle 603 can be taken into account accurately and efficiently.
[0119] In addition, method 910 may involve determining focus information in relation to a field of vision 701 and / or a focus of attention 702 of the listener. Petition 870260038484, dated 04 / 24 / 2026, page 47 / 132 38 / 50 181. The target audio signal can then be determined based on focus information. In particular, a spectral composition of an audio signal can be tailored depending on the focus information. By doing so, a listener's VR experience 181 can be further enhanced.
[0120] Furthermore, method 910 may comprise determining that audio source 311, 312, 313 is an ambient audio source. In this context, an indication (e.g., a flag) may be received within a bitstream 140 of an encoder 130, wherein the indication indicates that audio source 311, 312, 313 is an ambient audio source. An ambient audio source typically provides a background audio signal. The source position of an ambient audio source may be maintained as the destination source position. Alternatively or in addition, the source audio signal strength of the ambient audio source may be maintained as the destination audio signal strength. By doing so, ambient audio sources may be handled efficiently and consistently in the context of a local transition 192.
[0121] The above-mentioned aspects are applicable to audio scenes 111, comprising a plurality of audio sources 311, 312, 313. In particular, method 910 may comprise rendering a plurality of source audio signals from a corresponding plurality of audio sources 311, 312, 313 from a plurality of different source source positions in the source sphere 114. Furthermore, method 910 may comprise determining a plurality of target source positions for the corresponding plurality of audio sources 311, 312, 313 in the target sphere 114 based on the plurality of source source positions, respectively. Furthermore, method 910 may comprise determining a plurality of target audio signals from the corresponding plurality of audio sources 311, 312, 313 based on the plurality of source audio signals, respectively. The plurality of destination audio signals of plurality Petition 870260038484, dated 04 / 24 / 2026, page 48 / 132 39 / 50 corresponding audio sources 311, 312, 313 can then be rendered from the corresponding plurality of target source positions in the target sphere 114 around the target listener position 302.
[0122] In addition, a virtual reality audio renderer 160 is described for rendering an audio signal in a virtual reality rendering environment 180. The audio renderer 160 is configured to render an audio signal from an audio source 311, 312, 313 from an origin source position in an origin sphere 114 around an origin listener position 301 of a listener 181 (notably using a 3D audio renderer 162 of the VR audio renderer 160).
[0123] Furthermore, VR audio renderer 160 is configured to determine that listener 181 moves from source listener position 301 to destination listener position 302. In response to this, VR audio renderer 160 can be configured (e.g., within a preprocessing unit 161 of VR audio renderer 160) to determine a destination source position of audio source 311, 312, 313 on a destination sphere 114 around destination listener position 302 based on the source source position and to determine a destination audio signal of audio source 311, 312, 313 based on the source audio signal.
[0124] In addition, VR audio renderer 160 (e.g., 3D audio renderer 162) can be configured to render the target audio signal from audio source 311, 312, 313 from the target source position on the target sphere 114 around the target listener position 302.
[0125] Therefore, the virtual reality audio renderer 160 may comprise a pre-processing unit 161 that is configured to determine the target source position and the target audio signal from audio sources 311, 312, 313. Furthermore, the VR audio renderer 160 may Petition 870260038484, dated 04 / 24 / 2026, page 49 / 132 40 / 50 comprise a 3D audio renderer 162 that is configured to render the target audio signal from audio source 311, 312, 313. The 3D audio renderer 162 can be configured to adapt the rendering of an audio signal from audio source 311, 312, 313 on a sphere (unit) 114 around a listener position 301, 302 of a listener 181, subject to a rotational movement of a listener's head 181 (to provide 3DoF within a rendering environment 180). Conversely, the 3D audio renderer 162 may not be configured to adapt the rendering of the audio signal from audio source 311, 312, 313, subject to a translational movement of the listener's head 181. Therefore, the 3D audio renderer 162 may be limited to 3 DoFs. Translational DoFs can then be efficiently provided using the 161 preprocessing unit, thus providing a global VR audio renderer 160 with 6 DoFs.
[0126] Furthermore, an audio encoder 130 configured to generate a bitstream 140 is described. The bitstream 140 is generated such that the bitstream 140 is indicative of an audio signal from at least one audio source 311, 312, 313 and indicative of a position of at least one audio source 311, 312, 313 within a rendering environment 180. In addition, the bitstream 140 may be indicative of environment data 193 with respect to an audio propagation property of the audio within the rendering environment 180. By signaling environment data 193 with respect to audio propagation properties, local transitions 192 within the rendering environment 180 may be precisely activated.
[0127] In addition, a bit stream 140 is described, which is indicative of an audio signal from at least one audio source 311, 312, 313; of a position of at least one audio source 311, 312, 313 within a rendering environment 180; and of environment data 193 indicative of a propagation property of Petition 870260038484, dated 04 / 24 / 2026, page 50 / 132 41 / 50 audio from the audio within the 180 rendering environment. Alternatively or in addition, the 140 bitstream may be indicative of whether or not the 311, 312, 313 audio source is an 801 environment audio source.
[0128] Fig. 9d shows an example flowchart of method 920 for generating a bitstream 140. Method 920 comprises determining 921 an audio signal from at least one audio source 311, 312, 313. In addition, method 920 comprises determining 922 position data relative to a position of at least one audio source 311, 312, 313 within a rendering environment 180. In addition, method 920 may comprise determining 923 environment data 193 indicative of an audio propagation property of the audio within the rendering environment 180. Method 920 further comprises inserting 934 the audio signal, the position data, and the environment data 193 into the bitstream 140. Alternatively or in addition, the indication may be of interest in the bitstream 140 to know whether the audio source 311, Is 312, 313 an audio source in the 801 environment or not?
[0129] Therefore, in the present document a virtual reality audio renderer 160 (a corresponding method) is described for rendering an audio signal in a virtual reality rendering environment 180. The audio renderer 160 comprises a 3D audio renderer 162 that is configured to render an audio signal from an audio source 113, 311, 312, 313 from a source position on a sphere 114 around a listener position 301, 302 of a listener 181 within the virtual reality rendering environment 180.Furthermore, the virtual reality audio renderer 160 comprises a pre-processing unit 161 that is configured to determine a new listener position 301, 302 from listener 181 within the virtual reality rendering environment 180 (within the same or within a different audio scene 111, 112). Additionally, the pre-processing unit 161 is configured to update the audio signal and the source position of the audio source 113, 311, 312, 313 in relation to each other. Petition 870260038484, dated 04 / 24 / 2026, page 51 / 132 42 / 50 to a sphere 114 around the new listener position 301, 302. The 3D audio renderer 162 is configured to render the updated audio signal from audio source 311, 312, 313 from the updated source position on sphere 114 around the new listener position 301, 302.
[0130] The methods and systems described in this document can be implemented as software, firmware, and / or hardware. Certain components can, for example, be implemented as software running on a digital signal processor or microprocessor. Other components can, for example, be implemented as hardware and / or as application-specific integrated circuits. The signals found in the described methods and systems can be stored in media such as random access memory or optical storage media. They can be transferred via networks, such as radio networks, satellite networks, wireless networks, or wired networks, for example, the Internet. Typical devices that utilize the methods and systems described in this document are portable electronic devices or other consumer equipment that are used to store and / or render audio signals.
[0131] Enumerated examples (EE) in this document are: EE 1) A method (900) for rendering audio in a virtual reality rendering environment (180), the method (900) comprising, - render (901), an audio signal from an audio source (113) from an audio scene (111) from a source position on a sphere (114) around a listener position (201) from a listener (181); - determine (902) that the listener (181) moves from listener position (201) within the source audio scene (111) to a listener position (202) within a different destination audio scene (112); - apply (903) a gradual output gain of the source audio signal to Petition 870260038484, dated 04 / 24 / 2026, page 52 / 132 43 / 50 determine a modified source audio signal; and - render (904) the modified source audio signal from source audio source (113) from the source source position on the sphere (114) around the listener position (201, 202). EE 2) The method (900) according to EE 1, wherein the method (900) comprises, - determine that the listener (181) moves from the source audio scene (111) to the destination audio scene (112) during a transition time interval; - determine an intermediate time instant (213) within the transition time interval; and - determine the gradual output gain based on a relative location of the intermediate time instant (213) within the transition time interval. EE 3) The method (900) according to EE 2, where - the method (900) comprises providing a phased output function (211) indicating the phased output gain at different intermediate time instants (213) within the transition time interval; and - the gradual exit function (211) is such that the gradual exit gain decreases with the progression of intermediate time instants (213). EE 4) The method (900) according to EE 3, wherein the gradual exit function (211) is such that - the source audio signal remains unmodified at the beginning of the transition time interval; and / or - the source audio signal is increasingly attenuated at intermediate time points in progress (213); and / or The source audio signal is completely attenuated at the end of the transition time interval. Petition 870260038484, dated 04 / 24 / 2026, page 53 / 132 44 / 50 EE 5) Method (900) in accordance with any previous EEs, wherein method (900) comprises - maintain the position of the source audio source (113) in the sphere (114) around the listener position (201,202) as the listener (181) moves from the source audio scene (111) to the destination audio scene (112)); and / or - maintain listener position (201, 202) unchanged as listener (181) moves from source audio scene (111) to destination audio scene (112). EE 6) Method (900) in accordance with any previous EEs, wherein method (900) comprises - determine a target audio signal from a target audio source (113) of the target audio scene (112); - determine a target source position on the sphere (114) around the listener position (201, 202); - Apply a gradual input gain to the target audio signal to determine a modified target audio signal; and - render the modified target audio signal from the target audio source (113) from the target source position on the sphere (114) around the listener position (201, 202). EE 7) The method (900) according to EE 6, wherein the method (900) comprises, - determine that the listener (181) moves from the source audio scene (111) to the destination audio scene (112) during a transition time interval; - determine an intermediate time instant (213) within the transition time interval; and Petition 870260038484, dated 04 / 24 / 2026, page 54 / 132 45 / 50 - determine the gradual input gain based on a relative location of the intermediate time instant (213) within the transition time interval. EE 8) The method (900) according to EE 7, where - the method (900) comprises providing a step-in function (212) indicating the step-in gain at different intermediate time instants (213) within the transition time interval; and - the gradual input function (212) is such that the gradual input gain increases with the progression of intermediate time instants (213). EE 9) The method (900) according to EE 8, wherein the gradual input function (212) is such that - the destination audio signal remains unchanged at the end of the transition time interval; and / or - the destination audio signal is attenuated decreasingly at intermediate time instants in progress (213); and / or The target audio signal is fully attenuated at the beginning of the transition time interval. EE 10) The (900) method according to any of EEs 6 to 9, wherein the (900) method comprises - maintain the position of the target audio source (113) in the sphere (114) around the listener position (201,202) as the listener (181) moves from the source audio scene (111) to the target audio scene (112)); and - maintain listener position (201, 202) unchanged as listener (181) moves from source audio scene (111) to destination audio scene (112). EE 11) The method (900), according to EE 8, referring to EE 3, in which the gradual output function (211) and the gradual input function (212) in combination provide Petition 870260038484, dated 04 / 24 / 2026, p. 55 / 132 46 / 50 a constant gain for a plurality of different intermediate instants of time (213). EE 12) The method (900), according to EE 8, referring to EE 3 where the gradual output function (211) and / or the gradual input function (212) - are derived from a bit stream (140) that is indicative of the source audio signal and / or the destination audio signal; and / or - are derived from a storage unit of a virtual reality audio renderer (160) configured to render the source audio signal and / or the destination audio signal within the virtual reality rendering environment (180). EE 13) The method (900) in accordance with any previous EEs, wherein the method (900) comprises receiving an indication that the listener (181) moves from the source audio scene (111) to the destination audio scene (112). EE 14) The method (900) according to EE 13, wherein the indication comprises a signal. EE 15) The method (900) in accordance with any of the preceding EEs, wherein the method (900) comprises, sending an indication that the listener (181) moves from the source audio scene (111) to the destination audio scene (112) to an encoder (130); wherein the encoder (130) is configured to generate a bitstream (140) that is indicative of the source audio signal. EE 16) The method (900) in accordance with any of the preceding EEs, wherein the first audio signal is rendered using a 3D audio renderer (162), notably an MPEG-H audio renderer. EE 17) The method (900) in accordance with any previous EE, wherein the method (900) comprises, - render a plurality of source audio signals from a corresponding plurality of source audio sources (113) from a Petition 870260038484, dated 04 / 24 / 2026, page 56 / 132 47 / 50 plurality of different source positions in the sphere (114) around the listener position (201, 202); - Apply the gradual output gain to the plurality of source audio signals to determine a plurality of modified source audio signals; and - render the plurality of modified source audio signals from the source audio source (113) from the corresponding plurality of source source positions in the sphere (114) around the listener position (201,202). EE 18) The (900) method according to any of EEs 6 to 17, wherein the (900) method comprises, - determine a plurality of target audio signals from a corresponding plurality of target audio sources (113) from the target audio scene (112); - determine a plurality of source-destination positions in the sphere (114) around the listener position (201, 202); and - Apply the gradual input gain to the plurality of target audio signals to determine a corresponding plurality of modified target audio signals; and - render the plurality of modified target audio signals from the plurality of target audio sources (113) from the corresponding plurality of target source positions in the sphere (114) around the listener position (201, 202). EE 19) The method (900) in accordance with any of the previous EEs, wherein the source audio signal is an overlay of audio signals from a plurality of source audio sources (113). EE 20) A virtual reality audio renderer (160) for rendering audio in a virtual reality rendering environment (180), wherein the virtual reality audio renderer (160) is configured for - Render an audio signal from an audio source. Petition 870260038484, dated 04 / 24 / 2026, page 57 / 132 48 / 50 (113) of an audio source scene (111) from a source source position on a sphere (114) around a source listener position (201) of a listener (181); - determine that the listener (181) moves from listener position (201) within the source audio scene (111) to a listener position (202) within a different destination audio scene (112); - Apply a gradual output gain to the source audio signal to determine a modified source audio signal; and - render the modified source audio signal from the source audio source (113) from the source source position on the sphere (114) around the listener position (201,202). EE 21) An encoder (130) configured to generate a bitstream (140) indicative of an audio signal to be rendered within a virtual reality rendering environment (180); wherein the encoder (130) is configured to - determine an audio signal from an audio source (113) of an audio scene (111); - determine source position data relative to a source position of the source audio source (113); - generate a bitstream (140) comprising the source audio signal and the source position data; - receive an indication that a listener (181) moves from the source audio scene (111) to a destination audio scene (112) within the virtual reality rendering environment (180); - determine a target audio signal from a target audio source (113) of the target audio scene (112); - determine the destination position data relative to a position of Petition 870260038484, dated 04 / 24 / 2026, page 58 / 132 49 / 50 destination source of the destination audio source (113); and - generate a bitstream (140) comprising the target audio signal and the target position data. EE 22) A method (930) for generating a bitstream (140) indicative of an audio signal to be rendered within a virtual reality rendering environment (180); the method (930) comprising, - determine (931) an audio signal from an audio source (113) from an audio scene (111); - determine (932) source position data relative to a source source position of the source audio source (113); - generate (933) a bit stream (140) comprising the source audio signal and the source position data; - receive (934) an indication that a listener (181) moves from the source audio scene (111) to a destination audio scene (112) within the virtual reality rendering environment (180); - determine (935) a target audio signal from a target audio source (113) from the target audio scene (112); - determine (936) the target position data relative to a target source position of the target audio source (113); and - generate (937) a bit stream (140) comprising the target audio signal and the target position data. EE 23) A virtual reality audio renderer (160) for rendering an audio signal in a virtual reality rendering environment (180), wherein the audio renderer (160) comprises, - a 3D audio renderer (162) that is configured to render an audio signal from an audio source (113) from a source position on a sphere (114) around a listener position (201, 202) of a listener (181) Petition 870260038484, dated 04 / 24 / 2026, page 59 / 132 50 / 50 within the virtual reality rendering environment (180); - a preprocessing unit (161) configured for - determine a new listener position (201,202) of listener (181) within the virtual reality rendering environment (180); and - update the audio signal and the position of the audio source (201, 202) relative to a sphere (114) around the new position of the listener (201,202); wherein the 3D audio renderer (162) is configured to render the updated audio signal from the audio source (113) from the updated source position on the sphere (114) around the new listener position (201, 202). Petition 870260038484, dated 04 / 24 / 2026, page 60 / 132
Claims
1 / 1 CLAIMS 1. Encoder (130) configured to generate a bitstream (140) indicative of an audio signal to be rendered within a virtual reality rendering environment (180); the encoder (130) CHARACTERIZED in that it is configured to: - determine an audio signal from an audio source (113) of an audio scene (111); - determine source position data relative to a source position of the source audio source (113); - generate a bitstream (140) comprising the source audio signal and the source position data; - receive an indication that a listener (181) moves from the source audio scene (111) to a destination audio scene (112) within the virtual reality rendering environment (180); - determine a target audio signal from a target audio source (113) of the target audio scene (112); - determine the target position data relative to a target audio source position (113); and generate a bitstream (140) comprising the target audio signal and the target position data. Petition 870260038484, dated 04 / 24 / 2026, page 61 / 132