Method and system for controlling the directionality of audio sources in a virtual reality environment - Patents.com

JP2024521689A5Active Publication Date: 2025-05-19DOLBY INTERNATIONAL AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023571739
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-17
Filing Date
2022-05-10
Publication Date
2025-05-19
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

Existing audio rendering systems are limited in handling larger translational changes and associated degrees of freedom (DoF) in virtual reality (VR) environments, failing to efficiently manage audio source directionality and occlusion, leading to discontinuities and resource inefficiencies.

Method used

A method and system for rendering audio signals in VR environments that dynamically adjust directional patterns based on listener position and orientation, using directional control functions to apply directionality only when relevant, thereby optimizing resource usage and avoiding perceptual artifacts.

Benefits of technology

This approach enhances the quality of 3D audio rendering by reducing computational complexity and preventing discontinuities in sound intensity, while maintaining high perceptual quality by applying directionality only where necessary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method (700) for rendering audio signals of audio sources (211, 212, 213) in a virtual reality rendering environment (180) is described. The method (700) includes a step of determining (701) whether directivity patterns (232) of the audio sources (211, 212, 213) should be taken into account with respect to a listening situation of a listener (181) in the virtual reality rendering environment (180). Furthermore, the method (700) includes a step of rendering (702) the audio signals of the audio sources (211, 212, 213) without taking into account the directivity patterns (232) of the audio sources (211, 212, 213) if it is determined that the directivity patterns (232) of the audio sources (211, 212, 213) should not be taken into account with respect to the listening situation of the listener (181). On the other hand, the method (700) includes a step (703) of rendering the audio signals of the audio sources (211, 212, 213) depending on the directional patterns (232) of the audio sources (211, 212, 213) if it is determined that the directional patterns (232) should be taken into account for the listening situation of the listener (181).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: U.S. Provisional Application No. 63 / 189,269, filed May 17, 2021 (Reference No. D21027USP1), and European Application No. 21174024.6, filed May 17, 2021 (Reference No. D21027EP), which are incorporated by reference herein.

[0002] This specification relates to efficient and consistent handling of audio source directionality in virtual reality (VR) rendering environments. [Background technology]

[0003] Virtual reality (VR), augmented reality (AR) and / or mixed reality (MR) applications are rapidly evolving with increasingly sophisticated acoustic models of audio sources and scenes that can be enjoyed from different perspectives and / or viewpoints and different listening positions. For example, for VR applications, two different classes of flexible audio representations can be employed: sound field representations and object-based representations. Sound field representations are physically based approaches that encode the incident wavefront at the listening position. For example, approaches such as B-format or Higher Order Ambisonics (HOA) represent the spatial wavefront using spherical harmonic decomposition. Object-based approaches represent complex auditory scenes as a collection of individual elements that include audio waveforms or signals and associated parameters or metadata (possibly time-varying).

[0004] Enjoying VR, AR and / or MR applications may include users experiencing different auditory perspectives or viewpoints. For example, room-based virtual reality may be provided based on mechanisms using six degrees of freedom (DoF). 6DoF interaction may include translational movements (forward / backward, up / down and left / right) and rotational movements (pitch, yaw and roll). Unlike 3DoF spherical video experiences that are limited to head rotation, content generated for 6DoF interaction allows navigation in the virtual environment (e.g., physical walking in a room) in addition to head rotation. This can be achieved based on position trackers (e.g., camera-based) and orientation trackers (e.g., gyroscopes and / or accelerometers). 6DoF tracking technology may be available on desktop VR systems (e.g., PlayStation®VR, Oculus Rift, HTC Vive) and mobile VR platforms (e.g., Google Tango). The user's experience of the directionality and spatial extent of a sound or audio source is important to the realism of a 6DoF experience, particularly for navigation through a scene and around a virtual audio source.

[0005] Existing audio rendering systems (such as MPEG-H 3D audio renderers) are typically limited to rendering 3DoF (i.e., rotational movement of the audio scene caused by movement of the listener's head) or 3DoF+, which adds small translational changes of the listener's listening position but does not consider effects such as directionality or occlusion. Larger translational changes of the listener's listening position and the associated DoFs typically cannot be handled by those renderers. Summary of the Invention [Problem to be solved by the invention]

[0006] This specification relates to the technical problem of providing a resource-efficient method and system for handling translation in the context of audio rendering. In particular, this specification addresses the technical problem of handling the directionality of an audio source in a resource-efficient and consistent manner within 6DoF audio rendering. [Means for solving the problem]

[0007] overview According to one aspect, a method for rendering an audio signal of an audio source in a virtual reality rendering environment is described. The method includes a step of determining whether a directivity pattern of the audio source should be taken into account for a (current) listening situation of a listener in the virtual reality rendering environment. Furthermore, the method includes a step of rendering the audio signal of the audio source without taking into account the directivity pattern of the audio source if it is determined that the directivity pattern of the audio source should not be taken into account for the listening situation of the listener. In addition, the method includes a step of rendering the audio signal of the audio source depending on the directivity pattern of the audio source if it is determined that the directivity pattern should be taken into account for the listening situation of the listener.

[0008] According to a further aspect, a method for rendering an audio signal of a first audio source to a listener in a virtual reality rendering environment is described. It should be noted that the term virtual reality rendering environment should also include an augmented and / or mixed reality rendering environment. The method includes determining a control value for a listening situation of the listener in the virtual reality rendering environment based on a directivity control function. Furthermore, the method includes adjusting a directivity pattern, in particular a directivity gain, of the first audio source depending on the control value. In addition, the method includes rendering the audio signal of the first audio source to the listener in the virtual reality rendering environment depending on the adjusted directivity pattern of the first audio source, in particular depending on the adjusted directivity gain.

[0009] According to a further aspect, a virtual reality audio renderer for rendering an audio signal of an audio source in a virtual reality rendering environment is described. The audio renderer is configured to determine whether a directivity pattern of the audio source should be taken into account for a listening situation of a listener in the virtual reality rendering environment. In addition, the audio renderer is configured to render the audio signal of the audio source without taking into account the directivity pattern of the audio source if it is determined that the directivity pattern of the audio source should not be taken into account for the listening situation of the listener. Furthermore, the audio renderer is configured to render the audio signal of the audio source depending on the directivity pattern of the audio source if it is determined that the directivity pattern should be taken into account for the listening situation of the listener.

[0010] According to another aspect, a virtual reality audio renderer for rendering an audio signal of a first audio source to a listener in a virtual reality rendering environment is described. The audio renderer is configured to determine a control value for a listening situation of the listener in the virtual reality rendering environment based on a directivity control function (e.g., provided in a bitstream). Furthermore, the audio renderer is configured to adjust a directivity pattern of the first audio source (e.g., provided in a bitstream) depending on the control value. Furthermore, the audio renderer is configured to render the audio signal of the first audio source to the listener in the virtual reality rendering environment depending on the adjusted directivity pattern of the first audio source.

[0011] According to a further aspect, a method for generating a bitstream is described. The method includes determining an audio signal of at least one audio source and determining a source position in a virtual reality rendering environment of the at least one audio source. In addition, the method includes determining a (non-uniform) directivity pattern of the at least one audio source and determining a directivity control function for controlling the use of the directivity pattern for rendering the audio signal of the at least one audio source depending on a listening situation of a listener in the virtual reality rendering environment. Furthermore, the method includes inserting data related to the audio signal, the source position, the directivity pattern and the directivity control function into the bitstream.

[0012] According to a further aspect, an audio encoder configured to generate a bitstream is described. The bitstream may indicate an audio signal of at least one audio source and / or a source position in a virtual reality rendering environment of the at least one audio source. Furthermore, the bitstream may indicate a directivity pattern of the at least one audio source and / or a directivity control function for controlling the use of the directivity pattern for rendering the audio signal of the at least one audio source depending on a listening situation of a listener in the virtual reality rendering environment.

[0013] According to another aspect, a bitstream and / or a syntax for the bitstream is described. The bitstream may indicate an audio signal of at least one audio source and / or a source position in a virtual reality rendering environment of the at least one audio source. Furthermore, the bitstream may indicate a directivity pattern of the at least one audio source and / or a directivity control function for controlling the use of the directivity pattern for rendering the audio signal of the at least one audio source depending on a listening situation of a listener in the virtual reality rendering environment. The bitstream may include one or more data elements including data related to the above information.

[0014] According to a further aspect, a software program is described, which may be configured to run on a processor and, when executed on the processor, may be configured to perform the method steps outlined herein.

[0015] According to another aspect, a computer readable storage medium is described. The computer readable storage medium may comprise a software program configured to run on a processor (or computer) and configured to perform the method steps outlined herein when executed on the processor.

[0016] According to a further aspect, a computer program product is described. The computer program may comprise executable instructions for performing the method steps outlined herein when executed on a computer.

[0017] It should be noted that the methods and systems, including the preferred embodiments, outlined in this patent application may be used alone or in combination with other methods and systems disclosed herein. Furthermore, all aspects of the methods and systems outlined in this patent application may be combined in any manner. In particular, the features of the claims may be combined with each other in any manner. [Brief description of the drawings]

[0018] The invention is described below, by way of example, with reference to the accompanying drawings, in which:

[0019] [Figure 1a] FIG. 1a illustrates an example audio processing system for providing 6DoF audio. [Figure 1b] FIG. 1b illustrates an example situation in a 6DoF audio and / or rendering environment. [Diagram 2] FIG. 2 illustrates an example audio scene. [Figure 3a] FIG. 3a illustrates the remapping of audio sources in response to changes in listening position within the audio scene. [Figure 3b] FIG. 3b illustrates an example distance function. [Figure 4a] FIG. 4a illustrates an audio source with a non-uniform directivity pattern. [Figure 4b] FIG. 4b illustrates an example directivity function of an audio source. [Figure 5a] FIG. 5a illustrates an example setup for measuring the directivity pattern. [Figure 5b] FIG. 5b illustrates an example directional gain when passing a virtual audio source. [Figure 6a] FIG. 6a illustrates an example directional control function. [Figure 6b] FIG. 6b illustrates an example decay function. [Figure 7a] FIG. 7a illustrates a flowchart of an exemplary method for rendering a 3D audio signal of an audio source within an audio scene. [Figure 7b] FIG. 7b illustrates a flowchart of an exemplary method for rendering a 3D audio signal of an audio source within an audio scene. [Figure 7c]FIG. 7c illustrates a flowchart of an exemplary method for generating a bitstream for a virtual reality audio scene. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0020] Detailed Description As outlined above, the present specification relates to efficiently and consistently providing 6DoF in a 3D (three-dimensional) audio environment. FIG. 1a illustrates a block diagram of an exemplary audio processing system 100. An acoustic environment 110, such as a stadium, may comprise a variety of different audio sources 113. Exemplary audio sources 113 in a stadium are individual spectators, stadium speakers, performers on the field, etc. The acoustic environment 110 may be subdivided into different audio scenes 111, 112. As an example, a first audio scene 111 may correspond to a cheering block for the home team, and a second audio scene 112 may correspond to a cheering block for the guest team. Depending on where the listener is located in the audio environment, the listener will perceive either the audio source 113 from the first audio scene 111 or the audio source 113 from the second audio scene 112.

[0021] The different audio sources 113 of the audio environment 110 may be captured using audio sensors 120, in particular microphone arrays. One or more audio scenes 111, 112 of the audio environment 110 may be described using multi-channel audio signals, one or more audio objects and / or Higher Order Ambisonics (HOA) and / or First Order Ambisonics signals. In the following, it is assumed that the audio source 113 is associated with audio data captured by one or more audio sensors 120. Here, the audio data indicates the audio signal (transmitted by the audio source 113) and the position of the audio source 113 as a function of time (at a certain sampling rate, e.g., 20 ms).

[0022] A 3D audio renderer, such as an MPEG-H 3D audio renderer, typically assumes that a listener 181 is located at a specific (fixed) listening position 182 within the audio scenes 111, 112. Audio data for different audio sources 113 of the audio scenes 111, 112 is typically provided assuming that the listener 181 is located at this specific listening position 182. The audio encoder 130 may comprise a 3D audio encoder 131 configured to encode audio data of one or more audio sources 113 of one or more audio scenes 111, 112 of the audio environment 110.

[0023] Additionally, virtual reality (VR) metadata may be provided that allows the listener 181 to change a listening position 182 within the audio scene 111, 112 and / or move between different audio scenes 111, 112. The encoder 130 may comprise a metadata encoder 132 configured to encode the VR metadata. The encoded VR metadata and the encoded audio data of the audio sources 113 may be combined in a combiner 133 to provide a bitstream 140 indicative of the audio data and the VR metadata. The VR metadata may include, for example, environmental data describing the acoustics of the audio environment 110.

[0024] The bitstream 140 may be decoded using the decoder 150 to provide (decoded) audio data and (decoded) VR metadata. An audio renderer 160 for rendering audio in a rendering environment 180 enabling 6DoF may comprise a preprocessing unit 161 and a (conventional) 3D audio renderer 162 (such as an MPEG-H 3D audio renderer). The preprocessing unit 161 may be configured to determine a listening position 182 in the listening environment 180 of the listener 181. The listening position 182 may indicate an audio scene 111 in which the listener 181 is located. Furthermore, the listening position 182 may indicate an exact position in the audio scene 111. The preprocessing unit 161 may further be configured to determine a 3D audio signal for the current listening position 182 based on the (decoded) audio data and possibly based on the (decoded) VR metadata. The 3D audio signal may then be rendered using the 3D audio renderer 162. The 3D audio signal may include audio signals of one or more audio sources 113 of the audio scene 111.

[0025] It should be noted that the concepts and techniques described herein may be specified in a frequency-varying manner, may be defined either globally or object / media-dependently, may be applied directly in the spectral or time domain, and / or may be hard-coded into the VR renderer 160 or may be specified via a corresponding input interface.

[0026] 1b shows an exemplary rendering environment 180. A listener 181 may be located in an origin audio scene 111. For rendering purposes, audio sources 113, 194 may be assumed to be located at different rendering positions on a (unity) sphere 114 around the listener 181, in particular around a listening position 182. The rendering positions of the different audio sources 113, 194 may change over time (according to a given sampling rate). The positions of the different audio sources 113, 194 may be indicated in the VR metadata. Different situations may occur within the VR rendering environment 180. The listener 181 may make a global transition 191 from the origin audio scene 111 to the destination audio scene 112. Alternatively or additionally, the listener 181 may make a local transition 192 to a different listening position 182 within the same audio scene 111. Alternatively or in addition, the audio scene 111 may indicate acoustically relevant environmental properties (e.g., walls), which may be described using environmental data 193 and which should be taken into account when changes occur to the listening position 182. The environmental data 193 may be provided as VR metadata. Alternatively or in addition, the audio scene 111 may include one or more ambient audio sources 194 (e.g., for background noise), which should be taken into account when changes occur to the listening position 182.

[0027] 2 illustrates an exemplary local transition 192 from a source listening position B 201 to a destination listening position C 202 within the same audio scene 111. The audio scene 111 includes different audio sources or objects 211, 212, 213. The different audio sources or objects 211, 212, 213 may have different directivity profiles 232 (also referred to herein as directivity patterns). The directivity profiles 232 of one or more audio sources or objects 211, 212, 213 may be indicated as VR metadata. Furthermore, the audio scene 111 may have environmental characteristics, in particular one or more obstacles, that affect the propagation of audio within the audio scene 111. The environmental characteristics may be described using environmental data 193. In addition, the relative distances 221, 222 from the audio object 211 to the different listening positions 182, 201, 202 may be known (eg, based on one or more sensors of the VR renderer 160, such as a gyroscope or accelerometer).

[0028] 3a and 3b illustrate an approach for handling the effect of a local transition 192 on the intensity of different audio sources or objects 211, 212, 213. As outlined above, the audio sources 211, 212, 213 of an audio scene 111 are typically assumed by the 3D audio renderer 162 to be located on a sphere 114 around the listening position 201. Thus, at the start of the local transition 192, the audio sources 211, 212, 213 may be located on the source sphere 114 around the source listening position 201, and at the end of the local transition 192, the audio sources 211, 212, 213 may be located on the destination sphere 114 around the destination listening position 202. The audio sources 211, 212, 213 may be remapped from the source sphere 114 to the destination sphere 114. To do so, a ray may be considered from the destination listening position 202 to the source positions of the audio sources 211, 212, 213 on the source sphere 114. The audio sources 211, 212, 213 may be located on the intersection of the ray with the destination sphere 114.

[0029] The intensity F of the audio sources 211, 212, 213 on the destination sphere 114 is typically different from the intensity on the source sphere 114. The intensity F may be modified using an intensity gain function or distance function 315 (also referred to herein as attenuation function) that provides a distance gain 310 (also referred to herein as attenuation gain) as a function of the distance 320 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202. The distance function 315 typically exhibits a cutoff distance 321 beyond which a zero distance gain 310 is applied. The source distance 221 from the audio source 211 to the source listening position 201 provides the source gain 311. Furthermore, the destination distance 222 from the audio source 211 to the destination listening position 202 provides the destination gain 312. The intensity F of the audio source 211 may be rescaled using the source gain 311 and the destination gain 312, thereby giving the intensity F of the audio source 211 on the destination sphere 114. In particular, the intensity F of the source audio signal of the audio source 211 on the source sphere 114 may be divided by the source gain 311 and multiplied by the destination gain 322 to give the intensity F of the destination audio signal of the audio source 211 on the destination sphere 114.

[0030] Thus, the position of the audio source 211 after the local transition 192 is C i =source_remap_function(B i , C) (e.g., using a geometric transformation). Furthermore, the intensity of the audio source 211 after the local transition 192 may be determined as F(C i )=F(B i )*distance_function(B i ,C i , C). Thus, the distance attenuation can be modeled by a corresponding intensity gain 315 given by the distance function 315.

[0031] 4a and 4b illustrate an audio source 212 having a non-uniform directivity profile 232. The directivity profile may be defined using a directivity gain 410 that indicates gain values ​​for different directions or directivity angles 420. In particular, the directivity profile 232 of the audio source 212 may be defined using a directivity gain function 415 that indicates the directivity gain 410 as a function of the directivity angle 420 (the angle 420 may range from 0° to 360°). Note that for a 3D audio source 212, the directivity angle 420 is typically a two-dimensional angle that includes an azimuth angle and an elevation angle. Thus, the directivity gain function 415 is typically a two-dimensional function of the quadratic source directivity angle 420.

[0032] The directional profile 232 of the audio source 212 may be considered with respect to the local transition 192 by determining a source directional angle 421 of a source ray between the audio source 212 and the source listening position 201 (where the audio source 212 is located on the source sphere 114 around the source listening position 201) and a destination directional angle 422 of a destination ray between the audio source 212 and the destination listening position 202 (where the audio source 212 is located on the destination sphere 114 around the destination listening position 202). Using a directional gain function 415 of the audio source 212, the source directional gain 411 and the destination directional gain 412 may be determined as function values ​​of the directional gain function 415 for the source directional angle 421 and the destination directional angle 422, respectively (see FIG. 4b). The strength F of the audio source 212 at the source listening position 201 may then be divided by the source directional gain 411 and multiplied by the destination directional gain 412 to determine the strength F of the audio source 212 at the destination listening position 202.

[0033] Thus, the sound source directivity may be parameterized by a directional factor or gain 410, indicated by a directional gain function 415. The directional gain function 415 may indicate the strength of the audio source 212 at a specified distance as a function of the angle 420 relative to the listening position 182, 201, 202. The directional gain 410 may be defined as a ratio to the gain of an audio source 212 at the same distance and with the same total power radiated uniformly in all directions. The directional profile 232 may be parameterized by a set of gains 410 corresponding to vectors starting from the center of the audio source 212 and ending at points distributed on a unit sphere around the center of the audio source 212.

[0034] The audio intensity of the audio source 212 at the destination listening position 202 obtained as above is expressed as F(C i )=F(B i )*Distance_function()*Directivity_gain_function(C i ,C,Directivity_parametrization), where Directivity_gain_function depends on the directional profile 232 of the audio source 212. Distance_function() takes into account the modified strength due to the change in distance 321, 322 of the audio source 212 due to the transition of the listening positions 201, 202.

[0035] 5a shows an example setup for measuring a directivity profile 232 of an audio source 211 at a given distance 320. For this purpose, the audio sensors 120, in particular the microphones, may be arranged on a circumference centered on the audio source 211, where the circumference has a radius corresponding to the given distance 320. The strength or magnitude of the audio signal may be measured for different corners 420 on the circumference to give a directivity pattern P(D) 232 for a distance D 320. Such a directivity pattern or profile may be determined for a number of different distance values ​​D 320.

[0036] The directional pattern data may represent: sound emission characteristics resulting from an audio source or object 211 generating a sound field; The effect of local sound occlusion caused by nearby obstacles and affecting the sound field, and / or · The intent of the content creator. Directivity pattern data is typically measured (and only available) for a particular range of distances from the audio source 211, as illustrated in FIG. 5a.

[0037] If the directional gain from the directional pattern 232 is applied directly to the rendered audio signal, one or more problems may arise: directional considerations may often be perceptually irrelevant for listening positions 182, 201, 202 that are relatively far from the audio object 211. This may be because: Relatively low total object sound energy (e.g. due to dominant distance attenuation effects), Psychoacoustic masking effects (caused by one or more other audio sources 212, 213 in close proximity to the listening position 201, by reverberance and / or by early reflections), a relatively low subjective importance of the audio object 211 to the listener's attention, and / or · Lack of clear visual indication of the orientation of the audio object 211 corresponding to the directionality of the object 211.

[0038] A further problem may be that applying directivity results in a discontinuity in sound intensity at the origin 500 (i.e., source location) of the audio source 211, as illustrated in Fig. 5b. This acoustic artifact may be perceived when a listener passes the origin 500 of the audio source 211 (especially an audio source 211 represented by a volumetric object) in the virtual audio scene 111. While passing such a virtual object, the user may perceive an abrupt sound level change at the center point 500 of the audio source 211. This is illustrated in Fig. 5b, which shows a sound level 501 when approaching the center point 500 of the audio source 211 from the front, and a sound level 502 when moving away from the center point 500 behind the audio source 211. In Fig. 5b, it can be seen that a discontinuity occurs in the sound levels 501, 502 at the center point 500.

[0039] For a given audio source 211, a set of N acoustic source directivity patterns P i =P(D i ) (where i∈{1,...,N}) are different distances D i (where N is 1 or more, 2 or more, or 3 or more). All distances D between these directional patterns, i.e., min(D i )≦D≦max(D i ), a spherical interpolation technique may be applied to determine the directivity pattern for a particular distance D. As an example, linear interpolation may be used.

[0040] In this specification, the relatively small distance (D <min(D i A technique is described for determining an extrapolated (and possibly optimized) directional gain value P=P(D) for a relatively large distance (D>max(D i)) (i.e., beyond the greatest distance over which the directional pattern exists), a technique is described for determining an extrapolated (and possibly optimized) directional gain value P=P(D).

[0041] The techniques described herein are configured to prevent discontinuities in sound level at the origin 500, i.e., D=0, and / or to avoid directional gain calculations that cause irrelevant changes in the perception of the sound field for relatively large distances, i.e., D→∞. D→0 and / or D>D * (Here, D * is a specified distance threshold), the effect of applying directionality may be assumed to be negligibly small.

[0042] For a particular listening situation, a directivity control value (directivity_control_gain) may be calculated. The directivity control value may indicate the relevance of the corresponding directivity data to a particular listening situation of the listener 181. The directivity control value may be a function of the directivity (e.g., min(D i ) or max(D i ) based on the value of the user-object distance D (distance) 320 and a given reference distance D i The reference distance D i is the distance D at which the directional pattern 232 is available. i The directivity control value (also referred to herein as directivity control gain) can be calculated using the directivity control function (shown in FIG. 6a), i.e. directivity_control_gain = get_directivity_control_gain(distance, reference_distance) can be determined using

[0043] The value of the directivity control value (directivity_control_gain) is set to a predetermined directivity control value threshold D *(directivity_control_threshold). The directivity control value is compared with the threshold D * If it is greater than , a (distance independent) directivity gain value (directivity_gain_tmp) may be determined based on the directivity pattern 232, and the directivity gain value may be modified according to the directivity control value (directivity_control_gain). The modification of the directivity gain value may be performed as shown by the following pseudo code: if directive_control_gain > directive_control_threshold directivity_gain_tmp = get_directivity_gain(); directive_gain = directive_control_gain*(directivity_gain_tmp - 1) + 1; else directivity_gain = 1; end

[0044] As shown in the pseudo code above, the application of the directional data is performed when the directional control value exceeds a threshold D * If it is less than 1, it may be omitted.

[0045] The above obtained (distance dependent) directivity_gain can be applied to the corresponding distance attenuation gain as follows: distance_attenuation_gain = get_attenuation_gain(); gain = directivity_gain*distance_attenuation_gain;

[0046] FIG. 6a shows an example directivity control function 600 get_directivity_control_gain(). The most significant (i.e., subjective relevance) point of the directivity effect is located in the area around the marked position or a reference distance 610 (reference_distance). The reference distance 610 may be the distance 320 at which the directivity pattern 232 is measured. The directivity effect starts for values ​​smaller than the reference distance 610 and disappears for relatively large distances above the reference distance 610. The calculation for applying the directivity is performed when the directivity control value is greater than a threshold value D * 3. Thus, directionality is not applied near and / or far from the origin 500.

[0047] 6a illustrates how a directional control function 600 may be determined based on a combination of ("s" shaped) functions 602. The function 602 decreases monotonically towards 0 for relatively small distances and stays close to 1 for distances 610 and above (to address discontinuity issues at the origin 500). Another function 601 may decrease monotonically towards 0 for relatively large distances 320 and have a maximum at the reference distance 610 (to address perceptual relevance issues). The function 600, which represents the multiplication of both functions 601, 602, considers both issues simultaneously and may be used as the directional control function.

[0048] 6b illustrates an example distance attenuation function 650 get_attenuation_gain(). The distance attenuation function 650 illustrates the distance gain 651 as a function of distance 320 and represents the attenuation of the audio signal emitted by the audio signal 211. The distance attenuation function 650 may be identical to the distance function 315 described with respect to FIG. 3b.

[0049] For directional adaptation control, different types of directional control functions 600 and / or distance attenuation functions 650 may be considered, defined and applied. Their shapes and values ​​may be derived from physical considerations, measurement data and / or the intent of the content creator. The directional control function 600 may depend on the particular listening situation for the listener 181. The listening situation may be described by one or more of the following parameters, i.e. the directional control function 600 may depend on one or more of the following parameters: the frequency of the audio signal emitted by the audio source 211; the time at which the audio signal is rendered, Listener-object orientation, listener-view direction and / or listener 181 trajectory, User interactions, other conditions and / or scene events, and / or · System-related conditions (e.g. rendering workload).

[0050] The techniques described herein allow to improve the quality of 3D audio rendering, in particular by avoiding discontinuities in the audio volume close to the origin 500 of the audio source 211. Furthermore, the complexity of the directivity application may be reduced, in particular by avoiding the application of object directivity where it is perceptually irrelevant. In addition, the possibility to control the directivity application via the configuration of the encoder 130 is provided.

[0051] Herein, a general approach is described for increasing 6DoF audio rendering quality, reducing computational complexity, and achieving directivity adaptive control via the bitstream 140 (without modifying the directivity data itself). Additionally, a decoder interface is described for enabling the decoders 150, 160 to perform the directivity related processing outlined herein. Additionally, a bitstream syntax is described for enabling the bitstream 140 to carry directivity control data. The directivity control data, in particular the directivity control function 600, may be provided in a parameterized and sampled manner and / or as a predefined function.

[0052] 7a shows a flowchart of an example method 700 for rendering audio signals of audio sources 211, 212, 213 in a virtual reality rendering environment 180. The method 700 may be performed by the VR audio renderer 160.

[0053] The method 700 may include a step 701 of determining whether the directivity patterns 232 of the audio sources 211, 212, 213 should be considered with respect to a listening situation of a listener 181 in the virtual reality rendering environment 180. The listening situation may describe a context in which the listener 181 perceives the audio signals of the audio sources 211, 212, 213. The context may depend on a distance between the audio sources 211, 212, 213 and the listener 181. Alternatively, or in addition, the context may depend on whether the listener 181 is facing towards the audio sources 211, 212, 213 or whether the listener 181 is facing away from the audio sources 211, 212, 213. Alternatively or additionally, the context may depend on the listening position 182, 201, 202 of a listener 181 within the virtual reality rendering environment 180, and in particular within the audio scene 111 to be rendered.

[0054] In particular, a listening situation may be described by one or more parameters, where different listening situations may differ with respect to at least one of said one or more parameters. Examples of parameters are: Listening positions 182, 201, 202 of a listener 181 within a virtual reality rendering environment 180, and in particular within an audio scene 111; the distance 320 between a source position 500 (in particular a centre point) of the audio source 211, 212, 213 and a listening position 182, 201, 202 of a listener 181 within the virtual reality rendering environment 180, in particular within the audio scene 111; the frequency and / or spectral composition of the audio signal, The time at which the audio signal should be rendered, the orientation and / or gaze direction and / or movement trajectory of the listener 181 relative to the audio sources 211, 212, 213 within the virtual reality rendering environment 180 (i.e. within the virtual audio scene 111), e.g. the listening situation may be different depending on whether the listener 181 is moving closer or further away from the audio sources 211, 212, 213, the requirements of the renderer 160 for rendering the audio signal, in particular the requirements for the (available) computing resources, and / or Actions of the listener 181 relative to the virtual reality rendering environment 180

[0055] The method 700 may comprise a step 701 of determining whether the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account based on one or more parameters describing the listening situation. For this purpose, a (predetermined) directivity control function 600 may be used. Here, the directivity control function 600 may be configured to indicate whether the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account for different listening situations (in particular for different combinations of one or more parameters). In particular, the directivity control function 600 may be configured to identify listening situations in which the directivity of the audio sources 211, 212, 213 is perceptually irrelevant and / or in which the directivity of the audio sources 211, 212, 213 may lead to perceptual artifacts.

[0056] Furthermore, the method 700 includes a step 702 of rendering the audio signals of the audio sources 211, 212, 213 without considering the directivity patterns 232 of the audio sources 211, 212, 213 if it is determined that the directivity patterns 232 of the audio sources 211, 212, 213 should not be taken into account for the listening situation of the listener 181. Thereby, the directivity patterns 232 may be ignored by the renderer 160. In particular, the renderer 160 may omit the calculation of the directivity gain 410 based on the directivity patterns 232. Furthermore, the renderer 160 may omit applying the directivity gain 410 to the audio signal in order to render the audio signal. Thereby, a resource-efficient rendering of the audio signal may be achieved without affecting the perceptual quality.

[0057] On the other hand, the method 700 includes a step 703 of rendering the audio signals of the audio sources 211, 212, 213 depending on the directivity patterns 232 of the audio sources 211, 212, 213 if it is determined that the directivity patterns 232 should be taken into account for the listening situation of the listener 181. In this case, the renderer 160 may determine a directivity gain 410 to be applied to the audio signals (based on the directivity angle 420 between the source positions 500 of the audio sources 211, 212, 213 and the listening positions 182, 201, 202 of the listener 181). The directivity gain 410 may be applied to the audio signals before rendering the audio signals. This allows the audio signals to be rendered with a high perceptual quality (in listening situations where directivity is relevant).

[0058] Thus, a method 700 is described that, prior to processing an audio signal for rendering, determines in advance whether the use of directionality is relevant and / or perceptually advantageous in the current listening situation of the listener 181 (e.g., using a directivity control function 600). Directivity is only calculated and applied if the use of directionality is determined to be relevant and / or perceptually advantageous. This achieves resource-efficient rendering of high perceptual quality of the audio signal.

[0059] The directional patterns 232 of the audio sources 211, 212, 213 may indicate the strength in different directions of the audio signal. Alternatively, or in addition, the directional patterns 232 may indicate a direction-dependent directional gain 410 to be applied to the audio signal to render it (as outlined in relation to FIGS. 4a and 4b).

[0060] In particular, the directivity pattern 232 may exhibit a directivity gain function 415. The directivity gain function 415 may exhibit a directivity gain 410 as a function of a directivity angle 420 between a source position 500 of the audio source 211, 212, 213 and a listening position 182, 201, 202 of the listener 181. The directivity angle 420 may vary from 0° to 360° as the listening position 182, 201, 202 moves (circumferentially) around the source position 500. In the case of a non-uniform directivity pattern 232, the directivity gain 410 varies as a function of the directivity angle 420 (e.g., as shown in FIG. 4b).

[0061] Rendering 703 the audio signals of the audio sources 211, 212, 213 depending on the directivity patterns 232 of the audio sources 211, 212, 213 may include determining a directivity gain 410 (for rendering the audio signals in a particular listening situation) based on the directivity patterns 232 and based on a directivity angle 420 between the source positions 500 of the audio sources 211, 212, 213 and the listening positions 182, 201, 202 of the listener 181 (as outlined with respect to Figs. 4a and 4b). The audio signals may then be rendered depending on the directivity gain 410 (in particular by applying the directivity gain 410 to the audio signals before rendering). This allows a high perceptual quality to be achieved in the virtual reality rendering environment 180 (when the listening situation is changed by the listener 181 moving around in the rendering environment 180).

[0062] It should be noted that the method 700 described herein is typically repeated in a sequence of time instances (e.g. periodically at some repetition rate, such as every 20 ms). At each time instant, the currently valid listening situation is determined (e.g. by determining current values ​​for one or more parameters for describing the listening situation). Furthermore, at each time instant, it is determined whether the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account. Furthermore, at each time instant, rendering of the audio signals is performed depending on said determination. Thereby, continuous rendering of the audio signals in the virtual reality rendering environment 180 can be achieved.

[0063] Furthermore, it should be noted that typically multiple audio signals from multiple different audio sources 211, 212, 213 are rendered simultaneously within the virtual reality rendering environment 180 (e.g., as outlined with respect to Figures 3a and 3b). The method 700 may be performed for each of the different audio signals and / or audio sources 211, 212, 213. Further parameters for describing the listening situation may be the number, location and / or intensity of the different audio sources 211, 212, 213 active within the virtual reality rendering environment 180 (at a particular time).

[0064] The method 700 may include determining an attenuation or distance gain 310, 651 depending on the distance 320 between the source position 500 of the audio source 211, 212, 213 and the listening position 182, 201, 202 of the listener 181. The attenuation or distance gain 310, 651 may be determined using an attenuation or distance function 315, 650 indicating the attenuation or distance gain 310, 651 as a function of the distance 320. The audio signal may be rendered in dependence on the attenuation or distance gain 310, 651 (as described with respect to Figures 3a and 3b) to further increase the perceptual quality.

[0065] As described above, the distance 320 of the source positions 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 of the listener 181 within the virtual reality rendering environment 180 can be determined (in determining the listening situation of the listener 181).

[0066] The method 700 may comprise a step 701 of determining whether the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account based on the determined distance 320. For this purpose, a predefined directivity control function 600 may be used. The directivity control function 600 is configured to indicate the relevance and / or appropriateness of using the directivity of the audio sources 211, 212, 213 in the current listening situation, in particular with respect to the current distance 320 between the source position 500 and the listening positions 182, 201, 202. The distance 320 between the source position 500 and the listening positions 182, 201, 202 is a particularly important parameter of the listening situation and, as a result, has a particularly large impact on the resource efficiency and / or the perceived quality when rendering the audio signals of the audio sources 211, 212, 213.

[0067] It may be determined that the distance 320 of the source location 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 is smaller than a near field distance threshold. Based on this, it may be determined that the directivity patterns 232 of the audio sources 211, 212, 213 should not be considered. On the other hand, it may be determined that the distance 320 of the source location 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 is larger than a near field threshold. Based on this, it may be determined that the directivity patterns 232 of the audio sources 211, 212, 213 should be considered. The near field threshold may be, for example, 0.5 m or less. By suppressing the use of directivity at relatively small distances, perceptual artifacts may be prevented when the listener 181 crosses the (virtual) audio sources 211, 212, 213 in the virtual reality rendering environment 180.

[0068] Further, it may be determined that the distance 320 of the source location 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 is greater than a far field distance threshold (greater than the near field threshold). Based on this, it may be determined that the directivity patterns 232 of the audio sources 211, 212, 213 should not be considered. On the other hand, it may be determined that the distance 320 of the source location 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 is less than the far field threshold. Based on this, it may be determined that the directivity patterns 232 of the audio sources 211, 212, 213 should be considered. The far field threshold may be 5m or more. By suppressing the use of directivity at relatively large distances, the resource efficiency of the renderer 160 may be improved without affecting the perceived quality.

[0069] The near-field threshold and / or the far-field threshold may depend on a directivity control function 600. The directivity control function 600 may be configured to provide a control value as a function of the distance 320 between the source position 500 and the listening positions 182, 201, 202. The control value may indicate the extent to which the directivity pattern 232 should be taken into account. In particular, the control value may indicate whether the directivity pattern 232 should be taken into account (e.g., if the control value is greater than the control threshold By using the directional control function 600, the application of the directional pattern 232 can be efficiently and reliably controlled.

[0070] The method 700 may include determining a control value for a listening situation based on the directivity control function 600. The directivity control function 600 may be configured to provide different control values ​​for different listening situations, in particular for different distances 320 between the source position 500 and the listening positions 182, 201, 202. Then, based on the control values, it may be reliably determined whether the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account.

[0071] In particular, the method 700 may include setting the control value to a control threshold. JPEG2024521689000003.jpg54. The directivity control function 600 may be configured to provide a control value between a minimum value (e.g., 0) and a maximum value (e.g., 1). A control threshold may be between the minimum and maximum values ​​(e.g., 0.5). Based on said comparison, it may then be reliably determined whether the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account, in particular depending on whether the control value is greater or less than the control threshold. In particular, the method 700 may include determining that the directivity patterns 232 of the audio sources 211, 212, 213 should not be taken into account if the control value for the listening situation is less than the control threshold. Alternatively or additionally, the method 700 may include determining that the directivity patterns 232 of the audio sources 211, 212, 213 should be taken into account if the control value for the listening situation is greater than the control threshold.

[0072] The directivity control function 600 may be configured to provide a control value lower than the control threshold in listening situations where the distance 320 of the source positions 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 of the listener 181 is less than a near-field threshold (thereby preventing perceptual artifacts when the user crosses the virtual audio sources 211, 212, 213 in the virtual reality rendering environment 180). Alternatively or additionally, the directivity control function 600 may be configured to provide a control value lower than the control threshold in listening situations where the distance 320 of the source positions 500 of the audio sources 211, 212, 213 from the listening positions 182, 201, 202 of the listener 181 is greater than a far-field threshold (thereby increasing resource efficiency of the renderer 160 without affecting perceptual quality).

[0073] Thus, the method 700 may include determining a control value for a listening situation based on the directivity control function 600, where the directivity control function 600 may provide different control values ​​for different listening situations of the listener 181 in the virtual reality rendering environment 180. As mentioned above, the control value may indicate the degree to which the directivity pattern 232 should be taken into account.

[0074] Furthermore, the method 700 may include a step of adjusting the directional pattern 232 of the audio source 211, 212, 213, in particular the directional gain 410 of the directional pattern 232, in dependence on the control value (in particular if it is determined that the directional pattern 232 should be taken into account). The audio signals of the audio sources 211, 212, 213 may then be rendered in dependence on the adjusted directional pattern 232 of the audio source 211, 212, 213, in particular in dependence on the adjusted directional gain 410. Thus (in addition to determining whether to use the directional pattern 232) the directivity control function 600 may be used to control the degree to which the directivity is taken into account when rendering the audio signals. The degree may vary (continuously) in dependence on the listening situation of the listener 181 (in particular in dependence on the distance 320 between the source position 500 and the listening position 182, 201, 202). Doing so may further improve the perceived quality of the audio rendering within the virtual reality rendering environment 180.

[0075] The step of adjusting the directional pattern 232 of the audio sources 211, 212, 213 (depending on the control value) may include determining a weighted sum of the (non-uniform) directional patterns 232 of the audio sources 211, 212, 213 with a uniform directional pattern, where the weights for determining the weighted sum may depend on the control value. The adjusted directional pattern may be a weighted sum. In particular, the adjusted directional gain may be determined as a weighted sum of the source directional gain 410 and a uniform gain (typically 1 or 0 dB). By adjusting (smoothly) the directional pattern 232 for a distance 320 approaching the near-field or far-field threshold, a smooth transition between application and suppression of the directional pattern 232 may be achieved, so that the perceptual quality may be further improved.

[0076] The directivity patterns 232 of the audio sources 211, 212, 213 may be applicable to a reference listening situation, in particular to a reference distance 610 between the source positions 500 of the audio sources 211, 212, 213 and the listening position of the listener 181. In particular, the directivity patterns 232 may be measured and / or designed for the reference listening situation, in particular for the reference distance 610.

[0077] The directivity control function 600 may be such that the directivity pattern 232, and in particular the directional gain 410, is not adjusted when the listening situation corresponds to the reference listening situation (in particular when the distance 320 corresponds to the reference distance 610). By way of example, the directivity control function 600 may provide a maximum value for the control value (e.g., 1) when the listening situation corresponds to the reference listening situation.

[0078] Furthermore, the directivity control function 600 may be such that the degree of adjustment of the directivity pattern 232 increases the more the listening situation deviates from the reference listening situation (particularly as the distance 320 deviates from the reference distance 610). In particular, the directivity control function 600 may be such that the more the listening situation deviates from the reference listening situation (particularly as the distance 320 deviates from the reference distance 610), the more the directivity pattern 232 gradually moves toward a uniform directivity pattern (i.e., the directivity gain 410 gradually moves toward 1 or 0 dB). This may further increase the perceived quality.

[0079] 7b shows a flowchart of an example method 710 for rendering an audio signal of a first audio source 211 to a listener 181 in a virtual reality rendering environment 180. The method 710 may be performed by the renderer 160. It should be noted that all aspects described herein, and in particular with respect to the method 700, are also applicable to the method 710 (either alone or in combination).

[0080] The method 710 includes a step 711 of determining a control value for a listening situation of a listener 181 in a virtual reality rendering environment 180 based on the directivity control function 600. As mentioned above, the directivity control function 600 may provide different control values ​​for different listening situations, where the control value may indicate the degree to which the directionality of the audio sources 211, 212, 213 should be taken into account.

[0081] The method 710 may further comprise a step 712 of adjusting the directivity pattern 232 of the first audio source 211, 212, 213, in particular the directivity gain 410, depending on the control value. In addition, the method 710 comprises a step 713 of rendering the audio signal of the first audio source 211, 212, 213 to the listener 181 in the virtual reality rendering environment 180, in particular in dependence on the adjusted directivity pattern 232 of the first audio source 211, 212, 213, in dependence on the directivity gain 410. By adjusting the degree of application of the directivity depending on the current listening situation, the perceived quality of the audio rendering in the virtual reality rendering environment 180 may be increased.

[0082] As outlined with respect to Fig. 1a, data for rendering an audio signal in a virtual reality rendering environment may be provided by the encoder 130 in a bitstream 140. Fig. 7c shows a flowchart of an example method 720 for generating the bitstream 140. The method 720 may be performed by the encoder 130. It should be noted that features described herein may be applied to the method 720 (alone and / or in combination).

[0083] The method 720 comprises a step 721 of determining an audio signal of at least one audio source 211, 212, 213, a step 722 of determining a source position 500 in the virtual reality rendering environment 180 of the at least one audio source 211, 212, 213, and / or a step 723 of determining a directivity pattern 232 of the at least one audio source 211, 212, 213. Furthermore, the method 720 comprises a step 724 of determining a directivity control function 600 for controlling the use of the directivity pattern 232 for rendering the audio signal of the at least one audio source 211, 212, 213 depending on a listening situation of a listener 181 in the virtual reality rendering environment 180. In addition, the method 720 comprises a step 725 of inserting data relating to the audio signal, the source position 500, the directivity pattern 232 and / or the directivity control function 600 into the bitstream 140.

[0084] Thus, the creator of the virtual reality environment is provided with a means to flexibly and precisely control the directionality of one or more audio sources 211, 212, 213.

[0085] Further described is a virtual reality audio renderer 160 for rendering audio signals of audio sources 211, 212, 213 in a virtual reality rendering environment 180. The audio renderer 160 may be configured to perform the method steps of method 700 and / or method 710.

[0086] Additionally, an audio encoder 130 configured to generate the bitstream 140 is described. The audio encoder 130 may be configured to perform the method steps of the method 720.

[0087] Further, a bitstream 140 is described. The bitstream 140 may indicate an audio signal of at least one audio source 211, 212, 213 and / or a source position 500 within the virtual reality rendering environment 180 (i.e. within the audio scene 111) of said at least one audio source 211, 212, 213. Furthermore, the bitstream 140 may indicate a directivity control function 600 for controlling the use of the directivity pattern 232 for rendering the audio signal of said at least one audio source 211, 212, 213 depending on a directivity pattern 232 of said at least one audio source 211, 212, 213 and / or a listening situation of a listener 181 within the virtual reality rendering environment 180. The directivity control function 600 may be indicated in a parameterized and / or sampled manner. The directional pattern 232 and / or the directional control function 600 may be provided as VR metadata in the bitstream 140 (as generally described with respect to FIG. 1a).

[0088] The methods and systems described herein may be implemented as software, firmware and / or hardware. Certain components may be implemented as software running on, for example, a digital signal processor or a microprocessor. Other components may be implemented as, for example, hardware and / or as an application specific integrated circuit. The signals appearing in the methods and systems described above may be stored in a medium such as a random access memory or an optical storage medium. They may be transferred over a network such as a radio network, a satellite network, a wireless network or a wired network, for example the Internet. A typical device utilizing the methods and systems described herein is a portable electronic device or other consumer device used to store and / or render audio signals.

[0089] Various aspects of the present invention can be understood from the enumerated example embodiments (EEE) set forth below.

[0090] 1) A method (700) for rendering an audio signal of an audio source (211, 212, 213) in a virtual reality rendering environment (180), said method (700) comprising: determining (701) whether a directional pattern (232) of the audio source (211, 212, 213) should be taken into account for a listening situation of a listener (181) in the virtual reality rendering environment (180); rendering (702) the audio signals of the audio sources (211, 212, 213) without taking into account the directional patterns (232) of the audio sources (211, 212, 213) if it is determined that the directional patterns (232) of the audio sources (211, 212, 213) should not be taken into account for the listening situation of the listener (181); rendering (703) the audio signals of the audio sources (211, 212, 213) depending on the directional patterns (232) of the audio sources (211, 212, 213) if it is determined that the directional patterns (232) should be taken into account for the listening situation of the listener (181); Includes: Method (700).

[0091] 2) The method (700) comprises: determining one or more parameters describing said listening situation; determining (701) whether the directional pattern (232) of the audio source (211, 212, 213) should be taken into account based on the one or more parameters; Includes: The method according to EEE1 (700).

[0092] 3) the one or more parameters are a distance (320) between a source position (500) of said audio source (211, 212, 213) and a listening position (182, 201, 202) of said listener (181); the frequency of the audio signal, the time at which the audio signal should be rendered; the orientation and / or gaze direction and / or trajectory of the listener (181) with respect to the audio sources (211, 212, 213) within the virtual reality rendering environment (180); the requirements of a renderer (160) for rendering the audio signal, in particular with regard to computational resources, and / or movement of the listener (181) with respect to the virtual reality rendered environment (180); Including, The method according to EEE2 (700).

[0093] 4) The method (700) comprises: determining a distance (320) of a source position (500) of said audio sources (211, 212, 213) from a listening position (182, 201, 202) of said listener (181) within said virtual reality rendering environment (180); determining (701) whether the directional pattern (232) of the audio source (211, 212, 213) should be taken into account based on the distance (320); Includes: 7. The method according to any preceding claim (700).

[0094] 5) The method (700) comprises: determining that the distance (320) of the source location (500) of the audio source (211, 212, 213) from the listening position (182, 201, 202) is less than a near field threshold; and in response thereto, determining that the directional pattern (232) of the audio source (211, 212, 213) should not be taken into account; and / or determining that the distance (320) of the source location (500) of the audio source (211, 212, 213) from the listening position (182, 201, 202) is greater than the near field threshold; and determining in response that the directional pattern (232) of the audio source (211, 212, 213) should be taken into account; Includes: The method according to EEE4 (700).

[0095] 6) The method (700) comprises: determining that the distance (320) of the source location (500) of the audio source (211, 212, 213) from the listening position (182, 201, 202) is greater than a far-field threshold; and in response thereto, determining that the directional pattern (232) of the audio source (211, 212, 213) should not be taken into account; and / or determining that the distance (320) of the source location (500) of the audio source (211, 212, 213) from the listening position (182, 201, 202) is less than the far-field threshold; and determining in response that the directional pattern (232) of the audio source (211, 212, 213) should be taken into account; Includes: The method according to claim 4 or 5 (700).

[0096] 7) the near-field threshold and / or the far-field threshold depend on a directivity control function (600); The directivity control function (600) provides a control value as a function of the distance (320); and the control value indicates the degree to which the directional pattern (232) should be taken into account; The method according to claim 5 or 6 (700).

[0097] 8) The method (700) comprises: determining a control value for the listening situation based on a directional control function (600), the directional control function (600) providing different control values ​​for different listening situations; determining, based on said control value, whether the directional pattern (232) of said audio source (211, 212, 213) should be taken into account; Includes: 7. The method according to any preceding claim (700).

[0098] 9) The method (700) comprises: comparing the control value with a control threshold; - determining, based on said comparison, in particular depending on whether said control value is greater or less than said control threshold, whether said directional pattern (232) of said audio source (211, 212, 213) should be taken into account; Includes: The method according to EEE8 (700).

[0099] 10) the directional control function (600) is configured to provide a control value between a minimum and a maximum value; In particular, said minimum value is 0 and / or said maximum value is 1, the control threshold is between the minimum and maximum values, and The method (700) comprises: determining that the directional pattern (232) of the audio source (211, 212, 213) should not be taken into account if the control value for the listening situation is less than the control threshold; and / or determining that the directional pattern (232) of the audio source (211, 212, 213) should be taken into account if the control value for the listening situation is greater than the control threshold; Includes: The method described in EEE9 (700).

[0100] 11) The directivity control function (600) is the distance (320) of the source location (500) of the audio source (211, 212, 213) from the listening position (182, 201, 202) of the listener (181) is less than a near-field threshold; and / or the distance (320) of the source location (500) of the audio source (211, 212, 213) from the listening position (182, 201, 202) of the listener (181) is greater than a far-field threshold; configured to provide a control value less than the control threshold in such a listening situation. The method according to EEE10 (700).

[0101] 12) The method (700) comprises: determining a control value for the listening situation based on a directional control function (600), the directional control function (600) providing different control values ​​for different listening situations, the control value indicating the degree to which the directional pattern (232) should be taken into account; adjusting the directional pattern (232) of the audio source (211, 212, 213) in dependence on the control value; rendering (703) the audio signals of the audio sources (211, 212, 213) depending on the adjusted directional patterns (232) of the audio sources (211, 212, 213); Includes: 7. The method according to any preceding claim (700).

[0102] 13) adjusting the directional patterns (232) of the audio sources (211, 212, 213) comprises determining a weighted sum of the directional patterns (232) of the audio sources (211, 212, 213) using a uniform directional pattern; and the weights for determining the weighted sum depend on the control values. The method according to EEE12 (700).

[0103] 14) the directivity pattern (232) of the audio source (211, 212, 213) is applicable to a reference listening situation, in particular a reference distance (610) between a source position (500) of the audio source (211, 212, 213) and a listening position of the listener (181); and The directivity control function (600) is if the listening situation corresponds to the reference listening situation, the directional pattern (232) is not adjusted; and / or In particular, the degree of adjustment of the directional pattern (232) increases as the deviation of the listening situation from the reference listening situation increases, such that the directional pattern (232) gradually tends towards a uniform directional pattern as the deviation of the listening situation from the reference listening situation increases. Such a function is The method according to any one of claims 1 to 7 (700) described in EEE12 or 13.

[0104] 15) the directional patterns (232) of the audio sources (211, 212, 213) indicate the strength of the audio signal in different directions; and / or the directional pattern (232) indicates a direction-dependent directional gain (410) to be applied to the audio signal to render the audio signal; 7. The method according to any preceding claim (700).

[0105] 16) The directional pattern (232) exhibits a directional gain function (415), and the directivity gain function (415) indicates the directivity gain (410) as a function of the directivity angle (420) between a source position (500) of the audio source (211, 212, 213) and a listening position (182, 201, 202) of the listener (181); The method according to EEE14 (700).

[0106] 17) The step of rendering (703) the audio signals of the audio sources (211, 212, 213) in dependence on the directional patterns (232) of the audio sources (211, 212, 213) comprises: determining a directional gain (410) based on the directional pattern (232) and based on a directional angle (420) between a source position (500) of the audio source (211, 212, 213) and a listening position (182, 201, 202) of the listener (181); rendering the audio signal in dependence on the directional gain (410); Includes: 7. The method according to any preceding claim (700).

[0107] 18) The method (700) comprises: determining an attenuation gain (651) depending on a distance (320) between a source position (500) of said audio source (211, 212, 213) and a listening position (182, 201, 202) of said listener (181) using an attenuation function (650) that indicates said attenuation gain (651) as a function of said distance (320); rendering the audio signal in dependence on the attenuation gain (651); Includes: 7. The method according to any preceding claim (700).

[0108] 19) A method (710) for rendering an audio signal of a first audio source (211) to a listener (181) in a virtual reality rendering environment (180), the method (710) comprising: - determining (711) a control value for a listening situation of the listener (181) in the virtual reality rendering environment (180) based on a directivity control function (600), the directivity control function (600) providing different control values ​​for different listening situations, the control value indicating the degree to which the directionality of the audio sources (211, 212, 213) should be taken into account; adjusting (712) a directional pattern (232) of the first audio source (211, 212, 213) depending on the control value; rendering (713) the audio signal of the first audio source (211, 212, 213) to the listener (181) within the virtual reality rendering environment (180) in dependence on the adjusted directivity pattern (232) of the first audio source (211, 212, 213); Includes: Method (710).

[0109] 20) A virtual reality audio renderer (160) for rendering audio signals of audio sources (211, 212, 213) in a virtual reality rendering environment (180), said audio renderer (160) comprising: determining whether the directional patterns (232) of the audio sources (211, 212, 213) should be taken into account with respect to the listening situation of a listener (181) within the virtual reality rendering environment (180); rendering the audio signals of the audio sources (211, 212, 213) without taking into account the directional patterns (232) of the audio sources (211, 212, 213) if it is determined that the directional patterns (232) of the audio sources (211, 212, 213) should not be taken into account for the listening situation of the listener (181); and rendering the audio signals of the audio sources (211, 212, 213) depending on the directional patterns (232) of the audio sources (211, 212, 213) if it is determined that the directional patterns (232) should be taken into account for the listening situation of the listener (181); It was configured as follows: Virtual reality audio renderer (160).

[0110] 21) A virtual reality audio renderer (160) for rendering an audio signal of a first audio source (211) to a listener (181) within a virtual reality rendering environment (180), said audio renderer (160) comprising: determining a control value for a listening situation of the listener (181) in the virtual reality rendering environment (180) based on a directivity control function (600), the directivity control function (600) providing different control values ​​for different listening situations, the control value indicating the degree to which the directionality of the audio sources (211, 212, 213) should be taken into account; adjusting the directional pattern (232) of the first audio source (211, 212, 213) in dependence on the control value; and rendering the audio signal of the first audio source (211, 212, 213) to the listener (181) within the virtual reality rendering environment (180) in dependence on the adjusted directivity pattern (232) of the first audio source (211, 212, 213); It was configured as follows: Virtual reality audio renderer (160).

[0111] 22) an audio signal from at least one audio source (211, 212, 213); a source location (500) within the virtual reality rendering environment (180) of the at least one audio source (211, 212, 213); a directional pattern (232) of the at least one audio source (211, 212, 213); a directivity control function (600) for controlling the use of the directivity pattern (232) for rendering the audio signals of the at least one audio source (211, 212, 213) depending on a listening situation of a listener (181) within the virtual reality rendering environment (180); An audio encoder (130) configured to generate a bitstream (140) indicative of

[0112] 23) an audio signal from at least one audio source (211, 212, 213); a source location (500) within the virtual reality rendering environment (180) of the at least one audio source (211, 212, 213); a directional pattern (232) of the at least one audio source (211, 212, 213); a directivity control function (600) for controlling the use of the directivity pattern (232) for rendering the audio signals of the at least one audio source (211, 212, 213) depending on a listening situation of a listener (181) within the virtual reality rendering environment (180); A bitstream (140) showing:

[0113] 24) A method (720) for generating a bitstream (140), the method (720) comprising: determining (721) an audio signal of at least one audio source (211, 212, 213); determining (722) a source location (500) within a virtual reality rendering environment (180) of the at least one audio source (211, 212, 213); determining (723) a directional pattern (232) of said at least one audio source (211, 212, 213); determining (724) a directivity control function (600) for controlling the use of the directivity pattern (232) for rendering the audio signals of the at least one audio source (211, 212, 213) depending on a listening situation of a listener (181) within the virtual reality rendering environment (180); inserting (725) data relating to said audio signal, said source position (500), said directivity pattern (232) and said directivity control function (600) into said bitstream (140); Includes: Method (720).

Claims

1. 1. A method for rendering an audio signal of an audio source in a virtual reality rendering environment, the method comprising: determining a distance of a source location of the audio source from a listening position of a listener within the virtual reality rendered environment; determining whether a directional pattern of the audio source should be taken into account for a listening situation of the listener in the virtual reality rendering environment based on the distance; - rendering the audio signal of the audio source without taking into account the directional pattern of the audio source if it is determined that the directional pattern of the audio source should not be taken into account for the listening situation of the listener; - rendering the audio signal of the audio source depending on the directional pattern of the audio source if it is determined that the directional pattern should be taken into account for the listening situation of the listener; The method includes:

2. The method comprises: determining one or more parameters describing said listening situation; determining, based on the one or more parameters, whether the directional pattern of the audio source should be taken into account; The method of claim 1 , comprising:

3. The one or more parameters are: the distance between the source position of the audio source and the listening position of the listener; the frequency of the audio signal, the time at which the audio signal should be rendered; the orientation and / or gaze direction and / or trajectory of the listener with respect to the audio source within the virtual reality rendered environment; the requirements of a renderer for rendering said audio signal, in particular with regard to computational resources, and / or movement of the listener with respect to the virtual reality rendered environment; The method of claim 2 , comprising:

4. The method comprises: determining that the distance of the source location of the audio source from the listening position is less than a near field threshold; and and / or in response determining that the directional pattern of the audio source should not be taken into account. determining that the distance of the source location of the audio source from the listening position is greater than the near field threshold; and in response thereto, determining that the directional pattern of the audio source should be taken into account; The method according to any one of claims 1 to 3, comprising:

5. The method comprises: determining that the distance of the source location of the audio source from the listening position is greater than a far-field threshold; and In response, the directional pattern of the audio source should not be taken into account. and / or determining that determining that the distance of the source location of the audio source from the listening position is less than the far-field threshold; and in response thereto, determining that the directional pattern of the audio source should be taken into account; The method according to any one of claims 1 to 3, comprising:

6. the near field threshold and / or the far field threshold are dependent on a directivity control function; the directivity control function providing a control value as a function of the distance; and The method according to claim 1 , wherein the control value indicates the degree to which the directional pattern should be taken into account.

7. The method comprises: determining a control value for the listening situation based on a directivity control function, the directivity control function providing different control values ​​for different listening situations; determining based on said control value whether the directional pattern of said audio source should be taken into account; The method according to any one of claims 1 to 3, comprising:

8. The method comprises: comparing the control value with a control threshold; - determining, based on said comparison, in particular depending on whether said control value is greater or less than said control threshold, whether said directional pattern of said audio source should be taken into account; The method of claim 7, comprising:

9. the directional control function is configured to provide a control value between a minimum and a maximum value; the minimum value is 0 and / or the maximum value is 1; the control threshold is between the minimum and maximum values, and The method comprises: - determining that the directional pattern of the audio source should not be taken into account if the control value for the listening situation is less than the control threshold; and / or determining that the directional pattern of the audio source should be taken into account if the control value for the listening situation is greater than the control threshold; The method of claim 8, comprising:

10. The directivity control function is the distance of the source position of the audio source from the listening position of the listener is less than a near-field threshold; and / or the distance of the source location of the audio source from the listening position of the listener is greater than a far-field threshold; 10. The method of claim 9, configured to provide a control value below the control threshold in such listening situations.

11. The method comprises: determining a control value for the listening situation based on a directivity control function, the directivity control function providing different control values ​​for different listening situations, the control value indicating the degree to which the directivity pattern should be taken into account; adjusting the directional pattern of the audio source in dependence on the control value; rendering the audio signal of the audio source depending on the adjusted directional pattern of the audio source; The method according to any one of claims 1 to 3, comprising:

12. adjusting the directional patterns of the audio sources comprises determining a weighted sum of the directional patterns of the audio sources using a uniform directional pattern; and The method of claim 11 , wherein weights for determining the weighted sum depend on the control value.

13. the directivity pattern of the audio source is applicable to a reference listening situation, in particular a reference distance between the source position of the audio source and the listening position of the listener, and The directivity control function is if the listening situation corresponds to the reference listening situation, the directional pattern is not adjusted; and / or In particular, the degree of adjustment of the directional pattern increases with increasing deviation of the listening situation from the reference listening situation, such that the directional pattern gradually tends towards a uniform directional pattern with increasing deviation of the listening situation from the reference listening situation. The method of claim 11, wherein the function is:

14. the directional pattern of the audio source indicating the strength of the audio signal in different directions; and / or The method of claim 1 , wherein the directional pattern indicates a direction-dependent directional gain to be applied to the audio signal to render it.

15. The directional pattern exhibits a directional gain function; and The method of claim 13 , wherein the directivity gain function indicates directivity gain as a function of a directivity angle between the source position of the audio source and the listening position of the listener.

16. Rendering the audio signal of the audio source in dependence on the directional pattern of the audio source comprises: determining a directivity gain based on the directivity pattern and based on a directivity angle between the source position of the audio source and the listening position of the listener; rendering the audio signal in dependence on the directional gain; The method according to any one of claims 1 to 3, comprising:

17. The method comprises: determining an attenuation gain in dependence on a distance between a source position of the audio source and a listening position of the listener using an attenuation function that indicates the attenuation gain as a function of the distance; rendering the audio signal in dependence on the attenuation gain; The method according to any one of claims 1 to 3, comprising:

18. 1. A method for rendering an audio signal of a first audio source to a listener in a virtual reality rendering environment, the method comprising: determining a control value for a listening situation of the listener in the virtual reality rendering environment based on a directivity control function, the directivity control function providing different control values ​​for different listening situations, the control values ​​being determined based on a position of the first audio source and the listener. and said control value indicates the degree to which the directionality of the audio source should be taken into account; adjusting a directional pattern of the first audio source in dependence on the control value; rendering the audio signal of the first audio source to the listener within the virtual reality rendering environment in dependence on the adjusted directivity pattern of the first audio source; The method includes:

19. 1. A virtual reality audio renderer for rendering an audio signal of an audio source in a virtual reality rendering environment, said audio renderer comprising: determining a distance of a source location of the audio source from a listening position of a listener within the virtual reality rendering environment; determining whether a directional pattern of the audio source should be taken into account for a listening situation of the listener within the virtual reality rendering environment based on the distance; - rendering the audio signal of the audio source without taking into account the directional pattern of the audio source if it is determined that the directional pattern of the audio source should not be taken into account for the listening situation of the listener; and rendering the audio signal of the audio source depending on the directional pattern of the audio source if it is determined that the directional pattern should be taken into account for the listening situation of the listener. A virtual reality audio renderer configured to:

20. 1. A virtual reality audio renderer for rendering an audio signal of a first audio source to a listener within a virtual reality rendering environment, the audio renderer comprising: determining a control value for a listening situation of the listener in the virtual reality rendering environment based on a directivity control function, the directivity control function providing different control values ​​for different listening situations, the control value being a function of a distance between a listening position of the first audio source and a position of the listener, the control value indicating the degree to which a directionality of an audio source should be taken into account; adjusting a directional pattern of the first audio source in dependence on the control value; and rendering the audio signal of the first audio source dependent on the adjusted directivity pattern of the first audio source to the listener within the virtual reality rendering environment. A virtual reality audio renderer configured to:

21. a source location of at least one audio source within the virtual reality rendering environment; a directional pattern of the at least one audio source; and a directivity control function for controlling the use of the directivity pattern for rendering audio signals of the at least one audio source depending on a listening situation of a listener within the virtual reality rendering environment, the directivity control function being configured to provide a control value as a function of a distance between a source position and a listening position of the listener; 16. An audio encoder configured to generate a bitstream representing:

22. A method for rendering an audio signal of an audio source in a virtual reality rendering environment, the method comprising: determining an angle and / or distance of a source position of the audio source from a listening position of a listener within the virtual reality rendered environment; determining a use of a directional pattern based at least on the angle of the source location from the listening position; determining a directional gain based on the directional pattern and further based on an angle and / or distance relative to the source position of the audio source and the listening position of the listener; rendering the audio signal of the audio source based on the directional gain; The method includes: